LingBot-World Fast, int8 transformer

An int8 weight-only version of the transformer of robbyant/lingbot-world-fast-diffusers at revision 5221cb0454fbf46047288ba7d01b5186df626f6f, by the Robbyant team. The text encoder, decoder, tokenizer, scheduler and pipeline index are the original's, unchanged, so this folder is complete on its own.

Original This repository
Transformer 74.2 GB (FP32) 18.6 GB
Whole folder 86.1 GB 30.5 GB

Method

  • Weight-only int8, one scale per output row. For each quantized linear layer, scale = max(|W|, row) / 127 and q = round(W / scale) clamped to [-127, 127]. The layer stores <name>.weight as int8 [out, in] and <name>.weight_scale as float32 [out]. Dequantized, W ≈ q * scale[:, None].
  • Which layers: every 2-D .weight with at least 1,048,576 values whose two sides are multiples of 32 (566 layers): self- and cross-attention, feed-forward, camera conditioning, text embedding and time projection.
  • Kept as the original, in bf16: the timestep embedder (time_embedding.0, time_embedding.2). Common loaders cast the timestep embedder's input to the dtype of its first layer's weight; as int8 that input is destroyed, and every frame comes out black.
  • Everything else (norms, biases, modulation tables, the patch embedding, the output head) is the original's, in bf16 where it was FP32.
  • The conversion is deterministic: the same original gives byte-identical files.
  • transformer/int8_quantization.json records the method, the kept layers, the layer count and the size.

Loading

The checkpoint is not a stock diffusers format. A loader must:

  1. build the transformer, then give each quantized layer an int8 weight and a float32 weight_scale of the shapes above before loading weights, so the int8 values are not cast to a floating type;
  2. compute each quantized layer as x @ (q * scale[:, None]).T + bias, either by dequantizing per layer or with a weight-only int8 matmul such as PyTorch's torch._weight_int8pack_mm(x, q, scale) (which needs both sides to be multiples of 32, as every quantized layer here is);
  3. leave the timestep embedder's layers in bf16.

Check

Run at 640×352 against the bf16 original on an Apple M3 Ultra, from the same picture, prompt and seed, and looked at frame by frame. Over the first chunk (9 frames), where the camera input was the same, both show the same scene, and each frame's mean brightness matches the original's to within 2 levels out of 255 (int8: 96, 87, 81, 81, 82, 85, 86, 90, 90; original: 95, 89, 80, 81, 81, 85, 86, 89, 90). The second chunk was run with a different camera input in each, so it is not compared here. Not measured: any benchmark score.

Memory while running there, at 640×352, transformer and text encoder resident: about 31 GB loaded, 44–46 GB while generating.

Files

File Bytes SHA-256
LICENSE 11,357 c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4
model_index.json 595 0a439c3869d30998c60e030abb3dc80830ffe7a15af286db8151e5ebd74d6c19
scheduler/scheduler_config.json 820 40624e058848feddf4bc2da5e8a232668b9a6f4a4939365810bebd9ce0166578
text_encoder/config.json 855 a2bcb24699f6c009a2427432bdd483ef8b2b42a712abc9503759cdc77d171f07
text_encoder/model-00001-of-00003.safetensors 4,935,812,536 a8e861969c7433e707cc5a74065d795d36cca07ec96eb6763eb4083df7248f58
text_encoder/model-00002-of-00003.safetensors 4,983,103,192 d57d948ece4837d850b7a859a4415121d57cacf8b9ee1d4db200c67f592902d7
text_encoder/model-00003-of-00003.safetensors 1,442,935,480 0da9ee284e21d1406df708788db1d502d95d75f69faa25cd26151bf8829b7c5f
text_encoder/model.safetensors.index.json 22,476 31c4c7bcce679eaa0dd4667462394ddb013dc2f748e0bffc893dc9146a320dab
tokenizer/special_tokens_map.json 7,079 456b58fd240a06c743a7c2cf8008bec501240d68ebd1fc4018ea569505fea270
tokenizer/spiece.model 4,548,313 e3909a67b780650b35cf529ac782ad2b6b26e6d1f849d3fbb6a872905f452458
tokenizer/tokenizer.json 16,837,459 20a46ac256746594ed7e1e3ef733b83fbc5a6f0922aa7480eda961743de080ef
tokenizer/tokenizer_config.json 61,758 1d8d2a216bf8e70ac15b7ddcea566c4dd0433c024b39a58ca5e4c66bd78defbd
transformer/config.json 617 e9d6eca86aa86a96b0de9c21c6303a309dd34ae6fa6fd56416a8c751d8d8dfdb
transformer/diffusion_pytorch_model-00001-of-00016.safetensors 1,235,224,464 a76bbd3859a634d642699454d31640182f579dc7c373755596ffde3199b08704
transformer/diffusion_pytorch_model-00002-of-00016.safetensors 1,123,485,808 6c10b00c93818025ea9b4fea22373304fc3af9ad10a60856affe67f85190613a
transformer/diffusion_pytorch_model-00003-of-00016.safetensors 1,238,942,944 6c8fd3775a67cd84f8bb59cf72d5691e2017b55ed93a46d0ab3acd466256f97b
transformer/diffusion_pytorch_model-00004-of-00016.safetensors 1,186,442,216 51af611e4863f23a19bc31e8ed548104fd3e0b15ef3c6bab296e7986a3d563e7
transformer/diffusion_pytorch_model-00005-of-00016.safetensors 1,123,537,168 36a003ddc4c32d187876e0f0f5629c650fcfd956c0c9fd5b8cffd4093a76cd2d
transformer/diffusion_pytorch_model-00006-of-00016.safetensors 1,212,667,120 3dd6ac58a57e44f371789b9b68c93481e6b0bc77135b3f742c56b99393dfff35
transformer/diffusion_pytorch_model-00007-of-00016.safetensors 1,149,761,896 9c13bb7c5a3bbb9b5c9a603bf448e98b67d687038aeac647d3f5f343fa0f7633
transformer/diffusion_pytorch_model-00008-of-00016.safetensors 1,186,493,584 836ea76441a603901f0f507b87620aa01ae6815ea3aebcfe42f9398470aa538a
transformer/diffusion_pytorch_model-00009-of-00016.safetensors 1,123,485,936 01b31f28443de8afc8b25b95a6d4177918c94628e1a56ef19fbfe4d70f76ee4c
transformer/diffusion_pytorch_model-00010-of-00016.safetensors 1,238,943,072 f81833fc631232dcae8dd6d4419557d1ef0f9df53d770e0b531a7dea4741bc61
transformer/diffusion_pytorch_model-00011-of-00016.safetensors 1,186,442,344 fc106a14c590a41e7749b93364d0e109d3b817935ae1390cdff35b0dabf35654
transformer/diffusion_pytorch_model-00012-of-00016.safetensors 1,123,537,168 9e9896a1937320d4b0d684bbc26bd6b1570db2f139d74fcaec3826b06545bb5b
transformer/diffusion_pytorch_model-00013-of-00016.safetensors 1,212,667,120 9e5b9f520f30f4e4339a625761507c2e56a15a4373972945d690c951f6d7e2d7
transformer/diffusion_pytorch_model-00014-of-00016.safetensors 1,149,761,896 c475bddd0c934545262781b0d65d513a63b6a15c2c5651ad6aefd035cce89d3a
transformer/diffusion_pytorch_model-00015-of-00016.safetensors 1,186,493,584 1c5cfc8fa40a303c2a46ae53c0b453bdffadf1f7387d992a86787740e77388dd
transformer/diffusion_pytorch_model-00016-of-00016.safetensors 914,075,312 a66812f617f28103b1452e2f519a9df6247e6fd92bebadba0d8e404efc90515a
transformer/diffusion_pytorch_model.safetensors.index.json 182,519 1fb31fa73e6866e4171253ffc5afb4225b6a54d55a02ca35889e17a17fd46cae
transformer/int8_quantization.json 128 1f6f6f3253cc416d8f49502ea089eda3914e9d9df05db9067b94e4ee2df811ea
vae/config.json 724 47e8bcf55e93e9c182e1962a8c7a0650faeb34ea0f66826d6f8aaa9f73e08ec9
vae/diffusion_pytorch_model.safetensors 507,591,892 d6e524b3fffede1787a74e81b30976dce5400c4439ba64222168e607ed19e793

Licence

Apache-2.0, as the original (LICENSE). The model is the work of the Robbyant team; see the original repository and its paper, arXiv:2601.20540.

Downloads last month
-
Safetensors
Model size
19B params
Tensor type
F32
·
BF16
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ogtsvc/lingbot-world-fast-diffusers-int8

Quantized
(1)
this model

Paper for ogtsvc/lingbot-world-fast-diffusers-int8