Instructions to use ogtsvc/lingbot-world-fast-diffusers-int8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use ogtsvc/lingbot-world-fast-diffusers-int8 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline from diffusers.utils import load_image, export_to_video # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("ogtsvc/lingbot-world-fast-diffusers-int8", dtype=torch.bfloat16, device_map="cuda") pipe.to("cuda") prompt = "A man with short gray hair plays a red electric guitar." image = load_image( "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/diffusers/guitar-man.png" ) output = pipe(image=image, prompt=prompt).frames[0] export_to_video(output, "output.mp4") - Notebooks
- Google Colab
- Kaggle
LingBot-World Fast, int8 transformer
An int8 weight-only version of the transformer of
robbyant/lingbot-world-fast-diffusers
at revision 5221cb0454fbf46047288ba7d01b5186df626f6f, by the Robbyant team. The
text encoder, decoder, tokenizer, scheduler and pipeline index are the original's,
unchanged, so this folder is complete on its own.
| Original | This repository | |
|---|---|---|
| Transformer | 74.2 GB (FP32) | 18.6 GB |
| Whole folder | 86.1 GB | 30.5 GB |
Method
- Weight-only int8, one scale per output row. For each quantized linear layer,
scale = max(|W|, row) / 127andq = round(W / scale)clamped to [-127, 127]. The layer stores<name>.weightas int8[out, in]and<name>.weight_scaleas float32[out]. Dequantized,W ≈ q * scale[:, None]. - Which layers: every 2-D
.weightwith at least 1,048,576 values whose two sides are multiples of 32 (566 layers): self- and cross-attention, feed-forward, camera conditioning, text embedding and time projection. - Kept as the original, in bf16: the timestep embedder (
time_embedding.0,time_embedding.2). Common loaders cast the timestep embedder's input to the dtype of its first layer's weight; as int8 that input is destroyed, and every frame comes out black. - Everything else (norms, biases, modulation tables, the patch embedding, the output head) is the original's, in bf16 where it was FP32.
- The conversion is deterministic: the same original gives byte-identical files.
transformer/int8_quantization.jsonrecords the method, the kept layers, the layer count and the size.
Loading
The checkpoint is not a stock diffusers format. A loader must:
- build the transformer, then give each quantized layer an int8
weightand a float32weight_scaleof the shapes above before loading weights, so the int8 values are not cast to a floating type; - compute each quantized layer as
x @ (q * scale[:, None]).T + bias, either by dequantizing per layer or with a weight-only int8 matmul such as PyTorch'storch._weight_int8pack_mm(x, q, scale)(which needs both sides to be multiples of 32, as every quantized layer here is); - leave the timestep embedder's layers in bf16.
Check
Run at 640×352 against the bf16 original on an Apple M3 Ultra, from the same picture, prompt and seed, and looked at frame by frame. Over the first chunk (9 frames), where the camera input was the same, both show the same scene, and each frame's mean brightness matches the original's to within 2 levels out of 255 (int8: 96, 87, 81, 81, 82, 85, 86, 90, 90; original: 95, 89, 80, 81, 81, 85, 86, 89, 90). The second chunk was run with a different camera input in each, so it is not compared here. Not measured: any benchmark score.
Memory while running there, at 640×352, transformer and text encoder resident: about 31 GB loaded, 44–46 GB while generating.
Files
| File | Bytes | SHA-256 |
|---|---|---|
LICENSE |
11,357 | c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4 |
model_index.json |
595 | 0a439c3869d30998c60e030abb3dc80830ffe7a15af286db8151e5ebd74d6c19 |
scheduler/scheduler_config.json |
820 | 40624e058848feddf4bc2da5e8a232668b9a6f4a4939365810bebd9ce0166578 |
text_encoder/config.json |
855 | a2bcb24699f6c009a2427432bdd483ef8b2b42a712abc9503759cdc77d171f07 |
text_encoder/model-00001-of-00003.safetensors |
4,935,812,536 | a8e861969c7433e707cc5a74065d795d36cca07ec96eb6763eb4083df7248f58 |
text_encoder/model-00002-of-00003.safetensors |
4,983,103,192 | d57d948ece4837d850b7a859a4415121d57cacf8b9ee1d4db200c67f592902d7 |
text_encoder/model-00003-of-00003.safetensors |
1,442,935,480 | 0da9ee284e21d1406df708788db1d502d95d75f69faa25cd26151bf8829b7c5f |
text_encoder/model.safetensors.index.json |
22,476 | 31c4c7bcce679eaa0dd4667462394ddb013dc2f748e0bffc893dc9146a320dab |
tokenizer/special_tokens_map.json |
7,079 | 456b58fd240a06c743a7c2cf8008bec501240d68ebd1fc4018ea569505fea270 |
tokenizer/spiece.model |
4,548,313 | e3909a67b780650b35cf529ac782ad2b6b26e6d1f849d3fbb6a872905f452458 |
tokenizer/tokenizer.json |
16,837,459 | 20a46ac256746594ed7e1e3ef733b83fbc5a6f0922aa7480eda961743de080ef |
tokenizer/tokenizer_config.json |
61,758 | 1d8d2a216bf8e70ac15b7ddcea566c4dd0433c024b39a58ca5e4c66bd78defbd |
transformer/config.json |
617 | e9d6eca86aa86a96b0de9c21c6303a309dd34ae6fa6fd56416a8c751d8d8dfdb |
transformer/diffusion_pytorch_model-00001-of-00016.safetensors |
1,235,224,464 | a76bbd3859a634d642699454d31640182f579dc7c373755596ffde3199b08704 |
transformer/diffusion_pytorch_model-00002-of-00016.safetensors |
1,123,485,808 | 6c10b00c93818025ea9b4fea22373304fc3af9ad10a60856affe67f85190613a |
transformer/diffusion_pytorch_model-00003-of-00016.safetensors |
1,238,942,944 | 6c8fd3775a67cd84f8bb59cf72d5691e2017b55ed93a46d0ab3acd466256f97b |
transformer/diffusion_pytorch_model-00004-of-00016.safetensors |
1,186,442,216 | 51af611e4863f23a19bc31e8ed548104fd3e0b15ef3c6bab296e7986a3d563e7 |
transformer/diffusion_pytorch_model-00005-of-00016.safetensors |
1,123,537,168 | 36a003ddc4c32d187876e0f0f5629c650fcfd956c0c9fd5b8cffd4093a76cd2d |
transformer/diffusion_pytorch_model-00006-of-00016.safetensors |
1,212,667,120 | 3dd6ac58a57e44f371789b9b68c93481e6b0bc77135b3f742c56b99393dfff35 |
transformer/diffusion_pytorch_model-00007-of-00016.safetensors |
1,149,761,896 | 9c13bb7c5a3bbb9b5c9a603bf448e98b67d687038aeac647d3f5f343fa0f7633 |
transformer/diffusion_pytorch_model-00008-of-00016.safetensors |
1,186,493,584 | 836ea76441a603901f0f507b87620aa01ae6815ea3aebcfe42f9398470aa538a |
transformer/diffusion_pytorch_model-00009-of-00016.safetensors |
1,123,485,936 | 01b31f28443de8afc8b25b95a6d4177918c94628e1a56ef19fbfe4d70f76ee4c |
transformer/diffusion_pytorch_model-00010-of-00016.safetensors |
1,238,943,072 | f81833fc631232dcae8dd6d4419557d1ef0f9df53d770e0b531a7dea4741bc61 |
transformer/diffusion_pytorch_model-00011-of-00016.safetensors |
1,186,442,344 | fc106a14c590a41e7749b93364d0e109d3b817935ae1390cdff35b0dabf35654 |
transformer/diffusion_pytorch_model-00012-of-00016.safetensors |
1,123,537,168 | 9e9896a1937320d4b0d684bbc26bd6b1570db2f139d74fcaec3826b06545bb5b |
transformer/diffusion_pytorch_model-00013-of-00016.safetensors |
1,212,667,120 | 9e5b9f520f30f4e4339a625761507c2e56a15a4373972945d690c951f6d7e2d7 |
transformer/diffusion_pytorch_model-00014-of-00016.safetensors |
1,149,761,896 | c475bddd0c934545262781b0d65d513a63b6a15c2c5651ad6aefd035cce89d3a |
transformer/diffusion_pytorch_model-00015-of-00016.safetensors |
1,186,493,584 | 1c5cfc8fa40a303c2a46ae53c0b453bdffadf1f7387d992a86787740e77388dd |
transformer/diffusion_pytorch_model-00016-of-00016.safetensors |
914,075,312 | a66812f617f28103b1452e2f519a9df6247e6fd92bebadba0d8e404efc90515a |
transformer/diffusion_pytorch_model.safetensors.index.json |
182,519 | 1fb31fa73e6866e4171253ffc5afb4225b6a54d55a02ca35889e17a17fd46cae |
transformer/int8_quantization.json |
128 | 1f6f6f3253cc416d8f49502ea089eda3914e9d9df05db9067b94e4ee2df811ea |
vae/config.json |
724 | 47e8bcf55e93e9c182e1962a8c7a0650faeb34ea0f66826d6f8aaa9f73e08ec9 |
vae/diffusion_pytorch_model.safetensors |
507,591,892 | d6e524b3fffede1787a74e81b30976dce5400c4439ba64222168e607ed19e793 |
Licence
Apache-2.0, as the original (LICENSE). The model is the work of the Robbyant
team; see the original repository and its paper,
arXiv:2601.20540.
- Downloads last month
- -
Model tree for ogtsvc/lingbot-world-fast-diffusers-int8
Base model
robbyant/lingbot-world-fast-diffusers