Instructions to use dasilva333/rwkv7-g1-webgpu-prefabs with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RWKV
How to use dasilva333/rwkv7-g1-webgpu-prefabs with RWKV:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
RWKV-7 G1 WebGPU Pre-Quantized Prefabs
Pre-quantized CBOR .prefab models for WebGPU inference in the browser via @cryscan/web-rwkv-wasm and AIRI.
These prefabs are pre-quantized offline directly from DanielClough/rwkv7-g1-safetensors for fast, zero-degradation browser execution.
Available Models
| Model | Variant | Size | Format | Direct HF Download Link |
|---|---|---|---|---|
| RWKV-7 G1 0.4B | Int8 |
581 MB | .prefab |
rwkv7-g1d-0.4b-int8.prefab |
| RWKV-7 G1 0.4B | NF4 |
437 MB | .prefab |
rwkv7-g1d-0.4b-nf4.prefab |
| RWKV-7 G1 1.5B | Int8 |
1.88 GB | .prefab |
rwkv7-g1d-1.5b-int8.prefab |
| RWKV-7 G1 1.5B | NF4 |
1.28 GB | .prefab |
rwkv7-g1d-1.5b-nf4.prefab |
Why .prefab instead of SafeTensors?
- Instant browser startup: Deserializes directly into WebGPU buffers via
Session.from_prefab(data, SessionType.Chat)in ~5.8 seconds (vs 60s+ for in-browser quantization). - Zero in-browser compute degradation: Completely bypasses the on-the-fly shader quantization bug in browser WebGPU runtimes.
- Smaller downloads:
rwkv7-g1d-0.4b-int8.prefab: 581 MB (-35% vs 902 MB raw FP16).rwkv7-g1d-0.4b-nf4.prefab: 437 MB (-51% vs 902 MB raw FP16).rwkv7-g1d-1.5b-int8.prefab: 1.88 GB (-38% vs 3.06 GB raw FP16).rwkv7-g1d-1.5b-nf4.prefab: 1.28 GB (-58% vs 3.06 GB raw FP16).
Usage with @cryscan/web-rwkv-wasm
import init, { Session, SessionType } from '@cryscan/web-rwkv-wasm'
await init()
const res = await fetch('https://huggingface.co/dasilva333/rwkv7-g1-webgpu-prefabs/resolve/main/rwkv7-g1d-1.5b-nf4.prefab')
const buffer = new Uint8Array(await res.arrayBuffer())
const session = await Session.from_prefab(buffer, SessionType.Chat)
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support