muse-glimmer-30b
A tt-model container package: the serving platform ships as a Docker image, so a consumer needs only Docker and a Tenstorrent card β no tt-metal, no vLLM, no venv on the host.
Serve it
tt-model pull tt-hous/muse-glimmer-30b
tt-model serve tt-hous/muse-glimmer-30b
Quickstart
Verify it is up
curl -s localhost:8000/v1/models
First boot JIT-compiles kernels (~4 min of weight loading and compilation) before
Application startup complete.
Serve profiles
One image serves every profile below; pick one with --profile.
| profile | hardware | mesh | max_num_seqs | max_model_len |
|---|---|---|---|---|
default (default) |
p300x2 | P300x2 | 32 | 131072 |
Concurrency capacity
The KV pool holds 1,050,624 tokens (16,416 blocks x 64, 14.86 GB/device, BFLOAT8_B).
Concurrent sequences are bounded by that pool, and capped at 32 by the decode op.
| context per sequence | blocks/seq | max concurrent |
|---|---|---|
| 640 | 10 | 32 |
| 4,608 | 72 | 32 |
| 16,896 | 264 | 32 |
| 33,280 | 520 | 31 |
| 66,048 | 1,032 | 15 |
| 131,072 | 2,048 | 8 |
Performance
Measured on P300x2 against the pinned tt-metal commit. Medians. Batch-1 rows use the in-process generator; concurrency rows use the vLLM OpenAI server.
Batch 1
| ISL | OSL | conc | TTFT | TPOT | E2EL | t/s/u |
|---|---|---|---|---|---|---|
| 128 | 512 | 1 | 65.0 ms | 23.59 ms | 12.12 s | 42.38 |
| 1,024 | 512 | 1 | 145.9 ms | 24.97 ms | 12.91 s | 40.05 |
| 4,096 | 512 | 1 | 458.1 ms | 26.64 ms | 14.07 s | 37.54 |
| 8,192 | 512 | 1 | 918.7 ms | 27.88 ms | 15.17 s | 35.87 |
| 16,384 | 512 | 1 | 2.08 s | 30.33 ms | 17.58 s | 32.97 |
| 32,768 | 512 | 1 | 4.51 s | 35.24 ms | 22.52 s | 28.37 |
| 65,536 | 512 | 1 | 10.27 s | 45.12 ms | 33.32 s | 22.16 |
| 130,560 | 512 | 1 | 25.59 s | 64.79 ms | 58.70 s | 15.43 |
Concurrency scaling β ISL 1,024, OSL 512
| ISL | OSL | conc | TTFT | TPOT | E2EL | t/s/u | out tok/s |
|---|---|---|---|---|---|---|---|
| 1,024 | 512 | 1 | 150.9 ms | 23.84 ms | 12.33 s | 41.94 | 41.5 |
| 1,024 | 512 | 2 | 294.5 ms | 23.85 ms | 12.48 s | 41.92 | 82.0 |
| 1,024 | 512 | 4 | 582.8 ms | 23.84 ms | 12.77 s | 41.94 | 160.4 |
| 1,024 | 512 | 8 | 1.17 s | 23.86 ms | 13.37 s | 41.91 | 306.5 |
| 1,024 | 512 | 16 | 2.33 s | 24.06 ms | 14.63 s | 41.56 | 559.8 |
| 1,024 | 512 | 32 | 4.67 s | 24.88 ms | 17.39 s | 40.19 | 942.0 |
Maximum concurrency per context β OSL 512
| ISL | OSL | conc | TTFT | TPOT | E2EL | t/s/u | out tok/s |
|---|---|---|---|---|---|---|---|
| 128 | 512 | 32 | 2.16 s | 23.55 ms | 14.19 s | 42.47 | 1,154.5 |
| 4,096 | 512 | 32 | 14.87 s | 26.48 ms | 28.43 s | 37.77 | 576.3 |
| 16,384 | 512 | 32 | 53.11 s | 59.05 ms | 83.22 s | 16.94 | 196.9 |
| 32,768 | 512 | 31 | 96.93 s | 125.37 ms | 161.00 s | 7.98 | 98.6 |
| 65,536 | 512 | 15 | 116.37 s | 116.74 ms | 176.03 s | 8.57 | 43.6 |
| 130,560 | 512 | 8 | 145.17 s | 169.17 ms | 231.62 s | 5.91 | 17.7 |
Concurrent requests prefill serially: chunked prefill is disabled for this model, so TTFT at concurrency N includes the prefills queued ahead of it.
What is inside
- weights:
meta-models/Muse-Glimmer-30Bβ downloaded to your HF cache at pull time, never baked into the image - arch: blackhole
- serving stack:
vllm-plugin
Provenance
Everything below is pinned; the image was built from exactly these.
| component | pinned to |
|---|---|
| tt-metal | 0dd37ce6ee33826ebb8ce23a5d83a45bca7d6b29 (dirty tree) |
| plugin | a48857ac68b17c31303e4809f348caaebbf10f74 |
| code digest | 55af55ae2e769dc2 |
| built | 2026-08-31T22:16:34+00:00 by tt-model 0.1.0 |
Shipped code
code/ in this repo is byte-identical to what runs inside the image.
models/common/models/tt_transformers/models/autoports/meta_models_muse_glimmer_30b/tt/models/autoports/meta_models_muse_glimmer_30b/doc/datatype_sweep/