muse-glimmer-30b

A tt-model container package: the serving platform ships as a Docker image, so a consumer needs only Docker and a Tenstorrent card β€” no tt-metal, no vLLM, no venv on the host.

Serve it

tt-model pull  tt-hous/muse-glimmer-30b
tt-model serve tt-hous/muse-glimmer-30b

Quickstart

Verify it is up

curl -s localhost:8000/v1/models

First boot JIT-compiles kernels (~4 min of weight loading and compilation) before Application startup complete.

Serve profiles

One image serves every profile below; pick one with --profile.

profile hardware mesh max_num_seqs max_model_len
default (default) p300x2 P300x2 32 131072

Concurrency capacity

The KV pool holds 1,050,624 tokens (16,416 blocks x 64, 14.86 GB/device, BFLOAT8_B). Concurrent sequences are bounded by that pool, and capped at 32 by the decode op.

context per sequence blocks/seq max concurrent
640 10 32
4,608 72 32
16,896 264 32
33,280 520 31
66,048 1,032 15
131,072 2,048 8

Performance

Measured on P300x2 against the pinned tt-metal commit. Medians. Batch-1 rows use the in-process generator; concurrency rows use the vLLM OpenAI server.

Batch 1

ISL OSL conc TTFT TPOT E2EL t/s/u
128 512 1 65.0 ms 23.59 ms 12.12 s 42.38
1,024 512 1 145.9 ms 24.97 ms 12.91 s 40.05
4,096 512 1 458.1 ms 26.64 ms 14.07 s 37.54
8,192 512 1 918.7 ms 27.88 ms 15.17 s 35.87
16,384 512 1 2.08 s 30.33 ms 17.58 s 32.97
32,768 512 1 4.51 s 35.24 ms 22.52 s 28.37
65,536 512 1 10.27 s 45.12 ms 33.32 s 22.16
130,560 512 1 25.59 s 64.79 ms 58.70 s 15.43

Concurrency scaling β€” ISL 1,024, OSL 512

ISL OSL conc TTFT TPOT E2EL t/s/u out tok/s
1,024 512 1 150.9 ms 23.84 ms 12.33 s 41.94 41.5
1,024 512 2 294.5 ms 23.85 ms 12.48 s 41.92 82.0
1,024 512 4 582.8 ms 23.84 ms 12.77 s 41.94 160.4
1,024 512 8 1.17 s 23.86 ms 13.37 s 41.91 306.5
1,024 512 16 2.33 s 24.06 ms 14.63 s 41.56 559.8
1,024 512 32 4.67 s 24.88 ms 17.39 s 40.19 942.0

Maximum concurrency per context β€” OSL 512

ISL OSL conc TTFT TPOT E2EL t/s/u out tok/s
128 512 32 2.16 s 23.55 ms 14.19 s 42.47 1,154.5
4,096 512 32 14.87 s 26.48 ms 28.43 s 37.77 576.3
16,384 512 32 53.11 s 59.05 ms 83.22 s 16.94 196.9
32,768 512 31 96.93 s 125.37 ms 161.00 s 7.98 98.6
65,536 512 15 116.37 s 116.74 ms 176.03 s 8.57 43.6
130,560 512 8 145.17 s 169.17 ms 231.62 s 5.91 17.7

Concurrent requests prefill serially: chunked prefill is disabled for this model, so TTFT at concurrency N includes the prefills queued ahead of it.

What is inside

  • weights: meta-models/Muse-Glimmer-30B β€” downloaded to your HF cache at pull time, never baked into the image
  • arch: blackhole
  • serving stack: vllm-plugin

Provenance

Everything below is pinned; the image was built from exactly these.

component pinned to
tt-metal 0dd37ce6ee33826ebb8ce23a5d83a45bca7d6b29 (dirty tree)
plugin a48857ac68b17c31303e4809f348caaebbf10f74
code digest 55af55ae2e769dc2
built 2026-08-31T22:16:34+00:00 by tt-model 0.1.0

Shipped code

code/ in this repo is byte-identical to what runs inside the image.

  • models/common/
  • models/tt_transformers/
  • models/autoports/meta_models_muse_glimmer_30b/tt/
  • models/autoports/meta_models_muse_glimmer_30b/doc/datatype_sweep/
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support