Blaster-Think

Blaster-Think is a snapshot of the v35 hearing ship tip (checkpoint-116000) plus the Loop 2 typed-think decode contract. Weights are the same merged serving export as live Blaster at that tip. Thinking is not extra LoRA in this dump: the gateway prefills Answer: <think> on cued math and strips the span before TTS.

This repo is a separate Hub model so later GRPO / dual-LoRA work cannot overwrite the hearing ship:

SERVE_TIP on the live box stays outputs/molmo-audio-lora-diar-d-v35/checkpoint-116000. Think CE at 117500 was aborted (hearing dual-eval HEARING_NO_SHIP). Do not serve 117500.

What’s inside

Same files as the diar-d serving merge (model-*-of-*.safetensors, audio_modules.pt, tokenizer, remote-code sources), plus:

Path Purpose
serving/think_scaffold.py Loop 2 P0/P1 cue, 192-token body, force-close, loop XOR
eval/think_eval_116000_gsm_192.json Scaffolded GSM heldout at this tip
SNAPSHOT.json Pin: tip id, parent, abort notes

Think eval (this tip, scaffolded GSM, n=20)

Gate Result
Format (<think></think>) 20/20
Exact 8/20 (worded finals and a few real misses; not a ship gate for think)
Unearned correct 0
Uncued leak 0

Exact is a format contract + weak arithmetic, not a reason to invert training from base Molmo-7B-D.

Intended use

  • Restore this exact 116000 merge if GRPO or a second think LoRA goes wrong
  • Research on cued typed thinking without replacing live Blaster

License

CC BY-NC-SA 4.0. Private unless the owner changes visibility.

Marketing name: Blaster-Think. Technical Hub id: molmo-audio-serving-blaster-think.

Downloads last month
6
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 0x8badbeef/molmo-audio-serving-blaster-think

Base model

Qwen/Qwen2-7B
Finetuned
(4)
this model