MiraLM-47M

A 47,640,968-parameter hybrid language model trained from scratch โ€” no pretrained weights, no distillation โ€” under a hard 50,000,000 cap, on a single Tesla T4. Interleaved Mamba โˆฅ attention blocks (Jamba-style) with a sparse top-2-of-8 mixture of experts whose router is domain-seeded by a guide loss.

This repository holds the step-11,000 best checkpoint (train loss 2.2113, perplexity 9.13) โ€” the run that carries the measured results below. The step-17,000 "last" checkpoint diverged (loss 2.587) and is deliberately not published.

Measured results

Model HellaSwag ARC-Easy PIQA WinoGrande WikiText-103 PPL
MiraLM-47M 33.0 27.0 50.0 50.2 2837.6

Read honestly: at 47.6M parameters and 45,056,000 tokens โ€” roughly 4.7% of the Chinchilla-optimal budget for this size โ€” the model is undertrained by construction. Multiple-choice commonsense sits near its random floor, and the WikiText-103 perplexity reflects a domain shift, since the training corpus is code / math / web text, not prose. The claim is not "we win on HellaSwag". It is that the whole pipeline โ€” architecture, budget enforcement, data, routing, structured SFT, evaluation โ€” runs end to end, reproduces from one config, and fits the ceiling.

Load it

git clone https://github.com/qtttyr/MiraLM && cd MiraLM
from src.model.hf_interface import MiraLMForCausalLM  # registers model_type "mira"
from transformers import AutoTokenizer

repo = "vaprooll/MiraLM-47M"
model = MiraLMForCausalLM.from_pretrained(repo).eval()
tok   = AutoTokenizer.from_pretrained(repo)

prompt = "<|json|> Return a JSON object with product, price and stock: product mug, price 500, stock 23."
ids = tok(prompt, return_tensors="pt")
out = model.generate(**ids, max_new_tokens=64, do_sample=False)
print(tok.decode(out[0][ids["input_ids"].shape[1]:]))

No trust_remote_code needed: importing hf_interface registers the mira model type locally, so from_pretrained resolves the checkpoint.

Downloads last month
204
Safetensors
Model size
48M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support