Maba Logo

Maba-v1-600M

Sub-quadratic hybrid causal language model with linear recurrence and sliding GQA

Architecture Repository • Benchmarks • Quickstart

Model Architecture Parameters Context License


Maba-v1-600M is a 752M-parameter causal language model (600M active non-embedding backbone) trained on ~150 billion tokens. It uses a 3:1 hybrid design: 75% Gated DeltaNet-2 linear recurrence layers and 25% Grouped-Query Attention layers. This structure provides O(N) context scaling and cuts KV-cache memory requirements by 75% compared to standard Transformers.

Hardware implementations, CUDA kernels, and theoretical derivations are documented in the Maba Architecture Repository.


Benchmarks

Evaluations were conducted with lm-evaluation-harness on a single Tesla T4 (16 GB VRAM) in float16 precision with batch size 4. Total runtime was 2 hours 1 minute across 104,876 test items.

Maba-v1-600M Benchmark Results

Note on dual-pass inference: Dual-pass execution was disabled in this release due to a parameter bug in the training script. This was caught when training had already practically finished. All benchmark numbers reported above reflect standard single-pass inference. Dual-pass evaluation will be released in an upcoming checkpoint.

Benchmark Comparison

Benchmark Maba-v1-600M Qwen3-0.6B SmolLM2-360M Llama-3.2-1B
MMLU (Academic Knowledge, 0-shot) 49.2% 47.2% 35.8% 49.3%
GSM8K (Math Reasoning, 5-shot) 32.5% 43.0% 3.2% 26.2%
ARC-Challenge (Science Reasoning, 0-shot) 38.1% 42.3% 35.8% 46.2%
Winogrande (Context & Logic, 0-shot) 58.6% 59.2% 52.5% 61.2%
HellaSwag (Common Sense, 0-shot) 50.5% 53.8% 54.5% 68.8%

All Maba evaluations were executed via lm-evaluation-harness in float16 on a single Tesla T4 GPU.

MMLU Performance by Domain

Domain Accuracy
Social Sciences 55.31%
Applied & Professional 53.04%
STEM 45.54%
Humanities 45.14%

Top MMLU Disciplines

Discipline Accuracy
Marketing 77.78%
International Law 74.38%
US Foreign Policy 69.00%
Management 67.96%
High School Psychology 67.89%

Specifications

Attribute Value
Total Parameters 752M
Active Backbone Parameters 600M
Layers 24 (18 Gated DeltaNet-2 + 6 GQA)
Hidden Size 1024
Intermediate Size (SwiGLU) 2816
Attention Heads 16 Query / 8 Key-Value (GQA)
Vocabulary Size 248,320
Context Length Up to 32,768 tokens
Pretraining Volume ~150B tokens

Quickstart

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AndrewThompson1233/maba-v1-600m"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.float16,
    device_map="auto",
    trust_remote_code=True
)

messages = [
    {"role": "system", "content": "You are Maba, an assistant built on the Maba v1 architecture."},
    {"role": "user", "content": "Explain the difference between linear attention and quadratic attention."}
]

prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

outputs = model.generate(
    **inputs,
    max_new_tokens=128,
    do_sample=True,
    temperature=0.7,
    top_p=0.9
)

response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response)

Architecture Reference

For architecture design, training dynamics, benchmarks against quadratic baselines, and implementation details, visit the Maba Architecture Repository.

Downloads last month
290
Safetensors
Model size
0.8B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including AndrewThompson1233/maba-v1-600m

Evaluation results