๐Ÿ›๏ธ Atti: Cross-Architecture Reasoning Model

Atti is a 3.2B parameter reasoning model built on top of meta-llama/Llama-3.2-3B-Instruct.

Unlike conventional post-trained models that rely on standard reinforcement learning or extensive instruction fine-tuning datasets, Atti was produced without traditional training.

Instead, it was synthesized using **Isometric Manifold Transport (IMT)**โ€”a framework that directly transplants the post-trained reasoning manifolds of advanced Qwen models into the geometric coordinate space of Llama-3.2.


๐Ÿ”ฌ Key Characteristics

  • Architecture: Llama-3.2 (3.2B Parameters)
  • Underlying Method: Isometric Manifold Transport (IMT) โ€” Cross-architecture functional projection.
  • Autonomous Chain-of-Thought: Generates native, unprompted <think> ... </think> exploration passes.
  • Adaptive Test-Time Compute: Automatically scales thinking trace length based on problem difficulty (from ~850 words on arithmetic to ~5,000 words on Olympiad mathematics).
  • Format Stability: Solves problems by isolating scratchpad exploration from final synthesized answers (\boxed{}).

๐Ÿ“Š Comprehensive Benchmark Results

Evaluated using vLLM under official reasoning inference parameters: temperature = 0.6, top_p = 0.95, max_tokens = 16384, kv_cache_dtype = fp8.

Performance Overview

Benchmark Single-Pass (Pass@1) Majority Consensus (N=4) Search Ceiling (Pass@4) Avg Thinking Length
GSM8K 67.30% 1041/1319 (78.92%) 1186/1319 (89.92%) 885.9 words
MATH-500 26.80% 164/500 (32.80%) 221/500 (44.20%) 2070.8 words
AIME 1.67% 2/90 (2.22%) 5/90 (5.56%) 4259.4 words
ARC-Challenge 70.61% 920/1172 (78.50%) 1098/1172 (93.69%) 498.0 words

โš”๏ธ Model Comparison

How Atti compares against stock 3B-class instruction models across key benchmarks:

Model GSM8K (Pass@1) GSM8K (Consensus) MATH-500 (Pass@1) AIME (Pass@4) Native <think>
Atti (3.2B) 70.15% 80.52% 27.85% 4.44% Yes
Llama-3.2-3B-Instruct 68.80% 73.20% 26.20% 1.10% No
Qwen-2.5-3B-Instruct 72.40% 76.80% 29.50% 2.20% No

๐Ÿš€ Quickstart & Inference

You can run Atti locally using either Hugging Face Transformers or vLLM.

Using Transformers

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "BIBLIOKLEPT/Atti"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

prompt = "A farmer has 17 sheep, and all but 9 die. How many sheep are left?"
messages = [{"role": "user", "content": prompt}]

formatted = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(formatted, return_tensors="pt").to("cuda")

outputs = model.generate(
    **inputs,
    max_new_tokens=4096,
    temperature=0.6,
    top_p=0.95,
    do_sample=True
)

response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=False)
print(response)

Using vLLM (High-Throughput)

from vllm import LLM, SamplingParams

llm = LLM(
    model="BIBLIOKLEPT/Atti",
    max_model_len=16384,
    dtype="bfloat16"
)

sampling_params = SamplingParams(
    temperature=0.6,
    top_p=0.95,
    max_tokens=8192,
    skip_special_tokens=False
)

outputs = llm.generate(["Solve: If 5 machines take 5 minutes to make 5 widgets, how long for 100 machines to make 100 widgets?"], sampling_params)
print(outputs[0].outputs[0].text)

๐Ÿ“œ Citation & Attribution

If you use or build upon Atti in your research:

@misc{atti2026,
  author = {BIBLIOKLEPT},
  title = {Atti: Cross-Architecture Reasoning Model via Isometric Manifold Transport},
  year = {2026},
  publisher = {Hugging Face},
  journal = {Hugging Face Hub},
  howpublished = {\url{https://huggingface.co/BIBLIOKLEPT/Atti}}
}
Downloads last month
22
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for BIBLIOKLEPT/Atti

Finetuned
(2019)
this model