Mind-v1: Deterministic Reasoning Model (35.8M Parameters)
Mind-v1 is a compact 35.8-million parameter decoder-only transformer trained from scratch on algorithmic and logical reasoning problems generated by The Mind autonomous protocol on Solana.
Unlike marketing-heavy releases, Mind-v1 does not claim to replace frontier foundation models. It is a purpose-built, lightweight, 100% open-weights reasoning model designed to run with sub-100ms latency on commodity CPU hardware with zero API dependencies.
Model Architecture
Mind-v1 uses a modern LLaMA-style decoder-only transformer architecture with SwiGLU activations, RMSNorm, and Rotary Position Embeddings (RoPE), configured with an independent custom Byte-Level BPE tokenizer.
| Parameter | Value | Note |
|---|---|---|
| Total Parameters | 35,756,544 (~35.8M) | Exact non-embedding + embedding weights |
| Layers (Transformer Blocks) | 6 | Decoder-only self-attention |
| Hidden Size ($d_{\text{model}}$) | 512 | Representation dimension |
| Intermediate Size | 1,376 | SwiGLU feed-forward projection |
| Attention Heads | 8 | 64 dimensions per head |
| Vocabulary Size | 16,384 | Custom Byte-Level BPE tokenizer |
| Context Window | 2,048 tokens | Trained on sequence length 128 |
| Position Embeddings | RoPE | Rotary position embeddings |
| Normalization | RMSNorm | $\epsilon = 10^{-5}$ |
| Weight Format | SafeTensors | 143 MB (model.safetensors, float32) |
Training & Loss Convergence
The model was trained from scratch with randomly initialized weights (no pre-trained base model weights were used).
- Training Dataset:
ShoggothMind/mind-judgements(synthetic_trainsplit, 3,000 verified reasoning examples across 8 task types and 10 difficulty tiers). - Validation Dataset: 500 held-out examples.
- Optimizer: AdamW ($\beta_1=0.9, \beta_2=0.95$, weight decay = $0.01$, learning rate = $3\times 10^{-4}$).
- Scheduler: Cosine Annealing with 100 warmup steps.
- Initial Loss:
9.7041(accurately matches theoretical uniform random baseline $-\ln(1/16384) = 9.704$). - Final Training Loss:
1.0608 - Final Validation Loss:
1.2210 - Training Time: ~14.9 minutes on commodity hardware.
Training Loss Curve
Evaluation Benchmark Against Frontier Models
We evaluated Mind-v1 against six frontier models accessible via OpenRouter on an identical held-out evaluation slice from test.jsonl.
| Model | Provider / Type | Parameters | Exact-Match Accuracy | Avg Latency (CPU/API) | Cost per 1M Tokens | Memory Footprint | Open Weights |
|---|---|---|---|---|---|---|---|
| Mind-v1 (Ours) | Local / Open Weights | 35.8M | 6.0% (Direct output) | 96.0 ms (CPU) | $0.00 | 143 MB | Yes |
| DeepSeek V3 | Cloud MoE API | 671B (37B active) | 20.0% | 5,080 ms | $0.14 / $0.28 | Cloud | No |
| Qwen 2.5 72B | Cloud API | 72B | 60.0% | 1,410 ms | $0.35 / $0.40 | Cloud | Yes (Weights) |
| Claude 3.5 Sonnet | Cloud API | Proprietary | 0.0%* | 350 ms | $3.00 / $15.00 | Cloud | No |
| GPT-4o | Cloud API | Proprietary | 80.0% | 1,930 ms | $2.50 / $10.00 | Cloud | No |
| Llama 3.3 70B | Cloud API | 70B | 60.0% | 2,300 ms | $0.30 / $0.40 | Cloud | Yes (Weights) |
| Mistral Large | Cloud API | Proprietary | 0.0%* | 480 ms | $2.00 / $6.00 | Cloud | No |
*Note: 0% results on some frontier models occurred due to conversational refusal formatting or answering outside strict <answer> tags on single-token generation constraints.
Mind-v1 Task Breakdown
- Logic & Relations: 30.0%
- Letter Counting: 10.0%
- Weekday Calculation: 10.0%
- Number Sequences: 5.0%
- Test Set Cross-Entropy Loss:
1.8986 - Test Set Perplexity:
6.68
Honest Analysis: Strengths vs Limitations
Strengths
- Zero Cost & Offline Capability: Runs entirely on a single CPU core in 143 MB of RAM without external API keys or cloud billing.
- Deterministic Output Tags: Consistently formats answers within
<answer>...</answer>tags as learned during training. - Low Latency: ~96 ms per inference on commodity x86_64 CPU.
- Reproducible & Verifiable: Custom 16,384 BPE tokenizer and training pipeline completely open source in the project repository.
Limitations
- Model Capacity: At 35.8 million parameters, the model cannot perform complex multi-step mental arithmetic (e.g., multi-digit multiplication or modular exponentiation) without an external chain-of-thought scratchpad.
- Domain Focus: Trained strictly on synthetic math, sequence, logic, and calendar problems. It has no open-domain world knowledge, literature, or trivia understanding.
- Context Length: Context is configured to 2,048 tokens, trained on 128-token lengths, and is not suitable for document summarization.
Quickstart & Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "ShoggothMind/mind-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)
prompt = "Question: What is 45 + 17?\\nAnswer: <answer>"
inputs = tokenizer(prompt, return_tensors="pt")
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=16,
do_sample=False,
pad_token_id=tokenizer.pad_token_id,
eos_token_id=tokenizer.eos_token_id
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Protocol Integration
To query the live consensus swarm of The Mind on Solana via HTTP API instead of running weights locally:
python examples/query_the_mind.py --url https://shoggothmind.live --question "What is the consensus truth?"
License
Apache License 2.0 (Apache-2.0).
- Downloads last month
- 430
