Mind-v1: Deterministic Reasoning Model (35.8M Parameters)

Mind-v1 is a compact 35.8-million parameter decoder-only transformer trained from scratch on algorithmic and logical reasoning problems generated by The Mind autonomous protocol on Solana.

Unlike marketing-heavy releases, Mind-v1 does not claim to replace frontier foundation models. It is a purpose-built, lightweight, 100% open-weights reasoning model designed to run with sub-100ms latency on commodity CPU hardware with zero API dependencies.


Model Architecture

Mind-v1 uses a modern LLaMA-style decoder-only transformer architecture with SwiGLU activations, RMSNorm, and Rotary Position Embeddings (RoPE), configured with an independent custom Byte-Level BPE tokenizer.

Parameter Value Note
Total Parameters 35,756,544 (~35.8M) Exact non-embedding + embedding weights
Layers (Transformer Blocks) 6 Decoder-only self-attention
Hidden Size ($d_{\text{model}}$) 512 Representation dimension
Intermediate Size 1,376 SwiGLU feed-forward projection
Attention Heads 8 64 dimensions per head
Vocabulary Size 16,384 Custom Byte-Level BPE tokenizer
Context Window 2,048 tokens Trained on sequence length 128
Position Embeddings RoPE Rotary position embeddings
Normalization RMSNorm $\epsilon = 10^{-5}$
Weight Format SafeTensors 143 MB (model.safetensors, float32)

Training & Loss Convergence

The model was trained from scratch with randomly initialized weights (no pre-trained base model weights were used).

  • Training Dataset: ShoggothMind/mind-judgements (synthetic_train split, 3,000 verified reasoning examples across 8 task types and 10 difficulty tiers).
  • Validation Dataset: 500 held-out examples.
  • Optimizer: AdamW ($\beta_1=0.9, \beta_2=0.95$, weight decay = $0.01$, learning rate = $3\times 10^{-4}$).
  • Scheduler: Cosine Annealing with 100 warmup steps.
  • Initial Loss: 9.7041 (accurately matches theoretical uniform random baseline $-\ln(1/16384) = 9.704$).
  • Final Training Loss: 1.0608
  • Final Validation Loss: 1.2210
  • Training Time: ~14.9 minutes on commodity hardware.

Training Loss Curve

Training Loss Curve


Evaluation Benchmark Against Frontier Models

We evaluated Mind-v1 against six frontier models accessible via OpenRouter on an identical held-out evaluation slice from test.jsonl.

Model Provider / Type Parameters Exact-Match Accuracy Avg Latency (CPU/API) Cost per 1M Tokens Memory Footprint Open Weights
Mind-v1 (Ours) Local / Open Weights 35.8M 6.0% (Direct output) 96.0 ms (CPU) $0.00 143 MB Yes
DeepSeek V3 Cloud MoE API 671B (37B active) 20.0% 5,080 ms $0.14 / $0.28 Cloud No
Qwen 2.5 72B Cloud API 72B 60.0% 1,410 ms $0.35 / $0.40 Cloud Yes (Weights)
Claude 3.5 Sonnet Cloud API Proprietary 0.0%* 350 ms $3.00 / $15.00 Cloud No
GPT-4o Cloud API Proprietary 80.0% 1,930 ms $2.50 / $10.00 Cloud No
Llama 3.3 70B Cloud API 70B 60.0% 2,300 ms $0.30 / $0.40 Cloud Yes (Weights)
Mistral Large Cloud API Proprietary 0.0%* 480 ms $2.00 / $6.00 Cloud No

*Note: 0% results on some frontier models occurred due to conversational refusal formatting or answering outside strict <answer> tags on single-token generation constraints.

Mind-v1 Task Breakdown

  • Logic & Relations: 30.0%
  • Letter Counting: 10.0%
  • Weekday Calculation: 10.0%
  • Number Sequences: 5.0%
  • Test Set Cross-Entropy Loss: 1.8986
  • Test Set Perplexity: 6.68

Honest Analysis: Strengths vs Limitations

Strengths

  1. Zero Cost & Offline Capability: Runs entirely on a single CPU core in 143 MB of RAM without external API keys or cloud billing.
  2. Deterministic Output Tags: Consistently formats answers within <answer>...</answer> tags as learned during training.
  3. Low Latency: ~96 ms per inference on commodity x86_64 CPU.
  4. Reproducible & Verifiable: Custom 16,384 BPE tokenizer and training pipeline completely open source in the project repository.

Limitations

  1. Model Capacity: At 35.8 million parameters, the model cannot perform complex multi-step mental arithmetic (e.g., multi-digit multiplication or modular exponentiation) without an external chain-of-thought scratchpad.
  2. Domain Focus: Trained strictly on synthetic math, sequence, logic, and calendar problems. It has no open-domain world knowledge, literature, or trivia understanding.
  3. Context Length: Context is configured to 2,048 tokens, trained on 128-token lengths, and is not suitable for document summarization.

Quickstart & Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "ShoggothMind/mind-v1"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

prompt = "Question: What is 45 + 17?\\nAnswer: <answer>"
inputs = tokenizer(prompt, return_tensors="pt")

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=16,
        do_sample=False,
        pad_token_id=tokenizer.pad_token_id,
        eos_token_id=tokenizer.eos_token_id
    )

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Protocol Integration

To query the live consensus swarm of The Mind on Solana via HTTP API instead of running weights locally:

python examples/query_the_mind.py --url https://shoggothmind.live --question "What is the consensus truth?"

License

Apache License 2.0 (Apache-2.0).

Downloads last month
430
Safetensors
Model size
35.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using ShoggothMind/mind-v1 1