Ather 1 (22.1M Parameters)

Ather 1 is a compact, high-efficiency causal language model trained for code synthesis, algorithm reasoning, and software engineering. Developed by Atherium Labs for Milo.

Model Summary

  • Parameters: 22,146,048
  • Architecture: Decoder-only Transformer with Rotary Position Embeddings (RoPE)
  • Attention: Grouped-Query Attention (GQA) with 8 query heads and 2 key-value heads
  • Feed-Forward: SwiGLU Gated Multi-Layer Perceptron
  • Sequence Length: 512 tokens
  • Vocabulary: 576 byte-level BPE tokens (guarantees 0% out-of-vocabulary rate)
  • Trained With: PyTorch, AdamW, Cosine Learning Rate Schedule, Target Masking SFT

Architecture Specifications

Hyperparameter Value
Embedding Dimension (dim) 512
Hidden Dimension (hidden_dim) 1,360
Transformer Layers (n_layers) 8
Query Heads (n_heads) 8
Key-Value Heads (n_kv_heads) 2
Position Embeddings RoPE (base = 10,000)
Normalization RMSNorm (eps = 1e-5)
Tied Embeddings Yes

Inference Example

You can run Ather 1 directly using the included standalone scripts:

import torch
import json
from safetensors.torch import load_file
from standalone_model import Ather1Model
from standalone_tokenizer import StandaloneAtherTokenizer

# 1. Load configuration and weights
with open("config.json", "r") as f:
    config = json.load(f)

tokenizer = StandaloneAtherTokenizer("vocab.json")
model = Ather1Model(config)
model.load_state_dict(load_file("model.safetensors"))
model.eval()

# 2. Encode prompt
prompt = "<|user|>\nWrite a Python function to reverse a string\n<|assistant|>\n"
prompt_tokens = tokenizer.encode(prompt, add_bos=True)

# 3. Generate completion
out_tokens = model.generate(
    prompt_tokens=prompt_tokens,
    max_new_tokens=256,
    temperature=0.7,
    eos_token_id=tokenizer.eos_token_id,
)

response = tokenizer.decode(out_tokens[len(prompt_tokens):], skip_special_tokens=True)
print(response)

API Access

Ather 1 is hosted live on Cloudflare Pages and through the Atherium AI platform:

Downloads last month
23
Safetensors
Model size
22.3M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support