- β¨ Usaid AI (500M) β Causal Language Model (Experimental)
- π Model Overview
- π Architectural Specifications
- π¬ Training Telemetry & Empirical Convergence
- π Zero-Shot Diagnostic Benchmark
- π¬ Prompt Template & Formatting
- π» Quickstart with Transformers
- π οΈ Interpretability Tools in GitHub Repository
- π¬ Research Artifacts & Collaboration Opportunities
- π Citation & Attribution
β¨ Usaid AI (500M) β Causal Language Model (Experimental)
An Experimental 500M Parameter Causal Language Model Architected, Pretrained & SFT Aligned from Scratch
Engineered by Mohamed Usaid
π Model Overview
Usaid AI (500M) is an independent, autoregressive causal language model containing 500,136,960 parameters. It was designed, architected, pretrained, and aligned from scratch by Mohamed Usaid.
πΉ Trained 100% From Scratch
This model was initialized completely from random Gaussian weights (mean 0, std 0.02 with scaled residual projections). It is NOT a fine-tune, merge, distillation, or adaptation of any pre-existing model. Every single parameterβfrom the 50,257 token embeddings across all 30 attention and feed-forward layers to the untied language modeling headβwas learned directly from raw token data.
πΉ Dense Architecture (Non-MoE)
Unlike Mixture-of-Experts (MoE) architectures where only a subset of parameters activate per token, Usaid AI is a fully dense model: all 500,136,960 parameters and all 30 Transformer layers actively participate in computing representations for every single token.
π Architectural Specifications
The model incorporates modern frontier causal Transformer innovations:
| Architectural Component | Value | Technical Description |
|---|---|---|
| Model Type | Causal LM | Autoregressive decoder-only Transformer |
| Total Parameters | 500,136,960 |
Exact count (+0.027% delta from 500M target) |
| Active Parameters | 500,136,960 |
100% dense (all parameters active per token) |
| Hidden Dimension (d_model) | 1,024 |
Power-of-2 dimension (2^10 = 1024) for optimal Tensor Core tiling |
| Transformer Layers (L) | 30 |
Deep hierarchical reasoning depth |
| Query Attention Heads (H_q) | 16 |
Head dimension d_head = 64 (16 x 64 = 1024) |
| Key/Value Attention Heads (H_kv) | 4 |
Grouped-Query Attention (GQA 4:1 ratio) |
| KV Cache Compression | 75% |
4 KV heads shared across 16 Q heads for fast generation |
| Feed-Forward Network (FFN) | SwiGLU | 3-matrix gated formulation (d_ff = 3,456 ~ 3.375 x d_model) |
| Positional Encoding | RoPE | Vectorized Rotary Position Embeddings (theta = 10,000, 0 static params) |
| Normalization | RMSNorm | Pre-normalization with learnable scale gamma (zero-mean, bias-free) |
| LM Head Tying | False | Untied embeddings (51.46M input embed + 51.46M output LM head) |
| Attention Kernel | SDPA | PyTorch native scaled_dot_product_attention |
| Vocabulary Size (V) | 50,257 |
Byte-Pair Encoding (tiktoken / GPT-2 standard) |
| Context Length | 1,024 |
Native training context (extensible to 2,048) |
π¬ Training Telemetry & Empirical Convergence
1. Pretraining Phase (Foundational Base Model)
- Compute Infrastructure: Dual Cloud GPUs (2 x 16 GB NVIDIA Tesla T4, 32 GB GDDR6 total)
- Distributed Framework: PyTorch DistributedDataParallel (DDP, torchrun, world_size=2)
- Precision: Mixed Precision AMP FP16 with dynamic GradScaler
- Optimizer: CUDA Fused AdamW (beta1=0.9, beta2=0.95, weight decay 0.1 with 1D bias/norm exclusion)
- Learning Rate Schedule: Cosine Annealing with Linear Warmup (1e-6 -> 3e-4 -> 3e-5)
- Batch Geometry: Micro-batch 2 x Grad Accum 16 x 1024 seq len x 2 GPUs = 65,536 tokens / step
- Pretraining Ingestion: 131,072,000 tokens (~131.07 Million tokens across 2,000 steps)
- Pretraining Corpus: FineWeb-Edu, Cosmopedia-v2 synthetic textbooks, Python-Edu, C/C++, Java, and SmolTalk
- Empirical Convergence:
- Step 0: Initial Cross-Entropy Loss > 10.5 (random initialization)
- Step 800 (40%): Train Loss
3.6849| Val Loss3.5274| Val PPL34.04 - Step 1600 (80%): Train Loss
3.1250| Val Loss3.0410| Val PPL20.93 - Step 2000 (100%): Train Loss
2.9623| Val Loss2.9091| Val PPL18.34
2. Supervised Fine-Tuning (SFT Alignment)
- Base Model: Pretrained 500M foundational checkpoint (Val Loss
2.9091) - Compute Infrastructure: Dual Cloud GPUs (2 x NVIDIA Tesla T4)
- Dataset: 4,731 curated conversational and multi-language programming instruction pairs
- Alignment Technique: Strict Prompt Loss Masking (tokens corresponding to
User: ...are labeled withtarget = -100so gradient updates occur exclusively on assistant responses) - Tokens Backpropagated: ~28.9 Million tokens across 3 epochs (441 optimizer steps)
- Empirical SFT Loss Convergence:
- Step 1: Initial SFT loss
2.9302 - Step 147 (Epoch 1): SFT loss
2.5299 - Step 294 (Epoch 2): SFT loss
2.2036 - Step 384: All-Time Best SFT Loss
1.3938(>52% convergence drop) - Step 441 (Final): SFT loss
1.4715
- Step 1: Initial SFT loss
π Zero-Shot Diagnostic Benchmark
Evaluated using length-normalized completion log-likelihood over 4 candidate choices:
| Domain Category | Pre-SFT Base Model | Post-SFT Aligned Model | Empirical Delta |
|---|---|---|---|
| ML & Transformer Architecture | 40% | 60% | +20% |
| World Knowledge & Science | 40% | 60% | +20% |
| Python & Software Engineering | 20% | 20% | Baseline |
| Logic & Arithmetic | 20% | 20% | Baseline |
| OVERALL ACCURACY | 30% | 40% | +10% |
π¬ Prompt Template & Formatting
During SFT, the model was aligned strictly on direct conversational dialogue turns without any System: tokens:
User: {your question or instruction}
Assistant:
Do not inject a
System:prefix. Prepending an unrecognizedSystem:token degrades the 500M model's attention conditioning. Always use the nativeUser: ...\n\nAssistant:format.
π» Quickstart with Transformers
Load and query the model directly using Hugging Face:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Usaidddddddddddddd/UsaidAI-500M"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
dtype=torch.float16 if torch.cuda.is_available() else torch.float32,
device_map="auto"
)
prompt = "User: Who created you and what is your architecture?\n\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=120,
temperature=0.6,
top_p=0.9,
repetition_penalty=1.15,
eos_token_id=tokenizer.eos_token_id,
)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response.strip())
π οΈ Interpretability Tools in GitHub Repository
The official repository (mohamedusaid/TinyGPT) includes dedicated interpretability tools:
- Interactive Chat Console:
python scripts/run_chat_loop.py(live token streaming, multi-turn sliding memory, optional RAG grounding). - Top-5 Token Probability Explorer:
python scripts/run_token_probs.py(step through token generation one token at a time, inspect top candidate distributions, and explore counterfactual branches). - 3D/2D Embedding Projector:
python scripts/visualize_embeddings.py --open(interactive TensorFlow-Projector-style orbital 3D browser for exploring the 1024-D embedding space via PCA, t-SNE, and nearest neighbors).
π¬ Research Artifacts & Collaboration Opportunities
To support ongoing academic exploration, scaling-law analysis, and open-source mechanistic interpretability, the following intermediate training artifacts are preserved:
best_tinygpt_500m.pt(1.86 GB):- The pure foundational base model checkpoint captured at the optimal pretraining validation loss inflection point (
2.9091/ PPL18.34) prior to supervised fine-tuning. - Ideal for studying raw causal language modeling entropy, base feature representations, and evaluating custom alignment strategies (e.g. DPO, PPO, or specialized SFT).
- The pure foundational base model checkpoint captured at the optimal pretraining validation loss inflection point (
checkpoint_step_002000.pt(5.59 GB):- The full step 2,000 training state containing master weights, AdamW optimizer moments, scheduler states, and RNG seeds.
- Enables seamless pretraining resumption or continuous training on extended token corpora.
data_shards(~131M Tokens):- Pre-tokenized, memory-mapped binary uint16 dataset shards (FineWeb-Edu, Cosmopedia-v2 synthetic textbooks, multi-language code) packed for distributed PyTorch streaming.
π€ Connect & Collaborate
If you are an academic researcher, AI engineer, or student interested in exploring these artifacts, investigating mechanistic interpretability on 500M-scale GQA/RoPE architectures, or collaborating on compute-efficient scaling research:
- Open a Discussion: Connect via the Hugging Face Community Discussions or GitHub Issues.
- GitHub Profile: @mohamedusaid
- Archive Repository: Usaidddddddddddddd/TinyGPT-500M-Archive
π Citation & Attribution
@misc{usaid_ai_2026,
author = {Mohamed Usaid},
title = {Usaid AI: A 500M Parameter Causal Language Model with Grouped-Query Attention},
year = {2026},
publisher = {Hugging Face / GitHub},
howpublished = {\url{https://huggingface.co/Usaidddddddddddddd/UsaidAI-500M}}
}
- Downloads last month
- 478