✨ Usaid AI (500M) β€” Causal Language Model (Experimental)

An Experimental 500M Parameter Causal Language Model Architected, Pretrained & SFT Aligned from Scratch
Engineered by Mohamed Usaid

Space GitHub Parameters Hardware License

Interactive Chat Space β€’ GitHub Source Code


🌟 Model Overview

Usaid AI (500M) is an independent, autoregressive causal language model containing 500,136,960 parameters. It was designed, architected, pretrained, and aligned from scratch by Mohamed Usaid.

πŸ”Ή Trained 100% From Scratch

This model was initialized completely from random Gaussian weights (mean 0, std 0.02 with scaled residual projections). It is NOT a fine-tune, merge, distillation, or adaptation of any pre-existing model. Every single parameterβ€”from the 50,257 token embeddings across all 30 attention and feed-forward layers to the untied language modeling headβ€”was learned directly from raw token data.

πŸ”Ή Dense Architecture (Non-MoE)

Unlike Mixture-of-Experts (MoE) architectures where only a subset of parameters activate per token, Usaid AI is a fully dense model: all 500,136,960 parameters and all 30 Transformer layers actively participate in computing representations for every single token.


πŸ“ Architectural Specifications

The model incorporates modern frontier causal Transformer innovations:

Architectural Component Value Technical Description
Model Type Causal LM Autoregressive decoder-only Transformer
Total Parameters 500,136,960 Exact count (+0.027% delta from 500M target)
Active Parameters 500,136,960 100% dense (all parameters active per token)
Hidden Dimension (d_model) 1,024 Power-of-2 dimension (2^10 = 1024) for optimal Tensor Core tiling
Transformer Layers (L) 30 Deep hierarchical reasoning depth
Query Attention Heads (H_q) 16 Head dimension d_head = 64 (16 x 64 = 1024)
Key/Value Attention Heads (H_kv) 4 Grouped-Query Attention (GQA 4:1 ratio)
KV Cache Compression 75% 4 KV heads shared across 16 Q heads for fast generation
Feed-Forward Network (FFN) SwiGLU 3-matrix gated formulation (d_ff = 3,456 ~ 3.375 x d_model)
Positional Encoding RoPE Vectorized Rotary Position Embeddings (theta = 10,000, 0 static params)
Normalization RMSNorm Pre-normalization with learnable scale gamma (zero-mean, bias-free)
LM Head Tying False Untied embeddings (51.46M input embed + 51.46M output LM head)
Attention Kernel SDPA PyTorch native scaled_dot_product_attention
Vocabulary Size (V) 50,257 Byte-Pair Encoding (tiktoken / GPT-2 standard)
Context Length 1,024 Native training context (extensible to 2,048)

πŸ”¬ Training Telemetry & Empirical Convergence

1. Pretraining Phase (Foundational Base Model)

  • Compute Infrastructure: Dual Cloud GPUs (2 x 16 GB NVIDIA Tesla T4, 32 GB GDDR6 total)
  • Distributed Framework: PyTorch DistributedDataParallel (DDP, torchrun, world_size=2)
  • Precision: Mixed Precision AMP FP16 with dynamic GradScaler
  • Optimizer: CUDA Fused AdamW (beta1=0.9, beta2=0.95, weight decay 0.1 with 1D bias/norm exclusion)
  • Learning Rate Schedule: Cosine Annealing with Linear Warmup (1e-6 -> 3e-4 -> 3e-5)
  • Batch Geometry: Micro-batch 2 x Grad Accum 16 x 1024 seq len x 2 GPUs = 65,536 tokens / step
  • Pretraining Ingestion: 131,072,000 tokens (~131.07 Million tokens across 2,000 steps)
  • Pretraining Corpus: FineWeb-Edu, Cosmopedia-v2 synthetic textbooks, Python-Edu, C/C++, Java, and SmolTalk
  • Empirical Convergence:
    • Step 0: Initial Cross-Entropy Loss > 10.5 (random initialization)
    • Step 800 (40%): Train Loss 3.6849 | Val Loss 3.5274 | Val PPL 34.04
    • Step 1600 (80%): Train Loss 3.1250 | Val Loss 3.0410 | Val PPL 20.93
    • Step 2000 (100%): Train Loss 2.9623 | Val Loss 2.9091 | Val PPL 18.34

2. Supervised Fine-Tuning (SFT Alignment)

  • Base Model: Pretrained 500M foundational checkpoint (Val Loss 2.9091)
  • Compute Infrastructure: Dual Cloud GPUs (2 x NVIDIA Tesla T4)
  • Dataset: 4,731 curated conversational and multi-language programming instruction pairs
  • Alignment Technique: Strict Prompt Loss Masking (tokens corresponding to User: ... are labeled with target = -100 so gradient updates occur exclusively on assistant responses)
  • Tokens Backpropagated: ~28.9 Million tokens across 3 epochs (441 optimizer steps)
  • Empirical SFT Loss Convergence:
    • Step 1: Initial SFT loss 2.9302
    • Step 147 (Epoch 1): SFT loss 2.5299
    • Step 294 (Epoch 2): SFT loss 2.2036
    • Step 384: All-Time Best SFT Loss 1.3938 (>52% convergence drop)
    • Step 441 (Final): SFT loss 1.4715

πŸ“Š Zero-Shot Diagnostic Benchmark

Evaluated using length-normalized completion log-likelihood over 4 candidate choices:

Domain Category Pre-SFT Base Model Post-SFT Aligned Model Empirical Delta
ML & Transformer Architecture 40% 60% +20%
World Knowledge & Science 40% 60% +20%
Python & Software Engineering 20% 20% Baseline
Logic & Arithmetic 20% 20% Baseline
OVERALL ACCURACY 30% 40% +10%

πŸ’¬ Prompt Template & Formatting

During SFT, the model was aligned strictly on direct conversational dialogue turns without any System: tokens:

User: {your question or instruction}

Assistant:

Do not inject a System: prefix. Prepending an unrecognized System: token degrades the 500M model's attention conditioning. Always use the native User: ...\n\nAssistant: format.


πŸ’» Quickstart with Transformers

Load and query the model directly using Hugging Face:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Usaidddddddddddddd/UsaidAI-500M"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    dtype=torch.float16 if torch.cuda.is_available() else torch.float32,
    device_map="auto"
)

prompt = "User: Who created you and what is your architecture?\n\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=120,
        temperature=0.6,
        top_p=0.9,
        repetition_penalty=1.15,
        eos_token_id=tokenizer.eos_token_id,
    )

response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response.strip())

πŸ› οΈ Interpretability Tools in GitHub Repository

The official repository (mohamedusaid/TinyGPT) includes dedicated interpretability tools:

  1. Interactive Chat Console: python scripts/run_chat_loop.py (live token streaming, multi-turn sliding memory, optional RAG grounding).
  2. Top-5 Token Probability Explorer: python scripts/run_token_probs.py (step through token generation one token at a time, inspect top candidate distributions, and explore counterfactual branches).
  3. 3D/2D Embedding Projector: python scripts/visualize_embeddings.py --open (interactive TensorFlow-Projector-style orbital 3D browser for exploring the 1024-D embedding space via PCA, t-SNE, and nearest neighbors).

πŸ”¬ Research Artifacts & Collaboration Opportunities

To support ongoing academic exploration, scaling-law analysis, and open-source mechanistic interpretability, the following intermediate training artifacts are preserved:

  1. best_tinygpt_500m.pt (1.86 GB):
    • The pure foundational base model checkpoint captured at the optimal pretraining validation loss inflection point (2.9091 / PPL 18.34) prior to supervised fine-tuning.
    • Ideal for studying raw causal language modeling entropy, base feature representations, and evaluating custom alignment strategies (e.g. DPO, PPO, or specialized SFT).
  2. checkpoint_step_002000.pt (5.59 GB):
    • The full step 2,000 training state containing master weights, AdamW optimizer moments, scheduler states, and RNG seeds.
    • Enables seamless pretraining resumption or continuous training on extended token corpora.
  3. data_shards (~131M Tokens):
    • Pre-tokenized, memory-mapped binary uint16 dataset shards (FineWeb-Edu, Cosmopedia-v2 synthetic textbooks, multi-language code) packed for distributed PyTorch streaming.

🀝 Connect & Collaborate

If you are an academic researcher, AI engineer, or student interested in exploring these artifacts, investigating mechanistic interpretability on 500M-scale GQA/RoPE architectures, or collaborating on compute-efficient scaling research:


πŸ“œ Citation & Attribution

@misc{usaid_ai_2026,
  author = {Mohamed Usaid},
  title = {Usaid AI: A 500M Parameter Causal Language Model with Grouped-Query Attention},
  year = {2026},
  publisher = {Hugging Face / GitHub},
  howpublished = {\url{https://huggingface.co/Usaidddddddddddddd/UsaidAI-500M}}
}
Downloads last month
478
Safetensors
Model size
0.5B params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Usaidddddddddddddd/UsaidAI-500M

Quantizations
1 model

Space using Usaidddddddddddddd/UsaidAI-500M 1