TR-HASH MoE 100M - 70B Agentic Refinement

This repository is the public checkpoint archive for the lexical refinement stage of the 100,366,720-parameter TR-HASH MoE language model.

Status: waiting for pretraining completion. Refinement starts only after the 125B-token pretraining run exits successfully and its final checkpoint has been verified locally.

Training lineage

  1. Pretraining: 70B unique tokens plus 55B proportional replay, totaling 124,999,598,080 trained tokens.
  2. Refinement: one clean pass over the exact 70B unique-token core, initialized from the completed pretraining weights with a fresh optimizer.
  3. Instruction tuning: a separate full-parameter SFT stage after refinement.

The refinement keeps the same 32,000-token vocabulary and deterministic routing topology. It does not add new tokens or change the expert assignments.

Architecture

  • Parameters: 100,366,720
  • Layers: 10
  • Hidden size: 640
  • Attention: causal GQA, 10 query heads and 2 KV heads
  • Experts: 4, deterministic top-2 token-ID multi-hash routing
  • Context length: 2,048 tokens
  • Precision: BF16

Resources

Intermediate checkpoints are research artifacts, not instruction-tuned chat models or production-ready releases. Final evaluation results will be added after refinement completes.

License

Released under CC BY-NC 4.0.

Downloads last month
1,340
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AETHORIA-AI/TR-HASH-MoE-100M-70B-Agentic-Refinement

Finetuned
(1)
this model
Finetunes
1 model

Dataset used to train AETHORIA-AI/TR-HASH-MoE-100M-70B-Agentic-Refinement