Nool-Alpha-100M-Chat

Nool-Alpha-100M-Chat is an efficient bilingual (Indonesian & English) and Python code causal language model based on an innovative hybrid architecture:

  • GSLA (Grouped-Subspace Latent Attention): Compresses KV-cache by 87.8% compared to standard dense MHA.
  • HFK-MoE (Heterogeneous Factorized MoE): Dense SwiGLU shared anchor + 8 low-rank factorized experts (r=96, Top-2 routing) saving 21.3% FLOPs.
  • Global Residual Highway: Stabilizes deep residual pathways with $\tanh(\alpha) \cdot \text{RMSNorm}(x_0)$.
  • Logit Soft-Capping: 30.0 tanh threshold preventing logit divergence.

Official Code Repository: GitHub - Ch3nOff/Nool-Alpha

Model Details

  • Architecture: nool_alpha
  • Total Parameters: 149.8M (111M)
  • Active Parameters / Token: ~97.9M
  • Vocabulary Size: 50257 (GPT-2 BPE)
  • Context Length: 2048 tokens
  • Training Step: 900
  • Recorded Loss: 2.633549153804779

Files Included

  • model.safetensors: Model weights in safe, zero-copy format.
  • config.json: Complete architectural hyper-parameters.
  • Tokenizer files: vocab.json, merges.txt, tokenizer.json, etc.

Quick Start (PyTorch)

import torch
from safetensors.torch import load_file
from nool_alpha.config import NoolAlphaConfig
from nool_alpha.model import NoolAlphaForCausalLM

# Load configuration and weights
config = NoolAlphaConfig.from_dict({
  "architectures": [
    "NoolAlphaForCausalLM"
  ],
  "model_type": "nool_alpha",
  "vocab_size": 50257,
  "d_model": 768,
  "n_layers": 10,
  "num_heads": 12,
  "head_dim": 64,
  "d_c": 192,
  "d_pe": 32,
  "rope_theta": 500000.0,
  "max_position_embeddings": 2048,
  "sliding_window": 512,
  "swa_interval": 4,
  "shared_ffn_dim": 1536,
  "num_experts": 8,
  "top_k_experts": 2,
  "expert_rank": 96,
  "moe_aux_loss_coeff": 0.01,
  "highway_alpha_init": 0.05,
  "logit_soft_cap": 30.0,
  "rms_norm_eps": 1e-06,
  "tie_word_embeddings": true,
  "checkpoint_step": 900,
  "checkpoint_loss": 2.633549153804779
})
model = NoolAlphaForCausalLM(config)
state_dict = load_file("model.safetensors")
model.load_state_dict(state_dict)
model.eval()
Downloads last month
112
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support