J.A.R.V.I.S. Titan 14.8B MoE — Milestone M4 Standalone Merged Model

Official 100% standalone merged weights combining J.A.R.V.I.S. Titan 14.8B DeepSeekMoE with the Milestone M4 Calibrated Tri-Brid Adapter.

Architecture & Specifications

  • Backbone: DeepSeekMoE 14.8B (1 Shared Expert + 8 Routed Experts, Top-2 active routing).
  • Tri-Brid Strategic Bridge Layers: Layers [3, 7, 11, 15, 19, 23, 27] (7 memory bridge checkpoints).
  • Tier 1 (Local SWA): Sliding Window Attention ($W = 2048$) with zero-copy GQA.
  • Tier 2 (Salient Reservoir): Exact KV retrieval subspace ($D = 512$).
  • Tier 3 (Titans Neural Memory): Bounded test-time learning recurrence ($\eta=10^{-3}$, $\rho=10^{-4}$, $\mu=0.95$, $||M||_F \le 50.0$).
  • MAG-3 Adaptive Gating: Calibrated routing distribution targeting [$L$: ~55%, $R$: ~25%, $M$: ~20%].
  • 100% Zero-Loss Passthrough: Unbroken backbone residual stream for stable autoregressive generation.
Downloads last month
681
Safetensors
Model size
15B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dhanesh-hf/Jarvis-Titan-M4-Merged

Finetuned
(6)
this model