J.A.R.V.I.S. Titan 14.8B MoE — Milestone M4 Standalone Merged Model
Official 100% standalone merged weights combining J.A.R.V.I.S. Titan 14.8B DeepSeekMoE with the Milestone M4 Calibrated Tri-Brid Adapter.
Architecture & Specifications
- Backbone: DeepSeekMoE 14.8B (1 Shared Expert + 8 Routed Experts, Top-2 active routing).
- Tri-Brid Strategic Bridge Layers: Layers
[3, 7, 11, 15, 19, 23, 27](7 memory bridge checkpoints). - Tier 1 (Local SWA): Sliding Window Attention ($W = 2048$) with zero-copy GQA.
- Tier 2 (Salient Reservoir): Exact KV retrieval subspace ($D = 512$).
- Tier 3 (Titans Neural Memory): Bounded test-time learning recurrence ($\eta=10^{-3}$, $\rho=10^{-4}$, $\mu=0.95$, $||M||_F \le 50.0$).
- MAG-3 Adaptive Gating: Calibrated routing distribution targeting [$L$: ~55%, $R$: ~25%, $M$: ~20%].
- 100% Zero-Loss Passthrough: Unbroken backbone residual stream for stable autoregressive generation.
- Downloads last month
- 681
Model tree for dhanesh-hf/Jarvis-Titan-M4-Merged
Base model
dhanesh-hf/Jarvis-Titan-V14-MoE-Merged