J.A.R.V.I.S. TITAN 14.8B MoE โ€” MILESTONE M3 ULTRA-LONG ADAPTER

Official Milestone M3 (Phase 3) weights for J.A.R.V.I.S. Titan 14.8B DeepSeekMoE + Tri-Brid Memory Architecture, distilled from full dense attention to Sliding Window Attention ($W=2048$) on Google Cloud TPU v5e-8.

Distillation & Training Specifications

  • Base Model: dhanesh-hf/Jarvis-Titan-V14-MoE-Merged (14.75B MoE, 100% Frozen)
  • Adapter Initialization: dhanesh-hf/Jarvis-Titan-M2-TriBrid-Adapter (Phase 2 Tri-Brid)
  • Dataset Source: dhanesh-hf/jarvis-v10-rft-dataset (100% Real Non-NIAH Peer-Reviewed Papers & Code)
  • Sliding Window Size: $W = 2048$ tokens (capping KV cache at ~115 MB)
  • Strategic Layers: [3, 7, 11, 15, 19, 23, 27] (7 Memory Bridges)
  • Tier 2 Salient Reservoir: $R=1024$ slots ($H_Q=28, H_{KV}=4$)
  • Tier 3 Titans Neural Memory: $d=512$, Google Pallas TPU VMEM SRAM kernel
  • Triple-Gated Adaptive Fusion: MAG-3 ($g_{\text{local}}, g_{\text{res}}, g_{\text{mem}}$)
  • Distillation Loss: $(1 - 0.5) \mathcal{L}{\text{CE}} + 0.5 T^2 \mathcal{D}{\text{KL}}$ ($T=2.0$)
  • Total Adapter Parameters: 90,044,458 (90.04M)
  • Total Tokens Trained: 20,012,495
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for dhanesh-hf/Jarvis-Titan-M3-UltraLong-Adapter

Finetuned
(4)
this model