DragonData-FinMoE-8x3B (ROADMAP)

8 experts, ~3B active per token, one public finance brain. Built on the DragonData-FinDense-3B base: experts specialize by language/domain (incl. a CN expert), routed per sample. Same training budget as dense 3B, far more capacity.

Status: PLANNED (S5 β€” starts after flagship pretrain + distill). This repo is the public roadmap so the trajectory is auditable from day one.

Locked method

  1. Dense 3B pretrains first (no SFT/DPO before MoE).
  2. 8 experts freshly init on the dense base; MoE pretraining with sample-expert-masking (each sample routes to a subset; one permissive-license sub-domain per expert).
  3. After MoE pretrain + evals: joint SFT β†’ DPO (masking OFF, all experts + base update).
  4. Defect scan β†’ safetensors + GGUF β†’ full token/GPU-hour records on this card.

Where to find things (lands at S5)

β”œβ”€β”€ README.md β”œβ”€β”€ RUN.md β”œβ”€β”€ eval/ β”œβ”€β”€ checkpoints/ └── gguf/

Base: DragonData-FinDense-3B. Data: DragonData-Finance-Corpus.

License

Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support