DragonData-FinMoE-8x3B (ROADMAP)
8 experts, ~3B active per token, one public finance brain. Built on the
DragonData-FinDense-3B base: experts specialize by language/domain (incl. a CN
expert), routed per sample. Same training budget as dense 3B, far more capacity.
Status: PLANNED (S5 β starts after flagship pretrain + distill). This repo is the public roadmap so the trajectory is auditable from day one.
Locked method
- Dense 3B pretrains first (no SFT/DPO before MoE).
- 8 experts freshly init on the dense base; MoE pretraining with sample-expert-masking (each sample routes to a subset; one permissive-license sub-domain per expert).
- After MoE pretrain + evals: joint SFT β DPO (masking OFF, all experts + base update).
- Defect scan β safetensors + GGUF β full token/GPU-hour records on this card.
Where to find things (lands at S5)
βββ README.md βββ RUN.md βββ eval/ βββ checkpoints/ βββ gguf/
Base: DragonData-FinDense-3B. Data: DragonData-Finance-Corpus.
License
Apache 2.0.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support