DragonData-FinDense-1.2B
The pilot that proves the pipeline. Same architecture family, tokenizer, and data
as the flagship β at 1.2B it iterates in days, de-risks every recipe decision, and
graduates into DragonData-FinDense-3B.
Status: PLANNED (S2 starts after tokenizer v1 lands). Targets with gates; checked = measured.
Why a pilot
- Proves tokenizer β pretrain β eval loop end-to-end before burning flagship compute.
- Recipe lab: LR schedules, batch mix, grounding-triple ratios are tuned here, frozen for 3B.
- Promotion rule: pilot must beat a size-matched baseline on FinQA slice + hold citation rate, or 3B does not start.
Spec (planned)
- ~18 layers, hidden 2048, GQA; ctx 4096. Random-init, own 64k finance BPE.
- 10β20B tokens, overtrained past Chinchilla minimum.
Where to find things
βββ README.md β you are here
βββ RUN.md β repro (lands with first checkpoint)
βββ eval/ β score sheets
βββ checkpoints/β dated releases
How to use
Releases will ship bf16 weights + GGUF quants + tokenizer + RUN.md. Nothing yet β watch this space. Data: DragonData-Finance-Corpus. Flagship: DragonData-FinDense-3B.
License
Apache 2.0.
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support