DragonData-FinDense-1.2B

The pilot that proves the pipeline. Same architecture family, tokenizer, and data as the flagship β€” at 1.2B it iterates in days, de-risks every recipe decision, and graduates into DragonData-FinDense-3B.

Status: PLANNED (S2 starts after tokenizer v1 lands). Targets with gates; checked = measured.

Why a pilot

  • Proves tokenizer β†’ pretrain β†’ eval loop end-to-end before burning flagship compute.
  • Recipe lab: LR schedules, batch mix, grounding-triple ratios are tuned here, frozen for 3B.
  • Promotion rule: pilot must beat a size-matched baseline on FinQA slice + hold citation rate, or 3B does not start.

Spec (planned)

  • ~18 layers, hidden 2048, GQA; ctx 4096. Random-init, own 64k finance BPE.
  • 10–20B tokens, overtrained past Chinchilla minimum.

Where to find things

β”œβ”€β”€ README.md   ← you are here
β”œβ”€β”€ RUN.md      ← repro (lands with first checkpoint)
β”œβ”€β”€ eval/       ← score sheets
└── checkpoints/← dated releases

How to use

Releases will ship bf16 weights + GGUF quants + tokenizer + RUN.md. Nothing yet β€” watch this space. Data: DragonData-Finance-Corpus. Flagship: DragonData-FinDense-3B.

License

Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support