FinModernBERT-embed-large-v1

A 395M-parameter finance text-embedding model that beats the 7B Fin-E5 on FinMTEB Summarization (+0.109) and STS (+0.010) at 1/18th the size — built with ≈120 GPU-hours (~5 days) on a single 2× RTX 3090 workstation, through a fully decontaminated four-step pipeline: measured baseline → 5.65B-token domain MLM (DAPT, 85 of those GPU-hours) → 270k synthetic contrastive pairs → multi-task InfoNCE (8 h) → WiSE-FT weight interpolation (0 training).

Now on the official FinMTEB leaderboard (18 EN models): STS #2 (the top embedding model — only a lexical BOW baseline scores higher) and Summarization #2 (behind only voyage-3-large), with the single best FINDsum score on the board. See §2 for the full placement.

This card documents the entire journey: the goal, every training stage, every measured number, the failure that almost sank the release, and the trick that fixed it.

1. The goal

Build the strongest defensible finance embedder on a single encoder of ModernBERT-large size, evaluated on the full English FinMTEB benchmark (35 tasks, 7 task types), against the published state of the art — Fin-E5 (FinMTEB paper), a finance-adapted e5-mistral-7B. Target: win STS and Summarization outright, make Retrieval competitive.

"Defensible" means: every training stage runs behind a hard decontamination gate against all FinMTEB English eval sets (word-shingle overlap index; eval-source datasets and contaminated lineages like FiQA/ConvFinQA excluded outright), and the eval harness is the official FinMTEB fork end to end.

2. Headline results — full FinMTEB-EN, 35 tasks

Task type n This model (395M) Fin-E5 (7B, published) e5-mistral base (7B) Δ vs Fin-E5
Summarization 3 0.588 0.480 0.528 +0.109 👑
STS 2 0.444 0.434 0.380 +0.010 👑
Reranking 3 0.961 0.990 0.988 −0.028
Clustering 6 0.513 0.565 0.578 −0.052
Classification 8 0.623 0.757 0.645 −0.133
PairClassification 3 0.617 0.801 0.739 −0.184
Retrieval 10 0.502 0.711 0.675 −0.208
Overall (7-type mean) 35 0.607 0.677 0.648 −0.070

Fin-E5 / e5-mistral rows are the published Table-1 numbers (weights are not public); this model was scored locally on the same FinMTEB fork with identical task mains (nDCG@10 retrieval, Spearman STS/Summ, MAP reranking, accuracy/v-measure/AP elsewhere).

Official FinMTEB leaderboard placement

The model is now listed on the official FinMTEB leaderboard (English board: 18 models, including voyage-3-large, OpenAI text-embedding-3-large/small, NV-Embed v2, e5-mistral-7B, gte-Qwen1.5-7B, bge-en-icl and Fin-E5). Position as of 2026-07-19, ranks computed from the leaderboard's own per-task scores by averaging tasks within each type:

Task type (EN) Score Rank Note
STS 0.444 #2 / 18 #1 among all embedding models — only the lexical bag-of-words baseline (0.485) sits higher; ahead of Fin-E5 (0.434), voyage-3-large (0.415), NV-Embed v2 (0.374)
Summarization 0.588 #2 / 18 behind only voyage-3-large (0.648); ahead of text-embedding-3-large (0.567) and Fin-E5 (0.480). FINDsum 0.745 is the single best score on the whole board (next: voyage 0.700)
Classification 0.623 #11 / 18
Reranking 0.961 #13 / 18
Clustering 0.513 #13 / 18
Retrieval 0.502 #14 / 18
PairClassification 0.617 #15 / 18
Overall (7-type mean) 0.607 #11 / 18 −0.070 behind the #1 (Fin-E5 0.677) at 1/18th its size

Only four models on the board achieve two or more top-2 task-type finishes: voyage-3-large and text-embedding-3-large (closed-source APIs), Fin-E5 (7B, weights not public) — and this 395M open-weights model. The leaderboard's numbers for this model match the locally-measured scores reported below exactly.

3. The journey, stage by stage

Stage 0 — Measure the real gap (2026-06-29)

Everything starts from a measured baseline of the untrained base model against the finance SOTA, on a 15-task FinMTEB subset (2 STS + 3 Summarization + 10 Retrieval):

task type ModernBERT-large (untrained) Fin-Retriever-base (110M, trained) Fin-E5 (published)
STS 0.450 0.294 0.434
Summarization 0.118 0.281 0.480
Retrieval 0.051 0.402 0.711

Two findings shaped the whole project: (a) untrained ModernBERT-large already matches Fin-E5 on STS — that lead must survive training; (b) Retrieval (0.051) and Summarization (0.118) are where the crown is won or lost.

Stage 1 — Domain-adaptive MLM pretraining, 5.65B tokens

Released separately as FinModernBERT-large-DAPT.

  • Corpus: 185 GB raw finance pool (SEC filings, financial news, finance web) → cleaning + decontamination gate → 5,520,194 packed windows × 1,024 tokens ≈ 5.65B tokens.
  • Run: 1 epoch, 43,127 steps (≈131k tokens/step), mask rate 0.30, peak LR 5e-5 cosine, bf16, full finetune, 2× RTX 3090 DDP. MLM loss 1.134 → 0.888.
  • Measured honestly: DAPT alone barely moves zero-shot embedding scores (STS 0.446 / Summ 0.120 / Retr 0.056) — an MLM objective doesn't reshape sentence geometry. Its value appears downstream, and later it becomes the WiSE-FT anchor.

Interlude — the long-document discovery (+0.185 Summarization for free)

FinMTEB Summarization documents are huge (FNS averages ~290k chars ≈ 70k+ tokens). A zero-training experiment showed head-truncation was destroying the signal: chunk the document into ≤16 windows of ~506 body tokens, embed each, L2-normalize, mean-pool, re-normalize → Summarization 0.120 → 0.305 with no training at all (FNS alone: 0.223 → 0.707). This became the official doc-side embedding strategy.

Stage 2 — 270k synthetic contrastive pairs

  • Teacher: an open-weights 35B-A3B MoE instruction model (AWQ, vLLM, tensor-parallel over both GPUs) generating from SEC full-text chunks.
  • Generated: 100k+ query→passage and graded-STS pairs; 16,696 document↔summary positives + 16,700 mismatched (score-0) summary negatives — built because FinMTEB Summarization requires exactly this separation, and Fin-E5 never trained doc↔summary pairs at all (its own finance adaptation regressed Summarization 0.528 → 0.480); 1,452 term→definition pairs (early-stopped when the distinct-term pool saturated).
  • Filtering: faithfulness checks, dedup, length gates, contamination gate (dropped: 2,628 dup / 308 length / 81 contaminated / 9 unfaithful) → 115,238 finance pairs.
  • Blend with open general data (Investopedia 45k, HotpotQA 30k, NQ 30k, GooAQ 20k, AllNLI 30k) → 270,238 pairs, 59% finance: query_passage 198,870 · sts 36,520 · doc_summary 16,696 · doc_summary_neg 16,700 · term_definition 1,452.

Stage 3 — Contrastive training

  • Hard-negative mining (positive-aware): top-4 negatives per finance anchor mined with intfloat/e5-large-v2, exact top-50 then filtered by sim(cand) ≤ 0.95 · sim(positive) to avoid false negatives; ≥95% of anchors got 4 negatives.
  • Trainer: SentenceTransformer (DAPT encoder + mean pooling, max_seq 512). Multi-dataset training: CachedMultipleNegativesRankingLoss (InfoNCE, scale 20, cross-device in-batch negatives — 255 negatives/device) for retrieval-type data, CoSENT for graded STS pairs. Datasets: qp_finance 108,870 (4 mined negs each) · qp_general 90,000 · doc_summary 16,696 triplets · sts 36,520 · term_definition 1,446.
  • Run: 2 epochs, 1,976 steps, effective 256 anchors/step, lr 2e-5 cosine, bf16, gradient checkpointing, 2× RTX 3090 DDP, ~8h wall. Epoch-2 mean loss 2.18 < epoch-1 2.77. Fully checkpointed/resumable (the run survived two deliberate pauses).

The crisis — and the WiSE-FT fix

Pure contrastive training delivered Retrieval 0.569 and Summarization 0.625 — but STS collapsed 0.450 → 0.371, below the pre-registered release gate (≥ 0.434 = Fin-E5 parity). Both epoch checkpoints failed it.

Instead of retraining, the zero-cost fix: interpolate encoder weights back toward the DAPT initialization (WiSE-FT): w = α·contrastive + (1−α)·DAPT, swept on the eval:

α (contrastive share) STS Summarization Retrieval gate ≥0.434
1.00 (pure) 0.371 0.625 0.569 ✗
0.80 0.402 — — ✗
0.65 → released 0.444 0.588 0.502 ✓
0.50 0.465 0.457 0.327 ✓ (retrieval collapses)
0.30 0.446 — — ✓
0.00 (DAPT) 0.446 0.305 0.056 ✓

Two lessons: α=0.5 scores above both endpoints on STS (the classic WiSE-FT bump), and retrieval skill decays steeply toward the DAPT end — α=0.65 is the balance point that passes the gate while keeping the Summarization crown and most of the retrieval gain.

4. All 35 task scores

Task Type Score
FINAL STS 0.5885
FinSTS STS 0.3003
FNS2022sum Summarization 0.8528
FINDsum Summarization 0.7453
Ectsum Summarization 0.1667
Apple10KRetrieval Retrieval 0.8808
TradeTheEventEncyclopediaRetrieval Retrieval 0.8129
TradeTheEventNewsRetrieval Retrieval 0.7540
FinanceBenchRetrieval Retrieval 0.5853
USNewsRetrieval Retrieval 0.5369
HC3Retrieval Retrieval 0.4216
TheGoldmanEnRetrieval Retrieval 0.3891
FiQA2018Retrieval Retrieval 0.2886
TATQARetrieval Retrieval 0.1885
FinQARetrieval Retrieval 0.1633
FinFactReranking Reranking 0.9752
HC3Reranking Reranking 0.9672
FiQA2018Reranking Reranking 0.9412
ESGClassification Classification 0.8144
FinancialPhraseBankClassification Classification 0.7815
FinancialFraudClassification Classification 0.6392
FiQAClassification Classification 0.6074
FinSentClassification Classification 0.5854
FLSClassification Classification 0.5648
SemEva2017Classification Classification 0.5618
FOMCClassification Classification 0.4319
PiiClustering Clustering 0.8619
MInDS14EnClustering Clustering 0.8268
WikiCompany2IndustryClustering Clustering 0.6831
ComplaintsClustering Clustering 0.2791
FinanceArxivS2SClustering Clustering 0.2146
FinanceArxivP2PClustering Clustering 0.2134
HeadlinePDDPairClassification PairClassification 0.6371
HeadlineACPairClassification PairClassification 0.6073
HeadlinePDUPairClassification PairClassification 0.6073

5. The journey in one table

Milestone STS Summarization Retrieval
Stage 0: untrained ModernBERT-large 0.450 0.118 0.051
Stage 1: + 5.65B-token DAPT 0.446 0.120 0.056
+ chunk+mean-pool eval strategy 0.446 0.305 0.056
Stage 3: + contrastive (pure) 0.371 0.625 0.569
+ WiSE-FT α=0.65 (this model) 0.444 0.588 0.502
Fin-E5 7B (the bar) 0.434 0.480 0.711

6. Usage

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("BatuhanECB/FinModernBERT-embed-large-v1")

# Asymmetric retrieval: prefix queries and passages
queries = ["query: What drove the increase in operating expenses?"]
passages = [
    "passage: Operating expenses rose 12% year-over-year, driven primarily by "
    "increased headcount in R&D and higher cloud infrastructure costs.",
]
q = model.encode(queries, normalize_embeddings=True)
p = model.encode(passages, normalize_embeddings=True)
print(q @ p.T)

# Symmetric use (STS / clustering): prefix both sides with "passage: "
sents = ["passage: Net revenue increased 8%.", "passage: Sales grew by eight percent."]
emb = model.encode(sents, normalize_embeddings=True)

Prefixes matter — training baked literal query: / passage: prefixes in. Long documents (>512 tokens): reproduce the eval numbers by chunking into ≤16 windows of ~506 body tokens, embedding each, L2-normalizing, mean-pooling, re-normalizing.

Embedding dim 1024 · mean pooling · cosine similarity · max_seq 512 (native ModernBERT window is 8,192; 512 is the training configuration).

7. Limitations (measured, not hypothetical)

  • Table/numerical-QA retrieval is the weakest axis (FinQA 0.163, TATQA 0.189): no table-derived pairs in this release's training mix.
  • Scientific-topic clustering (FinanceArxiv S2S/P2P ≈ 0.21): scientific abstracts absent from the mix. General clustering is fine (Pii 0.86, MInDS 0.83).
  • Headline pair-classification (≈0.62): finance-NLI-style pairs absent from the mix.
  • Ectsum (0.167) stays hard for all dense models in this size class.
  • English-only, finance-tuned; general-domain quality not guaranteed.

Planned next iteration: round-2 self-mining with this model, STS-preserving retrain (dropping the need for interpolation), synthetic table-QA pairs, finance-NLI headline pairs, and label-based contrastive for classification.

8. Licensing & data

Model weights: Apache-2.0. The DAPT corpus and synthetic pair set are not released (source licensing does not permit redistribution); the encoder does not memorize/regurgitate corpus text in embedding use. Every training stage was decontaminated against all FinMTEB English eval sets before training.

Related

Downloads last month
101
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BatuhanECB/FinModernBERT-embed-large-v1

Finetuned
(330)
this model

Paper for BatuhanECB/FinModernBERT-embed-large-v1