FinLongformer

FinLongformer is an English finance/quant encoder obtained by continuing masked-language pretraining of allenai/longformer-base-4096. It retains the masked-language-modeling head and supports up to 4,096 input tokens, including special tokens. The encoder has 12 layers and 768-dimensional token representations.

The model provides a reusable starting point for financial text tasks. It has no downstream task-specific head, and it was not trained with trading or return-prediction labels.

Training

The source corpus contains English SEC 10-K narrative sections, Federal Reserve speeches/policy statements/minutes/Beige Books, and arXiv quant-finance metadata and selected openly licensed full texts. Original document and company/paper-group assignments were preserved across training and validation partitions. Source sampling targeted SEC/Fed/quant at 80%/10%/10%. Source attribution and reuse terms are recorded in source_notices.json and training_source_attributions.json; corpus text is not included in this repository.

The full run processed 300,085,534 content-token exposures over 5,123 optimizer updates. These are token exposures, not a count of unique corpus tokens. The released weights are the checkpoint selected at update 5,000, after 292,995,933 content-token exposures, using the lowest fixed-mask financial validation loss.

Training used dynamic 15% masking, a maximum sequence length of 4,096, global attention on the first CLS token, an effective batch of 32 sequences, AdamW with an initial learning rate of 2e-5, 5% warmup and linear decay, and BF16 autocast with FP32 model weights. The original tokenizer and vocabulary were retained. Detailed settings and the distinction between the selected checkpoint and the final training step are in training_summary.json.

Available evaluation

The released checkpoint's fixed-mask validation loss is 0.898456, calculated as an 80%/10%/10% weighted mean of SEC/Fed/quant group losses. The group losses are:

Group MLM loss
SEC 0.854611
Federal Reserve 0.923689
Quant research 1.223984

These are development diagnostics from sampled financial validation windows. Both the company/paper-group and later-period validation partitions were used for checkpoint selection. They are not an independent locked test, and this release does not establish improvement over the unmodified base on downstream tasks. Full metrics are in evaluation.json.

Usage

import torch
from transformers import AutoTokenizer, AutoModelForMaskedLM

repo = "LeoDingggg/FinLongformer"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForMaskedLM.from_pretrained(repo).eval()

inputs = tokenizer(
    "The company's operating margin increased during the year.",
    return_tensors="pt", truncation=True, max_length=4096,
)
global_attention_mask = torch.zeros_like(inputs["input_ids"])
global_attention_mask[:, 0] = 1

with torch.no_grad():
    outputs = model.longformer(
        **inputs, global_attention_mask=global_attention_mask,
    )
    token_representations = outputs.last_hidden_state

The saved model can also be used for masked-token prediction with AutoModelForMaskedLM. Pooling and any downstream head must be chosen and validated for the intended application; this checkpoint was not contrastively trained as a sentence-embedding model.

Limitations

General-language retention, downstream financial benchmarks, and downstream benchmark contamination have not been independently evaluated. SEC dates in the source corpus are report-year buckets rather than exact acceptance timestamps, source cutoffs differ, and original table layout is not reliably reconstructed. The checkpoint is intended for reusable text representations; its MLM scores do not demonstrate numerical reasoning or trading performance.

License and attribution

The model is released under Apache-2.0, following the original Longformer model. Training-source licenses remain source-specific and are documented separately. See LICENSE and NOTICE.

Original Longformer: Iz Beltagy, Matthew E. Peters and Arman Cohan, Longformer: The Long-Document Transformer, 2020.

Downloads last month
25
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LeoDingggg/FinLongformer

Finetuned
(145)
this model

Paper for LeoDingggg/FinLongformer