reign-tiny-l1_gn-bge-large_val-selected

REIGN tiny-l1 cross-chunk encoder trained on GoodWiki-Long-Synthetic over a frozen BGE-large guidance network — a long-document bi-encoder that reads a sequence of cached chunk embeddings instead of tokens.

Configuration

REIGN encoder tiny-l1 — 1 layer, d = 192, 3 heads, FFN 768, 0.56M trainable parameters
Guidance network (frozen) BAAI/bge-large-en-v1.5 — BGE-large, 335M
Chunk size K 512 (the guidance network's context window)
Training stride S 384 — gn_stride recorded in this checkpoint's config.json
Chunk-position signal none — the encoder is a permutation-equivariant set function
Pooling mean over the chunk sequence
Training data devrim/goodwiki_long_synthetic_ir
Training recipe released cosine recipe (below)
Checkpoint selection best validation nDCG@10 on the val qrels split

Reported results

The paper reports no table row for this exact checkpoint — it is released for completeness of the configuration sweep. For the numbers of the checkpoints that are reported, see the model zoo and the paper.

Usage

The checkpoint holds only the REIGN cross-chunk encoder. The guidance network is loaded separately and stays frozen, so both must be named at construction time.

pip install git+https://github.com/devrimcavusoglu/reign.git
import numpy as np
from huggingface_hub import snapshot_download
from reign.encoders.reign import ReignBaselineEncoder

checkpoint_path = snapshot_download("devrim/reign-tiny-l1_gn-bge-large_val-selected")

encoder = ReignBaselineEncoder(
    checkpoint_path=checkpoint_path,
    gn_model="BAAI/bge-large-en-v1.5",
    chunk_size=512,
    stride=384,
)
docs = [open("doc_a.txt").read(), open("doc_b.txt").read()]
emb = encoder.encode(docs, batch_size=8)   # (2, hidden_size), L2-normalised
print(float(np.dot(emb[0], emb[1])))       # cosine similarity

ReignBaselineEncoder returns L2-normalised vectors, so the cosine is a dot product. chunk_size is the guidance network's sliding-window size — 512 for every released checkpoint, matching its context window — and stride controls the overlap, with stride == chunk_size giving non-overlapping chunking. The evaluation-time stride is a runtime argument, and the paper's headline tables report the best-performing stride per guidance network.

For the lower-level surface, ReignModel (a PreTrainedModel consuming inputs_embeds) and ReignFeatureExtractor (the guidance-network wrapper, with the on-disk embedding cache) are importable from reign and reign.feature_extractor.

Operating regime

REIGN targets multi-chunk inputs and primarily document-to-document retrieval. Inputs shorter than the chunk size collapse to a single chunk embedding, leaving the cross-chunk encoder nothing to aggregate — that regime is served by the guidance network alone, and this checkpoint should not be used for it.

Training recipe

The released checkpoints use a three-way cosine embedding loss over document pairs carrying graded targets s ∈ {1, 0, −1} — positive, partial, negative — with partial weight λ = 0.5.

Setting Value
Objective three-way cosine embedding loss, λ = 0.5
Batch construction 18 anchors × (1 positive + 2 partials + 17 in-batch negatives) = 360 pairs / step
Optimiser AdamW, lr 1e-5, weight decay 1e-4, cosine annealing
Epochs 50, validating every 4
Selection best validation nDCG@10
Precision 16-mixed
Seed 42
Guidance-network embeddings precomputed and cached
Hardware single 24 GB consumer GPU

A second, separate protocol (InfoNCE at τ = 0.07, batch 48, 20 epochs) is used only for the paper's positional-encoding and training-objective ablations. Those arms sit below this released operating point by construction and are not comparable to it. See docs/TRAINING.md in the code repository.

Because 16-mixed training is not bit-reproducible even at a fixed seed, a retrained checkpoint will not match these weights bit-for-bit; compare metrics, not weights.

Files

  • config.jsonReignModel configuration
  • model.safetensors — encoder weights (float32)

Links

Citation

@inproceedings{cavusoglu2026reign,
  title     = {{REIGN}: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling},
  author    = {{\c{C}}avu{\c{s}}o{\u{g}}lu, Devrim and Akba{\c{s}}, Emre},
  booktitle = {Findings of the Association for Computational Linguistics: {EMNLP} 2026},
  year      = {2026},
  publisher = {Association for Computational Linguistics},
  note      = {To appear}
}

License

Apache License 2.0. The devrim/goodwiki_long_synthetic_ir dataset is released under CC BY-SA 4.0, preserving the share-alike licensing and attribution of GoodWiki and the underlying Wikipedia text.

Downloads last month
-
Safetensors
Model size
681k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for devrim/reign-tiny-l1_gn-bge-large_val-selected

Finetuned
(96)
this model

Dataset used to train devrim/reign-tiny-l1_gn-bge-large_val-selected

Collection including devrim/reign-tiny-l1_gn-bge-large_val-selected