YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Cross-Scriptural Verse Matcher

A unified framework for finding theologically relevant verses across the Old Testament (OT), New Testament (NT), and the Quran based on semantic meaning and latent thematic guidance.

Architecture

  • Base Encoder: intfloat/multilingual-e5-base โ€” multilingual sentence embeddings, strong zero-shot retrieval
  • Latent Guidance Head: Multi-label thematic classifier (45 theological themes) on top of embeddings
  • Scripture-Type Embedding: Learned embedding for OT/NT/Quran to condition the shared space
  • Loss: Combined contrastive (InfoNCE with hard negatives) + thematic BCE classification

Training Recipe

Component Value
Base model intfloat/multilingual-e5-base
Fine-tuning LoRA (r=16, alpha=32) on Q,K,V,Dense
Max length 256
Batch size 32
Learning rate 3e-4
Temperature 0.05
Theme loss weight (ฮป) 0.3
Epochs 5
Optimizer AdamW with cosine schedule
Precision bf16

Dataset

Cross-scriptural verse pairs generated via LLM annotation with:

  • Similarity scores (0.0โ€“1.0)
  • Relationship types: thematic, narrative, prophetic, lexical, ethical, cosmological
  • Hard negatives: same-theme verses with different meaning
  • 45 theological themes for latent guidance

Source datasets:

  • Quran: freococo/quran_multilingual_parallel (English)
  • Bible: davidguzmanr/open-bible-resources (English Standard, verse-level)

Usage

Training

python train.py

Inference

python inference.py --query "For God so loved the world" --top_k 5

Repositories

Citation

Built on:

  • SimCSE (Gao et al., 2021) โ€” contrastive sentence embeddings
  • E5 (Wang et al., 2022) โ€” weakly-supervised text embedding
  • multilingual-e5 (XLM-R backbone) โ€” cross-lingual alignment
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support