Instructions to use ysmeta/EVE-Embed-1.0-Legal with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use ysmeta/EVE-Embed-1.0-Legal with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("ysmeta/EVE-Embed-1.0-Legal") sentences = [ "The weather is lovely today.", "It's so sunny outside!", "He drove to the stadium." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [3, 3] - Notebooks
- Google Colab
- Kaggle
EVE-Embed-1.0-Legal
Superseded by EVE-Embed-1.1-Legal, which adds LLM-verified hard negatives and scores 0.7872 nDCG@10 / 0.6700 recall@1 against this model's 0.7800 / 0.6617 on the same benchmark. Use 1.1 unless you specifically need to reproduce these numbers.
Korean embedding model for legal precedent retrieval, fine-tuned from dragonkue/BGE-m3-ko.
It is built for the query style practitioners actually use β plain questions like "λ νΈμΉ΄ νμ¬ μ§μ μ°¨μ£Όκ° μκΈ° μ°¨λ‘ λ λ°κ³ μμ νλ©΄ μ΄μμ¬μ λ² μλ°μΈκ°μ?" β rather than the formal νμμ¬ν phrasing courts write.
| baseline (BGE-m3-ko) | EVE-Embed-1.0-Legal | change | |
|---|---|---|---|
| nDCG@10 | 0.6982 | 0.7800 | +11.7% |
| recall@1 | 0.5647 | 0.6617 | +17.2% |
| general Korean retrieval | 0.8698 | 0.8695 | β0.03% |
Retrieval is over the full 59,786-document corpus β no candidate pre-filtering, no reduced pool. The last row is the check that matters as much as the first: the model gained in its domain without losing general Korean retrieval ability.
Evaluation
Scored on EVE-Bench-Legal-1.0: 6,000 practitioner-style queries against 59,786 Korean court precedents.
| model | nDCG@10 | recall@1 | recall@10 |
|---|---|---|---|
| EVE-Embed-1.0-Legal | 0.7800 | 0.6617 | 0.8978 |
| dragonkue/BGE-m3-ko | 0.6982 | 0.5647 | 0.8342 |
| nlpai-lab/KURE-v1 | 0.6916 | 0.5562 | 0.8308 |
| BAAI/bge-m3 | 0.6472 | 0.5083 | 0.7932 |
| intfloat/multilingual-e5-large | 0.6288 | 0.4897 | 0.7788 |
Why the benchmark uses rewritten queries
The corpus ships a natural query-document pair per case: νμμ¬ν (the legal question) and νκ²°μμ§ (the holding). Public models score ~0.87 nDCG@10 on that pairing β but both fields are written by the same court about the same issue and share most of their wording, so the score largely measures lexical overlap. Rewriting the query into ordinary practitioner language drops every public model by 0.17β0.20, which is the part of the original score that was not comprehension. All numbers above are on the rewritten (hard) set.
Usage
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("ysmeta/EVE-Embed-1.0-Legal")
query = "κ³μ½κΈμ μ΄λ―Έ λ°μλλ°λ μλλ°©μ΄ κ³μ½μ μ·¨μν μ μλμ?"
docs = [
"μμ κΆμ΄μ λ±κΈ°μ²κ΅¬μ¬κ±΄ λΉμ¬μμ μΌλ°©μ΄ κ³μ½μ΄νμ μ°©μν νμλ ν΄μ κΆ νμ¬λ₯Ό ν μ μμΌλ―λ‘ β¦",
"μλμ°¨κ³μ½μ 묡μμ κ°±μ μ΄ μΈμ λλ κ²½μ° β¦",
]
q = model.encode(query, normalize_embeddings=True)
d = model.encode(docs, normalize_embeddings=True)
print(q @ d.T)
No instruction prefix is required.
Training
- Base: dragonkue/BGE-m3-ko (Apache-2.0)
- Data: 134,996 (query, precedent) pairs β practitioner-style queries generated from 45,000 precedents, see EVE-Train-Legal-1.0
- Loss: CachedMultipleNegativesRankingLoss, in-batch negatives only
- Effective batch: 512 (256 Γ 2 GPUs, DDP) β batch size is the negative count here
- LR: 1e-5, 1 epoch, max_seq_len 384, bf16
- Hardware: 2 Γ NVIDIA H200
What did not work
Published for reuse, since two of the three levers moved the score the wrong way:
| configuration | nDCG@10 | vs baseline |
|---|---|---|
| mined hard negatives (skip top-3), lr 2e-5, 65% formal queries | 0.4434 | β36.5% |
| mined hard negatives (skip top-10), lr 5e-6, synthetic only | 0.6457 | β7.5% |
| no mined hard negatives, batch 64 | 0.7535 | +7.9% |
| no mined hard negatives, batch 512 | 0.7643 | +9.5% |
| above + 4.3Γ more data | 0.7800 | +11.7% |
Mined hard negatives are counter-productive in this domain. Legal precedents cluster tightly by issue, so the top-ranked non-gold documents for a legal question are usually genuinely relevant cases. Training the model to push them away damages its legal semantics. Skipping the top 10 hits did not help β only removing mined negatives entirely did.
A related trap: the first run gained 5.7% on the formal-query set while losing 36.5% on the practitioner set. Evaluating only on the easy set would have reported that run as a success.
Limitations
- νλ‘ only. The corpus contains court precedents, not statute text (λ²λ Ή μ‘°λ¬Έ), so queries that should resolve to a specific article are not covered.
- Case distribution of the 59,786-document corpus: λ―Όμ¬ 46.3%, νμ¬ 21.9%, μΌλ°νμ 14.1%, μΈλ¬΄ 11.9%, νΉν 4.3%, κ°μ¬ 1.6%. νΉν/κ°μ¬ coverage is thin.
- Training queries are LLM-generated, not collected from real users. They were written to imitate practitioner phrasing, but real query logs would differ.
- Not legal advice. Retrieval surfaces precedents; it does not interpret them.
License
Apache-2.0, inherited from the base model. The training corpus derives from
joonhok-exo-ai/korean_law_open_data_precedents (OpenRAIL); the dataset repos carry
that licence.
- Downloads last month
- -