err2fix-py-v1 Model Card
Model Details
- Model name:
err2fix-py-v1 - Base model:
thenlper/gte-small - Task: English retrieval embeddings for error-to-fix matching
- Domain: Python troubleshooting (imports, packaging, runtime)
- Architecture family: Sentence Transformers-compatible encoder
Intended Use
err2fix-py-v1 is intended for retrieval workflows that map Python error strings or troubleshooting queries to the most relevant fix snippets, docs excerpts, or known solutions.
This model is retrieval-only. It is not a generator, reranker, multilingual model, or production-serving package.
Training Data
- Local curated dataset:
468manually reviewed examples - Splits:
train=320,val=75,test=73 - Source files:
dataset/train.jsonl,dataset/val.jsonl,dataset/test.jsonl - Data emphasis: hard negatives and same-ecosystem wrong-fix separation
Evaluation Summary (468 Snapshot)
Primary metrics are Recall@1, Recall@5, MRR, and NDCG@10.
| Model | Recall@1 | Recall@5 | MRR | NDCG@10 |
|---|---|---|---|---|
Baseline (thenlper/gte-small) |
0.8312 | 1.0000 | 0.9056 | 0.9298 |
Candidate v1.1 (training/runs/20260420T091003Z/model) |
1.0000 | 1.0000 | 1.0000 | 1.0000 |
Public Evaluation Notes
The model performs strongly on the curated Python troubleshooting evaluation split, but broad-domain retrieval transfer is limited. Published MTEB metadata currently reflects the available official-style retrieval task results in the model card metadata.
Future updates should be selected using both the Python troubleshooting split and a broader English Retrieval validation slice.
Reproducibility
Setup and validation:
uv sync --group dev
uv run --group dev pytest
uv run --group dev ruff check .
uv run --group dev mypy src
uv build
Baseline evaluation:
uv run python scripts/evaluate_baseline.py
Tuned checkpoint evaluation:
uv run python scripts/evaluate_checkpoint.py --model training/runs/20260420T091003Z/model --source submission-candidate-v1.1 --output evaluation/runs/err2fix-py-v1-phase11-sweep-b-margin03-on-468.json
Limitations
- Local Python troubleshooting evaluation currently saturates for tuned checkpoints on the 468-example snapshot.
- Performance claims are scoped to the repository dataset and domain slice, not broad MTEB generalization.
- This model is not intended for JavaScript troubleshooting, multilingual retrieval, reranking, or generation.
Ethical and Safety Notes
- This model assists troubleshooting retrieval and can return incomplete or context-specific fixes.
- Retrieved guidance should be validated against environment specifics (OS, Python version, dependency set) before application.
- Downloads last month
- 2