err2fix-py-v1 Model Card

Model Details

  • Model name: err2fix-py-v1
  • Base model: thenlper/gte-small
  • Task: English retrieval embeddings for error-to-fix matching
  • Domain: Python troubleshooting (imports, packaging, runtime)
  • Architecture family: Sentence Transformers-compatible encoder

Intended Use

err2fix-py-v1 is intended for retrieval workflows that map Python error strings or troubleshooting queries to the most relevant fix snippets, docs excerpts, or known solutions.

This model is retrieval-only. It is not a generator, reranker, multilingual model, or production-serving package.

Training Data

  • Local curated dataset: 468 manually reviewed examples
  • Splits: train=320, val=75, test=73
  • Source files: dataset/train.jsonl, dataset/val.jsonl, dataset/test.jsonl
  • Data emphasis: hard negatives and same-ecosystem wrong-fix separation

Evaluation Summary (468 Snapshot)

Primary metrics are Recall@1, Recall@5, MRR, and NDCG@10.

Model Recall@1 Recall@5 MRR NDCG@10
Baseline (thenlper/gte-small) 0.8312 1.0000 0.9056 0.9298
Candidate v1.1 (training/runs/20260420T091003Z/model) 1.0000 1.0000 1.0000 1.0000

Public Evaluation Notes

The model performs strongly on the curated Python troubleshooting evaluation split, but broad-domain retrieval transfer is limited. Published MTEB metadata currently reflects the available official-style retrieval task results in the model card metadata.

Future updates should be selected using both the Python troubleshooting split and a broader English Retrieval validation slice.

Reproducibility

Setup and validation:

uv sync --group dev
uv run --group dev pytest
uv run --group dev ruff check .
uv run --group dev mypy src
uv build

Baseline evaluation:

uv run python scripts/evaluate_baseline.py

Tuned checkpoint evaluation:

uv run python scripts/evaluate_checkpoint.py --model training/runs/20260420T091003Z/model --source submission-candidate-v1.1 --output evaluation/runs/err2fix-py-v1-phase11-sweep-b-margin03-on-468.json

Limitations

  • Local Python troubleshooting evaluation currently saturates for tuned checkpoints on the 468-example snapshot.
  • Performance claims are scoped to the repository dataset and domain slice, not broad MTEB generalization.
  • This model is not intended for JavaScript troubleshooting, multilingual retrieval, reranking, or generation.

Ethical and Safety Notes

  • This model assists troubleshooting retrieval and can return incomplete or context-specific fixes.
  • Retrieved guidance should be validated against environment specifics (OS, Python version, dependency set) before application.
Downloads last month
2
Safetensors
Model size
33.4M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results