JevCodeLocator-0.6B
A 0.6B code localizer. Given a repository and a natural-language query (an issue, a symptom, a
symbol, an error message, or a described behavior), it reranks candidate code locations and returns a
ranked list of path:start-end chunks together with a file-level ranking.
The model is a reranker, not a retriever: a cheap sparse retriever (BM25 plus symbol/path channels) produces a shortlist of K candidates and the model scores all of them in a single forward pass. There are no embeddings and no vector index. On an entry-level 6 GB GPU it runs in roughly 0.4 s at K=8 and 0.9 s at K=16, which makes it usable as a local first-stage localizer for coding agents.
Architecture
state (task / repo / candidate_count / query) + candidate texts ("path:start-end: summary")
β
ββ one leaf-token path per candidate βββΊ Qwen3-0.6B backbone
β (last hidden state at the EOS of each path)
ββ LayerNorm β Linear(d, 1) per-candidate logit z_i
ββ set-attention over the K candidates:
u = Linear(d + 1, 128)([h_i ; log K])
mixed = MultiheadAttention(u, u, u) (key-padding mask for empty slots)
z_i += Linear(128, 1)(tanh(u + mixed))
β
ββ softmax over the K candidates
The scalar head alone can rank candidates independently; the zero-initialized set-attention correction is what makes a candidate's score depend on the whole shortlist (for example, suppressing a cluster of look-alike test files in favour of the implementation file).
Training
Three consecutive fine-tuning stages on top of Qwen/Qwen3-0.6B:
| stage | training data | steps | notes |
|---|---|---|---|
| 1 | Python programmatic queries, Kβ€16, file-level auxiliary loss Ξ»=0.30 | 600 | base localizer |
| 2 | + Go / TypeScript-JavaScript / C++ / Rust / C# / PHP programmatic queries; 12,274 rows, K=16 BM25 pools with the gold candidate forced in | 600 | doc-comment queries oversampled 2Γ, backbone LR 1e-5 |
| 3 | continuation of stage 2 on 18,688 rows | 600 | final checkpoint |
- Objective: cross-entropy against the gold candidate (
gold_distribution) plus the file-level auxiliary loss (weight 0.30). - Candidate text is identical at training and serving time: a compact summary consisting of the path, the signature and the first lines of the chunk.
- A Python replay set (2,286 rows) is kept in every round so that Python behaviour does not regress.
- Training queries are generated programmatically from doc comments, symbol names and constants. No human-written query set was used for training.
Evaluation
Frozen held-out pools: 200 queries per repository, K=8, gold candidate forced into the pool.
file@1 is the share of queries whose gold file is ranked first. The baseline is the same model
before the multilingual stages (stage 1 only).
| evaluation set | language | baseline | this model |
|---|---|---|---|
| service repository, ~1.3k files | Go | 0.700 | 0.945 |
| ML framework core, ~2k chunks | C++ | 0.730 | 0.870 |
| agent CLI, ~20k chunks | Rust | 0.900 | 0.925 |
| API service, ~5k chunks | TypeScript | 0.795 | 0.855 |
| language server, ~600 chunks | C++ | 0.974 | 0.989 |
| held-out dev set, 460 queries | Python | 0.839 | 0.843 |
Restricted to semantic (doc-comment) queries, file@1: Go 0.524 β 0.913, C++ 0.629 β 0.814,
TypeScript 0.667 β 0.758, Rust 0.833 β 0.875.
On a set of five hand-checked, real reverse-proxy questions about a large Go service (K=16, FP32): gold-file recall@1 = 1.000, recall@3 = 1.000, recall@5 = 1.000. Latency on an entry-level 6 GB GPU: 0.40β0.44 s (K=8), ~0.83 s (K=16).
Usage
pip install torch transformers safetensors
python inference_example.py # runs on this directory
import torch
from modeling_jev import load_jev_model, score_candidates
model, tok = load_jev_model(".", device="cuda" if torch.cuda.is_available() else "cpu",
dtype=torch.float32)
candidates = {
"src/auth/session.py:41-88": "def create_session(user, password) | validates credentials ...",
"src/db/pool.py:12-60": "def connect(dsn, max_size) | opens the database pool ...",
}
for key, prob in score_candidates(model, tok, "where are credentials validated?", candidates):
print(f"{prob:.4f} {key}")
score_candidates returns the candidates ranked by probability. Summing the probabilities of all
candidates that belong to the same file gives the file-level ranking.
Files
| file | contents |
|---|---|
model.safetensors |
full model: 322 tensors (backbone.* plus the decision head); sha256 fb06ba6ac6d30c395912ebd156723a0739afdd260f9b44f85c56b074c68b8a98 |
config.json |
model configuration (architectures: JevDecisionModel) and training hyperparameters |
backbone_config/config.json |
Qwen3-0.6B backbone configuration |
tokenizer/ |
tokenizer files (tokenizer_config.json is in the transformers>=5 format) |
head.safetensors |
the 12 decision-head tensors only (803 KB), for custom runtimes |
modeling_jev.py |
self-contained model class, loader and scoring helper |
inference_example.py |
runnable reranking example |
patch_tokenizer_for_transformers4.py |
converts tokenizer_config.json to the 4.x format |
LICENSE, NOTICE |
Apache-2.0 and attribution |
Intended use and limitations
- Use it as a reranker. It selects among the candidates it is given; if the retriever's shortlist does not contain the gold file, the model cannot recover it. Pool recall is the ceiling.
- Trained on programmatically generated queries (doc comments, symbols, constants) with a small amount of issue-style data mixed in. Free-form questions work but are the hardest distribution.
- Language coverage: Python, Go, TypeScript/JavaScript, C, C++, Rust, C#, PHP, Java, Ruby, Bash. The chunker used to build candidates is tree-sitter based; other languages are untested.
- Candidate text is a compact summary, so very long functions are effectively truncated.
- Not instruction-tuned, not a chat model, no safety tuning. Do not use it to make decisions about code it cannot see.
Training data and licensing
The weights in this repository are released under Apache-2.0, the same license as the base model
Qwen/Qwen3-0.6B (Apache-2.0, Copyright 2025 Alibaba Cloud).
The inference code here (modeling_jev.py, inference_example.py,
patch_tokenizer_for_transformers4.py) is original work under the same license; see LICENSE and
NOTICE.
Training rows are generated programmatically: a tree-sitter chunker extracts code locations, queries are derived from doc comments, symbol names and constants, and the candidate text is a compact summary of the chunk. No source file from any repository is redistributed in this repository.
Licenses of the source repositories used for training and evaluation data:
| license | sources |
|---|---|
| MIT | 6 repositories (Go / TypeScript / Python tooling) |
| Apache-2.0 | 3 repositories (C++ / Go / Rust) |
| GPL-3.0 | 2 repositories |
| LGPL-3.0 | 2 repositories |
| AGPL-3.0 | 1 repository |
| non-commercial custom license | 1 repository |
| license not declared in the checkout used | 6 repositories |
A detailed list of the source repositories is available on request.
Notes:
- The weights are not a redistribution of those repositories, and the model does not reproduce training code verbatim: it only scores candidate locations supplied by the caller. Whether copyleft-licensed training data extends to model weights is legally unsettled in most jurisdictions.
- Commercial use: the training mix includes one AGPL-3.0 project and one non-commercial project. For a commercially clean release, retrain without those two (and preferably without the GPL/LGPL projects); the data pipeline is deterministic and reproducible.
- You remain responsible for complying with the licenses of the code you run the model on.
- This section is provenance information, not legal advice.
- Downloads last month
- 14