YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
PR quality/multilingual/tags fastText classifiers
Four fastText classifiers that score or tag a GitHub PR from its diff + touched-file
content, distilled from Jev (typesafe/jev-1.13) annotations on mined GitHub Archive
PRs. Each subfolder is self-contained: model.bin (weights), train.txt/val.txt
(the exact fastText training data), train.py (the exact training call — rerun in
place to reproduce model.bin from train.txt), and infer.py (standalone scoring
script — fetches a live PR via the GitHub API and classifies it, no dependency on the
original mining pipeline).
Usage
pip install fasttext==0.9.3 "numpy<2" requests # numpy<2 is required -- see note below
export GITHUB_PAT=... # any GitHub token (read-only is fine)
python quality/infer.py https://github.com/owner/repo/pull/123
python multilingual/infer.py https://github.com/owner/repo/pull/123
python research_novelty/infer.py https://github.com/owner/repo/pull/123
python tags/infer.py https://github.com/owner/repo/pull/123 # multi-label
python tags/infer.py --threshold 0.3 https://github.com/owner/repo/pull/123
To retrain a model from its own train.txt in place: cd <subfolder> && python train.py.
quality/multilingual/research_novelty each print a soft score (probability-
weighted expected value, the better-correlating of the two — use this) and a hard
score (argmax level), both on a 0-10 scale. tags instead prints every tag (of 424)
whose predicted probability clears --threshold (default 0.5), sorted by confidence —
see its own section below for why it's one multi-label model, not 424 separate ones.
Note on fastText + NumPy 2.x: the prebuilt fasttext-wheel can silently corrupt
batch-predict() probabilities under NumPy 2.x, and even a from-source build's Python
wrapper crashes outright on np.array(..., copy=False). Pin numpy<2.
quality/ — Jev's strict 0-10 quality rubric
Rates the PR's diff on a strict, sensibility-gated 0-10 scale: 0-5 is gated on whether
the codebase/diff is coherent and credible (a simple-but-clean change tops out at 5); 6+
requires genuine engineering or conceptual complexity on top of that. See
annotate_pr_quality.py's QUALITY_LEVELS for the full rubric text.
- Hyperparameters: dim=64, epoch=10, lr=0.3, wordNgrams=2, minCount=3, bucket=2,000,000
- Train: 254,143 PRs (
train.txt) - Val: 28,238 PRs, 1,766 at level≥6 (
val.txt)
This is the plain dim64/epoch10 baseline — every variant tried (2-4x capacity, class- rebalanced train sets, a ~1.55x-larger combined corpus with a separately-mined higher-credibility PR pool) came out flat or slightly worse on a held-out 10,000-PR tracking set (685 at quality≥6). More data/capacity does not fix this classifier; treat it as a known ceiling, not a bug to chase.
Metrics (10,000-PR tracking holdout, rule = quality >= 6):
| metric | value |
|---|---|
| rule accuracy | 93.73% |
| precision | 74.17% |
| recall | 12.99% |
| exact-level accuracy | 78.01% |
| MAE | 0.531 |
| soft Pearson / Spearman | 0.7252 / 0.6339 |
Read this before using it as a filter: it misses ~87% of genuinely good PRs
(quality≥6) to keep false positives to 0.33% of negatives, and its own predictions never
exceed ~6.8-7 on any PR we've tested — it is a high-precision, very-low-recall filter,
not a general-purpose quality scorer. It also gave near-identical predictions (4.4-5.3)
to real, high-quality ML/research PRs (e.g. lucidrains' x-transformers, imagen-pytorch)
as to trivial personal-repo PRs — it is not sensitive to genuine technical depth. Lowering
the decision threshold to >=5 makes this worse, not better (accuracy drops to 37%,
since ~80% of real PRs are already >=5 and the model's own predicted mass sits below
that line).
multilingual/ — non-English comment/string content
Purely descriptive 0-10 scale: how much non-English natural-language text (comments,
docstrings, string literals) appears in the diff and touched files, independent of the
programming language itself. See annotate_pr_quality.py's MULTILINGUAL_LEVELS.
- Hyperparameters: dim=128, epoch=25, lr=0.3, wordNgrams=2, minCount=3, bucket=2,000,000 (best of a capacity sweep vs dim64/epoch10 baseline and dim256/epoch25 — dim128 won on every axis, dim256 added nothing further for 2x the model size)
- Train: 254,143 PRs (
train.txt) - Val: 28,238 PRs (
val.txt)
Metrics (100-PR live holdout, 79 true positives, rule = multilingual <= 1):
| metric | value |
|---|---|
| rule accuracy | 96.00% |
| precision | 96.30% |
| recall | 98.73% |
| MAE | 0.586 |
| Pearson / Spearman | 0.906 / 0.721 |
This one is genuinely reliable as a pre-filter — high recall and precision. If you only trust one of these two classifiers for a filtering decision, trust this one.
research_novelty/ — research/IP content (⚠️ still recall≈0 at the >=6 rule)
0-10 scale rating the extent to which the PR implements technical algorithms, ideas from
research papers, or original design/IP, as opposed to routine engineering — independent
of how difficult or high-quality the engineering itself is. See annotate_pr_quality.py's
RESEARCH_NOVELTY_LEVELS.
- Hyperparameters: dim=64, epoch=10 (same baseline as
quality/multilingual) - Train: 96,634 PRs, 297 at level≥6 (
train.txt) - Val: 10,772 PRs, 35 at level≥6 (
val.txt) - Same train/val PR ids as
tags/(both come from the same annotation batch; see Provenance) — comparable across the two.
Updated 2026-09-30: retrained on a new 107,406-PR annotation batch (up from the
original 39,081-PR/98-positive pass), joined back to quality/multilingual's existing
text field so the same 254K+-PR pool's diff/file content is reused rather than
re-mined. Labels use the field's own fixed 0-10 scale (int truncation), same as
quality/multilingual — not sample quantiles.
Metrics (own val.txt, 10,772 PRs):
| metric | value |
|---|---|
| exact-level accuracy | 82.45% |
| within-1-level accuracy | 94.06% |
| MAE | 0.2745 |
| hard Pearson / Spearman | 0.5679 / 0.5069 |
| soft Pearson / Spearman | 0.6662 / 0.5470 |
rule (>=6) accuracy |
99.68% |
rule (>=6) precision |
undefined (0 predicted positive) |
rule (>=6) recall |
0.00% (0/35) |
Read this before using it as a filter: correlation improved modestly over the old
diagnostic pass (soft Pearson 0.60→0.67) and the accuracy/MAE numbers look good, but that's
almost entirely the near-100%-easy "correctly predict low novelty" majority class — at the
actual decision boundary that matters (>=6), it now has enough val data to say
confidently (35 true positives, not 2) that it misses every single one, the same
recall-collapse pattern as quality. Its own predictions rarely if ever clear ~6-7 (see
the Megatron-LM spot-check below). Treat it as a continuous novelty signal (the soft
score correlates reasonably, 0.67) rather than a working >=6 binary classifier.
Spot-check: NVIDIA/Megatron-LM#7762 (an Engram/CUDA synchronization change) scores
soft=2.46, hard=1 — plausible as "some genuine ML-systems depth but not paper-derived,"
though still likely underrating it given the classifier's known insensitivity to real
technical depth (see quality's note above; the same failure mode shows up here).
tags/ — multi-label topic tagging (424 tags)
Multi-label classification over the topic vocabulary in /ansh/harbor-tags.txt (424
distinct tags after dedup — e.g. python, machine-learning, web-security,
gpu, formal methods) — any number of tags can apply to one PR, independently.
Each tag is an independent Jev "noul" (boolean-probability) claim ("this PR's diff or
touched files involve code/content related to topic X"), bundled into one Jev call per
PR alongside research_novelty (see annotate_tags_research.py).
One joint multi-label model, not 424 separate binary classifiers. fastText's native
multi-label mode (loss="ova", one-vs-all) shares a single embedding/vocabulary across
all 424 tags and trains in one pass — dramatically cheaper than 424x the training
time/storage for what is fundamentally the same text signal, at the cost of slightly
less per-tag threshold flexibility (mitigated by tuning --threshold in infer.py
per use case; default 0.5).
- Hyperparameters: dim=64, epoch=10, lr=0.3, wordNgrams=2, minCount=3, bucket=2,000,000, loss=ova
- Train: 96,604 PRs (
train.txt) / Val: 10,769 PRs (val.txt) — same PR ids asresearch_novelty/'s updated split (both from the same annotation batch) - 33 PRs with zero tags above threshold were dropped (fastText needs ≥1 label/line)
Metrics (10,769-PR val set, pooled across all 10,769×424 PR-tag pairs, threshold=0.5):
| metric | value |
|---|---|
| precision | 81.95% |
| recall | 57.17% |
| f1 | 67.35% |
pooled Pearson r (vs. Jev's raw tag_probs) |
0.8055 |
| pooled Spearman ρ | 0.7145 |
Precision stays high across common tags (0.79-0.99) with recall degrading gracefully as per-tag support drops — normal multi-label behavior, not a red flag. This is the best-behaved of the four classifiers relative to its task difficulty (a 424-way multi-label problem vs. a single 0-10 scale).
Known data-hygiene issue: harbor-tags.txt has a couple of near-duplicate tag pairs
that survived as distinct labels (data processing/data-processing, data engineering/data_engineering — the latter pair also collided post-slugification and
one was suffixed _dup in tag_vocab.json to disambiguate). Worth deduping the source
file; not fixed here since it would invalidate the current label vocabulary/model.
Spot-check: vercel/next.js#99499 (a tag-revalidation fix) → 30 tags incl.
typescript, react, nextjs, frontend, routing, concurrency, web-server.
NVIDIA/Megatron-LM#7762 → 36 tags incl. machine-learning, pytorch, cuda, gpu,
distributed-systems, performance-optimization, deeply-technical-ip. Both match
what a human skim of the PRs would tag.
Provenance
- Distilled from
typesafe/jev-1.13(via OpenRouter) on GitHub Archive-mined PRs. - Full annotation/training pipeline:
/ansh/prism-datagen/annotate_pr_quality.py(rubric + Jev calls, incl.HARBOR_TAGS/TAG_CLAIMS),/ansh/fasttext-stuff/run.py/train_sweep.py(quality/multilingual training — each subfolder'strain.pyis the relevant one of these, trimmed to a single model),/ansh/fasttext-stuff/evaluate.py/evaluate_sweep.py/compare_pre_post.py(eval). quality's known ceiling was diagnosed, not guessed: a dim/epoch capacity sweep (64→128→256), a class-rebalancing ablation (10%/15% minority oversampling), and a ~1.55x larger combined training corpus (adding a separately GitHub-credibility-filtered PR pool skewed toward higher quality) were all tried and none moved recall past ~14%.research_novelty/tagscome from a separate, lighter-weight annotation pass:/ansh/prism-datagen/annotate_tags_research.pyre-streams the same S3 sourcequality/multilingualwere mined from, filtered to PR ids already in that 254K+-PR corpus, and asks Jev only forresearch_novelty+ the 424 tag claims (bundled into one call per PR, not 424 separate calls) — reusing that corpus's existingtextfield rather than re-annotating everything from scratch. Dataset build:/ansh/fasttext-stuff/build_research_novelty_v2_dataset.pyandbuild_tags_dataset.py(both replicatequality/multilingual's exact train/val split — same seed, same 10%, same PR ids — for cross-classifier comparability). Training/eval: each subfolder's owntrain.py/evaluate.py.