YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

PR quality/multilingual/tags fastText classifiers

Four fastText classifiers that score or tag a GitHub PR from its diff + touched-file content, distilled from Jev (typesafe/jev-1.13) annotations on mined GitHub Archive PRs. Each subfolder is self-contained: model.bin (weights), train.txt/val.txt (the exact fastText training data), train.py (the exact training call — rerun in place to reproduce model.bin from train.txt), and infer.py (standalone scoring script — fetches a live PR via the GitHub API and classifies it, no dependency on the original mining pipeline).

Usage

pip install fasttext==0.9.3 "numpy<2" requests   # numpy<2 is required -- see note below
export GITHUB_PAT=...                             # any GitHub token (read-only is fine)
python quality/infer.py https://github.com/owner/repo/pull/123
python multilingual/infer.py https://github.com/owner/repo/pull/123
python research_novelty/infer.py https://github.com/owner/repo/pull/123
python tags/infer.py https://github.com/owner/repo/pull/123               # multi-label
python tags/infer.py --threshold 0.3 https://github.com/owner/repo/pull/123

To retrain a model from its own train.txt in place: cd <subfolder> && python train.py.

quality/multilingual/research_novelty each print a soft score (probability- weighted expected value, the better-correlating of the two — use this) and a hard score (argmax level), both on a 0-10 scale. tags instead prints every tag (of 424) whose predicted probability clears --threshold (default 0.5), sorted by confidence — see its own section below for why it's one multi-label model, not 424 separate ones.

Note on fastText + NumPy 2.x: the prebuilt fasttext-wheel can silently corrupt batch-predict() probabilities under NumPy 2.x, and even a from-source build's Python wrapper crashes outright on np.array(..., copy=False). Pin numpy<2.

quality/ — Jev's strict 0-10 quality rubric

Rates the PR's diff on a strict, sensibility-gated 0-10 scale: 0-5 is gated on whether the codebase/diff is coherent and credible (a simple-but-clean change tops out at 5); 6+ requires genuine engineering or conceptual complexity on top of that. See annotate_pr_quality.py's QUALITY_LEVELS for the full rubric text.

  • Hyperparameters: dim=64, epoch=10, lr=0.3, wordNgrams=2, minCount=3, bucket=2,000,000
  • Train: 254,143 PRs (train.txt)
  • Val: 28,238 PRs, 1,766 at level≥6 (val.txt)

This is the plain dim64/epoch10 baseline — every variant tried (2-4x capacity, class- rebalanced train sets, a ~1.55x-larger combined corpus with a separately-mined higher-credibility PR pool) came out flat or slightly worse on a held-out 10,000-PR tracking set (685 at quality≥6). More data/capacity does not fix this classifier; treat it as a known ceiling, not a bug to chase.

Metrics (10,000-PR tracking holdout, rule = quality >= 6):

metric value
rule accuracy 93.73%
precision 74.17%
recall 12.99%
exact-level accuracy 78.01%
MAE 0.531
soft Pearson / Spearman 0.7252 / 0.6339

Read this before using it as a filter: it misses ~87% of genuinely good PRs (quality≥6) to keep false positives to 0.33% of negatives, and its own predictions never exceed ~6.8-7 on any PR we've tested — it is a high-precision, very-low-recall filter, not a general-purpose quality scorer. It also gave near-identical predictions (4.4-5.3) to real, high-quality ML/research PRs (e.g. lucidrains' x-transformers, imagen-pytorch) as to trivial personal-repo PRs — it is not sensitive to genuine technical depth. Lowering the decision threshold to >=5 makes this worse, not better (accuracy drops to 37%, since ~80% of real PRs are already >=5 and the model's own predicted mass sits below that line).

multilingual/ — non-English comment/string content

Purely descriptive 0-10 scale: how much non-English natural-language text (comments, docstrings, string literals) appears in the diff and touched files, independent of the programming language itself. See annotate_pr_quality.py's MULTILINGUAL_LEVELS.

  • Hyperparameters: dim=128, epoch=25, lr=0.3, wordNgrams=2, minCount=3, bucket=2,000,000 (best of a capacity sweep vs dim64/epoch10 baseline and dim256/epoch25 — dim128 won on every axis, dim256 added nothing further for 2x the model size)
  • Train: 254,143 PRs (train.txt)
  • Val: 28,238 PRs (val.txt)

Metrics (100-PR live holdout, 79 true positives, rule = multilingual <= 1):

metric value
rule accuracy 96.00%
precision 96.30%
recall 98.73%
MAE 0.586
Pearson / Spearman 0.906 / 0.721

This one is genuinely reliable as a pre-filter — high recall and precision. If you only trust one of these two classifiers for a filtering decision, trust this one.

research_novelty/ — research/IP content (⚠️ still recall≈0 at the >=6 rule)

0-10 scale rating the extent to which the PR implements technical algorithms, ideas from research papers, or original design/IP, as opposed to routine engineering — independent of how difficult or high-quality the engineering itself is. See annotate_pr_quality.py's RESEARCH_NOVELTY_LEVELS.

  • Hyperparameters: dim=64, epoch=10 (same baseline as quality/multilingual)
  • Train: 96,634 PRs, 297 at level≥6 (train.txt)
  • Val: 10,772 PRs, 35 at level≥6 (val.txt)
  • Same train/val PR ids as tags/ (both come from the same annotation batch; see Provenance) — comparable across the two.

Updated 2026-09-30: retrained on a new 107,406-PR annotation batch (up from the original 39,081-PR/98-positive pass), joined back to quality/multilingual's existing text field so the same 254K+-PR pool's diff/file content is reused rather than re-mined. Labels use the field's own fixed 0-10 scale (int truncation), same as quality/multilingual — not sample quantiles.

Metrics (own val.txt, 10,772 PRs):

metric value
exact-level accuracy 82.45%
within-1-level accuracy 94.06%
MAE 0.2745
hard Pearson / Spearman 0.5679 / 0.5069
soft Pearson / Spearman 0.6662 / 0.5470
rule (>=6) accuracy 99.68%
rule (>=6) precision undefined (0 predicted positive)
rule (>=6) recall 0.00% (0/35)

Read this before using it as a filter: correlation improved modestly over the old diagnostic pass (soft Pearson 0.60→0.67) and the accuracy/MAE numbers look good, but that's almost entirely the near-100%-easy "correctly predict low novelty" majority class — at the actual decision boundary that matters (>=6), it now has enough val data to say confidently (35 true positives, not 2) that it misses every single one, the same recall-collapse pattern as quality. Its own predictions rarely if ever clear ~6-7 (see the Megatron-LM spot-check below). Treat it as a continuous novelty signal (the soft score correlates reasonably, 0.67) rather than a working >=6 binary classifier.

Spot-check: NVIDIA/Megatron-LM#7762 (an Engram/CUDA synchronization change) scores soft=2.46, hard=1 — plausible as "some genuine ML-systems depth but not paper-derived," though still likely underrating it given the classifier's known insensitivity to real technical depth (see quality's note above; the same failure mode shows up here).

tags/ — multi-label topic tagging (424 tags)

Multi-label classification over the topic vocabulary in /ansh/harbor-tags.txt (424 distinct tags after dedup — e.g. python, machine-learning, web-security, gpu, formal methods) — any number of tags can apply to one PR, independently. Each tag is an independent Jev "noul" (boolean-probability) claim ("this PR's diff or touched files involve code/content related to topic X"), bundled into one Jev call per PR alongside research_novelty (see annotate_tags_research.py).

One joint multi-label model, not 424 separate binary classifiers. fastText's native multi-label mode (loss="ova", one-vs-all) shares a single embedding/vocabulary across all 424 tags and trains in one pass — dramatically cheaper than 424x the training time/storage for what is fundamentally the same text signal, at the cost of slightly less per-tag threshold flexibility (mitigated by tuning --threshold in infer.py per use case; default 0.5).

  • Hyperparameters: dim=64, epoch=10, lr=0.3, wordNgrams=2, minCount=3, bucket=2,000,000, loss=ova
  • Train: 96,604 PRs (train.txt) / Val: 10,769 PRs (val.txt) — same PR ids as research_novelty/'s updated split (both from the same annotation batch)
  • 33 PRs with zero tags above threshold were dropped (fastText needs ≥1 label/line)

Metrics (10,769-PR val set, pooled across all 10,769×424 PR-tag pairs, threshold=0.5):

metric value
precision 81.95%
recall 57.17%
f1 67.35%
pooled Pearson r (vs. Jev's raw tag_probs) 0.8055
pooled Spearman ρ 0.7145

Precision stays high across common tags (0.79-0.99) with recall degrading gracefully as per-tag support drops — normal multi-label behavior, not a red flag. This is the best-behaved of the four classifiers relative to its task difficulty (a 424-way multi-label problem vs. a single 0-10 scale).

Known data-hygiene issue: harbor-tags.txt has a couple of near-duplicate tag pairs that survived as distinct labels (data processing/data-processing, data engineering/data_engineering — the latter pair also collided post-slugification and one was suffixed _dup in tag_vocab.json to disambiguate). Worth deduping the source file; not fixed here since it would invalidate the current label vocabulary/model.

Spot-check: vercel/next.js#99499 (a tag-revalidation fix) → 30 tags incl. typescript, react, nextjs, frontend, routing, concurrency, web-server. NVIDIA/Megatron-LM#7762 → 36 tags incl. machine-learning, pytorch, cuda, gpu, distributed-systems, performance-optimization, deeply-technical-ip. Both match what a human skim of the PRs would tag.

Provenance

  • Distilled from typesafe/jev-1.13 (via OpenRouter) on GitHub Archive-mined PRs.
  • Full annotation/training pipeline: /ansh/prism-datagen/annotate_pr_quality.py (rubric + Jev calls, incl. HARBOR_TAGS/TAG_CLAIMS), /ansh/fasttext-stuff/run.py/ train_sweep.py (quality/multilingual training — each subfolder's train.py is the relevant one of these, trimmed to a single model), /ansh/fasttext-stuff/evaluate.py/ evaluate_sweep.py/compare_pre_post.py (eval).
  • quality's known ceiling was diagnosed, not guessed: a dim/epoch capacity sweep (64→128→256), a class-rebalancing ablation (10%/15% minority oversampling), and a ~1.55x larger combined training corpus (adding a separately GitHub-credibility-filtered PR pool skewed toward higher quality) were all tried and none moved recall past ~14%.
  • research_novelty/tags come from a separate, lighter-weight annotation pass: /ansh/prism-datagen/annotate_tags_research.py re-streams the same S3 source quality/multilingual were mined from, filtered to PR ids already in that 254K+-PR corpus, and asks Jev only for research_novelty + the 424 tag claims (bundled into one call per PR, not 424 separate calls) — reusing that corpus's existing text field rather than re-annotating everything from scratch. Dataset build: /ansh/fasttext-stuff/build_research_novelty_v2_dataset.py and build_tags_dataset.py (both replicate quality/multilingual's exact train/val split — same seed, same 10%, same PR ids — for cross-classifier comparability). Training/eval: each subfolder's own train.py/evaluate.py.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support