- Auto-Researching, Not Hyperparameter Tuning β NeurIPS 2026 Anonymous Reproduction Package
- Contents
- Reproducing the headline numbers
- Nexar deployed champion β 0.910 mAP_ALL (Section 5)
- Nexar search-policy ceiling β 0.727 (Table 6)
- Cross-backbone transfer (paper Table 4; V-JEPA 2 0.906 vs retrained MViTv2-S 0.692 at matched HPs β per-run data in
mvitv2s_matched_hp/) - ANOVA decomposition β primary $\eta^2_{\mathrm{arch}} = 0.74$ (cell-capped, $\omega^2=0.716$; sensitivity range 0.20β0.79) (Section 4)
- E2E LoRA boundary ablation β $\eta^2_{\mathrm{arch}} \downarrow 0.04$ (Section 9)
- ASR 8B LoRA β 5.30% WER (Section 7)
- Experiment log sample
- Citation
- License
- Contents
Auto-Researching, Not Hyperparameter Tuning β NeurIPS 2026 Anonymous Reproduction Package
This repository is the anonymous reproduction bundle for the NeurIPS 2026 submission "Auto-Researching, Not Hyperparameter Tuning: Convergence Analysis of 10,000 Experiments".
It contains the checkpoints, representative run logs, and precomputed analysis artifacts required to reproduce every headline number in the paper. Real-name hosting will be provided upon acceptance.
Contents
checkpoints/
champion-vjepa2-deploy/ # Deployed single-model champion (0.910 mAP_ALL on Nexar)
search-best-vjepa2/ # LLM search-policy best (0.727 mAP on Nexar)
best-non-vjepa2/ # Strongest non-V-JEPA 2 baseline (Table 3 cross-backbone)
asr-8b-lora/ # 8B LoRA adapter (5.30% WER, #1 on Open ASR Leaderboard)
experiments_sample/ # 256 Nexar + 50 ASR run logs sampled from ~3,190 + ~2,643
computed_values/ # JSON artifacts driving all figures and tables
data/
anova.json # Section 4 ANOVA decomposition
convergence.json # Figure 1 search-policy trajectories
e2e_anova.json # Section 6 E2E LoRA boundary ablation
ablation.json # Obfuscated-names ablation
cost_efficiency_deep.json # Appendix cost table
deployable_analysis/ # Cross-checks for Section 5 (SMAC / LLM test-set results)
oracle_correction.json # Sample-size-matched comparison
results/
best_model.pt # Full V-JEPA 2 champion (weights + head)
soup_best.pt # Model-soup variant used in ablation
Reproducing the headline numbers
Commands below use a standard torch + huggingface_hub + peft environment.
See the requirements.txt in the anonymous code repository
(https://anonymous.4open.science/r/nips-2026-submission-1C52) for pinned versions.
Nexar deployed champion β 0.910 mAP_ALL (Section 5)
The champion result of 0.910 mAP_ALL requires 12-view TTA with cv_mix
aggregation (alpha=0.95, tuned on public split). See paper Table 6 for the full
TTA/aggregation grid; the paper headline is the 5-seed 12-TTA mean 0.9093 Β± 0.0026. The cv_mix strategy uses
alpha * last_clip + (1 - alpha) * top6_mean to combine temporal position
with confidence ranking.
Without cv_mix (i.e., using default mean+max aggregation with 4-TTA), the same checkpoint achieves 0.906 mAP_ALL β this is the 4-TTA baseline, not the deployed champion.
# Download full V-JEPA 2 checkpoint + head
huggingface-cli download nips2026-reviewer-artifacts/nips2026-artifacts-3894 \
computed_values/results/best_model.pt \
--repo-type model --local-dir ./ckpts
# Champion: 12-TTA + cv_mix (0.910 mAP_ALL)
python scripts/nexar/eval_e2e.py \
--checkpoint ckpts/computed_values/results/best_model.pt \
--tta 12 --aggregation cv_mix
# Expected: mAP_ALL = 0.910 (Public 0.923 / Private 0.900)
# Baseline: 4-TTA + mean aggregation (0.906 mAP_ALL)
python scripts/nexar/eval_e2e.py \
--checkpoint ckpts/computed_values/results/best_model.pt \
--tta 4 --aggregation mean
# Expected: mAP_ALL = 0.906
Nexar search-policy ceiling β 0.727 (Table 6)
huggingface-cli download nips2026-reviewer-artifacts/nips2026-artifacts-3894 \
checkpoints/search-best-vjepa2 --repo-type model --local-dir ./ckpts/search_best
python scripts/nexar/eval_e2e.py --checkpoint ckpts/search_best/best_model.pt --tta 1
# Expected: mAP β 0.727
Cross-backbone transfer (paper Table 4; V-JEPA 2 0.906 vs retrained MViTv2-S 0.692 at matched HPs β per-run data in mvitv2s_matched_hp/)
huggingface-cli download nips2026-reviewer-artifacts/nips2026-artifacts-3894 checkpoints/best-non-vjepa2 \
--repo-type model --local-dir ./ckpts/non_vjepa
python scripts/nexar/cross_backbone_transfer.py \
--vjepa ckpts/search_best/best_model.pt \
--alt ckpts/non_vjepa/best_model.pt \
--output xfer.json
ANOVA decomposition β primary $\eta^2_{\mathrm{arch}} = 0.74$ (cell-capped, $\omega^2=0.716$; sensitivity range 0.20β0.79) (Section 4)
The ANOVA numbers are fully reproducible from the precomputed JSON without re-running any experiments:
python3 -c "
import json
d = json.load(open('computed_values/data/anova.json'))
print('Nexar eta^2_arch:', d['nexar']['eta_squared']['architecture'])
print('UCF-101 eta^2_arch:', d['ucf101']['eta_squared']['architecture'])
"
# Nexar eta^2_arch: 0.74 cell-capped primary (0.51 adaptive sensitivity) | UCF-101 eta^2_arch: 0.148
To reproduce from scratch, run the analysis script from the anonymous code repo against the sampled (or full) experiment logs:
python analyze_predictions.py \
--experiments experiments_sample/ \
--output anova_from_sample.json
E2E LoRA boundary ablation β $\eta^2_{\mathrm{arch}} \downarrow 0.04$ (Section 9)
python3 -c "
import json
d = json.load(open('computed_values/data/e2e_anova.json'))
print(d['eta_squared'])
# {'backbone': 0.038, 'encoder': 0.0006, 'backbone_x_encoder': 0.043, 'learning_rate': 0.823}
"
ASR 8B LoRA β 5.30% WER (Section 7)
huggingface-cli download nips2026-reviewer-artifacts/nips2026-artifacts-3894 checkpoints/asr-8b-lora \
--repo-type model --local-dir ./ckpts/asr_8b
# Merge adapter on base model and evaluate on Open ASR test bundle
python asr_eval.py --adapter ckpts/asr_8b/ --benchmark open_asr
# Expected WER: 5.30%
Experiment log sample
experiments_sample/ contains 150 Nexar + 50 ASR runs sampled uniformly at
random (seed 20260421) from the full campaign. Each run directory contains:
idea_config.yamlβ the configuration proposed by the agentmetrics.jsonβ the resulting test-set metricsclaim.jsonβ the agent's rationale (where logged)
This sample is sufficient to independently re-run the ANOVA script and verify the variance decomposition to within sampling error. The complete ~10,000-run campaign log will be released upon acceptance.
Citation
The paper is under double-blind review; citation metadata will be added upon acceptance.
License
Apache 2.0. Third-party model weights (V-JEPA 2; the publicly released ASR base model, name withheld for double-blind review) remain subject to their original licenses.