LogLens: can you tell whether a tiny vision model will learn a task — before training it?

One number, computed from task imagery alone (no network, no gradients, no training), predicts whether a ~10⁵-parameter model can learn a visual change-dating task. Blind-validated at Spearman ρ = 0.988 on ten pre-registered conditions. This repo contains the predictor, the mechanism that explains why it works, the measured boundary where it stops working, and a one-command harness so you can reproduce — or kill — the result on your own machine and your own scene.

Why this might interest you

  • It is task-side, not architecture-side. Zero-cost NAS proxies (synflow, GradNorm, ZiCo, NTK scores) rank architectures for a fixed dataset using the network's gradients at init. This asks the complementary question: given a fixed small model, is this task learnable at all? The pixel-space predictor uses no network whatsoever.
  • There is a mechanism, not just a correlation. The predictor works because an untrained encoder preserves and amplifies the pixel signal-to-clutter ratio into the feature space its head learns from (ρ = 0.981 across 28 task cells; ρ = 0.932 against trained accuracy, 5 seeds per config).
  • The failure boundary is measured, not assumed. On natural-image classification the predictor inverts (ρ = −1.00) and no trainless variant rescues it. We explain why, and state the scope precisely.
  • Everything was pre-registered, and the failures are published. Five claims were falsified by their own controls during this work; all are reported.

Quickstart — reproduce (or refute) it in one command

pip install torch numpy opencv-python scipy
python replicate.py --bank <folder-of-~1000-photos-of-any-room> --full --seeds 3

It builds a 64px bank from your photos, computes the predictor, trains a size grid blind, and prints a PASS/FAIL scorecard:

test bar our result
P1 — predictor ↔ feature-space predictor ρ ≥ 0.90 0.981 (home), 0.969 (2nd scene)
P2 — the law: predictor ↔ learnability ρ ≥ 0.85 0.988 (10-cell blind); 0.932 / 0.850 at 5 seeds
P3 — ignition present (learner escapes, floor pinned at log K) present CNN + ViT, both scenes

Compare against EXPECTED_RESULTS.md. Disagreement is a result — please open a discussion with your results_replicate.json and scene hash.

Read this before you run it (hard-won)

  1. The learn/fail boundary is seed-bimodal and slow. A cell can look "collapsed" at 8k steps and escape at 18k. Escape times spread 6k–18k. Always use ≥ 2 seeds (ideally 3) and --full. Single-seed results near the boundary are noise.
  2. Collapse = loss ≈ log(num_classes) exactly (uniform guessing). Use it to detect dead runs with zero eval cost.
  3. Compare ordering across scenes, never absolute thresholds. Ordering is scene-universal (ρ = 0.996 across two scenes); the threshold is scene-local, so a cell that floors on our scene may learn on yours. That is expected.
  4. Don't use the raw signal/clutter ratio to judge escape mid-training — it is confounded (it spikes when the encoder degenerates and clutter → 0). Track the signal component or the head's margin.

What's here

replicate.py              one-command harness (PASS/FAIL scorecard)
EXPECTED_RESULTS.md       our numbers + the caveats above
REPLICATION_KIT.md        long-form protocol (Windows OK, CPU OK)
DRAFT_paper_v3.md         the paper draft
LITERATURE_CHECK.md       what is already published (and what we therefore do NOT claim)
triclock/, simreal/       the ~128K-param model + the task generator (imported by replicate.py)
code/compute_R.py         the pixel-space predictor, standalone
code/compute_R_grid.py    predictor over a size × contrast grid
code/compute_featR.py     the same quantity through an untrained encoder (the mechanism)
code/compiler_v2.py       automated measure → route → escape-and-select
results/                  the JSONs behind every number in the paper

Scope — please read before generalizing

This works where the discriminative signal already lives in a fixed representation: our change-dating task (ρ = 0.988) and algorithmic phase-transition tasks. It provably fails for natural-image classification, where within-class variance is the learnable material — we tested pooled, high-pass, edge, random-conv, gated, and k-NN forms; all fail (best ρ ≈ +0.04 against a ground-truth ordering, with the plain predictor at ρ = −1.00).

Grokking results in the paper's §7 are reported as supporting evidence only; that ground is already covered by OmniGrok (ICLR 2023) and 2026 work on norm-separation delay laws and interventional/predictive grokking frameworks. We claim no novelty there. See LITERATURE_CHECK.md.

Limitations

One scene family, one architecture class, one training recipe, n = 10 blind conditions, and no external replication yet. That last one is the point of this release.

Citation

@misc{loglens2026,
  title  = {Predicting Task Learnability for Nano Vision Models --- and Why the Prediction Works},
  author = {Cox, Joe and {Claude (Anthropic)}},
  year   = {2026},
  note   = {Draft. Code and pre-registrations: this repository.}
}

Experiments were run on a single Jetson AGX Orin. Every prediction in the paper was registered to a timestamped file before the corresponding run.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support