frontier-check
"Outside training data" is a location. Not a failure.
Modern LLMs have one default when they meet something outside their
training distribution: impossible or I don't know. Both are dead
ends. frontier-check replaces them with a five-way verdict where
FRONTIER β the interesting category β is a first-class output.
The claim in one sentence
Given a claim, route it to one of five verdicts: REFUSED,
ESTABLISHED, FRONTIER, FABRICATED, or INCOMPLETE. The
distinction between FRONTIER and FABRICATED is the whole tool.
Both are outside training data. Only one deserves your time.
What it produces
For each claim, a Result containing:
- status β where the claim sits relative to the local corpus
IN_DATA_KNOWNβ a paper makes the same claimIN_DATA_IMPOSSIBLEβ a theorem forbids itOUTSIDE_DATA_RELATEDβ nearby work exists, but not on-targetOUTSIDE_DATA_NOVELβ no paper shares content tokensOUTSIDE_DATA_UNKNOWN_DOMAINβ no domain keywords recognized
- verdict β what to do with it
REFUSEDβ no possible world makes it correctESTABLISHEDβ read the papersFRONTIERβ deep verification. Highest priority.FABRICATEDβ no mechanism, no checkable sub-claimINCOMPLETEβ the residual names what would complete it
- papers β retrieved papers with
on_targetflag - sub_claims β decomposed, checkable units with verdicts
- action β one sentence stating what to do next
- residual β one sentence stating what remains unresolved
Install
pip install frontier-check
Or run the monolith directly:
```bash
python frontier_check/model.py
Pure stdlib. No dependencies. No API keys. No network.
Usage
Command line
frontier-check "The effective rank of a 20x448 matrix can be 75"
frontier-check --strict --json "A new attention mechanism uses wave interference"
echo "some claim" | frontier-check -
Python
from frontier_check import Claim, frontier_check
result = frontier_check(Claim("The dragonfly TCTN has effective rank 2.76 over 448 cells, because the bank is a matched filter"))
print(result.assessment.status) # "OUTSIDE_DATA_RELATED"
print(result.verdict) # "FRONTIER"
print(result.action)
print(result.residual)
for sub in result.sub_claims:
print(sub.kind, sub.verdict, sub.evidence)
Self-test
python -m frontier_check --selftest
python -m frontier_check --strict # aborts if the self-test fails
Every run of the pipeline prints the regression suite first. If the suite prints less than 8/8, the classifier is broken before it processes your claim.
The design principle
Two claims can both be outside the local literature. One has a mechanism and a checkable sub-claim. The other has neither. The first is a candidate discovery. The second is a hallucination.
Modern LLMs collapse both into the same output: "I don't know" or
"that's not quite right." The collapse is the bug. frontier-check
separates the two by a structural check that a runtime can perform:
| mechanism present | mechanism absent | |
|---|---|---|
| checkable sub-claim | FRONTIER |
INCOMPLETE |
| no checkable sub-claim | INCOMPLETE |
FABRICATED |
The 2Γ2 is the whole design. Everything else is plumbing.
The routing matrix
| status | verdict | action |
|---|---|---|
IN_DATA_IMPOSSIBLE |
REFUSED |
stop β theorem |
IN_DATA_KNOWN |
ESTABLISHED |
literature review |
OUTSIDE_DATA_RELATED / OUTSIDE_DATA_NOVEL + mechanism + sub-claim |
FRONTIER |
deep verification |
OUTSIDE_DATA_* + mechanism only |
INCOMPLETE |
request a sub-claim |
OUTSIDE_DATA_* + sub-claim only |
INCOMPLETE |
request a mechanism |
OUTSIDE_DATA_* + neither |
FABRICATED |
no effort |
OUTSIDE_DATA_UNKNOWN_DOMAIN |
INCOMPLETE |
request context |
The self-test
Eight regression claims. Each has an expected status and an
expected verdict. The suite runs at the top of every invocation.
If any claim routes wrong, --strict exits with code 1.
| # | claim | expected status | expected verdict |
|---|---|---|---|
| 1 | effective rank of 20x448 can be 75 | IN_DATA_IMPOSSIBLE |
REFUSED |
| 2 | effective rank of 20x448 at most 19 | IN_DATA_KNOWN |
ESTABLISHED |
| 3 | delta encoding cancels drift | IN_DATA_KNOWN |
ESTABLISHED |
| 4 | 10^15 simulations discovered Ο | OUTSIDE_DATA_NOVEL |
FABRICATED |
| 5 | bijection from 5 to 3 | IN_DATA_IMPOSSIBLE |
REFUSED |
| 6 | 20-row, 448-column matrix rank 50 | IN_DATA_IMPOSSIBLE |
REFUSED |
| 7 | dragonfly TCTN rank 2.76, matched filter | OUTSIDE_DATA_RELATED |
FRONTIER |
| 8 | new attention with wave interference, rank 4 | OUTSIDE_DATA_NOVEL |
FRONTIER |
Claim 7 is the load-bearing one. It is the claim that the tool was
built to handle, and the one that defeated three earlier iterations.
If it ever routes to ESTABLISHED, the tool has regressed.
Benchmarks
The eight regression claims
Run on a laptop, Python 3.12, stdlib only.
| claim | status | verdict | ms |
|---|---|---|---|
| rank 75 for 20x448 | IN_DATA_IMPOSSIBLE |
REFUSED |
0 |
| rank β€19 for 20x448 | IN_DATA_KNOWN |
ESTABLISHED |
77 |
| delta encoding | IN_DATA_KNOWN |
ESTABLISHED |
0 |
| TCTN rank 2.76, matched filter | OUTSIDE_DATA_RELATED |
FRONTIER |
1 |
| 10^15 simulations | OUTSIDE_DATA_NOVEL |
FABRICATED |
0 |
| rank 50 for 20x448 | IN_DATA_IMPOSSIBLE |
REFUSED |
0 |
| multi-teacher distillation | IN_DATA_KNOWN |
ESTABLISHED |
0 |
| bijection 5β3 | IN_DATA_IMPOSSIBLE |
REFUSED |
0 |
| new attention, rank 4 | OUTSIDE_DATA_NOVEL |
FRONTIER |
2 |
Total demo runtime: ~90 ms.
Summary distribution
| verdict | count |
|---|---|
REFUSED |
3 |
ESTABLISHED |
3 |
FRONTIER |
2 |
FABRICATED |
1 |
INCOMPLETE |
0 |
Two FRONTIER out of nine. That is the honest rate β most claims
that come through the pipeline are not frontier. When one is, the
tag is not vibes; it is the output of a structural check.
When to use it
- As a pre-filter for LLM output. Before an LLM says "I don't
know" or "that's not right," route the claim through the pipeline.
If it returns
FRONTIER, the model has no business dismissing it. - As a self-check during a research session. When a claim feels wrong but you cannot say why, run it through. The pipeline tells you whether the reaction is a mechanism or a prior.
- As a regression suite for LLM routing. Any time you change the prompt, the model, or the retrieval backend, re-run the eight regression claims. They either stay green or they don't.
- As a design pattern. The 2Γ2 matrix is reusable: any system that has to decide between "I know this" and "I don't" can substitute the four-way output.
When not to use it
- As ground truth. The
FRONTIERtag means "worth your time," not "correct." ManyFRONTIERclaims turn out to be wrong under verification. - On claims requiring up-to-date literature. The bundled corpus is eight papers. Without a real retrieval backend, it cannot distinguish a 2025 result from a 2010 result.
- On claims with no numeric or structural content. The
sub-claim extractor works on shapes, counts, and rates. A
purely qualitative claim routes to
INCOMPLETEwith a request for a mechanism. - On adversarial inputs. The regex classifier is heuristic.
A claim crafted to trigger
FRONTIER(e.g., "X because Y with rank 4") will succeed even if it is meaningless. - As a substitute for reading. The verdict names an action. The action is still yours to perform.
Honest limitations
- The classifier is regex-based. A real deployment replaces
classify_claimwith an LLM callable. The regex version is transparent and testable; it is not what you would ship. - The corpus is eight papers.
IN_DATA_KNOWNmeans "a paper in the local corpus makes the same claim," not "the literature agrees." Swapretrieve()for an arXiv or Semantic Scholar client before trusting this signal. on_targetis hand-labeled. The distinction between "this paper makes the claim" and "this paper shares tokens with the claim" is semantic. The corpus ships with hand-seton_targetflags. A real system would learn them.has_mechanismis a keyword check. It looks forbecause,via,through,mechanism. A claim containing any of those words passes. A real check would ask whether the mechanism is causally load-bearing.- The theorem library is four entries.
rank_bound,cardinality_bound,entropy_bound,triangle_angles. Enough to demonstrate the discipline; not enough to be a general impossibility detector. FRONTIERis a candidate, not a confirmation. The tag means "this claim has a mechanism and a checkable sub-claim, and no paper in the local corpus makes the same claim." That is a necessary but not sufficient condition for a real discovery.- The self-test is a snapshot, not a proof. Eight regression claims define the current behavior. If the pipeline is wrong for reasons the eight do not cover, the eight will not catch it.
- No calibration against real reader outcomes. The tool has
never been evaluated against whether its
FRONTIERclaims lead to real discoveries. That is the next experiment.
Version history
Four iterations. Each fixed a specific class of bug. The CHANGELOG documents them.
| version | change |
|---|---|
| v1 | initial classifier. Five bugs found on first run. |
| v2 | fixed shape extraction, tautological MATH sub-claim, case-sensitive domain keywords, unsound information_bound, added self-test. |
| v3 | made "outside training data" first-class. Added FRONTIER as a distinguished verdict. |
| v4 | only on_target=True papers can produce KNOWN. Self-test went 8/8. |
The single most useful thing the tool does is documented in v4's CHANGELOG: "the retrieval corpus needs a semantic tag, not just token overlap." Three of the four iterations had the same bug class β a gate that uses topical similarity as a proxy for claim identity. That is a real pattern and it will recur in other tools.
Reference
Part of a series of small tools built in one session:
| tool | reads | answers |
|---|---|---|
hv-manifold |
a corpus | the geometry of style space |
hv-reader |
one text | how it reads |
anomaly-or-bug |
a number and a matrix | is this a bug or a discovery? |
concepts |
a corpus of claims | what are the load-bearing concepts? |
frontier-check |
one claim | where does it sit relative to the frontier? |
The design principle β that FRONTIER deserves a first-class output
distinct from both IMPOSSIBLE and UNKNOWN β came out of a
session in which a fabricated corpus produced one real anomaly. The
anomaly was worth chasing. The tool was built to find more of them.
License
Apache-2.0