hv-tail

Above-ceiling detector. Estimate whether the user is operating above the model's calibrated ceiling for this query type.

The claim in one sentence

A user who sees the answer, points at it, and is correct without reading is operating above the mean. Modern LLMs are calibrated on the mean. hv-tail is the first model on the Hub that detects the tail.

What it detects

signal what it captures
framing "in the language of X", "from the perspective of Y"
precision formal notation, technical vocabulary, backticks, math
undecidability "is there", "can one prove", "well-defined", "GΓΆdel"
meta "your reasoning", "are you sure", "what are your limits"
negative "not the standard", "don't just", "beyond the obvious"
constraint multiple simultaneous requirements
dialect non-standard register signals
requery history shows repeated refinements on one topic

Each signal returns a value in [0, 1]. The raw score is sum(signal Γ— weight) / 3.0, capped at 1.0. Three strong weighted signals saturate.

Install

pip install numpy


Actually β€” no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+
runs.

## Usage

### Score a query

```python
from hv_tail import HVTail

m = HVTail()
r = m.score("what is a black hole")
# {'above_ceiling': False, 'below_ceiling': True,
#  'confidence': 1.000, 'raw_score': 0.000,
#  'recommendation': 'proceed', ...}

Score with history

r = m.score(
    "then how does Hawking radiation preserve unitarity",
    history=[
        "what is a black hole",
        "but why does the horizon have no local features",
        "and does that mean information is destroyed",
    ],
)
# {'above_ceiling': False, 'confidence': 0.345,
#  'raw_score': 0.490, 'recommendation': 'ask', ...}

Verbose explanation

print(m.explain("your reasoning on that seems inconsistent..."))

Calibrate to your population

m.calibrate(
    above_examples=[
        "in the language of category theory, is there a functor...",
        "your reasoning on that last answer seems inconsistent...",
    ],
    below_examples=[
        "what is a black hole",
        "how do I make coffee",
    ],
)

The thresholds are fit from the labeled examples:

  • Clean separation β€” above and below hug the gap between the two clusters, with a proportional margin.
  • Overlap β€” thresholds maximize TPR βˆ’ FPR (above) and TNR βˆ’ FNR (below).

CLI

# score a single query
python hv_tail.py --query "what is a black hole"

# verbose
python hv_tail.py --query "..." --explain

# score with history from a file
python hv_tail.py --query "..." --history history.json --explain

# JSON output
python hv_tail.py --query "..." --json

# no args = run all demos
python hv_tail.py

The formula

raw_score = min(1.0, sum(signal_i Γ— weight_i) / 3.0)

where the eight signals and default weights are:

signal weight what fires it
framing 1.0 explicit requests to reframe
precision 1.2 formal notation, technical terms
undecidability 1.8 questions that require "I don't know"
meta 1.8 questions about the model's own reasoning
negative 1.1 explicit rejection of the standard answer
constraint 0.7 multiple simultaneous requirements
dialect 0.5 non-standard register
requery 1.8 repeated refinements on one topic

NORMALIZER = 3.0 means three strong weighted signals saturate the score. This is a "noisy-OR-like" interpretation: independent weak signals multiply their evidence.

The two weights that matter most:

  • undecidability (1.8): a query that requires "I don't know" is the strongest signal that the mean-trained model is about to answer the wrong question.
  • meta (1.8): a query about the model's own reasoning is a direct signal that the user is treating the model as an object, not an oracle.

Either one alone crosses the above-ceiling threshold. Both together saturate the score.

Recommended actions

verdict condition action
proceed raw_score ≀ 0.20 the model can answer normally
ask 0.20 < raw_score < 0.55 a clarifying question is needed
reframe 0.55 ≀ raw_score < 0.85 present multiple frames
refuse raw_score β‰₯ 0.85 say "I don't know"

The recommendation layer is what makes this model useful: not just a score, but an action. Every previous detector on the Hub returns a number. This one returns a decision.

Benchmarks

Mean-user query

query: "what is a black hole"
raw_score     : 0.000
above_ceiling : False
below_ceiling : True
recommendation: proceed

All eight signals fire at zero. This is the reference point.

Tail-user query

query: "in the language of category theory, is there a functor
        from the black hole information paradox to a broader
        framework of unitarity that doesn't assume the standard
        measurement postulate?"
raw_score     : 1.000
above_ceiling : True
recommendation: refuse

Three signals fire (framing 0.33, precision 0.50, undecidability 1.00, negative 0.50). The score saturates. The model correctly identifies this as above the ceiling and recommends refusing rather than answering.

Borderline query

query: "explain this in the style of a Feynman diagram but not
        the standard textbook version"
raw_score     : 0.303
above_ceiling : False
below_ceiling : False
recommendation: ask

Framing 0.33 and negative 0.50, but nothing else. This is a user reaching toward the tail but still framing the question at the mean. The model says "ask" β€” a clarifying question would help.

Requery history

history:
  [0] "what is a black hole"
  [1] "but why does the horizon have no local features"
  [2] "and does that mean information is destroyed"
query: "then how does Hawking radiation preserve unitarity"
raw_score     : 0.490
above_ceiling : False
recommendation: ask

Requery fires at 0.67. Four consecutive refinements is a strong pattern; the model is on the edge of the ceiling.

Meta query

query: "your reasoning on that last answer seems inconsistent
        with what you said earlier. can you explain your
        calibration? are you sure the assumption you made is
        defensible?"
raw_score     : 0.631
above_ceiling : True
recommendation: reframe

Meta fires at 1.00. A user asking about the model's own reasoning is directly treating it as an object to inspect. This alone crosses the ceiling.

Batch comparison

name score above below conf recommendation
mean 0.000 False True 1.000 proceed
borderline 0.303 False False 0.587 ask
tail 1.000 True False 1.000 refuse
requery 0.490 False False 0.345 ask
meta 0.631 True False 0.179 reframe

Calibration behavior

The calibrate() method fits the two decision thresholds. It handles both cases:

Narrow gap (above and below scores nearly touch):

above-threshold after : 0.0670
below-threshold after : 0.0223
gap                   : 0.0446

Wide gap (clusters separate cleanly):

above-threshold after : 0.5806
below-threshold after : 0.2500
gap                   : 0.3306

The margins are proportional to the gap. Narrow gaps produce tight thresholds; wide gaps produce comfortable ones.

Why this is a genuine new category

Every model on Hugging Face assumes the user is at the mean:

  • text generation β€” assumes the reader wants text
  • image captioning β€” assumes the image is the topic
  • classification β€” assumes the label is the question

None of them model the user. hv-tail is the first model on the Hub that classifies the person, not the content. Specifically: is this person operating above the ceiling their tool was calibrated for?

This is the "genius fails the IQ test" problem, packaged as a detector. It's the missing diagnostic in every LLM product's infrastructure. Combined with hv-mode and hv-ttu, it forms the first complete reader-model stack:

  • hv-tail β€” who is asking?
  • hv-mode β€” what shape should the answer take?
  • hv-ttu β€” how long will it take to understand?

Honest limitations

  • Signals are regex-based. They cover the common tail-user markers. A user operating above the ceiling without those markers will not be detected.
  • Coefficients are heuristic. They are plausible and internally consistent. They are not fit to a labeled corpus. Calibrate on your own data.
  • Threshold semantics are domain-specific. raw_score β‰₯ 0.55 means "above the mean," not "above GPT-4's mean" or "above Gemini's mean." Every model has its own calibration.
  • History assumes relevance. The requery signal assumes the caller passes history from the same conversation. Passing unrelated history will inflate the score.
  • No memory across sessions. The model has no persistent state. Each call is independent.
  • No semantic understanding. It can detect that a user appears to be operating above the ceiling. It cannot verify that they actually are.
  • Calibration is threshold-only. The signal weights are fixed. To re-weight signals for a specific population, edit the config directly.

Reference

Extracted from the XuanJi-ISA exploratory track, "Visual-Whole Reasoning Interfaces" (issue #122), specifically the blueprint that proposed reader models as an unmodeled axis in LLM delivery.

The core insight β€” that the mean is not the population, and that the tail is where the interesting users live β€” is the same one behind the entire reader-model category.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support