token-provenance

Every token count carries the encoding that produced it.

A prompt is budgeted with len(text) // 4. The number is plausible. It propagates into a cost estimate. A user sees a bill. The bill is wrong. Nothing in the pipeline flagged it.

Concrete example: a prompt was budgeted with len(text) // 4 = 44 tokens. The serving stack used a ChatML template that adds 11 special tokens, and the actual BPE tokenizer produced 43 tokens for the body. Real count: 54. Under the heuristic the count fit; under the real encoding it did not. The server truncated the prompt silently. The answer was wrong.

token-provenance fixes this by making the encoding a first-class property of the count.

The claim in one sentence

Attach encoding metadata to every token count at the point of production; propagate it through every operation; refuse to format or consume it unless the caller explicitly opts in.

What it produces

Every tokenizer output is a TCount carrying:

  • value โ€” the raw int
  • encoding.tokenizer โ€” which tokenizer produced the count
  • encoding.template โ€” which chat template was applied
  • encoding.special_tokens โ€” how many special tokens the template adds
  • encoding.exact โ€” True iff the count came from a real tokenizer, False for heuristics like len // 4
  • encoding.version โ€” tokenizer version string

Install

pip install token-provenance

Or run the module directly:

python token_provenance.py

Pure stdlib. No dependencies.

Usage

Basic

from token_provenance import (
    tokenizer, serving, report, assert_serving,
    TCount, Encoding, InexactError, ServingMismatchError,
)

SERVING = Encoding(tokenizer="gpt2", template="chatml",
                   special_tokens=11, version="v1", exact=True)

@tokenizer("gpt2", template="chatml", special_tokens=11,
           version="v1", exact=True)
def count_gpt2_chatml(text):
    """Return (count, info)."""
    ids = gpt2_tokenizer.encode(text)
    return len(ids) + 11, {"notes": ("chatml template applied",)}

t = count_gpt2_chatml(prompt)

# Sanctioned reporting.  Refuses inexact or mismatched encodings.
print(report(t, expected=SERVING))

# Sanctioned assertion.  Fails loudly on mismatch.
assert_serving(t, SERVING, "startup")

# Arithmetic propagates encoding.
result = (t + 100) * 2
print(report(result, expected=SERVING))

Strict consumers

@serving(SERVING)
def cost_estimate(count, price_per_1k=0.002):
    """Refuses counts from the wrong encoding."""
    return count.value * price_per_1k / 1000.0

try:
    cost_estimate(naive_count)
except ServingMismatchError as e:
    print(f"refused: {e}")

Arrays

from token_provenance import TArray

# A batch of independent token counts.
counts = TArray([count_gpt2_chatml(p) for p in prompts])

# Elementwise arithmetic preserves the mixed state.
doubled = counts * 2 + 1

# Reductions merge provenance.
total = counts.sum()
if total.exact:
    print(f"total = {total.value}")
else:
    print(f"refused: {report(total)}")

Cross-checks

from token_provenance import cross_check, Mismatch

result = cross_check([count_gpt2, count_llama, count_heuristic],
                     prompt, rel_tol=0.05)

if isinstance(result, TCount):
    print(f"all agree: {report(result)}")
else:  # Mismatch
    print(f"mismatch: {report(result)}")
    idx, name, dist = result.outlier()
    print(f"outlier: {name}")

Adaptive budgets

from token_provenance import safe_context_budget

# Instead of hardcoding 4096 - 512 = 3584:
budget = safe_context_budget(4096, SERVING, reserve_for_output=512)
# = 4096 - 11 - 512 = 3573

The six defenses

The module provides six levels of escalation for the same bug class:

level mechanism action
1. tag @tokenizer every output is a TCount, not an int
2. propagate arithmetic taint spreads through + - * //
3. refuse report() refuses to format inexact or mismatched counts
4. raise @serving raises ServingMismatchError on wrong encoding
5. assert assert_serving() AssertionError at pipeline startup
6. bypass .unwrap(allow_inexact=True) requires explicit opt-in

Any pipeline can choose its level. The choice is explicit at every boundary.

Benchmarks

The demo (python token_provenance.py)

part concept outcome
0 naive pipeline prints 36 with no flag
1 TCount carries exact=False, tokenizer=heuristic:chars//4
2 taint propagation (naive + 100) * 2 = 272, still inexact
3 report() refuses inexact, formats exact as 54 [toy-bpe+chatml(+11)]
4 @serving raises on heuristic and on llama2; accepts chatml
5 escape hatch .unwrap(allow_inexact=True) = 36; .unwrap() raises
6 exact path 54 with tokenizer=toy-bpe, special_tokens=11
7 assert_serving fails on llama2 and heuristic; returns raw 54 on chatml
8 TArray mixed batch 3/4 exact, sum=72; exact batch sum=95
9 adaptive budget none=3584, llama2=3577, chatml=3573
10 cross_check A=mixed TCount, B/C=Mismatch, D=clean TCount
11 discipline the rule set, in one place

Total runtime: <1 second on a laptop CPU.

Template overhead and context budget

safe_context_budget(model_max, encoding, reserve_for_output):

model_max reserve template special budget
4096 512 none 0 3584
4096 512 llama2 7 3577
4096 512 chatml 11 3573

The naive answer 4096 - 512 = 3584 silently overshoots by 11 tokens under ChatML. That is enough to truncate the last few tokens of a prompt โ€” often the JSON closing brace or the instruction suffix.

The outlier rules

When cross_check returns a Mismatch, the outlier is identified by three rules in order:

  1. Inexact wins. A heuristic encoder is the outlier by construction, regardless of where its value landed.
  2. Most special tokens wins. If all encoders are exact but templates differ, the one with the largest template overhead is less trustworthy. Ties break by distance from the min-special-tokens group mean.
  3. Distance from mean. If all encodings are identical, fall back to the value farthest from the mean. The mean is unambiguous for both odd and even counts, unlike the median.

The order matters. A rule based purely on distance-from-mean gets the outlier wrong in two cases:

  • A heuristic that happens to land near the exact encoders is not flagged, even though it was produced by a different process.
  • A chatml count that happens to land near the raw token count is flagged as the outlier when the real problem is the missing template on the other solvers.

Rules 1 and 2 fix both cases. Rule 3 is the fallback for the case where every encoder used the same encoding and still disagreed โ€” which means the disagreement is real, not a provenance artifact.

Value agreement is not encoding equivalence

This is the subtlest point in the module.

cross_check returning a TCount means the values agreed within tolerance. It does not mean the counts are interchangeable with the serving config. A merged count of {toy-bpe, toy-bpe+chatml, toy-bpe+llama2} is not a chatml count, even if its mean value happens to land near one.

assert_serving enforces this: a merged encoding matches the serving config only if every constituent matches. This is why demo Case A is refused even though the three encoders agreed within rel_tol=0.30:

A (refused)    : AssertionError
  A: encoding 'mix[toy-bpe, toy-bpe+chatml(+11), toy-bpe+llama2(+7)]'
     != serving 'toy-bpe+chatml(+11)'

Case D passes because all three tokenizers used the same chatml template, so the merged encoding is trivially a chatml encoding:

D (passes)     : 54

Value agreement is necessary but not sufficient.

When to use it

  • Any pipeline that bills or truncates by token count. Cost estimation, context-window budgeting, rate limiting. If a count decides what the user sees or pays, wrap it.
  • Multi-tokenizer pipelines. When two encoders should agree, cross_check makes disagreement an output rather than a silent pass-through.
  • Serving-config-aware pipelines. When the model, the template, or the special-token set can drift between training and serving, @serving catches the drift at the first call site.
  • Publication-grade methodology. A token count that goes into a paper should carry its encoding. report() refuses to format anything else.

When not to use it

  • When the encoder is always the same and always exact. A single fixed tokenizer at a single fixed template has no provenance to track. The overhead is small but nonzero.
  • When performance is critical and profiling shows the wrapper dominates. The TCount layer is cheap (a few microseconds per operation), but arithmetic on many small counts in a tight loop will pay for the provenance. Unwrap at the boundary and rewrap the result.
  • When downstream code cannot accept the wrapper. Third-party libraries will coerce TCount to a plain int at the call boundary. Unwrap before calling, rewrap after.
  • As a substitute for using the real tokenizer. The module makes the heuristic visible; it does not make it accurate. If your budget allows, run the real tokenizer.

Honest limitations

  • Trusts the tokenizer's self-report. A @tokenizer function can claim exact=True while doing len // 4. The module has no way to verify the claim. The decorator's exact argument is a declaration, not a proof.
  • Integers only, plus TArray. No float counts (subword fractions, weighted averages). Costs must be computed on the unwrapped int.
  • Arithmetic covers + - * //. No %, no **, no math.log for cost formulas that need it. Add the ones you need by extending _arith.
  • Comparisons strip provenance. tc < 4096 returns a plain bool. If you need to propagate the fact that a comparison was made on an inexact count, use assert_serving first.
  • cross_check runs tokenizers sequentially. No parallelism. Tokenizers are usually fast enough that this does not matter; for a remote tokenizer API, run them yourself and construct Mismatch manually.
  • safe_context_budget is an accounting formula, not a guarantee. Real serving stacks may inject additional tokens (system prompts, tool schemas, RAG prefixes) that are invisible to the tokenizer used to produce the count. The budget is a floor on what fits, not a ceiling.
  • No calibration against real tokenizers. The demo uses toy-bpe, a synthetic counter. The numbers in the demo are illustrative, not measured against GPT-2, Llama, or any production tokenizer.

The bug it prevents

The silent truncation from the introduction would have looked like this in a provenance-aware pipeline:

naive budget   : 4096 - 512 = 3584 tokens available
real encoding  : toy-bpe+chatml(+11)
real count     : 54 tokens for a 143-char prompt

report(naive)  : <refused: 'heuristic:chars//4' is a heuristic;
                  no exact tokenizer was used>
report(exact)  : 54 [toy-bpe+chatml(+11)]
assert_serving : AssertionError: encoding 'heuristic:chars//4'
                 != serving 'toy-bpe+chatml(+11)'

The heuristic count never reaches a budget. The prompt is never truncated silently. The failure is caught at the point of production, not at the point of billing.

Version history

version change
0.1.0 Encoding, TCount, @tokenizer, report(), assert_serving()
0.2.0 added TArray for sequences
0.3.0 added cross_check and Mismatch
0.4.0 safe_context_budget accounts for template overhead
0.4.1 Mismatch.outlier() fixed: inexact first, then most-special-tokens, then distance from mean
0.4.2 Encoding.merge flattens nested merges; serving_matches requires every constituent to match

Reference

Part of a series of small tools built in one session:

tool reads answers
hv-manifold a corpus the geometry of style space
hv-reader one text how it reads
anomaly-or-bug a number and a matrix is this a bug or a discovery?
frontier-check one claim where does it sit relative to the frontier?
numerical-provenance a numeric pipeline can I trust this number?
token-provenance an LLM pipeline can I trust this token count?

The design principle โ€” that a token count should carry its own encoding rather than relying on the caller to remember which tokenizer produced it โ€” came out of a session in which a len(text) // 4 heuristic silently disagreed with a ChatML-wrapped BPE tokenizer by 18 tokens. The tool was built to make that class of bug impossible.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support