Piko

A closed decision model. Give it a state and a closed question; it returns the decision β€” a chosen option, a full probability distribution, and a calibrated right to abstain.

Piko does not generate text. It reads the next-token distribution of a single step β€” one chat/completions request, max_tokens: 1, logprobs on β€” keeps only the probability mass on the option letters (A., B., C. …), and renormalises. What comes back is the decision itself: the chosen option, the distribution over the options, the margin between the top two, and a calibrated abstention when that margin falls under a threshold.

Because the answer is confined to the menu it was handed, Piko cannot invent an option that was not offered β€” and when the read is too close it declines instead of guessing. There is nothing to parse and nothing to hallucinate.

Every modality is served by one unified multimodal decoder: text, image, video and audio ride the same model, the same request and the same endpoint, with no routing between modalities and no separate audio or vision model. Media is read as raw pixels and raw waveform, never as captions or transcripts.

System 1 / System 2. Piko is the fast System 1 read β€” one token, a probability. The same model in full chat mode is the System 2 answer. Escalation is a route change, not a second dependency: decide first with one token, then send the uncertain cases (margin under theta) to the same model as a normal chat call.

Playground Website API License Python Dependencies


What it is

A normal model answers a closed question by writing an answer and hoping the text parses. Piko answers it by reading a distribution. The difference is not cosmetic:

generate the answer read the decision (Piko)
cost a completion you pay for one token, completions are free
latency as long as the text you write one prefill + one decode step
output free text β†’ a parser the verdict, already typed
confidence none, unless you ask again the margin, in the same call
failure a plausible wrong string a named DEGRADED / UNDECIDED, no guess
off-menu answers possible impossible β€” the codomain is closed

Quickstart

The client is stdlib only β€” no build step, no dependencies.

git clone https://huggingface.co/Agnosia-Piko/piko
cd piko
export PIKO_API_KEY=...          # a key from https://www.agnosia.tech
from piko import Piko

piko = Piko()                    # https://www.agnosia.tech/v1 by default

r = piko.choice(
    state    = "The minister announced a plan on Tuesday.",
    question = "event type",
    options  = ["ANNOUNCE", "MEET", "VOTE"],
)

print(r.choice)                  # ANNOUNCE    the verdict, already typed
print(r.p1, r.margin, r.band)    # 0.999… 0.999… high
print(r.probabilities)           # [0.9999, 3.1e-07, 9.7e-08]  over the menu
if r.undecided:
    print(r.explain())           # what a thin margin looks like

Three closed forms, one wire:

form you pass you get back
choice options=["A_code", …] (≀ 26, your order) r.choice β€” a code, or None
noul a boolean question r.value β€” True / False, any language
score scale=5 r.value β€” a grade; r.expectation β€” the mean over the scale
b = piko.noul("The minister announced a plan.", "is this about policy?")
print(b.value, b.p1)                                   # True 0.97

s = piko.score("The minister announced a plan with three named measures.", "how concrete is it?", 5)
print(s.value, s.expectation)                          # 4 4.31

The wire β€” drop-in OpenAI

There is no bespoke SDK on the server. Piko is the OpenAI Chat Completions contract, plus the provider document. Any OpenAI client works unchanged.

curl -s https://www.agnosia.tech/v1/chat/completions \
  -H "Authorization: Bearer $PIKO_API_KEY" -H 'content-type: application/json' -d '{
  "model": "piko",
  "messages": [{
    "role": "user",
    "content": "Choose the correct option. Reply with only its letter.\n\nContext:\nThe minister announced a plan on Tuesday.\n\nQuestion: event type\nA. ANNOUNCE\nB. MEET\nC. VOTE"
  }],
  "max_tokens": 1,
  "logprobs": true,
  "top_logprobs": 20
}'

The prompt layout is fixed and it is the format instruction β€” Choose the correct option. Reply with only its letter., then Context:, the state, Question: <name>, then one X. option per line. The library builds it for you (piko.prompt(...) returns it without sending anything β€” the first rung of every diagnosis).

What comes back

field meaning
choice the option code, or null β€” null is the absence of a verdict, not a fallback
probabilities the distribution over your options; sums to 1 by construction
p1, p2, margin top-1, runner-up, and p1 - p2 β€” the margin is the product
band low < 0.5 ≀ med < 0.75 ≀ high < 0.9 ≀ certain β€” the portable form of p1
coverage the option mass before renormalisation; low β‡’ the model wanted to answer something else
degraded true when no option mass was seen β€” a named failure, never a silent uniform
undecided true when margin < theta β€” nothing is written
letter the raw letter, kept for indexing your own options

Two properties worth stating plainly, because they are the whole point:

  • degraded is not a verdict. Without option mass the model answered outside the menu; a uniform has a margin of zero, and the library says so instead of picking the first letter.
  • p1 is not portable; band is. Moving engine changes 2.3–10.4 % of verdicts but changes the band ~0–0.1 % above 0.9. Store the band, or store the raw float with its provenance β€” never 0.9137 as if it were a precision the measurement does not carry.

Abstention β€” theta

theta is the margin under which Piko refuses to answer. It is the only tunable the model is not allowed to swallow: the abstention stays in your hands, and an UNDECIDED is an outcome a caller can price and route, not a failure to disguise.

r = piko.choice(state, question, options, theta=0.5)
# margin >= 0.5  ->  a verdict
# margin  < 0.5  ->  choice is None, undecided is True, nothing is written

An UNDECIDED question stays open and is re-asked when the state changes β€” the natural escalation route to System 2 on the same model.

Modalities β€” one decoder, four units

Media enters the same messages array as an OpenAI content part; the readout treats it as the state. Nothing about the menu changes.

modality content part how it is read
text a plain string per token
image {"type":"image_url","image_url":{"url":"data:image/jpeg;base64,…"}} per megapixel
video {"type":"video_url","video_url":{"url":…}} per second
audio {"type":"input_audio","input_audio":{"data":"<base64>","format":"wav"}} per second
state = [
    {"kind": "image", "url": data_uri("receipt.jpg"), "width": 1200, "height": 1600},
    {"text": "The attached photograph is the only evidence."},
]
r = piko.choice(state, "what dominates the picture?", ["NATURE", "PEOPLE", "TEXT", "OBJECT"])

The engine reads raw pixels and raw waveform. Audio is base64 (input_audio carries no URL field) β€” a caller that sends a bare URL gets a named refusal, not an answer computed from the text alone.

Pricing

One flat rate for every modality: $0.025 per million input tokens, applied to the unit that was actually read β€” text per token, image per megapixel, video and audio per second. Completions are free: the readout generates one token, so the output is a free line, never a hidden cost.

unit the engine reads input limit
text per token β€”
image 258 tokens / megapixel 179 MP (max_area)
video 80 tokens / second 30 s (max_duration)
audio 65 tokens / second 30 s (max_duration)

A 4 MP picture is ~1032 tokens; a 10 s clip ~800 tokens; a 5 s recording ~325 tokens. r.cost returns the line-item price of the state, per unit, in USD.

Self-host it

The reference engine is Gemma 4 12B quantised to W8A8-INT8 (compressed-tensors: weight and activation 8-bit, dynamic), served by vLLM.

./serve.sh                                  # serve on :8000, model name `piko`
PORT=8100 ./serve.sh                        # elsewhere
MODEL_ID=google/gemma-4-12B-it-qat-w4a16-ct ./serve.sh   # a different quantisation

Why W8A8 and not W4A16: a read leans on prefill, and the INT8 tensor cores give ~1.9Γ— the prefill of W4A16 on the same card (4 446 vs 2 388 tok/s at concurrency 1, RTX 3090) at the cost of a slower decode β€” which never matters, because the readout generates exactly one token. Measured, not assumed; the full table ships in the deployment notes.

export PIKO_BASE_URL=http://127.0.0.1:8000/v1
python examples/quickstart.py

Reproduce

Try it live first: the Playground Space runs the curated examples (text, image, video, audio, and real recorded cases with a gold answer) against the API β€” one token, one verdict, nothing generated.

The readout is deterministic: temperature: 0, no sampling, one token. Re-running a call returns the same distribution, so a decision can be recorded and replayed offline with no engine at all.

python -m piko decide --state @state.txt --question "event type" \
  --option ANNOUNCE --option MEET --option VOTE --theta 0.5 --json

Exit codes: 0 a verdict Β· 3 UNDECIDED Β· 1 an error. A shell can tell "it said ANNOUNCE" from "it said nothing".

Model & architecture

  • Backbone: Gemma 4 12B, a unified multimodal decoder β€” one transformer reads text, image, video and audio with no separate towers at the decision point.
  • Quantisation: W8A8-INT8 via compressed-tensors (dynamic per-channel activation scaling).
  • Readout: the next-token distribution of one step, restricted to the option letters and renormalised. The letter regime caps a menu at 26 (one letter = one token) β€” and that cap is a ranking, never a silent truncation.
  • Mode: thinking is forced off at the template level so the first generated token is the answer, not a channel marker.
  • Envelope: a context of up to 8 192 tokens per decision.

An honest note on the probabilities

The distribution is ordered β€” a higher p1 really is a stronger read β€” but the raw float is not calibrated to a frequency across every engine and every menu size, which is exactly why two things exist: the band (portable) and theta (yours). Read p1 as a margin to threshold, not as a guarantee; and when a decision matters and the margin is thin, escalate to System 2 on the same model.

The family

Piko is one readout of the System 1 / System 2 decision family. Related open work in the same space: JEV, Clef, Laya.

Links

Piko is a stateless relay hosted in the European Union, in Germany; prompt and completion content is not logged, nothing is retained for training, and API keys are stored only as hashes.

License

Apache-2.0. The client is stdlib-only Python; the reference engine is the public checkpoint named in serve.sh.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using Agnosia-Piko/piko 1