Payo-0.8B-Core-Preview

A small specialized model for keeping structured state up to date.

Given a schema, the current structured state, and a natural-language observation, Payo emits only the smallest state change justified by that observation.

Instead of regenerating the full state, it produces a compact, executable patch — or [] when no change to the current state is justified.

No-op is a first-class output, not a fallback.

This release is the frozen RR08 LoRA adapter for Qwen/Qwen3.5-0.8B-Base.

Research preview. Payo proposes state changes; it is not an autonomous state-management system. It can still make false mutations, and its patches should be validated deterministically or reviewed before being applied to important state. See Limitations.

Example

Current state:

{
  "status": "scheduled"
}

Schema:

/status: enum[scheduled, live, finished]

For the observation:

The match is underway.

Payo emits:

[
  {
    "op": "set",
    "path": "/status",
    "value": "live"
  }
]

For the observation:

The match starts tomorrow.

the correct output is:

[]

The second observation describes a future event, not a change that has already happened to the current state.

Quickstart

The reference example was smoke-tested with:

  • Python 3.13.15
  • PyTorch 2.11.0
  • Transformers 5.16.1
  • PEFT 0.20.0

In an activated virtual environment:

python3 -m pip install \
  torch==2.11.0 \
  transformers==5.16.1 \
  peft==0.20.0

Run the baseline example:

python3 examples/inference_baseline.py

By default, the example loads the adapter from:

paaya/Payo-0.8B-Core-Preview

To use a local adapter directory instead:

python3 examples/inference_baseline.py \
  --adapter ./adapter-dir

The example loads the Qwen base model at the exact frozen revision, applies this LoRA adapter, and decodes greedily.

It prints the raw generated output. It intentionally does not repair, normalize, or coerce generations, so validate or review each patch before applying it to any state.

Model Details

Model type LoRA adapter (PEFT) for causal language modeling
Base model Qwen/Qwen3.5-0.8B-Base
Base revision dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68
Tokenizer revision dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68
Checkpoint RR08 (frozen)
License Apache-2.0

This repository contains the LoRA adapter only, not the Qwen base-model weights.

Inputs and outputs

The model expects four inputs:

  1. a Payo Schema View (PSV) describing writable JSON Pointer paths, types, constraints, and allowed operations,
  2. the current JSON state,
  3. an optional context object,
  4. and a new natural-language observation.

The reference inference example demonstrates the serialization expected by this checkpoint.

Payo outputs a JSON array containing the smallest justified set of Payo Patch IR operations.

Payo Patch IR is not RFC 6902 JSON Patch. See docs/PATCH_IR.md for the complete compact format and application rules.

set

Creates or replaces a value at a schema-authorized path.

{
  "op": "set",
  "path": "/status",
  "value": "live"
}

delete

Removes a field where deletion is allowed.

{
  "op": "delete",
  "path": "/temporary_note"
}

append

Adds one item to an append-only array.

{
  "op": "append",
  "path": "/events",
  "value": {
    "type": "goal",
    "player": "Benjamin"
  }
}

No-op

Returned when the observation does not justify any change to the current state.

[]

null, delete, and no-op are different

Output Meaning
[{"op": "set", "path": "/foo", "value": null}] The field exists and its value is null.
[{"op": "delete", "path": "/foo"}] The field is removed.
[] No state change is justified.

Open-world state updates

Values written by Payo do not need to come from a predefined choice list. For example, in a support-ticket workflow:

Current state

{
  "assignee": null
}

Schema /assignee: string | null

Observation

Assign this ticket to Maya Chen.

Expected Payo output

[
  {
    "op": "set",
    "path": "/assignee",
    "value": "Maya Chen"
  }
]

“Maya Chen” need not appear in a predefined choice list. Payo can extract an open-world value from the observation and write it into state. The operation type, schema-authorized path, and enum values where applicable remain constrained.

Typed decision models typically answer predefined questions over constrained choices. Payo instead determines whether the observation justifies a state change, which schema path is affected, and what value should be written. Some parts of the output remain closed-world, while strings, numbers, and structured objects may be open-world.

Intended Use

This release is intended for:

  • research on sparse state updates,
  • schema-conditioned state mutation,
  • structured-output modeling,
  • state-tracking experiments,
  • and evaluation of small specialized language models.

It may also serve as a baseline for systems that wrap model-generated state changes in additional validation, routing, abstention, or human review.

Out-of-scope use

This preview should not be relied on for:

  • autonomous modification of important production state,
  • safety-critical state tracking,
  • unattended database mutation,
  • high-confidence long-horizon tracking,
  • or any application where an incorrect state mutation could cause significant harm.

When experimenting with the model, use deterministic validation and application-specific safeguards.

Limitations

The model can still produce false mutations. Known weaknesses include:

  • relative arithmetic,
  • binding an observation to the correct same-type entity or record,
  • mapping everyday language to canonical schema values,
  • composing several related field changes into one complete transition,
  • noisy or indirect natural language,
  • and long-horizon or free-running state tracking.

Future plans, past-only events, hypotheticals, questions, hedges, and observations about another entity should generally not modify the current state unless the observation clearly establishes that the tracked state has changed.

The results below do not establish production readiness or safety for unattended updates.

Optional validation

A deterministic, fail-closed integration layer is documented in docs/OPTIONAL_VALIDATION.md.

It can reject outputs that violate structural or schema constraints, such as malformed patches or invalid paths.

A rejected output is treated as an abstention and is never converted to [].

This validator is a safety and integration mechanism. It does not repair semantic mistakes or improve the model's underlying language understanding.

Evaluation

Payo is evaluated by applying the predicted patch to the current state and comparing the resulting state exactly against the gold state.

FSEM (Final State Exact Match) is the fraction of examples whose predicted patch produces exactly the expected final state.

Invalid generations count as failures. Raw evaluation does not retry, repair, or silently coerce model output.

Evaluation Examples RR08 baseline FSEM Additional results Interpretation
Untouched capability-structured synthetic Preview Gate 312 309/312 (99.04%) NOOP 185/186; false mutations 3/312; mutation precision 98.30%, recall 98.86%, F1 98.58% Controlled evaluation across 12 predeclared semantic categories. Not an estimate of real-world accuracy.
Frozen realistic holdout 200 129/200 (64.5%) Six gold cases were flagged as ambiguous during the original audit. All 200 examples remain in the reported denominator. Hand-authored across 10 domains; considerably more realistic in language and harder than the controlled synthetic set.

Why the two scores differ

The 99.04% synthetic score should not be read as real-world reliability.

The synthetic Preview Gate is capability-structured: it deliberately covers well-defined semantic cases under controlled generation.

The realistic holdout contains more varied, indirect, and natural language, and it exposes much harder state-tracking behavior.

Evaluation protocol notes

  • The synthetic Preview Gate was frozen before RR08 inference, and its category distribution was fixed before any outputs were generated.
  • The realistic holdout was frozen before its original evaluation and used once for the final RR08 comparison. It is therefore no longer an untouched holdout and should not be reused for future model-selection claims.
  • Future Payo model comparisons will require a new, independent evaluation set.

Prompt comparison

The recommended research default is the baseline prompt in examples/inference_baseline.py.

A fixed four-shot prompt was also tested on the untouched 312-example synthetic Preview Gate:

Prompt FSEM
Baseline 309/312
Fixed four-shot 303/312

The four-shot prompt also used:

  • 4.01× the mean input tokens,
  • and 3.55× the median MPS wall-clock latency.

It is therefore not the recommended default for this release.

The improvement observed on a small development probe did not reproduce on untouched evaluation data.

Training

RR08 is a LoRA fine-tune of Qwen/Qwen3.5-0.8B-Base.

Adapter configuration

  • Trainable parameters: 20,447,232
  • LoRA rank: 32
  • LoRA alpha: 64
  • LoRA dropout: 0.05
  • Training precision: bfloat16
  • Epochs: 2
  • Optimizer steps: 762
  • Seed: 2026092806

Training data

The frozen training input contains:

Split Rows
Train 6,086
Validation 465
Test 460
Total 7,011

Training used the 6,086-row train split. Each example pairs:

  • a schema,
  • the current structured state,
  • a natural-language observation,
  • and a deterministic minimal-patch target.

The corpus spans 12 application domains.

All examples are synthetic and contain no real customer records.

The training corpus is not redistributed in this adapter repository.

Realism data

RR08 incorporates a frozen realism-v0.2 source artifact containing 2,400 examples.

Its run configuration records:

  • DeepSeek V4.1 Flash as writer/auditor,
  • NVIDIA Nemotron 3 Ultra as verifier.

These models were used for model-assisted naturalization and verification around deterministically defined state transitions.

Provider and model identifiers, artifact hashes, and source links are recorded in PROVENANCE.json.

The generated source corpus itself is not redistributed here.

Reproducibility

The frozen evaluation used:

  • bfloat16,
  • greedy decoding,
  • batch size 1,
  • max_new_tokens=768,
  • and no reusable prefix KV cache.

The reference MPS run used:

  • Apple M3,
  • PyTorch 2.11.0,
  • Transformers 5.16.1,
  • PEFT 0.20.0.

The Preview Gate evaluator was version 0.1.0 at source commit:

d8458a1e6d13f8c351dfcc528672808e9a742a79

The realistic holdout is identified as:

realuse-holdout-200-v1

Exact provenance for the model, adapter, prompt, training input, evaluator, datasets, predictions, and comparisons is recorded in PROVENANCE.json.

Frozen artifact hashes (SHA-256)
Artifact SHA-256
Base weights c2b1e5a17d9c1e27685d92ed9b382911ebb99955ecd89052d1721241adfbab6c
Base config b90b86f35c8e6925ef74ee04d0e758f0a845c83a42089ad82bbaa948de9b4204
Adapter weights 3ab32a97248379cbb058e12c503df935b36f11fdb484b3b6c37fe83041aba9d6
Published adapter config 7c6153953b8fa62a44d90cddcfcab4c9ec135a7e05cf053be030b0a07d3531ae
Baseline static prompt prefix d4dd0dd89022a46bd4a713bfb163ca889587a752807ff11e8f16de7cb33863a7
Full frozen training input 08476a2006a003176e326d9b00e4ef2bce44b5d8e0352664e3ea21a0f6cdc404
Train split only 5bf403204719bc1066629d7f3dc3fc612e23ca6a50a76f711f286137055f0348
Synthetic Preview Gate 1ecd290a192b9457d47aa5d413f680df8a52e6b757d0fae87f5394433f28f8f5
Realistic holdout 1b16090618b11c899b8d5907b4ded5bdce01df2000b355c4cf0cfc0016483107

License

This adapter is distributed under Apache-2.0.

The base model — Apache-2.0 at the revision above — remains subject to its upstream license and notices.

The upstream license file is included in this repository for reference.

Downloads last month
17
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for paaya/Payo-0.8B-Core-Preview

Adapter
(31)
this model

Collection including paaya/Payo-0.8B-Core-Preview