Ayaka large

Ayaka large is an open, MIT-licensed decision model for Jev-style structured
decisions. You give it a state and typed questions (noul / choice /
score), and it returns a calibrated probability for every candidate label.
It is a LoRA adapter plus a small decision head on
google/gemma-4-12B-it
(text stack only, base revision 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7).
It is trained only on license-clean data.
Naming. This is a Gemma 4 model. "Electra" is the project's earlier (pre-Gemma) model name. It survives in internal identifiers such as
electra-largeand inelectra_config.json, which is kept as a legacy alias ofayaka_config.json. Both files hold the same config.
Code, training pipeline and documentation: https://github.com/alice-noa-chan/ayaka
Ayaka family
JevBench public tiers. The single-pass p50 is on an A100, one request at a time.
| model | backbone | easy | original | hard | hard ECE | p50 |
|---|---|---|---|---|---|---|
| ayaka-large + worked-steps route | Gemma 4 12B | 48/48 | 72/72 | 88/111 | 0.090 | 0.29 s* |
| ayaka-large, single pass | Gemma 4 12B | 48/48 | 72/72 | 72/111 | 0.178 | 0.18 s |
| ayaka-base | Gemma 4 E4B | 48/48 | 68/72 | 61/111 | 0.223 | 0.18 s |
| ayaka-small | Gemma 4 E2B | 48/48 | 65/72 | 55/111 | 0.238 | 0.07 s |
* Measured through the HTTP server. Routed hard questions take about 10 s. Choose large for accuracy, and small or base when latency, memory or cost matter more.
Other Ayaka models
All share the same code, training data, prompt format, server and API. They differ in backbone size, LoRA rank (r64 for large, r32 for base and small) and training length (860 steps for large, 1,200 for base and small). The worked-steps route is released only for large.
- Ayaka base (Gemma 4 E4B, about 16 GB in bf16): the middle ground, single pass. It gets 48/48, 68/72 and 61/111 on the public tiers, with hard ECE 0.223 and p50 0.18 s on A100.
- Ayaka small (Gemma 4 E2B, about 10 GB in bf16): the lightest member, single pass. It gets 48/48, 65/72 and 55/111 on the public tiers, with p50 0.07 s on A100. Use it for the lowest latency, memory and cost.
Use (pinned)
pip install -c https://raw.githubusercontent.com/alice-noa-chan/ayaka/7727e9ec5cec062c05de34ad40ade6fc1b6866d3/constraints.txt \
"git+https://github.com/alice-noa-chan/ayaka@7727e9ec5cec062c05de34ad40ade6fc1b6866d3"
python -m ayaka.serve --ckpt alice-noa-chan/ayaka-large --revision <this repo revision> \
--reasoning --device cuda --port 8000
A Dockerfile with the exact base image, pinned by digest, is in the repository.
Verified environment:
- Python 3.11.11 and torch
2.8.0.dev20250319+cu128(imagerunpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04). - transformers 5.17.0, peft 0.21.0, accelerate 1.15.0, safetensors 0.8.0, tokenizers 0.23.2.
- NVIDIA driver 570 or newer. A 550-driver host cannot initialise CUDA 12.8 torch.
This repository holds only the adapter (bf16 safetensors, 525 MB), the
head (head.safetensors) and the config. The base model is fetched from
Google's repository at the pinned revision.
The server speaks TypeSafe's POST /v1/systemone format and returns a
probability for every label:
noul: P(true);choice: per-label probabilities;score: per-level probabilities and the expected score.
usage.input_tokens and usage.output_tokens are reported per request.
Worked-steps route (--reasoning)
A question is routed only when all three conditions hold:
- it asks about a quantity. Rubric level numbers such as "0: not helpful" do not count;
- the state holds at least three numbers;
- the single-pass top probability is at most 0.9.
A routed question takes two steps:
- The base model, with the adapter off, writes at most 384 greedy tokens of worked steps.
- The trained model re-reads the state with those notes.
The policy was selected on non-public data and then frozen. Without the flag, every question takes one forward pass.
Results (JevBench public tiers)
| easy | original | hard | |
|---|---|---|---|
| Ayaka large + worked-steps route | 48/48 | 72/72 | 88/111 |
| Ayaka large, single pass | 48/48 | 72/72 | 72/111 |
| Jev 1.13.0 (reference) | 48/48 | 71/72 | 81/111 |
- Hard-tier ECE: 0.152 (single pass) โ 0.090 (with the route).
- Independent procedural test (500 questions, never used for any choice): 348 โ 417 correct.
- Open-Jev test (300 ordinary decisions): 4% routed, 273 โ 274 correct.
These public numbers will likely overstate sealed-tier results (see the disclosure below).
Latency
A100 SXM 80GB, one request at a time.
| tier | single pass p50 / p95 | with --reasoning p50 / p95 |
|---|---|---|
| easy | 0.18 / 0.19 s | 0.29 / 0.30 s |
| original | 0.18 / 0.18 s | 0.29 / 0.30 s |
| hard | 0.19 / 0.80 s | 9.9 / 31.4 s |
The single-pass numbers were measured in-process. The --reasoning
numbers went through the HTTP server, which adds about 0.1 s. On H100 the
single-pass p50 is 0.15 s. Easy and original questions are never routed,
so their tail stays flat. Only routed hard questions pay for generation.
Tokens and cost per decision
| cohort | mean input tokens | mean output tokens | $ / 1,000 decisions* |
|---|---|---|---|
| Open-Jev test (300) | 348 | 10 (236 when routed, 4% of questions) | 0.017 |
| public original (72) | 146 | 0 | 0.007 |
| public hard (111) | 1,242 | 160 (312 when routed, 51% of questions) | 0.062 |
* Priced at the JevBench leaderboard's 12B reference: $0.05 per million input tokens, output not billed. With output at $0.10 per million, the figures are 0.018, 0.007 and 0.078.
Training
Adapter: LoRA r 64, alpha 128, dropout 0.05, on q/k/v/o and gate/up/down projections.
Schedule: 860 steps ร 64 questions (LR 1e-4, head 5e-4, seed 0) on one H100.
Calibration: temperatures fitted per question type and prompt length (short/long at 1,024 tokens) on held-out data:
type short long noul 0.838 1.409 choice 1.270 1.401 score 1.088 1.187 Adapter storage: bf16. Parity against the fp32 training copy gives the same accuracy, with a median probability difference of 0.001.
Training data (35 sources)
No commercial-LLM outputs and no Jev API labels were used. Every source was checked against the JevBench public items for shared 13-grams, and 0 samples matched.
| source id | dataset | licence |
|---|---|---|
| jev_open | SargeDev/jev-distill-corpus-v3 | CC0-1.0 (Open-Jev release-v2-redistributable stream) |
| open_jev_bde | ZefanCai/Open-Jev/browser-drone-expansion-v1-redistributable | CC0-1.0 (Open-Jev redistributable config) |
| vitaminc | tals/vitaminc | CC BY-SA 3.0 |
| helpsteer2 | nvidia/HelpSteer2 | CC BY 4.0 |
| hh_rlhf | Anthropic/hh-rlhf | MIT |
| aegis_safety | nvidia/Aegis-AI-Content-Safety-Dataset-2.0 | CC BY 4.0 |
| aqua_rat | deepmind/aqua_rat/raw | Apache-2.0 |
| hotpot_decisions | hotpotqa/hotpot_qa/distractor | CC BY-SA 4.0 |
| squad_v2_answerable | rajpurkar/squad_v2 | CC BY-SA 4.0 |
| strategyqa | ChilleD/StrategyQA | MIT |
| arc_challenge | allenai/ai2_arc/ARC-Challenge | CC BY-SA 4.0 |
| commonsense_qa | tau/commonsense_qa | MIT |
| legalbench | nguha/legalbench | CC BY 4.0 (only tasks whose README states CC BY 4.0) |
| synth_temporal | generated: ayaka/data/synthetic.py | generated by this repository (MIT) |
| synth_numeric | generated: ayaka/data/synthetic.py | generated by this repository (MIT) |
| synth_policy | generated: ayaka/data/synthetic.py | generated by this repository (MIT) |
| synth_long_rules | generated: ayaka/data/hard_synthetic.py | generated by this repository (MIT) |
| synth_calendar | generated: ayaka/data/hard_synthetic.py | generated by this repository (MIT) |
| synth_probability | generated: ayaka/data/hard_synthetic.py | generated by this repository (MIT) |
| snli | stanfordnlp/snli | CC BY-SA 3.0 |
| multi_nli | nyu-mll/multi_nli | mixed by genre: OANC, CC BY 3.0, CC BY-SA 3.0 (see card) |
| boolq | google/boolq | CC BY-SA 3.0 |
| banking77 | mteb/banking77/default | CC BY 4.0 |
| clinc_oos | clinc/clinc_oos/plus | CC BY 3.0 |
| klue_nli | klue/klue/nli | CC BY-SA 4.0 |
| klue_ynat | klue/klue/ynat | CC BY-SA 4.0 |
| kor_nli_multi | kakaobrain/kor_nli/multi_nli | CC BY-SA 4.0 |
| massive_ja | AmazonScience/massive/ja-JP | CC BY 4.0 |
| massive_ko | AmazonScience/massive/ko-KR | CC BY 4.0 |
| jglue_jnli | shunk031/JGLUE/JNLI | CC BY-SA 4.0 (JGLUE) |
| jglue_jsts | shunk031/JGLUE/JSTS | CC BY-SA 4.0 (JGLUE) |
| jglue_commonsense | shunk031/JGLUE/JCommonsenseQA | CC BY-SA 4.0 (JGLUE) |
| go_emotions | google-research-datasets/go_emotions/simplified | Apache 2.0 |
| quality | emozilla/quality | CC BY 4.0 (QuALITY) |
| quality_dev | emozilla/quality | CC BY 4.0 (QuALITY) |
CC BY-SA and CC BY sources were used as training data only. Their licence and attribution are listed here. Whether ShareAlike terms extend to model weights is not settled, and we do not claim otherwise.
Benchmark disclosure
No JevBench items or labels were included in the training corpus. All 35 training sources were checked against the public items for shared 13-grams, and 0 samples matched. However, public JevBench results were used as a development signal. They helped identify weak task families and informed synthetic-data design, model-size selection, and the decision to develop a reasoning path. Public-item outputs were also read to diagnose failures (a zero-shot calculation-plan format, a calibration bug). One prompt example that structurally mirrored a public item was removed before the reported runs.
Not selected on public results:
- Reasoning-route threshold, gating and fusion. This covers the calculation gate rules, the confidence cutoff of 0.9 and the fusion weight. They were selected on a procedural dev set (500 questions) and Open-Jev test (300 questions) only. The policy was frozen and hash-recorded before any public scoring.
- Worked-steps writer. Using the base model with the adapter off, rather than with the adapter on, was decided on the dev set.
- Checkpoint. Each model is the final checkpoint of its run. No intermediate checkpoint was compared on public items.
- Training length and hyperparameters. Steps were set by budget (860 for large, 1,200 for small and base). Learning rate, LoRA rank and the data-mixture quotas were fixed before training and not re-tuned on public scores. The synthetic sources themselves were designed as described above.
- Calibration temperatures. They were fitted on held-out training-distribution data.
- Freeze. The released models are frozen for this submission, and no further changes will be made in response to public results.
License
MIT for code and fine-tuned weights. The base model
google/gemma-4-12B-it is Apache-2.0, and its notices apply to the base
portion.