Ayaka base

Ayaka base is the mid-size member of the open,
MIT-licensed Ayaka decision-model family. You give it a state and typed
questions (noul / choice / score), and it returns a calibrated
probability for every candidate label in one forward pass.
It is a LoRA adapter plus a small decision head on
google/gemma-4-E4B-it (text stack only, base
revision ee0ef6023621cff504d758262d4e04895a5af4a2). It is trained only on license-clean data. Use it
when speed, memory or cost matter more than hard-tier accuracy.
Ayaka large is the
accurate option.
Naming. This is a Gemma 4 model. "Electra" is the project's earlier model name. It survives in internal identifiers (
electra-base) and inelectra_config.json, a legacy alias ofayaka_config.jsonwith the same content.
Code and documentation: https://github.com/alice-noa-chan/ayaka
Ayaka family
JevBench public tiers. The single-pass p50 is on an A100, one request at a time.
| model | backbone | easy | original | hard | hard ECE | p50 |
|---|---|---|---|---|---|---|
| ayaka-large + worked-steps route | Gemma 4 12B | 48/48 | 72/72 | 88/111 | 0.090 | 0.29 s* |
| ayaka-large, single pass | Gemma 4 12B | 48/48 | 72/72 | 72/111 | 0.178 | 0.18 s |
| ayaka-base | Gemma 4 E4B | 48/48 | 68/72 | 61/111 | 0.223 | 0.18 s |
| ayaka-small | Gemma 4 E2B | 48/48 | 65/72 | 55/111 | 0.238 | 0.07 s |
* Measured through the HTTP server. Routed hard questions take about 10 s. Choose large for accuracy, and small or base when latency, memory or cost matter more.
Other Ayaka models
All share the same code, training data, prompt format, server and API. They differ in backbone size, LoRA rank (r64 for large, r32 for base and small) and training length (860 steps for large, 1,200 for base and small). The worked-steps route is released only for large.
- Ayaka large (Gemma 4 12B, about 24 GB in bf16): the
accurate member. It gets 72/111 on the public hard tier in one pass and
88/111 with the optional worked-steps route (
--reasoning), which writes short calculation notes only for low-confidence quantitative questions. Single-pass p50 is 0.18 s on A100. - Ayaka small (Gemma 4 E2B, about 10 GB in bf16): the lightest member, single pass. It gets 48/48, 65/72 and 55/111 on the public tiers, with p50 0.07 s on A100. Use it for the lowest latency, memory and cost.
Use (pinned)
pip install -c https://raw.githubusercontent.com/alice-noa-chan/ayaka/042d6869e7167976e921662bac2a70a1f50295aa/constraints.txt \
"git+https://github.com/alice-noa-chan/ayaka@042d6869e7167976e921662bac2a70a1f50295aa"
python -m ayaka.serve --ckpt alice-noa-chan/ayaka-base --revision <this repo revision> \
--device cuda --port 8000
This model is released as single pass. The worked-steps route
(--reasoning) was measured and frozen only for Ayaka large, and it is
unevaluated for this size.
Verified environment:
- Python 3.11.11 and torch
2.8.0.dev20250319+cu128(runpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04). - transformers 5.17.0, peft 0.21.0, accelerate 1.15.0.
- NVIDIA driver 570 or newer.
A Dockerfile pinned by digest is in the repository.
The server speaks TypeSafe's POST /v1/systemone format and returns a
probability for every label:
noul: P(true);choice: per-label probabilities;score: per-level probabilities and the expected score.
Results (JevBench public tiers)
| easy | original | hard | |
|---|---|---|---|
| Ayaka base (single pass) | 48/48 | 68/72 | 61/111 |
| Ayaka large (single pass) | 48/48 | 72/72 | 72/111 |
| Ayaka large + worked-steps route | 48/48 | 72/72 | 88/111 |
- Hard-tier ECE: 0.223.
- Held-out mixture: accuracy 0.891, ECE 0.024.
- Open-Jev test (1000 questions): accuracy 0.912, ECE 0.019.
Public numbers will likely overstate sealed-tier results (see the disclosure below).
Latency and cost
A100 SXM 80GB, one request at a time, in-process.
| tier | p50 | p95 |
|---|---|---|
| easy | 0.194 s | 0.201 s |
| original | 0.180 s | 0.196 s |
| hard | 0.190 s | 0.401 s |
The model generates no tokens. Input tokens per decision are the same as for large (the prompt format is shared):
| cohort | mean input tokens | $ / 1,000 decisions* |
|---|---|---|
| Open-Jev test | 348 | 0.0070 |
| public original | 146 | 0.0029 |
| public hard | 1,242 | 0.0248 |
* Priced at the leaderboard's reference for this size class: $0.02 per million input tokens.
Training
Adapter: LoRA r 32, alpha 64, dropout 0.05, on q/k/v/o and gate/up/down projections. Stored as bf16 safetensors.
Schedule: 1200 steps × 64 questions (LR 0.0001, head 0.0005, seed 0) on one A100 80GB, 3.1 h.
Calibration: temperatures fitted per question type and prompt length on held-out data. Long means ≥ 1,024 tokens.
type short long noul 0.809 1.758 choice 1.158 1.086 score 1.189 1.048 Training: direct, on the same data and code as large. No distillation.
Training data (35 sources)
No commercial-LLM outputs and no Jev API labels were used. Every source was checked against the JevBench public items for shared 13-grams, and 0 samples matched.
| source id | dataset | licence |
|---|---|---|
| jev_open | SargeDev/jev-distill-corpus-v3 | CC0-1.0 (Open-Jev release-v2-redistributable stream) |
| open_jev_bde | ZefanCai/Open-Jev/browser-drone-expansion-v1-redistributable | CC0-1.0 (Open-Jev redistributable config) |
| vitaminc | tals/vitaminc | CC BY-SA 3.0 |
| helpsteer2 | nvidia/HelpSteer2 | CC BY 4.0 |
| hh_rlhf | Anthropic/hh-rlhf | MIT |
| aegis_safety | nvidia/Aegis-AI-Content-Safety-Dataset-2.0 | CC BY 4.0 |
| aqua_rat | deepmind/aqua_rat/raw | Apache-2.0 |
| hotpot_decisions | hotpotqa/hotpot_qa/distractor | CC BY-SA 4.0 |
| squad_v2_answerable | rajpurkar/squad_v2 | CC BY-SA 4.0 |
| strategyqa | ChilleD/StrategyQA | MIT |
| arc_challenge | allenai/ai2_arc/ARC-Challenge | CC BY-SA 4.0 |
| commonsense_qa | tau/commonsense_qa | MIT |
| legalbench | nguha/legalbench | CC BY 4.0 (only tasks whose README states CC BY 4.0) |
| synth_temporal | generated: ayaka/data/synthetic.py | generated by this repository (MIT) |
| synth_numeric | generated: ayaka/data/synthetic.py | generated by this repository (MIT) |
| synth_policy | generated: ayaka/data/synthetic.py | generated by this repository (MIT) |
| synth_long_rules | generated: ayaka/data/hard_synthetic.py | generated by this repository (MIT) |
| synth_calendar | generated: ayaka/data/hard_synthetic.py | generated by this repository (MIT) |
| synth_probability | generated: ayaka/data/hard_synthetic.py | generated by this repository (MIT) |
| snli | stanfordnlp/snli | CC BY-SA 3.0 |
| multi_nli | nyu-mll/multi_nli | mixed by genre: OANC, CC BY 3.0, CC BY-SA 3.0 (see card) |
| boolq | google/boolq | CC BY-SA 3.0 |
| banking77 | mteb/banking77/default | CC BY 4.0 |
| clinc_oos | clinc/clinc_oos/plus | CC BY 3.0 |
| klue_nli | klue/klue/nli | CC BY-SA 4.0 |
| klue_ynat | klue/klue/ynat | CC BY-SA 4.0 |
| kor_nli_multi | kakaobrain/kor_nli/multi_nli | CC BY-SA 4.0 |
| massive_ja | AmazonScience/massive/ja-JP | CC BY 4.0 |
| massive_ko | AmazonScience/massive/ko-KR | CC BY 4.0 |
| jglue_jnli | shunk031/JGLUE/JNLI | CC BY-SA 4.0 (JGLUE) |
| jglue_jsts | shunk031/JGLUE/JSTS | CC BY-SA 4.0 (JGLUE) |
| jglue_commonsense | shunk031/JGLUE/JCommonsenseQA | CC BY-SA 4.0 (JGLUE) |
| go_emotions | google-research-datasets/go_emotions/simplified | Apache 2.0 |
| quality | emozilla/quality | CC BY 4.0 (QuALITY) |
| quality_dev | emozilla/quality | CC BY 4.0 (QuALITY) |
CC BY-SA and CC BY sources were used as training data only. Their licence and attribution are listed here. Whether ShareAlike terms extend to model weights is not settled, and we do not claim otherwise.
Benchmark disclosure
No JevBench items or labels were included in the training corpus. All 35 training sources were checked against the public items for shared 13-grams, and 0 samples matched. However, public JevBench results were used as a development signal. They helped identify weak task families and informed synthetic-data design, model-size selection, and the decision to develop a reasoning path. Public-item outputs were also read to diagnose failures (a zero-shot calculation-plan format, a calibration bug). One prompt example that structurally mirrored a public item was removed before the reported runs.
Not selected on public results:
- Reasoning-route threshold, gating and fusion. This covers the calculation gate rules, the confidence cutoff of 0.9 and the fusion weight. They were selected on a procedural dev set (500 questions) and Open-Jev test (300 questions) only. The policy was frozen and hash-recorded before any public scoring.
- Worked-steps writer. Using the base model with the adapter off, rather than with the adapter on, was decided on the dev set.
- Checkpoint. Each model is the final checkpoint of its run. No intermediate checkpoint was compared on public items.
- Training length and hyperparameters. Steps were set by budget (860 for large, 1,200 for small and base). Learning rate, LoRA rank and the data-mixture quotas were fixed before training and not re-tuned on public scores. The synthetic sources themselves were designed as described above.
- Calibration temperatures. They were fitted on held-out training-distribution data.
- Freeze. The released models are frozen for this submission, and no further changes will be made in response to public results.
License
MIT for code and fine-tuned weights. The base model
google/gemma-4-E4B-it is Apache-2.0, and its notices apply to the base
portion.