Ayaka large

Ayaka

Ayaka large is an open, MIT-licensed decision model for Jev-style structured decisions. You give it a state and typed questions (noul / choice / score), and it returns a calibrated probability for every candidate label.

It is a LoRA adapter plus a small decision head on google/gemma-4-12B-it (text stack only, base revision 707f0a3b8a3c7ad586ed01e27eafbad8a27dd0f7). It is trained only on license-clean data.

Naming. This is a Gemma 4 model. "Electra" is the project's earlier (pre-Gemma) model name. It survives in internal identifiers such as electra-large and in electra_config.json, which is kept as a legacy alias of ayaka_config.json. Both files hold the same config.

Code, training pipeline and documentation: https://github.com/alice-noa-chan/ayaka

Ayaka family

JevBench public tiers. The single-pass p50 is on an A100, one request at a time.

model backbone easy original hard hard ECE p50
ayaka-large + worked-steps route Gemma 4 12B 48/48 72/72 88/111 0.090 0.29 s*
ayaka-large, single pass Gemma 4 12B 48/48 72/72 72/111 0.178 0.18 s
ayaka-base Gemma 4 E4B 48/48 68/72 61/111 0.223 0.18 s
ayaka-small Gemma 4 E2B 48/48 65/72 55/111 0.238 0.07 s

* Measured through the HTTP server. Routed hard questions take about 10 s. Choose large for accuracy, and small or base when latency, memory or cost matter more.

Other Ayaka models

All share the same code, training data, prompt format, server and API. They differ in backbone size, LoRA rank (r64 for large, r32 for base and small) and training length (860 steps for large, 1,200 for base and small). The worked-steps route is released only for large.

  • Ayaka base (Gemma 4 E4B, about 16 GB in bf16): the middle ground, single pass. It gets 48/48, 68/72 and 61/111 on the public tiers, with hard ECE 0.223 and p50 0.18 s on A100.
  • Ayaka small (Gemma 4 E2B, about 10 GB in bf16): the lightest member, single pass. It gets 48/48, 65/72 and 55/111 on the public tiers, with p50 0.07 s on A100. Use it for the lowest latency, memory and cost.

Use (pinned)

pip install -c https://raw.githubusercontent.com/alice-noa-chan/ayaka/7727e9ec5cec062c05de34ad40ade6fc1b6866d3/constraints.txt \
  "git+https://github.com/alice-noa-chan/ayaka@7727e9ec5cec062c05de34ad40ade6fc1b6866d3"
python -m ayaka.serve --ckpt alice-noa-chan/ayaka-large --revision <this repo revision> \
  --reasoning --device cuda --port 8000

A Dockerfile with the exact base image, pinned by digest, is in the repository.

Verified environment:

  • Python 3.11.11 and torch 2.8.0.dev20250319+cu128 (image runpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04).
  • transformers 5.17.0, peft 0.21.0, accelerate 1.15.0, safetensors 0.8.0, tokenizers 0.23.2.
  • NVIDIA driver 570 or newer. A 550-driver host cannot initialise CUDA 12.8 torch.

This repository holds only the adapter (bf16 safetensors, 525 MB), the head (head.safetensors) and the config. The base model is fetched from Google's repository at the pinned revision.

The server speaks TypeSafe's POST /v1/systemone format and returns a probability for every label:

  • noul: P(true);
  • choice: per-label probabilities;
  • score: per-level probabilities and the expected score.

usage.input_tokens and usage.output_tokens are reported per request.

Worked-steps route (--reasoning)

A question is routed only when all three conditions hold:

  1. it asks about a quantity. Rubric level numbers such as "0: not helpful" do not count;
  2. the state holds at least three numbers;
  3. the single-pass top probability is at most 0.9.

A routed question takes two steps:

  1. The base model, with the adapter off, writes at most 384 greedy tokens of worked steps.
  2. The trained model re-reads the state with those notes.

The policy was selected on non-public data and then frozen. Without the flag, every question takes one forward pass.

Results (JevBench public tiers)

easy original hard
Ayaka large + worked-steps route 48/48 72/72 88/111
Ayaka large, single pass 48/48 72/72 72/111
Jev 1.13.0 (reference) 48/48 71/72 81/111
  • Hard-tier ECE: 0.152 (single pass) โ†’ 0.090 (with the route).
  • Independent procedural test (500 questions, never used for any choice): 348 โ†’ 417 correct.
  • Open-Jev test (300 ordinary decisions): 4% routed, 273 โ†’ 274 correct.

These public numbers will likely overstate sealed-tier results (see the disclosure below).

Latency

A100 SXM 80GB, one request at a time.

tier single pass p50 / p95 with --reasoning p50 / p95
easy 0.18 / 0.19 s 0.29 / 0.30 s
original 0.18 / 0.18 s 0.29 / 0.30 s
hard 0.19 / 0.80 s 9.9 / 31.4 s

The single-pass numbers were measured in-process. The --reasoning numbers went through the HTTP server, which adds about 0.1 s. On H100 the single-pass p50 is 0.15 s. Easy and original questions are never routed, so their tail stays flat. Only routed hard questions pay for generation.

Tokens and cost per decision

cohort mean input tokens mean output tokens $ / 1,000 decisions*
Open-Jev test (300) 348 10 (236 when routed, 4% of questions) 0.017
public original (72) 146 0 0.007
public hard (111) 1,242 160 (312 when routed, 51% of questions) 0.062

* Priced at the JevBench leaderboard's 12B reference: $0.05 per million input tokens, output not billed. With output at $0.10 per million, the figures are 0.018, 0.007 and 0.078.

Training

  • Adapter: LoRA r 64, alpha 128, dropout 0.05, on q/k/v/o and gate/up/down projections.

  • Schedule: 860 steps ร— 64 questions (LR 1e-4, head 5e-4, seed 0) on one H100.

  • Calibration: temperatures fitted per question type and prompt length (short/long at 1,024 tokens) on held-out data:

    type short long
    noul 0.838 1.409
    choice 1.270 1.401
    score 1.088 1.187
  • Adapter storage: bf16. Parity against the fp32 training copy gives the same accuracy, with a median probability difference of 0.001.

Training data (35 sources)

No commercial-LLM outputs and no Jev API labels were used. Every source was checked against the JevBench public items for shared 13-grams, and 0 samples matched.

source id dataset licence
jev_open SargeDev/jev-distill-corpus-v3 CC0-1.0 (Open-Jev release-v2-redistributable stream)
open_jev_bde ZefanCai/Open-Jev/browser-drone-expansion-v1-redistributable CC0-1.0 (Open-Jev redistributable config)
vitaminc tals/vitaminc CC BY-SA 3.0
helpsteer2 nvidia/HelpSteer2 CC BY 4.0
hh_rlhf Anthropic/hh-rlhf MIT
aegis_safety nvidia/Aegis-AI-Content-Safety-Dataset-2.0 CC BY 4.0
aqua_rat deepmind/aqua_rat/raw Apache-2.0
hotpot_decisions hotpotqa/hotpot_qa/distractor CC BY-SA 4.0
squad_v2_answerable rajpurkar/squad_v2 CC BY-SA 4.0
strategyqa ChilleD/StrategyQA MIT
arc_challenge allenai/ai2_arc/ARC-Challenge CC BY-SA 4.0
commonsense_qa tau/commonsense_qa MIT
legalbench nguha/legalbench CC BY 4.0 (only tasks whose README states CC BY 4.0)
synth_temporal generated: ayaka/data/synthetic.py generated by this repository (MIT)
synth_numeric generated: ayaka/data/synthetic.py generated by this repository (MIT)
synth_policy generated: ayaka/data/synthetic.py generated by this repository (MIT)
synth_long_rules generated: ayaka/data/hard_synthetic.py generated by this repository (MIT)
synth_calendar generated: ayaka/data/hard_synthetic.py generated by this repository (MIT)
synth_probability generated: ayaka/data/hard_synthetic.py generated by this repository (MIT)
snli stanfordnlp/snli CC BY-SA 3.0
multi_nli nyu-mll/multi_nli mixed by genre: OANC, CC BY 3.0, CC BY-SA 3.0 (see card)
boolq google/boolq CC BY-SA 3.0
banking77 mteb/banking77/default CC BY 4.0
clinc_oos clinc/clinc_oos/plus CC BY 3.0
klue_nli klue/klue/nli CC BY-SA 4.0
klue_ynat klue/klue/ynat CC BY-SA 4.0
kor_nli_multi kakaobrain/kor_nli/multi_nli CC BY-SA 4.0
massive_ja AmazonScience/massive/ja-JP CC BY 4.0
massive_ko AmazonScience/massive/ko-KR CC BY 4.0
jglue_jnli shunk031/JGLUE/JNLI CC BY-SA 4.0 (JGLUE)
jglue_jsts shunk031/JGLUE/JSTS CC BY-SA 4.0 (JGLUE)
jglue_commonsense shunk031/JGLUE/JCommonsenseQA CC BY-SA 4.0 (JGLUE)
go_emotions google-research-datasets/go_emotions/simplified Apache 2.0
quality emozilla/quality CC BY 4.0 (QuALITY)
quality_dev emozilla/quality CC BY 4.0 (QuALITY)

CC BY-SA and CC BY sources were used as training data only. Their licence and attribution are listed here. Whether ShareAlike terms extend to model weights is not settled, and we do not claim otherwise.

Benchmark disclosure

No JevBench items or labels were included in the training corpus. All 35 training sources were checked against the public items for shared 13-grams, and 0 samples matched. However, public JevBench results were used as a development signal. They helped identify weak task families and informed synthetic-data design, model-size selection, and the decision to develop a reasoning path. Public-item outputs were also read to diagnose failures (a zero-shot calculation-plan format, a calibration bug). One prompt example that structurally mirrored a public item was removed before the reported runs.

Not selected on public results:

  • Reasoning-route threshold, gating and fusion. This covers the calculation gate rules, the confidence cutoff of 0.9 and the fusion weight. They were selected on a procedural dev set (500 questions) and Open-Jev test (300 questions) only. The policy was frozen and hash-recorded before any public scoring.
  • Worked-steps writer. Using the base model with the adapter off, rather than with the adapter on, was decided on the dev set.
  • Checkpoint. Each model is the final checkpoint of its run. No intermediate checkpoint was compared on public items.
  • Training length and hyperparameters. Steps were set by budget (860 for large, 1,200 for small and base). Learning rate, LoRA rank and the data-mixture quotas were fixed before training and not re-tuned on public scores. The synthetic sources themselves were designed as described above.
  • Calibration temperatures. They were fitted on held-out training-distribution data.
  • Freeze. The released models are frozen for this submission, and no further changes will be made in response to public results.

License

MIT for code and fine-tuned weights. The base model google/gemma-4-12B-it is Apache-2.0, and its notices apply to the base portion.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for alice-noa-chan/ayaka-large

Adapter
(97)
this model