Ayaka base

Ayaka

Ayaka base is the mid-size member of the open, MIT-licensed Ayaka decision-model family. You give it a state and typed questions (noul / choice / score), and it returns a calibrated probability for every candidate label in one forward pass.

It is a LoRA adapter plus a small decision head on google/gemma-4-E4B-it (text stack only, base revision ee0ef6023621cff504d758262d4e04895a5af4a2). It is trained only on license-clean data. Use it when speed, memory or cost matter more than hard-tier accuracy. Ayaka large is the accurate option.

Naming. This is a Gemma 4 model. "Electra" is the project's earlier model name. It survives in internal identifiers (electra-base) and in electra_config.json, a legacy alias of ayaka_config.json with the same content.

Code and documentation: https://github.com/alice-noa-chan/ayaka

Ayaka family

JevBench public tiers. The single-pass p50 is on an A100, one request at a time.

model backbone easy original hard hard ECE p50
ayaka-large + worked-steps route Gemma 4 12B 48/48 72/72 88/111 0.090 0.29 s*
ayaka-large, single pass Gemma 4 12B 48/48 72/72 72/111 0.178 0.18 s
ayaka-base Gemma 4 E4B 48/48 68/72 61/111 0.223 0.18 s
ayaka-small Gemma 4 E2B 48/48 65/72 55/111 0.238 0.07 s

* Measured through the HTTP server. Routed hard questions take about 10 s. Choose large for accuracy, and small or base when latency, memory or cost matter more.

Other Ayaka models

All share the same code, training data, prompt format, server and API. They differ in backbone size, LoRA rank (r64 for large, r32 for base and small) and training length (860 steps for large, 1,200 for base and small). The worked-steps route is released only for large.

  • Ayaka large (Gemma 4 12B, about 24 GB in bf16): the accurate member. It gets 72/111 on the public hard tier in one pass and 88/111 with the optional worked-steps route (--reasoning), which writes short calculation notes only for low-confidence quantitative questions. Single-pass p50 is 0.18 s on A100.
  • Ayaka small (Gemma 4 E2B, about 10 GB in bf16): the lightest member, single pass. It gets 48/48, 65/72 and 55/111 on the public tiers, with p50 0.07 s on A100. Use it for the lowest latency, memory and cost.

Use (pinned)

pip install -c https://raw.githubusercontent.com/alice-noa-chan/ayaka/042d6869e7167976e921662bac2a70a1f50295aa/constraints.txt \
  "git+https://github.com/alice-noa-chan/ayaka@042d6869e7167976e921662bac2a70a1f50295aa"
python -m ayaka.serve --ckpt alice-noa-chan/ayaka-base --revision <this repo revision> \
  --device cuda --port 8000

This model is released as single pass. The worked-steps route (--reasoning) was measured and frozen only for Ayaka large, and it is unevaluated for this size.

Verified environment:

  • Python 3.11.11 and torch 2.8.0.dev20250319+cu128 (runpod/pytorch:2.8.0-py3.11-cuda12.8.1-cudnn-devel-ubuntu22.04).
  • transformers 5.17.0, peft 0.21.0, accelerate 1.15.0.
  • NVIDIA driver 570 or newer.

A Dockerfile pinned by digest is in the repository.

The server speaks TypeSafe's POST /v1/systemone format and returns a probability for every label:

  • noul: P(true);
  • choice: per-label probabilities;
  • score: per-level probabilities and the expected score.

Results (JevBench public tiers)

easy original hard
Ayaka base (single pass) 48/48 68/72 61/111
Ayaka large (single pass) 48/48 72/72 72/111
Ayaka large + worked-steps route 48/48 72/72 88/111
  • Hard-tier ECE: 0.223.
  • Held-out mixture: accuracy 0.891, ECE 0.024.
  • Open-Jev test (1000 questions): accuracy 0.912, ECE 0.019.

Public numbers will likely overstate sealed-tier results (see the disclosure below).

Latency and cost

A100 SXM 80GB, one request at a time, in-process.

tier p50 p95
easy 0.194 s 0.201 s
original 0.180 s 0.196 s
hard 0.190 s 0.401 s

The model generates no tokens. Input tokens per decision are the same as for large (the prompt format is shared):

cohort mean input tokens $ / 1,000 decisions*
Open-Jev test 348 0.0070
public original 146 0.0029
public hard 1,242 0.0248

* Priced at the leaderboard's reference for this size class: $0.02 per million input tokens.

Training

  • Adapter: LoRA r 32, alpha 64, dropout 0.05, on q/k/v/o and gate/up/down projections. Stored as bf16 safetensors.

  • Schedule: 1200 steps × 64 questions (LR 0.0001, head 0.0005, seed 0) on one A100 80GB, 3.1 h.

  • Calibration: temperatures fitted per question type and prompt length on held-out data. Long means ≥ 1,024 tokens.

    type short long
    noul 0.809 1.758
    choice 1.158 1.086
    score 1.189 1.048
  • Training: direct, on the same data and code as large. No distillation.

Training data (35 sources)

No commercial-LLM outputs and no Jev API labels were used. Every source was checked against the JevBench public items for shared 13-grams, and 0 samples matched.

source id dataset licence
jev_open SargeDev/jev-distill-corpus-v3 CC0-1.0 (Open-Jev release-v2-redistributable stream)
open_jev_bde ZefanCai/Open-Jev/browser-drone-expansion-v1-redistributable CC0-1.0 (Open-Jev redistributable config)
vitaminc tals/vitaminc CC BY-SA 3.0
helpsteer2 nvidia/HelpSteer2 CC BY 4.0
hh_rlhf Anthropic/hh-rlhf MIT
aegis_safety nvidia/Aegis-AI-Content-Safety-Dataset-2.0 CC BY 4.0
aqua_rat deepmind/aqua_rat/raw Apache-2.0
hotpot_decisions hotpotqa/hotpot_qa/distractor CC BY-SA 4.0
squad_v2_answerable rajpurkar/squad_v2 CC BY-SA 4.0
strategyqa ChilleD/StrategyQA MIT
arc_challenge allenai/ai2_arc/ARC-Challenge CC BY-SA 4.0
commonsense_qa tau/commonsense_qa MIT
legalbench nguha/legalbench CC BY 4.0 (only tasks whose README states CC BY 4.0)
synth_temporal generated: ayaka/data/synthetic.py generated by this repository (MIT)
synth_numeric generated: ayaka/data/synthetic.py generated by this repository (MIT)
synth_policy generated: ayaka/data/synthetic.py generated by this repository (MIT)
synth_long_rules generated: ayaka/data/hard_synthetic.py generated by this repository (MIT)
synth_calendar generated: ayaka/data/hard_synthetic.py generated by this repository (MIT)
synth_probability generated: ayaka/data/hard_synthetic.py generated by this repository (MIT)
snli stanfordnlp/snli CC BY-SA 3.0
multi_nli nyu-mll/multi_nli mixed by genre: OANC, CC BY 3.0, CC BY-SA 3.0 (see card)
boolq google/boolq CC BY-SA 3.0
banking77 mteb/banking77/default CC BY 4.0
clinc_oos clinc/clinc_oos/plus CC BY 3.0
klue_nli klue/klue/nli CC BY-SA 4.0
klue_ynat klue/klue/ynat CC BY-SA 4.0
kor_nli_multi kakaobrain/kor_nli/multi_nli CC BY-SA 4.0
massive_ja AmazonScience/massive/ja-JP CC BY 4.0
massive_ko AmazonScience/massive/ko-KR CC BY 4.0
jglue_jnli shunk031/JGLUE/JNLI CC BY-SA 4.0 (JGLUE)
jglue_jsts shunk031/JGLUE/JSTS CC BY-SA 4.0 (JGLUE)
jglue_commonsense shunk031/JGLUE/JCommonsenseQA CC BY-SA 4.0 (JGLUE)
go_emotions google-research-datasets/go_emotions/simplified Apache 2.0
quality emozilla/quality CC BY 4.0 (QuALITY)
quality_dev emozilla/quality CC BY 4.0 (QuALITY)

CC BY-SA and CC BY sources were used as training data only. Their licence and attribution are listed here. Whether ShareAlike terms extend to model weights is not settled, and we do not claim otherwise.

Benchmark disclosure

No JevBench items or labels were included in the training corpus. All 35 training sources were checked against the public items for shared 13-grams, and 0 samples matched. However, public JevBench results were used as a development signal. They helped identify weak task families and informed synthetic-data design, model-size selection, and the decision to develop a reasoning path. Public-item outputs were also read to diagnose failures (a zero-shot calculation-plan format, a calibration bug). One prompt example that structurally mirrored a public item was removed before the reported runs.

Not selected on public results:

  • Reasoning-route threshold, gating and fusion. This covers the calculation gate rules, the confidence cutoff of 0.9 and the fusion weight. They were selected on a procedural dev set (500 questions) and Open-Jev test (300 questions) only. The policy was frozen and hash-recorded before any public scoring.
  • Worked-steps writer. Using the base model with the adapter off, rather than with the adapter on, was decided on the dev set.
  • Checkpoint. Each model is the final checkpoint of its run. No intermediate checkpoint was compared on public items.
  • Training length and hyperparameters. Steps were set by budget (860 for large, 1,200 for small and base). Learning rate, LoRA rank and the data-mixture quotas were fixed before training and not re-tuned on public scores. The synthetic sources themselves were designed as described above.
  • Calibration temperatures. They were fitted on held-out training-distribution data.
  • Freeze. The released models are frozen for this submission, and no further changes will be made in response to public results.

License

MIT for code and fine-tuned weights. The base model google/gemma-4-E4B-it is Apache-2.0, and its notices apply to the base portion.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alice-noa-chan/ayaka-base

Adapter
(362)
this model