Tulana eval harness

Status: experimental · Canonical source: github.com/kinyoubi-atelier/tulana-eval — release v0.1.0-pilot-experimental. Code, issues, and contributions live on GitHub; this Hub repo is the Tulana artifact family's stable address for the harness.

The evaluation harness for Tulana: instruction-following and generation-faithfulness measurement for compact open language models under native-script, romanized, and Odia-English code-mixed Odia conditions. Runner, fake+MLX engines, a byte-faithful IndicIFEval checker port (with a zero-divergence report), the studio's kinyoubi.* Odia checker layer, constructed-condition pipelines (romanization, seeded variability, code-mixing at 25/50/75%), and aggregation that never prints a point score without a 95% CI.

No benchmark data ships in this repo — see the license-separation reasoning in the GitHub README. Data lives in the sibling dataset repos.

Five-minute quickstart (zero-download)

git clone --branch v0.1.0-pilot-experimental https://github.com/kinyoubi-atelier/tulana-eval
cd tulana-eval
python3 -m venv .venv && ./.venv/bin/pip install -r requirements.txt  # recent Python required; verified on 3.14
./.venv/bin/python -m harness.cli run --config examples/configs/hello-fake.yaml

Tulana artifact family

Artifact Repo Status
Eval harness (this repo) GitHub · Hub pointer experimental, published v0.1.0-pilot-experimental
Benchmark: studio items, Class F (Pool B) KinyoubiAtelier/tulana-items-classf experimental, published; renderings PROVISIONAL
Faithfulness protocol pack KinyoubiAtelier/tulana-protocol not yet published
Benchmark: anchor pool (Pool A) KinyoubiAtelier/tulana-bench-anchor not yet published
Condition-viewer Space KinyoubiAtelier/tulana-space not yet published

Maintained by Kinyoubi Atelier & Co. · Tulana collection

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including KinyoubiAtelier/tulana-eval