barga
barga (Nepali बर्ग, "category") is a small system-one decision model for English and Nepali. It reads a state (a conversation, a call transcript, a document, a game position) and one or more questions, each with its own candidate options, and returns a probability for every option in a single forward pass. It does not generate text. It decides.
- Size: 140.6M parameters (ModernBERT / mmBERT-small encoder + a small option scorer), fp32 safetensors, 562 MB.
- Languages: English, Nepali (Devanagari and romanized), and code-mixed Nepali–English.
- Architecture: split-layer. The lower 11 of 22 encoder layers read the state once; the top 11 layers read each question jointly with the state. Every question about the same state shares one state encoding.
- Speed: 0.87 s per turn (median; 1.38 s p95) on an Apple M3 Pro CPU with 4 threads, for simulator call turns of about 6 questions over states of up to 1,024 tokens; a short input with one question takes 0.05–0.2 s. Peak process memory 2.6 GB.
- Licence: Apache-2.0. Built on Julia-1 (Apache-2.0), which is built on mmBERT-small (MIT).
- Made by Ampixa.
- Try it: play against barga in the Kirtipur courtyard: Bagh-Chal, chess, Ludo, Marriage, Dhumbal and Call Break, with every option it weighed and the exact request and reply. Launch video: English · नेपाली.
- Faster than a chat model: on 40 synthetic customer-service decisions, barga on a laptop CPU took 0.24 s per decision (median) and Gemini 3.8 Flash over its API took 2.25 s; both answered 38 of 40 correctly. The deployments differ (local model vs. hosted API including network).
Quickstart
pip install barga
import barga
model = barga.load("ampixa/barga") # device="cuda" for a GPU
state = ("Caller: Hello, I want to book a check-up for my dog tomorrow morning.\n"
"Agent: Sure. We have 9:30 or 11:00 tomorrow.\n"
"Caller: 9:30 is fine.")
for d in model.decide(state, [
{"id": "intent", "text": "What does the caller want?",
"options": [{"id": "book", "description": "Book an appointment"},
{"id": "cancel", "description": "Cancel an appointment"},
{"id": "info", "description": "Only ask for information"}]},
{"id": "intent_ne", "text": "कल गर्नेले के चाहनुहुन्छ?",
"options": [{"id": "book", "description": "अपोइन्टमेन्ट बुक गर्न"},
{"id": "cancel", "description": "अपोइन्टमेन्ट रद्द गर्न"},
{"id": "info", "description": "जानकारी मात्र सोध्न"}]},
]):
print(d.question_id, d.choice, d.probs)
# intent book {'book': 0.989, 'cancel': 0.009, 'info': 0.003}
# intent_ne book {'book': 0.983, 'cancel': 0.002, 'info': 0.015}
Questions can be "choice" (default), "boolean" (two options valued True/False) or "ordinal" (numeric
option values; the result also carries an expected_value). Each takes 2–20 options and an optional rubric. The
model reads only the state, the question, the rubric and each option's description, never the ids. The input
format is documented in the barga package.
How to read the output
choice is the most probable option; probs gives every option's probability (they sum to 1 within a question).
Probabilities compare options within one question. A flat distribution (say a top probability below 0.6) means the
model is unsure, which is often correct when the state does not answer the question. Set your own threshold on
your own data before acting on a decision automatically.
Intended uses
- Turn-by-turn decisions in voice and chat agents: intent, slot values, whether the caller confirmed, what to do next.
- Classification, routing and yes/no checks over English or Nepali text, with the label set given at run time.
- Reading-comprehension style checks (does the text support this statement? is this question answerable?).
- Rule and strategy questions about Nepali-played card and board games (see the limits below).
Evaluation
All numbers are accuracy (%) on held-out test sets that were never used for training or checkpoint selection. The released checkpoint is seed 17; two-seed is the mean of seed 17 and an identically trained seed-18 replicate, shown to indicate seed variance. Sampled sets use a fixed seed (17).
Calls and agent decisions
| test set | questions | released | two-seed |
|---|---|---|---|
| Held-out real customer-service calls, set A (English questions) | 2,589 | 77.0 | 76.6 |
| same calls, Nepali questions | 2,589 | 77.1 | 76.7 |
| same calls rendered in the output style of a speech recognizer (English / Nepali questions) | 2,589 | 72.1 / 71.6 | 72.1 / 71.9 |
| Held-out real calls, set B (English) | 135 | 64.4 | 64.8 |
| set B in speech-recognizer style | 135 | 62.2 | 60.7 |
| Receptionist simulator, template phrasing (300 calls) | 1,779 | 94.4 | 94.4 |
| Receptionist simulator, LLM phrasing (300 calls) | 1,678 | 96.1 | 96.1 |
| typed-decisions test split | 2,000 | 76.3 | 76.0 |
Reading comprehension and classification (300-question samples unless noted)
| test set | released | two-seed |
|---|---|---|
| MultiNLI | 66.0 | 65.5 |
| MNLI-Nepali | 64.3 | 63.0 |
| SQuAD 2.0 | 61.0 | 60.2 |
| BoolQ | 60.7 | 62.3 |
| ShARC | 62.7 | 64.2 |
| PAWS | 56.7 | 56.8 |
| CLINC150 intents | 90.7 | 91.3 |
| AG News (100) | 95.0 | 94.0 |
| Emotion (100) | 81.0 | 84.5 |
| Hard cases (1,500; the hard slice of an Ampixa support-ticket decision set) | 26.3 | 26.1 |
Games (rule and strategy questions; positions from engines and self-play)
| game | questions | released | two-seed |
|---|---|---|---|
| Dhumbal | 1,440 | 62.0 | 62.1 |
| Ludo | 1,592 | 62.9 | 62.7 |
| Marriage | 1,494 | 47.7 | 47.6 |
| Call Break | 1,828 | 47.8 | 46.7 |
| Chess | 1,674 | 42.8 | 41.3 |
| Bagh-Chal | 3,726 | 24.0 | 24.1 |
Game questions have 2–20 options, so chance differs by set: about 14% on Bagh-Chal, and up to 50% on yes/no families.
Against Julia-1 on identical questions
Same questions and the same scoring for both models; a question a model cannot encode counts as wrong.
| Julia-1 | barga (released) | |
|---|---|---|
| typed-decisions test | 72.5 | 76.3 |
| CLINC150 | 55.7 | 90.7 |
| MultiNLI / MNLI-Nepali | 30.7 / 31.0 | 66.0 / 64.3 |
| ShARC / BoolQ / SQuAD 2.0 | 31.0 / 45.7 / 51.0 | 62.7 / 60.7 / 61.0 |
| PAWS | 54.3 | 56.7 |
| AG News (100) | 94.0 | 95.0 |
| Emotion (100) | 86.0 | 81.0 |
| Hard cases | 33.7 | 26.3 |
| Julia-1's own evaluation families (12 Open-Jev families, macro) | 87.8 | 58.5 |
On the real-call and simulator sets barga scores 77 and 94–96, against Julia-1's 22 and 23. Read that gap with care:
- Encoding: Julia-1 cannot encode some of those questions; on the questions both models answer it scores 17.7–24.8.
- Training data: barga was trained on the training splits of those same sources, and Julia-1 was not.
Julia-1 leads on its own evaluation families: game-playing agents (tile platformer, ViZDoom, Snake), workflow and customer controls, and reasoning controls. barga was not trained on those families.
Limitations
- Near chance on some game families. Chess move legality (yes/no), tic-tac-toe, Bagh-Chal, Marriage discards, Ludo safe squares and Call Break card play score within a few points of chance. Do not use barga as a game engine; use a rules engine for legality and barga, at most, to rank legal moves.
- Out of domain. Julia-1's agent-control families (above) and the hard support-ticket cases are weak spots.
- Real-call accuracy is about 77%. Speech-recognizer errors cost about 5 points (72%). Keep a human or a confirmation step in the loop for anything with consequences.
- Calibration. Probabilities are informative but not calibrated to your data; threshold them on your own data.
- Context. 1,024 state tokens (when a state is longer, barga keeps its beginning and its most recent end) and 768 tokens per question with its options. Longer questions are rejected, not cut.
- Languages. Tested on English, Nepali (Devanagari and romanized) and code-mixed text. Other languages inherit some ability from mmBERT but were not evaluated.
- Seed variance. Two identically trained seeds differ by up to 3.5 points on small sets (135–300 questions); differences smaller than that between models are not meaningful.
Out-of-scope uses
Generating text; open-ended question answering without candidate options; medical, legal, financial or safety-critical decisions without human review; profiling or surveillance of people; any use that violates the licences of the training data sources listed below.
Training data
barga was fine-tuned from the encoder of Julia-1 on a weighted mix of 32 training sets:
| source | share | licence / terms |
|---|---|---|
| Nepali/English customer-service calls from partners, used under agreement; not released | 13.1% | private; used under agreement |
| Receptionist simulator: synthetic calls; template phrasing and LLM-written phrasing | 13.9% | generated by Ampixa |
| typed-decisions (train split) | 8.8% | Apache-2.0 |
| MultiNLI, SQuAD 2.0, ShARC, CLINC150, BoolQ, PAWS (English) | 19.5% | CC BY 3.0 / CC BY-SA 3.0 / MIT / other (MultiNLI), CC BY-SA 4.0, CC BY-SA 3.0, CC BY 3.0, CC BY-SA 3.0, "may be freely used for any purpose" |
| MultiNLI, SQuAD 2.0, ShARC, CLINC150, BoolQ machine-translated to Nepali (Sarvam-Translate) | 14.1% | as the originals |
| MNLI-Nepali (IRIIS-RESEARCH) | 4.8% | see the dataset |
| Bagh-Chal, Marriage, Dhumbal, Ludo, Call Break positions with rule-checked labels | 12.4% | generated by Ampixa |
| Chess positions from the Lichess open database with engine labels | 3.0% | CC0 (Lichess database) |
| Ampixa support-ticket decision sets (routing, negation, urgency, emotion, abstention) | 4.8% | Ampixa |
| Typed-decision replays with teacher soft labels | 2.4% | generated by Ampixa |
| AG News, Emotion (dair-ai), Banking77, MASSIVE (scenario) | 3.2% | research / non-commercial (AG News), educational and research purposes only (Emotion), CC BY 4.0, CC BY 4.0 |
Protected test calls and the typed-decisions test split were never used for training. No call audio or transcripts are released.
Files
backbone/: ModernBERT encoder (config.json, model.safetensors)heads.pt: option scorer weightstokenizer/: tokenizer (the mmBERT/Gemma-style 256k vocabulary)bundle.json: architecture, serialization config and the sha256 of every file above;barga.loadverifies them before loading. Its backbone reference was rewritten from a local path to its public source; every hashed file is byte-identical to the evaluated checkpoint.
Citation
@misc{barga2026,
title = {barga: a system-one decision model for English and Nepali},
author = {Ampixa},
year = {2026},
url = {https://huggingface.co/ampixa/barga}
}
barga builds on Julia-1 (Supersonic Labs) and mmBERT (Johns Hopkins University CLSP).
Model tree for ampixa/barga
Evaluation results
- Accuracy (released checkpoint) on typed-decisions (test split)test set self-reported76.300
- Accuracy (released checkpoint) on MultiNLI (300-question sample)self-reported66.000
- Accuracy (released checkpoint) on MNLI-Nepali (300-question sample)self-reported64.300
- Accuracy (released checkpoint) on SQuAD 2.0 (300-question sample)self-reported61.000
- Accuracy (released checkpoint) on CLINC150 (300-question sample)self-reported90.700