Instructions to use TypeSafeAI/Qyvos with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use TypeSafeAI/Qyvos with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="TypeSafeAI/Qyvos")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("TypeSafeAI/Qyvos", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Qyvos
Qyvos is a decision / classification model built from SupersonicLabs/Julia-1 by head-only fine-tuning on the ZefanCai/Open-Jev dataset (config release-v2-redistributable).
It maps a state + question + options context to a probability distribution over the options, for three decision kinds:
| qtype | kind | meaning |
|---|---|---|
| 0 | choice |
pick one option (categorical decision) |
| 1 | score |
grade/distribute mass over options (soft scoring) |
| 2 | noul |
binary yes/no decision (options ordered [no, yes]) |
Complete backbone retention
The entire Julia-1 backbone is retained bit-exact:
encoderβ ModernBERT-small (22 layers, hidden 384, vocab 256k), 140,530,432 params β frozen, untouchedact_headβ auxiliary action head β frozen, untouched
Training updated only the decision head (3,699,073 trainable params = 2.56% of 144,292,867):
headβ 2ΓTransformerEncoderLayer(384, nhead=6, dim_ff=1536, norm_first=True)type_embβEmbedding(3, 384)(one embedding per decision kind)scorerβLayerNorm β Linear(384,384) β GELU β Linear(384,1)per-option scorer
Weights are float32 end-to-end (the Julia checkpoint format has no FP16 variant; nothing was quantized or pruned).
Training
- Data:
ZefanCai/Open-Jev, configrelease-v2-redistributable,trainsplit (79,116 rows available), streamed 30,000 rows in deterministic shuffled order - Objective: soft-target cross-entropy
-(target Β· log_softmax(scores)).sum(-1).mean()β trains calibrated probability distributions, exactly matching Open-Jev's softtargetvectors - Encoding: identical
julia.data.sequence()marker-token serialization as the official Julia runtime (train/inference parity) - Optimizer: AdamW, lr 1e-4 (linear warmup 100 steps β linear decay to 10%), Ξ²=(0.9,0.999), weight decay 0.01 (no decay on norms/embeddings), grad clip 1.0
- Batch: 1 micro-batch Γ 8 accumulation (effective batch 8), 3,750 optimizer steps
- Sequence budget:
max_length=1024,head_length=512 - Seed: 17
Low-RAM training constraints (3.9 GB system RAM)
No full fine-tuning was possible (or needed). The run was designed for an 8 GB-class machine:
- frozen encoder β no gradients/optimizer state for 140.5M params; encoder forward runs under
no_grad - bs=1 micro-batches,
num_workers=0, fp32 head-only optimizer state - parquet read via pyarrow mmap, one row-group (~1,000 rows, a few MB) at a time; lazy per-row tokenization; no dataset-wide RAM cache
- head-only checkpoints (~45 MB, incl. optimizer) with deterministic skip fast-forward resume
- psutil RAM guard: auto-checkpoint + abort if system available RAM < 350 MB
- Measured peak RSS: 1.64 GB
Results (honest shuffled-mixture evaluation)
Evaluation samples are drawn round-robin across all parquet row-groups (each region contributes equally), fixed seed 999, length-bucketed batches. Accuracy = argmax match vs the target's argmax; softCE = soft-target cross-entropy (lower is better).
| model | val acc (n=900) | val softCE | test acc (n=900) | test softCE |
|---|---|---|---|---|
| Julia-1 base (no training) | 83.67% | 0.4542 | 83.56% | 0.4764 |
| Qyvos (30,000 rows) | 84.11% | 0.4282 | 83.11% | 0.4447 |
Larger-sample confirmation for Qyvos on the test split: n=1500 β acc 81.67%, softCE 0.4587.
Per-kind (Qyvos, n=900):
| kind | val acc | test acc |
|---|---|---|
| choice | 66.5% | 68.2% |
| score | 79.4% | 79.3% |
| noul | 92.7% | 90.2% |
Honest reading: Julia-1's pretrained decision head is already strong on this benchmark family, and top-1 accuracy differences between base and Qyvos are within evaluation noise (~Β±1.2pt stderr at n=900). What fine-tuning did deliver, consistently across both splits and sample sizes, is better-calibrated probabilities (softCE β5.7% val, β6.7% test) β which is precisely the objective that matters for Open-Jev's soft, often high-entropy targets, and for any downstream consumer of the full probability vector rather than a single argmax. The accuracy ceiling here is largely data-intrinsic: many Open-Jev targets are intentionally soft/ambiguous.
Exact inference (official runtime)
import julia # from the julia/ package shipped in this repo (pip install -e julia --no-deps)
engine = julia.load_model("<path-to-Qyvos>", device="cpu", max_length=1024, head_length=512)
pred = engine.predict([{
"state": "User chat: my login is blocked after the password reset.",
"question": "Determine the broad category of this support ticket.",
"options": ["account: Login, permissions, profile, security",
"billing: Charges, invoices, refunds, subscriptions"],
"type": "choice", # choice | score | noul
}])[0]
print(pred["index"], pred["probabilities"])
Or with the bundled CLI (see training/):
python3 scripts/infer_qyvos.py --demo # per-kind examples from the test split
python3 scripts/infer_qyvos.py --eval-test 900 # honest shuffled test accuracy
python3 scripts/infer_qyvos.py --row '{"state":"...","question":"...","options":["a","b"],"type":"choice"}'
Compatibility note: on
transformers >= 5.17, julia's optional fast inference path (julia/router/encoder.py) referencesModernBertModel._update_attention_mask, which was removed upstream. The bundled scripts disable that optimization via a two-line shim (specialize_decision_encoder β False); inference then uses the standard forward with numerically equivalent results. Ontransformers 5.0.x(the version this checkpoint was built against) no shim is needed.
Repository layout
Qyvos/
βββ config.json # architecture pointer file (Julia format v1)
βββ julia_config.json # Qyvos identity + head settings
βββ model.safetensors # fp32 weights (backbone bit-exact + trained head)
βββ encoder/config.json # ModernBERT-small config
βββ tokenizer/ # tokenizer (copied bit-exact from Julia-1)
βββ label_map.json # qtype id -> kind
βββ qyvos_training_config.json # full training config + final metrics
βββ provenance.json # base/dataset revisions + SHA-256 hashes
βββ training/ # all scripts used (download/inspect/train/build/infer)
Provenance
- Base model:
SupersonicLabs/Julia-1(fp32, weights SHA-256 recorded inprovenance.json) - Dataset:
ZefanCai/Open-Jev@ configrelease-v2-redistributable(per-shard SHA-256 inprovenance.json) - Backbone tensors: bit-exact vs base; only
head,type_emb,scorerdiffer - Final
model.safetensorsSHA-256:1d03ac9a0c7fbc66f1bde7ecb158741b02c9f898078c0a2f0cc270fb6326f920
- Downloads last month
- 38