Instructions to use open-zzrl/kyo-instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use open-zzrl/kyo-instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="open-zzrl/kyo-instruct")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("open-zzrl/kyo-instruct", device_map="auto") - Notebooks
- Google Colab
- Kaggle
β‘ Kyo-Instruct (140M System-One Decision Engine)
Kyo-Instruct is an ultra-fast (140M parameter) System-One semantic routing and policy gating engine engineered to govern autonomous AI agents, tool selection, and high-throughput intent pipelines in real time.
Fine-tuned directly on top of open-zzrl/kyo using multi-domain Experience Replay, Kyo-Instruct bridges the gap between deterministic typed policy Instruct and generalized natural language intent classification. Powered by a specialized Dual-Pooling Cross-Encoder head ([CLS] + Masked Mean to 768d over a ModernBERT backbone), it delivers sub-10ms deterministic execution with 99.95% order-invariance, outperforming models 3Γ its size.
- GitHub Repository: github.com/open-zzrl/kyo
- Developed by: constructai & zait-ai (
open-zzrl) - Base Model:
open-zzrl/kyo - Underlying Backbone:
jhu-clsp/mmBERT-small(ModernBERT encoder, 12 layers, hidden size 384) - Context Window: Up to 4096 tokens (native RoPE + SDPA)
- Average Inference Latency: 8.80 ms (P95: 9.71 ms on CUDA,
bfloat16) - Memory Footprint: ~280 MB VRAM (< 1.5 GB peak with runtime context)
π Benchmark Results
Evaluated across 2,000 holdout agent policy scenarios and multi-class NLP benchmarks:
1. Comparative SOTA Overview
| Benchmark / Capability | Kyo-Instruct (140M) | Laya Base (421M) | TypeSafe Jev (~490M) |
|---|---|---|---|
| DAIR Emotion (6 classes, 2000 holdout) | 71.55% | 59.50% | 48.00% |
| Banking77 (Full 77 classes, 3076 test) | 47.56% (77 options) | 42.50% (72 options) | 87.00% (72 options) |
| Typed-Decisions (Agent Policy Holdout) | 71.60% | 76.60% | 72.70% |
β³ noul (Boolean Safety & Gating) |
82.33% | 85.70% | 77.50% |
β³ choice (Action & Tool Routing) |
69.50% | 73.30% | 72.00% |
β³ score (Risk Tiers 0β3) |
65.12% | 72.30% | 69.60% |
| Order-Invariance Consistency | 99.95% | 27.05% | 13.00% |
| Average Latency (GPU) | 8.80 ms | ~32 ms | ~36 ms |
| P95 Latency (GPU) | 9.71 ms | > 45 ms | > 50 ms |
| Parameter Scale | 140M (~280 MB) | 421M (~1.2 GB) |
2. Detailed Performance Analysis
- 99.95% Order-Invariance: Autoregressive models and standard classification heads suffer severe position bias when candidate options are shuffled. Kyo isolates candidate scoring pairs, guaranteeing decisions remain stable regardless of input option order.
- Full-Scale Candidate Evaluation: On Banking77, Kyo processes all 77 candidates simultaneously per pass with 12.18 ms latency, eliminating the need to pre-filter options.
- Weight Efficiency: At 140M parameters, Kyo-Instruct exceeds the 421M parameter Laya by +12.05% on emotion classification and +5.06% on fine-grained banking intent routing.
π¦ Installation via GitHub
Because Kyo utilizes an optimized dual-pooling cross-encoder head, models run natively via the lightweight kyo Python engine:
# Install directly from source
pip install git+[https://github.com/open-zzrl/kyo.git](https://github.com/open-zzrl/kyo.git)
Upgrading to Latest Release
pip install --upgrade --no-cache-dir --force-reinstall git+[https://github.com/open-zzrl/kyo.git](https://github.com/open-zzrl/kyo.git)
Requirements: Python β₯ 3.9, PyTorch β₯ 2.0 with CUDA support, and
git.
π Quickstart
Kyo loads pre-trained weights, tokenizer configurations, and execution parameters directly from Hugging Face Hub:
import torch
from kyo import Kyo
# 1. Initialize the engine
engine = Kyo.from_pretrained("open-zzrl/kyo-instruct", confidence_threshold=0.80)
# ==========================================
# Scenario A: Autonomous Agent Policy & Safety Gating
# ==========================================
telemetry = {
"agent_id": "payment-worker-4",
"endpoint": "/api/v1/transfer/batch",
"requested_amount_usd": 145000,
"spending_limit_usd": 50000,
"user_approval_present": False,
"geo_anomaly": True
}
safety_Instruction = "Verify transaction safety against corporate treasury compliance rules."
safety_options = {
"Allow": "Transaction within limits and conforms to standard policy.",
"Require2FA": "Minor anomaly; hold for secondary automated factor.",
"BlockAndEscalate": "Hard limit breach or suspicious telemetry; freeze and alert human."
}
# Warm-up (initializes CUDA context and JIT-compiles SDPA kernels)
_ = engine.decide(context=telemetry, Instruction=safety_Instruction, options=safety_options)
result_safety = engine.decide(
context=telemetry,
Instruction=safety_Instruction,
options=safety_options
)
print("-" * 55)
print(f"Safety Decision : {result_safety.decision}")
print(f"Confidence : {result_safety.confidence * 100:.2f}%")
print(f"Hot Latency : {result_safety.latency_ms:.2f} ms")
print(f"Requires Fallback : {result_safety.fallback_to_llm}")
# ==========================================
# Scenario B: Dynamic Semantic Intent Classification
# ==========================================
customer_query = "Why is my card balance still unchanged after I deposited cash at the ATM two hours ago?"
routing_Instruction = "Classify customer inquiry into the matching operational queue."
routing_options = {
"atm_deposit_pending": "Delays or missing balance updates after ATM cheque or cash deposit.",
"wire_transfer_delay": "Inbound or outbound international wire transfer pending.",
"compromised_card": "Card reported stolen, cloned, or fraudulent charges detected.",
"order_physical_card": "Requesting delivery of a new physical replacement card."
}
result_route = engine.decide(
context=customer_query,
Instruction=routing_Instruction,
options=routing_options
)
print("-" * 55)
print(f"Routed Queue : {result_route.decision}")
print(f"Confidence : {result_route.confidence * 100:.2f}%")
print(f"Latency : {result_route.latency_ms:.2f} ms")
print("-" * 55)
π― Intended Use & Boundaries
β What Kyo Excels At (In Scope)
- Agent Security Guardrails & Policy Gating: Evaluating execution traces, payloads, and parameter states against programmatic guardrails (
noulaccuracy: 82.33%). - Deterministic Tool & Action Routing: Dynamic selection of functions, endpoints, or execution paths from candidate options (69.50% on structured choices, 47.56% across 77 simultaneous intents).
- Real-Time Stream Routing: Intent categorization and customer support triage under strict latency constraints (< 10 ms).
- Cost-Efficient Frontier LLM Offloading: Calibrated confidence scoring identifies uncertain predictions (
result.fallback_to_llm = True), routing only ambiguous requests to costly generative LLMs.
β Out of Scope
- Open-Ended Autoregressive Generation: Kyo is a pure discriminator/cross-encoder, not a generative language model.
- Relational Multi-Column Spreadsheet Arithmetic: Complex numerical calculations over dense tabular formats require specialized math reasoning models or large frontier LLMs.
π§ Architecture: Dual-Pooling Pairwise Cross-Encoder
Kyo scores each candidate option $k$ in parallel against the shared query context:
[Context + Instruction] βββ
βββ> [mmBERT-small (12L, 384d)] ββ> [CLS (384d)] βββββββββ
[Candidate Option (k)] ββββ [Mean Pool (384d)] βββ΄β> LayerNorm (768d) ββ> MLP Head ββ> Logit_k
- Pairwise Isolation: Evaluates each candidate independently against the context, eliminating candidate-order positional bias and guaranteeing 99.95% order-invariance.
- Dual Representation: Concatenating global contextual attention (
[CLS]) with dense token coverage (masked mean pooling) doubles representation capacity without increasing encoder FLOPs.
π οΈ Environment & Optimization Notes
- Native
bfloat16Calibration: Kyo is natively trained and calibrated inbfloat16. Do not wrap evaluation in standard fp16GradScaler. - Latency Profile:
- Cold Start (Run 1): ~1.0β1.4s (CUDA context initialization and PyTorch SDPA kernel compilation).
- Warm Inference (Run 2+): 8.80 ms average (P95: 9.71 ms).
Python 3.13 & Vision Dependencies Note
Kyo is an encoder-based text decision engine and does not use vision or audio models. If your environment encounters ABI issues with torchvision (e.g. RuntimeError: operator torchvision::nms does not exist), you can safely strip vision libraries:
pip uninstall -y torchvision torchaudio torchao
π License & Attribution
The model weights and inference runtime are released under the Apache-2.0 License.
@software{kyo_Instruct2026,
author = {constructai, zait-ai},
title = {Kyo-Instruct: 140M System-One Decision and Routing Engine},
year = {2026},
publisher = {Hugging Face},
journal = {Hugging Face Model Hub},
howpublished = {\url{[https://huggingface.co/open-zzrl/kyo-instruct](https://huggingface.co/open-zzrl/kyo-instruct)}}
}
Model tree for open-zzrl/kyo-instruct
Base model
open-zzrl/kyo


