Ornith-1.0-9B oQ5 Text-Only (Optimized for Apple Silicon)

This repository contains a custom-quantized, text-only configuration of the Ornith-1.0-9B dense model, optimized explicitly for local repository-level agentic coding on Apple Silicon using the oMLX inference engine.

🎯 Why This Was Created

Ornith-1.0-9B is a state-of-the-art, self-improving dense model specialized for agentic coding. It jointly optimizes search scaffolds and solution rollouts via Reinforcement Learning to discover superior code trajectories.

This specific build was converted using oMLX v0.4.5.dev1 to address long-context deployment constraints on Apple Silicon:

  • The Dense Precision Balance: As a dense 9B model, Ornith-1.0-9B provides outstanding agentic coding performance in a very compact footprint. Quantizing it to oQ5 (5-bit) delivers a notable step up in code syntax retention and complex instruction adherence over 4-bit alternatives, while keeping the active footprint low enough to maximize headroom for deep KV caches during repository-wide multi-file edits.
  • Overcoming the 128k Context Wall: Traditional backends often choke or suffer severe latency degradation when context sizes scale out. Moving to oMLX's specialized two-tier caching eliminates this overhead, allowing you to fluidly ingest giant codebases.
  • Prefill Speedup via float16: While the base weights are distributed in bfloat16, this build explicitly targets Apple Silicon hardware by using float16 for non-quantized weights, unlocking a ~20% faster prefill speed on M1/M2 Max chips.
  • MTP Note: Unlike some Qwen base models, the original Ornith-1.0 weights do not contain Multi-Token Prediction (mtp.*) headers. As a result, native MTP decoding is not available for this model.

🚀 Key Differences

Feature / Attribute Standard Ornith-1.0-9B This Custom Build (oQ5-fp16)
Native MTP Heads Not present in base architecture Not Available (No base MTP tensors)
Vision Model (VLM) N/A (Text-only coding agent) Stripped/Text-Only
Quantization Method Standard Uniform / HF / Unsloth oQ5 (Dynamic mixed-precision calibration)
Non-Quant Weight DType bfloat16 float16 (~20% faster prefill on M1/M2 Silicon)

💻 Hardware & RAM Recommendations

Mac Hardware Configuration RAM Recommendation Status / Performance Expectation
Apple Silicon (Base 8GB / 16GB / 24GB) 16GB Unified Memory Supported — 8GB systems can run it with tight context constraints. 16GB is recommended as the baseline for comfortable use.
M1 / M2 / M3 / M4 (16GB / 24GB / 36GB / 48GB) 24GB / 36GB Unified Memory Recommended — Excellent response latency and solid space for long context windows.
M1 / M2 / M3 / M4 Max / Ultra (64GB+) 64GB+ Unified Memory Optimal / Best Experience — Unlocks the full 262k context boundary, keeping the model running entirely in fast cache.

🛠️ Quantization Settings

This model was quantized using oMLX v0.4.5.dev1 with the following specification:

  • Source Model: deepreinforce-ai/Ornith-1.0-9B
  • Sensitivity Model: None
  • oQ Level: oQ5
  • Text Only: Enabled
  • Preserve MTP weights: Disabled (Not present in source architecture)
  • Non-quant weight dtype: float16

⚙️ Optimized oMLX Settings (v0.4.5)

To seamlessly route this model through agentic development workspaces like OpenCode, apply the following server specifications in your oMLX dashboard:

Model Basic Settings

  • Reasoning Parser: qwen3 (Isolates the <think> ... </think> blocks securely away from IDE syntax parsers)
  • Tool Call Parser: qwen3_xml
  • CTX Window: 262,144
  • Max Tokens: 32,768
  • Temperature: 0.6 (Use 1.0 if attempting to perfectly replicate official benchmark environments)
  • Top P / Top K: 0.95 / 20
  • Min P: 0
  • Repetition / Presence Penalty: 1 / 0

Model Advanced Settings

  • Enabled Thinking: Checked (True)
  • Chat Template Kwargs: enable_thinking: true, preserve_thinking: true
  • Native MTP: Unchecked (False)

Resource Management & Cache

  • Memory Guard: Aggressive (Forces strict macOS memory/swap cleanup cycles)
  • Hot Cache Limit (RAM): 40GB (Allocated for hyper-speed Unified Memory history)
  • Cold Cache Limit (SSD): 371GB (Serialized safetensors storage for context overflow handles)
  • Max Concurrent Requests: 2 (Protects the 400 GB/s bandwidth bus from degradation)
  • Embedding Batch Size: 32
  • Chunked Prefill: Enabled (Prevents instantaneous out-of-memory crashes on massive project context ingestion)
  • Burst Decode: Aggressive (Coalesces tokens for maximized typing speeds)
  • Initial Cache Blocks: 256
  • SSE Keepalive Mode: Chunk

🌡️ Thermal Optimization Notice

Sustained execution across massive context windows heavily taxes the Apple Silicon GPU/CPU complexes, causing rapid heat buildup. Because Apple's default fan profiles emphasize near-silent operation, they delay ramping up system fans until thermal throttling has already begun to affect generation tokens-per-second (TPS).

To protect performance integrity during prolonged coding sessions, it is highly recommended to run a custom fan utility to enforce proactive, aggressive cooling curves:

Downloads last month
36
Safetensors
Model size
9B params
Tensor type
U32
·
F16
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Noctalin/Ornith-1.0-9B-oQ5-fp16

Quantized
(103)
this model