NeoHorse-1-9B-GGUF

NeoHorse-1-9B is a 9-billion-parameter causal language model from TokenRhythm, post-trained from Qwen3.5-9B as an initial prototype on the path toward recursive self-improvement (RSI), targeting text-based agent harnesses, tool use, coding, and instruction following. Its core innovation is a routing harness that assigns tasks to a heterogeneous model pool, records tool interactions and outcomes, estimates capability demand, and feeds that signal back into shaping the next training mixture — using routing-guided curriculum SFT and routing-guided on-policy distillation to turn execution trajectories into training data, supported by rigorous deduplication, decontamination, and six-dimensional semantic evaluation of training data. This release contains language-model weights only (vision weights excluded, with configuration and tensor keys repackaged for text-only inference) and retains a 262,144-token native context window extensible to 1,010,000. Across a ten-benchmark evaluation against Granite-4.2-8B, Qwen3.5-9B, Ornith-1.5-9B, Gemma-4-12B-it, and the much larger Muse-Glimmer-30B, NeoHorse-1-9B posts the best overall macro average (69.04 vs. 65.60 for its Qwen3.5-9B base, a +3.44 gain), with standout improvements on agentic benchmarks like VitaBench (+11.00), PinchBench (+7.70), and QwenClawBench (+4.69), alongside strong tau2-Bench (90.82) and BFCL v4 (67.43) scores, though instruction-following gains were mixed. It's servable via SGLang or vLLM with Qwen3-style reasoning and tool-call parsers, and is released under the Apache License 2.0.

Model Files

File Name Quant Type File Size File Link
NeoHorse-1-9B.BF16.gguf BF16 17.9 GB Download
NeoHorse-1-9B.Q3_K_L.gguf Q3_K_L 4.93 GB Download
NeoHorse-1-9B.Q3_K_M.gguf Q3_K_M 4.62 GB Download
NeoHorse-1-9B.Q3_K_S.gguf Q3_K_S 4.26 GB Download
NeoHorse-1-9B.Q4_0.gguf Q4_0 5.31 GB Download
NeoHorse-1-9B.Q4_K_M.gguf Q4_K_M 5.63 GB Download
NeoHorse-1-9B.Q4_K_S.gguf Q4_K_S 5.35 GB Download
NeoHorse-1-9B.Q5_0.gguf Q5_0 6.31 GB Download
NeoHorse-1-9B.Q5_K_M.gguf Q5_K_M 6.47 GB Download
NeoHorse-1-9B.Q5_K_S.gguf Q5_K_S 6.31 GB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
1,317
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/NeoHorse-1-9B-GGUF

Finetuned
Qwen/Qwen3.5-9B
Quantized
(13)
this model

Collection including prithivMLmods/NeoHorse-1-9B-GGUF