ThoxAir-16M-role

On-device conversational model for the ThoxAir RV1103 (Cortex-A7 ARMv7-A + NEON). BitNet b1.58 ternary, 15,737,088 params, 86.6% ternary.

Lineage β€” verified

Fine-tuned from Thox-ai/ThoxMicro-1bit-16M, a THOX model trained from scratch. No external base. Training initialised from that run's checkpoint (deep-16m-ternary/best.pt); the Hub repo publishes the same run as GGUF.

ThoxAir is a dual-chip device

This is the RV1103 half. The ESP32-C6 half is Thox-ai/ThoxMesh-Head-C6, an int8 TFLite triage classifier that decides which radio/sensor events are worth waking this model for.

Needs an armv7-neon llama.cpp build β€” the Pi Zero arm64 binaries will not run on it.

Budget: 5.51 MB weights, 13.89 MB resident of ~33 MB usable, ~11.0 tok/s. 1-bit inference is already proven on this silicon at 10.7–11.2 tok/s.

val_loss 1.6163. Ship-then-test: not measured on target. Trained locally at $0.

Downloads last month
55
GGUF
Model size
15.7M params
Architecture
llama
Hardware compatibility
Log In to add your hardware

1-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Thox-ai/ThoxAir-16M-role

Quantized
(1)
this model

Space using Thox-ai/ThoxAir-16M-role 1