Tellama model distributions

The distribution below passed Tellama's stated model qualification gates. App release and update-installation acceptance are tracked separately.

TM Qwen3 4B Q4_K_M โ€” Tellama quantization

Tellama quantized the original Qwen3 4B weights through F16 into Q4_K_M, retaining Q8_0 token embeddings and output tensors. No fine-tuning or calibration was used. Qwen/Alibaba Cloud retains authorship of the base model. Tellama's validation does not imply Qwen endorsement or guarantee correct answers.

  • Exact file: models/qwen3-4b-tellama-thinking-q4_k_m.gguf
  • Bytes: 2,591,481,472 (approximately 2.59 GB download).
  • SHA-256: 01e739de8a200c769e72a676d038b84b1c042bd160a153a348fbb443bd713b85.
  • Base revision: 1cfa9a7208912126459214e8b04321603b3df60c.
  • llama.cpp revision: cb295bf59663cd3577389315636772f4060bd1f5.
  • License and attribution: licenses/qwen3-4b-tellama-thinking-q4_k_m/.
  • Reproducible recipe: recipes/qwen3-4b-tellama-thinking-q4_k_m.json.

Measured support and limitations

Measured on Samsung Galaxy Z Fold6 SM-F956N / Android 16, CPU, Tellama DEEP mode, 4,096-token context, F16 KV cache, default 768-token output budget, temperature 0.7, top-p 0.9 and top-k 40. Explicit user profiles can differ from these settings.

Screen Result
Basic, KV reset each turn 21/21
Basic, reuse allowed 20/21
Extended, multi-turn and long context 13/13

All critical checks passed; 54 of 55 total checks passed. The noncritical failure translated the requested Korean JSON city value into English (Busan). Do not interpret this small synthetic suite as universal accuracy or structured-output certification. Review important answers independently.

Minimum measured decode speed was 3.76 tokens/s. Visible responses sometimes took over two minutes because of internal reasoning. Basic short prompts were fully re-prefilled (no actual KV reuse); the extended suite reused up to 1,122 tokens. Peak observed Android thermal status was 2. Memory/thermal guards remain enabled; these measurements do not guarantee sustained performance under every condition.

The original quality reports truthfully identify an internal modified 1.3.8 benchmark build. They are not reports from the previously published stock 1.3.8. Internal 1.3.9 hosting acceptance completed a full download, interruption/resume, process restart, SHA-256 verification and installation on Android16/16KB emulator. Fold6 1.3.9 runtime replacement from a loaded Qwen2.5 1.5B to this 4B and back passed at 4,096 context, without changing user settings. These are distinct from the final signed APK/update release checks.

FAST mode, GPU execution, other devices, image/audio input, and the original upstream maximum context are not qualified by these results. Tellama requires explicit consent to enable DEEP for this distribution. An ordinary Qwen3 base model can technically run without thinking; this curated variant is offered only under the settings that passed Tellama's checks.

Previous experimental files

The previous Qwen3.5 4B file remains available for provenance/history, but its extended arithmetic checks failed. It is not a TM-approved recommendation and is not offered for new downloads in the new curated catalog. Existing users' local models and conversations are not automatically deleted.

Downloads last month
118
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for LaraAI-Labs/tellama

Finetuned
Qwen/Qwen3-4B
Quantized
(335)
this model