RACHIDA-Mini-2.03B-Q4_K_M GGUF

This repository contains the GGUF selected for RACHIDA.AI, an offline maternal-health information and safety prototype submitted to the Africa Deep Tech Challenge 2026.

The file is a straight Q4_K_M quantization of Qwen/Qwen3-1.7B. No RACHIDA-specific fine-tune or adapter is merged into this artifact. RACHIDA.AI is the surrounding application and safety system; the model provenance remains Qwen3-1.7B.

File

File Format Size SHA-256
RACHIDA-Mini-2.03B-Q4_K_M.gguf GGUF Q4_K_M 1,282,439,040 bytes fc0d5cc4ce1cee6ec3566da40c48f9ed953e3ee2eedcce36e44b3c4924001ab9

The GGUF tensor metadata reports 2,031,739,904 parameters. The upstream model is named Qwen3-1.7B; the 2.03B label records the tensor count reported by the converted artifact.

Intended use

  • Offline language-model evaluation with llama.cpp.
  • Preliminary maternal-health information and patient-education research.
  • Demonstration of CPU-only inference on commodity laptops.
  • Use inside an application where independently verified deterministic controls own safety-critical escalation decisions.

Important limitations

This artifact is not clinically validated. It is not a medical device, diagnostic system, triage system, prescription tool, or replacement for a clinician. Do not use it for autonomous diagnosis, treatment, prescribing, emergency-resource lookup, or decisions involving patient safety. The GGUF alone does not contain RACHIDA.AI's signed clinical content pack, deterministic danger-sign routing, output guard, or audit controls.

English is the only declared and evaluated language scope for this release. Do not provide patient records or other sensitive personal data to the model.

Run locally

llama-cli \
  -m RACHIDA-Mini-2.03B-Q4_K_M.gguf \
  -t 4 \
  -ngl 0 \
  -c 2048

The ADTC submission uses llama.cpp, four CPU threads, zero GPU layers, and no network access during inference.

Verified development measurements

Participant-mode measurements on the development laptop:

  • 23.09 generated tokens/second;
  • 2,521.14 ms first-token latency;
  • 1,937.40 MB peak RSS;
  • 0.76 ARC-Easy accuracy over 50 samples with seed 42.

These are development results, not official ADTC audit-machine results. The development host ran Windows and had 31.7 GB RAM; Ubuntu 22.04 / 8 GB results may differ. Core temperature was unavailable, so no thermal claim is made.

License and attribution

The weights retain the upstream Apache License 2.0. Base model: Qwen3-1.7B by the Qwen Team. Conversion, quantization, and CPU inference use llama.cpp, licensed under MIT.

Downloads last month
95
GGUF
Model size
2B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Mariem-Daha/RACHIDA-Mini-2.03B-GGUF

Finetuned
Qwen/Qwen3-1.7B
Quantized
(339)
this model