NIM-1 (3B)

NIM-1 is a high-efficiency 3-billion parameter local reasoning model engineered for rapid, deterministic task execution, code intelligence, and structured agentic workflows. Built to deliver flagship reasoning density within consumer hardware limits, NIM-1 runs completely offline with ultra-low latency.

Hugging Face Ollama License: Apache-2.0


Highlights

  • High-Density Reasoning: Tuned on high-signal chain-of-thought demonstrations to handle multi-step logic, Python algorithms, and complex instruction following.
  • Consumer-Ready Edge Inference: Consumes under 4 GB of VRAM in 8-bit quantization, delivering 60โ€“90 tokens/sec on entry-level GPUs (such as RTX 3060/4060) or Apple Silicon.
  • Zero-Degradation Precision: Packaged with Q8_0 GGUF quantization for maximum output stability and numerical fidelity.
  • Autonomous Agent Ready: Built with native support for tool orchestration, structured schema outputs, and multi-turn conversational memory.

Quickstart with Ollama

Run NIM-1 instantly from Hugging Face:

ollama run hf.co/NIM-AI/NIM-1-3B:NIM-1-3B-Q8_0.gguf

Or build and run directly from the local repository:

ollama create nim-1 -f Modelfile
ollama run nim-1

Model Specifications

Parameter Specification
Model Architecture Dense Transformer (Decoder-only)
Total Parameters 3.09 Billion
Context Window 2,048 tokens (extensible to 32k)
Quantization Format GGUF (Q8_0)
Inference Footprint ~3.4 GB VRAM / System Memory
Chat Template ChatML format (`<

Architectural & Training Methodology

NIM-1 was trained using parameter-efficient fine-tuning (QLoRA) with custom Triton-accelerated backpropagation kernels:

  1. Curated High-Signal Supervision: Trained against an curated dataset of algorithmic logic puzzles, step-by-step math derivations, and software architecture patterns.
  2. Quantized Fine-Tuning: Trained using 4-bit base model weight caching, rank-16 target projection adapters (q, k, v, o, gate, up, down), and fused cross-entropy loss.
  3. High-Fidelity Export: LoRA adapters merged back into 16-bit floating-point space before single-pass quantization to Q8_0 GGUF format to avoid quantization degradation.

Hardware Compatibility

Environment Performance VRAM / Memory
NVIDIA RTX 4060 (8 GB) ~70โ€“90 tok/s ~3.4 GB
Apple Silicon (M-Series, 16 GB) ~60โ€“80 tok/s ~3.6 GB Unified
Modern x86 CPU (AVX-512) ~18โ€“28 tok/s ~4.2 GB RAM

Reproducing Training & Packaging

# Clone the repository
git clone [https://github.com/NIM-AI/NIM-1.git](https://github.com/NIM-AI/NIM-1.git)
cd NIM-1

# Install requirements
pip install -r requirements.txt

# Run the training pipeline
python scripts/train_student.py

# Export weights to GGUF format
python scripts/export_gguf.py
Downloads last month
41
GGUF
Model size
3B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for NIM-AI/NIM-1-3B

Base model

Qwen/Qwen2.5-3B
Quantized
(298)
this model