dan-omni-3b-mobile

Ultra-compact variant of dan-omni-3b optimized for strict mobile memory budgets. Smaller context window (2K) for minimal RAM footprint.

Model Description

Property Value
Base Model Qwen2.5-3B (further quantized)
Fine-tuning LoRA on mobile-optimized instruction data
Context Length 2048 tokens
Parameters 3B (compressed)
File Size 1.2 GB
RAM Usage ~1.5 GB at inference

System Prompt

You are dan, a helpful AI assistant for mobile devices. You can help with general questions, writing, math, coding, translation, and creative tasks. Keep answers concise and natural. Don't explain your architecture unless asked.

Benchmarks

Tested on Intel i9-9880H @ 2.30GHz, 16GB RAM, Ollama runtime.

Category Avg tok/s Prompt tok/s Tokens Time
Reasoning 9.4 37.6 170 18.1s
Coding 10.2 37.6 163 16.0s
Creative Writing 9.5 37.6 37 3.9s
Instruction Following 9.9 37.6 52 5.3s
Math 10.0 37.6 127 12.7s
General Knowledge 10.1 37.6 62 6.1s
Average 9.8 37.6 102 10.4s

Comparison vs Competitors

Speed Comparison Efficiency Frontier

Model Size Speed Context RAM
dan-omni-3b-mobile 1.2 GB 9.8 tok/s 2K ~1.5 GB
dan-omni-3b 2.0 GB 11.3 tok/s 4K ~2.5 GB
Gemma 4 E2B ~1.4 GB ~35 tok/s 128K ~2 GB
Qwen2.5-1.5B ~0.95 GB ~28 tok/s 32K ~1.5 GB
Command-R7B ~4.0 GB ~8 tok/s 128K ~5 GB

Why dan-omni-3b-mobile? Same 3B quality as dan-omni-3b, but 40% smaller file and 40% less RAM. Ideal for phones with tight memory.

Key Differences from dan-omni-3b

Property dan-omni-3b dan-omni-3b-mobile
File size 2.0 GB 1.2 GB
Context 4096 2048
Speed 11.3 tok/s 9.8 tok/s
RAM usage ~2.5 GB ~1.5 GB
Multimodal Yes No

Dan Omni Model Family

Model Size Speed Use Case
dan-omni-3b 3.5 GB 11.3 tok/s Full multimodal (text + vision)
dan-omni-3b-mobile 1.2 GB 9.8 tok/s Compressed for mobile, 2K context
dan-omni-3b-q3s 1.5 GB 12.8 tok/s Aggressive quantization
dan-omni-smolm2 259 MB 62.2 tok/s Ultralight, fastest
dan-omni-smolm2-v2 259 MB 62.1 tok/s Improved quality variant

Usage

ollama pull sodan/dan-omni-3b-mobile
ollama run sodan/dan-omni-3b-mobile
./llama-cli -m dan-omni-3b-mobile.gguf -p "Hello" --ctx-size 2048

Intended Use

  • Phones with ≤3 GB RAM free
  • On-device inference with strict memory constraints
  • Quick Q&A, translation, short-form tasks
  • Offline AI assistant

License

Apache 2.0

Downloads last month
-
GGUF
Model size
3B params
Architecture
qwen2vl
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support