DAVID โ€” LFM2-1.2B 8-bit MLX (Apple Silicon)

8-bit MLX quantisation of freelion/DAVID-lfm2-1.2b-full
for on-device inference on Apple Silicon Macs.

Latency (benchmarked on Apple M2 Pro)

Metric Value
Median latency 0.57 s
p95 latency 0.98 s
Sub-1s rate 96%
Decode speed 125 tok/s

Usage with MLX

from mlx_lm import load, generate                                                                                                                       
                                                                                                                                                        
model, tokenizer = load("freelion/DAVID-lfm2-1.2b-8bit-mlx")                                                                                            

Or via the browser extension:
DAVID Extension (https://[extension-url]) โ€” supports ChatGPT, Claude, and Gemini.

Full-precision version

freelion/DAVID-lfm2-1.2b-full (https://huggingface.co/freelion/DAVID-lfm2-1.2b-full)

Downloads last month
-
Safetensors
Model size
1B params
Tensor type
U32
ยท
BF16
ยท
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support