stablelm-2-1_6b-mlx-16bit

16-bit MLX conversion of stabilityai/stablelm-2-1_6b for Apple Silicon.

Converted by: SirSahOl Source model: stabilityai/stablelm-2-1_6b Framework: MLX by Apple Quantization: 16-bit Format: safetensors License: other


Quick Start

Installation

pip install mlx-lm

CLI Usage

# Chat interactively
mlx_lm.chat --model SirSahOl/stablelm-2-1_6b-chat-mlx-16bit

# Generate text
mlx_lm.generate --model SirSahOl/stablelm-2-1_6b-chat-mlx-16bit --prompt "Your prompt here"

Python Usage

from mlx_lm import load, generate

model, tokenizer = load("SirSahOl/stablelm-2-1_6b-chat-mlx-16bit")
response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256)
print(response)

Performance Benchmarks

| Metric | 4-bit | 8-bit | 16-bit | |--------|--------|--------|--------|| Tokens/sec | 47.84 | 29.85 | 16.53 | | TTFT | 20.9 ms | 33.52 ms | 60.49 ms | | Peak Memory | 915.6 MB | 40.4 MB | 43.9 MB |

Benchmarked on Apple M1 with 8GB unified memory. Average over 5 runs with 256 max tokens.


Who Should Use This?

Your Hardware Recommended Quantization
M1/M2 (8GB) 4-bit โ€” Best balance of quality and memory usage
M1/M2 Pro/Max (16-32GB) 8-bit โ€” Higher quality with reasonable memory
M2/M3/M4 Ultra (64GB+) 16-bit โ€” Full precision, no quality loss

General guidance:

  • Use 4-bit if you want to run this model alongside other applications
  • Use 8-bit if you have the memory and want better quality
  • Use 16-bit for research, evaluation, or if memory isn't a concern

Other Quantization Variants


Conversion Details

Property Value
Source Model stabilityai/stablelm-2-1_6b
Quantization 16-bit
mlx-lm Version 0.31.3
Conversion Time 11.43s
Output Size 3.1 GB
Date 2026-09-10T22:51:53.477974+00:00

Reproduction

To reproduce this conversion:

pip install mlx-lm==0.31.3
python3 -m mlx_lm.convert --hf-path stabilityai/stablelm-2-1_6b --mlx-path output/stablelm-2-1_6b-mlx-16bit

Limitations & Known Issues

  • Performance may degrade with very long contexts (>8K tokens) at lower quantization levels.
  • This is a weight-only conversion; the model architecture and behavior are inherited from the source model.
  • Quantization introduces a small quality loss compared to the original model. Lower bit counts = more loss.
  • This model requires Apple Silicon (M1 or later) to run with MLX.

License

This model conversion inherits the license of the source model: other.

See the original model card for full license details.


Changelog

Version Date Changes
v1.0 2026-09-10 Initial conversion

Converted with MLX Foundry โ€” a professional pipeline for converting models to Apple MLX format.

Downloads last month
311
Safetensors
Model size
2B params
Tensor type
F16
ยท
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for SirSahOl/stablelm-2-1_6b-chat-mlx-16bit

Finetuned
(4)
this model

Collection including SirSahOl/stablelm-2-1_6b-chat-mlx-16bit