Phi-4-mini-instruct-mlx-16bit

16-bit MLX conversion of microsoft/Phi-4-mini-instruct for Apple Silicon.

Converted by: SirSahOl Source model: microsoft/Phi-4-mini-instruct Framework: MLX by Apple Quantization: 16-bit Format: safetensors License: mit


Quick Start

Installation

pip install mlx-lm

CLI Usage

# Chat interactively
mlx_lm.chat --model SirSahOl/Phi-4-mini-instruct-chat-mlx-16bit

# Generate text
mlx_lm.generate --model SirSahOl/Phi-4-mini-instruct-chat-mlx-16bit --prompt "Your prompt here"

Python Usage

from mlx_lm import load, generate

model, tokenizer = load("SirSahOl/Phi-4-mini-instruct-chat-mlx-16bit")
response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256)
print(response)

Performance Benchmarks

Benchmarks coming soon.


Who Should Use This?

Your Hardware Recommended Quantization
M1/M2 (8GB) 4-bit โ€” Best balance of quality and memory usage
M1/M2 Pro/Max (16-32GB) 8-bit โ€” Higher quality with reasonable memory
M2/M3/M4 Ultra (64GB+) 16-bit โ€” Full precision, no quality loss

General guidance:

  • Use 4-bit if you want to run this model alongside other applications
  • Use 8-bit if you have the memory and want better quality
  • Use 16-bit for research, evaluation, or if memory isn't a concern

Other Quantization Variants


Conversion Details


Limitations & Known Issues

  • Performance may degrade with very long contexts (>8K tokens) at lower quantization levels.
  • This is a weight-only conversion; the model architecture and behavior are inherited from the source model.
  • Quantization introduces a small quality loss compared to the original model. Lower bit counts = more loss.
  • This model requires Apple Silicon (M1 or later) to run with MLX.

License

This model conversion inherits the license of the source model: mit.

See the original model card for full license details.


Changelog

Version Date Changes
v1.0 N/A Initial conversion

Converted with MLX Foundry โ€” a professional pipeline for converting models to Apple MLX format.

Downloads last month
291
Safetensors
Model size
4B params
Tensor type
BF16
ยท
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for SirSahOl/Phi-4-mini-instruct-chat-mlx-16bit

Finetuned
(120)
this model

Collection including SirSahOl/Phi-4-mini-instruct-chat-mlx-16bit