Hi, I'm Saheed Olaide (@SirSahOl) π
I build and optimize AI models for local edge inference on Apple Silicon, specializing in Apple MLX quantization, benchmarking, and deployment tooling.
π Featured Collection
Check out my curated collection of native MLX conversions: π MLX Models β Optimized for Apple Silicon
π οΈ Open Source Tooling
- mlx-foundry: A professional CLI pipeline for converting, benchmarking, and publishing HuggingFace models to Apple MLX format with publication-grade model cards.
π Models by SirSahOl
Qwen3 Series (Native MLX Conversions)
- Qwen3-0.6B-chat-mlx-4bit β 109.6 tok/s on M1 Apple Silicon (~550 MB RAM)
- Qwen3-0.6B-chat-mlx-8bit β 68.3 tok/s on M1 Apple Silicon
- Qwen3-0.6B-chat-mlx-16bit β 39.9 tok/s on M1 Apple Silicon
GLM Edge Series
- GLM-Edge-1.5B-Chat-mlx-16bit
- GLM-Edge-1.5B-Chat-mlx-8bit
- GLM-Edge-1.5B-Chat-mlx-4bit
- GLM-Edge-4B-Chat-mlx-4bit
- GLM-Edge-4B-Chat-mlx-8bit
π» Hardware & Benchmarking Methodology
All models are verified and benchmarked on Apple M1 (8GB unified memory) and profiled for:
- Token generation throughput (
tok/s) - Time to first token (
TTFT) - Peak unified memory usage (
RSS MB)
π¬ Connect
- GitHub: github.com/sirsahol
- HuggingFace: huggingface.co/SirSahOl
Inference Providers NEW
This model isn't deployed by any Inference Provider. π Ask for provider support