Instructions to use Irfanuruchi/MedPsy-4B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Irfanuruchi/MedPsy-4B-MLX-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Irfanuruchi/MedPsy-4B-MLX-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Irfanuruchi/MedPsy-4B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Irfanuruchi/MedPsy-4B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Irfanuruchi/MedPsy-4B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Irfanuruchi/MedPsy-4B-MLX-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Irfanuruchi/MedPsy-4B-MLX-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Irfanuruchi/MedPsy-4B-MLX-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Irfanuruchi/MedPsy-4B-MLX-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Irfanuruchi/MedPsy-4B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Irfanuruchi/MedPsy-4B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Irfanuruchi/MedPsy-4B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Irfanuruchi/MedPsy-4B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Irfanuruchi/MedPsy-4B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Irfanuruchi/MedPsy-4B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
MedPsy-4B MLX 4-bit
This repository contains an MLX 4-bit conversion of
qvac/MedPsy-4B, optimized for
inference on Apple silicon.
MedPsy-4B is a text-only medical and healthcare reasoning model built on Qwen3-4B-Thinking-2507 and post-trained by QVAC using supervised fine-tuning and reinforcement learning.
Quantization
- Format: MLX
- Quantization mode: affine
- Bits: 4
- Group size: 64
- Effective bits per weight: 4.501
- Approximate model directory size: 2.1 GB
- Conversion tool: MLX-LM 0.31.3
- MLX version: 0.32.1
Available MLX variants
| Variant | Approximate size | Local generation speed | Peak memory |
|---|---|---|---|
| 4-bit | 2.1 GB | 47.054 tok/s | 2.468 GB |
| 6-bit | 3.1 GB | 35.439 tok/s | 3.473 GB |
| 8-bit | 4.0 GB | 28.149 tok/s | 4.459 GB |
Installation
pip install -U mlx-lm
Command-line usage
mlx_lm.generate \
--model Irfanuruchi/MedPsy-4B-MLX-4bit \
--prompt "Explain the difference between sensitivity and specificity." \
--max-tokens 1024 \
--temp 0
Python usage
from mlx_lm import load, generate
model, tokenizer = load("Irfanuruchi/MedPsy-4B-MLX-4bit")
messages = [
{
"role": "user",
"content": "Explain the difference between sensitivity and specificity.",
}
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
response = generate(
model,
tokenizer,
prompt=prompt,
max_tokens=1024,
)
print(response)
MedPsy may emit a <think>...</think> reasoning section before its final
answer. This behavior is inherited from the source model.
Local validation
Validated using a deterministic medical-domain prompt on an Apple M3 Pro MacBook Pro with 18 GB unified memory.
- Python: 3.12.14
- MLX: 0.32.1
- MLX-LM: 0.31.3
- Prompt tokens: 48
- Generated tokens: 356
- Prompt processing: 53.435 tokens/second
- Generation: 47.054 tokens/second
- Peak unified memory: 2.468 GB
- Natural EOS termination: passed
- Exactly two requested final bullets: passed
- Coherent medical-domain generation: passed
These figures represent one local inference run and are not clinical-quality or benchmark evaluations. Performance varies by device, operating conditions, prompt length, and software version.
Important medical limitation
This model is not a medical device and is not a substitute for professional medical judgment, diagnosis, or treatment. It can produce incorrect, incomplete, or misleading outputs that appear authoritative. Medical outputs must be independently reviewed by appropriately qualified professionals.
Source evaluation
Benchmark results reported by QVAC belong to the source model and were not independently reproduced for this quantized conversion. See the source model card and MedPsy research overview.
License and attribution
The source repository identifies MedPsy-4B under the Apache 2.0 license for
research and educational use. The original LICENSE and ATTRIBUTIONS.md
files are included in this repository. Users should review those files and
the source model card before redistribution or deployment.
- Downloads last month
- 10
4-bit