MedPsy-4B MLX 4-bit

This repository contains an MLX 4-bit conversion of qvac/MedPsy-4B, optimized for inference on Apple silicon.

MedPsy-4B is a text-only medical and healthcare reasoning model built on Qwen3-4B-Thinking-2507 and post-trained by QVAC using supervised fine-tuning and reinforcement learning.

Quantization

  • Format: MLX
  • Quantization mode: affine
  • Bits: 4
  • Group size: 64
  • Effective bits per weight: 4.501
  • Approximate model directory size: 2.1 GB
  • Conversion tool: MLX-LM 0.31.3
  • MLX version: 0.32.1

Available MLX variants

Variant Approximate size Local generation speed Peak memory
4-bit 2.1 GB 47.054 tok/s 2.468 GB
6-bit 3.1 GB 35.439 tok/s 3.473 GB
8-bit 4.0 GB 28.149 tok/s 4.459 GB

Installation

pip install -U mlx-lm

Command-line usage

mlx_lm.generate \
  --model Irfanuruchi/MedPsy-4B-MLX-4bit \
  --prompt "Explain the difference between sensitivity and specificity." \
  --max-tokens 1024 \
  --temp 0

Python usage

from mlx_lm import load, generate

model, tokenizer = load("Irfanuruchi/MedPsy-4B-MLX-4bit")

messages = [
    {
        "role": "user",
        "content": "Explain the difference between sensitivity and specificity.",
    }
]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

response = generate(
    model,
    tokenizer,
    prompt=prompt,
    max_tokens=1024,
)

print(response)

MedPsy may emit a <think>...</think> reasoning section before its final answer. This behavior is inherited from the source model.

Local validation

Validated using a deterministic medical-domain prompt on an Apple M3 Pro MacBook Pro with 18 GB unified memory.

  • Python: 3.12.14
  • MLX: 0.32.1
  • MLX-LM: 0.31.3
  • Prompt tokens: 48
  • Generated tokens: 356
  • Prompt processing: 53.435 tokens/second
  • Generation: 47.054 tokens/second
  • Peak unified memory: 2.468 GB
  • Natural EOS termination: passed
  • Exactly two requested final bullets: passed
  • Coherent medical-domain generation: passed

These figures represent one local inference run and are not clinical-quality or benchmark evaluations. Performance varies by device, operating conditions, prompt length, and software version.

Important medical limitation

This model is not a medical device and is not a substitute for professional medical judgment, diagnosis, or treatment. It can produce incorrect, incomplete, or misleading outputs that appear authoritative. Medical outputs must be independently reviewed by appropriately qualified professionals.

Source evaluation

Benchmark results reported by QVAC belong to the source model and were not independently reproduced for this quantized conversion. See the source model card and MedPsy research overview.

License and attribution

The source repository identifies MedPsy-4B under the Apache 2.0 license for research and educational use. The original LICENSE and ATTRIBUTIONS.md files are included in this repository. Users should review those files and the source model card before redistribution or deployment.

Downloads last month
10
Safetensors
Model size
0.6B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Irfanuruchi/MedPsy-4B-MLX-4bit

Finetuned
qvac/MedPsy-4B
Quantized
(6)
this model