GLM-Edge-1.5B-Chat (8-bit MLX Quantized)

This repository contains Zhipu AI's GLM-Edge-1.5B-Chat quantized to 8-bit precision. It is compiled natively for Apple Silicon under the MLX framework.

8-bit quantization offers a practically lossless alternative to the unquantized model. It balances speed with strict logical precision.

Performance Benchmarks

  • Inference Speed: ~48.50 tokens per second (base M1 Apple Silicon)
  • VRAM Footprint: ~1.56 GB
  • Memory Efficiency: Low overhead allows side-by-side execution with memory-heavy developer tools.

Installation

Install the MLX LM package.

pip install mlx-lm

Usage

Command Line Interface

Chat with the model in your terminal.

mlx_lm.chat --model SirSahOl/glm-edge-1.5b-chat-mlx-8bit

Python API

Load and generate text programmatically.

from mlx_lm import load, generate

model, tokenizer = load("SirSahOl/glm-edge-1.5b-chat-mlx-8bit")

messages = [{"role": "user", "content": "Explain quantum superposition."}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)

response = generate(model, tokenizer, prompt=prompt, verbose=True)

Multi-Quantization Comparison

Evaluate your hardware budget and choose the optimal precision:

Variant Disk Size VRAM Footprint M1 Speed Key Advantage
4-bit MLX ~800 MB ~0.87 GB ~72.5 tokens/sec Maximum speed, lowest RAM.
8-bit MLX (This Repo) ~1.56 GB ~1.56 GB ~48.5 tokens/sec Lossless balance, highly stable reasoning.
16-bit MLX ~2.94 GB ~3.00 GB ~32.5 tokens/sec Raw full-precision, absolute peak quality.

Limitations

  • Requires slightly more memory than the 4-bit variant. Recommended for users who prioritize logical accuracy over maximum token generation speed.

LM Studio Configuration (Universal Preset Fix)

If you load this model in LM Studio, you must configure custom Stop Strings to prevent the model from entering an infinite self-dialogue loop.

Option A: Automatic Preset (Recommended)

You can create a custom prompt preset to configure all settings automatically. Create a JSON file named GLM-Edge.json inside your LM Studio config directory:

  • macOS / Linux: ~/.lmstudio/config-presets/GLM-Edge.json
  • Windows: %USERPROFILE%\.lmstudio\config-presets\GLM-Edge.json

Add the following JSON content:

{
  "name": "GLM-Edge",
  "inference_params": {
    "pre_prompt": "You are a helpful, direct, and honest assistant.",
    "input_prefix": "<|user|>\n",
    "input_suffix": "\n<|assistant|>\n",
    "pre_prompt_prefix": "<|system|>\n",
    "pre_prompt_suffix": "\n",
    "antiprompt": [
      "<|user|>",
      "<|observation|>",
      "<|endoftext|>"
    ],
    "stopStrings": [
      "<|user|>",
      "<|observation|>",
      "<|endoftext|>"
    ],
    "temperature": 0.7,
    "max_tokens": 2048
  }
}

Restart LM Studio, open a Chat session, and select "GLM-Edge" from the Prompt Template dropdown.

Option B: Manual Configuration

Alternatively, configure the settings manually in the Advanced Configuration sidebar:

  1. Stop Strings (Antiprompts / stopStrings): Add <|user|>, <|observation|>, and <|endoftext|>
  2. Prompt Formatting:
    • User Prefix: <|user|>\n
    • Assistant Suffix: \n<|assistant|>\n
    • System Prefix: <|system|>\n
    • System Suffix: \n
Downloads last month
66
Safetensors
Model size
0.4B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SirSahOl/glm-edge-1.5b-chat-mlx-8bit

Quantized
(12)
this model