Instructions to use sixstringzen/Hemmingway-1-oQ3e-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use sixstringzen/Hemmingway-1-oQ3e-mtp with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("sixstringzen/Hemmingway-1-oQ3e-mtp") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use sixstringzen/Hemmingway-1-oQ3e-mtp with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sixstringzen/Hemmingway-1-oQ3e-mtp"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "sixstringzen/Hemmingway-1-oQ3e-mtp" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use sixstringzen/Hemmingway-1-oQ3e-mtp with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "sixstringzen/Hemmingway-1-oQ3e-mtp"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "sixstringzen/Hemmingway-1-oQ3e-mtp" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "sixstringzen/Hemmingway-1-oQ3e-mtp", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use sixstringzen/Hemmingway-1-oQ3e-mtp with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sixstringzen/Hemmingway-1-oQ3e-mtp"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default sixstringzen/Hemmingway-1-oQ3e-mtp
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use sixstringzen/Hemmingway-1-oQ3e-mtp with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "sixstringzen/Hemmingway-1-oQ3e-mtp"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "sixstringzen/Hemmingway-1-oQ3e-mtp" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Hemmingway-1 oQ3e with MTP
This repository contains an enhanced oQ3e quantization of Altworld/Hemmingway-1 for MLX and oMLX on Apple silicon. The conversion preserves the model's multi-token prediction (MTP) tensors.
Altworld developed and published the source model. sixstringzen performed this conversion and published the converted weights with their quantization report. The original model, its intended use, and its training details remain documented in the source model card.
Quantization set
This repository is part of the Hemmingway-1 oMLX oQe Quantizations collection. Every build in the set uses the same source revision, group size, non-quantized dtype, calibration pass, and MTP preservation policy.
| Build | Base precision | Output size |
|---|---|---|
| oQ2e | 2-bit | 10.14 GiB |
| oQ3e | 3-bit | 12.22 GiB |
| oQ3.5e | 3-bit with additional higher-precision overrides | 13.19 GiB |
| oQ4e | 4-bit | 15.21 GiB |
| oQ6e | 6-bit | 21.39 GiB |
| oQ8e | 8-bit | 27.10 GiB |
Quantization details
| Item | Value |
|---|---|
| Source model | Altworld/Hemmingway-1 |
| Source revision | 4d711aac0f0043075ae334d2a3de3db3e10135c9 |
| Quantizer | oMLX 0.7.0.dev2 |
| Method | Enhanced oQ3e mixed-precision affine quantization |
| Base precision | 3-bit |
| Group size | 64 |
| Non-quantized dtype | bfloat16 |
| Higher-precision tensors | 8 tensors at 4-bit, 123 at 5-bit, and language_model.lm_head at 6-bit |
| Calibration dataset | oqe_code_multilingual |
| Calibration shape | 128 samples at 512 tokens |
| Imatrix entries | 504 |
| Imatrix cache | Reused from the matching source-model sensitivity pass |
| MTP tensors | 29 preserved tensors |
| Output size | 13,123,999,645 bytes (12.22 GiB) |
oQe uses activation importance to assign additional precision to sensitive tensors. This build uses 3-bit weights as its base, with mixed-precision overrides ranging from 4 to 6 bits. The quantization report records no matrix-shape mismatches and no missing weight shards.
The included oq_imatrix_report.json records the sensitivity pass, calibration settings, tensor coverage, and fallback. Strict imatrix coverage was disabled for the known language_model.lm_head fallback.
Compatibility
This model was created with oMLX 0.7.0.dev2. The source model identifies its text architecture as qwen3_5_text; the converted artifact uses qwen3_5, which matches the architecture name supported by this oMLX build.
The weights use MLX safetensors and are not GGUF files. Compatibility with other MLX runtimes or earlier oMLX releases has not been verified.
Use with oMLX
Download sixstringzen/Hemmingway-1-oQ3e-mtp from the oMLX model browser, then load it as an LLM. Set enable_thinking to false when you want direct prose without visible planning. Runtime defaults and the registered model identifier can vary with the local oMLX installation.
Verification
The finished artifact passed local structural checks on 2026-09-20. It contains three safetensors shards, 1,876 indexed tensors, and 29 MTP tensors. The index references no missing shards.
These checks confirm that the artifact is complete and internally consistent. A generation smoke test has not been recorded for this quantization, and the checks do not establish quality parity with the BF16 source model.
Quality evaluation (v1)
This model is part of the Hemmingway-1 oMLX quantization collection and was compared with a clean BF16 reference in a bounded blind A/B writing study.
Local measurement provenance
Generations were produced by real MLX/oMLX software on an Apple M5 Max with 128 GB unified memory under a deterministic quant-isolation profile. MTP and mixed-runtime speculative accelerators were disabled for the baseline. Captured local manifests, generation records, and oMLX/macmon telemetry are the measurement evidence. Claude monitored or orchestrated a subset of the local tests; Claude is workflow provenance, not the inference engine or measurement source.
Blind-judge result
The study used 14 writing tasks, three judge lanes, and normal/swapped response order. Results below pool the candidate comparisons across Claude Opus 5, Gemini 3.8 Flash, and Grok 4.7:
| Build | Win | Loss | Tie |
|---|---|---|---|
| oQ2e | 47.8% | 51.1% | 1.1% |
| oQ3.5e | 36.7% | 58.9% | 4.4% |
| oQ3e | 55.6% | 43.3% | 1.1% |
| oQ4e | 32.2% | 48.9% | 18.9% |
| oQ6e | 27.8% | 31.1% | 41.1% |
| oQ8e | 31.1% | 23.3% | 45.6% |
The hosted reference row in the local report is a cap-matched subset of 11 prompts, not another quantization condition. Raw inter-rater agreement was 83.7% across 606 pairings. Normal-versus-swapped agreement was 39.6% for Claude, 48.5% for Gemini, and 42.6% for Grok, so the results should be read as subjective and order-sensitive.
Interpretation for this build
Highest point estimate at 55.6%, but inconclusive because of order sensitivity and uncertainty.
This pass does not establish mathematical equivalence, token-level fidelity, or general benchmark superiority. KLD, top-1/top-k agreement, KV-cache divergence, and broader benchmark task scores remain separate pending measurements. The public task and metadata package is the companion benchmark dataset.
Limitations
Quantization can change word choice, coherence, and instruction following. A controlled BF16 comparison has not been published for this build.
The sensitivity pass used oMLX's oqe_code_multilingual calibration dataset. No prose-specific calibration dataset was used. The MTP tensors are present in the artifact, but MTP-assisted decoding has not been benchmarked separately.
The original model's documented limitations and acceptable-use guidance also apply to this quantized release.
License
The source model is released under the Apache 2.0 license. This quantized derivative uses the same license; refer to the source repository for the upstream model card and attribution.
Feedback
Send compatibility reports through this repository's Community tab and include your oMLX version and Apple hardware.
- Downloads last month
- 376
3-bit