Instructions to use miutti/intel-mac-local-llm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use miutti/intel-mac-local-llm with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: llama cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: llama cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: ./llama-cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf miutti/intel-mac-local-llm:UD-Q2_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Use Docker
docker model run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- LM Studio
- Jan
- vLLM
How to use miutti/intel-mac-local-llm with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "miutti/intel-mac-local-llm" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "miutti/intel-mac-local-llm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- Ollama
How to use miutti/intel-mac-local-llm with Ollama:
ollama run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- Unsloth Desktop
- Pi
How to use miutti/intel-mac-local-llm with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "miutti/intel-mac-local-llm:UD-Q2_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use miutti/intel-mac-local-llm with Docker Model Runner:
docker model run hf.co/miutti/intel-mac-local-llm:UD-Q2_K_XL
- Lemonade
How to use miutti/intel-mac-local-llm with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull miutti/intel-mac-local-llm:UD-Q2_K_XL
Run and chat with the model
lemonade run user.intel-mac-local-llm-UD-Q2_K_XL
List all available models
lemonade list
- Hermes Agent
How to use miutti/intel-mac-local-llm with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default miutti/intel-mac-local-llm:UD-Q2_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use miutti/intel-mac-local-llm with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf miutti/intel-mac-local-llm:UD-Q2_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "miutti/intel-mac-local-llm:UD-Q2_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Intel Mac Local LLM — 古い Mac で動く「育つ AI」
16GB・GPU なしの Intel Mac で、手元の AI に Mac の仕事(ファイルの整理・読み書き・状態の確認・調べもの)をさせるための 頭脳(重み) と 仕組み(カーネル・デスクトップアプリ) 一式です。合言葉は「早い・安い・賢い」。
- GitHub(同じ中身・最新): https://github.com/vhpctnyd5x-lab/intel-mac-local-llm
- このページの
source/は、その公開リポジトリの写しです。
English summary: An expert-pruned GGUF of Qwen3.6-35B-A3B (256 → 160 experts per layer, 8.29 GB) that fits a 16 GB Intel Mac without GPU, plus the agent kernel (sandboxed tools, approval gate, self-learning) and the macOS desktop app source. Measured on an i7-9750H / 16 GB.
頭脳: Qwen3.6-35B-A3B-UD-Q2_K_XL-k160.gguf(8.29GB / 7.7GiB)
| 元(256人) | この版(160人) | |
|---|---|---|
| 大きさ | 12.57GB(11.7GiB) | 8.29GB(7.7GiB) |
| 常駐メモリ(実測) | 9.6〜10.0GB | 8.0〜8.2GB |
| 書く速さ(300字×3回) | 8.3〜8.6 字/秒 | 8.1〜9.1 字/秒(変わらない) |
| 知識の問い 25問 | 24/25 | 23/25 |
| PC 操作の回帰テスト 41問 | — | 36〜39/41(回ごとに±2問ぶれる) |
すべて Intel Core i7-9750H・16GB・GPU なしで測った数字です。1回の数字は1回の数字として見てください。
作り方: 元は unsloth/Qwen3.6-35B-A3B-GGUF の UD-Q2_K_XL。
校正文で imatrix を取り、層ごとに使われる順で専門家を160人だけ残し、MTP 層も外しました(手順は source/dougu/kezuru_tejun.py)。
重みの値そのものは変えていません(量子化は元のまま)。
動かし方: 動作を確かめたのは、llama.cpp に source/llama_patch/ を当てて建てた llama-server です
(語彙を 99.99% に絞る改造が入っています)。手順は source/README.md。素の llama.cpp で読み込めるかは確かめていません。
llama-server -m Qwen3.6-35B-A3B-UD-Q2_K_XL-k160.gguf -t 6 -ngl 0 -c 32768 -np 2 -kvu --no-cache-idle-slots \
-cb -cram 512 --cache-reuse 16 -fa off --reasoning-format none
仕組み(source/)
- カーネル(
source/kernel/): 頭脳が呼ぶ道具と、危ない操作を止める門番。命令は macOS のsandbox-exec(Linux は bubblewrap)で隔離し、 消す・送るなどは本人の承認が要ります。会話ごとに「触ってよいフォルダ」を選べます。 - デスクトップアプリ:
source/kernel/tools/kernel_window.swift(窓)とsource/kernel/ui.html(画面)。会話・動きの欄・文脈の量・/compact・/loop・/整理。 - 事前学習: 空いた時間に Wikipedia・教科書・文学・法令・論文を読み、知識の箱と学習ノートに貯めて答えに使います。興味が出た事を重点的に読みます。
- 育つ AI: 芯・感情の記憶・日記・相手の記憶・自作スキル(要る/要らないを自分で決める)。
ライセンス
重みは元と同じ Apache License 2.0(Qwen3.6 の著作権表示は元の模型に従います。変更点は上の「作り方」)。
コード(source/)は MIT License(source/LICENSE)。
- Downloads last month
- 53
2-bit
Model tree for miutti/intel-mac-local-llm
Base model
Qwen/Qwen3.6-35B-A3B