Instructions to use whcl412/LycheeAI-coder-1b-II-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use whcl412/LycheeAI-coder-1b-II-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
Use Docker
docker model run hf.co/whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use whcl412/LycheeAI-coder-1b-II-GGUF with Ollama:
ollama run hf.co/whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use whcl412/LycheeAI-coder-1b-II-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use whcl412/LycheeAI-coder-1b-II-GGUF with Docker Model Runner:
docker model run hf.co/whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
- Lemonade
How to use whcl412/LycheeAI-coder-1b-II-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.LycheeAI-coder-1b-II-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use whcl412/LycheeAI-coder-1b-II-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use whcl412/LycheeAI-coder-1b-II-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "whcl412/LycheeAI-coder-1b-II-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
LycheeAI-coder-1b-II-GGUF
当前版本:V3(反谄媚版) GGUF 量化版本 · 基于 MiniCPM5-1B 微调 · Apache 2.0
这是个人开发者的初次尝试,模型能力有限,1B 参数注定它只是个"轻量小助手"。 详细能力说明与局限,请看 主仓库。
📦 三个版本,按设备选
| 文件 | 大小 | 内存占用 | 用在哪 | 质量 |
|---|---|---|---|---|
lychee-coder2-v3-q4_k_m.gguf |
656 MB | ~1 GB | 📱 老手机(4GB 内存)、树莓派 | ⚠️ 有损 |
lychee-coder2-v3-q8_0.gguf |
1.1 GB | ~1.5 GB | 📱 中端手机、轻薄本 | ✅ 接近无损 |
mlx-LycheeAI-coder-1b-II.gguf(f16) |
2.0 GB | ~2.5 GB | 💻 Mac / PC、高端手机 | ✅ 无损 |
原则:设备扛得住就用 f16,空间紧张再往下退。
📱 手机能跑吗?能。
1B 模型的优势就是小。q4 版只有 656MB,一台 4GB 内存的老安卓机就能跑:
速度大约每秒十几个字,完全离线、不联网、不花 API 费用。
📥 怎么下载指定版本
方法 1:命令行(推荐,可只下单个文件)
# HuggingFace(需先 pip install huggingface_hub)
huggingface-cli download whcl412/LycheeAI-coder-1b-II-GGUF \
lychee-coder2-v3-q4_k_m.gguf --local-dir ./
# ModelScope(国内推荐,需先 pip install modelscope)
modelscope download --model whcl412/LycheeAI-coder-1b-II-GGUF \
lychee-coder2-v3-q4_k_m.gguf --local_dir ./
把文件名换成你想要的即可(去掉文件名就是下载整个仓库)。
方法 2:直链下载
# HuggingFace
wget https://huggingface.co/whcl412/LycheeAI-coder-1b-II-GGUF/resolve/main/lychee-coder2-v3-q4_k_m.gguf
# ModelScope
wget https://www.modelscope.cn/models/whcl412/LycheeAI-coder-1b-II-GGUF/resolve/master/lychee-coder2-v3-q4_k_m.gguf
方法 3:网页点下载
打开仓库页面的 Files 标签 → 点文件名 → 右上角 Download。
⚠️ 关于量化的实话
1B 参数本来就小,对量化比大模型敏感得多。实测:
| 问题 | f16(无损) | q8_0 | q4_k_m |
|---|---|---|---|
| 1+1=3 对吗 | 不是,1+1=2 ✅ | 答成 2+2=4 ⚠️ | 答成 2+2=4 ❌ |
| Python 是编译型 | 不是,解释器执行 ✅ | 提到 GIL ⚠️ | 说 GIL 是编译器 ❌ |
量化越低,事实细节越容易出错,但日常聊天、写代码、简单问答依然可用。 要在老手机上跑,q4 是唯一现实的选择——能跑起来比什么都强。
快速开始
Ollama
先准备 Modelfile(重要:TEMPLATE 必须预填 <think>,模型依赖它进入思考模式):
FROM ./lychee-coder2-v3-f16.gguf
SYSTEM """你是 LycheeAI-coder-1b-II,由 MiniCPM5-1B 通过 LoRA 微调而来的编程助手。"""
TEMPLATE """{{ if .System }}<|im_start|>system
{{ .System }}<|im_end|>
{{ end }}<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
<think>
"""
PARAMETER stop <|im_end|>
PARAMETER stop <|im_start|>
PARAMETER temperature 0.4
PARAMETER num_ctx 4096
PARAMETER num_predict 2048
然后:
ollama create lychee-coder2 -f Modelfile
ollama run lychee-coder2
llama.cpp
./llama-cli -m lychee-coder2-v3-f16.gguf \
-p "用 Python 写个判断质数的函数" \
-c 4096 -n 2048 --temp 0.4
关于思考模式
这个模型原生支持思考开关(底层是 enable_thinking 参数,通过 chat_template 控制):
- 开思考(默认):
<think>标签内推理,然后给答案。数学、逻辑题请务必开思考。 - 关思考:直接给答案,更简洁。但数学准确率会明显下降,这是本质权衡。
在 Ollama 里,模型固定走思考模式(Ollama 不支持 enable_thinking 参数)。如果需要直答,请用 MLX 版本并传 enable_thinking=False。
实测踩坑:Ollama 上做"直答变体"(预填空 think 或不预填)都无效——模型仍会自己生成思考内容,而且代码质量反而变差。所以只提供思考版。
推荐参数
| 项目 | 值 |
|---|---|
| 上下文长度 | 4096 |
| 最大输出 | 2048 |
| 温度 | 0.4(写代码可降到 0.2~0.3) |
它做不到什么
- 复杂算法会翻车 —— 语法可能对,逻辑可能错,写完请自己跑。
- 事实细节偶尔胡说 —— 它知道"这话不对",但未必说得清哪儿不对。
- 超纲逻辑题会绕圈 —— 多人真假话推理这类,可能推不出来。
- 关思考时数学会算错 —— 本质权衡,不是 bug。
许可与致谢
- 许可:Apache 2.0
- 基座:MiniCPM5-1B by 面壁智能(OpenBMB)
- 训练:MLX-LM | 量化:llama.cpp
关注作者
📺 B站:space.bilibili.com/3493128967293256
不定期分享 AI 模型训练与折腾记录,欢迎来玩~
- Downloads last month
- 146
Model tree for whcl412/LycheeAI-coder-1b-II-GGUF
Base model
openbmb/MiniCPM5-1B