Instructions to use ebenezerdon/curious-2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ebenezerdon/curious-2b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ebenezerdon/curious-2b:Q4_0 # Run inference directly in the terminal: llama cli -hf ebenezerdon/curious-2b:Q4_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ebenezerdon/curious-2b:Q4_0 # Run inference directly in the terminal: llama cli -hf ebenezerdon/curious-2b:Q4_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ebenezerdon/curious-2b:Q4_0 # Run inference directly in the terminal: ./llama-cli -hf ebenezerdon/curious-2b:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ebenezerdon/curious-2b:Q4_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf ebenezerdon/curious-2b:Q4_0
Use Docker
docker model run hf.co/ebenezerdon/curious-2b:Q4_0
- LM Studio
- Jan
- vLLM
How to use ebenezerdon/curious-2b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ebenezerdon/curious-2b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ebenezerdon/curious-2b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ebenezerdon/curious-2b:Q4_0
- Ollama
How to use ebenezerdon/curious-2b with Ollama:
ollama run hf.co/ebenezerdon/curious-2b:Q4_0
- Unsloth Desktop
- Pi
How to use ebenezerdon/curious-2b with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ebenezerdon/curious-2b:Q4_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ebenezerdon/curious-2b:Q4_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ebenezerdon/curious-2b with Docker Model Runner:
docker model run hf.co/ebenezerdon/curious-2b:Q4_0
- Lemonade
How to use ebenezerdon/curious-2b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ebenezerdon/curious-2b:Q4_0
Run and chat with the model
lemonade run user.curious-2b-Q4_0
List all available models
lemonade list
- Hermes Agent
How to use ebenezerdon/curious-2b with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ebenezerdon/curious-2b:Q4_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ebenezerdon/curious-2b:Q4_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ebenezerdon/curious-2b with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ebenezerdon/curious-2b:Q4_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ebenezerdon/curious-2b:Q4_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Curious 2B
Curious 2B is the on-device model used by CuriousLM. It is Qwen3.5 2B fine-tuned on the tasks CuriousLM's assistant performs on Android: answering questions from the user's calendar, adding calendar events through a tool call, reading incoming messages and notifications (what a message asks, whether a reply is needed, bills, bookings, deliveries and scams), and writing the assistant's morning and evening notes and reply drafts.
The model runs locally with llama.cpp. CuriousLM supplies the calendar and notification data and runs the tools; nothing is sent to a server.
Releases
| Release | File | Size | SHA-256 |
|---|---|---|---|
curious-2b-261010 (10 October 2026, current) |
curious-2b-261010-Q4_0.gguf |
1.21 GB | e106382634f3ec5319fb7fa5a6d0bd28e8c5fb1d031d8ae29b8e8866fc743f0f |
curious-2b-261009 (9 October 2026) |
curious-2b-261009-Q4_0.gguf |
1.21 GB | 42b4c223434842288c999a447472df33d118169b60eeda55622d0c01866ca333 |
curious-2b-2610 (8 October 2026) |
curious-2b-2610-Q4_0.gguf |
1.21 GB | 45840b88cc63279df6b42c857832068d85c6f9a13745cebcb64fc792af45320d |
Evaluation
Held-out cases, both releases under the same prompts and settings.
Phone and calendar
| 261009 | 261010 | |
|---|---|---|
| Turns passed (665) | 88.0% | 90.1% |
| Requests with times said in words, marked and cleared days, day of the month, people in titles (580) | 58.6% | 79.8% |
| Questions about a past day (150) | 56.7% | 75.3% |
| Reminder and alarm times in English, Spanish, French and German, from MTOP (1,050) | 83.7% | 88.1% |
| Says it cannot read the phone | 1.1% | 0.8% |
| Claims an action that was not taken | 0.5% | 0.3% |
Notification reading
| 261009 | 261010 | |
|---|---|---|
| Standard set (519 tasks) | 96.5% | 98.5% |
| Hard set: receipts, deliveries, friends' plans, scams | 93.5% | 95.2% |
| Real text messages (300) | 53.7% | 78.7% |
| Real scams flagged (of 100) | 94 | 99 |
| Promotions flagged as scams (of 100) | 47 | 21 |
The assistant's writing
| 261009 | 261010 | |
|---|---|---|
| Morning and evening notes, what came in, reply drafts | 54.0% | 61.2% |
| Home briefing leads with what needs the user (120) | 40.8% | 58.3% |
General chat
| 261009 | 261010 | |
|---|---|---|
| Knowledge questions | 93.0% | 94.5% |
| Identifies itself correctly | 94.4% | 98.6% |
| Judged false claims per open answer (105) | 2.6 | 2.9 |
Use
In CuriousLM: Models, then Curious 2B.
With llama.cpp:
llama-server -m curious-2b-261010-Q4_0.gguf --jinja
The tool definitions and prompts the model was tuned on are CuriousLM's own. For general function calling outside the app, Qwen3.5 2B scores higher.
Details
| Base model | Qwen3.5 2B |
| Method | LoRA fine-tune, merged |
| Format | GGUF, Q4_0 |
| Languages | English, Spanish, French, Portuguese, German, Italian, Dutch, Chinese |
Licence
Apache 2.0. Fine-tuned from Qwen3.5-2B, Copyright 2026 Alibaba Cloud, Apache 2.0. Training data includes MASSIVE (Amazon) and When2Call (NVIDIA), both CC BY 4.0; the SMS Phishing Dataset (Mishra and Soni, Mendeley Data, doi:10.17632/f45bkkt8pr.1) and the SMS Spam Collection (Almeida and Hidalgo, UCI), both CC BY 4.0; and spoken date and time phrases from Recognizers-Text (Microsoft, MIT) and Duckling (Meta, BSD 3-Clause).
- Downloads last month
- 11
4-bit