Instructions to use NGDtuanh/abook-analyzer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use NGDtuanh/abook-analyzer with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf NGDtuanh/abook-analyzer:Q8_0 # Run inference directly in the terminal: llama cli -hf NGDtuanh/abook-analyzer:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf NGDtuanh/abook-analyzer:Q8_0 # Run inference directly in the terminal: llama cli -hf NGDtuanh/abook-analyzer:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf NGDtuanh/abook-analyzer:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf NGDtuanh/abook-analyzer:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf NGDtuanh/abook-analyzer:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf NGDtuanh/abook-analyzer:Q8_0
Use Docker
docker model run hf.co/NGDtuanh/abook-analyzer:Q8_0
- LM Studio
- Jan
- vLLM
How to use NGDtuanh/abook-analyzer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "NGDtuanh/abook-analyzer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "NGDtuanh/abook-analyzer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/NGDtuanh/abook-analyzer:Q8_0
- Ollama
How to use NGDtuanh/abook-analyzer with Ollama:
ollama run hf.co/NGDtuanh/abook-analyzer:Q8_0
- Unsloth Desktop
- Pi
How to use NGDtuanh/abook-analyzer with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf NGDtuanh/abook-analyzer:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "NGDtuanh/abook-analyzer:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use NGDtuanh/abook-analyzer with Docker Model Runner:
docker model run hf.co/NGDtuanh/abook-analyzer:Q8_0
- Lemonade
How to use NGDtuanh/abook-analyzer with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull NGDtuanh/abook-analyzer:Q8_0
Run and chat with the model
lemonade run user.abook-analyzer-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use NGDtuanh/abook-analyzer with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf NGDtuanh/abook-analyzer:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default NGDtuanh/abook-analyzer:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use NGDtuanh/abook-analyzer with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf NGDtuanh/abook-analyzer:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "NGDtuanh/abook-analyzer:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
abook-analyzer:v3
Vietnamese audiobook analysis model for ABook: labels every segment of a Vietnamese novel with who speaks, narration / dialogue / thought, emotion, intensity, pace, volume and gender, as JSON under ABook's own prompt and schema. Qwen3-4B-Instruct-2507 + LoRA (r=16), merged, GGUF Q8_0 (4.28 GB). Apache-2.0.
Model phân tích truyện của dây chuyền làm sách nói ABook: với mỗi đoạn văn, ai nói (-> giọng đọc), đoạn kể / thoại / nội tâm, cảm xúc, cường độ, nhịp, âm lượng, giới tính. Chạy được trên card NVIDIA 8 GB, nhanh hơn qwen3:8b 1,5-1,7 lần.
Dùng
Studio của ABook tự tải khi cài hoặc bấm "Cập nhật Studio" - không phải làm gì.
Dùng tay: tải abook-analyzer-v3.Q8_0.gguf và Modelfile vào cùng một thư mục, kiểm SHA-256
(9545ce0bf921f3b771a796272b736d043dc4082cb14d1a51fcc93c55e4c53778), rồi:
ollama create abook-analyzer:v3 -f Modelfile
Đã đo với Ollama 0.33.2. Ollama 0.34 cho dòng qwen3 "suy nghĩ" trước khi trả JSON kể cả khi request có format: gửi
"think": false như ABook.
Model chỉ học prompt và schema của ABook (ebook_reader/analysis.py); hỏi kiểu khác thì không có gì bảo đảm.
Huấn luyện
- Nền:
Qwen/Qwen3-4B-Instruct-2507(Apache-2.0). - LoRA r=16, alpha 32, dropout 0,05 trên mọi lớp chiếu (q, k, v, o, gate, up, down); 1 epoch; gộp vào model nền rồi
xuất GGUF Q8_0 (
scripts/model_eval/train_lora.py,serve_lora.py). - Loss chỉ tính trên câu trả lời (
assistant_only_loss): model học gán nhãn, không học sinh lại văn bản truyện. - Dữ liệu: nhãn đáp án chuẩn do chính dự án làm (
scripts/model_eval/gold/, quy ước trongdocs/GOLD_GUIDE.md), phát lại qua đúng bộ phân tích sản xuất (gold_replay.py->build_training_set.py). Không dùng bộ dữ liệu ngoài nào. Văn bản truyện không được phát hành. - Chương dùng để đo không bao giờ vào dữ liệu huấn luyện.
Kết quả đo
Thước chính: F1 giọng B-cubed (người nghe nghe thấy đúng một giọng cho một người) / người nói đúng tuyệt đối. Cùng máy,
cùng mã host, khởi đầu lạnh. Chi tiết: docs/ANALYSIS_RESEARCH.md.
| bộ đo | câu | qwen3:8b gốc | abook-analyzer:v3 |
|---|---|---|---|
| Light novel Nhật/Hàn, 6 truyện, chương chưa học | 519 | 52,9 / 54,9 | 55,7 / 62,4 |
| Young Master's PoV 248 (Hàn), ngôi thứ nhất | 47 | 56,1 / 61,7 | 85,5 / 93,6 |
| Tam quốc diễn nghĩa, hồi 50-52 | 219 | 74,8 / 79,5 | 85,6 / 89,5 |
| Tắt đèn XX, XXI, XXIV | 112 | 72,7 / 75,9 | 73,9 / 82,1 |
| Throne of Magical Arcana, 4 chương đo | 156 | 55,8 / 65,4 | 60,1 / 69,9 |
Lỗi còn nhiều nhất: truyện ngôi thứ nhất có hai người cùng tên gọi quen, lời trong 『』 (linh thể, giọng trong đầu), và chuỗi thoại hai người không lời dẫn bị lệch một nhịp.
Giấy phép
Apache-2.0, như model nền Qwen3-4B-Instruct-2507.
- Downloads last month
- 29
8-bit
Model tree for NGDtuanh/abook-analyzer
Base model
Qwen/Qwen3-4B-Instruct-2507