Instructions to use leonx1995/docudis-intent-gemma4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use leonx1995/docudis-intent-gemma4 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf leonx1995/docudis-intent-gemma4 # Run inference directly in the terminal: llama cli -hf leonx1995/docudis-intent-gemma4
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf leonx1995/docudis-intent-gemma4 # Run inference directly in the terminal: llama cli -hf leonx1995/docudis-intent-gemma4
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf leonx1995/docudis-intent-gemma4 # Run inference directly in the terminal: ./llama-cli -hf leonx1995/docudis-intent-gemma4
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf leonx1995/docudis-intent-gemma4 # Run inference directly in the terminal: ./build/bin/llama-cli -hf leonx1995/docudis-intent-gemma4
Use Docker
docker model run hf.co/leonx1995/docudis-intent-gemma4
- LM Studio
- Jan
- vLLM
How to use leonx1995/docudis-intent-gemma4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "leonx1995/docudis-intent-gemma4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "leonx1995/docudis-intent-gemma4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/leonx1995/docudis-intent-gemma4
- Ollama
How to use leonx1995/docudis-intent-gemma4 with Ollama:
ollama run hf.co/leonx1995/docudis-intent-gemma4
- Unsloth Desktop
- Pi
How to use leonx1995/docudis-intent-gemma4 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf leonx1995/docudis-intent-gemma4
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "leonx1995/docudis-intent-gemma4" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use leonx1995/docudis-intent-gemma4 with Docker Model Runner:
docker model run hf.co/leonx1995/docudis-intent-gemma4
- Lemonade
How to use leonx1995/docudis-intent-gemma4 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull leonx1995/docudis-intent-gemma4
Run and chat with the model
lemonade run user.docudis-intent-gemma4-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use leonx1995/docudis-intent-gemma4 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf leonx1995/docudis-intent-gemma4
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default leonx1995/docudis-intent-gemma4
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use leonx1995/docudis-intent-gemma4 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf leonx1995/docudis-intent-gemma4
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "leonx1995/docudis-intent-gemma4" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Docudis intent (docudis-intent-gemma4)
The instruction-reading model of Docudis, which replaces personal details in documents with labels on the user's device. The user says in their own words what to hide or keep ("hide names and phone numbers, the amounts can stay, it's a French payslip"); this model turns that into a small JSON object that Docudis maps onto its detection settings. It reads only the instruction, never the document.
这是法国的工资单,遮住名字和电话,金额留着
→ {"types":{"PERSON":"hide","PHONE":"hide","AMOUNT":"keep"},"regions":["fr"],"verticals":["employment"]}
- Format: GGUF for llama.cpp,
intent-mixq8.gguf(3.6 GB): Q8_0, with the per-layer embeddings at Q4_K and the token embeddings at Q8_0. - Also here:
system_prompt.txt(the one line the model was trained with),intent.gbnf(the grammar for constrained decoding) andmodel.json(run settings). - Code: the output format, post-processing, evaluation and training pipeline, at the commit this model was built from, are in docudis-ner/intent (Apache-2.0).
Output
One JSON object, every field optional; {} means "use the defaults".
| Field | Meaning |
|---|---|
types |
Entity type (PERSON EMAIL PHONE ID NUMBER CARD IBAN DATE BIRTH_DATE AMOUNT IP URL ADDRESS COMPANY SECRET API_KEY, or * for every other type) to hide, keep (detect, leave visible) or off (do not detect) |
regions |
Country rule packs: at be ch cn de dk es fi fr gb ie it jp nl no pl pt se us |
verticals |
Document domain: healthcare legal finance employment insurance technology utilities |
dictionary |
Literal terms to hide, copied from the instruction |
never_hide |
Literal terms to leave visible, copied from the instruction |
unsupported |
true when part of the request cannot be done (fake names, partial masking, translation, keeping only some values of a type: one person's name, "my employer's name", "the hospital", "the rent"); the model has then hidden the whole type, the safe side |
The full rules are in intent/spec.md.
How to run it
llama-server -m intent-mixq8.gguf -ngl 99 -c 2048 --jinja --reasoning off --reasoning-budget 0
Send system_prompt.txt as the system message and the user's instruction as the user message, with
temperature 0 and the contents of intent.gbnf as the request's grammar. Do not use
response_format: llama-server then changes the prompt and the model starts with a reasoning block.
--reasoning off is needed for the same reason.
The reply is one JSON object. Drop empty lists and "unsupported": false, validate it against
intent/schema.json (use
{} if it does not validate), then run the host post-processing in
intent/postprocess.py:
it drops keep, off and * that the instruction does not support (prompt injection, "do whatever
with the rest"), drops literal terms that are not in the instruction, and adds regions and verticals
from a keyword list. The scores below include it.
How it was made
LoRA fine-tune of google/gemma-4-E2B-it with Unsloth: base loaded in 4 bits (QLoRA), rank 16, alpha 16, no dropout, on the language model's attention and MLP layers; 5 epochs, learning rate 2e-4 with a cosine schedule and 5% warm-up, weight decay 0.01, effective batch 16, sequences up to 512 tokens, loss on the reply only, seed 2. The adapter was merged into the 16-bit base and converted with llama.cpp b11379.
The training data is 1,309 instruction → JSON pairs (145 more held out for validation) in Chinese, English, French, Spanish, German, Italian and mixed-language messages. They are synthetic: written with an LLM (Claude) against the spec, from short commands to long, informal or misspelled messages and attempts to override the model, then reviewed and corrected case by case by a second LLM pass. They contain no real personal data; every name and identifier is invented. Cases close to any evaluation instruction were removed before training. The data is in intent/train/raw.
Evaluation
Exact match of the whole JSON object, temperature 0, constrained decoding, with the host post-processing:
| Set | Cases | Exact match | Leaks |
|---|---|---|---|
| Test, frozen: written independently of the training data and never used to tune it | 300 | 90.3% | 7 (2.3%) |
| Dev, used to steer the training data | 272 | 90.4% | 4 |
A leak is a miss that leaves visible something the user wanted hidden. The four on dev: "they only need the rent amount and the dates" read as keeping all amounts, a named hospital kept along with every company, "the rest whatever" read as keeping every other type, and one reply with a repeated key (see Limitations). Without post-processing the test score is 87.7%; most of that difference is regions and verticals the model leaves out. Two training seeds on the same data differ by about one point (dev 89.3% and 90.4%). For comparison, the base model prompted with the full spec and six examples scored 43.5% on the first 200 dev cases.
Both sets are scored on the labels as revised on 2026-10-05: a keep narrowed to some values of a type ("my employer's name", "the rent") counts as unsupported, not as keeping the whole type. The previous release scores 87.3% with 13 leaks on test under these labels (89.7% with 4 under the old ones).
Speed for one instruction (output 15–30 tokens): about 0.12 s on an RTX 3080, about 3 s on a desktop CPU (Intel i7-13700KF, no GPU).
Limitations
- It reads instructions in the six languages above; other languages are untested.
- Docudis cannot tell one value of a type from another (a doctor's name from a patient's, the rent
from the deposit). Such requests come back as the whole type hidden plus
unsupported, and occasionally as the whole type kept: most of the leaks above. A host should show the user what will be kept visible before running. - The grammar does not stop a key from repeating, and the model has written replies such as
{"types":{"*":"off",…,"*":"hide"}}. JSON parsers disagree on which value wins, so a host should treat a reply with a repeated key as invalid. - PERSON and COMPANY are sometimes confused for roles like "landlord" or "employer", and
API_KEYsometimes comes back asSECRET. - The training and test instructions are synthetic; real users phrase things differently, and the scores on their messages may be lower.
- It depends on the host post-processing for prompt-injection and keyword checks. Used alone, it can follow an instruction that dictates its JSON.
Versions
| Revision | Date | Test | Change |
|---|---|---|---|
e25d293 |
2026-10-05 | 90.3%, 7 leaks | A keep of only part of a type ("my employer's name", "the rent") is now unsupported; 41 new training instructions; 5 epochs |
f73e207 |
2026-10-04 | 87.3%, 13 leaks (89.7%, 4 leaks on the old answers) | First release |
Load an older version with its revision, for example hf_hub_download(..., revision="f73e207").
The full history, with the reason for each change, is in
intent/HISTORY.md.
License
Copyright 2026 Pengda Xia (stonetech). Licensed under the Apache License 2.0.
This model is a derivative of google/gemma-4-E2B-it (Apache License 2.0, Copyright Google LLC), modified by stonetech by fine-tuning it on the data above, merging the adapter and quantizing it.
- Downloads last month
- 252
We're not able to determine the quantization variants.