Instructions to use Ammonix/AmmonixRCM-Writer-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Ammonix/AmmonixRCM-Writer-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Ammonix/AmmonixRCM-Writer-9B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Ammonix/AmmonixRCM-Writer-9B") model = AutoModelForMultimodalLM.from_pretrained("Ammonix/AmmonixRCM-Writer-9B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Ammonix/AmmonixRCM-Writer-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Ammonix/AmmonixRCM-Writer-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ammonix/AmmonixRCM-Writer-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Ammonix/AmmonixRCM-Writer-9B
- SGLang
How to use Ammonix/AmmonixRCM-Writer-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Ammonix/AmmonixRCM-Writer-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ammonix/AmmonixRCM-Writer-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Ammonix/AmmonixRCM-Writer-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Ammonix/AmmonixRCM-Writer-9B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Ammonix/AmmonixRCM-Writer-9B with Docker Model Runner:
docker model run hf.co/Ammonix/AmmonixRCM-Writer-9B
AmmonixRCM-Writer-9B
The frozen local language model of the Ammonix RCM Agent: it writes the claim paperwork after each decision. The decision path makes no language-model call — this model only writes, it never decides.
This is the exact artifact evaluated in the RCM paper (https://doi.org/10.5281/zenodo.22871212), a LoRA (r=32) trained on 1,339 self-generated pairs rewarded by the payer's verdict — no teacher model — merged into the base weights. Under the paper's payer, this trained writer loses the least money of all four writers tested (the writer comparison table).
Evaluation
Four writers were swapped into the same seat under identical decisions (200 demonstration claims, 525 decision points; the payer's reviewer reads every appeal letter against the records it holds — a letter with an unsupported statement comes back unprocessed):
| Writer | Rejected by payer | Money on rejected claims |
|---|---|---|
| Qwen 27B | 48 | $27,039 |
| Ornith 1.5 9B (base) | 44 | $23,933 |
| This model | 16 | $7,709 |
| Claude Opus 5 | 37 | $21,157 |
Decisions were identical under all four writers (525/525) — the agent decides from recorded outcomes; the writer only affects whether the paperwork survives the payer's reading.
Training
- Data: fully synthetic. Every training pair comes from the Cardessa synthetic claims world; no real patient, provider, or payer data was used anywhere in training or evaluation.
- Rejection-sampling SFT with no teacher model: the base 9B drafts each claim's paperwork 8× at temperature 0.8; each draft is scored by the world's own payer, and the best draft per claim becomes a training pair (1,503 claims → 1,339 pairs).
- LoRA r=32, 2 epochs, ~3 h on one consumer GPU, then merged into the base weights. The pairs and training script ship in the code repository.
Limitations and out-of-scope use
- Trained and evaluated only on the synthetic world; it has never seen a real claim, real payer rules, or real denial codes. Its writing style is fitted to one simulated payer's reviewer.
- It is a writer, not a decision-maker: outside the agent's harness (schema-constrained decoding, scrubber, escalation rules) its outputs are unchecked.
- Not for production billing, clinical use, or any submission to a real payer.
License
Released under the Ammonix Research License (LICENSE.md): research, educational, and
evaluation use is free — including evaluation by a commercial organization deciding whether
to seek a commercial license. Any commercial use requires a separate license from Ammonix —
licensing@ammonix.ai.
Verify your download
Every shard's SHA-256 is pinned in the code repository
(runs/manifests/writer_pin.json) — your download should hash identically.
Use
Serve with vLLM (the agent's configuration: temperature 0, seed 0, JSON-schema-constrained
decoding, thinking disabled) and point the RCM agent's writer at it — see the code repo's
REPRODUCING.md. The adapter/ folder carries the un-merged LoRA for use on the pinned
base (ornith-ai/Ornith-1.5-9B @ c927ad73b7eb); training pairs and scripts ship in the
code repository, so the adapter can also be retrained from scratch.
Disclaimer: research demonstration on a fully synthetic claims world — not a medical device, not for clinical or production billing use.
Cite
See the code repository's CITATION.cff (paper: https://doi.org/10.5281/zenodo.22871212). Commercial licensing:
licensing@ammonix.ai
- Downloads last month
- 157