Instructions to use VertexAGI/amethyst-2-mini-failed with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use VertexAGI/amethyst-2-mini-failed with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("VertexAGI/amethyst-2-mini-failed") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use VertexAGI/amethyst-2-mini-failed with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf VertexAGI/amethyst-2-mini-failed:Q4_K_M # Run inference directly in the terminal: llama cli -hf VertexAGI/amethyst-2-mini-failed:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf VertexAGI/amethyst-2-mini-failed:Q4_K_M # Run inference directly in the terminal: llama cli -hf VertexAGI/amethyst-2-mini-failed:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf VertexAGI/amethyst-2-mini-failed:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf VertexAGI/amethyst-2-mini-failed:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf VertexAGI/amethyst-2-mini-failed:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf VertexAGI/amethyst-2-mini-failed:Q4_K_M
Use Docker
docker model run hf.co/VertexAGI/amethyst-2-mini-failed:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use VertexAGI/amethyst-2-mini-failed with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VertexAGI/amethyst-2-mini-failed" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAGI/amethyst-2-mini-failed", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VertexAGI/amethyst-2-mini-failed:Q4_K_M
- Ollama
How to use VertexAGI/amethyst-2-mini-failed with Ollama:
ollama run hf.co/VertexAGI/amethyst-2-mini-failed:Q4_K_M
- Unsloth Desktop
- Pi
How to use VertexAGI/amethyst-2-mini-failed with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAGI/amethyst-2-mini-failed"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "VertexAGI/amethyst-2-mini-failed" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use VertexAGI/amethyst-2-mini-failed with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "VertexAGI/amethyst-2-mini-failed"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "VertexAGI/amethyst-2-mini-failed" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAGI/amethyst-2-mini-failed", "messages": [ {"role": "user", "content": "Hello"} ] }' - Docker Model Runner
How to use VertexAGI/amethyst-2-mini-failed with Docker Model Runner:
docker model run hf.co/VertexAGI/amethyst-2-mini-failed:Q4_K_M
- Lemonade
How to use VertexAGI/amethyst-2-mini-failed with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull VertexAGI/amethyst-2-mini-failed:Q4_K_M
Run and chat with the model
lemonade run user.amethyst-2-mini-failed-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use VertexAGI/amethyst-2-mini-failed with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAGI/amethyst-2-mini-failed"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default VertexAGI/amethyst-2-mini-failed
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use VertexAGI/amethyst-2-mini-failed with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "VertexAGI/amethyst-2-mini-failed"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "VertexAGI/amethyst-2-mini-failed" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
⚠️ DO NOT USE THIS MODEL — IT IS NON-FUNCTIONAL
Amethyst 2 Mini is not a working model. It is unreliable, performs poorly, and contains serious defects. We do not recommend it for any purpose, including testing, evaluation, or production use.
Status: superseded — Amethyst 2.5 is coming soon
Amethyst 2.5 is currently in development and will be released in three size variants. It replaces this model entirely. Please wait for Amethyst 2.5 rather than building on this release.
Known problems
- Tool calling is fragile. It only works with one specific prompt format. With thinking disabled (a common default in many runtimes), the model stops calling the search tool and answers from memory, sometimes inventing citations.
- Inconsistent training format. The training data rendered tool-call turns and answer turns differently. This is the root cause of the behavior above and cannot be fixed without retraining.
- Unreliable answers. Responses can contain factual errors and oddly phrased claims, even on simple questions.
- Limited testing. The evaluation numbers below come from a small, hand-written test set. They do not reflect real-world reliability, and multi-turn tool use was not tested end to end.
The rest of this card is preserved for the historical record only. Its benchmark figures should not be taken as evidence that the model works.
Amethyst 2 Mini (Failed)
A general-purpose chat model on Qwen3-4B that knows when to call a web-search tool, and when not to.
Part of the Amethyst family. Amethyst 2 Mini (Failed) keeps a direct, conversational voice and adds reliable web-search tool calling: search when the answer depends on current information, answer from knowledge when it does not, decompose multi-part questions into several queries, and summarize retrieved snippets into a cited answer.
Important: use the default chat template
Do not disable thinking (enable_thinking=False, --reasoning off, or a hand-built prompt that ends in an empty <think></think> block). The training data taught tool calls in turns that have no empty think block, so with one present the model tends to answer in prose instead of calling the tool (15/28 correct tool decisions instead of 28/28 in our test, sometimes inventing citations). Use the tokenizer's default apply_chat_template(..., add_generation_prompt=True) and the default llama.cpp --jinja template.
Training
- Base: Qwen3-4B, trained via LoRA on
mlx-community/Qwen3-4B-4bit - Data: 10,000 synthetic conversations distilled from
nvidia/nemotron-3-super-120b-a12bandnvidia/nemotron-3-ultra-550b-a55bthrough NVIDIA NIM, filtered to 9,156 train / 796 validation. Mix: general chat 36%, search-positive 48%, no-search (answerable from knowledge) 8%, follow-up turns 8%. - Method: LoRA rank 8, scale 20, 16 layers, lr 1e-5, batch 2, sequence length 2,048, 13,800 iterations (run in four resumed sessions).
- Tool format: the model emits
<tool_call>{"name": "web_search", "arguments": {"queries": ["..."]}}</tool_call>and receives<tool_result>...</tool_result>, then answers with[1]-style citations. The tool only exists if you describe it in the system prompt, exactly as in training (see below).
Evaluation
48 hand-authored held-out prompts (zero overlap with the training templates): 28 tool-decision prompts (14 should search, 14 should not) and 20 general-chat prompts.
| Base Qwen3-4B | Amethyst 2 Mini | |
|---|---|---|
| Correct tool decisions (28) | 24 | 28 (0 false calls, 0 misses) |
| General chat, blind pairwise judge, both A/B orders | 9 wins | 19 wins (9 ties, 3 unjudged) |
| Held-out loss (60 validation examples) | 3.528 | 0.585 |
The base was run with enable_thinking=False (its normal non-thinking mode). Chat quality was judged on the LoRA-adapter form by nvidia/nemotron-3-super-120b-a12b; the MLX build below matches the adapter's held-out loss.
System prompts (use these exactly)
For plain chat (no tool):
You are Amethyst, a helpful, direct conversational assistant. Answer the actual question asked. Match your length to the question -- a one-line question gets a one-line answer, a substantive question gets real depth. Use plain language, skip filler openers, and don't restate the question back before answering. Format with markdown only when structure genuinely helps (lists, tables, code); prose questions get prose answers.
To enable web search, use the chat prompt followed by a blank line and the tool description (this is what the model was trained with):
You are Amethyst, a helpful, direct conversational assistant. Answer the actual question asked. Match your length to the question -- a one-line question gets a one-line answer, a substantive question gets real depth. Use plain language, skip filler openers, and don't restate the question back before answering. Format with markdown only when structure genuinely helps (lists, tables, code); prose questions get prose answers.
You have access to one tool:
web_search(queries: list[str]) -- search the live web. Pass 1-3 short keyword queries. Use it when the answer depends on current, local, or frequently changing information (news, prices, scores, releases, schedules, weather, "latest"/"current"/"today"). Do NOT use it for stable knowledge you already have -- arithmetic, definitions, settled history, language questions, or code.
To call it, emit exactly this and nothing else:
<tool_call>
{"name": "web_search", "arguments": {"queries": ["query one", "query two"]}}
</tool_call>
Results come back as:
<tool_result>
[1] Title -- snippet text (source.com)
[2] Title -- snippet text (source.com)
</tool_result>
After results arrive, answer in prose and cite the snippets you used as [1], [2], etc. If the results don't actually answer the question, say so plainly rather than guessing.
Formats in this repo
| File | Format | Size | Tool decisions (28) |
|---|---|---|---|
model.safetensors (+ config, tokenizer) |
MLX, mixed precision | 2.9 GB | 28/28 (held-out loss 0.586) |
amethyst-2-mini-Q8_0.gguf |
GGUF Q8_0 | 4.3 GB | 28/28 |
amethyst-2-mini-Q4_K_M.gguf |
GGUF Q4_K_M | 2.5 GB | 25/28 (3 missed searches) |
Why the MLX build is mixed precision: merging a LoRA into a 4-bit model re-quantizes it, and a plain 4-bit merge measurably damaged tool use (21/28 decisions, held-out loss 0.731). So the 16 layers the LoRA changed are merged at 8-bit and the remaining layers keep their original 4-bit weights. Q4_K_M loses some tool accuracy for the same reason; prefer Q8_0 when tool reliability matters.
MLX
from mlx_lm import load, generate
model, tokenizer = load("VertexAGI/amethyst-2-mini-failed")
system = """<paste the tool system prompt from above>"""
messages = [{"role": "system", "content": system},
{"role": "user", "content": "What's the current price of gold per ounce?"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=200))
GGUF (llama.cpp)
llama-server -m amethyst-2-mini-Q8_0.gguf --jinja -c 4096 # send the system prompt above in your chat request
Limitations
A 4B model: it can still miss a search on unusual phrasing, and the evaluation set is small (48 prompts) and hand-written. The tool must be wired up by you (the model only emits the call). Q4_K_M is measurably less reliable at tool decisions than Q8_0 and the MLX build.
License
Apache 2.0, inherited from Qwen3.
- Downloads last month
- 352
4-bit