⚠️ DO NOT USE THIS MODEL — IT IS NON-FUNCTIONAL

Amethyst 2 Mini is not a working model. It is unreliable, performs poorly, and contains serious defects. We do not recommend it for any purpose, including testing, evaluation, or production use.

Status: superseded — Amethyst 2.5 is coming soon

Amethyst 2.5 is currently in development and will be released in three size variants. It replaces this model entirely. Please wait for Amethyst 2.5 rather than building on this release.

Known problems

  • Tool calling is fragile. It only works with one specific prompt format. With thinking disabled (a common default in many runtimes), the model stops calling the search tool and answers from memory, sometimes inventing citations.
  • Inconsistent training format. The training data rendered tool-call turns and answer turns differently. This is the root cause of the behavior above and cannot be fixed without retraining.
  • Unreliable answers. Responses can contain factual errors and oddly phrased claims, even on simple questions.
  • Limited testing. The evaluation numbers below come from a small, hand-written test set. They do not reflect real-world reliability, and multi-turn tool use was not tested end to end.

The rest of this card is preserved for the historical record only. Its benchmark figures should not be taken as evidence that the model works.

Amethyst 2 Mini (Failed)

A general-purpose chat model on Qwen3-4B that knows when to call a web-search tool, and when not to.

Part of the Amethyst family. Amethyst 2 Mini (Failed) keeps a direct, conversational voice and adds reliable web-search tool calling: search when the answer depends on current information, answer from knowledge when it does not, decompose multi-part questions into several queries, and summarize retrieved snippets into a cited answer.

Important: use the default chat template

Do not disable thinking (enable_thinking=False, --reasoning off, or a hand-built prompt that ends in an empty <think></think> block). The training data taught tool calls in turns that have no empty think block, so with one present the model tends to answer in prose instead of calling the tool (15/28 correct tool decisions instead of 28/28 in our test, sometimes inventing citations). Use the tokenizer's default apply_chat_template(..., add_generation_prompt=True) and the default llama.cpp --jinja template.

Training

  • Base: Qwen3-4B, trained via LoRA on mlx-community/Qwen3-4B-4bit
  • Data: 10,000 synthetic conversations distilled from nvidia/nemotron-3-super-120b-a12b and nvidia/nemotron-3-ultra-550b-a55b through NVIDIA NIM, filtered to 9,156 train / 796 validation. Mix: general chat 36%, search-positive 48%, no-search (answerable from knowledge) 8%, follow-up turns 8%.
  • Method: LoRA rank 8, scale 20, 16 layers, lr 1e-5, batch 2, sequence length 2,048, 13,800 iterations (run in four resumed sessions).
  • Tool format: the model emits <tool_call>{"name": "web_search", "arguments": {"queries": ["..."]}}</tool_call> and receives <tool_result>...</tool_result>, then answers with [1]-style citations. The tool only exists if you describe it in the system prompt, exactly as in training (see below).

Evaluation

48 hand-authored held-out prompts (zero overlap with the training templates): 28 tool-decision prompts (14 should search, 14 should not) and 20 general-chat prompts.

Base Qwen3-4B Amethyst 2 Mini
Correct tool decisions (28) 24 28 (0 false calls, 0 misses)
General chat, blind pairwise judge, both A/B orders 9 wins 19 wins (9 ties, 3 unjudged)
Held-out loss (60 validation examples) 3.528 0.585

The base was run with enable_thinking=False (its normal non-thinking mode). Chat quality was judged on the LoRA-adapter form by nvidia/nemotron-3-super-120b-a12b; the MLX build below matches the adapter's held-out loss.

System prompts (use these exactly)

For plain chat (no tool):

You are Amethyst, a helpful, direct conversational assistant. Answer the actual question asked. Match your length to the question -- a one-line question gets a one-line answer, a substantive question gets real depth. Use plain language, skip filler openers, and don't restate the question back before answering. Format with markdown only when structure genuinely helps (lists, tables, code); prose questions get prose answers.

To enable web search, use the chat prompt followed by a blank line and the tool description (this is what the model was trained with):

You are Amethyst, a helpful, direct conversational assistant. Answer the actual question asked. Match your length to the question -- a one-line question gets a one-line answer, a substantive question gets real depth. Use plain language, skip filler openers, and don't restate the question back before answering. Format with markdown only when structure genuinely helps (lists, tables, code); prose questions get prose answers.

You have access to one tool:

web_search(queries: list[str]) -- search the live web. Pass 1-3 short keyword queries. Use it when the answer depends on current, local, or frequently changing information (news, prices, scores, releases, schedules, weather, "latest"/"current"/"today"). Do NOT use it for stable knowledge you already have -- arithmetic, definitions, settled history, language questions, or code.

To call it, emit exactly this and nothing else:
<tool_call>
{"name": "web_search", "arguments": {"queries": ["query one", "query two"]}}
</tool_call>

Results come back as:
<tool_result>
[1] Title -- snippet text (source.com)
[2] Title -- snippet text (source.com)
</tool_result>

After results arrive, answer in prose and cite the snippets you used as [1], [2], etc. If the results don't actually answer the question, say so plainly rather than guessing.

Formats in this repo

File Format Size Tool decisions (28)
model.safetensors (+ config, tokenizer) MLX, mixed precision 2.9 GB 28/28 (held-out loss 0.586)
amethyst-2-mini-Q8_0.gguf GGUF Q8_0 4.3 GB 28/28
amethyst-2-mini-Q4_K_M.gguf GGUF Q4_K_M 2.5 GB 25/28 (3 missed searches)

Why the MLX build is mixed precision: merging a LoRA into a 4-bit model re-quantizes it, and a plain 4-bit merge measurably damaged tool use (21/28 decisions, held-out loss 0.731). So the 16 layers the LoRA changed are merged at 8-bit and the remaining layers keep their original 4-bit weights. Q4_K_M loses some tool accuracy for the same reason; prefer Q8_0 when tool reliability matters.

MLX

from mlx_lm import load, generate

model, tokenizer = load("VertexAGI/amethyst-2-mini-failed")
system = """<paste the tool system prompt from above>"""
messages = [{"role": "system", "content": system},
            {"role": "user", "content": "What's the current price of gold per ounce?"}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=200))

GGUF (llama.cpp)

llama-server -m amethyst-2-mini-Q8_0.gguf --jinja -c 4096   # send the system prompt above in your chat request

Limitations

A 4B model: it can still miss a search on unusual phrasing, and the evaluation set is small (48 prompts) and hand-written. The tool must be wired up by you (the model only emits the call). Q4_K_M is measurably less reliable at tool decisions than Q8_0 and the MLX build.

License

Apache 2.0, inherited from Qwen3.

Downloads last month
352
Safetensors
Model size
4B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for VertexAGI/amethyst-2-mini-failed

Finetuned
Qwen/Qwen3-4B
Adapter
(1170)
this model