PhoneLLM Alpha 1 — ROCmFP4 STRIX_LEAN (ftype 106)

The model here is not our work. PhoneLLM Alpha 1 is by Daily / the Pipecat teampipecat-ai/phonellm-alpha-1 — a full-parameter fine-tune of NVIDIA Nemotron 3 Nano 30B-A3B. This repository adds only the ROCmFP4/ROCmFPX quantisation ladder for AMD Strix Halo and the measurements below. Go star their repo. Pipecat also ship an official NVFP4 build for NVIDIA Blackwell: pipecat-ai/phonellm-alpha-1-nvfp4.

The tier to start with. 15.91 GiB, and nothing larger measured better.

One file: PhoneLLM-Alpha-1-Q4_0_ROCMFP4_STRIX_LEAN.gguf — the flagship 4-bit tier of our six-tier ladder for PhoneLLM Alpha 1 on AMD Strix Halo (gfx1151). Runs on both HIP (ROCm) and Vulkan from one binary — the backend is a runtime -dev flag, not a rebuild.

Size 15.91 GiB (from a 58.8 GiB BF16 source)
ftype 106 · Q4_0_ROCMFP4_STRIX_LEAN
Head q8_0 protected (output.weight), read-back verified
Tool probe 3/5 — identical to the 30 GiB Q8 tiers, and above the BF16 control (1/5)
Loads in ~10 s
Why this one Smallest tier with no measurable loss vs anything larger; leaves ~110 GB free on a 128 GB box for your ASR + TTS

⚠️ PhoneLLM is text-in / text-out. It is the LLM stage of a voice pipeline, not a speech model — you still need STT/ASR in front and TTS behind. Pipecat's own voice-to-voice budget is ~1500 ms, of which the LLM target is ~650 ms time-to-first-token.

How this tier compares to the rest of the ladder

Tier Size Tool probe
STRIX_LEAN (this file) 15.91 GiB 3/5
FAST 15.83 GiB 3/5
COHERENT 16.91 GiB 2/5
Q6_0_ROCMFPX_AGENT 27.26 GiB 3/5
Q8_0_ROCMFPX 30.37 GiB 3/5
Q8_0_ROCMFPX_AGENT 30.84 GiB 2/5
BF16 source (control) 58.8 GiB 1/5

Paying 2× the disk for Q8 bought nothing measurable here. Full ladder, with every tier's numbers: PhoneLLM-Alpha-1-ROCmFP4-GGUF.

Quick start

llama-server -m PhoneLLM-Alpha-1-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
  -dev ROCm0 -fa on -ngl 999 -fit off -np 1 \
  -c 32768 -b 4096 -t 8 --jinja \
  --host 0.0.0.0 --port 8080

Run it the way Pipecat recommend the source model: temperature=0 and thinking disabled.

{"chat_template_kwargs": {"enable_thinking": false}}

llama.cpp resolves this model's chat format as peg-native; tool calls come back as proper tool_calls on /v1/chat/completions with --jinja.

⛔ You need a ROCmFPX build — stock llama.cpp will NOT load these files

ROCmFP4/ROCmFPX use ggml tensor types 100–119; upstream's table stops at 43. Build with both backends:

cmake -S . -B build-hipvk -DCMAKE_BUILD_TYPE=Release \
  -DGGML_HIP=ON -DGGML_VULKAN=ON -DGGML_NATIVE=ON \
  -DVulkan_GLSLC_EXECUTABLE=/usr/bin/glslc -DAMDGPU_TARGETS=gfx1151
cmake --build build-hipvk -j
build-hipvk/bin/llama-server --list-devices   # must list ROCm0 AND Vulkan0

Head protection — why every tier uses a q8_0 head

Our usual ladder protects output.weight with q6_K on the 4-bit tiers. That is impossible on this model. hidden_size is 2688, and K-quants use 256-element superblocks:

2688 % 256 = 128        →  ggml.c:8024: GGML_ASSERT(start % type_traits[type].blck_size == 0) failed
[1/401] output.weight - [2688, 131072], bf16, converting to q6_K ..   SIGABRT

q8_0 uses 32-element blocks and 2688 % 32 == 0, so every tier here carries a q8_0 headhigher precision than our usual q6_K, at a cost of roughly +88 MB on the head tensor.

We caught this as a clean natural experiment in a single run: the three q6_K-head tiers aborted in ~2 s while the q8_0-head tier built normally, same source, same binary, same moment. Note that llama-quantize --dry-run does not catch it — the dry run planned all 401 tensors and printed a clean 60247 MiB → 17223 MiB (4.58 BPW) summary. The assert only fires once real data is written.


Verification

Every tier is checked for load, coherence, and — because this is the whole point of PhoneLLM — tool calling, using the vendor-recommended mode (temperature=0, enable_thinking: false).

The tool probe is deliberately adversarial: it includes a case where the model must not call anything, and a multi-turn case, because the failure Pipecat built this model to avoid is an agent that says "yes, I've booked that" without emitting a call.

All tiers, greedy (temperature=0), enable_thinking: false, --jinja, chat format peg-native.

Tier Size Loads Coherent Tool probe
Q4_0_ROCMFP4_STRIX_LEAN 15.91 GiB ✅ ~10 s 3/5
Q4_0_ROCMFP4_FAST 15.83 GiB ✅ ~10 s 3/5
Q4_0_ROCMFP4_COHERENT 16.91 GiB ✅ ~10 s 2/5
Q6_0_ROCMFPX_AGENT 27.26 GiB ✅ ~20 s 3/5
Q8_0_ROCMFPX_AGENT 30.84 GiB ✅ ~25 s 2/5
Q8_0_ROCMFPX 30.37 GiB ✅ ~20 s 3/5
BF16 source (control) 58.8 GiB 1/5

6/6 tiers load and stay coherent. There is no precision-dependent degradation: the 30 GiB Q8 tiers score the same as the 16 GiB 4-bit tiers, and every quantised tier scores at or above the BF16 control. If quantisation were damaging tool calling, the Q8 tiers would lead. They do not — so pick on size.

Flagship detail (STRIX_LEAN), 5 adversarial cases:

PASS  booking       book_table {"name":"Chen","party_size":2,"time":"19:00"}   ← normalised "7pm" → 19:00
PASS  escalate      transfer_to_human {"reason":"Customer has called multiple times ..."}
PASS  no-tool       (correctly emitted NO call)
FAIL  availability  (no call — the one unambiguous miss)
FAIL  multiturn     check_availability {"date":"Saturday","party_size":4}

Read 3/5 carefully — the rubric is strict and opinionated. The multiturn "failure" is the model checking availability before booking, which is defensible agent behaviour; we counted it wrong because our expected answer demanded a booking. The no-tool pass matters most: the model declined to invent a call when none was warranted, which is the failure mode PhoneLLM exists to avoid. Treat these as a smoke test that the tool path survives quantisation, not as a benchmark score — for a real score use Pipecat's PhoneBench.

⚠️ An honest limitation: we could not establish a BF16 baseline on this hardware

We ran the BF16 GGUF as a control arm and it misbehaves on gfx1151 when tools are attached — the same prompt that a quantised tier answers with a correct book_table call returns, from BF16, either a degenerate repetition loop or an unrelated non-sequitur. Without tools, BF16 is coherent.

So we can report what the quantised tiers do, but we cannot publish a "delta vs BF16" the way Pipecat report NVFP4 (PhoneBench 72.06 → 71.51). Anyone quoting a quality delta for these files against BF16 on this hardware would be quoting a broken control. We have not root-caused it (candidates: gfx1151 bf16 compute, or the peg-native tool-template path); it is flagged here rather than papered over.


Reproduction block

A number without its binary is a rumour.

Host Ryzen AI Max+ 395 (Strix Halo), Radeon 8060S, gfx1151, 128 GB unified
Build ROCmFPX fork @ e7712358806055c70a9753b070202b0cc7c637e3
GGML_HIP=ON GGML_VULKAN=ON GGML_NATIVE=ON, Release, AMDGPU_TARGETS=gfx1151
llama-server sha256 e860b763d5e8f496ddb4b01236d8a8f37ae09dda456b5d076ab3abe2d9deb9fc
Source pipecat-ai/phonellm-alpha-1, 13 safetensors shards, 58.8 GiB
Converted convert_hf_to_gguf.py --outtype bf16 → 401 tensors, 63.18 GB, arch nemotron_h_moe
Quantise llama-quantize --output-tensor-type q8_0 <bf16> <out> <ftype> 12
Serve (verification) -dev ROCm0 -fa on -ngl 999 -fit off -np 1 -c 32768 -b 4096 -t 8 --jinja

Not measured

  • Perplexity (the source is a voice-agent fine-tune; wikitext PPL is a poor proxy and we would rather publish nothing than a misleading number).
  • PhoneBench — that is Pipecat's harness; we did not run it.
  • Vision — text-only model, no projector.
  • Context beyond 32768 (the source supports 262144).
  • Decode throughput per tier.

License and attribution

Released under BSD 2-Clause, matching the source. The source is itself a derivative of an NVIDIA Nemotron Open Model License work — see LICENSE_NVIDIA.txt in the upstream repo.

Acknowledgements

Daily / Pipecat for PhoneLLM and for publishing an honest PhoneBench methodology. NVIDIA for Nemotron 3 Nano and the hybrid Mamba-Transformer architecture. The ROCmFPX project for the FP4/FPX tensor types and the Strix Halo kernels that make these files possible.

Downloads last month
-
GGUF
Model size
32B params
Architecture
nemotron_h_moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kingjones777/PhoneLLM-Alpha-1-ROCmFP4-STRIX-LEAN-GGUF

Quantized
(68)
this model