Qwen3.8-Flash-Next HFQ โ experimental hipfire artifact
This repository publishes the branch-compatible HFQ artifact produced for hipfire's Qwen4 bring-up. It is not an official Qwen checkpoint and is not a uniform 4-bit conversion. The artifact contains mixed-format records, including QT44/QT53 routed expert payloads and BF16 tensors, and must be read by a compatible hipfire HFQ loader.
Artifact identity
| File | Size | SHA-256 | MD5 |
|---|---|---|---|
qwen3.8-flash-next.mq4r |
178001249816 bytes | 7dcbceb4f501a66abef81cc624b86a4da250850ba79acf8dcc6a5204a5c8303f |
fda74d3760dc803e778e9b30a2fe0ebd |
The recorded full tensor payload is 177986811416 bytes; the remaining bytes
are HFQ container/index metadata. Hub metadata and this model card may become
visible before the 178.0 GB weight file is committed. Treat the repository as
not download-ready until qwen3.8-flash-next.mq4r appears in the Files view at
178001249816 bytes and a downloaded copy matches the SHA-256 above; metadata
alone is insufficient. The source artifact was not modified or duplicated for
publication; the upload reads the frozen local path recorded in the provenance
file.
This is the MQ4R SKU-named publication of the artifact. The same bytes were
first published in this repository on 2026-09-19 under the filename
qwen3.8-flash-next.hfq (the production run's output name); the file was
renamed on 2026-09-20 with no change to its content โ the two names are the
same 178001249816-byte file, as recorded in the local artifact identity
(sha256 above, and provenance publication_history). The superseded
.hfq name no longer exists in this repository; use the name in the table.
Upstream provenance and license
The source checkpoint is
Qwen/Qwen3.8-Flash-Next
at pinned revision
de4b8e4d43b917e7706784d8bb445c9af86a3540.
The upstream model card and license were reviewed before publication. The
model remains under the Qwen Community License 1.0; this repository
does not relicense the upstream weights.
The quantized container and runtime integration are hipfire work. The exact
runtime/source reference used for this artifact is
hipfire commit 8c6be1d7d045e34f972ac2dc3541927b3fd7a97d
from the public PR #772.
Hipfire's software is licensed separately under
Apache-2.0
and
MIT, as
applicable to the referenced software files. No hipfire source code is
included in this model repository.
The immutable upstream model card is
Qwen3.8-Flash-Next at revision de4b8e4d43b917e7706784d8bb445c9af86a3540.
Validation boundary
This is an experimental branch-compatible artifact, not a ship-ready or quality-certified model release.
- Tested on AMD Strix Halo
gfx1151with ordinary HIP inference. - Native greedy MTP was exercised on the same frozen artifact.
- The validation boundary does not include retained AQL/PM4/Redline admission or physical EP2/EP4 validation.
- The latest fixed fixture measured 28.8 prefill tok/s and 8.9 decode tok/s
with a 291-token prompt,
max_tokens=16, ten warmups, and three fresh processes under contention. These are fixture-bound measurements, not a product throughput claim; the 500 prefill / 25 decode targets were not met. - Exact engine-oracle checks plus a deterministic Paris AR/MTP smoke were used for the scoped runtime boundary. No model-quality certificate is implied; no KLD/PPL claim is made for this artifact.
See the hipfire Qwen4 measurement checkpoint for the dated evidence boundary and non-claims.
Intended use
Use only with a hipfire build that supports this Qwen4 HFQ layout. Treat the
artifact as research/experimental software and preserve the upstream license
notice and attribution when redistributing it. The upload provenance,
container metadata, and exact local identity are recorded in
provenance.json.
Model tree for hipfire-models/qwen3.8-flash-next
Base model
Qwen/Qwen3.8-Flash-Next