Qwen3.8-Flash-Next HFQ โ€” experimental hipfire artifact

This repository publishes the branch-compatible HFQ artifact produced for hipfire's Qwen4 bring-up. It is not an official Qwen checkpoint and is not a uniform 4-bit conversion. The artifact contains mixed-format records, including QT44/QT53 routed expert payloads and BF16 tensors, and must be read by a compatible hipfire HFQ loader.

Artifact identity

File Size SHA-256 MD5
qwen3.8-flash-next.mq4r 178001249816 bytes 7dcbceb4f501a66abef81cc624b86a4da250850ba79acf8dcc6a5204a5c8303f fda74d3760dc803e778e9b30a2fe0ebd

The recorded full tensor payload is 177986811416 bytes; the remaining bytes are HFQ container/index metadata. Hub metadata and this model card may become visible before the 178.0 GB weight file is committed. Treat the repository as not download-ready until qwen3.8-flash-next.mq4r appears in the Files view at 178001249816 bytes and a downloaded copy matches the SHA-256 above; metadata alone is insufficient. The source artifact was not modified or duplicated for publication; the upload reads the frozen local path recorded in the provenance file.

This is the MQ4R SKU-named publication of the artifact. The same bytes were first published in this repository on 2026-09-19 under the filename qwen3.8-flash-next.hfq (the production run's output name); the file was renamed on 2026-09-20 with no change to its content โ€” the two names are the same 178001249816-byte file, as recorded in the local artifact identity (sha256 above, and provenance publication_history). The superseded .hfq name no longer exists in this repository; use the name in the table.

Upstream provenance and license

The source checkpoint is Qwen/Qwen3.8-Flash-Next at pinned revision de4b8e4d43b917e7706784d8bb445c9af86a3540. The upstream model card and license were reviewed before publication. The model remains under the Qwen Community License 1.0; this repository does not relicense the upstream weights.

The quantized container and runtime integration are hipfire work. The exact runtime/source reference used for this artifact is hipfire commit 8c6be1d7d045e34f972ac2dc3541927b3fd7a97d from the public PR #772. Hipfire's software is licensed separately under Apache-2.0 and MIT, as applicable to the referenced software files. No hipfire source code is included in this model repository.

The immutable upstream model card is Qwen3.8-Flash-Next at revision de4b8e4d43b917e7706784d8bb445c9af86a3540.

Validation boundary

This is an experimental branch-compatible artifact, not a ship-ready or quality-certified model release.

  • Tested on AMD Strix Halo gfx1151 with ordinary HIP inference.
  • Native greedy MTP was exercised on the same frozen artifact.
  • The validation boundary does not include retained AQL/PM4/Redline admission or physical EP2/EP4 validation.
  • The latest fixed fixture measured 28.8 prefill tok/s and 8.9 decode tok/s with a 291-token prompt, max_tokens=16, ten warmups, and three fresh processes under contention. These are fixture-bound measurements, not a product throughput claim; the 500 prefill / 25 decode targets were not met.
  • Exact engine-oracle checks plus a deterministic Paris AR/MTP smoke were used for the scoped runtime boundary. No model-quality certificate is implied; no KLD/PPL claim is made for this artifact.

See the hipfire Qwen4 measurement checkpoint for the dated evidence boundary and non-claims.

Intended use

Use only with a hipfire build that supports this Qwen4 HFQ layout. Treat the artifact as research/experimental software and preserve the upstream license notice and attribution when redistributing it. The upload provenance, container metadata, and exact local identity are recorded in provenance.json.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for hipfire-models/qwen3.8-flash-next

Finetuned
(55)
this model