Gingerlabs Qwen3.8 prompt-enhancer runtime mirror

This immutable deployment mirror contains only the two GGUF artifacts used by the Gingerlabs Runpod prompt-enhancer worker:

  • Q6_K_P target model
  • HauhauCS FastMTP 32K draft-vocabulary sidecar

The artifacts are byte-identical copies from upstream revision 993a5971fda8f30dd1b7eb2654792ba4415c7460. Signed upstream provenance and the FastMTP runtime patch are included. This repository intentionally omits all other quantizations and the vision projector so Runpod Cached Models does not prepare unused files.

The worker is text-only. The 32K in the FastMTP filename refers to the draft vocabulary, not the serving context length.

See THIRD_PARTY_NOTICES.md and LICENSE before redistribution.

Downloads last month
168
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for aperire0402/qwen38-prompt-enhancer-runtime