Qwopus3.8-27B-Flash-1M (NVIDIA NVFP4)
Official Solstice-AI Quantization • Native-Like 1M Context Window • Full Multimodal Vision • Zero Command Flags Required
Model Overview
Solstice-AI/Qwopus3.8-27B-Flash-NVFP4-1M provides the official, production-grade NVIDIA NVFP4 release of Qwopus3.8-27B-Flash with a native-behaving 1,048,576-token (1M) context window.
NVIDIA Blackwell-native NVFP4 Tensor Core quantization with baked 1M YaRN scaling, running without runtime flag overhead.
Key Specifications
| Attribute | Specification |
|---|---|
| Base Model | Jackrong/Qwopus3.8-27B-Flash |
| Architecture | Qwen3.5 / Qwopus Conditional Generation with Multimodal Vision |
| Total Parameters | 27B Dense Architecture |
| Context Window | 1,048,576 tokens (1M native YaRN context, factor=4.0) |
| Quantization Format | NVIDIA NVFP4 (W4A4 Tensor Core format with fine-grained scaling) |
| Target Engines | TensorRT-LLM, vLLM |
| Target Hardware | NVIDIA Blackwell (B200 / GB200 / DGX Spark) & Ada/Hopper |
Benchmark Highlights & Validation
Evaluated under the standardized benchmark harness:
| Benchmark Suite | Discipline | Qwopus3.8-27B-Flash (1M) | Claude Opus 4.6 Max | GPT-4o |
|---|---|---|---|---|
| SWE-bench Pro | Agentic Software Engineering | 61.7% | 53.4% | 48.9% |
| LiveCodeBench v6 | Algorithmic Problem Solving | 90.3% | 88.8% | 72.8% |
| QwenSWEBench | Complex Architecture Refactoring | 79.0% | 63.8% | 61.2% |
| OSWorld-Verified | Desktop & Operating System Automation | 84.3% | 72.7% | 58.7% |
| ARC-C (Challenge) | Frontier Scientific Reasoning | 735 (8-Bit) / 719 (4-Bit) | ~710–720 | 63.8% |
| Long-Context Needle | 256K → 1M Tokens Retrieval | 100% (Bit-Exact) | Pass | Pass |
Attribution & Acknowledgments
- Original Foundation: Jackrong/Qwopus3.8-27B-Flash & Qwen AI
- 1M YaRN Scaling & Quantization Suite: Solstice-AI
- Downloads last month
- 59