Ling 3.0 Flash — CIRU IU4
Current runtime release: v1.1.0
Ciru builds for Kairic.ai to bring the best inference optimizations to production.
Ling-3.0-Flash-CIRU-IU4 is CIRU's Ling-specific INT4 checkpoint and AMD Strix Halo serving build for InclusionAI's 124B-parameter, 5.1B-active Ling 3.0 Flash model.
Custom runtime required for native IU4: use the
ciru/ling3-gfx1151-iu4branch ofciru-ai/vllm. Stock vLLM can load these weights through W4A16 compatibility, but it does not provide the IU4 path or reproduce the performance figures below.
This release combines three complementary changes:
- Better S4/G32 weights. CIRU recalibrated all 61,440 routed-expert projections in layers 2–41 using a Ling activation-importance objective. The checkpoint remains symmetric signed INT4, group size 32, in the same compressed-tensors format as the official model.
- Native S4 × S4 prefill on AMD IU4. Prefill uses a two-plane residual activation representation and the native
gfx1151integer WMMA lane. The packed checkpoint bytes remain canonical and safe for the normal W4 path. - Production W4 decode with native MTP K1. Sparse decode stays on the optimized W4 path, where Ling's 512-expert/top-8 routing is better matched to small token batches.
The result is one production profile: native IU4 where it earns throughput, optimized W4 where sparse decode wins, and no user-facing mode selection.
Weight fidelity
Against the unchanged official INT4 checkpoint on CIRU's locked 92 × 4096-token BF16-reference corpus:
| Checkpoint | Mean KL | Top-token agreement |
|---|---|---|
| Official Ling INT4 | 0.040795 | 95.7063% |
| CIRU IU4 weights | 0.035464 | 95.8784% |
| Improvement | −13.07% | +0.172 percentage points |
The recalibration changed weights, not the model architecture or tensor schema. The final next-token-prediction layer remains byte-identical to the official checkpoint because it was outside the calibration coverage.
Measured serving performance
Validated on one Ryzen AI MAX+ 395 / Radeon 8060S Strix Halo APU with 128 GB unified memory, native Linux, ROCm/HIP 7.15, Python 3.12, Torch 2.13 ROCm, Triton 3.8, and the CIRU gfx1151 vLLM runtime.
Six simultaneous 4K requests
Matched warmed run: six unique uncached 4096-token prompts, 256 forced output tokens each, MTP K1, 16K scheduler budget, prefix reuse disabled.
| Metric | Official W4 control | CIRU IU4 | Change |
|---|---|---|---|
| Aggregate prefill | 504.89 tok/s | 644.94 tok/s | +27.74% |
| Aggregate decode | 70.92 tok/s | 81.90 tok/s | +15.48% |
| Median per-request decode | 11.92 tok/s | 14.51 tok/s | +21.73% |
| Median TTFT | 48.67 s | 38.10 s | −21.72% |
| Batch wall time | 70.25 s | 55.71 s | −20.69% |
| MTP acceptance | 98.19% | 98.32% | +0.13 pp |
Both arms completed all requests with zero cached prompt tokens, zero prefix hits, zero preemptions, and no faults.
Dynamic long-context scheduling
The current runner automatically chooses the scheduler shape that fits the live Ling workload. One runnable request—or a request that began with more than 8K uncached prompt tokens—is held to the faster 8K prefill shape until its prompt completes. When two or more short/medium prefills are runnable, the budget rises to 16K. Decode keeps its existing priority.
Against the previous static-16K public launcher configuration, an order-balanced two-mirror warmed comparison on the same diagnostic source (policy disabled for control) at C1 with exactly 32K uncached input tokens and 128 output tokens produced:
| Metric | Static 16K | Dynamic (two-run range) | Pooled change |
|---|---|---|---|
| Prefill | 424.37–425.17 tok/s | 555.67–558.65 tok/s | +31.17% |
| Decode | 22.73–22.79 tok/s | 22.71–22.73 tok/s | −0.17% |
| TTFT | 77.07–77.22 s | 58.66–58.97 s | −23.76% |
| Wall time | 82.65–82.80 s | 64.25–64.56 s | −22.15% |
At C6 with six 8K uncached prompts, the same mirrored comparison retained the static-16K path: pooled PP was −0.59%, median request TG was −0.25%, and wall time was +0.58%. All measured requests completed without cache reuse, preemptions, or device faults. Both C1 token sequences were exact; a natural HumanEval/0–5 check passed 24/24 executions across the four independent server launches.
The raw order-balanced matched rows were:
| Position | C1 32K PP | C1 TG | C1 TTFT | C1 wall | C6 8K PP | C6 median TG | C6 E2E output/s | C6 wall | C6 accept |
|---|---|---|---|---|---|---|---|---|---|
| A1 static 16K | 424.369 | 22.727 | 77.216 | 82.805 | 612.935 | 13.232 | 8.585 | 89.460 | 96.401% |
| C1 dynamic | 558.645 | 22.707 | 58.656 | 64.250 | 608.033 | 12.828 | 8.488 | 90.486 | 95.652% |
| C2 dynamic | 555.673 | 22.731 | 58.970 | 64.558 | 611.529 | 13.255 | 8.565 | 89.670 | 95.652% |
| A2 static 16K | 425.172 | 22.786 | 77.070 | 82.645 | 613.897 | 12.916 | 8.566 | 89.653 | 95.652% |
An earlier dedicated-profile oracle check is also retained. It compared the dynamic profile with separately tuned static 8K C1 and static 16K C6 servers, so it is not the deployed-default comparison above:
| Boundary | Profile | PP | Median TG | E2E output/s | TTFT s | Wall s | Accept |
|---|---|---|---|---|---|---|---|
| C1 16K | Static 8K oracle | 645.290 | 22.430 | 4.120 | 25.390 | 31.052 | 96.92% |
| C1 16K | Dynamic | 632.705 | 22.610 | 4.060 | 25.895 | 31.513 | 96.92% |
| C1 32K | Static 8K oracle | 567.140 | 22.760 | 2.020 | 57.780 | 63.359 | 100.00% |
| C1 32K | Dynamic | 555.525 | 22.797 | 1.983 | 58.986 | 64.557 | 100.00% |
| C6 4K | Static 16K oracle | 638.240 | 14.880 | 16.290 | 38.510 | 47.157 | 96.91% |
| C6 4K | Dynamic | 630.710 | 14.260 | 15.980 | 38.970 | 48.062 | 97.16% |
| C6 8K | Static 16K oracle | 601.760 | 11.310 | 8.460 | 79.050 | 90.832 | 96.40% |
| C6 8K | Dynamic | 599.920 | 12.920 | 8.410 | 81.360 | 91.273 | 96.65% |
The dedicated static-8K oracle was 1.95% and 2.05% faster in C1 wall time at 16K and 32K. Dynamic C6 wall was +1.92% at 4K and +0.49% at 8K. These small oracle gaps do not justify a second serving profile.
Dynamic C1 context sweep
One excluded 512-token warm request was followed by one uncached measured pass at each context, with exactly 128 forced output tokens:
| Input | PP tok/s | TG tok/s | Wall s | MTP acceptance |
|---|---|---|---|---|
| 2K | 720.688 | 24.688 | 7.990 | 98.44% |
| 4K | 438.732 | 23.828 | 14.704 | 96.92% |
| 8K | 675.071 | 24.360 | 17.356 | 95.38% |
| 16K | 519.984 | 22.725 | 37.136 | 98.44% |
| 32K | 559.334 | 21.819 | 64.449 | 95.38% |
| 64K | 437.604 | 18.568 | 156.650 | 92.42% |
| 128K | 342.860 | 15.963 | 390.310 | 96.92% |
These are one-pass boundary measurements, not best-of results. The isolated 4K row did not reproduce after a same-concurrency warm batch: warmed C1/4K measured 700.276 PP tok/s in the ladder below. Both values are reported so the one-pass anomaly remains visible.
Dynamic 4K concurrency ladder
Each row used one excluded same-concurrency warm batch followed by one measured batch of unique uncached 4096-token prompts and exactly 128 output tokens per request:
| C | Aggregate PP | Median/request TG | Aggregate TG | E2E output/s | Wall s | MTP acceptance |
|---|---|---|---|---|---|---|
| 1 | 700.276 | 23.578 | 23.578 | 11.392 | 11.236 | 96.92% |
| 2 | 704.722 | 19.851 | 39.458 | 14.173 | 18.062 | 97.67% |
| 3 | 688.311 | 17.704 | 53.107 | 15.343 | 25.027 | 96.91% |
| 4 | 671.804 | 16.416 | 65.186 | 15.910 | 32.181 | 97.29% |
| 5 | 661.500 | 15.088 | 74.887 | 16.228 | 39.439 | 96.60% |
| 6 | 651.611 | 14.202 | 84.525 | 16.435 | 46.730 | 96.15% |
| 7 | 641.725 | 13.345 | 92.634 | 16.508 | 54.276 | 96.48% |
| 8 | 642.692 | 12.743 | 100.015 | 16.804 | 60.937 | 95.03% |
| 9 | 645.142 | 9.140 | 80.513 | 17.025 | 67.665 | 95.24% |
| 10 | 647.727 | 7.060 | 68.972 | 17.168 | 74.558 | 95.40% |
All 55 measured requests completed with exact prompt/output accounting, zero cache reuse, zero preemptions, and no runtime faults. C8 is peak aggregate decode. C9 crosses from two to three 16K prefill waves. The launcher stays at six default slots because C8 improves end-to-end output by only 2.25% over C6 while increasing batch wall time by 30.4%.
Full-agent dynamic-build validation
The thinking-enabled HermesAgent-20 C6 run measured:
| Metric | Result |
|---|---|
| Score | 90 |
| Outcomes | 16 pass / 1 partial / 3 fail |
| Wall time | 1,387.878 s (23:07.9) |
| Prefill | 279.977 tok/s |
| Median/request decode | 4.817 tok/s |
| Active aggregate decode | 18.413 tok/s |
| Suite output / wall | 17.337 tok/s |
| End-to-end throughput | 52.904 tok/s |
| Mean TTFT | 11.259 s |
| MTP acceptance | 90.989% |
| Prompt / computed / cached tokens | 484,645 / 355,774 / 126,720 |
| Generated tokens / requests | 24,062 / 121 |
There were zero preemptions, zero device/runtime faults, and no scenario timeout. This is a standalone dynamic-build validation, not a matched comparison to an older run with different timeout or sampling conditions.
Full HumanEval and HumanEval+
The production profile completed a six-worker full-suite run with workers replenished continuously:
| Evaluation | Result |
|---|---|
| HumanEval pass@1 | 159 / 164 — 96.95% |
| HumanEval+ pass@1 | 156 / 164 — 95.12% |
| Aggregate generated throughput | 64.01 tok/s |
| Completion tokens | 200,550 |
| Generation wall time | 3,133.06 s |
| MTP acceptance | 91.88% |
| Runtime faults / preemptions | 0 / 0 |
Sampling used InclusionAI's recommended live settings (temperature=0.6, top_p=0.95, top_k=20). The harness imposed a 4096-token ceiling per task; seven tasks reached that ceiling, so these are the measured capped scores rather than an extrapolated score.
Model details
| Property | Value |
|---|---|
| Architecture | BailingMoeV3ForCausalLM / hybrid KDA–MLA MoE |
| Parameters | 124B total, 5.1B active |
| Experts | 512 routed experts, top-8 |
| Native context | 262,144 tokens |
| Weight format | compressed-tensors pack-quantized |
| Quantization | symmetric signed INT4, group size 32, BF16 scales |
| Source checkpoint | inclusionAI/Ling-3.0-flash-int4 |
Usage
The custom runner, gfx1151 kernel sources, extension builder, and production launcher are published in tools/ling_iu4. This scheduler release is pinned to tag ling-iu4-v1.1.0.
1. Build the runner
Ubuntu 24.04 / Debian-family setup:
sudo apt update
sudo apt install -y build-essential git libnuma-dev libdrm-dev libelf-dev \
ninja-build pkg-config python3.12 python3.12-dev python3.12-venv xxd
curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
git clone --branch ling-iu4-v1.1.0 \
https://github.com/ciru-ai/vllm.git "$HOME/vllm-ling-ciru-iu4"
cd "$HOME/vllm-ling-ciru-iu4"
bash tools/ling_iu4/install_runtime.sh
source .venv/bin/activate
The installer rebuilds the wheel-pinned ROCr userspace library with the host
compiler. This prevents the distributed ROCm 7.15 runtime from leaving a CPU
core continuously busy while the gfx1151 server is idle. The launcher
checksum-verifies and preloads the corrected library automatically.
Existing installs made from commit 92ccee356 can apply only the idle-runtime
hotfix without rebuilding vLLM, the IU4 extension, or the model:
git fetch origin ciru/ling3-gfx1151-iu4 --tags
git checkout ling-iu4-v1.1.0
bash tools/ling_iu4/install_rocr_idle_fix.sh "$PWD"
Restart the server afterward. CIRU_DISABLE_ROCR_IDLE_FIX=1 is an emergency
diagnostic rollback to the wheel runtime and may restore the idle CPU spin.
2. Download the model
hf download jcbtc/Ling-3.0-Flash-CIRU-IU4 \
--local-dir "$HOME/models/Ling-3.0-Flash-CIRU-IU4"
3. Launch the production profile
From the CIRU vLLM checkout:
source .venv/bin/activate
MODEL_PATH="$HOME/models/Ling-3.0-Flash-CIRU-IU4" \
bash tools/ling_iu4/run_ling_ciru_iu4.sh
The launcher provides the measured configuration: native 262,144-token context, six scheduler slots, a dynamic 8K/16K Ling batched-token policy, chunked prefill, prefix caching, Ling tool/reasoning parsers, optimized W4 sparse execution, native IU4 execution for occupied expert work, and MTP K1. MAX_NUM_BATCHED_TOKENS remains the upper bound. Override HOST, PORT, MAX_MODEL_LEN, MAX_NUM_SEQS, MAX_NUM_BATCHED_TOKENS, or GPU_MEMORY_UTILIZATION only when needed.
To use a conventional static budget for diagnosis or a deliberately fixed deployment, disable the policy explicitly:
VLLM_LING_DYNAMIC_BATCHED_TOKENS=0 MAX_NUM_BATCHED_TOKENS=8192 \
MODEL_PATH="$HOME/models/Ling-3.0-Flash-CIRU-IU4" \
bash tools/ling_iu4/run_ling_ciru_iu4.sh
The kernel build and arithmetic smoke were reproduced from the clean public branch on CIRU's Radeon 8060S (gfx1151). The performance measurements were validated on CIRU's native-Linux NixOS host. The Ubuntu commands above are the supported mainstream build path, but are not presented as a separate clean-host benchmark validation. Native Windows vLLM is unsupported; Windows users require WSL2 with AMD ROCDXG.
Request example
curl -fsS http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "Ling-3.0-Flash-CIRU-IU4",
"messages": [{"role": "user", "content": "Write a Python function that merges overlapping intervals."}],
"chat_template_kwargs": {"enable_thinking": true},
"temperature": 0.6,
"top_p": 0.95,
"top_k": 20,
"max_tokens": 1024
}'
Intended use and limitations
This release targets local, OpenAI-compatible text generation and coding-agent workloads on high-memory AMD Strix Halo systems. It preserves Ling's native reasoning and tool-call behavior.
- The native IU4 acceleration is hardware-specific to
gfx1151; other devices use their supported W4A16 kernels. - The checkpoint is approximately 77 GB before runtime caches and KV cache. A 128 GB unified-memory system is the validated target.
- Benchmark results are single-host measurements and should not be generalized to different ROCm, vLLM, firmware, power, or memory configurations.
- Generated code must be reviewed and sandboxed before production execution.
Credits and license
- Base model, tokenizer, architecture, and official INT4 checkpoint: InclusionAI, MIT license.
- Inference engine: vLLM, Apache-2.0.
- AMD ROCm, HIP, and
gfx1151integer WMMA support: AMD and their respective upstream projects. - Weight calibration, native IU4 kernels, runtime integration, profiling, and validation: CIRU.
The original MIT license governs the model weights and accompanying model code. Review the upstream model card before deployment.
- Downloads last month
- 454
Model tree for jcbtc/Ling-3.0-Flash-CIRU-IU4
Base model
inclusionAI/Ling-3.0-flash-int4