llm-jp-4-8b-instruct-NVFP4 DFlash Drafter

This repository contains a DFlash draft model for speculative decoding with kel-jp/llm-jp-4-8b-instruct-NVFP4.

This is not a standalone text-generation model. Use it as the draft/speculator model while serving the NVFP4 model as the verifier in stock vLLM.

Intended Use

  • Verifier model: kel-jp/llm-jp-4-8b-instruct-NVFP4
  • Drafting method: dflash
  • Draft architecture: block size 4, 1 transformer layer, 28k draft vocabulary
  • Proposal setting used in the released config: greedy, speculative_tokens=3
  • Runtime target: stock vLLM with DFlash support

vLLM Example

vllm serve kel-jp/llm-jp-4-8b-instruct-NVFP4 \
  --trust-remote-code \
  --reasoning-parser llmjp4 \
  --speculative-config '{"method":"dflash","model":"kel-jp/llm-jp-4-8b-instruct-NVFP4-speculator.dflash","num_speculative_tokens":3}'

If you use an environment wrapper, set the same JSON as SPECULATIVE_CONFIG. The verifier model still needs the serving requirements documented on the NVFP4 model page, including the LLM-jp remote code/plugin setup.

Benchmark (DGX Spark / GB10, SM121)

Representative Japanese streaming decode benchmark on ELYZA-100 prompts, measured on NVIDIA DGX Spark / GB10 (SM121) with vLLM 0.24.0, stock vLLM DFlash, temperature 0, serial streaming requests, and exactly 128 generated tokens per request. Decode speed excludes time-to-first-token.

Setup Mean decode tok/s Speedup vs BF16 Notes
BF16 base model 14.42 1.00x llm-jp/llm-jp-4-8b-instruct, no speculative decoding
NVFP4 verifier 29.73 2.06x kel-jp/llm-jp-4-8b-instruct-NVFP4, no speculative decoding
NVFP4 + this DFlash drafter 49.04 3.40x stock vLLM DFlash, num_speculative_tokens=3

The DFlash row is 1.65x faster than the NVFP4 verifier baseline by ratio of standalone means. In the paired per-prompt run against the same NVFP4 baseline, the DFlash mean was 48.76 tok/s and the ratio of mean decode throughput was 1.64x.

Paired speedup summary:

  • Ratio of mean decode throughput: 1.64x
  • Mean paired request speedup: 1.64x
  • Median paired request speedup: 1.63x
  • p05/p95 paired speedup: 1.37x / 1.97x
  • Requests at least 1.5x faster: 70 / 100
  • Requests at least 2.0x faster: 5 / 100
  • Requests slower than baseline: 0 / 100

The raw benchmark artifacts are included under benchmark/:

  • bf16-elyza100-streaming.json
  • bf16-elyza100-streaming.csv
  • baseline-elyza100-streaming.json
  • baseline-elyza100-streaming.csv
  • dflash20k-b4-l1-v28k-elyza100-streaming.json
  • dflash20k-b4-l1-v28k-elyza100-streaming.csv
  • elyza100-bf16-nvfp4-dflash-speedup-summary.json
  • elyza100-bf16-nvfp4-dflash-speedup.png
  • elyza100-baseline-vs-dflash-summary.json
  • elyza100-baseline-vs-dflash-speedup.png
  • dflash20k-b4-l1-elyza100-tps-histogram.png

BF16, NVFP4, and DFlash speedup distribution

Validation Metrics

The released checkpoint's held-out validation metrics:

{
  "loss_epoch": 2.047496609403255,
  "full_acc_epoch": 0.4021220020978507,
  "position_1_acc_epoch": 0.5539027137301571,
  "position_2_acc_epoch": 0.38056573750461736,
  "position_3_acc_epoch": 0.27111557473966125
}

Files

  • model.safetensors: released drafter weights
  • config.json: DFlash/speculators configuration consumed by vLLM
  • config.py: custom config class required by the drafter
  • val_metrics.json: validation metrics from the selected checkpoint
  • benchmark/: benchmark summaries, per-request CSV/JSON, and plots

Limitations

  • This artifact is useful only with a compatible verifier model and DFlash-capable vLLM runtime.
  • The published benchmark is throughput-oriented and uses Japanese ELYZA-100 prompts. It is not a general quality evaluation of the verifier model.
  • Speedup depends on prompt mix, max token settings, batching, hardware, and vLLM version.

License

Apache-2.0, following the verifier model release.

Downloads last month
24
Safetensors
Model size
1B params
Tensor type
I64
BF16
BOOL
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for kel-dx/llm-jp-4-8b-instruct-NVFP4-speculator.dflash

Finetuned
(1)
this model