Spark-X2.5

Slack Discord YouTube dev.to Bluesky X Zhihu WeChat

This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format.

Introduction

We are introducing Spark-X2.5-4B and Spark-X2.5-1.7B, two compact, general-purpose language models designed to make capable AI more practical, efficient, and accessible. The models deliver strong performance across a broad range of everyday tasks—including conversation, writing, translation, reasoning, coding, tool use, and agentic workflows—achieving leading results among open-source models of comparable size. Spark-X2.5 combines an efficiency-oriented architecture with native context windows of up to 1M tokens, and support for more than 200 languages.

Technical Highlights:

  • Efficient Architecture and Native 1M-token Context: The models use a hybrid attention architecture that combines one full-attention layer with three sliding-window attention layers. This design substantially reduces the computational overhead typically associated with long-context models while natively supporting a context window of up to 1M tokens.
  • Strong Coding and Agent Capabilities: The models are deeply integrated with popular agent harnesses, including Codex, Claude Code, OpenClaw, and Hermes. They deliver state-of-the-art performance among models of comparable size across everyday coding, agentic workflows, reasoning, and instruction-following tasks.
  • Broad Hardware and Software Compatibility: The models support a wide range of hardware platforms, including NVIDIA, Huawei, Hygon, HOUMO.AI, etc. It is compatible with leading inference frameworks such as vLLM, SGLang, llama.cpp, MLX, and can be deployed quickly through platforms including Ollama and LM Studio. The models can also be customized using popular fine-tuning frameworks such as LLaMA-Factory. Across multiple hardware platforms, they deliver superior TTFT, TOPT, and overall inference efficiency compared with similarly sized models.
  • Advanced Training Algorithms: The models were trained on Huawei Ascend clusters. Large-scale reinforcement learning and post-training techniques such as MOPD significantly enhance its reasoning, coding, agentic, and instruction-following capabilities.

Spark-X2.5 benchmark comparison

Model Overview

For agent tasks, balancing performance, inference speed, and cache usage has long been a key bottleneck limiting model performance. Spark-X2.5 systematically integrates and optimizes mature attention technologies, combining sliding-window attention (SWA) with a hybrid full-attention architecture. This approach leverages the strengths of both mechanisms while avoiding the limitations of relying on a single structure, achieving an effective balance among performance, inference efficiency, and KV-cache size—thereby improving its practicality and effectiveness across real-world deployment scenarios.

Spark-X2.5 hybrid architecture

Training Methods

Spark-X2.5 is pretrained on approximately 20 trillion tokens from a diverse corpus spanning web pages, books, academic publications, code, and encyclopedic materials. Particular attention is paid to data quality, domain coverage, and the sampling weights assigned to different data categories. Extensive data-mixture studies are conducted to determine an effective balance among mathematics, logic, code, and other high-value domains. This enables the models to acquire broad general knowledge while developing stronger capabilities in complex reasoning and code generation. Long-context capability is developed through a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens.

Post-training begins with supervised fine-tuning on a carefully curated corpus. This stage establishes robust instruction following, structured generation, and task-completion, while providing a stable policy initialization for reinforcement learning. We subsequently apply large-scale reinforcement learning across several capability domains, including language understanding, reasoning, programming, tool-augmented agentic behavior, and instruction following. This process yields a set of domain-specialized teacher policies, whose complementary strengths are consolidated into a single deployable model through MOPD.

Spark-X2.5 hybrid architecture

Benchmarks

We evaluate our models and compare them with leading on-device models of similar size across a broad range of tasks, including agent, code, math, general and knowledge.

Benchmark Spark‑X2.5‑4B Spark‑X2.5‑1.7B Qwen3.5‑9B Qwen3.5‑4B Qwen3.5‑2B Gemma4‑12B Gemma4‑E4B Gemma4‑E2B
Agent
BFCL‑V465.146.966.1*50.3*43.6*37.436.930.2
τ²‑bench75.165.379.1*79.9*48.8*69.0*42.2*24.5*
τ³‑bench30.420.19.36.74.113.310.18.8
MCP‑Atlas54.623.447.4*40.8*14.830.5*15.0*12.6
MCP‑Mark14.22.313.412.5
Workspace Bench31.218.925.521.37.7
VitaBench2.025.28.315.618.25.212.44.84.4
BrowseComp40.929.78.314.33.110.08.33.7
Code
SWE‑Bench Pro44.410.433.8*29.4*1.921.9*4.0*
SWE‑Bench Verified41.628.353.1*38.8*6.844.2*14.0*
SWE‑Bench Multilingual53.323.343.327.75.032.5*
SciCode34.718.232.7*24.06.039.827.520.5
Math
Gaokao 2026133.4114.8135.5130.394.0130.6102.481.8
AIME 202690.769.488.283.030.882.1*42.5*37.5*
HMMT Feb 202681.248.470.869.721.565.634.220.5
IMO‑AnswerBench74.245.469.868.557.226.922.6
General & Knowledge
IFEval93.089.591.5*89.8*78.6*94.845.334.8
IFBench75.066.364.559.241.3*73.5*44.0*22.7
AA‑LCR56.324.363.0*57.0*25.6*55.3*34.718.3
HLE12.36.314.38.62.113.13.92.5
GPQA67.443.877.267.244.672.854.543.8
  • * denotes reported results from publicly‑released model cards / papers and - denotes scores not yet available.
  • All evaluations are conducted in thinking mode. The recommended sampling parameters for Spark-X2.5 are temperature=1.0, top_p=0.95, and top_k=-1.
  • Gaokao 2026 consists of the five 2026 Chinese GAOKAO examinations (National I,National II, Beijing, Shanghai, Tianjin), each graded out of 150 points.

Quickstart

SGLang

Install SGLang

Use the pre-built image that tracks the Spark-X2.5 runtime:

docker pull lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1

Run Inference

The following command can be used to start an OpenAI-compatible API server on a single GPU with maximum context length 1,048,576 tokens.

Server

docker run -it \
  --gpus '"device=0"' \
  --ipc=host \
  -p 30000:30000 \
  -v "$MODEL_PATH":/root/Spark-X2.5-4B \
  lmsysorg/sglang:nightly-dev-cu13-20260827-20621aa1 \
  python -m sglang.launch_server \
    --model-path /root/Spark-X2.5-4B \
    --served-model-name spark2.5 \
    --tool-call-parser spark25 \
    --reasoning-parser qwen3 \
    --tp-size 1 \
    --mem-fraction-static 0.8 \
    --context-length 1048576 \
    --chat-template /root/Spark-X2.5-4B/chat_template.jinja \
    --host 0.0.0.0 \
    --port 30000

Client

Thinking is enabled by default by both the chat template and the qwen3 reasoning parser. To disable thinking for a specific request, set "chat_template_kwargs": {"enable_thinking": false}.

curl -s http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spark2.5",
    "messages": [
      {
        "role": "user",
        "content": "安徽的省会在哪里?"
      }
    ],
    "max_tokens": 131072,
    "temperature": 1,
    "top_k": -1,
    "top_p": 0.95,
    "repetition_penalty": 1,
    "presence_penalty": 0,
    "frequency_penalty": 0
  }'

vLLM

Install vLLM

pip install uv
uv venv ~/spark2_5
source ~/spark2_5/bin/activate
git clone https://github.com/XHToken/Spark-plugin.git
cd ./Spark-plugin
uv pip install .

Server

vllm serve "./Spark-X2.5-4B" \
 --port "30000" \
 --trust-remote-code \
 --served-model-name spark25 \
 --tensor-parallel-size 1 \
 --gpu-memory-utilization 0.7 \
 --enable-prefix-caching \
 --chat-template Spark-X2.5-4B/chat_template.jinja

Client

curl -s http://127.0.0.1:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "spark25",
    "messages": [{"role": "user", "content": "安徽的省会在哪里?"}],
    "temperature": 1.0, 
    "top_k": -1,
    "top_p": 0.95
  }'

MLX

Spark-MLX-LLM runs the original Spark-X2.5 Hugging Face checkpoints locally. It supports Apple silicon GPU, Linux CPU, and NVIDIA CUDA on Linux. No GGUF conversion is required.

Installation

git clone https://github.com/XHToken/Spark-MLX-LLM.git
cd Spark-MLX-LLM

python3 -m venv .venv
source .venv/bin/activate

# Apple silicon
python -m pip install -e .
# Linux cpu
python -m pip install -e '.[cpu]'
# Linux with cuda12
python -m pip install -e '.[cuda12]' 
# Linux with cuda13
python -m pip install -e '.[cuda13]'

Run Spark-X2.5

spark-mlx-generate \
  --device gpu \
  --dtype bfloat16 \
  --model XHToken/Spark-X2.5-4B \
  --prompt "安徽的省会在哪里?" \
  --max-tokens 512 \
  --temp 0

Ollama

Build

git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
git clone https://github.com/ollama/ollama.git ollama-spark
cd ollama-spark
export OLLAMA_LLAMA_CPP_SOURCE="$(cd ../llama.cpp-spark && pwd)"
cmake -S . -B build
cmake --build build --parallel 8

Create and Run

printf 'FROM /absolute/path/to/your.gguf\n' > ./Modelfile.spark
./ollama serve
./ollama create Spark-X2.5-4B -f ./Modelfile.spark
./ollama run Spark-X2.5-4B

LM Studio

Build

git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
cd llama.cpp-spark
cmake -S . -B build
cmake --build build --parallel 8

Set Up LM Studio

  1. Close LM Studio.

  2. Back up the selected runtime directory:

    <LM_STUDIO_HOME>/extensions/backends/<selected-runtime>/
    
  3. Copy the llama.cpp-spark build output into the selected runtime directory, overwriting the existing files.

  4. Place the GGUF model in the following directory:

    <LM_STUDIO_HOME>/models/<org>/<name>/
    

Example runtime directory on macOS:

./build/bin/* -> ~/.lmstudio/extensions/backends/llama.cpp-mac-arm64-apple-metal-advsimd-<version>/

Run with LM Studio

Open My Models, select the Spark-X2.5 model, click Load, then start a new Chat.

Run with lms cli

# Replace `<model>` with a model listed by `lms ls`
lms load <model>
lms chat <model>

Finetuning

We advise you to use Llama-Factory to finetune your models.

License

The Spark-X2.5 model series is licensed under the Apache 2.0 License.

Citation

If you find our work helpful, feel free to give us a cite.

@misc{sparkx2.5,
    title  = {Spark-X2.5 4B&1.7B: Pushing the Limits of Agentic Capabilities in On-Device Models},
    author = {SparkLLM Team},
    year   = {2026}
}
Downloads last month
429
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 2 Ask for provider support

Model tree for XHToken/Spark-X2.5-4B

Finetuned
(1)
this model
Quantizations
1 model

Collection including XHToken/Spark-X2.5-4B