Instructions to use eewer/qwen35-v0-8-step22000 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use eewer/qwen35-v0-8-step22000 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="eewer/qwen35-v0-8-step22000") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("eewer/qwen35-v0-8-step22000") model = AutoModelForCausalLM.from_pretrained("eewer/qwen35-v0-8-step22000", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use eewer/qwen35-v0-8-step22000 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "eewer/qwen35-v0-8-step22000" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "eewer/qwen35-v0-8-step22000", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/eewer/qwen35-v0-8-step22000
- SGLang
How to use eewer/qwen35-v0-8-step22000 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "eewer/qwen35-v0-8-step22000" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "eewer/qwen35-v0-8-step22000", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "eewer/qwen35-v0-8-step22000" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "eewer/qwen35-v0-8-step22000", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use eewer/qwen35-v0-8-step22000 with Docker Model Runner:
docker model run hf.co/eewer/qwen35-v0-8-step22000
Qwen3.5-35B-A3B terminal agent — step 22,000
This is the step-22,000 intermediate checkpoint of the terminal agent run.
It is a supervised fine-tune of Qwen/Qwen3.5-35B-A3B-Base for agentic text generation.
Terminal-agent SFT over Terminus-2-focused trajectories from two modern teacher lineages, with sequence packing and supervision on every assistant output.
This repository contains BF16 Hugging Face safetensors exported from the finalized Megatron distributed checkpoint. It is not quantized.
Checkpoint identity
| Field | Value |
|---|---|
| Hugging Face repository | eewer/qwen35-v0-8-step22000 |
| Run name | terminal |
| Training iteration | 22,000 |
| Planned iterations / epoch | 2,961 |
| Position in planned epoch | 742.99% |
| Base model | Qwen/Qwen3.5-35B-A3B-Base |
| Source checkpoint | /KRAFTON/WORKSPACE/wbl-workspace/posttraining-2606/areal_runs/qwen35_b300/reasoning-repair/sft/v0_8-ep3-1node-65k-tp2-ep8-gbs16-lr1e5/checkpoints/iter_0022000 |
| Export format | BF16 Hugging Face safetensors |
| Context length used for SFT | 32,768 tokens |
Training configuration
- Hardware: one node with 8 NVIDIA B300 GPUs.
- Framework: Megatron-Bridge / Megatron-Core in the NVIDIA NeMo 26.06 environment.
- Parallelism: TP=1, PP=1, CP=1, EP=8, DP=8.
- Micro batch size: 1 per data-parallel rank.
- Global batch size: 32.
- Activation recomputation: full, uniform recomputation.
- MoE router fusion: enabled; shared-expert overlap, gradient-reduce overlap, and parameter-gather overlap disabled.
- Precision: BF16 training/model tensors with the recipe's precision-aware optimizer state; exported weights are BF16.
- Learning-rate schedule: cosine decay from a peak of
1e-5to1e-6. - Warmup: 158 optimizer steps for every run.
- Checkpoint interval: 500 optimizer steps.
- Sequence packing: Yes; 94,742 packs, 93.8153% packing efficiency.
- Objective: assistant-only causal language modeling. System, user, and tool messages are context rather than prediction targets.
Training data
The materialized source mixture has 168,646 rows; 168,646 rows enter training after the 32K boundary check. The complete training representation contains 2,912,501,244 effective/rendered tokens.
| Source dataset | Selected rows |
|---|---|
nvidia/Nemotron-Terminal-Corpus |
79,201 |
open-thoughts/OpenThoughts-Agent-SFT-100K |
89,445 |
The source rows are the same deterministic shuffled union used by the unpacked terminal run. All 168,646 rows pass the 32K boundary check with zero drops. Sequences are packed to 32,768 tokens while preserving trajectory boundaries, and every assistant output is supervised without per-turn loss masks.
The tokenizer is from Qwen3.5, with a Qwen3.6 thinking-preserving chat template. The template keeps reasoning content in assistant messages and preserves tool declarations when the source supplies a validated schema.
Intended use
This checkpoint is intended for research on terminal, software-engineering, tool-use, and interactive agents. Use the included chat template and supply tool schemas expected by the target harness. It can be served with Transformers-compatible runtimes that support the Qwen3.5 MoE architecture.
This is a training checkpoint, not a polished instruction model. Compare checkpoints using held-out loss and task-level agent evaluations before selecting one for deployment.
Limitations and safety
- No complete benchmark suite is claimed in this model card.
- Agent trajectories can produce destructive shell commands, modify files, call tools, or expose secrets. Run the model in an isolated environment with scoped credentials.
- Tool names and schemas vary across source harnesses. A caller must provide the schema appropriate for its own environment; do not assume a generated call is valid or safe.
- Dataset filtering and deduplication reduce, but cannot guarantee removal of, benchmark contamination, incorrect reasoning, insecure code, or teacher-model artifacts.
- The model can hallucinate successful tool execution and should not be trusted without checking actual environment observations.
- This checkpoint inherits the base model's limitations and license.
Reproducibility notes
The B300 SFT recipes use the same execution topology, optimizer family, 158-step warmup, and checkpoint cadence. The checkpoint-specific learning-rate range and packing mode are reported above. The terminal_swe and terminal_swe_general runs differ from the terminal runs primarily in their source mixtures. The mixed datasets were validated for unique conversation hashes, deterministic shuffle order, non-empty reasoning in every assistant turn, maximum rendered length, and tool-call/schema consistency.
- Downloads last month
- 253
Model tree for eewer/qwen35-v0-8-step22000
Base model
Qwen/Qwen3.5-35B-A3B-Base