Qwen3.5-35B-A3B terminal agent — step 25,500

This is the step-25,500 intermediate checkpoint of the terminal agent run. It is a supervised fine-tune of Qwen/Qwen3.5-35B-A3B-Base for agentic text generation. Terminal-agent SFT over Terminus-2-focused trajectories from two modern teacher lineages, with sequence packing and supervision on every assistant output.

This repository contains BF16 Hugging Face safetensors exported from the finalized Megatron distributed checkpoint. It is not quantized.

Checkpoint identity

Field Value
Hugging Face repository eewer/qwen35-v0-7-step25500
Run name terminal
Training iteration 25,500
Planned iterations / epoch 2,961
Position in planned epoch 861.20%
Base model Qwen/Qwen3.5-35B-A3B-Base
Source checkpoint /KRAFTON/WORKSPACE/wbl-workspace/posttraining-2606/areal_runs/qwen35_b300/reasoning-repair/sft/v0_7-ep3-1node-65k-tp2-ep8-gbs16-lr1e5/checkpoints/iter_0025500
Export format BF16 Hugging Face safetensors
Context length used for SFT 32,768 tokens

Training configuration

  • Hardware: one node with 8 NVIDIA B300 GPUs.
  • Framework: Megatron-Bridge / Megatron-Core in the NVIDIA NeMo 26.06 environment.
  • Parallelism: TP=1, PP=1, CP=1, EP=8, DP=8.
  • Micro batch size: 1 per data-parallel rank.
  • Global batch size: 32.
  • Activation recomputation: full, uniform recomputation.
  • MoE router fusion: enabled; shared-expert overlap, gradient-reduce overlap, and parameter-gather overlap disabled.
  • Precision: BF16 training/model tensors with the recipe's precision-aware optimizer state; exported weights are BF16.
  • Learning-rate schedule: cosine decay from a peak of 1e-5 to 1e-6.
  • Warmup: 158 optimizer steps for every run.
  • Checkpoint interval: 500 optimizer steps.
  • Sequence packing: Yes; 94,742 packs, 93.8153% packing efficiency.
  • Objective: assistant-only causal language modeling. System, user, and tool messages are context rather than prediction targets.

Training data

The materialized source mixture has 168,646 rows; 168,646 rows enter training after the 32K boundary check. The complete training representation contains 2,912,501,244 effective/rendered tokens.

Source dataset Selected rows
nvidia/Nemotron-Terminal-Corpus 79,201
open-thoughts/OpenThoughts-Agent-SFT-100K 89,445

The source rows are the same deterministic shuffled union used by the unpacked terminal run. All 168,646 rows pass the 32K boundary check with zero drops. Sequences are packed to 32,768 tokens while preserving trajectory boundaries, and every assistant output is supervised without per-turn loss masks.

The tokenizer is from Qwen3.5, with a Qwen3.6 thinking-preserving chat template. The template keeps reasoning content in assistant messages and preserves tool declarations when the source supplies a validated schema.

Intended use

This checkpoint is intended for research on terminal, software-engineering, tool-use, and interactive agents. Use the included chat template and supply tool schemas expected by the target harness. It can be served with Transformers-compatible runtimes that support the Qwen3.5 MoE architecture.

This is a training checkpoint, not a polished instruction model. Compare checkpoints using held-out loss and task-level agent evaluations before selecting one for deployment.

Limitations and safety

  • No complete benchmark suite is claimed in this model card.
  • Agent trajectories can produce destructive shell commands, modify files, call tools, or expose secrets. Run the model in an isolated environment with scoped credentials.
  • Tool names and schemas vary across source harnesses. A caller must provide the schema appropriate for its own environment; do not assume a generated call is valid or safe.
  • Dataset filtering and deduplication reduce, but cannot guarantee removal of, benchmark contamination, incorrect reasoning, insecure code, or teacher-model artifacts.
  • The model can hallucinate successful tool execution and should not be trusted without checking actual environment observations.
  • This checkpoint inherits the base model's limitations and license.

Reproducibility notes

The B300 SFT recipes use the same execution topology, optimizer family, 158-step warmup, and checkpoint cadence. The checkpoint-specific learning-rate range and packing mode are reported above. The terminal_swe and terminal_swe_general runs differ from the terminal runs primarily in their source mixtures. The mixed datasets were validated for unique conversation hashes, deterministic shuffle order, non-empty reasoning in every assistant turn, maximum rendered length, and tool-call/schema consistency.

Downloads last month
260
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for eewer/qwen35-v0-7-step25500

Finetuned
(68)
this model

Datasets used to train eewer/qwen35-v0-7-step25500