The ultimate guide to multi-harness RL

Qwen3.5-2B-opencode-standalone-RL

A full fine-tune of Qwen/Qwen3.5-2B for agentic data-analysis tasks, trained with asynchronous GRPO (TRL Async GRPO) using OpenCode. This release is step 1,000 on main, scoring 29.8% pass@1 across four evaluation harnesses.

Article · Collection · Evaluation tasks · Training dashboard

Training and evaluation curves

Animated training, pass@1 and tool-use curves through step 1,000

Training curves use a trailing 50-step mean (at least 10 observations). Evaluation markers show measured checkpoints; hollow markers and dotted segments indicate incomplete coverage. Faint background lines show the full recorded trajectory; colored lines reveal the measured checkpoints. The first frame shows the completed chart before replaying, so previews also contain the full curves. The tool-use panel shows absolute calls per graded evaluation rollout; this Qwen run used correctness-only rewards. The animation stops at this revision's step 1,000.

Static chart · Plotted data · Interactive article

Training

TRL Async GRPO with binary correctness only, initialized from the base model. All training rollouts use OpenCode. The standalone OpenCode environment runs outside Harbor in Daytona sandboxes.

Setting Value
Training task pool 1,000 tasks: 150 easy / 600 medium / 250 hard
Checkpoint 1,000 optimizer steps
Learning rate 3e-6
Rollouts per GRPO group / maximum staleness 8 / 4 optimizer steps
Optimizer / precision paged AdamW 8-bit / bfloat16
Training harnesses OpenCode
Sampling temperature / top-p 0.8 / 1.0
Per-call output budget, training / evaluation 16,384 / 4,096 tokens

These are the Qwen ablations from the article. Their reward has no tool-efficiency bonus. The output-budget mismatch and other limitations are discussed in the article; later Harbor checkpoints declined, so this release preserves the selected checkpoint rather than substituting the last one.

Evaluation

250 fixed SmolDataEnvs test tasks (33 easy, 118 medium, 99 hard), each evaluated under four harnesses: 1,000 graded task/harness cells. Pass@1 uses the first graded attempt per cell; infrastructure retries do not turn it into pass@k. All cells below are graded.

Harness Correct / evaluated Pass@1
OpenCode 51 / 250 20.4%
Claude Code 83 / 250 33.2%
Codex 74 / 250 29.6%
Mini-SWE-Agent 90 / 250 36.0%
Overall 298 / 1,000 29.8%

Harness versions: OpenCode 1.18.31, Claude Code 2.1.270, Codex 0.154.0, Mini-SWE-Agent 2.4.6. Full scores, including difficulty breakdowns, are in eval_results.json.

This is the run's best observed checkpoint on the same test set (also the final checkpoint for standalone OpenCode). Selection on test performance can inflate the reported result.

These are single-run results on a specific task set and harness versions. Harness mix, training exposure and compute differ across runs; the scores do not isolate a causal effect of the harness or objective.

Load the checkpoint

The repository contains full saved model weights and the saved tokenizer/chat template, not a LoRA adapter. Training used Transformers 5.14.1; use a compatible Transformers release.

import torch
from transformers import AutoTokenizer, AutoModelForImageTextToText

model_id = "FineEnvs/Qwen3.5-2B-opencode-standalone-RL"
revision = "main"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForImageTextToText.from_pretrained(
    model_id, revision=revision, dtype=torch.bfloat16, device_map="auto"
)

Reproducing the task scores requires the agent harness and tools described in the article; a plain chat prompt is not the same evaluation. This Qwen release was fine-tuned and evaluated for text/tool interaction; no image benchmark is claimed.

Related releases

All models, datasets, environments and the article are linked in the multi-harness RL collection.

License and provenance

Derived from Qwen/Qwen3.5-2B under Apache License 2.0. The base model's license is included unchanged in LICENSE. FineEnvs modified the weights by reinforcement learning; this is not an official Qwen release. See NOTICE and release_manifest.json for modification notices, the pinned base revision, checkpoint identity and file checksums. Optimizer, scheduler, RNG and trainer state are excluded.

Citation

For the experiment, methodology and interpretation, cite the main article:

@misc{kolavi2026multiharnessrl,
  author = {Adithya S Kolavi},
  title = {The ultimate guide to multi-harness RL},
  year = {2026},
  url = {https://huggingface.co/spaces/AdithyaSK/multi-harness-rl}
}
Downloads last month
-
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FineEnvs/Qwen3.5-2B-opencode-standalone-RL

Finetuned
Qwen/Qwen3.5-2B
Finetuned
(427)
this model

Dataset used to train FineEnvs/Qwen3.5-2B-opencode-standalone-RL

Collection including FineEnvs/Qwen3.5-2B-opencode-standalone-RL

Evaluation results