GRWO: Guided Reasoning Window Optimization

Guided Reasoning Window Optimization (GRWO) is a training strategy for improving reasoning models by optimizing local reasoning windows instead of full long-form solution traces.

The core idea is simple:

Given a problem and a partial reasoning prefix, train the model to prefer a better next reasoning window over a weaker next reasoning window.

Instead of forcing a model to generate and train on a full 5k–10k token solution, GRWO focuses on the next N tokens of reasoning, usually around a meaningful branch point.

This makes reasoning training cheaper, more targeted, and more aligned with how models actually fail.


1. Motivation

Long reasoning traces are expensive.

A hard math problem may require thousands of tokens of exploration, correction, theorem selection, calculation, and verification. Training directly on full traces has several problems:

  • It is costly to generate.
  • It is costly to train.
  • It may overwrite the model's natural reasoning style.
  • It gives loss on many tokens that are not the real failure point.
  • It can teach verbosity instead of better reasoning decisions.

Most reasoning failures happen at local branch points:

  • The model chooses the wrong theorem.
  • The model treats an inconclusive test as proof.
  • The model chooses bad coordinates.
  • The model forgets a Jacobian.
  • The model keeps exploring instead of finishing.
  • The model gives a confident but unsupported conclusion.

GRWO targets these local branch points directly.


2. One-Line Thesis

Train reasoning models by correcting local self-generated reasoning windows instead of supervising full solution traces.


3. Core Definition

A GRWO sample contains:

problem + partial reasoning prefix

Then the model generates two possible continuations:

rejected = unguided continuation
chosen   = guided or corrected continuation

Only the next local reasoning window is optimized.

prompt   = problem + partial reasoning prefix
chosen   = better next reasoning window
rejected = weaker next reasoning window

The key constraint:

The keypoints or guided answer are used only to create the chosen continuation. They are not exposed in the student model's prompt.


4. Why Windowed Reasoning?

Instead of this:

Question -> full 10k-token reasoning trace -> answer

GRWO trains this:

Question + reasoning prefix -> next 512–1500 reasoning tokens

This changes the objective from:

Solve the entire problem from scratch.

to:

Given this current reasoning state, choose the better next reasoning move.

That is the behavior weak reasoning models need most.


5. Name

Recommended method name:

Guided Reasoning Window Optimization (GRWO)

Variants:

GRWO-DPO   = DPO over local reasoning windows
GRWO-SFT   = SFT on corrected local reasoning windows
GRWO-GRPO  = RL/GRPO over local reasoning windows

Primary direction:

GRWO-DPO: preference training over local self-generated reasoning continuations.

6. Dataset Structure

6.1 DPO Format

{
  "prompt": "Problem statement...\n<think>\nPartial model-generated reasoning prefix...",
  "chosen": "Guided/correct next reasoning window...",
  "rejected": "Unguided/weaker next reasoning window..."
}

The prompt contains the problem and the partial reasoning prefix.

The chosen and rejected completions should continue from the exact same prefix.

6.2 SFT Format

{
  "prompt": "Problem statement...\n<think>\nPartial model-generated reasoning prefix...",
  "completion": "Correct next reasoning window..."
}

SFT is optional, but useful when the model lacks a reasoning move entirely.


7. Core Training Modes

7.1 GRWO-DPO

GRWO-DPO is the preferred main method when the base model already has decent reasoning ability and follows the desired format.

It teaches:

  • This reasoning direction is better than that one.
  • This theorem choice is better than that theorem choice.
  • This continuation is more grounded than the weaker continuation.
  • This recovery path is better than wandering.
  • This confidence is justified; that confidence is not.

GRWO-DPO preserves the model's natural language style better than pure SFT because it does not force exact imitation.

7.2 GRWO-SFT

GRWO-SFT is useful when the model does not know how to produce a required reasoning move.

It teaches:

  • How to perform a missing transformation.
  • How to recover from a known trap.
  • How to continue after an inconclusive test.
  • How to set up a correct local calculation.

SFT should be used carefully because it can overwrite the model's native style if the teacher traces are too polished or unnatural.

7.3 GRWO-GRPO

GRWO-GRPO is a later-stage method.

It can be used when DPO/SFT are not enough and you want online sampling with a reward function.

Possible reward components:

+ correct final answer, if available
+ correct theorem/test selection
+ valid transformation
+ keypoint coverage
+ forward progress
+ calibrated confidence
- invalid conclusion
- theorem misuse
- wandering
- fake final answer
- malformed protocol tags

8. Recommended First Pipeline

Start with GRWO-DPO-first if the model already follows the format and can produce plausible reasoning.

1. Select a hard problem.
2. Let the model generate a partial reasoning prefix.
3. Stop around a meaningful branch point.
4. Generate an unguided continuation for the next N tokens.
5. Generate a guided/corrected continuation using keypoints, verifier, or teacher.
6. Build a DPO pair:
   prompt   = problem + partial reasoning prefix
   chosen   = guided continuation
   rejected = unguided continuation
7. Train with DPO on only this local continuation window.

Optional:

8. Add a small SFT bucket for cases where the model cannot produce the desired move at all.

9. Window Length

A good starting range:

prefix length:       500–1500 tokens
local window length: 768–1500 tokens

A practical default:

prefix length:       ~1000 tokens
continuation window: ~1500 tokens

The model does not automatically learn to stop at 1500 tokens just because generation is capped there. The cap is an external data-generation/training budget.

However, avoid making accepted samples look like naturally finished answers if they are actually truncated.

For process-window training, it is acceptable for the chosen continuation not to finish the whole problem. The objective is local reasoning continuation, not full solution completion.


10. Loss Scope

The loss should only apply to the local continuation window.

problem tokens:              no loss
partial reasoning prefix:    no loss
next N local window tokens:  loss active
future tokens after window:  not included / no loss

For SFT, this means:

labels for prompt/prefix tokens = -100
labels for continuation tokens  = token IDs

For DPO, the row should be structured so that:

prompt   = prefix
chosen   = continuation window only
rejected = continuation window only

Do not include future tokens after the supervised window, because that leaks information.


11. Good Chosen vs Rejected Design

The chosen and rejected continuations should be as similar as possible except for reasoning quality.

Good pair design:

same prompt
same prefix
similar length
similar format
similar style
similar confidence level visually
only reasoning direction differs

Bad pair design:

chosen = polished teacher solution
rejected = messy model rambling

That would teach surface style instead of reasoning direction.

Better pair:

chosen = model-like continuation, corrected at the branch
rejected = model-like continuation, wrong branch

12. Example Reasoning Trap

Problem type: series convergence.

Partial reasoning prefix:

The absolute series behaves like the harmonic series, so it does not converge absolutely.
Now I need to determine whether the original alternating series converges conditionally.
The alternating series test is awkward because monotonicity is not obvious...

Rejected continuation:

Since the alternating series test fails, the series diverges.

Chosen continuation:

But failure of the alternating series test would only be inconclusive, not proof of divergence.
I should try another method. A useful approach is to rewrite the term as an alternating harmonic component plus an absolutely convergent correction...

This pair teaches the model:

A failed test is inconclusive, not proof of divergence.

That is a local reasoning correction.


13. Reasoning Categories to Target

Good GRWO windows should target common reasoning branch errors.

Series

  • Ratio test equals 1 means inconclusive.
  • Root test equals 1 means inconclusive.
  • Alternating series test failure is inconclusive.
  • Absolute convergence must be checked separately.
  • Conditional convergence requires convergence without absolute convergence.
  • nth-term test only proves divergence when term limit is nonzero.
  • Limit comparison requires positive terms.

Double/Triple Integrals

  • Wrong coordinate bounds.
  • Missing Jacobian.
  • Bad region interpretation.
  • Wrong polar/spherical conversion.
  • Wrong order-change bounds.
  • Ignoring symmetry incorrectly.

General Calculus

  • Invalid theorem use.
  • Endpoint checks missed.
  • Domain restrictions ignored.
  • Sign mistakes in substitution.
  • Confusing necessary and sufficient conditions.

Reasoning Control

  • Wandering after solution is already known.
  • Premature final answer.
  • Overconfident unsupported conclusion.
  • Excessively hesitant correct conclusion.
  • Failure to verify.

14. Dataset Buckets

Recommended first 1000-sample mix:

700 GRWO-DPO branch-point pairs
150 GRWO-DPO finish-vs-wander pairs
100 weak-correct vs strong-correct pairs
50 optional GRWO-SFT missing-move examples

Alternative conservative mix:

80–90% GRWO-DPO
10–20% GRWO-SFT

Use SFT only when the model cannot produce the desired reasoning move at all.


15. Unguided Correct Samples

Do not automatically reject all unguided continuations.

If the unguided continuation is correct, either:

keep it as a positive SFT sample
use it as chosen against a weaker continuation
or skip it

Do not train the model to believe all of its natural reasoning is bad.

The rule:

unguided correct  -> positive or skip
unguided wrong    -> rejected against guided correction
unguided wandering -> rejected against progress continuation

16. Confidence Calibration

DPO can teach confidence in language even when both samples are technically correct.

Example rejected:

This probably converges because it looks alternating.

Example chosen:

The absolute series behaves like the harmonic series, so it is not absolutely convergent.
The remaining signed series can be rewritten as an alternating harmonic term plus an absolutely convergent correction, so it converges conditionally.

Both may reach the same answer, but the chosen one teaches:

  • Better justification.
  • More grounded confidence.
  • Cleaner theorem use.
  • Stronger conclusion control.

The goal is not confidence alone.

The goal is:

calibrated confidence

The model should be decisive when the logic supports it and cautious when a test is inconclusive.


17. Training Configuration Starting Point

For GRWO-DPO on a small high-quality dataset:

learning_rate: 5e-6 to 1e-5
beta: 0.03 to 0.05
num_train_epochs: 1 to 2
prompts_per_step: 5 to 10
dpo_epochs_per_step: 3 to 4
max_length: 3072 to 4096
max_prompt_length: 1024 to 1536
max_completion_length: 1536 to 3072
per_device_train_batch_size: 1
gradient_accumulation_steps: 4 to 8
save_steps: 25
logging_steps: 1

For GRWO-SFT, if used:

learning_rate: 5e-5 to 1e-4
num_train_epochs: 1 to 2
max_length: 3072 to 4096

DPO should usually be softer than SFT because it can over-steer quickly.


18. Evaluation

Evaluate GRWO by behavior, not just loss.

Test whether the model:

  • Avoids known reasoning traps.
  • Recovers from inconclusive tests.
  • Chooses better next steps.
  • Keeps protocol format.
  • Does not hallucinate unsupported conclusions.
  • Does not become too short or robotic.
  • Maintains mathematical accuracy.
  • Knows when to continue and when to finish.

Useful benchmark prompts:

1. Series trap where AST failure is inconclusive.
2. Ratio/root test equals 1 trap.
3. Double integral with coordinate-bound trap.
4. Spherical-coordinate Jacobian trap.
5. Problem where early answer is tempting but wrong.
6. Problem where reasoning should finish instead of wander.

19. Expected Benefits

GRWO should reduce training/data cost because it avoids full-trace generation.

Approximate reduction:

full trace: 10,000 tokens
GRWO window: 1,500 tokens
cost reduction: ~85% target-token reduction

It also produces more training examples per hard problem.

One hard problem can become several local windows:

window 1: prefix 0–600    -> train 600–1500
window 2: prefix 0–1300   -> train 1300–2500
window 3: prefix 0–2200   -> train 2200–3500
finish:   near-end prefix -> train final closeout

Do not over-sample one problem too much. Use roughly 2–4 windows per problem.


20. Risks

20.1 Style Shortcut Learning

DPO may learn cheap differences if chosen and rejected differ too much in formatting, length, or tone.

Mitigation:

make chosen/rejected similar in style and length

20.2 Hint Dependence

If keypoints are visible in the student prompt, the model may depend on information that will not exist at inference.

Mitigation:

keypoints only guide teacher/chosen generation
not student prompt

20.3 Endless Continuation Bias

If only mid-reasoning windows are trained, the model may improve continuation but not termination.

Mitigation:

include 10–20% finish-window samples

20.4 Overconfidence

DPO can accidentally reward assertive but unsupported language.

Mitigation:

reward calibrated confidence, not raw confidence

20.5 Overfitting to Local Windows

The model may learn local moves but still fail globally.

Mitigation:

include multi-window coverage and final-answer evaluation

21. Implementation Sketch

for problem in problems:
    prefix = generate_partial_reasoning(
        model=model,
        problem=problem,
        max_new_tokens=prefix_budget,
        stop_at_branch=True,
    )

    rejected = generate_continuation(
        model=model,
        prompt=problem + prefix,
        max_new_tokens=window_budget,
        guided=False,
    )

    score = evaluate_continuation(
        problem=problem,
        prefix=prefix,
        continuation=rejected,
        answer=gold_answer,
        keypoints=keypoints,
    )

    if score.is_good:
        keep_as_positive_or_skip(problem, prefix, rejected)
        continue

    chosen = teacher_generate_guided_continuation(
        problem=problem,
        prefix=prefix,
        rejected=rejected,
        gold_answer=gold_answer,
        keypoints=keypoints,
        max_new_tokens=window_budget,
    )

    add_dpo_row(
        prompt=problem + prefix,
        chosen=chosen,
        rejected=rejected,
    )

22. Minimal First Experiment

Start small.

problems: 200
windows per problem: 1–2
DPO pairs: 300–400
window length: 768–1500 tokens
prefix length: 500–1500 tokens

Train:

Experiment A: GRWO-DPO only
Experiment B: small GRWO-SFT + GRWO-DPO

Compare:

base/SFT model vs GRWO-DPO vs GRWO-SFT+DPO

Main question:

Does the model choose better reasoning branches under the same token budget?

23. Current Working Definition

GRWO is a process-preference training method where a reasoning model is optimized on local continuation windows from its own partial traces. A guided continuation is preferred over an unguided continuation under the same prefix, with loss applied only to the local window. The goal is to improve reasoning trajectory decisions while preserving the model's native reasoning style and reducing full-trace training cost.

24. Research Notes

GRWO is promising because it combines three useful ideas:

  1. Self-generated prefixes preserve the model's natural reasoning distribution.
  2. Guided continuations inject correct reasoning direction.
  3. Windowed optimization reduces token cost and targets local failure points.

This is especially suitable for math reasoning models where errors often occur at identifiable branch points.

The most important design principle:

Train the next reasoning move, not the entire solution.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support