logo

1. Introduction

We're introducing GRM-3.2-Cliff, our intermediate model built for long-horizon agentic tasks and extremely difficult reasoning problems in local environments. GRM-3.2-Cliff marks a substantial leap in long-horizon task capability over its predecessor, GRM-2.5-Plus, and is designed to serve as a dependable engine for complex, multi-step local workflows.

The model is purpose-built for long-horizon agentic tasks and problems that are simply hard — difficult coding challenges, advanced mathematics, and rigorous logical reasoning. GRM-3.2-Cliff aims to sustain coherent, goal-directed behavior over extended interactions while remaining optimized for resource-constrained hardware, making it ideal for developers and researchers who need local execution without sacrificing multi-step planning and self-correction performance.

2. Key Capabilities

  • Long-Horizon Agentic Mastery: GRM-3.2-Cliff is specifically optimized to maintain coherence, planning quality, and task fidelity across long, multi-step agentic workflows, representing a major upgrade over GRM-2.5-Plus.
  • Local Workflow Efficiency: Engineered to run smoothly in lower GPU environments while delivering high-tier reasoning performance.
  • Elite Reasoning on Hard Problems: Strong performance on difficult coding, advanced mathematics, and logical reasoning tasks with careful, structured step-by-step problem-solving.
  • Robust Coding Ability: Handles complex, multi-file coding tasks, debugging, refactoring, and long-running terminal sessions locally.
  • Consistent Logical Reasoning: Built to reason carefully through multi-constraint logic problems without losing track of intermediate steps over extended execution runs.

3. Performance

GRM-3.2-Cliff is designed as our premier mid-sized model for local, long-horizon agentic work. It builds directly on the strengths of GRM-2.5-Plus while targeting common edge-case failures in smaller models — contextual drift, multi-step degradation, and loss of initial goal states — delivering strong reliability across extended sessions.

Agentic Performance Evaluation

Detailed Benchmarks

GRM-3.2-Cliff GRM-2.5-Plus GPT-5.6-Luna Sonnet 5 Gemini 3 Pro
Knowledge & STEM
MMLU-Pro 83.3 84.2 89.8
GPQA Diamond 82.4 82.7 92.3 91.9
Reasoning & Coding
LiveCodeBench v6 69.3 67.2 82.9
General Agent
SWE-bench Verified 70.3 85.2 76.2
SWE-bench Pro 43.4 62.7 63.2
Terminal-Bench 2.1 45.3 84.7 80.4
NL2Repo 28.5

Scores are taken from each provider's own published model card, blog post, or system card where available; "—" indicates a score was not publicly reported by that provider at the time of writing. Different labs may use different agent scaffolds when reporting SWE-bench and Terminal-Bench results, so cross-provider comparisons should be read with that caveat.

4. Family

The GRM-3.2 family is available in various sizes to suit every use case.

Model Size Domain
GRM-3.2-Sky 35B-A3B Flagship model for long-horizon tasks
GRM-3.2-Cliff 9B Capable model for low GPU environments
GRM-3.2-Turf 1.2B Lightweight model for practical reasoning

5. Architecture

GRM-3.2-Cliff is built on the Ornith-1.0-9B architecture, a 9B-parameter model optimized for long-horizon agentic workflows, complex coding tasks, advanced mathematics, and logical reasoning, structured to run efficiently in low-to-mid GPU hardware environments.


GRM-3.2-Cliff is developed by OrionLLM and released under the Apache 2.0 License.

Downloads last month
41
Safetensors
Model size
1.47M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 2 Ask for provider support

Model tree for OrionLLM/GRM-3.2-Cliff

Finetuned
(30)
this model

Collection including OrionLLM/GRM-3.2-Cliff