M1: Towards a Generalist Agentic Model for Autonomous Machine Learning Engineering

🤗 Hugging Face

M1 is a 35B Mixture-of-Experts agentic model for autonomous machine learning engineering. It organizes experiments, writes and revises code, and improves solutions through sustained interaction with execution environments. Hierarchical agentic SFT develops general tool-use and domain-specific optimization skills, while Windowed Interaction Reinforcement Learning (WIRL) trains multi-turn decisions using feedback from completed experiments.

Highlights

  • Long-horizon optimization: Plans experiments, interprets execution feedback, revisits previous solutions, and selects submission candidates within a persistent workspace.
  • Strong MLE performance: Achieves a 72.7% medal rate on MLE-Bench Lite, matching GLM-5.3 and Kimi-K3 under the M1 Harness.
  • General-purpose harness compatibility: Achieves a 69.7% medal rate with Claude Code, compared with 42.4% for its base model.
  • Scientific task optimization: Matches published scientific results on seven of ten NatureBench Lite tasks and surpasses them on five.

Performance

MLE-Bench Lite

We primarily evaluate our trained models on MLE-Bench Lite, a 22-task subset of MLE-Bench consisting of historical Kaggle competitions for end-to-end machine learning engineering. We report the percentage of tasks awarded each medal tier (Gold, Silver, and Bronze), the percentage receiving any medal (Any Medal), and the percentage exceeding the median human score (Med+), with Any Medal as the primary metric. All five metrics are computed over the 22 tasks and reported as percentages.

The first five metric columns use the M1 Harness; the last two columns report the general-purpose harness and its medal rate. All values are percentages.

MLE-Bench Lite main results (Table 1 from the paper)

M1 achieves the highest medal rate among same-scale models and reaches the performance level of larger frontier models. Under the unified M1 Harness, M1 raises the medal rate of its Qwen3.6-35B-A3B base model from 37.9% to 72.7%, achieving the highest rate among the same-scale models evaluated and performance comparable to GLM-5.3 and Kimi-K3. These results show that long-horizon agentic post-training substantially improves the base model's MLE capability, bringing a 35B model to a performance level comparable to larger frontier models.

M1 retains its advantage over the base model under a general-purpose coding harness. With Claude Code, M1 achieves a 69.7% medal rate, compared with 42.4% for its base model, and exceeds the strongest evaluated same-scale baseline by 22.7 percentage points. The consistent advantage over the base model across both harnesses suggests that the improvement from post-training is not specific to the M1-Harness and remains effective in a general-purpose agentic coding environment.

NatureBench Lite

We also evaluate on NatureBench, following the ten-task Lite subset used in Frontis-MA1. M denotes Match-SOTA (g ≥ 0), while S denotes Surpass-SOTA (g > 0.1), which serves as a key metric for directly assessing agents' ability to surpass published scientific results.

NatureBench Lite main results (Table 2 from the paper)

On NatureBench Lite, M1 improves scientific task performance over the base model. With Claude Code, M1 raises the Surpass-SOTA rate from 0.0% to 50.0%, exceeding Nex-N2.5-mini at 20.0% and matching GLM-5.3 at 50.0%. Its Match-SOTA rate also improves from 40.0% to 70.0%. These results show that post-training improves scientific task optimization, enabling M1 to match or exceed published reference results on more tasks.

Usage

The commands below load DubbyDu/M1 and serve an API endpoint at http://localhost:8000/v1.

SGLang

Install SGLang with uv:

uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
uv pip install sglang

Text Generation

python -m sglang.launch_server \
    --model-path DubbyDu/M1 \
    --host 127.0.0.1 \
    --port 8000 \
    --tp-size 2 \
    --context-length 262144 \
    --reasoning-parser qwen3 \
    --language-model-only

Tool Use

python -m sglang.launch_server \
    --model-path DubbyDu/M1 \
    --host 127.0.0.1 \
    --port 8000 \
    --tp-size 2 \
    --context-length 262144 \
    --reasoning-parser qwen3 \
    --tool-call-parser qwen3_coder \
    --language-model-only

vLLM

Install vLLM with uv:

uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
uv pip install vllm --torch-backend=auto

Text Generation

vllm serve DubbyDu/M1 \
    --host 127.0.0.1 \
    --port 8000 \
    --tensor-parallel-size 2 \
    --max-model-len 262144 \
    --reasoning-parser qwen3 \
    --language-model-only

Tool Calling

vllm serve DubbyDu/M1 \
    --host 127.0.0.1 \
    --port 8000 \
    --tensor-parallel-size 2 \
    --max-model-len 262144 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --language-model-only

Recommended Sampling Parameters

The two M1 inference profiles use the same sampling parameters and differ in whether thinking is enabled. The machine-readable profiles are provided in inference_profiles.json.

Parameter Thinking Non-thinking
temperature 1.0 1.0
top_p 0.95 0.95
top_k 20 20
min_p 0.0 0.0
presence_penalty 1.5 1.5
repetition_penalty 1.0 1.0
enable_thinking true false

Use thinking mode for coding and experimental reasoning. Non-thinking mode provides direct responses for planner or structured-output calls.

Thinking Modes

Set chat_template_kwargs.enable_thinking to true for thinking mode or false for non-thinking mode. The deployment commands above use the qwen3 reasoning parser and, for tool use, the qwen3_coder tool-call parser.

Downloads last month
-
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support