Instructions to use Atlas-AI-research/reflex-reason-2b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Atlas-AI-research/reflex-reason-2b with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
Reflex Reason 2B
Reflex Reason is the thinking model behind Reflex, a local browser agent by Atlas AI. It looks at a screenshot of the page together with its accessibility tree, reasons step by step about what to do next, and then gives exactly one action.
In Reflex, the fast decision model Reflex Instinct 0.6B handles every step it is confident about, in about 0.2 s. When Instinct's calibrated confidence falls below a threshold, the step goes to Reason, which takes a few seconds and sees the page.
- Sees the page: it reads a screenshot and the accessibility tree together
- Thinks first: it writes its reasoning inside
<think>, then gives one action in triple backticks - Runs locally: a LoRA adapter on Qwen3-VL-2B-Instruct that fits on a consumer GPU
- Better together: routing Instinct and Reason scored 52.4% step accuracy, against 26.8% for Instinct alone
Code: github.com/leemadov/reflex: the Reflex browser agent that runs this model (Chrome extension and desktop app), plus the training, evaluation and JevBench scripts.
Results
Measured on held-out steps after 1,244 training steps (6.9 hours on one RTX 5070):
| Set | Steps | Operation | Element | Full step | Latency p50 |
|---|---|---|---|---|---|
| Multimodal-Mind2Web (screenshot + tree) | 100 | 98% | 77% | 76% | 3.6 s |
| NNetNav live web (tree only) | 81 | 55.6% | 42.0% | 34.6% | 6.9 s |
| NNetNav WebArena sites (tree only) | 69 | 43.5% | 33.3% | 29.0% | 5.5 s |
Routing (the way Reflex uses it). Instinct decides first, and any step where its confidence is below 0.2 goes to Reason. On the same 250 held-out steps, this scored 52.4% step accuracy against 26.8% for Instinct alone, calling Reason on 59% of steps. A threshold of 0.2 tied for the best accuracy, and higher thresholds only add Reason calls.
How to read this:
- Mind2Web is the easiest of the three. The test set is 5% of Multimodal-Mind2Web's tasks, held out from its train split, so its websites can overlap with training. Pages were also cut to a window of the tree that contains the target. That is not Mind2Web's official cross-website benchmark.
- The NNetNav labels are noisy. They come from an LLM explorer, so a step often has several reasonable next actions and the logged one is only one of them.
Usage
The model loads with standard transformers and peft. This example runs as written (tested on CPU):
import torch
from peft import PeftModel
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
BASE, ADAPTER = "Qwen/Qwen3-VL-2B-Instruct", "Atlas-AI-research/reflex-reason-2b"
processor = AutoProcessor.from_pretrained(BASE)
model = Qwen3VLForConditionalGeneration.from_pretrained(BASE, dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(model, ADAPTER).merge_and_unload().eval()
SYSTEM = ("You are a web agent. Given the task, the URL, your previous actions, the page's accessibility tree and "
"possibly a screenshot, think step by step about what to do next, then give exactly one action in triple "
"backticks. Actions: click [id], type [id] [text], select [id] [option], hover [id], press [key_comb], "
"scroll [down|up], goto [url], go_back, go_forward, new_tab, tab_focus [index], close_tab, "
"stop [answer] (stop [N/A] if the task is impossible).")
state = ("task: Find reviews for a product in the Electronics category.\n"
"url: http://shop.local/\n"
"history: none\n"
"page:\nRootWebArea 'One Stop Market'\n\t[227] link 'My Account'\n\t[815] menuitem 'Electronics'\n\t[272] combobox 'Search'")
prompt = (f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n"
f"{state}<|im_end|>\n<|im_start|>assistant\n<think>\n")
inputs = processor(text=[prompt], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=384, do_sample=False)
print(processor.decode(out[0, inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Output:
Let's think step-by-step. The current webpage is the main page of the One Stop Market, and the objective is to find
reviews for a product in the Electronics category. The current webpage has a menuitem 'Electronics' with id [815].
This menuitem seems to be related to the Electronics category. I will click on this menuitem to navigate to the
Electronics category page.
</think>
```click [815]```
With a screenshot. Resize the browser window's screenshot to fit inside 1024×640, which is what the model was trained on (640 image tokens), padding with black instead of stretching. Then put the image marker at the start of the user turn and pass the image to the processor:
from PIL import Image
image = Image.open("page.png").convert("RGB") # already fitted to 1024x640
prompt = (f"<|im_start|>system\n{SYSTEM}<|im_end|>\n<|im_start|>user\n"
f"<|vision_start|><|image_pad|><|vision_end|>\n{state}<|im_end|>\n<|im_start|>assistant\n<think>\n")
inputs = processor(text=[prompt], images=[image], return_tensors="pt").to(model.device)
Input format:
- The state is
task:,url:,history:(one previous action per line as- ..., ornone) andpage:. - The page should be a WebArena/BrowserGym accessibility tree: one element per line as
[id] role 'name' properties, tab-indented by depth. - Prompts were up to 7,168 tokens in training. Trim long pages by whole lines.
Parsing the answer: take the last triple-backtick block, for example click [815] or type [272] [usb hub].
As a decision model (Jev wire format)
serve.py in this repo serves Reason as a Jev-style decision model on /v1/systemone (TypeSafe-compatible):
pip install torch transformers peft fastapi uvicorn
python serve.py --port 8766
For each question it thinks first (greedy, up to 384 tokens), then letters the options A, B, C... and reads its score
for each letter after </think> and the opening triple backticks, in one forward pass. A softmax at temperature 3.176 turns those into a probability
for every option. That temperature was fitted (lowest log-loss) on JevBench's 231 public items; --temperature 1
gives the raw model.
On those 231 public items, through JevBench's own typesafe adapter, it scored easy 100%, standard 87.5% and hard
47.7%, at a median 5 s per decision on an RTX 5070. To check the temperature, it was fitted on one half of the items
and tested on the other half, and the other way round. These are our own measurements, not JevBench's.
Training
- Base model:
Qwen/Qwen3-VL-2B-Instructwith a LoRA adapter (rank 32, alpha 64) on the language model's attention and MLP projections. The vision tower is frozen. - Text steps: NNetNav (nnetnav-wa, nnetnav-live), using each step's step-by-step reasoning as the target thought.
- Screenshot steps: Multimodal-Mind2Web, 5,660 steps. Screenshots were cropped to a 1280×800 window holding the target and resized to 1024×640. They made up 30% of training steps.
- Target format: the reasoning, then
</think>, then the action in triple backticks, after a prompt that ends in<think>. The loss was on the answer only. - Run: 1,244 optimizer steps (8 sequences each, learning rate 1e-4) in 6.9 hours on one RTX 5070.
Limitations
- Short template on screenshot steps. On Mind2Web-style screenshot pages its reasoning often copies Mind2Web's short template ("the next step on this page is to click [id]") instead of really reasoning.
- Slow by design. It takes 3–7 s per step, so it is meant for the uncertain steps, not every step.
- Learned an explorer's habits. The training labels come from an LLM explorer and from Mind2Web annotators, so it learned their habits along with the task.
- English only.
- Keep a person in the loop. Never let an agent type passwords or payment details on its own, and keep a person in the loop for anything irreversible.
License
OpenRAIL. Reason was trained partly on Multimodal-Mind2Web, which is released under OpenRAIL, so its use restrictions carry over to this model. The base model (Qwen3-VL-2B-Instruct) and NNetNav are Apache 2.0.
Built by Atlas AI.
- Downloads last month
- 30
Model tree for Atlas-AI-research/reflex-reason-2b
Base model
Qwen/Qwen3-VL-2B-Instruct