Reflex-1

Fast local decisions over dynamic candidate sets.

Reflex-1 is a 421M-parameter decision model from Good AI Labs. Given a textual state, a question, and candidate answers, it selects an answer and returns probabilities in a single forward pass. Candidates can change with every request, supporting classification, tool routing, and action selection through one interface.

The public download includes the weights, both tokenizers, and the Transformers implementation. The FP32 Safetensors weights are 1.68 GB.

Quickstart

Tested with Python 3.12:

python -m pip install "torch==2.8.0" "transformers==5.17.0" "safetensors==0.8.0"
import torch
from transformers import AutoModel

torch.set_num_threads(1)
model = AutoModel.from_pretrained("gai-labs/reflex-1", trust_remote_code=True)

decision = model.predict(
    state="The customer was charged twice for one card payment.",
    question="Choose the matching issue.",
    options=["duplicate charge", "lost card", "unknown fee", "cash withdrawal"],
)[0]

print(decision.choice)         # duplicate charge
print(decision.probabilities)  # probabilities in candidate order

No access token is required. Load once and reuse the model. Inference defaults to CPU FP32; model.to("cuda") enables CUDA FP32. The included example.py covers single requests and batches. Pin revision= for reproducible loading.

Architecture

Reflex-1 has 420,778,370 parameters across two text encoders, a candidate projection, and a shared scoring head.

Component Parameters Structure and role
Context encoder 394,781,696 28 bidirectional Transformer layers, 1,024 hidden dimensions, and 16 attention heads. Local attention with a global-attention layer every third layer encodes the state and questions.
Candidate encoder 22,713,216 6 Transformer layers, 384 hidden dimensions, and 12 attention heads. Masked mean pooling produces a separate representation for each candidate.
Candidate projection 394,240 A 384 → 1,024 affine projection aligns normalized candidate vectors with the context space, followed by normalization.
Scoring head 2,889,218 A context transform and scaled cosine scores, plus a 256-dimensional evidence path with four-head cross-attention and one four-head Transformer layer over each question's candidate set.

State and questions share one context pass. The scoring head reads evidence from the context and models interactions among candidates, then adds a learned correction to their cosine scores. A softmax returns probabilities over the supplied candidates; the highest-scoring candidate is the decision. Multiple questions can share the same state, with a separate distribution per question.

Demos

Recordings captured on October 3 with an earlier checkpoint. The evaluation below describes the current default weights. Recording details.

Chromium Dino — Native game and keyboard actions, with visible obstacle geometry supplied as text. This run ends in a collision at 42.3 seconds.

ViZDoom — Defend the Center, skill 5. Inputs describe visible objects, health, and ammunition. This episode records 30 kills and ends at 34.5 seconds. Playback follows simulation time, excluding inference delays.

Browser workflow — A synthetic WebGym restaurant task in Chromium, replayed at 4× speed. Reflex selects actions and targets; Qwen3-1.7B supplies text and dropdown values.

Evaluation

The default checkpoint was fixed before these reserved evaluation subsets were scored. The task families are represented in training; subset sizes vary.

Reflex-1 classification accuracy on eight reserved evaluation subsets.

Task Correct Accuracy
AG News 465/500 93.00%
SST-5 299/500 59.80%
Emotion 458/500 91.60%
Banking77 460/495 92.93%
BoolQ 171/202 84.65%
MNLI 436/500 87.20%
RACE 287/500 57.40%
SciQ 123/128 96.09%

Native MiniWoB++ browser tasks completed 21/30 episodes with the fixed controller. Chromium Dino averaged 29.87 seconds across 16 native attempts; none reached the 60-second target. These are specific controller evaluations.

The public 231-item JevBench diagnostic scored 117/231 (50.65%) using a candidate-scoring adaptation. These previously inspected public items are separate from the official private benchmark.

Evaluation data and scope include per-task regressions and incomplete game/WebGym checks.

CPU latency

Intel Xeon Platinum 8558, FP32, batch size 1. Each thread setting covers 480 calls over 120 inputs in two fresh processes. Timings include tokenization and inference, with loading and warmup excluded and cross-request caches disabled.

Reflex-1 CPU latency with median and p95 at one and four compute threads.

2.17 GiB peak process RSS, including imports, loading, warmup, and inference. Runs used one or four PyTorch intra-op threads, one inter-op thread, and matching BLAS limits of one or four. Two and eleven total OS threads, respectively, were observed after scoring; those counts include runtime helpers.

Latency data · Resource measurements. These shared-host timings use the evaluation runtime; the public Transformers interface is verified separately for loading and inference parity.

Usage notes

Reflex-1 scores English text; the application handles observations, tool arguments, and action execution. Each question accepts 1–255 candidates, with a 512-token limit per candidate and a 2,048-token request budget. State text is capped at 512 tokens and may be shortened further; check state_truncated. Probabilities are relative to the supplied candidates and require calibration before use as confidence thresholds.

License

Code and weights: Apache 2.0. Training used RACE and SciQ, whose source terms include noncommercial restrictions; commercial-use clearance has not been established. See the source terms, training sources, and attributions.

Release metadata · Checksums · Runtime 1.0.0 · Good AI Labs

Downloads last month
109
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train gai-labs/reflex-1

Evaluation results