ARC-1

ARC-1 is a small decision model. You give it some context and a question, and it picks an option, gives a score or answers yes/no, with a probability for every answer. It has 1.7B parameters, runs on a normal gaming GPU and answers a short request in about 16 ms.

I built it over 11 days on a single RTX 4060 Ti, to see how close a small local model can get to hosted decision APIs like Jev. It is a lot faster and runs offline, but it is not as accurate. The numbers are below.

Usage

The model has its own small inference package (it is not a plain transformers model):

pip install git+https://github.com/realslapout/arc-1
from arc1 import ARC1Predictor

model = ARC1Predictor("realslapout/ARC-1", cuda_graphs=True)

state = {"ticket": "I was charged twice for the same order and I want my money back."}
questions = {
    "team": {"type": "choice", "instructions": "Which team should handle `ticket`?",
             "criteria": {"billing": "payments, charges and refunds",
                          "shipping": "delivery problems",
                          "tech": "app or login problems"}},
    "urgency": {"type": "score", "instructions": "How urgent is `ticket`?",
                "criteria": ["not urgent", "normal", "urgent"]},
    "angry": {"type": "noul", "instructions": "Is the customer angry?"},
}
answers = model.predict(state, questions)["answers"]
# team -> billing (0.992), urgency -> 1.03 on a 0-2 scale, angry -> 0.97

Question types: choice (pick one of the labels in criteria), score (a level from a list ordered from lowest to highest) and noul (probability that a yes/no statement is true). state can be a string, a dict or a list. The input window is 1,024 tokens.

Results

Decision benchmarks that ARC-1 was not trained on (accuracy, %):

model size JevBench public 231 DecideBench v1.1 speed
ARC-1 1.7B, open 68.4 75.5 25 ms median on an RTX 4060 Ti (local)
Jev 1.13.0 undisclosed, API 86.6 98.0 ~620 ms median (API, network included)
Strands Decider 2B 1.9B, open 72.3 * – 115 ms median on an RTX 3090 *
decider-2b 1.9B, open 71.0 – –
Laya 421M, open 58.4 59.8 –

The other models' JevBench numbers come from the JevBench repository (results v1.2, the same 231 items), and their DecideBench numbers from DecideBench. ARC-1 is measured on DecideBench v1.1 (400 items). * = reported by the authors. Jev is clearly more accurate. ARC-1's advantages are speed, running locally, and being free and open.

On held-out items of tasks that are in the training data: Banking77 87.2, CLINC150 92.0, MASSIVE (English) 84.2, AG News 90.6, MNLI 83.0, BoolQ 82.6, ToxicChat 89.4 (balanced accuracy), prompt injection 82.3 (balanced accuracy). On XSTest, which is not in the training data, it gets 88.9% balanced accuracy. The full table is on GitHub.

Speed (this package, RTX 4060 Ti, batch 1, bf16, CUDA graphs): 16 ms for a short request, 25 ms median on 4-option JevBench items, about 100 ms for a full 1,024-token input. On a CPU it takes about 0.4 s per short request.

Training

Qwen3-1.7B-Base with LoRA (rank 64, merged). Each option is scored in its own branch after a shared encoding of the state and question, so the order of the options doesn't change the result. Training used about 3.8 million examples from around 330 public datasets: intents, topics, sentiment, inference and logic, moderation and safety, prompt injection, support tickets, tool selection and preference data, plus generated decision tasks. Training rows that overlapped evaluation items were removed. This checkpoint interpolates two runs (run 6 plus 0.3 of the difference to run 9). Full data list: DATA.md.

Limitations

Much weaker than large models on hard multi-step decisions. Mostly English. A 1,024-token window. The probabilities are calibrated on our dev data, so check thresholds on your own data. Not for high-stakes decisions without a human in the loop.

License

CC BY-NC 4.0 (non-commercial). About 7% of the training data only allows non-commercial use, and many sources don't state a licence. The base model, Qwen3-1.7B-Base, is Apache 2.0. The inference code on GitHub is Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for realslapout/ARC-1

Finetuned
(444)
this model