ARC-1
ARC-1 is a small decision model. You give it some context and a question, and it picks an option, gives a score or answers yes/no, with a probability for every answer. It has 1.7B parameters, runs on a normal gaming GPU and answers a short request in about 16 ms.
I built it over 11 days on a single RTX 4060 Ti, to see how close a small local model can get to hosted decision APIs like Jev. It is a lot faster and runs offline, but it is not as accurate. The numbers are below.
- Code and full write-up: github.com/realslapout/arc-1
- Try it on a free GPU: Colab notebook
Usage
The model has its own small inference package (it is not a plain transformers model):
pip install git+https://github.com/realslapout/arc-1
from arc1 import ARC1Predictor
model = ARC1Predictor("realslapout/ARC-1", cuda_graphs=True)
state = {"ticket": "I was charged twice for the same order and I want my money back."}
questions = {
"team": {"type": "choice", "instructions": "Which team should handle `ticket`?",
"criteria": {"billing": "payments, charges and refunds",
"shipping": "delivery problems",
"tech": "app or login problems"}},
"urgency": {"type": "score", "instructions": "How urgent is `ticket`?",
"criteria": ["not urgent", "normal", "urgent"]},
"angry": {"type": "noul", "instructions": "Is the customer angry?"},
}
answers = model.predict(state, questions)["answers"]
# team -> billing (0.992), urgency -> 1.03 on a 0-2 scale, angry -> 0.97
Question types: choice (pick one of the labels in criteria), score (a level from a list ordered from lowest
to highest) and noul (probability that a yes/no statement is true). state can be a string, a dict or a list.
The input window is 1,024 tokens.
Results
Decision benchmarks that ARC-1 was not trained on (accuracy, %):
| model | size | JevBench public 231 | DecideBench v1.1 | speed |
|---|---|---|---|---|
| ARC-1 | 1.7B, open | 68.4 | 75.5 | 25 ms median on an RTX 4060 Ti (local) |
| Jev 1.13.0 | undisclosed, API | 86.6 | 98.0 | ~620 ms median (API, network included) |
| Strands Decider 2B | 1.9B, open | 72.3 * | – | 115 ms median on an RTX 3090 * |
| decider-2b | 1.9B, open | 71.0 | – | – |
| Laya | 421M, open | 58.4 | 59.8 | – |
The other models' JevBench numbers come from the JevBench repository (results v1.2, the same 231 items), and their DecideBench numbers from DecideBench. ARC-1 is measured on DecideBench v1.1 (400 items). * = reported by the authors. Jev is clearly more accurate. ARC-1's advantages are speed, running locally, and being free and open.
On held-out items of tasks that are in the training data: Banking77 87.2, CLINC150 92.0, MASSIVE (English) 84.2, AG News 90.6, MNLI 83.0, BoolQ 82.6, ToxicChat 89.4 (balanced accuracy), prompt injection 82.3 (balanced accuracy). On XSTest, which is not in the training data, it gets 88.9% balanced accuracy. The full table is on GitHub.
Speed (this package, RTX 4060 Ti, batch 1, bf16, CUDA graphs): 16 ms for a short request, 25 ms median on 4-option JevBench items, about 100 ms for a full 1,024-token input. On a CPU it takes about 0.4 s per short request.
Training
Qwen3-1.7B-Base with LoRA (rank 64, merged). Each option is scored in its own branch after a shared encoding of the state and question, so the order of the options doesn't change the result. Training used about 3.8 million examples from around 330 public datasets: intents, topics, sentiment, inference and logic, moderation and safety, prompt injection, support tickets, tool selection and preference data, plus generated decision tasks. Training rows that overlapped evaluation items were removed. This checkpoint interpolates two runs (run 6 plus 0.3 of the difference to run 9). Full data list: DATA.md.
Limitations
Much weaker than large models on hard multi-step decisions. Mostly English. A 1,024-token window. The probabilities are calibrated on our dev data, so check thresholds on your own data. Not for high-stakes decisions without a human in the loop.
License
CC BY-NC 4.0 (non-commercial). About 7% of the training data only allows non-commercial use, and many sources don't state a licence. The base model, Qwen3-1.7B-Base, is Apache 2.0. The inference code on GitHub is Apache 2.0.
Model tree for realslapout/ARC-1
Base model
Qwen/Qwen3-1.7B-Base