NextAction: connect4

A small, auditable classifier that predicts the column to play in Connect Four. Given the current state, it returns one action from a fixed vocabulary with a probability for each, so an agent can act alone when it is confident and hand the decision to a person when it is not. It ships with a decision policy (config.toml) that blocks impossible actions and sets a minimum confidence per action.

Demo

connect4 demo

In 200 games against a depth-4 search, the model alone wins 19%. With moves below 0.5 confidence reviewed by a depth-6 search it wins 40%, at its teacher's level (38%), and 13 games more than when the depth-4 teacher reviews instead; the 95% intervals overlap, so the last difference is suggestive. Scripts, benchmark and analysis: demos/connect4.

Model details

Developed by Luiz Araujo
Model type TF-IDF features + multinomial logistic regression (scikit-learn)
Task next-action prediction for task-oriented agents (text classification)
Language English
License mit
Size 222 kB
Code github.com/leduardoaraujo/NextAction

Uses

Direct use. Choose a Connect Four column in the NextAction demo, as an example of a model whose unsure moves are reviewed by a stronger search.

Out of scope. Competitive play: alone the model is much weaker than the search it imitates; it needs the reviewer to play at that level.

How to get started

from huggingface_hub import hf_hub_download
import skops.io as sio

model = sio.load(hf_hub_download("ludolua/nextaction-connect4", "model.skops"))  # no types outside skops' trusted list

state = "[blocked none] [danger none] [features c1_better c2_better c3_better c4_good c5_better c6_better c7_better hc1_0 hc2_0 hc3_0 hc4_0 hc5_0 hc6_0 hc7_0]"
probabilities = dict(zip(model.classes_, model.predict_proba([state])[0]))
print(max(probabilities, key=probabilities.get), probabilities)

Dependencies: scikit-learn, skops and huggingface_hub. Use a scikit-learn version compatible with the training one. The full toolkit (CLI, REST API, decision log, evaluation, policy) is in the NextAction repository.

Input format

A free-text description of the current situation in English, for example "[blocked none] [danger none] [features c1_better c2_better c3_better c4_good c5_better c6_better c7_better hc1_0 hc2_0 hc3_0 hc4_0 hc5_0 hc6_0 hc7_0]".

States are built by features.describe() in the demos/connect4 folder; use it to turn a position into the text this model expects.

Actions

Action Meaning
c1 drop a disc in column 1 (leftmost)
c2 drop a disc in column 2
c3 drop a disc in column 3
c4 drop a disc in column 4 (center)
c5 drop a disc in column 5
c6 drop a disc in column 6
c7 drop a disc in column 7 (rightmost)

Training details

Data. 32642 positions labeled with the move of a depth-4 alpha-beta search, from games generated by demos/connect4/generate_data.py. The training data is not redistributed here; download it from the original source.

Preprocessing. States are normalized to the English tag format, then vectorized with TF-IDF over word unigrams and bigrams with accents stripped.

Hyperparameter Value
ngram_range (1, 2)
min_df 2
C 1.0
class_weight None
max_iter 1000
calibration none

The solver is L-BFGS, run in warm-started chunks so training progress can be reported. Hyperparameters live in config.toml and can be searched with nextaction tune.

Evaluation

Method. separate test split (connect4_test.csv), with 6716 examples evaluated.

Results. Accuracy 0.668 and macro F1 0.63; always choosing the most frequent action would give 0.282. Expected calibration error: 0.042.

Action Precision Recall F1 Support
c1 0.625 0.533 0.575 614
c2 0.612 0.572 0.592 795
c3 0.639 0.655 0.647 1087
c4 0.765 0.866 0.812 1894
c5 0.626 0.58 0.602 1039
c6 0.642 0.586 0.613 742
c7 0.568 0.572 0.57 545

Review thresholds. Decisions with confidence below the threshold go to a person:

Threshold Decisions automated Accuracy on automated
0.3 96.1% 68.3%
0.4 82.0% 73.5%
0.5 62.7% 81.1%
0.6 48.3% 88.2%
0.7 37.2% 93.3%
0.8 27.7% 96.9%
0.9 19.1% 98.7%

Decision policy

config.toml holds the decision policy used with this model: rules that block impossible actions, a minimum confidence per action, risk levels and error costs, in units of one human review (review_cost = 1). nextaction calibrate-policy re-fits the thresholds from the costs.

Action Risk Min. confidence Unreviewed error cost Confirm with user
c1 low 0.5 1 no
c2 low 0.5 1 no
c3 low 0.5 1 no
c4 low 0.5 1 no
c5 low 0.5 1 no
c6 low 0.5 1 no
c7 low 0.5 1 no

Rules:

  • no_full_columns: deny when
  • no_gifting_a_win: deny when

Simulated on the evaluation data above, the policy automates 64.8% of decisions at 80.0% accuracy, with an expected cost of 0.4821 reviews per decision.

Apply it with the NextAction package (pip install git+https://github.com/leduardoaraujo/NextAction):

import tomllib
from nextaction.policy import apply, from_dict

with open(hf_hub_download("ludolua/nextaction-connect4", "config.toml"), "rb") as file:
    policy = from_dict(tomllib.load(file)["policy"])
decision = apply(probabilities, state, policy)  # probabilities and state from the snippet above
print(decision["action"], decision["needs_review"], decision["reasons"])

Bias, risks and limitations

  • The model matches word patterns. It does not understand business rules or verify facts, so the calling application must check preconditions before acting.
  • Logistic regression probabilities are not calibrated; pick thresholds from the tables above, and re-evaluate on your own traffic before automating.
  • Imitation accuracy is 0.67, and in a tactical game a single bad move can lose; several columns are often equally good, so accuracy understates move quality.
  • The action vocabulary is fixed; new actions require new labeled data.

Citation

@software{araujo2026nextaction,
  author = {Araujo, Luiz},
  title  = {NextAction: a small, auditable next-action model for AI agents},
  year   = {2026},
  url    = {https://github.com/leduardoaraujo/NextAction}
}

Model card contact

Open an issue in the NextAction repository.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results