3dttt-clm-ft1

Fine-tuned choice head for CLM at 4ร—4ร—4 (4-in-a-row) 3D tic-tac-toe.

This is the ft1 round (2026-10-01): the CLM_v0.1-8B encoder was frozen and only the state/action head was trained on teacher-oracle-labeled self-play.

This repository hosts only the head (best_head.pt). It is not a standalone model โ€” it must be served on top of the CLM base with the clm-serve tooling from the training workspace.

Files

file size md5 sha256
best_head.pt 75,557,470 B 911b68ca503fc6f1a326233b7107fac0 c37c2534b80d4ab5aac6be0ff9323e3b6430f784770b3ef4c61d49a99b648934

Training

  • Data: oracle-labeled 3D tic-tac-toe self-play (3dttt-v1): 44,762 unique positions (99,648 with block/win/threat boosting), 179,366 questions; 5,017 held-out test positions.
  • Config: frozen CLM encoder, head-only training; best epoch 20 (val acc 0.9412); ~51 min wall time on one GPU pod.
  • Selection: best held-out validation accuracy. The file was recovered from the training pod after the run; the hashes above pin its provenance.

Metrics (held-out test)

model test acc answer threat_check soft CE
init (zero-shot) 0.2925 0.2511 0.3339 1.5312
majority class 0.2973 โ€” โ€” โ€”
ft1 best_head 0.8917 0.7833 1.0000 0.3096

Targeted nuke double-threat probe: nuke probability 0.993.

Head-to-head vs JEV (valid re-run, 2026-10-04)

40 games, 10 seeds, fail-fast runner, 0 invalid moves:

Games 3dttt-clm-ft1 JEV
Overall 40 24W-13L-3D (win-equiv 64%, 95% CI 48โ€“77%) 13W-24L-3D
as X (first) 20 20W-0L-0D 0W-20L
as O (second) 20 4W-13L-3D 13W-4L-3D

An earlier ft1-vs-JEV series (2026-10-01) was invalidated (stale server environment silently fell back to random moves) and must not be cited.

Usage

wget https://huggingface.co/DeKodez/3dttt-clm-ft1/resolve/main/best_head.pt
# serve on top of CLM_v0.1-8B (training workspace tooling):
clm-serve --ckpt best_head.pt
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support