YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Model card: s1-workflow-tiny

Model details

  • Name: movahedi-ca/s1-workflow-tiny
  • Architecture: tiny transformer encoder (4 layers, d_model 128, 4 heads) with a pointer head over the action menu. 1,382,017 parameters.
  • Input: tokenized workflow state plus the valid-action menu, per the frozen specs/token-schema.json (encoding version 1.0.0). 16-bit token ids plus per-token field objects.
  • Output: a distribution over menu slots. The argmax is the chosen action. The model never sees a fixed action list: action names arrive in the input and are embedded through a stable string hash, so the same weights serve new recipes with new action names.
  • Format: ONNX (opset 17, no control-flow or custom ops), run with onnxruntime-web in the browser. CPU-friendly, no network at inference. Weights live in the companion file model.onnx.data (external data); keep both files together.
  • License: MIT.

Intended use

A System-1 reflex for guided workflow tools: given the current state and the valid actions, pick the next action. First instantiation: the Law 25 data mapping guide. The model is a reusable module, not a single-task brain.

Training data

Scripted-teacher demonstrations (Phase 4): the teacher is code, not an LLM. 526 labeled steps in 33 shards; 152 steps (29%) are adversarial recovery sessions (flag, undo, skip, abort) so the model learns to recover from corrupted state. Domains: 319 data-mapping/* steps plus 207 steps across 10 out-of-domain domains, so out-of-domain generalization is measurable. Dataset: movahedi-ca/ai-data-map-teacher (shard format 1.0.0, see training/SHARD-FORMAT.md in the code repo).

Training run

  • 12 epochs, batch 256, lr 3e-4, CPU. Final train loss 0.2145, validation accuracy 0.9375 (48 held-out steps, split by recipe).
  • ONNX export verified: max torch-vs-onnxruntime absolute difference 1.19e-06 on a sample batch.
  • Note: training ran on CPU rather than the planned Kaggle GPU because the Kaggle account requires phone verification to enable internet/GPU in notebook sessions. The model is small enough that CPU training completed in under a minute with no change to the training script.

Evaluation

Held-out results (eval-report.json):

  • In-domain menu-choice accuracy: 0.9561 on 319 steps (gate: 0.85)
  • Out-of-domain accuracy on unseen domains with unseen action names: 0.9420 on 207 steps (gate: 0.60) โ€” the modular/dynamic proof
  • Menu-permutation robustness: 0.9561 on 319 steps (drift 0.00, tolerance 0.10) โ€” the model reads the menu, it does not memorize positions

All three gates passed. See training/eval.py in the code repo for the evaluation harness.

Limitations

  • The model picks among the actions it is given; it cannot invent actions.
  • It is a reflex, not a planner: one step at a time, no lookahead.
  • Recovery behavior is only as good as the corrupted sessions in training.
  • Any change to the token encoding version requires retraining.

Files

  • model.onnx + model.onnx.data: the weights for in-browser inference
  • checkpoint.pt: PyTorch training checkpoint
  • config.json: architecture hyperparameters
  • eval-report.json: the gate results above, machine-readable
Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support