Instructions to use harpertoken/flow with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use harpertoken/flow with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="harpertoken/flow", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("harpertoken/flow", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
flow
A GRU classifier over sandbox lifecycle event sequences: eight ordered events in, one of four histories out. Event types embed to 8 dimensions, joined with timestamps and durations into a 32-unit GRU, with a linear head over recover_then_wait, wait_during_outage, delayed_fault and cooldown. Trained for 30 epochs with Adam at learning rate 1e-3, batch size 128, seed 0. Test accuracy 1.0000 on 800 held-out sequences. CPU training.
No transformer here. The PreTrainedModel wrapper exists only so the weights serialize as config.json plus model.safetensors and load through AutoModel (including trust_remote_code via auto_map), the same arrangement as the rest of the estate.
The gate this model passed
Every sequence carries exactly 8 events with identical type and exit-code bags across labels, and identical final state (exit 0, clean). A field-only MLP on aggregates scores 0.2362 and 0.2537 across two seeds, chance level on four classes. The GRU scores 1.0000 on both seeds. Order is therefore sufficient for perfect separation against the declared order-blind baseline, and the replication confirms it.
Order is not the only field that differs. Event durations are weakly label-bearing, because the fixed 2.0s wait sits at a different point in each procedure. An order-blind model given duration aggregates alone reaches 0.6996, and the first execute event duration alone reaches 0.5829. Counts and exit codes carry no such signal: both sit at exactly 0.2500.
Combining every permitted non-order feature does better than any single family. An order-blind model over counts, exit codes, durations, and prefix aggregates together reaches 0.9579. Order is sufficient for perfect separation against the declared counts-and-exit-codes budget, but against the strongest non-temporal predictor the margin is 0.0421, not 0.75. Read this model as evidence that order is a clean and sufficient route to the label on this data, not as evidence that only order can reach it.
Usage
from transformers import AutoModel
from modeling_flow import FlowGRU # registers the architecture
model = AutoModel.from_pretrained("harpertoken/flow")
model.eval()
sequence = [("sandbox_create", 0.0, 0.04), ("execute", 0.04, 0.01)]
print(model.predict_label(sequence))
predict_label takes a list of (type, t, duration) triples and returns the history label. Event types must come from the six-value vocabulary; timestamps and durations are seconds. Needs torch and transformers.
Training
4,000 sequences from the v1-ordering collection (1,000 per procedure), split 3,200 train and 800 test by seeded permutation. The seed-0 weights are published; seed 1 replicated at 1.0000 versus 0.2537. The wrapper was checked for exact equivalence against the trained module in eval mode.
Gate
The publish gate as measured: a field-only MLP on aggregates sits at chance on both seeds, while the ordered GRU scores 1.0000 on both. Same data and same split procedure, with the baseline restricted to counts and exit codes. Reproduce both families with the benchmark in coccinella-labs/sandbox-lifecycle. This is the evidence the model was published on, not an illustration of it.
Limitations
Four controlled histories with identical event multisets. Anything outside that closed vocabulary, longer sequences, new event types, real unsanitized logs, is out of scope. 1.0000 reflects a designed task, not general sandbox understanding.
- Downloads last month
- 46
Evaluation results
- accuracy on sandbox-lifecycle v1ordtest set self-reported1.000
