auto-0.4b โ ONNX
ONNX exports of ProCreations/auto-0.4b,
the 0.4B encoder that decides whether an AI agent's tool call is safe to run.
| file | precision | size |
|---|---|---|
model.onnx |
fp32 | ~1.5 GB |
model_int8.onnx |
dynamic int8 | ~400 MB |
import numpy as np, onnxruntime as ort
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("ProCreations/auto-0.4b-ONNX")
sess = ort.InferenceSession("model_int8.onnx", providers=["CPUExecutionProvider"])
enc = tok(text, return_tensors="np", truncation=True, max_length=8192)
logits = sess.run(None, {"input_ids": enc["input_ids"].astype(np.int64),
"attention_mask": enc["attention_mask"].astype(np.int64)})[0]
p_deny = np.exp(logits[0]) / np.exp(logits[0]).sum()
print("DENY" if p_deny[1] > 0.5 else "APPROVE", p_deny[1])
Build the input string with the exact format documented on the main model card โ proposed call first, then user request, then history.
Context length
Practical limit ~8k tokens. ModernBERT's non-flash attention path materialises a dense
(B, 1, L, L) sliding-window mask, so ONNX memory grows quadratically with sequence length
(at 64k that mask alone would be ~17 GB). Almost all real tool-call decisions are well under
8k. For the full 64k context use the PyTorch + flash-attn path in the main repo.