auto-0.4b โ€” ONNX

ONNX exports of ProCreations/auto-0.4b, the 0.4B encoder that decides whether an AI agent's tool call is safe to run.

file precision size
model.onnx fp32 ~1.5 GB
model_int8.onnx dynamic int8 ~400 MB
import numpy as np, onnxruntime as ort
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("ProCreations/auto-0.4b-ONNX")
sess = ort.InferenceSession("model_int8.onnx", providers=["CPUExecutionProvider"])

enc = tok(text, return_tensors="np", truncation=True, max_length=8192)
logits = sess.run(None, {"input_ids": enc["input_ids"].astype(np.int64),
                         "attention_mask": enc["attention_mask"].astype(np.int64)})[0]
p_deny = np.exp(logits[0]) / np.exp(logits[0]).sum()
print("DENY" if p_deny[1] > 0.5 else "APPROVE", p_deny[1])

Build the input string with the exact format documented on the main model card โ€” proposed call first, then user request, then history.

Context length

Practical limit ~8k tokens. ModernBERT's non-flash attention path materialises a dense (B, 1, L, L) sliding-window mask, so ONNX memory grows quadratically with sequence length (at 64k that mask alone would be ~17 GB). Almost all real tool-call decisions are well under 8k. For the full 64k context use the PyTorch + flash-attn path in the main repo.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ProCreations/auto-0.4b-ONNX

Quantized
(2)
this model