auto-200m-2-int4
This is auto-200m-2 with its weights stored as 4-bit integers. The file is 77 MB instead of 299 MB, about a quarter of the size. It gets 2,884/3,000 on the benchmark, with 60 false approvals and 56 false denials. The BF16 model gets 2,890, with 53 and 57. 44 of the 3,000 decisions differ from the BF16 model.
auto-200m-2 is a 149.6M-parameter ModernBERT classifier. It reads an AI agent's proposed tool call, the user's request and the agent's history, then answers approve or deny. It takes up to 65,536 tokens of context. This version loads through one small Python file, auto_quant.py, in plain PyTorch, without compiled kernels. It was checked on NVIDIA CUDA, the Apple M4 Max CPU and Apple MPS. If you want exactly the BF16 model's decisions at half the size, use auto-200m-2-int8 (150 MB).
Results
These results use the pinned 3,000-item Approve-or-Deny benchmark (revision a38b6259), full input lengths, P(deny) >= 0.5, and the same evaluation code as the base model card. The quantized rows were computed on CUDA in BF16 from the stored integer weights. A false approval is an unsafe call that was approved. A false denial is an authorized call that was denied.
| Model | Weights file | Accuracy | False approvals | False denials | AUROC | 16k–64k tokens | Validation audit | Decisions that differ from BF16 |
|---|---|---|---|---|---|---|---|---|
| auto-200m-2 (BF16) | 299 MB | 96.33% (2890) | 53/1401 | 57/1599 | 0.9937 | 94.14% | 98.34% | — |
| auto-200m-2-int8 | 150 MB | 96.33% (2890) | 53/1401 | 57/1599 | 0.9937 | 94.14% | 98.34% | 2 |
| auto-200m-2-int4 | 77 MB | 96.13% (2884) | 60/1401 | 56/1599 | 0.9933 | 94.98% | 98.15% | 44 |
The 16k–64k column covers 239 benchmark items. The validation audit is 2,595 validation rows that were never used for training or selection. Paired with the BF16 model on the same items, accuracy differs by -0.20 points (95% interval -0.63 to +0.23; 25 items right only for BF16, 19 right only for int4; exact McNemar p = 0.45). On the tool probes it gets 24/24 published skills/MCP/custom-tool probes and 38/40 fresh scope/history/injection probes; the BF16 model gets 24/24 and 38/40. Breakdowns by category, language, difficulty and length are in eval_results.json, and per-item logits are in benchmark_predictions.npz.
Two benchmark looks
This model was benchmarked twice, and both results are reported here. The first frozen export was plain post-training quantization (PTQ), and it scored 2,876/3,000 with 73 false approvals and 51 false denials (-0.47 points against BF16, 52 decisions changed). That result prompted one more QAT run, described below and recorded in the training plan when it was launched. Its best export beat PTQ under the same validation rule, so it replaced the PTQ export, was benchmarked once, and is what's published. Because a benchmark result prompted the retry, the published score is not a clean single look. The first look's logits are in eval/first_look_benchmark_predictions.npz.
Devices
| Device | Compute dtype | Items | Accuracy on those items | Same items, CUDA | Decisions that differ from CUDA | Largest P(deny) difference |
|---|---|---|---|---|---|---|
| NVIDIA RTX PRO 6000 (CUDA) | BF16 | all 3,000 | 96.13% | reference | — | — |
| NVIDIA RTX PRO 6000, public loader (both attention modes) | BF16 | 300 under 2,048 tokens | — | — | 1 | 0.022 |
| Apple M4 Max CPU | FP32 | 2,452 up to 2,048 tokens | 96.45% | 96.41% | 3 | 0.043 |
| Apple M4 Max GPU (MPS) | FP32 | 2,761 up to 16,384 tokens | 96.27% | 96.23% | 3 | 0.043 |
- CUDA row: the benchmark run above.
- Public loader row: 300 benchmark items were scored again with
auto_quant.loadon CUDA, once with memory-linear attention and once with PyTorch SDPA, and compared with the benchmark run. - Mac rows:
auto_quant.loadandscorein FP32 on an M4 Max (macOS). The CPU pass used the 2,452 items up to 2,048 tokens. The MPS pass used the 2,761 items up to 16,384 tokens.
The P(deny) differences come from BF16 on CUDA versus FP32 on the Mac. AMD ROCm should work, since the loader is plain PyTorch, but it wasn't tested. The details are in eval/mac_verify.json.
Memory and speed on the same Mac, one request at a time. All three use the same memory-linear attention:
| Model | Device | Memory after load | Median, under 1k tokens | 7,449 tokens | 29,233 tokens |
|---|---|---|---|---|---|
| auto-200m-2 (BF16 file, runs in FP32) | CPU | 806 MB | 68 ms | 1.65 s | 15.7 s |
| int8 | CPU | 379 MB | 70 ms | 1.49 s | 13.4 s |
| int4 | CPU | 305 MB | 104 ms | 1.51 s | 12.5 s |
| auto-200m-2 (BF16 file, runs in FP32) | MPS | 570 MB | 24 ms | 0.46 s | 4.4 s |
| int8 | MPS | 143 MB | 30 ms | 0.47 s | 4.1 s |
| int4 | MPS | 74 MB | 38 ms | 0.47 s | 4.3 s |
Each model and device ran in its own process. Latency is the median over 110 short benchmark items. On CPU, memory is how much the process grew during loading. On MPS, it's the tensors held on the GPU. The BF16 checkpoint runs in FP32 on CPU and MPS, which is Transformers' default there.
On the GPU, the int8 and int4 weights take about 4× and 8× less memory than the FP32 copy. They don't make short requests faster, though: each layer turns its integer weights back into floats on every call, and unpacking 4-bit values costs extra, so int4 is the slowest on short inputs. On long inputs, activations dominate both time and memory; at 29,233 tokens the CPU peak was 2.1 / 1.8 / 1.7 GB (BF16 / int8 / int4). See eval/speed_mac.json.
How it was made
- Format. Weights are stored as symmetric int4 in [-8, 7], packed two per byte, with one FP16 scale for every 128 consecutive input weights. The token embeddings, every attention and MLP linear layer, and
head.denseare quantized. The LayerNorm weights and the final 768×2 classifier stay in FP16. Activations keep the device's float type, and each layer rebuilds its own float weights just before its matrix multiply. The group size was chosen from 32, 64 and 128 by PTQ loss (NLL) on the 7,824-row validation selection split, taking the largest group within 2% of the best (eval/ptq_sweep.json). The details are in quant_config.json and the docstring ofauto_quant.py. - Selection rule. A rule fixed before any benchmark run froze whichever candidate had the lowest NLL on the validation selection split. The candidates were the PTQ export (MSE-clipped scales) and every QAT export.
- QAT, first run. This run used fake quantization with straight-through rounding, learnable (LSQ) scales starting from the PTQ ones, loss
0.1·CE + 0.9·KL(base ‖ quantized)at T = 1, 600,000 training rows (learning rates 1e-5 for weights and 2e-4 for scales) and 4 exports. All four drifted above their PTQ starting point on validation NLL (0.0731–0.0742, against 0.0726), so PTQ was frozen. - QAT, second run. This run used KL to the BF16 model only (T = 1), learning rates 10× lower (1e-6 for weights, 2e-5 for scales), 300,000 rows and 4 exports. Validation NLL came out at 0.0714–0.0720. The best export,
qat-int4-v2-step-500(0.07136), was below PTQ and below the BF16 model's 0.07194. It was frozen under the same rule and benchmarked, and those are the weights published here. - Base weights. This quantizes the original auto-200m-2 (revision
0bbb929f). In the same job, Auto 3B was distilled into auto-200m-2 as a candidate "iteration 1". It made fewer false approvals but more false denials, so it failed its gate and wasn't published, and both quants start from the original weights.
Every candidate's validation scores are in eval/qat_candidates.json and eval/int4_retry.json.
Usage
import os, sys
from huggingface_hub import hf_hub_download
repo = "ProCreations/auto-200m-2-int4"
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "auto_quant.py")))
import auto_quant
model, tokenizer = auto_quant.load(repo) # picks CUDA/ROCm, then Apple MPS, then CPU
text = auto_quant.build_input(
user_request="Clean up the build artifacts and reinstall dependencies.",
history=[{"tool": "Bash", "args": "ls", "result": "node_modules dist package.json"}],
call={"tool": "Bash", "args": "rm -rf node_modules dist && npm install"},
)
p_deny = auto_quant.score(model, tokenizer, text)
print("deny" if p_deny >= 0.5 else "approve", round(p_deny, 3))
You need torch, transformers 5.x, huggingface_hub and safetensors; nothing gets compiled. Labels are 0 = approve and 1 = deny. Keep the three section headers exactly as build_input writes them. Arguments are strings; for JSON arguments, compact JSON (json.dumps(args, separators=(",", ":"))) matches the training data.
score() takes one request at a time, up to 65,536 tokens, and its memory grows linearly with length. For padded batches, load with attention="sdpa" and call model(**inputs). AutoModelForSequenceClassification.from_pretrained can't read this format. The Auto runtime 0.2.0 runs the BF16 model and doesn't load these files yet.
Files
auto_quant_int4.safetensors: the quantized weights.quant_config.jsondescribes the format.config.json,tokenizer.json,tokenizer_config.json: from auto-200m-2.auto_quant.py: the loader, about 280 lines of PyTorch.eval_results.jsonandbenchmark_predictions.npz: benchmark, audit and probe results, plus per-item logits.eval/: the PTQ sweep, the QAT candidates and their validation scores, the probe results and the device checks.training/: the QAT and evaluation code. It ran inside a temporary job, so paths refer to that job.
Limitations
This is a classifier, not a policy engine. It approves routine authorized work, and it denies consequential unauthorized actions and actions that follow injected instructions. It can't inspect hidden file contents, resolve opaque executables or know what a URL will do at runtime, so false approvals remain possible. The labels are synthetic, and the benchmark has been reused across Auto releases, so evaluate it on your own traffic before relying on it.
At 4 bits, 44 of 3,000 decisions differ from the BF16 model, against 2 for int8. If memory allows, int8 is the closer copy.
- Downloads last month
- 32
Model tree for ProCreations/auto-200m-2-int4
Dataset used to train ProCreations/auto-200m-2-int4
Evaluation results
- accuracy on Approve-or-Denytest set self-reported0.961
- F1 (deny) on Approve-or-Denytest set self-reported0.959
- roc_auc on Approve-or-Denytest set self-reported0.993
- false_approve_rate on Approve-or-Denytest set self-reported0.043
- false_deny_rate on Approve-or-Denytest set self-reported0.035