mark

An IsolationForest anomaly detector over the stat macOS system-telemetry dataset. It learns what normal machine behaviour looks like from per-second CPU, memory, disk and network counters, and flags samples that do not fit. CPU only, scikit-learn only, no PyTorch and no transformer anywhere in the path. The saved bundle is a StandardScaler plus a 200-tree forest over eight features, and runs to about 1.9 MB.

Training used the first 80 percent of the session in time order (4,951 rows) and held out the last 20 percent (1,238 rows). The disk and network columns in stat are since-boot counters, so they enter as per-interval differences, exactly as that card recommends. cpu_temp is null throughout and is dropped; per-core cpu_usage enters as its mean and max.

Usage

No scikit-learn and no pickle. The forest ships as plain arrays in model.safetensors with parameters in config.json, and predict.py scores them with numpy only. It replicates scikit-learn's scoring: 99.9 to 100 percent predict agreement with the original estimator on train, holdout, and synthetic incidents, with scores matching to within 0.003.

import numpy as np
from huggingface_hub import hf_hub_download
from predict import MarkModel

m = MarkModel(".")
X = np.array([[46.8, 96.0, 3410.2, 33.0, 6.1, 5.8, 0.4, 0.3]])
print(m.predict(X))

The eight columns, in order, are cpu_mean, cpu_max, memory_used_mb, battery_status, disk_read_delta, disk_write_delta, net_sent_delta and net_recv_delta. Output is 1 for normal, -1 for flagged. Needs safetensors and numpy only.

Evaluation

Measured on the temporal holdout and on synthetic incidents built from holdout rows:

Case Flagged
Clean holdout (false alarms) 0.027
Stressed machine, CPU, memory and disk spiking together 0.917
Idle crash, CPU and memory dropping together 0.372
Network storm, sent and received spiking together 0.245

The contamination parameter is 0.01, so about one percent of training rows score as anomalous by construction; the holdout false-alarm rate of 0.027 reflects the battery recharge region the holdout covers, which the training window never saw.

Single-feature extremes score poorly, and that is a property of the algorithm rather than a bug in this fit. IsolationForest separates points by random splits, and a point that is extreme in one of eight dimensions waits several splits for that dimension to be picked, which is the same depth normal points reach. Standardizing the features changes nothing for the same reason: the splits are uniform in range, and rescaling only reparametrizes the draw. Incidents that move several counters at once, the way a real stressed machine does, are the ones it catches.

Limitations

This models one machine over one 1-hour-44-minute session. It has not seen another host, another workload, or a longer timescale, and there is no reason to expect it to transfer. The synthetic incidents are illustrative, not a benchmark. The stat card's unit caveat applies here too: the disk and network deltas are raw counter advances whose physical unit was never confirmed, so thresholds learned on them are in those raw units.

Downloads last month
-
Safetensors
Model size
105k params
Tensor type
I64
·
F64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train harpertoken/mark