AI & ML interests

Benchmarking AI agents on real-world infrastructure operations across the full system stack (hardware, distributed systems, storage) with fine-grained risk assessment

Recent Activity

xuanmiao-31  updated a Space 1 day ago
InfraBench/README
xuanmiao-31  updated a dataset 1 day ago
InfraBench/infrabench
xuanmiao-31  published a Space 2 days ago
InfraBench/README
View all activity

Organization Card
InfraBench
InfraBench
A benchmark for infrastructure agents

Twelve real operational incidents — from IPMI power recovery to silent data corruption — spanning hardware, local systems, distributed systems, and user applications. Beyond pass/fail: durable state, invariants, cleanup, and risk.

Website → Leaderboard Dataset GitHub

models 0

None public yet