good questions. the receipt has the machine at 27.4 GiB guest RAM, 18.1 GiB RSS, 7β9 GiB cache, and 4β6 GiB swap; nothing dropped caches between chunks. of the 2401 s wall, reads were 1025 s at 1.12 GB/s, verify 484 s at 2.38 GB/s, and madvise 79 s. so yes: page cache materially contributed--the 1.12 GB/s exceeds the 0.55 GB/s SATA device--but the cache was only 7β9 GiB against a ~43 GiB cold set, so the split was not measured. the NVMe rerun records device bytes beside engine reads and adds an O_DIRECT control
i also agree 1024 is the lever: our 120B traces put static top-32 at ~75% of picks and reselecting every 512 tokens at ~80β84%; the planner optimizes avoided stall, not hit rate. on verification, 121,952 Γ 3 + 4,608 is three tensors per expertβgate/up/down--each hashed once, not one slice hashed three times. we measured verification as the bottleneck and the next row moves SHA-256 onto the CPU SHA extensions with identical digests. distinct resident experts/chunk and the 380-vs-384 pick deficit are not receipt-backed yet, so i wont guess--they need instrumentation