QES — 1T-Parameter LLM on One 128 GB PC
1T-parameter Kimi K2.5 on one 128 GB Ryzen AI MAX+ PC
Independent applied AI research. Large language models, Mixture-of-Experts (MoE), local AI, bounded-memory execution, GPU inference, AMD Ryzen AI / ROCm, model integrity, high-assurance on-premises AI, and autonomous engineering systems.
Independent applied AI research focused on local large-model execution, Mixture-of-Experts inference, model integrity, high-assurance on-premises AI, and autonomous engineering systems.
Running a 1-trillion-parameter, ~375 GB Kimi K2.5 model on one 128 GB AMD Ryzen AI MAX+ 395 Windows PC.
QES investigates whether very large Mixture-of-Experts models can remain locally executable when their routed-expert capacity substantially exceeds available system memory.
The work has demonstrated cross-family execution across:
For Kimi K2.5, the published sustained validation reached 0.432091702 tokens/s over 128 generated tokens, with all 129 output token IDs exactly matching the frozen historical reference, 61,440 GPU expert executions, 7,680 authoritative correctness checks, and zero GPU errors.
QES also incorporates model-integrity measures including pre-execution verification, protected expert stores, continuous correctness checking, and fail-closed qualification behaviour.
DOI: https://doi.org/10.5281/zenodo.22730031
https://github.com/nigelhutch/nh-applied-qes
Technical discussion, research collaboration or commercial enquiries:
NH Applied — nhapplied@gmail.com
Nigel Hutchinson / NH Applied