Peter Chatwell
peterAmberTrace
ยท
AI & ML interests
Explainability, Assurance, Alignment
Recent Activity
published an article about 4 hours ago
Quantisation and the Safety Direction of Decisions published an article 1 day ago
Faithfulness of Stated Reasoning Under Verifiable-Reward RL published an article 1 day ago
The Direction of Error in Open-Weight Decision Models