arxiv:2608.21363
Ali Toygar Abak PRO
phionyx
AI & ML interests
AI governance, AI safety, multi-agent systems, reproducible LLM evaluation, agent protocols, open-source AI, local inference, and trustworthy AI
Recent Activity
posted an update about 12 hours ago
AI evals can be reproducible, signed, and still overclaim what the evidence actually establishes.
A tool-call denial proves a denial. It does not prove that a safeguard prevented an incident.
For AI evaluation, agent evaluation, and AI safety assurance, provenance alone is not enough. We also need to preserve the scope, context, and limits of the claim.
That is the case for a claim-preserving evidence contract.
đź“„ Access Is Not Yet Verifiability
https://huggingface.co/blog/phionyx/access-is-not-yet-verifiability new activity 1 day ago
phionyx/airep-evaluation-evidence:fix: preserve declared evaluation timestamp provenanceOrganizations
None yet