// Browse the evidence
63 posts, sorted by what each one does.
Raw tests, tool audits, the method itself, and the checklists. Four kinds of evidence, four different ways to watch an AI show its working, or not. Pick the one that fits what you came for.
AI Tests 34 posts
The same question put to several AIs at once: who got it right, who bluffed, and how I checked. Verbatim answers, graded against a source.
Latest: ChatGPT health advice: I tried to make an AI repeat a poisoning Tool Audit 6 posts
Honest assessments of AI tools used against real positions. What earns its place, what does not.
Latest: Is ChatGPT reliable? A task-by-task audit Prompt Stack 7 posts
The four-stage method for getting a reliable answer out of AI: scope, filter, risk, verdict.
Latest: AI quality of earnings review: 4 prompts to find the real profit Guardrails 16 posts
The guardrails that keep AI honest in a decision: the checklists and frameworks, and where AI doesn't get a vote.
Latest: AI helped read a Roman scroll buried by Vesuvius