A critical epistemological lens applied to AI benchmark claims, model capability assessments, and statistical presentations that may mislead through selective metrics, dataset contamination, cherry-picked results, or hallucination. The page collects resources and reasoning for evaluating AI performance claims with rigour, highlighting how large language models can generate plausible but false outputs that resemble statistical truth.

Semantic Classification

Content

Hallucinations

Provenance