An autonomous AI research agent was tasked with replicating nine published U.S. equity anomalies on clean, survivorship-free data. On a faithful build, none survive out-of-sample — and the lone apparent survivor turned out to be the agent’s own construction error. The real lesson is that an AI researcher is only as trustworthy as the guardrails…