Research

Sakana AI System Spots 73% of Critical Paper Errors

Sakana AI has launched an automated peer-review architecture called Multi-Layered Review, setting a new standard by detecting 73.43% of core logical errors in academic papers.

MarkTechPost1 day agoResearch
Image: MarkTechPost

Sakana AI has introduced Multi-Layered Review (MLR), an agentic AI system designed to analyze scientific literature and identify deliberate flaws. Built on standard API models without task-specific fine-tuning or custom GPU infrastructure, MLR relies on three specialized components. An Appendix Agent powered by Claude Haiku 3.5 summarizes implementation details, an optional Literature Review Agent running on Claude Sonnet 4 conducts web searches, and a Review Agent using Claude Sonnet 4 executes a three-pass chain directly on raw PDF documents to generate structured feedback.

To evaluate the system, researchers constructed a Contradiction Benchmark containing 1,164 planted errors across 257 papers from ACL, AISTATS, CVPR, ICML 2025, and NeurIPS 2024. Gemini 2.5 Pro mapped claim dependencies into knowledge graphs, GPT-4.1 generated contradictory edits, and an o3 model served as evaluator. Across four evaluation passes, MLR detected 73.43% of distance-0 core-claim errors and 40.95% of total errors, outperforming baselines like AgentReview, which caught only 14.81% of core mistakes. On a single review pass, MLR still identified 60.79% of severe errors.

When tested against real-world dataset WithdrarXiv-Check consisting of 211 retracted papers, MLR achieved a 26.07% similar match rate and 16.11% exact match rate. On ICLR 2025 submissions, MLR achieved a Pearson correlation of 0.586 with human reviewers, compared to a human-to-human reference of 0.742, while scoring 0.429 on ICML 2025 submissions versus AI Reviewer at 0.439. Running MLR consumes 189,062 input tokens and costs approximately $0.47 per paper, significantly cheaper than AgentReview at $0.81 and 310,964 tokens, though pricier than baseline LLM-Review at $0.01 and 6,517 tokens.

For developers building automated research tools, MLR demonstrates that structural prompt design and model selection drive performance more than brute-force sampling. Model ablation showed that swapping GPT-4.1 for Claude Sonnet 4 boosted baseline detection from 14.56% to 35.40%. However, practitioners should note that MLR focuses heavily on technical validity rather than presentation novelty, and remains vulnerable to hidden prompt injections embedded within manuscript files.

This is our own summary of reporting by MarkTechPost

More in Research