Winner, Ralphthon @ ICML 2026

We took 1st place in Track 2 (Review Agent) at Ralphthon @ ICML 2026 Auto-Research with MAC n CHEESE, an evidence-bound reviewer for scientific papers — and won $10,000 in OpenAI credits.

LLM reviewers are easy to fool and hard to trust, so the system is split in two. A deterministic core does the trustworthy work: reproducible checks for arithmetic, ledger-tracing, baseline-fairness, citation existence, and injection scanning that run offline and resist prompt injection.

An LLM committee of three specialist agents plus an area chair adds scientific judgment, but only after the deterministic audit is frozen. The models can enrich a review; they can never quietly rewrite the audit’s identity or verdict, nor be steered by instructions hidden in a paper.