OpenAI's math proof release falls short of new field standards
OpenAI this week published hundreds of claimed solutions to hard math problems (719 manuscripts, per the article), saying it consulted an advisory group of elite mathematicians to avoid the controversy its earlier result sparked. But only ten of the manuscripts included the model's chain of thought, and by the article's account roughly 42% of the proofs had not gone through formalization in Lean; the Advisory Group on Mathematics and Artificial Intelligence (AGMAI), hosted by Princeton's Institute for Advanced Study, said it is ultimately up to the mathematical community to judge whether its recommendations were followed — its first request being to stop testing advanced problems on proprietary models. A paper from Cambridge and King's College London mathematicians also documents at least two discrepancies between OpenAI's natural-language proof and the Lean code for a problem derived from the Navier-Stokes equations, concluding such autoformalized proofs should not be trusted without peer review.