OpenAI published AI-generated solutions to 372 families of math problems, totaling 719 manuscripts, on Oct. 6, then withdrew three of them the next day after finding a sign error, according to Implicator.ai and TechCrunch.

A sign mistake in a paper on Weil classes on split abelian eightfolds undid a key cancellation argument, invalidating two dependent manuscripts on the Hodge conjecture, Implicator.ai reported. OpenAI revised 14 more manuscripts to fix proofs and updated references in 13 others, according to the outlet.

An outside advisory group had set guidelines in September calling for machine-checked Lean proofs and published chains of reasoning when human understanding of a result is lacking, but only about 42% of OpenAI's results were formalized in Lean and just 10 of the 719 manuscripts included the model's chain of thought, TechCrunch reported. Mathematician Terence Tao said problems are "being solved autonomously by AI prompters who have no interest in the broader field itself once their initial target is solved," according to the outlet.

Set theorist Asaf Karagila reviewed OpenAI's paper on the century-old Partition Principle problem and called it "unclear, muddled" with "a strange structure," writing on his blog that it would warrant desk rejection from any reputable journal. He said OpenAI's approach amounts to a "denial of service" against the mathematical community by producing more output than the field can absorb.

Builders relying on AI for tasks that need independent verification should read this as a preview: generating hundreds of plausible outputs faster than experts can check them does not move a field forward, it moves the bottleneck.