OpenAI published a repository with novel solutions to a range of math problems solved by their next frontier model.
The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model.
On average, each result used three hours of ChatGPT Pro thinking compute with that model.
Over the course of the evaluation, the model was posed approximately 4,000 problems.