OpenAI adjusts Astra benchmark figures after launch blog post
OpenAI changed several evaluation metrics for its GPT-6 Astra model after publishing its launch announcement, with some scores favoring Astra and others fluctuating for rival models. The company says the adjustments reflect fixes to ensure accurate comparisons.
OpenAI has revised several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement on Sept. 3, with some updated figures showing Astra performing better while scores for rival Anthropic models worsened. The changes came amid an unusual rollout of the announcement, which was briefly published, retracted, and then republished with different metrics.
The blog post was originally scheduled to go live at 2 p.m. ET but took nearly two more hours before it was widely viewable. When OpenAI's X account shared the link at 3:32 p.m., it returned an error message. At 3:50 p.m., CEO Sam Altman posted the link himself, writing, «We hit a little snag getting the blog post deployed, but it is really great.» Multiple users, including Fortune, still could not access it at that time. The page became visible about an hour later.
OpenAI said the initial publication was retracted for reasons it could not disclose but which were unrelated to the benchmark figures. The company first attributed the issue to a bug in its content management system, then to an internet outage. Upon republishing, the blog contained different evaluation metrics that appeared to favor Astra, and some figures have continued to change since then.
Among the most notable changes was Astra's reported hallucination rate. The first archived snapshot of the blog post from 2:23 p.m. listed it at 4.2%, a figure that remained consistent through several snapshots. By the sixth archival snapshot taken at 5:20 p.m., the rate had been halved to 2% for Astra. The hallucination rate for Astra's predecessor, GPT-5.6 Sol, also dropped from 12.2% to 9.4%. As of this writing, both figures have reverted to their original levels of 4.2% and 12.2%.
OpenAI also boosted GPT-5.6 Sol's score on its internal version of the ExploitBench cybersecurity evaluation from 5.5% to 11.5% in later versions of the blog. The company said it is investigating reverting that number back to 5.5% because the higher result reflects a reasoning level that is not commercially available for Sol.
Mathematics scores also shifted. Astra's score on the FrontierMath Tier 4 (v2) evaluation remained steady at 97.6%, but OpenAI briefly altered the scores for GPT-5.6 Sol and Anthropic's Fable 5.1 model. In the first snapshot, Fable 5.1 scored 87.8%; by 5:17 p.m., it had dropped to 78%; today it stands at 83%. GPT-5.6 Sol's scores went from 83% down to 80.5% and back up to 83%.
The changes began before the blog was first published. An embargoed pre-publication draft provided to media outlets listed Astra's score on the ARC-AGI-3 evaluation as 98.6%. The live blog now shows 99.99%. An OpenAI spokesperson said, «We always verify evals before publication so adjustments between draft and final version are normal.» The company also noted that the Arc Prize Foundation, which created the benchmark, found Astra performed at 99.9% in its independent assessment when given a particularly powerful harness, and at 63% with the standard harness — still better than any other publicly released AI model.
OpenAI said different research teams oversee different metrics and are responsible for calculating and reporting them to a central team for publication. The company is open about the fact that the numbers are achieved under the best possible conditions and may differ from the models available in the production ChatGPT product. A disclaimer on the blog reads, «Evaluation scores are the maximum at any effort,» with further caveats included in footnotes.
«We care deeply about getting evaluations right,» an OpenAI spokesperson told Fortune. «Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.»
The episode highlights the challenges of measuring large language model performance through standardized benchmarks and raises concerns about the potential for manipulation and gamesmanship in reported specs. It comes amid intense competition in the AI industry, as companies release updates at a frenetic pace, each seeking to pull ahead of the other.



