GPT-6 Astra launched with two striking benchmark results: 97.6% on FrontierMath Tier 4 v2 and 99.9% on ARC-AGI-3. The evaluation setup matters, however. ARC Prize measured 62.7% under its Standard harness and 99.9% with OpenAI’s Provider Adapter, which preserves reasoning state between requests.
A separate mathematics announcement brought another distinction into focus. OpenAI’s proposed Navier-Stokes solution was generated by a substantially more capable, unreleased internal model, not GPT-6 Astra. According to OpenAI, Astra was used afterward for 17 hours of Lean formalization and verification.
That announcement also raised questions about research credit, the provenance of the method, and whether private work submitted through Codex could have influenced OpenAI’s effort. Mathematician Tristan Buckmaster says he does not know whether his team’s data was used. OpenAI says its investigation ruled out that influence. No independent audit addressing the question has been made public.
For enterprises using AI with confidential research or business information, the practical question begins before a result is generated: what original information enters the AI workflow, and what context must remain intact for the work to be useful? Provider assurances matter, but teams also need to understand and control the data path on their side.
The GPT-6 Astra scores depend on the system around the model
OpenAI introduced GPT-6 Astra on September 3, 2026. Its launch page reported 97.6% on FrontierMath Tier 4 v2 and 99.9% on ARC-AGI-3, alongside gains in coding, computer use, science, and other professional tasks.
Those numbers are real results under recorded configurations. They are not interchangeable measures of general mathematical ability.
| Headline | What the source actually reports | What it does not prove |
|---|---|---|
| 99.9% on ARC-AGI-3 | ARC Prize reports 99.9% with OpenAI’s Provider Adapter, which preserves opaque reasoning state and uses compaction. Its best observed Standard-harness result was 62.7%. | That Astra scores 99.9% under every harness or context-state policy. |
| 97.6% on FrontierMath Tier 4 v2 | OpenAI reports 97.6% on a benchmark of difficult expert-written problems. | That the same score applies to problems that remained open to mathematics in 2026. |
| 3% on FrontierMath Erdős | Epoch AI reports that a prerelease Astra solved 2 of 68 open research problems under one fixed attempt, a $300 budget, and a 72-hour limit per problem. | That the Tier 4 result was invalid. Erdős is a different problem set and protocol. |
| Ten advances in mathematics | OpenAI says an internal Astra version resolved or made substantial progress on ten heterogeneous problems. | That Astra completely solved all ten, or that these results were the Navier-Stokes result. |
Both ARC scores are valid within their recorded setups, but they answer different questions. The model’s performance changed with the harness, retained state, compaction, effort level, and cost. Those conditions belong in the result, not in a footnote.
The same principle matters outside benchmarks. When an enterprise compares two AI results, it needs more than the model name. It needs the data state, prompt, tools, permissions, retained context, evaluation method, and acceptance criterion that produced each result.
GPT-6 Astra did not produce the Navier-Stokes construction
On September 8, OpenAI announced a proposed solution to the three-dimensional incompressible Navier-Stokes Millennium Prize problem. According to OpenAI’s account, an unreleased internal model that is significantly more capable than GPT-6 Astra generated the forced blow-up construction.
The reported effort was far larger than a normal model interaction. OpenAI describes roughly 10,000 concurrent agents working for 88 hours and producing about 2.7 million messages and 130 billion output tokens. Astra’s stated role came afterward, when it spent 17 hours formalizing and checking the work in Lean.
The distinction matters because model identity is part of the evidence. A headline that credits Astra combines a private research system with a public product and obscures what was actually evaluated.
The mathematical status also remains provisional. On September 11, the Clay Mathematics Institute said the problem appears to have been settled, while emphasizing that the evaluation and assignment of credit would not be rushed. That is significant recognition, but it is not final acceptance or a prize award.
The dispute is about provenance as well as mathematical priority
The public dispute centers on work by NYU mathematician Tristan Buckmaster and Levent Alpöge, who works at Anthropic and pursued the research in a personal capacity. Their project extended a research direction associated with Diego Córdoba and Luis Martínez-Zoroa. Buckmaster says he and Alpöge used Claude, Codex with GPT-5.6 Sol, and later Astra while developing, writing, and auditing their work.
In a public statement, Buckmaster argues that OpenAI pursued a closely related forced route after learning that their team had made progress. He also raises the possibility that private Codex work might have influenced OpenAI’s effort. His statement is careful about the limit of that allegation: he says he does not know whether the data was used and is not accusing OpenAI of misconduct.
OpenAI’s public position became more definitive over the following two days. On September 8, Axios reported that OpenAI could not entirely rule out an indirect contribution from de-identified product-use data. OpenAI’s September 10 update now says Buckmaster’s Codex prompts could not have influenced the result, including through training, and that the proofs differ significantly. OpenAI also recognizes Buckmaster and Alpöge’s priority on the forced Euler result.
The current public record supports neither “AI stole the proof” nor “the data question has been independently settled.” It supports a narrower conclusion: AI-assisted research now creates provenance questions that provider policy, authorship convention, and mathematical review must address together.
What enterprise teams should learn from the dispute
Most companies are not trying to solve a Millennium Prize problem. They are still putting confidential work into systems operated by organizations that may improve models, build competing products, or change their data policies over time.
The sensitive asset may be a research draft, legal strategy, customer record, product roadmap, incident report, pricing model, or operating procedure. The same four questions apply before that asset enters an external AI workflow.
What original information leaves the controlled environment? Identify the exact fields, documents, relationships, and metadata the model path will receive. A workspace privacy setting does not answer this architectural question by itself.
What context must survive for the task to remain useful? Removing every identifier may protect values while destroying the links the model needs to reason about a contract, incident, account, or asset hierarchy.
What evidence can the organization retain? Record the active policy, protected working version, approved model path, model and tool configuration, result, reviewer, and acceptance decision. Provider logs and customer-held evidence are different control surfaces.
What remains outside the control? A protected data path cannot prove how a provider trained a model, settle intellectual-property ownership, or assign research credit. Those questions still require contractual, legal, and institutional controls.
This separation helps a cross-functional review. Security can inspect exposure and deployment boundaries. Data and workflow owners can test whether the protected version remains useful. Legal can evaluate contractual rights and retention terms. The business owner can decide whether the remaining evidence is sufficient for the intended action.
A provider assurance is not a customer-controlled data path
A provider may offer contractual protections, enterprise privacy terms, or technical options such as Zero Data Retention for eligible customers. Those controls matter. They still answer a different question from data-path design.
Provider assurance asks what the service says it will retain, process, or use. A controlled data path asks what original information the customer allows to leave its environment. Strong enterprise AI programs need both, because one cannot substitute for the other.
This distinction also avoids a common failure in privacy projects. If a team removes so much information that the model cannot follow the relationships in the work, the workflow may be protected but unusable. If it preserves every value to keep the task useful, the model may receive more original information than the organization intended to expose.
The design target is not maximum removal. It is the minimum exposure that still preserves the structure, relationships, and context required for one approved AI task.
Keep protected data useful for the AI task
LLM Capsule is CUBIG’s Context-Preserving Data Layer for AI. It is designed for AI work that is blocked because original sensitive values cannot travel through the approved model path.
In the customer-controlled deployment described here, original sensitive values and reconstruction mappings remain within the customer environment. LLM Capsule replaces sensitive values with consistent, context-preserving substitutes before the protected working version moves through the approved AI path. After the AI run, Business-ready Reconstruction uses that internal mapping to reconnect substituted names, codes, and values so the result can return to the business workflow.
Review retention, processing, and permitted use alongside the data path.
Keep task-relevant relationships consistent as sensitive values change.
The model receives the working context required for the approved task.
Reconnect substituted values using the internal mapping, then review the result.
Originals + reconstruction mapping stay inside.
The mapping is used when the result returns to the customer environment.
Substitution is sufficient when the values can change without altering the relationships the task needs. A contract review, for example, may require the same company and agreement references to remain consistent across a document, even though their original names and codes should not leave the environment.
Some tasks depend on sensitive patterns or distributions themselves. Replacing transaction amounts, timing, or clinical values one by one could erase the signal or invent a false one. When substitution cannot preserve both protection and target-task utility, embedded DTS can contribute AI-native Data Reconstruction. This is a conditional path, not a claim that every dataset needs synthetic data. Reconstructed data is also not reversed to the original records.
LLM Capsule does not determine whether OpenAI used a researcher’s private prompt, audit a provider’s training pipeline, or resolve ownership of an idea. Its role is narrower and operational: reduce the original sensitive information that enters an approved AI path while preserving enough context for the chosen work to remain usable.
Control the data path before confidential work reaches the model
Before confidential work reaches a frontier model, the organization should decide what must remain inside its environment, what context the AI task cannot lose, and what evidence it will need later. Those choices belong to the operating design of the workflow, not to model selection alone.
CUBIG defines this broader discipline as the AI-Ready Data Operating Layer. For sensitive workflows, the immediate job is to create an approved path that preserves the context the task requires while limiting how much original information leaves the customer environment.
Start with one blocked workflow and one representative payload. Identify the original values that must stay inside, the relationships the model needs, and the result the business must receive. Then test the complete round trip against the organization’s security, legal, and workflow criteria.
Explore LLM Capsule to see how a context-preserving data path can make sensitive enterprise data usable for AI.
