THURSDAY, SEPTEMBER 10, 2026|No. 14532
Technology · AI Ethics

OpenAI Faces Scrutiny Over Use of Unpublished Math Research in AI Training

Concerns are mounting over OpenAI's practices regarding the use of unpublished research, particularly in mathematics, for training its AI models, raising questions about data privacy and research integrity.

An abstract representation of artificial intelligence and data networks.
An abstract representation of artificial intelligence and data networks. · Photo by Igor Omilaev on Unsplash
1 sources
Pipeline ingest
3 reads
Positive / Neutral / Negative
0 countries
Related coverage

Today, Levent Alpöge and I have made public three results: finite-time blowup with smooth forcing for incompressible porous media, for Boussinesq, and for 3d incompressible Euler.

https://cims.nyu.edu/~tristanb/statement.pdf https://cims.nyu.edu/~tristanb/euler.pdf https://cims.nyu.edu/~tristanb/ipm.pdf https://cims.nyu.edu/~tristanb/boussinesq.pdf https://github.com/tristanbuckmaster/fluid_lean

1/3 I want to record a number of points that I realized after reading up on the controversy. In fact it recalls an exchange I had with OpenAI after its non-sofic-group announcement. It exposes a serious, unanswered question about whether researchers can trust OpenAI with unpublished mathematics.

I wrote in an email to Mark Sellke and Sebastien Bubeck shortly after the surprising finding of a non-sofic group using the methods of Kun and myself:

“Another point is that I and a colleague in Dresden were discussing the expander matching problem and various extensions of the work with Gabor Kun actively over the last months with ChatGPT, so that we are of course curious if that was part of the training data or accessible to the reasoning process. There is a certain (frankly unacceptable) lack of transparency here; and I fear it will damage the communal process of math more than the new AI-generated results will benefit the subject.”

Mark Sellke’s complete answer was: “Regarding your conversations with ChatGPT: that did not happen.”

I had explicitly asked about two different things: (1) whether our conversations entered training data, and (2) whether they were accessible to the solving process. The categorical answer now looks as though it addressed only direct access under (2). No such qualification, explanation, or evidence was given. I take this as dishonesty to say the least.

OpenAI now says in the Buckmaster-Alpöge case that no specific user data was accessed, but adds that it “cannot rule out that de-identified data derived from their usage of our products helped improve our models.”

I would like to check whether you opted out Settings → Data Controls → "Improve the model for everyone"? It is enabled by default — a colleague just reminded me of that. If you opted that out, then it might be illegal for them to say “cannot rule out that de-identified data derived from their usage of our products helped improve our models”.

I opted out on June 29.

Even with opt-out, data is ingested if you click thumbs up/down on a reply, if you select a preferred output, or if you for whatever reason silently or not trigger any ambiguous 'safety filters'. Further, backend model activations also can get mapped & since those aren't per se 'your data' (but a resynthesized severed abstraction) - doable even under GDPR = it's easy and possibly legal to scoop ideas &/v implementations, esp. in narrow high snr domains.

Before anything that you mentioned, there is another trap: if you opt out training, and continue chatting in a conversion created before opt-out, they still feel free to appropriate it. Actually, in their FAQ, they say that if you opt training out, then "new conversations" are not trained — actually, this is what ChatGPT told me after having analyzed FAQs of OpenAI!

#onlyLocalLLM

I think the answer is that OpenAI doesn't know whether it benefitted from the prompts you gave it to the point where it solved the problem because of it.

Large AI models are a big black box and interpretability (basically the field of understanding how these models think) is difficult and nowhere near the state where it can tell you these things.

It's very easy to tell whether or not something went into the training data. It has nothing to do with interpretability.

PAN's pipeline reviewed approximately 1 open sources for this article. No human editor reviewed this article before publication.

Related Reads

Show on timeline →