The trust debate around OpenAI and unpublished mathematics is aimed at the wrong target. Everyone wants to argue about leadership, motives, and whether a lab that wants artificial general intelligence can be relied on to sit politely on your unfinished theorem. That argument is unwinnable and mostly beside the point. The real issue is structural: these tools only get good when you feed them the thing you would least like to lose.
Greg Brockman said as much in his TIME interview, describing agent work as something close to micromanaging: you have to provide it with tasks, you have to give it the context, because it just doesn’t have it. That is an accurate description of how these systems earn their keep. It is also, read from a researcher’s chair, a description of the exact upload you are being asked to make. Vague prompts get vague answers. Your best results come from your least shareable material.
Why the usual reassurances don’t land
I review tools for a living, which means I spend a lot of time reading terms of service and a lot of time watching those terms change. The standard reassurance is that enterprise or API tiers don’t train on your data. Fine. That covers training. It does not cover retention windows, subprocessors, incident response, subpoena exposure, or what happens when a product team decides a new memory feature would be nice.
Regulatory attention is not settling this for you either. The PIPEDA joint investigation of OpenAI OpCo produced findings, and the report itself acknowledged that even after addressing privacy risks in how large language models are built and deployed, the technology raises many other open questions. That is a regulator saying the frame is still being drawn. Useful honesty, not a guarantee.
Then there is the internal weather. An AI researcher has publicly warned that companies are ignoring catastrophic risks. Concerns about leadership and ethical judgment at these labs have not gone away. None of that tells you your draft proof will leak. It tells you that the organization holding it is under strain and changing shape, which is a different kind of risk than a bad privacy policy.
What actually works right now
Being skeptical about custody is not the same as being skeptical about capability. OpenAI’s own account of research acceleration describes researchers using coding agents throughout the day, often in several concurrent sessions, with total usage climbing fast. Their daily work has changed substantially over the year. I believe that, because it matches what I see reviewing coding agents: the gains in agentic coding and experiment turnaround over the past year are the most real improvements in this category.
Nature ran the question directly on the deep research tool and whether it is useful for scientists, which is the right question to keep asking about any of these products. Nature also covered scientists flocking to DeepSeek. That second story matters more than it looks. The moment a capable model can run somewhere you control, the custody problem stops being a philosophical debate about a vendor’s soul and becomes an infrastructure choice.
A working policy for unpublished material
Here is the split I use, and I think it holds up for math and for anything else you haven’t published:
- Send freely: code that implements known methods, refactors, test scaffolding, literature triage, notation cleanup, plotting and data-wrangling glue.
- Send with care: problem statements you could plausibly present at a seminar next week. If you’d say it out loud to a room of strangers, the marginal risk is small.
- Keep local: the novel construction, the key lemma, the step that makes the whole thing work. Run those against a model you host, or don’t run them at all.
This is annoying, because the third bucket is where you most want help. That is the honest tension and I’m not going to pretend a settings toggle resolves it.
The cost angle nobody budgets for
Nature’s feature on the $1.5 million academia tax points at the part of this that gets lost in ethics discussions. Keeping sensitive work in-house means paying for compute, storage, and the people who maintain it. Labs with money can buy their way into a private setup. Labs without it face a plain choice: use the hosted tool and accept the custody terms, or work slower than the people who do. Framing this purely as a matter of individual researcher judgment quietly ignores that the safe option has a price tag most departments can’t cover.
My verdict
Use the coding agents. The productivity gains are documented and I’ve seen them hold up in testing. Treat hosted deep research as an assistant for material you could already discuss publicly. And stop waiting for a trust verdict from anyone, because the terms, the regulators, and the labs themselves are all still in motion. Trust is not the control you have. Scope is. Decide what leaves your machine, write it down, and hold the line even on the day the tool would be most useful if you didn’t.
🕒 Published: