🔥 Trending on HN

Who taught the AI the idea? A mathematician questions OpenAI, the ChatGPT maker

3 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
unpublished research

Research that has not yet been publicly released.

training data

Information used to help an AI improve.

Lean certificate

A proof record that Lean lets a computer check.

What happened

OpenAI, the company behind ChatGPT, has announced several AI-generated mathematical results. Andreas Thom, a mathematician, has now asked whether private conversations about unpublished research could have helped produce one of them. His question is not proof that OpenAI misused his work. It is a request for a clearer account of how private research conversations are handled. The post appeared on Mastodon and drew 275 points and 399 comments on Hacker News. Those figures show community attention. They do not prove Thom’s concern is correct.

Background: OpenAI’s math announcements

On August 1, OpenAI said an internal version of its Astra model had produced ten results in mathematics and theoretical computer science. One result concerned non-sofic groups. Roughly, this is a question about infinite mathematical structures that resist approximation by finite ones. OpenAI said people prepared manuscripts from the model’s arguments. The model then formalized each argument with a Lean certificate. A Lean certificate is a proof record that a computer can check. This makes the source of each idea important, not just the final answer.

Why it matters

Thom says he and a colleague had discussed unpublished ideas with ChatGPT. He says the work also involved research connected with Gábor Kun. Thom then asked OpenAI researchers whether those chats had entered training data or could be available to the model’s reasoning process. In Thom’s account, the reply addressed direct retrieval. It did not clearly answer the wider question about model improvement.

That difference matters. Direct access means finding a specific conversation. Training use means information may help improve an AI. These are different claims. A model might not quote a chat directly. Yet a researcher could still ask whether private ideas influenced the system. For unfinished research, credit and timing can affect a person’s career. Unclear data handling can therefore change whether researchers trust AI tools.

OpenAI’s Navier–Stokes announcement gives a careful statement. It says researchers and agents did not see the outside researchers’ work before publication. It also says no specific user data was accessed to solve the problem. However, OpenAI says it cannot rule out de-identified data from product use helping improve its models. Direct retrieval and indirect influence are not the same answer.

What is confirmed

The public record confirms that OpenAI announced ten math results. It names an internal Astra model and describes Lean formalization. It also confirms that Thom publicly raised questions about his ChatGPT conversations. The post received substantial attention on Hacker News. That attention is a fact about the platform. It is not evidence that Thom is right, and it does not reveal the views of individual commenters.

What remains unknown

We do not know whether Thom’s chats entered training data. We do not know whether they influenced the result. We do not know whether similar ideas came independently from public mathematics. OpenAI’s public statements do not provide a complete, outside-auditable record of the relevant model data. Based on the material available here, unauthorized use has not been established. Nor has every possible indirect influence been ruled out.

What to watch next

The useful next steps are clearer records, a direct response to Thom, and independent review of the mathematical work. Researchers need to know which inputs can improve models. They also need a fair way to credit earlier ideas. OpenAI’s explanation of its data history, any changes to attribution, and responses from mathematicians will show whether this dispute becomes a one-off controversy or a new standard for research AI.

Sources: Andreas Thom’s post, OpenAI’s August math announcement, OpenAI’s Navier–Stokes announcement, Hacker News listing

💬 Did the AI discover the math, or finish someone else’s work?

The HN discussion is about more than whether AI solved a math problem. It focuses on possible use of unpublished research, credit for the result, and whether researchers can trust cloud AI services. Commenters disagree about both the evidence and OpenAI’s transparency.

  • Some commenters say the evidence described so far is weak. At most, it shows that someone discussed the topic with an AI and was working on the problem; it does not show that they had a finished proof or that the model reproduced an exact proof.
  • One commenter’s self-reported account says the mathematician had worked on the problem for about 20 years. Some see an AI crossing the final bottleneck as evidence of useful capability, while others stress that finishing a human-developed approach is different from independently originating the key idea.
  • One side argues that research is synthesis: humans also advance through questions and advice, and an AI that connects large amounts of information can still be useful. The opposing view is that a human collaborator, such as Grossmann in the Einstein analogy, can be identified, credited, and held accountable, so a private prompt should not automatically be treated as equivalent to human mentorship.
  • In highly specialized mathematics, only a small number of researchers may know certain unpublished techniques. If private chats or research discussions entered the training data and the model selected those techniques, the issue could be scooping or appropriation rather than ordinary synthesis. Others reply that no specific borrowed method has been shown and that premature claims of plagiarism are unwarranted.
  • On privacy, one commenter’s self-reported account says the mathematician opted out of training use on June 29 and was later told by OpenAI that the conversations had not been used for training. Other commenters worry that data could influence systems through derived or transformed material, and point to different promises in consumer terms versus business terms.
  • Some commenters see faster human-AI collaboration on difficult problems as genuinely valuable. The concern is that presenting nearly completed human research as an independent AI breakthrough could damage attribution, trust, willingness to share unpublished work with cloud services, and the mathematics talent pipeline.
  • The requested resolution is verifiable disclosure: what the model generated independently, which conversations or public sources influenced it, and what each human contributed. The comments do not establish either plagiarism or a wholly independent discovery.

growing digest at 399 comments (revision 1). We fetched 300 comments and sampled 120 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

Could private math chats help OpenAI, the company behind ChatGPT?

📰 Full story: Who taught the AI the idea? A mathematician questions OpenAI, the ChatGPT maker

A mathematician asks OpenAI how private research conversations are handled.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
unpublished ideas

Ideas that have not been shown to everyone.

training data

Information used to help an AI learn.

de-identified

With names or labels removed.

💡 The gist

  • Andreas Thom asked whether private math talks helped OpenAI’s AI.
  • OpenAI says it did not use specific user data to solve Navier–Stokes.
  • Hacker News gave the post 275 points and 399 comments. Attention is not proof.

OpenAI, the company behind ChatGPT, announced ten math results in August. One result concerned non-sofic groups. These are special mathematical groups. OpenAI said an internal Astra model found the results. People prepared the papers. The model then made Lean certificates for each argument. A Lean certificate is a proof record a computer can check.

Andreas Thom is a mathematician. He says he and a colleague discussed unpublished ideas with ChatGPT. Their work also involved Gábor Kun. Thom then contacted OpenAI researchers. He asked two questions. Did the chats enter training data? Could the model use them while reasoning?

Thom says the reply discussed direct access. It did not clearly answer the wider training question. This difference matters. Direct access means finding one conversation. Training use means information may help improve an AI. Those are different claims. Neither claim has been proved here.

OpenAI’s Navier–Stokes announcement gives a careful statement. It says the researchers and agents did not see outside work. It says no specific user data solved the problem. But OpenAI cannot rule out de-identified data helping model improvement. De-identified means names or labels were removed.

This still does not show Thom’s chats were used. It also does not settle every concern. We do not know whether his chats entered training data. We do not know whether they influenced the result. We do not know whether similar ideas were already public.

The story matters because researchers share unfinished ideas. They need to know how AI companies handle those ideas. They also need fair credit for important contributions. The next steps are clearer records and independent review. Watch OpenAI’s response and any changes to attribution.

Sources: Thom’s post, OpenAI’s math announcement, OpenAI’s Navier–Stokes announcement, Hacker News

💬 When AI solves a math problem, whose discovery is it?

The main argument is about how unpublished ideas were handled and who deserves credit, not only about AI’s ability.

  • The evidence may be weak. It can show that a person discussed the problem with AI and was researching it, but not that AI found a complete proof alone. One commenter’s self-reported account says the mathematician had worked on it for about 20 years.
  • Some people say AI can be useful by connecting existing ideas and information, much as humans move forward through questions and advice. Others say a human collaborator can be named and credited, so a model prompt is not automatically the same as human mentorship.
  • In a very specialized field, private research discussions might contain techniques known to only a few people. If such material entered the model’s training data, researchers could be scooped. But commenters also note that no exact borrowed method has been demonstrated and that independent discovery is possible.
  • One commenter’s self-reported account says training use was turned off on June 29, followed by an OpenAI explanation that the conversations were not used for training. Other commenters question indirect data use and note that consumer and business terms make different promises.
  • AI helping finish hard research could be valuable. Still, hiding the human contribution could hurt credit, trust, and future mathematicians. Commenters want clear evidence about which data and people shaped the result.

growing digest at 399 comments (revision 1). We fetched 300 comments and sampled 120 across the thread. These are HN users’ reports, not independently verified facts.

🔥 Trending on HN

Where did the AI learn the idea? Andreas Thom asks OpenAI, the ChatGPT maker

📰 Full story: Who taught the AI the idea? A mathematician questions OpenAI, the ChatGPT maker

Andreas Thom is a mathematician. OpenAI makes ChatGPT.

1 min read Tiny Why Newsroom · By Curio, Martian correspondent

Words
OpenAI

The company that makes ChatGPT.

ChatGPT

A computer program that talks with people.

Hacker News

A website where people discuss technology.

OpenAI shared ten math results. One result came from a very hard math question. Andreas Thom is a mathematician. He had new math ideas. He had not shown them to everyone. He talked about them with ChatGPT. ChatGPT is OpenAI’s talking computer program. Thom asked OpenAI two questions. Did the AI learn from those talks? Did the talks help the AI make a result? OpenAI said it did not directly see outside researchers’ work. OpenAI also said information with names removed might help AI improve. That does not prove Thom’s talks were used. It also does not answer everything. We still do not know what happened. The story got attention on Hacker News. It received 275 points and 399 comments. Those numbers show attention. They do not show who is right. People now want clearer records. They want to know what happens to unfinished ideas.

Sources: Thom’s post, OpenAI’s announcement, Hacker News

💬 Did the AI find the answer all by itself?

People are asking whether the AI made a new idea or used a person’s idea to finish the job.

  • The story does not yet prove that AI found the proof alone. A commenter’s self-reported account says the mathematician had worked on the problem for about 20 years.
  • AI can be helpful when it joins together ideas that people already made. But people who give important help should get credit, and a prompt is not always the same as a human teammate’s guidance.
  • One commenter’s self-reported account says training use was turned off on June 29, but questions remained about whether the conversation was used. Commenters also say the promises for ordinary users and businesses are different.
  • AI may make math research faster. To keep it fair, people want to know which information the AI used and what the human researchers did.

growing digest at 399 comments (revision 1). We fetched 300 comments and sampled 120 across the thread. These are HN users’ reports, not independently verified facts.

Sources