Mathematician Accuses OpenAI of Using Unpublished Work
Sam Altman, chief executive officer of OpenAI, appeared on a media tour of the company’s Stargate AI data center this week, showcasing the infrastructure that underpins the firm’s rapidly expanding suite of language models. The visit, captured by Bloomberg and Getty Images, came amid growing scrutiny of the data sets that feed OpenAI’s systems and the methods used to train them.
Shortly after a public dispute over the use of unpublished research, mathematician Andreas Thom has publicly accused OpenAI of unethical and “dishonest” conduct, citing a lack of transparency about the origins of its training data. In a series of Mastodon posts, Thom argued that interactions he and colleagues had with the ChatGPT chatbot—prior to the company’s recent announcement of breakthrough mathematical results—may have influenced the AI’s performance in the field. Thom’s claims add to a broader debate about the provenance of the data powering advanced AI models and the responsibilities of developers to disclose their sources.
The controversy underscores the tension between rapid AI development and the need for clear, ethical data practices. While OpenAI has defended its training procedures and emphasized its commitment to responsible AI, the allegations from Thom and other researchers highlight the ongoing challenges in ensuring transparency and accountability in the industry. The company’s next steps will likely involve addressing these concerns and clarifying its data governance policies.