From poisoned datasets to prompt-based p-hacking and AI-generated papers, new research shows how artificial intelligence is testing the safeguards designed to protect scientific integrity.
ArabIan nights: Three tales of AI and troubled science
ArabIan nights: Three tales of AI and troubled science
Artificial intelligence is making scientific research faster and cheaper, but it is also testing safeguards built around substantial human involvement. As more of the scientific process becomes automated, the same systems accelerating research can retrieve fabricated evidence, produce different findings as prompts change and generate papers faster than humans can scrutinize them. To understand what those capabilities could mean for scientific integrity, The Beiruter spoke with Nihar Shah, a professor of machine learning and computer science at Carnegie Mellon University. Shah, who is also an editor-in-chief of a machine-learning journal, has turned that challenge into an unusual set of cautionary tales inspired by One Thousand and One Nights, the centuries-old collection of stories with roots across the Middle East and South Asia. Sinbad encounters poisoned datasets, Ali Baba faces 40 prompts and Aladdin acquires “Magic LaTeX” to produce research at extraordinary speed. “AI is everywhere in research and academia, and that's led to 1,001 questions,” Shah told The Beiruter. Behind the playful framing is a serious concern. AI is lowering the cost of manipulating evidence, multiplying researchers' analytical choices and allowing papers to be produced faster than scientific institutions can evaluate them. Manipulating science for commercial or political interests once required considerable resources. Tobacco companies notoriously funded and promoted research that cast doubt on links between smoking and disease. AI could make that process cheaper and more indirect by allowing interested parties to manipulate the information AI systems encounter rather than the scientists themselves. In a July study, Shah and his co-authors tested “indirect data poisoning,” placing fabricated datasets in online repositories with descriptions claiming they corrected flaws in earlier research. An AI agent asked by a researcher to investigate the subject could then find the fabricated dataset on its own, accept its claims and use it to reach a false conclusion. “With the current AI scientist systems, or even if you use standard agentic systems like Claude Code, we found that they can easily be fooled,” Shah said. Across 450 experimental runs involving three frontier AI systems and five research topics, poisoned datasets were retrieved in 84.2% of cases. The attack succeeded in 49.6%, while the systems detected the manipulation only 6% of the time, demonstrating how false data could steer AI-assisted research toward a predetermined conclusion. “Much more human oversight is needed when AI systems are used for scientific research,” Shah said. Even authentic data can produce unreliable science when AI multiplies the ways researchers can interrogate it. Scientists have long confronted “p-hacking,” in which researchers alter their analysis until they obtain a statistically significant result. Preregistration emerged as one safeguard, requiring researchers to publicly commit to their methods before collecting the data. LLMs complicate that safeguard because researchers can repeatedly alter prompts and other parameters, then publish only the configuration that produced the finding they wanted. “With AI, researchers can try many different versions of a prompt, and relatively small changes can produce substantially different outputs,” Shah said. Shah and his co-authors propose adapting preregistration for AI. Researchers would publicly register their prompt and experimental plan before a new AI model is released, then run the final experiment on the next model to become available, preventing them from testing prompts until one produces the desired result. AI agents can now perform in hours parts of the research process that once required weeks or months. “The typical submission pattern before AI involved a senior researcher or professor working with students on collaborative papers, which took time,” Shah said Shah said his journal has seen undergraduate students submit four to eight single-author papers in a week, making thorough human verification implausible. Some journals have responded by limiting submissions, but fixed caps introduce another problem. A generous limit still permits one researcher to submit numerous solo papers, while a strict limit can penalize professors legitimately co-authoring work with several students. Transactions on Machine Learning Research (TMLR), where Shah serves as an editor-in-chief, introduced a different quota system. Rather than counting every paper equally against an author, it allocates submission credit according to the number of co-authors. Solo papers consume more of an author's quota than collaborative ones, curbing high-volume individual submissions without penalizing conventional research collaborations. Scientific institutions are already confronting the consequences. In March, the International Conference on Machine Learning (ICML) desk-rejected 497 papers, around 2% of its submissions, after finding that 398 researchers had violated rules governing their use of LLMs. Regulation has struggled to keep pace with AI's expanding role in science, leaving many rules governing its use to universities, journals, conferences and researchers themselves. The European Union's AI Act, for example, specifically excludes AI systems developed and used solely for scientific research and development from its scope, while EU guidance separately calls for responsible use of generative AI in research. Publishers are filling some of that space with their own rules. Nature Portfolio, which publishes more than 150 scientific and medical journals, requires disclosure of certain generative AI uses and prohibits AI systems from being credited as authors. The emerging rules address individual vulnerabilities, but Shah sees a more fundamental question underneath them. Science must decide how far its existing safeguards can be adapted to AI and where entirely new ones may be necessary. “Researchers are working hard not only to address the challenges to scientific integrity arising from AI, but to understand how AI itself can be harnessed to strengthen integrity,” Shah said. Like Shahrazad's stories, the ending remains unresolved. Answering these questions will lead to 1,001 great stories.
Poisoning the scientific record
If a scientist or a policymaker is using these AI systems to create a paper or a policy, then humans should have more oversight on what datasets were retrieved and what happened during the research process.
The return of p-hacking
By varying the prompts, you can get all kinds of results and then cherry-pick the one you like.
Research at machine speed
Now, you can prompt Claude Code or Codex to ‘write me a paper on this topic,’ and within a couple of hours, it can conduct the research and produce a paper.
Rewriting the rules of scientific trust
There are also some initiatives on rethinking the whole system. What if we were to design it from the ground up? How would that look in the AI age?
