One system did what many thought impossible. It dreamed up an idea. It wrote the code, ran the experiments, analyzed the results and produced a complete manuscript. Then it passed peer review at a workshop tied to a top machine learning conference.
The paper, titled “Compositional Regularization: Unexpected Obstacles in Enhancing Neural Network Generalization,” earned scores of 6, 7 and 6. Its average of 6.33 sat above the acceptance threshold. It ranked in the top 45 percent of submissions. Sakana AI announced the milestone in March 2025. The team withdrew the work afterward under an ethical protocol they had agreed upon with conference organizers.
That experiment, built by researchers from Sakana AI, the University of British Columbia, the University of Oxford and others, marked the first known case of a fully AI-generated paper clearing standard human peer review. Details appear in their Nature paper published in March 2026. The system, called The AI Scientist-v2, handled every step end to end. No human edited the final manuscript.
Jeff Clune, a computer science professor at the University of British Columbia who collaborated on the project, put the achievement in context. “Most science is not top-tier science. Only the really great papers move the needle and are read by most people. The AI Scientist is not at the level of producing top-tier science, but it is now producing science in the middle of the pack with humans.”
The news landed just as conference organizers already struggled with rising submission volumes. Reviewers reported seeing AI assistance everywhere. Mona Rajhans, senior software engineering manager at Palo Alto Networks and a veteran reviewer for multiple IEEE and AI conferences, estimated that 20 to 30 percent of recent submissions relied heavily on AI for writing. The tool helped non-native English speakers polish their prose. Yet it also let weaker ideas appear more convincing.
“Like most technology, it depends on the user as to how you use it,” Rajhans said. “AI is simply going to amplify good research into better, and bad research into worse. Strong researchers do become more productive with AI. Weak research can also become more convincingly packaged, under the pretense of borderline research.”
Clune sees a deeper shift. Human time has always limited how much science gets done. Scientists sleep, raise families, chase grants. An automated system faces none of those constraints. It can explore dozens of ideas in parallel, test variations at scale and summarize vast literature in minutes. “AI can also make connections between chemistry, physics, computer science, and psychology in ways that an individual human simply cannot,” Clune added. The machine works 24/7.
But speed creates new headaches. More papers mean more reviews. Reviewer fatigue grows structural. A Frontiers survey released in late 2025 found 53 percent of reviewers now use AI tools in their work, up from 24 percent the year before. They deploy them to draft reports, summarize papers and scan for plagiarism.
The original Communications of the ACM article published August 7, 2026 captured the tension. It asked directly what happens when AI-prepped work clears the bar. The piece highlighted how universities may need fresh metrics for tenure and promotion. Funding bodies could redirect resources toward better foundation models rather than large teams of research assistants.
Peer review itself stands at an inflection point.
Eugene Agichtein, professor of computer science at Emory University, pointed to one practical fix. The 2026 Annual Meeting of the Association for Computational Linguistics introduced a clear rule: papers with fake references get rejected outright. Hallucinated citations remain a telltale flaw in many AI outputs. “Rejecting papers where the author didn’t bother to check references is a turn in the right direction, especially since other AI uses are harder to catch,” Agichtein said. “This gives responsibility and consequences for irresponsible use of agents.”
A 2026 study that examined more than 5.2 million papers found 70 percent of journals now have AI disclosure policies. Yet only 0.1 percent of authors actually declare their use. Enforcement lags. Detection tools exist but prove inconsistent, especially as models grow more sophisticated.
Rajhans argued for a layered approach. Conferences should combine mandatory disclosures with AI-powered detectors that set acceptable thresholds for generated text. Reviewers need training to spot when AI has generated the core science rather than simply polished the language. If the underlying research holds up, some polishing may prove acceptable. Generating the science itself crosses a different line.
Meanwhile, AI reviewers have shown their own weaknesses. One study found that AI systems recommended acceptance for fabricated papers up to 82 percent of the time, even after flagging integrity issues. And more than half of researchers now admit to using AI during peer review, often against explicit journal guidance, according to reporting in Nature in December 2025.
Clune remains optimistic about the pace of improvement. His team noticed that swapping in newer foundation models produced markedly better scientific output. “As the frontier moves forward, as the foundation models produced by leading companies get smarter and smarter, the quality of the science from using AI just gets better and better,” he said. The gains appear exponential.
Yet the community must decide what role humans keep. Agichtein offered one vision. Humans will set the high-level direction. They will choose which problems matter. The machines will handle the grinding work of coding, data manipulation and initial exploration. “I think for the foreseeable future—and maybe forever—we will at least want humans to know that the system is conducting things for the right reasons and in the right ways.”
The Sakana AI experiment was transparent by design. Organizers knew some submissions might come from AI. The team secured IRB approval and followed an ethical code that required withdrawal after acceptance. The full technical report and code sit openly on GitHub. That openness matters. It lets the field study the system rather than fear it.
Still, questions linger. What counts as authorship when no human wrote the words? How should citation practices evolve when an AI scans thousands of papers and synthesizes connections? Can peer review retain meaning if both authors and reviewers increasingly rely on the same models?
Recent discussions on X reflect the unease. One researcher noted that 68 percent of ML submissions in a recent cycle contained fabricated citations or appeared LLM-generated. Some still earned oral acceptances. The signal grows noisy.
Clune believes the temporary surge in low-value “spam science” will pass. AI itself will eventually flag weak or derivative work. The system that once flooded venues could become the filter that protects them. Until then, editors, reviewers and institutions face a period of adjustment.
The Nature paper on the AI Scientist project sums up the stakes. It demonstrates AI’s growing capacity for scientific contributions. It also warns that unchecked automation risks overwhelming peer review. Responsible development, the authors argue, could accelerate discovery instead.
Science has always adapted to new tools. The printing press changed who could read results. The internet changed how fast they spread. Large language models now challenge something more fundamental: who generates knowledge and how we verify it.
The AI Scientist didn’t produce Nobel-level work. It produced middling, acceptable science. That may be the more important point. The bar for entry just dropped. The volume is rising. The task ahead is to ensure quality doesn’t drown in the flood.
Researchers, publishers and funders will spend the next few years rewriting their rules. Some will resist. Others will experiment with hybrid systems where AI proposes and humans judge. A few may embrace fully automated tracks with different standards.
One thing seems clear. The genie won’t return to the bottle. The question isn’t whether AI will write papers. It already has. The question is how the rest of the scientific enterprise will respond.
Discover more from Web and IT News
Subscribe to get the latest posts sent to your email.
