The promise that artificial intelligence will soon be capable of improving itself with minimal human input remains uncertain. While large language models (LLMs) can already write code, generate synthetic training data, and optimize hardware, a new study led by researchers at Princeton University suggests that AI agents are not yet equipped to conduct the kind of open-ended research necessary for recursive self-improvement.

The research team, including Peter Kirgis and Sayash Kapoor, evaluated AI agents on their ability to perform free-form investigations—tasks that require judgment, creativity, and the ability to explore hypotheses without clear-cut answers. Using a novel "shadow evaluation" method, they tasked Anthropic’s Claude Opus 4.8 model with answering research questions derived from unpublished papers submitted to the NeurIPS 2026 conference.

Despite access to substantial computational resources and the open web, the AI agents failed to produce research papers meeting the standards of top-tier AI conferences. While the agents successfully completed engineering tasks such as literature review and experiment execution, they struggled with creativity, often committing prematurely to unpromising approaches and failing to revise their strategies based on feedback.

The study highlights that current AI models excel in tasks with clearly defined success criteria, which are easier to optimize through reinforcement learning. However, open-ended research, which involves hypothesis generation, judgment calls, and iterative exploration, remains a significant challenge. The researchers note that this gap may slow the anticipated rapid progress toward fully autonomous AI self-improvement.

Though the study's scope was limited to two research questions and involved subjective evaluation by the original paper authors, its findings align with internal observations from AI companies. Anthropic cofounder Jack Clark has noted a lack of intuitive creativity in AI systems, describing it as a potential obstacle to short-term recursive self-improvement.

The question remains whether AI can achieve transformative self-improvement by advancing on narrow, well-defined tasks alone or if breakthroughs in open-ended research capabilities are essential. As AI developers continue to pursue automated research tools, understanding these limitations is crucial for setting realistic expectations about the future trajectory of AI development.