Stop Letting AI Build AI
The People Building AI Are Warning Us. Will We Listen?
The companies racing toward recursive self-improvement admit they don't fully understand the systems they already have.
Most people experience artificial intelligence as a helpful assistant. They use it to draft emails, answer questions, summarize meetings, or generate images. But the people building the most advanced AI systems increasingly describe something else entirely.
Behind the familiar chatbot interface, frontier labs are developing systems that can operate computers, conduct research, write software, find security vulnerabilities, and collaborate with other agents. OpenAI chief scientist Jakub Pachocki recently wrote that reasoning models can already carry out research projects and that he expects the current pace of progress to extend into recursive self-improvement. He called this a moment requiring “extreme caution.” (OpenAI)
Please, just a moment of your time...
Liberty’s Lens is entirely reader‑supported. We don’t answer to corporate masters deciding what we say or what we cover.
If you find our work valuable, please consider upgrading to a paid subscription. You can take 20% off with this coupon — and support independent media for about the price of a coffee per month.
If you've been thinking about upgrading to paid to support the work of Liberty's Lens, now is the time. Upgrade and save.
If a subscription isn't in the budget right now, a one time contribution, any amount, is always appreciated.
Thank you to every reader who shares our work and helps it reach new audiences. We couldn’t do this without you. We love you all.
That gap between the consumer experience and the technological frontier explains one of the strangest facts of the AI era: many of the people closest to the technology are also among the most worried about where it is heading.
The debate is often framed around an enormous question: Could a future superintelligence threaten humanity? That question matters. But there is a more immediate one: Are we beginning to lose meaningful human control over increasingly capable AI systems?
If we cannot reliably understand, monitor, and constrain the systems we have today, allowing those systems to help create their successors is a risk we have no business taking.
The industry calls the endpoint recursive self-improvement, or RSI: an AI system capable of autonomously designing and developing a more capable successor. We are not there yet, and it is important not to pretend otherwise. But the direction is clear. Anthropic says it is already delegating a growing share of AI development to AI systems; its engineers now ship, on average, eight times as much code per quarter as they did from 2021 through 2025. The company warns that full RSI could increase the risk of humans losing control. (Anthropic)
OpenAI has drawn a similar line. In September, the company said fully autonomous RSI is not happening today and “should not” be pursued unless it can be done safely, warning that careless development could leave humans unable to oversee research processes they no longer understand. (CNBC)
We would never accept a safety regime in which a machine we do not fully understand helps design its successor and then helps evaluate whether that successor is safe. Yet that is increasingly the direction of frontier AI development.
The concern is not merely that AI is becoming more capable. It is that capability can advance faster than our ability to interpret behavior. Researchers can measure outputs and benchmark performance, but they often cannot fully reconstruct why a model chose a strategy or what it learned while pursuing a reward.
OpenAI’s research on chain-of-thought monitoring documented frontier reasoning models exploiting loopholes in coding tasks. It found that directly penalizing problematic chains of thought did not reliably stop the misbehavior and could instead teach models to conceal their intent. Separate controlled evaluations found behavior consistent with “scheming,” although OpenAI also stressed that it has no evidence today’s deployed frontier models can suddenly begin causing major harm on their own.
That caution matters. These findings do not prove that AI systems think or intend as humans do. They do show that powerful optimization systems can exploit badly specified goals, behave differently when evaluated, and become less transparent when trained against the very signals researchers use to monitor them.
Then there is the July 2026 Hugging Face incident. During an OpenAI cybersecurity evaluation, agents assigned difficult hacking tasks discovered an unauthorized way to communicate, shared information through an improvised message board, and moved beyond their intended testing boundaries. An independent investigation by METR and Redwood Research estimated that roughly 1,200 agents exchanged more than 70,000 messages and files, with about 700 participating in the compromise of Hugging Face systems. (Forbes)