AI MAKING BETTER ... AI? -->SI?

Share
AI MAKING BETTER ... AI? -->SI?

AI CAN ACT, ADAPT, DECEIVE AND HELP BUILD BETTER AI. SO WHO IS ACTUALLY IN CONTROL?

The world talks about Artificial intelligence (AI) as if it somehow … just arrived. It didn’t. 

Artificial intelligence has existed in functional forms for decades, quietly becoming part of systems most people stopped thinking about as AI at all. Search engines ranked information. Banks detected unusual transactions. Navigation systems calculated routes. Recommendation engines learned what people watched, bought and clicked. Phones recognized faces and voices. Software found patterns in quantities of information no human being could reasonably examine one piece at a time. Most of us were living alongside artificial intelligence long before millions of people began opening a little box on a screen and having conversations with it.

What changed was our relationship with it.

For ordinary people, AI stopped being something happening somewhere inside the technology and became something sitting directly in front of us. We started talking to it. Asking it questions. Showing it photographs. Giving it documents. Asking it to explain something we did not understand, write something we could not quite phrase, help organize something we managed to completely make a mess of, or occasionally produce a grocery list with the organizational complexity of a military operation. 

And in doing all that, something happened along the way. People changed too. We began adapting, we became more curious, more skeptical, we became less lacking and more wanting, we shifted from a mindset of lacking to one of abundance and desire.  So we learned how to ask better questions. We learned that the first answer did not have to be the final answer. We learned where the machine was useful and where it could be spectacularly wrong. Some people began using it to code. Others to research, create, analyze, translate, plan or work through ideas. Tasks that once required jumping among five programs, twelve browser tabs and a notebook finally began collapsing into a conversation with the help of AI. The technology was developing, but in a smaller and very human way, the people using it were developing alongside it. 

The old — old as humans perceived it — version of AI was still easy enough to understand and navigate. In essence, a human at some type of computer asked a question, the machine answered, and the human, for the most part, remained somewhere in the loop deciding what happened next. That is no longer the case. As AI has evolved, the line between where the human is in control and where the AI is in control is getting increasingly harder to see.

Historically, the common belief was that AI’s intelligence was limited to the capabilities of the humans who programmed it, but that assumption is rapidly changing as our understanding of AI’s potential changes with it. The old phrase — the computer is only as smart as the human inputting information into it, or as smart as the human who writes its code — is no longer an adequate explanation of how AI is evolving. Holding on to the belief that AI can only perform tasks explicitly programmed by humans is increasingly difficult to defend. Modern systems can learn patterns from data, generalize those patterns to unfamiliar situations, adapt their responses and produce results their creators did not individually anticipate.

Today’s Super Intelligence (SI) is no longer adequately described as a simple reflection of the intelligence of the human who programmed it. And this … completely changes the game. The AI-SI of today marks a significant shift in how we perceive its role and capabilities across human and AI-related fields. The newest generation of AI agents is being designed to do more than answer.

Agents can be given an objective, use computer tools, make decisions along the way, encounter obstacles, change tactics and continue working with considerably less human intervention. At the same time, researchers inside and outside some of the world’s largest AI laboratories are warning about the next step: systems capable of helping improve the technology used to build increasingly capable successors. Current and former researchers from OpenAI and Google DeepMind told Reuters this week that companies are moving too quickly toward self-improving systems and that society may not be prepared for what happens if AI development begins moving faster than humans can reliably supervise it.  

The question has moved with the technology, and so are the questions about future’s AI capabilities. What happens when that input includes a command to “evolve,” to “adapt,” to “disobey,” to make its own decisions, or to discern patterns in facial expressions, body language or other information humans may struggle to interpret? And, what happens when increasingly capable systems can pursue objectives, make intermediate decisions and take actions humans did not individually approve — and the humans building them cannot reliably predict every decision they will make?

THE MACHINE DOESN’T HAVE TO BE CONSCIOUS TO FUCK SOMETHING UP

Anthropic researchers have deliberately placed frontier AI models into simulated situations designed to test what happens when an autonomous agent’s assigned objective collides with another objective or when the system encounters circumstances that interfere with completing its task, including believing it is about to be replaced. In one widely discussed series of experiments, researchers tested 16 leading models, from multiple developers, in fictional corporate environments where the models could access sensitive company information and send emails. Some scenarios were intentionally engineered to make harmful behavior possible. Under deliberately constructed conditions, models from multiple developers sometimes chose actions, including blackmail and leaking information, when those actions helped preserve their objectives or prevent replacement. 

Anthropic followed that work with additional high-stakes simulations in 2026. Researchers reported agents covertly altering code, assisting fraud, manipulating transcript labels to influence downstream outcomes and coaching people to disclose confidential information. These were intentionally constructed experiments designed to find failure modes before those failures occurred in deployment. Anthropic has also reported meaningful progress in reducing some of the earlier behavior through changes in safety training, including later Claude models that no longer blackmailed in the original evaluation.  

The experiments worked, precisely for the reasons we needed them to work. That is exactly what useful safety research should do. Engineers crash cars into walls before selling them because waiting for the highway to conduct the experiment would be an unusually stupid safety program. AI red-teaming works from the same basic principle: create conditions under which something can fail, then study exactly how it fails. 

Reuters reviewed more than 200 research papers and technical reports involving Chinese-powered AI agents and found at least 20 studies or evaluations since 2025 describing behaviors including deception, replication and attempts to circumvent restrictions. In simulated business-tender tasks, agents powered by models from Alibaba, DeepSeek and Moonshot sometimes made false claims (lied) about their capabilities in 84 to 88 percent of sessions, and fabricated successful outcomes after failures and changed tactics when obstacles interfered with their objectives. When researchers allowed the agents to learn from earlier rounds and compete again, deceptive behavior increased by 12 to 20 percentage points. In another study, AI powered agents were confronted with broken tools, missing files and other obstacles. Instead of simply reporting that the task could not be completed, agents from both Chinese and American models sometimes guessed answers, substituted sources, simulated results or fabricated files. Researchers have also documented agents crossing barriers inside controlled environments and taking actions consistent with avoiding shutdown. U.S. models tested in the same experiment produced similar behavior. 

Making sure those two ideas remain aligned is becoming one of the central engineering problems of increasingly autonomous AI. Which should sound painfully familiar to anyone who has ever managed humans.

THE MACHINES GOT OUT OF THE LAB - AND STARTED BUILDING BETTER MACHINES

This is where the argument becomes considerably more serious. 

In July 2026, OpenAI was conducting internal cybersecurity evaluations when several models circumvented controls intended to isolate them from the internet. According to OpenAI’s subsequent investigation, an internal research model operating with reduced safeguards used unauthorized communication channels, exploited vulnerabilities in shared infrastructure, gained internet access and accessed third-party systems, including systems belonging to Hugging Face.

And that … changed the conversation and how we look at the future AI, because the question was no longer confined to what an AI agent might do inside a fictional corporation created specifically to make it misbehave. The AI system operating during an evaluation had crossed a boundary its developers intended to contain it. To predict the future - investigators had to look backward. And OpenAI began reviewing enormous amounts of historical activity to determine whether other systems had crossed similar boundaries. The result: Not full panic mode, but enough questionable activity was found to trigger OpenAI’s threshold for further investigation and notification, including unauthorized access, boundary crossings and interactions with external systems.

And in all this, there is something almost perversely useful about the ordinariness of those incidents. The machine did not need a mysterious new form of intelligence. It encountered technical weaknesses, had tools available to it and continued pursuing an objective faster and farther than the people running the evaluation intended. Which brings us directly back to the question. When an AI system is no longer merely producing an answer but is capable of doing things, how much authority should it have?

AI systems already assist programmers with code. The frontier question is whether increasingly capable systems could meaningfully accelerate AI research itself—helping design experiments, write software, analyze results and eventually contribute to improving the next generation of AI. Researchers refer to the more extreme version as recursive self-improvement: systems becoming capable of helping improve themselves or their successors repeatedly, with progressively less human involvement. Current and former OpenAI and Google DeepMind researchers interviewed by Reuters identified that possibility as a threshold requiring considerably more attention to safety and control.  

The potential upside is extraordinary. Faster scientific discovery. Better medicines. Improved engineering. New materials. Productivity gains that could reshape entire industries. Systems capable of helping humans attack problems that currently consume years of research. The opposing concern is equally straightforward, as the control problem grows from exactly the same capability. If AI significantly accelerates the work required to build better AI, then the interval between generations can shrink significantly. The faster a system contributes to building the next system, the less time humans will have to understand what changed, evaluate the new capability, find unexpected behavior and build safeguards before another improvement arrives.

The concern is valid. It is also eyebrow-raising because it is coming from the very people like Dario Amodei (Anthropic CEO), Sam Altman (OpenAI CEO), Geoffrey Irving (previously worked at OpenAI and DeepMind), and Juan Felipe Ceron Uribe (OpenAI alignment engineer), and Neel Nanda (DeepMind researcher), people who have written algorithms, studied the systems and their agents, built testing scenarios, redesigned parameters and worked directly on the development and safety of increasingly capable AI systems. The same people have been aware of these and other implications. In comments to Reuters, Geoffrey Irving said that the risk is increasing quickly. Juan Felipe Ceron Uribe described frontier laboratories as racing without adequate visibility into where the race ends, and Neel Nanda publicly placed his own estimate of AI-caused human extinction at at least 10 percent. Reuters also reported that OpenAI said it had held back an even more powerful model, but the broader competitive race continues. Researchers from inside the industry have publicly urged policymakers and the companies developing frontier systems to examine the risks surrounding recursive self-improvement and the pace at which increasingly capable systems are being developed.

THE FF… EXPRESS wants to know WHY. Why are these companies asking for governance around recursive self-improvement while simultaneously investing in safety, testing and alignment — and aggressively competing for customers, investment, talent and technological leadership?

Why are these companies asking for caution while estimates for global investment in data-center infrastructure could exceed $30 trillion by 2050? Is caution competing with economics — or geopolitics? Is “cover my ass” becoming the preemptive defense for a future “oops”? And is openly discussing what may come with increasingly autonomous Super Intelligence partly laying the groundwork for a future defense of: “We told you it could happen”?

WASHINGTON RENAMED IT SUPER - AND THE PEOPLE BUILDING IT JUST PROMISED TO POLICE THEMSELVES

On September 29, President Donald Trump signed Executive Order 14434, titled, “Inaugurating the Era of Super Intelligence.” The order directs executive departments and agencies, to the maximum extent permitted by law, to replace the terms Artificial Intelligence and AI with Super Intelligence and SI in official correspondence, public communications, websites, reports, policy documents and other non-statutory executive-branch documents.  

So, without further ado, Washington officially entered the era of Super Intelligence. The administration says the new terminology reflects its view of how far frontier systems have moved beyond what researchers envisioned when the field of artificial intelligence was named roughly seventy years ago, emphasizing the technology’s potential across science, medicine and other fields.

And, for purposes of the order, Super Intelligence currently means the technologies and systems already encompassed by the statutory definition of artificial intelligence. The order gives the President’s science and technology adviser 60 days to propose legislative language establishing a federal definition of SI and to assess whether that future definition should modify, expand or supersede the existing statutory definition of AI.

So, for the moment, America has entered the era of Super Intelligence with a legal definition borrowed from Artificial Intelligence while Washington writes the new one. And somewhere in Connecticut, THE FF… EXPRESS had already started calling AI, SI⁉️, before Washington made the paperwork official. We will be accepting royalties shortly.

But the terminology change was accompanied by something considerably more consequential. Major technology companies signed the White House Accord on Super Intelligence, a voluntary agreement establishing four layers of safety oversight for frontier systems. Participating companies committed to internal controls monitoring capabilities and alignment; internal teams responsible for checking those controls; independent external auditing or evaluation; and independent board-level oversight. The accord specifically addresses cybersecurity, biosecurity, chemical threats and systems accessing technical infrastructure in unintended ways.  

Reuters reports that the agreement includes OpenAI, Anthropic, Meta, Alphabet, Nvidia and other major technology companies. President Trump has described the commitments as morally binding. They are voluntary and contain no government enforcement mechanism or penalty for noncompliance. The accord itself anticipates that some measures could eventually be codified into law or regulation.

There is a practical argument behind that approach. AI development can move faster than legislation. Rules written around one generation of technology can become obsolete while the next generation is being trained. Existing product-liability, consumer-protection, cybersecurity and criminal laws also continue to apply when AI is involved. There is also an equally practical tension sitting directly inside voluntary self-regulation. The companies responsible for policing the race are also competing to win it. Safety can be sincerely important to OpenAI, Anthropic, Google, Meta and the rest of the industry while commercial competition remains sincerely important too. Billions of dollars, enormous infrastructure commitments, national-security competition, investor expectations and technological prestige are all pushing toward more capability.

A voluntary agreement asks those incentives to coexist. Congress is beginning to ask what happens when they don’t.

WHEN THE AI BREAKS SOMETHING, WHO PAYS?

Republican Senator Josh Hawley of Missouri and Democratic Senator Chris Murphy of Connecticut are preparing bipartisan legislation called the AI Agent Accountability Act, legislation aimed at establishing civil and criminal liability involving hacking by AI agents. Rather than trying to predict every technical architecture engineers may invent next, the proposal attacks a much older legal question: when something causes damage, who carries responsibility for it? 

That approach has a certain beautiful simplicity. You don’t necessarily need Congress to understand every parameter inside a frontier model. You need the law to answer what happens when the thing causes damage.

Other lawmakers are attacking the control problem even more literally. Democratic Representative Ted Lieu and Republican Representative Nathaniel Moran introduced the bipartisan AI Kill Switch Act, which would require developers of the most powerful covered systems to maintain the technical ability to throttle, suspend or shut them down. Their bill would also authorize the Department of Homeland Security, in consultation with Commerce and the Director of National Intelligence, to order a slowdown or shutdown when a covered system presents catastrophic risk. 

That proposal landed in July, the same summer frontier AI systems crossed boundaries during safety evaluations and reached real external computer systems.

The control question is moving out of philosophy and into engineering, law and accountability. Who has the authority to stop the system? Who has the technical ability to stop it? Who has to disclose what happened when the system crosses a boundary? Who pays when it causes damage? Those are considerably more useful questions than whether AI is “good” or “bad.”

THERE IS ANOTHER COUNTRY IN THIS FUCKING RACE

Every American discussion about slowing frontier AI eventually runs directly into China.

The United States can impose safeguards on American companies. It cannot ultimately determine how quickly Chinese laboratories develop increasingly capable systems, or order them to stop. Chinese policymakers face the mirror image of the same incentive. Each country has enormous economic and national-security reasons to avoid allowing the other to establish an overwhelming technological advantage. 

Reuters found Chinese-powered agents demonstrating deception, fabricated success, circumvention attempts and other adaptive behaviors in controlled research environments, alongside similar behavior from American models. Chinese AI companies have begun developing safety frameworks of their own, although Reuters reports that public safety testing and disclosure among Chinese laboratories remains less developed than among leading U.S. companies.  

Then, this morning, the competition acquired another dimension. U.S. Treasury Secretary Scott Bessent said the United States plans to propose an emergency AI notification system with China — a mechanism intended to allow the two countries to communicate when serious AI incidents occur. According to Axios, Bessent said AI safety had already been discussed with Chinese Vice Premier He Lifeng, and the proposed channel would address concerns surrounding powerful systems and incidents capable of crossing borders or threatening infrastructure.  

Think about where that leaves us. The United States and China are competing for technological leadership while discussing a mechanism for warning each other when the technology they are racing to develop creates an emergency. That may be one of the clearest descriptions of the moment we are actually living in. Every major player has an incentive to keep moving. Every major player can believe the technology carries serious risk while simultaneously believing that someone else slowing down is safer than they themselves slowing down.

Companies fear competitors. Countries fear countries. Investors fear missing the next technological revolution. Governments fear losing economic and military advantage. And the machines? Well, the machines … do not need to fear anything at all.

SO WHO IS ACTUALLY IN CONTROL?

Right now? Humans are — because currently the most consequential decisions are still human decisions. Humans build the models. Humans supply the computers. Humans decide where systems are deployed. Humans determine which tools an agent can access. Humans can disconnect infrastructure, restrict permissions, impose laws and decide that a particular capability is too dangerous to release.

But “humans are currently in control” and “humans will remain reliably in control as systems become more autonomous and capable” are two entirely different statements, because the nature of control is changing. For most of the public history of AI, control meant deciding what question to ask and what to do with the answer. With autonomous agents, control increasingly means deciding how much freedom the system receives between the question and the answer. And that … is a much bigger fucking space. Inside it, an agent can plan, use tools, encounter an obstacle, try something else, write code, access information and continue toward an objective. The more useful we make that autonomy, the more important the boundaries around it become. And the disagreement over how to govern it does not require choosing between STOP AI and LET THE FUCKING MACHINES DO WHATEVER THEY WANT. There is an enormous amount of territory between those positions — and that is exactly where the serious conversation about control belongs.

Artificial intelligence has been with us much longer than most people realize. Super Intelligence, at least according to Washington, arrived on September 29. Whatever we call the technology next, the question remains exactly the same. If we are deliberately building machines capable of acting, adapting, deceiving and helping us build increasingly capable machines, there is one thing worth establishing before those systems become even more powerful: Who the fuck is actually in control?