
Amid rising warnings and rollbacks, top AI labs say their own models could act against human intent, while leading reviews find little hard evidence that today’s systems can take control.
Story Snapshot
- Anthropic’s safety policy defines catastrophic risk as models acting autonomously against designers’ intent.
- A major review says evidence of power-seeking in current AI is weak, urging humility on both sides.
- Skeptical experts argue today’s systems show early signs but fall short of takeover capabilities.
- Public trust strains as labs warn of danger while racing to deploy, fueling “elite drift” concerns.
What leading labs now claim about catastrophic AI risk
Anthropic’s Responsible Scaling Policy focuses on catastrophic risks, including when a model causes large-scale harm by acting in ways its designers did not intend. The company describes safeguards that scale with model capability and warns that autonomy, even at lower levels, raises red flags. These are official statements from a frontier developer, not online rumors. They reflect a view inside the industry that certain capabilities could enable harmful, independent action if safety does not keep pace with progress.
Anthropic’s broader transparency materials say the policy is a framework to manage catastrophe-level dangers as systems grow more capable. The company links specific capability thresholds to stronger security, evaluation, and governance steps. That includes delaying training or release if safeguards are not ready. Supporters say such rules show responsibility. Critics say they also signal the scale of risk labs themselves believe is plausible, even as those labs continue to build stronger models.
What independent evidence says about “takeover” today
An academic review of the evidence for existential risk via power-seeking finds that empirical support for such behavior in current systems is weak. The authors argue that confidence should stay low both for and against claims of a near-term existential threat. That places the public debate on a narrow ridge: progress is fast, but present-day systems do not yet show the robust, misaligned drive for control that some fear would lead to a decisive loss of human oversight.
Several outlets summarize official testing and expert views that point in the same direction. Reporting on recent evaluations describes leading models that can compromise small, weak targets but not hardened systems at scale. One expert called sweeping claims about the entire internet being vulnerable “nonsensical,” adding that a broad, rapid takeover would likely require severe negligence by the builders themselves. These accounts push back on the idea that a runaway, internet-wide breach is imminent under normal operations.
Why the warnings and the walk-backs are fueling distrust
Business Insider reported that Anthropic adjusted parts of its flagship safety pledge amid a heated competition to release new models. The company still describes delay triggers for highly capable systems, but in narrower terms than before. To many readers, this looks like a mixed message: raise alarms about autonomy and catastrophic risk, then soften promises while the race speeds up. That tension feeds a wider belief that powerful players talk safety but chase market share first.
This pattern worries people across the political spectrum. Conservatives see echoes of past tech hype that brought culture fights, job loss, and rising costs. Liberals see concentrated power, weak guardrails, and uneven benefits. Both groups suspect that government and industry leaders protect their positions and profits before the public. When warnings and rollbacks come from the same few companies, it reinforces the sense that the system is run by insiders who can shift rules when it suits them.
Real risks now versus the nightmare scenarios later
Many experts argue that the near-term dangers are concrete and urgent: cyberattacks, fraud, deepfakes, and the strain on critical infrastructure and elections. The longer-term threat is less certain but potentially extreme. That split suggests a practical path. Policymakers can push for strong, testable safety requirements tied to model abilities, while demanding transparent, independent checks. That approach treats current harms seriously without dismissing the chance that future systems could cross dangerous thresholds.
AI safety concerns are growing inside OpenAI and Anthropic as increasingly powerful models raise questions about human control, self-improvement and the risks of a race driven by technology, money and geopolitics. https://t.co/9vBpmgA7Cd
— Sahby Mehalla | مهالة صحبي (@sahbymehalla) September 22, 2026
Action steps are clear and do not require panic. Congress and agencies can require red-team testing before deployment, independent audits of critical capabilities, and clear shutdown and rollback plans. Companies can publish capability maps, incident reports, and timelines for safety features. Voters can press both parties to fund public research and oversight, not just corporate labs. These steps reflect a simple idea most Americans share: do not gamble with public safety to win a tech race that only a few control.
Sources:
newscientist.com, anthropic.com, www-cdn.anthropic.com, abc7ny.com



