Who decides a frontier model is safe enough to release?
Developer accountability, independent evaluation and enforceable public authority address different parts of the release decision.
Testing is not release authority
Testing a model and deciding whether it may be released are different responsibilities. An evaluator can identify risks without having authority to stop deployment. This debate asks who sets acceptable risk, who examines the evidence and who can enforce the resulting decision—not simply whether development should move faster or slower. Anthropic Tom Wheeler · Brookings
The case for liability and targeted rules
Andreessen Horowitz’s December 2025 legislative roadmap emphasizes liability for harmful uses, targeted fixes to gaps in existing law and rules that preserve competition. It argues that fraud, discrimination and deceptive conduct should remain actionable when AI is involved. This is not a rejection of government involvement: the proposal also supports federal technical testing of national-security capabilities and a national model-transparency standard. It places weight on evidence about additional risk and on safeguards proportionate to a company’s size, rather than treating every model as the same regulatory problem. Andreessen Horowitz
The case for independent evaluation
Anthropic’s September 18 Accenture partnership illustrates a model of outside scrutiny alongside continuing developer responsibility. It proposes evaluators with employee-like access to training, deployment decisions and staff, able to assess safety commitments and report incidents. Anthropic says model safety remains its responsibility and it will continue training and releasing models. The announcement does not establish a government approval requirement or give evaluators a stated release veto. Anthropic
The case for enforceable public authority
Tom Wheeler argues that self-regulation leaves too much public risk in private hands. His September 16 proposal calls for enforceable decisions about evaluator selection, qualifications, funding and standards, including authority to delay or deny model releases. On this view, identifying a danger is insufficient if nobody outside the developer can require a response. These are proposed powers, not powers already granted to the evaluators. Tom Wheeler · Brookings
What makes oversight independent?
Independence depends on how the system works. Anthropic says it will initially pay Accenture directly, while preferring pooled or government funding over the longer term; access and reporting standards remain unsettled. Wheeler treats funding, qualifications and enforcement as essential design questions. The practical disagreement is therefore also about safeguards against conflicts of interest, what findings become public and who resolves a contested release—not just whether a company has hired an external tester. Anthropic Tom Wheeler · Brookings
Sources & attribution
- Commentary 17 Dec 2025A Roadmap for Federal AI Legislation: Protect People, Empower Builders, Win the Future
Andreessen Horowitz.
- First-party report 18 Sept 2026Partnering with Accenture on embedded evaluation
Anthropic.
- Commentary 16 Sept 2026Why AI safety requires more than industry self-regulation
Tom Wheeler · Brookings.