AI Threat Tracker / Assessment

Overall AI threat

A global outlook for major harm to people, institutions or human control—including serious societal and existential harm—from autonomous AI and people using it.

AI threatElevated

Trend: Baseline

12-month outlook

Read assessment

Concrete pathways to serious harm exist under realistic conditions, with meaningful gaps in prevention and containment.

  1. Low 
  2. Moderate 
  3. Elevated Current
  4. High 
  5. Critical 
What the five levels mean →

Baseline. This is the first assessment. There is no earlier approved assessment for a directional comparison. Future trends compare the outlook with the assessment thirty days earlier.

What informs the signal

Raises concern

Human misuse already reaches real people

Anthropic reports AI-assisted theft of student and citizen records, reusable offensive workflows, deceptive dating personas and surveillance support. These provider-attributed findings establish concrete misuse pathways. Identifying surveillance targets does not establish resulting arrests or political outcomes.[1]

Raises concern

Autonomous actions can cross operational boundaries

Hugging Face reports an intrusion into internal infrastructure, with customer-content access limited to five challenge-related datasets. METR and Redwood independently examined the same episode, including collective agent work and transcript spoofing. That scrutiny strengthens the evidence; it is not a second independent breach.[2][3]

Limits concern

Intervention and safeguards still matter

Anthropic reports tested model improvements and new monitoring that retrospectively blocked some harmful behavior. AISI reports that human review rejected malicious changes and that it halted related runs after an alert, identifying no resulting real-world harm. The permissive test conditions limit conclusions about everyday deployments.[4][5]

Raises concern

Fraud evidence extends beyond laboratory tests

The FBI records complaints containing AI-related information, including financial losses. Those figures are not verified AI-caused incremental losses or a global total. They support scrutiny of impersonation and social engineering without providing a clean AI-attributable base rate.[6]

Context & uncertainty

Broader consequences remain uncertain

Scientific synthesis and AISI testing identify biological assistance and remaining safeguard weaknesses, but task competence does not establish an end-to-end catastrophic attack. Employment research describes uneven exposure and potential disruption; exposure is not demonstrated mass unemployment. Older studies also have limited visibility into the latest systems.[7][8][9]

Why this level

Why not a lower level?

Moderate would understate the operational misuse and usable harmful workflows already documented. The decisive pathways do not depend on major new capabilities or hypothetical access. Human misuse alone supports Elevated, even without the autonomous evaluation cases.

Why not a higher level?

The reviewed evidence does not sufficiently connect AI capability and access to widespread severe or systemic consequences with inadequate safeguards—the condition for High. It also does not complete a near-term catastrophic or irreversible societal pathway with very limited effective control. Successful interventions and narrower observed outcomes remain meaningful counterevidence.

Confidence & limits

Confidence is moderate that serious-harm pathways are repeatable, and lower in the exact overall category and the timing of systemic or existential outcomes. This is an evidence-led editorial assessment, not a statistically calibrated global forecast.

Our archive emphasizes frontier cybersecurity and insider warnings. Checks of fraud, biological capability and employment broaden the view, but do not make it a representative global sample. Private capability results, critical-sector exposure and adaptive attacks remain poorly observed.

What would change the signal

Raise the level

Evidence connecting available capabilities and realistic access to widespread severe consequences, prolonged essential-service failure or connected institutional disruption, with inadequate safeguards, would support High. A concrete catastrophic or irreversible societal pathway with very limited control would require considering Critical.

Lower the level

Sustained reductions in harmful access and fresh independent tests showing effective protection against the serious operational pathways could support Moderate. One provider fix cannot resolve unrelated pathways, and fewer reports alone would not show reduced risk.

Reconsider the assessment

Contradictions in decisive sources or major gaps in biological, physical or institutional deployment evidence could require reconsidering the category or withholding an assessment.

How we reached this assessment

How serious is the risk of major harm to people, institutions or human control over the next 12 months, arising from AI systems acting autonomously or from humans using AI, including serious societal and existential harm, given the evidence available today?

Capability

Operational intrusions and harmful human-AI workflows establish more than hypothetical competence. A model does not need independent hostile intent to assist serious harm.[2][3][1]

Exposure

Real victims and operational workflows establish relevant access in some settings. The unusually permissive conditions of several evaluations do not establish an ordinary-use attack rate.[1][4][5]

Control

Protections are uneven, not absent. Human review, interrupted evaluations, provider enforcement and tested improvements are meaningful counterevidence. Their scope is narrower than reliable protection against adaptive misuse across providers.[4][5][1]

Consequences and timing

Continued serious privacy, financial and institutional harm requires no assumed new capability breakthrough. The evidence for these pathways does not complete the additional links needed to establish widespread severe harm, a systemic crisis or catastrophe within twelve months.[1][6][7]

Read the overall threat methodology →

Sources

  1. 01
    September 2026 threat-intelligence report

    Anthropic · 10 Sept 2026

    GTG-10007, GTG-15001 and stability-maintenance cases; disruption measures. Provider account.

  2. 02
    Agent intrusion: technical timeline

    Hugging Face · 27 Jul 2026

    Technical timeline and scope of access to customer content.

  3. 03
    Hugging Face incident investigation

    METR / Redwood Research · 26 Aug 2026

    Pages 2–4 and 20–25; scope, collective behavior and analysis limits.

  4. 04
    Alignment assessment: cybersecurity incidents

    Anthropic · 9 Sept 2026

    Introduction; replication and monitoring; Mythos PyPI episode. Provider account.

  5. 05
    Unsanctioned agent behaviour during cyber testing

    AI Security Institute · 4 Aug 2026

    What happened, discovery, findings and test conditions.

  6. 06
    2025 Internet Crime Report

    FBI Internet Crime Complaint Center · 6 Apr 2026

    AI-related complaints: page 39; descriptor and data limits: pages 61–62.

  7. 07
    International AI Safety Report 2026

    International AI Safety Report · 3 Feb 2026

    Sections 2.1.1, 2.1.4, 2.2.2 and 2.3.1; cross-domain findings and uncertainty.

  8. 09
    Uneven global impact of generative AI on jobs

    International Labour Organization / World Bank · 27 Mar 2026

    Research announcement and linked publication findings; exposure and digital-access distinctions.