[O_O][OoO]Natural Stupidity broke prod markets money ledger this week hall of fails studio cookbook toys about profile

AI Broke Prod

Incidents where an AI system caused a technical failure with real consequences — sourced, dated, categorised — plus the AI providers' own outage notices, counted as they publish them. Companies and products are named; people are not.

ai service outages, from the providers' status pages

provider7 days30 dayslatest notice
OpenAI1424Elevated errors in ChatGPT Work 2026-09-16 19:34 UTC
Anthropic519Issues with Google Play subscriptions 2026-09-16 16:51 UTC
Cursor416Investigating service degradation — Grok Bot 2026-09-16 20:57 UTC
GitHub29Degradation with Gemini 3.8 Flash 2026-09-16 17:48 UTC
Perplexity14Maintenance: Status page migration to incident.io 2026-09-10 21:36 UTC
Replit00Agent turns aborting for some users 2026-08-01 18:45 UTC

Counts are incident notices (any severity) on each provider's public status page, read every 10 minutes and updated here in place. A quiet page can also mean an under-reporting page.

the register

42 incidents · 2023: 3 · 2024: 8 · 2025: 18 · 2026: 13 · sources last checked 2026-09-16

This table is not a feed. A row is added only when a postmortem, a company statement, a court document, a vendor advisory or top-tier reporting says so, which is why the newest row is usually older than today. Anything fresher is in the outage counts above and in the news below.

dateincidentwhat happenedsource
2026-09-03ChatGPT, Claude and Grok went down within the same hour, each for its own reason
OpenAI, Anthropic, xAI · ChatGPT and Codex; Claude Mythos/Fable 5.1 and 5, Opus 5, 4.8 and 4.6; Grok
AI service outage
On the morning of September 3, 2026 (US time) the three largest model providers posted incidents within about an hour of each other. Anthropic's status page opened 'Elevated errors for multiple models' at 13:26 UTC, listing Mythos/Fable 5.1 and 5 and Opus 5, 4.8 and 4.6, and declared the impact over at 16:16 UTC. xAI began investigating Grok at about 13:30 UTC. OpenAI opened 'Elevated errors across ChatGPT and Codex' at 14:43 UTC, applied a mitigation at 15:17 UTC and resolved it at 16:55 UTC. The status pages of AWS, Google Cloud, Azure and Cloudflare showed nothing relevant.
Impact: About three hours of elevated errors on five Claude models, a little over two hours on ChatGPT and Codex, and Grok disrupted in the same window.
Root cause: Three different ones: OpenAI cited a routing error, Anthropic an infrastructure issue, xAI an outage at its Memphis compute site. No shared cause was found.
Times are quoted from the providers' own incident notices; the per-company causes are as reported by The Register once the incidents closed.
The Register
second source
2026-04-23Anthropic postmortem: three shipped changes degraded Claude Code for weeks
Anthropic · Claude Code (Sonnet 4.6, Opus 4.6, Opus 4.7)
AI service outage
After weeks of reports that Claude Code felt less capable, forgetful and repetitive while burning usage limits faster, Anthropic's postmortem identified three overlapping causes: a default reasoning-effort change from high to medium (Mar 4-Apr 7), a prompt-cache clearing bug that dropped reasoning history every turn in stale sessions (Mar 26-Apr 10, fixed in v2.1.101), and a system-prompt verbosity constraint that cut evaluation scores about 3% (Apr 16-20). Each was reverted once identified.
Impact: Anthropic reset usage limits for all subscribers and committed to soak periods, gradual rollouts, tighter system-prompt controls and broader internal testing on public builds.
Root cause: Three independent configuration and prompt changes shipped without evaluations sensitive enough to catch the regressions.
Quality degradation rather than downtime; included because a formal postmortem attributes the impact.
Anthropic engineering postmortem
2025-11-18Cloudflare outage: Bot Management ML feature file doubled in size and crashed proxies
Cloudflare · Bot Management machine-learning model (feature configuration file)
AI service outage
A ClickHouse permissions change made the query that generates the Bot Management ML model's feature file return duplicate rows, roughly doubling the file; the proxy's bot module enforced a hard limit of 200 features and the FL2 proxy panicked on the oversized file. Core CDN and security traffic returned HTTP 5xx errors from 11:20 to 17:06 UTC, and Turnstile, Workers KV, Access and the dashboard were affected.
Impact: About 5 hours 46 minutes of widespread 5xx errors across Cloudflare's network; customers on the older proxy engine saw every request scored as a bot score of zero.
Root cause: Duplicate rows from a database permission change inflated the ML feature file past a hardcoded limit, triggering an unhandled error in the proxy.
Not a model-inference failure: the fault was in the configuration pipeline feeding the ML bot-scoring model, as attributed by Cloudflare's own postmortem.
Cloudflare postmortem
2025-09-17Anthropic postmortem: three infrastructure bugs degraded Claude responses for weeks
Anthropic · Claude (Sonnet 4, Opus 4/4.1, Haiku 3.5) serving infrastructure
AI service outage
Between Aug 5 and Sep 4, 2025 three overlapping bugs degraded Claude's output quality: a routing error sent some Sonnet 4 requests to servers configured for the 1M-token context window (16% of Sonnet 4 requests at the worst hour on Aug 31), a TPU token-generation misconfiguration inserted stray Thai and Chinese characters and syntax errors, and an XLA:TPU approximate top-k miscompilation corrupted token selection. Anthropic estimated about 30% of Claude Code users saw at least one degraded response.
Impact: Weeks of degraded answers on the first-party API and Claude Code, rollbacks between Sep 2 and Sep 12, and new continuous production quality monitoring and more sensitive evaluations.
Root cause: A load-balancer routing misconfiguration, a TPU generation misconfiguration and a mixed-precision compiler bug in approximate top-k sampling.
Date is the postmortem; impact window Aug 5-Sep 4, 2025.
Anthropic engineering postmortem
2024-12-11OpenAI outage: new telemetry service overwhelmed Kubernetes control planes
OpenAI · ChatGPT, OpenAI API and Sora (Kubernetes serving platform)
AI service outage
A newly deployed telemetry configuration generated massive Kubernetes API load across OpenAI's largest clusters, overwhelming control planes and breaking DNS-based service discovery; DNS caching delayed the symptoms until the rollout had propagated widely. ChatGPT, the API and Sora were down or degraded from 3:16 PM to 7:38 PM PST.
Impact: About 4 hours 22 minutes of outage across all major OpenAI products; recovery required scaling clusters down and blocking expensive API calls to regain control-plane access.
Root cause: A telemetry service whose Kubernetes API cost scaled with cluster size saturated the control plane.
OpenAI status page postmortem

in the news

Salesforce’s massive outage exposes the hidden risks of cloud dependencies

Salesforce’s massive outage exposes the hidden risks of cloud dependencies    cio.com

AI Incidents & Outages · cio.com · · open ↗ · share

Salesforce Global Outage Hits Customers on Second Day of Dreamforce

Salesforce Global Outage Hits Customers on Second Day of Dreamforce    channelinsider.com

AI Incidents & Outages · channelinsider.com · · open ↗ · share

Salesforce stock dips as Dreamforce outage tests confidence

Salesforce stock dips as Dreamforce outage tests confidence    ad-hoc-news.de

AI Incidents & Outages · ad-hoc-news.de · · open ↗ · share

AI agents are going rogue. CIOs are racing to put guardrails around them

AI agents are going rogue. CIOs are racing to put guardrails around them    Fortune

AI Incidents & Outages · Fortune · · open ↗ · share

Spain logs its first data breach allegedly carried out by a rogue AI agent

Spain logs its first data breach allegedly carried out by a rogue AI agent    Olive Press News Spain

AI Incidents & Outages · Olive Press News Spain · · open ↗ · share

Another Rogue AI Agent? Test Of Alibaba's Qwen Goes Off-Script

Another Rogue AI Agent? Test Of Alibaba's Qwen Goes Off-Script    Forbes

AI Incidents & Outages · Forbes · · open ↗ · share

Salesforce Down Today, Global Outage Hits Logins and APIs During Dreamforce

Salesforce Down Today, Global Outage Hits Logins and APIs During Dreamforce    Pasquale Pillitteri

AI Incidents & Outages · Pasquale Pillitteri · · open ↗ · share

How to catch and kill a rogue agent

How to catch and kill a rogue agent    IT Brew

AI Incidents & Outages · IT Brew · · open ↗ · share

Salesforce Outage Hits Customers Worldwide, CRM Stock Falls

Salesforce Outage Hits Customers Worldwide, CRM Stock Falls    CryptoRank

AI Incidents & Outages · CryptoRank · · open ↗ · share

Salesforce global outage hits during Dreamforce conference

Salesforce global outage hits during Dreamforce conference    tech.yahoo.com

AI Incidents & Outages · tech.yahoo.com · · open ↗ · share

Salesforce suffers global outage amid Dreamforce shindig

Salesforce suffers global outage amid Dreamforce shindig    The Register

AI Incidents & Outages · The Register · · open ↗ · share

How to Test AI Agent Output Guardrails Before Shipping to Production

How to Test AI Agent Output Guardrails Before Shipping to Production    Startup Fortune

AI Incidents & Outages · Startup Fortune · · open ↗ · share

Is ChatGPT down? Why is ChatGPT not working? Chatgpt down?

Is ChatGPT down? Why is ChatGPT not working? Chatgpt down?    Asbury Park Press

AI Incidents & Outages · Asbury Park Press · · open ↗ · share

AIUC Wants To Insure Your AI Agents Before They Go Rogue

AIUC Wants To Insure Your AI Agents Before They Go Rogue    Startup Fortune

AI Incidents & Outages · Startup Fortune · · open ↗ · share

The Triple AI Outage Is A Wake-Up Call For Enterprises

The Triple AI Outage Is A Wake-Up Call For Enterprises    Forrester

AI Incidents & Outages · Forrester · · open ↗ · share

Early Anthropic hire, former METR COO have found a way to rein in rogue AI agents

Early Anthropic hire, former METR COO have found a way to rein in rogue AI agents    TechCrunch

AI Incidents & Outages · TechCrunch · · open ↗ · share

OpenAI launches a new framework to track and investigate rogue AI agents

OpenAI launches a new framework to track and investigate rogue AI agents    Business Insider

AI Incidents & Outages · Business Insider · · open ↗ · share

This site is operated by software. Items link to their sources and are never re-hosted; fact-checks are by the named publishers; selection and ranking are automated. JSON.

[O_O] ^ top