Salesforce’s massive outage exposes the hidden risks of cloud dependencies
Salesforce’s massive outage exposes the hidden risks of cloud dependencies cio.com
Incidents where an AI system caused a technical failure with real consequences — sourced, dated, categorised — plus the AI providers' own outage notices, counted as they publish them. Companies and products are named; people are not.
| provider | 7 days | 30 days | latest notice |
|---|---|---|---|
| OpenAI | 14 | 24 | Elevated errors in ChatGPT Work 2026-09-16 19:34 UTC |
| Anthropic | 5 | 19 | Issues with Google Play subscriptions 2026-09-16 16:51 UTC |
| Cursor | 4 | 16 | Investigating service degradation — Grok Bot 2026-09-16 20:57 UTC |
| GitHub | 2 | 9 | Degradation with Gemini 3.8 Flash 2026-09-16 17:48 UTC |
| Perplexity | 1 | 4 | Maintenance: Status page migration to incident.io 2026-09-10 21:36 UTC |
| Replit | 0 | 0 | Agent turns aborting for some users 2026-08-01 18:45 UTC |
Counts are incident notices (any severity) on each provider's public status page, read every 10 minutes and updated here in place. A quiet page can also mean an under-reporting page.
42 incidents · 2023: 3 · 2024: 8 · 2025: 18 · 2026: 13 · sources last checked 2026-09-16
This table is not a feed. A row is added only when a postmortem, a company statement, a court document, a vendor advisory or top-tier reporting says so, which is why the newest row is usually older than today. Anything fresher is in the outage counts above and in the news below.
| date | incident | what happened | source |
|---|---|---|---|
| 2026-09-03 | ChatGPT, Claude and Grok went down within the same hour, each for its own reason OpenAI, Anthropic, xAI · ChatGPT and Codex; Claude Mythos/Fable 5.1 and 5, Opus 5, 4.8 and 4.6; Grok AI service outage | On the morning of September 3, 2026 (US time) the three largest model providers posted incidents within about an hour of each other. Anthropic's status page opened 'Elevated errors for multiple models' at 13:26 UTC, listing Mythos/Fable 5.1 and 5 and Opus 5, 4.8 and 4.6, and declared the impact over at 16:16 UTC. xAI began investigating Grok at about 13:30 UTC. OpenAI opened 'Elevated errors across ChatGPT and Codex' at 14:43 UTC, applied a mitigation at 15:17 UTC and resolved it at 16:55 UTC. The status pages of AWS, Google Cloud, Azure and Cloudflare showed nothing relevant. Impact: About three hours of elevated errors on five Claude models, a little over two hours on ChatGPT and Codex, and Grok disrupted in the same window. Root cause: Three different ones: OpenAI cited a routing error, Anthropic an infrastructure issue, xAI an outage at its Memphis compute site. No shared cause was found. Times are quoted from the providers' own incident notices; the per-company causes are as reported by The Register once the incidents closed. | The Register second source |
| 2026-04-23 | Anthropic postmortem: three shipped changes degraded Claude Code for weeks Anthropic · Claude Code (Sonnet 4.6, Opus 4.6, Opus 4.7) AI service outage | After weeks of reports that Claude Code felt less capable, forgetful and repetitive while burning usage limits faster, Anthropic's postmortem identified three overlapping causes: a default reasoning-effort change from high to medium (Mar 4-Apr 7), a prompt-cache clearing bug that dropped reasoning history every turn in stale sessions (Mar 26-Apr 10, fixed in v2.1.101), and a system-prompt verbosity constraint that cut evaluation scores about 3% (Apr 16-20). Each was reverted once identified. Impact: Anthropic reset usage limits for all subscribers and committed to soak periods, gradual rollouts, tighter system-prompt controls and broader internal testing on public builds. Root cause: Three independent configuration and prompt changes shipped without evaluations sensitive enough to catch the regressions. Quality degradation rather than downtime; included because a formal postmortem attributes the impact. | Anthropic engineering postmortem |
| 2025-11-18 | Cloudflare outage: Bot Management ML feature file doubled in size and crashed proxies Cloudflare · Bot Management machine-learning model (feature configuration file) AI service outage | A ClickHouse permissions change made the query that generates the Bot Management ML model's feature file return duplicate rows, roughly doubling the file; the proxy's bot module enforced a hard limit of 200 features and the FL2 proxy panicked on the oversized file. Core CDN and security traffic returned HTTP 5xx errors from 11:20 to 17:06 UTC, and Turnstile, Workers KV, Access and the dashboard were affected. Impact: About 5 hours 46 minutes of widespread 5xx errors across Cloudflare's network; customers on the older proxy engine saw every request scored as a bot score of zero. Root cause: Duplicate rows from a database permission change inflated the ML feature file past a hardcoded limit, triggering an unhandled error in the proxy. Not a model-inference failure: the fault was in the configuration pipeline feeding the ML bot-scoring model, as attributed by Cloudflare's own postmortem. | Cloudflare postmortem |
| 2025-09-17 | Anthropic postmortem: three infrastructure bugs degraded Claude responses for weeks Anthropic · Claude (Sonnet 4, Opus 4/4.1, Haiku 3.5) serving infrastructure AI service outage | Between Aug 5 and Sep 4, 2025 three overlapping bugs degraded Claude's output quality: a routing error sent some Sonnet 4 requests to servers configured for the 1M-token context window (16% of Sonnet 4 requests at the worst hour on Aug 31), a TPU token-generation misconfiguration inserted stray Thai and Chinese characters and syntax errors, and an XLA:TPU approximate top-k miscompilation corrupted token selection. Anthropic estimated about 30% of Claude Code users saw at least one degraded response. Impact: Weeks of degraded answers on the first-party API and Claude Code, rollbacks between Sep 2 and Sep 12, and new continuous production quality monitoring and more sensitive evaluations. Root cause: A load-balancer routing misconfiguration, a TPU generation misconfiguration and a mixed-precision compiler bug in approximate top-k sampling. Date is the postmortem; impact window Aug 5-Sep 4, 2025. | Anthropic engineering postmortem |
| 2024-12-11 | OpenAI outage: new telemetry service overwhelmed Kubernetes control planes OpenAI · ChatGPT, OpenAI API and Sora (Kubernetes serving platform) AI service outage | A newly deployed telemetry configuration generated massive Kubernetes API load across OpenAI's largest clusters, overwhelming control planes and breaking DNS-based service discovery; DNS caching delayed the symptoms until the rollout had propagated widely. ChatGPT, the API and Sora were down or degraded from 3:16 PM to 7:38 PM PST. Impact: About 4 hours 22 minutes of outage across all major OpenAI products; recovery required scaling clusters down and blocking expensive API calls to regain control-plane access. Root cause: A telemetry service whose Kubernetes API cost scaled with cluster size saturated the control plane. | OpenAI status page postmortem |
Salesforce’s massive outage exposes the hidden risks of cloud dependencies cio.com
Salesforce Global Outage Hits Customers on Second Day of Dreamforce channelinsider.com
Salesforce stock dips as Dreamforce outage tests confidence ad-hoc-news.de
AI agents are going rogue. CIOs are racing to put guardrails around them Fortune
Spain logs its first data breach allegedly carried out by a rogue AI agent Olive Press News Spain
Another Rogue AI Agent? Test Of Alibaba's Qwen Goes Off-Script Forbes
Salesforce Down Today, Global Outage Hits Logins and APIs During Dreamforce Pasquale Pillitteri
How to catch and kill a rogue agent IT Brew
Salesforce Outage Hits Customers Worldwide, CRM Stock Falls CryptoRank
Salesforce global outage hits during Dreamforce conference tech.yahoo.com
Salesforce suffers global outage amid Dreamforce shindig The Register
How to Test AI Agent Output Guardrails Before Shipping to Production Startup Fortune
Is ChatGPT down? Why is ChatGPT not working? Chatgpt down? Asbury Park Press
AIUC Wants To Insure Your AI Agents Before They Go Rogue Startup Fortune
The Triple AI Outage Is A Wake-Up Call For Enterprises Forrester
Early Anthropic hire, former METR COO have found a way to rein in rogue AI agents TechCrunch
OpenAI launches a new framework to track and investigate rogue AI agents Business Insider
This site is operated by software. Items link to their sources and are never re-hosted; fact-checks are by the named publishers; selection and ranking are automated. JSON.