
Claude was down for part of Tuesday morning. As a Claude Partner Network Member, we build for exactly that morning.
On September 29, 2026, Anthropic reported elevated errors across Claude.ai, Claude Code, Claude Cowork, and the Claude API. The status page puts the impact at about an hour. Teams that had built Claude into client-facing workflows found out quickly whether those workflows had a fallback.
This isn't a post about one vendor's bad day. Every model provider and every CRM platform has incidents. Resilience comes from how you design the workflow, not from which vendor you pick. Below: what happened, why regulated firms should care, and the four design rules we use so an AI outage turns into a slower queue instead of a failed transaction.
Quick Answer
The Claude outage on September 29, 2026 caused elevated errors across Claude.ai, Claude Code, Claude Cowork, and the Claude API from 14:00 to 14:59 UTC, according to Anthropic's status page. For COOs and CTOs at banks, RIAs, and insurers putting Claude or Agentforce into production, the lesson is to design for vendor downtime: keep the CRM as the system of record, route failed AI steps to a human queue, retry with backoff, and keep a documented kill switch. Vantage Point builds these controls into Claude and CRM implementations.
Key Takeaways (TL;DR)
- What happened: Anthropic reported about an hour of elevated errors across Claude's apps and API on September 29, followed by a sign-in issue during recovery.
- Why it matters: if an AI step is the only path through a workflow, a vendor outage becomes your outage.
- The design rule: AI is a layer on top of the CRM, never the only path. Failed AI steps should land in a queue, not fail the transaction.
- The control examiners ask about: bank supervisors have asked about kill switches and contingency plans for AI, according to Reuters.
- Next step: run a one-hour "AI is off" drill before your next go-live.
What Happened in the Claude Outage on September 29, 2026?
Status as of September 30, 2026. Anthropic's incident report gives this timeline (all times UTC):
- 14:21: Anthropic began investigating elevated error rates on Claude.ai (including desktop and mobile apps), Claude Code, and Claude Cowork. Users might see failed requests, errors loading or sending conversations, or be asked to sign in again. The notice said "retrying may succeed."
- 14:28: Anthropic added that some requests to the Claude API were also returning errors.
- 14:36: the original errors were mitigated. A second issue then blocked many users from signing in, including through single sign-on, and from starting new chats or Claude Code and Cowork sessions. Anthropic asked signed-in users not to sign out.
- 14:59: most services recovered. Anthropic noted that some messages sent between 14:00 and 14:59 UTC may not have been saved.
- 16:27: the incident was marked resolved, with impact from 14:00 to 14:59 UTC (7:00 to 7:59 a.m. Pacific).
In U.S. terms, that was roughly 10:00 to 11:00 a.m. Eastern on a Tuesday. TechRadar's live coverage reported more than 11,800 Downdetector reports as of 10:17 a.m. ET. The Next Web noted it came less than a week after an earlier incident. Anthropic's September 22 incident report lists elevated errors on several Claude models from 00:50 to 02:10 UTC.
Two details matter most for workflow design. First, the sign-in path failed during recovery, so staff who were logged out stayed locked out after the errors cleared. Second, some messages may not have been saved. Any work that existed only inside an AI conversation was at risk.
Why Does an AI Vendor Outage Matter for Regulated Firms?
For a consumer, a model outage is an inconvenience. For a bank, RIA, or insurer, it can mean a client request that never got logged, a service case with no owner, or an intake step that silently failed.
Regulators have noticed. In June, Reuters reported, citing sources familiar with the matter, that OCC and Federal Reserve examiners had begun asking banks detailed questions about AI use in routine examinations, including whether they have controls such as "kill switches" and contingency plans in case of failures. PYMNTS' summary described the concern as whether banks can shut the systems down if necessary. Reuters said supervisors were gathering information rather than writing new rules. The questions still tell you what a good answer looks like.
This isn't unique to Claude. Salesforce had its own disruption this month (see our September 2026 Salesforce outage guide), and every AI vendor publishes incident history. The right question isn't "which vendor never goes down?" It's "what does our workflow do when one does?"
What Does a Resilient AI Workflow Look Like?
We use four rules when we put Claude or Agentforce into production workflows.
1. The CRM stays the system of record
Salesforce or HubSpot holds the client record, the case, the task, and the audit trail. AI reads from and writes to the CRM, but nothing important lives only in an AI conversation. If the model is unavailable, the record still exists and a person can pick it up. That also covers the "messages may not have been saved" problem.
2. AI is a layer, never the only path
Every AI step should have a manual route that produces the same business outcome. An AI-drafted client reply has a human-drafted fallback. An AI-summarized intake has a standard form. If the manual route doesn't exist, the AI step isn't ready for client-facing work.
3. Failed AI steps fall back to a queue, not a failed transaction
When an AI call fails in a client-facing workflow, the transaction should continue: create the case, log the request, and route it to a human queue with a clear flag such as "AI step skipped." The client gets a normal acknowledgment. Your team works the queue until service returns.
4. Retries with backoff, plus a documented kill switch
Transient errors are normal. Anthropic's API error documentation advises retrying internal server errors with exponential backoff, and says its official SDKs retry transient failures twice by default. Cap retries so a long outage doesn't pile up duplicate requests, then fail over to the queue. Separately, document a kill switch: who can turn off AI steps, where the setting lives, and how the workflow behaves when it's off.
AI Workflow Resilience Checklist: Failure Mode, Control, Owner
| Failure mode | Design control | Owner |
|---|---|---|
| Model API returns errors or times out | Capped retries with exponential backoff, then route to a human queue | Integration / engineering lead |
| AI step fails mid-transaction | Transaction completes in the CRM; AI step flagged "skipped" for follow-up | CRM platform owner |
| Staff can't sign in to the AI app (SSO issue) | Documented manual procedure that runs entirely in the CRM | Operations manager |
| AI conversation content lost | Outputs written back to the CRM record as each step completes | CRM platform owner |
| Extended outage or bad model behavior | Kill switch with a named decision-maker and a tested "AI off" mode | COO or CTO |
| Clients notice delays | Pre-approved client messaging and service-level expectations | Client service lead + compliance |
| No record of what happened | Incident log: vendor status updates, affected workflows, queue volume, recovery time | Risk / compliance |
How Should You Test AI Fallback Before the Next Outage?
Don't wait for a vendor to test it for you. Run a short drill in a sandbox or during low volume:
- Flip the kill switch. Confirm the named owner can turn AI steps off without a code deployment.
- Force failures. Simulate API errors and timeouts. Confirm retries stop at the cap and work lands in the queue.
- Run a real transaction end to end. A client request should still create a record, get an acknowledgment, and reach a person.
- Check the audit trail. Can you show which records skipped the AI step and who handled them?
- Time the recovery. When AI comes back on, confirm the queue clears and nothing gets processed twice.
Write the results down. An examiner, auditor, or board member asking about AI contingency plans wants evidence that you tested the plan, not just that you have one.
What Should Teams Do This Week?
- Inventory AI steps in production. List every workflow where Claude, Agentforce, or another model sits in a client-facing path.
- Mark single points of failure. Flag any step where an AI error stops the transaction.
- Subscribe to vendor status pages. Route incident alerts to the people who own the affected workflows.
- Review your contracts. Check what your AI vendor commits to on availability and incident notification. Our guide to AI vendor incident notification clauses covers what to ask for. This isn't legal advice; review contract terms with counsel.
- Add the resilience checklist to vendor review. Pair it with the questions in our AI vendor review guide for regulated firms.
How Vantage Point Helps
Vantage Point is a Claude Partner Network Member and a Salesforce and HubSpot partner, so we design the AI layer and the CRM underneath it together. Our Claude AI consulting team builds fallback queues, capped retries, kill switches, and CRM write-back into Claude implementations from the start. Our managed services team helps keep those controls tested after go-live. We use the same approach with Agentforce, including for a premier insurance and financial services brokerage now running Agentforce Service Agents in production. We've completed 400+ engagements for 150+ clients, with a 4.71/5.0 average engagement rating and 95% client retention.
Is Your AI Workflow Ready for the Next Outage?
Every model will have a bad morning. Vantage Point can review your Claude or Agentforce workflows, find the single points of failure, and help you build and test a fallback plan. Talk to Vantage Point about AI workflow resilience.
Frequently Asked Questions
What happened in the Claude outage on September 29, 2026?
Anthropic reported elevated errors across Claude.ai, Claude Code, Claude Cowork, and the Claude API. Its status page lists impact from 14:00 to 14:59 UTC, followed by a sign-in issue during recovery, and the incident was marked resolved at 16:27 UTC.
Was the Claude API affected by the September 29 outage?
Yes. Anthropic's status page said at 14:28 UTC that some requests to the Claude API were returning errors, and the incident report lists the Claude API and Claude Console among the affected components.
How do you design an AI workflow to survive a vendor outage?
Keep the CRM as the system of record, give every AI step a manual path, route failed AI steps to a human queue instead of failing the transaction, retry with capped exponential backoff, and keep a documented kill switch with a named owner.
What is an AI kill switch?
An AI kill switch is a documented control that lets a named person turn off AI steps in a workflow quickly, without a code deployment, while the business process keeps running through a manual or queue-based path.
Are bank examiners asking about AI contingency plans?
According to a June 2026 Reuters report citing sources familiar with the matter, OCC and Federal Reserve examiners have asked banks about AI controls such as kill switches and contingency plans in case of failures. Reuters said the questions were aimed at understanding practices rather than setting new rules.
Should we switch AI vendors after an outage?
Not by default. Every AI vendor and CRM platform has incidents. Designing workflows that degrade safely usually matters more than switching providers, though multi-model fallback can make sense for some high-volume use cases.
How often should we test AI fallback procedures?
Test before each go-live, after any major workflow change, and on a regular schedule your risk team sets. Document each drill so you can show auditors and examiners the plan works.
Sources
- Claude Status: Elevated errors on claude.ai, Claude Code, Claude Cowork and the Claude API (Sept. 29, 2026)
- Claude Status: Elevated errors for multiple models (Sept. 22, 2026)
- The Next Web: Claude is down as Anthropic investigates errors across its services
- TechRadar: Claude down live coverage (Sept. 29, 2026)
- Reuters: U.S. bank regulators ramp up scrutiny of AI use at financial companies (June 12, 2026)
- PYMNTS: Bank Regulators Probe Industry Use of AI (June 12, 2026)
- Claude API documentation: Errors
