Claude was down for part of Tuesday morning. As a Claude Partner Network Member, we build for exactly that morning.
On September 29, 2026, Anthropic reported elevated errors across Claude.ai, Claude Code, Claude Cowork, and the Claude API. The status page puts the impact at about an hour. Teams that had built Claude into client-facing workflows found out quickly whether those workflows had a fallback.
This isn't a post about one vendor's bad day. Every model provider and every CRM platform has incidents. Resilience comes from how you design the workflow, not from which vendor you pick. Below: what happened, why regulated firms should care, and the four design rules we use so an AI outage turns into a slower queue instead of a failed transaction.
The Claude outage on September 29, 2026 caused elevated errors across Claude.ai, Claude Code, Claude Cowork, and the Claude API from 14:00 to 14:59 UTC, according to Anthropic's status page. For COOs and CTOs at banks, RIAs, and insurers putting Claude or Agentforce into production, the lesson is to design for vendor downtime: keep the CRM as the system of record, route failed AI steps to a human queue, retry with backoff, and keep a documented kill switch. Vantage Point builds these controls into Claude and CRM implementations.
Status as of September 30, 2026. Anthropic's incident report gives this timeline (all times UTC):
In U.S. terms, that was roughly 10:00 to 11:00 a.m. Eastern on a Tuesday. TechRadar's live coverage reported more than 11,800 Downdetector reports as of 10:17 a.m. ET. The Next Web noted it came less than a week after an earlier incident. Anthropic's September 22 incident report lists elevated errors on several Claude models from 00:50 to 02:10 UTC.
Two details matter most for workflow design. First, the sign-in path failed during recovery, so staff who were logged out stayed locked out after the errors cleared. Second, some messages may not have been saved. Any work that existed only inside an AI conversation was at risk.
For a consumer, a model outage is an inconvenience. For a bank, RIA, or insurer, it can mean a client request that never got logged, a service case with no owner, or an intake step that silently failed.
Regulators have noticed. In June, Reuters reported, citing sources familiar with the matter, that OCC and Federal Reserve examiners had begun asking banks detailed questions about AI use in routine examinations, including whether they have controls such as "kill switches" and contingency plans in case of failures. PYMNTS' summary described the concern as whether banks can shut the systems down if necessary. Reuters said supervisors were gathering information rather than writing new rules. The questions still tell you what a good answer looks like.
This isn't unique to Claude. Salesforce had its own disruption this month (see our September 2026 Salesforce outage guide), and every AI vendor publishes incident history. The right question isn't "which vendor never goes down?" It's "what does our workflow do when one does?"
We use four rules when we put Claude or Agentforce into production workflows.
Salesforce or HubSpot holds the client record, the case, the task, and the audit trail. AI reads from and writes to the CRM, but nothing important lives only in an AI conversation. If the model is unavailable, the record still exists and a person can pick it up. That also covers the "messages may not have been saved" problem.
Every AI step should have a manual route that produces the same business outcome. An AI-drafted client reply has a human-drafted fallback. An AI-summarized intake has a standard form. If the manual route doesn't exist, the AI step isn't ready for client-facing work.
When an AI call fails in a client-facing workflow, the transaction should continue: create the case, log the request, and route it to a human queue with a clear flag such as "AI step skipped." The client gets a normal acknowledgment. Your team works the queue until service returns.
Transient errors are normal. Anthropic's API error documentation advises retrying internal server errors with exponential backoff, and says its official SDKs retry transient failures twice by default. Cap retries so a long outage doesn't pile up duplicate requests, then fail over to the queue. Separately, document a kill switch: who can turn off AI steps, where the setting lives, and how the workflow behaves when it's off.
| Failure mode | Design control | Owner |
|---|---|---|
| Model API returns errors or times out | Capped retries with exponential backoff, then route to a human queue | Integration / engineering lead |
| AI step fails mid-transaction | Transaction completes in the CRM; AI step flagged "skipped" for follow-up | CRM platform owner |
| Staff can't sign in to the AI app (SSO issue) | Documented manual procedure that runs entirely in the CRM | Operations manager |
| AI conversation content lost | Outputs written back to the CRM record as each step completes | CRM platform owner |
| Extended outage or bad model behavior | Kill switch with a named decision-maker and a tested "AI off" mode | COO or CTO |
| Clients notice delays | Pre-approved client messaging and service-level expectations | Client service lead + compliance |
| No record of what happened | Incident log: vendor status updates, affected workflows, queue volume, recovery time | Risk / compliance |
Don't wait for a vendor to test it for you. Run a short drill in a sandbox or during low volume:
Write the results down. An examiner, auditor, or board member asking about AI contingency plans wants evidence that you tested the plan, not just that you have one.
Vantage Point is a Claude Partner Network Member and a Salesforce and HubSpot partner, so we design the AI layer and the CRM underneath it together. Our Claude AI consulting team builds fallback queues, capped retries, kill switches, and CRM write-back into Claude implementations from the start. Our managed services team helps keep those controls tested after go-live. We use the same approach with Agentforce, including for a premier insurance and financial services brokerage now running Agentforce Service Agents in production. We've completed 400+ engagements for 150+ clients, with a 4.71/5.0 average engagement rating and 95% client retention.
Every model will have a bad morning. Vantage Point can review your Claude or Agentforce workflows, find the single points of failure, and help you build and test a fallback plan. Talk to Vantage Point about AI workflow resilience.
Anthropic reported elevated errors across Claude.ai, Claude Code, Claude Cowork, and the Claude API. Its status page lists impact from 14:00 to 14:59 UTC, followed by a sign-in issue during recovery, and the incident was marked resolved at 16:27 UTC.
Yes. Anthropic's status page said at 14:28 UTC that some requests to the Claude API were returning errors, and the incident report lists the Claude API and Claude Console among the affected components.
Keep the CRM as the system of record, give every AI step a manual path, route failed AI steps to a human queue instead of failing the transaction, retry with capped exponential backoff, and keep a documented kill switch with a named owner.
An AI kill switch is a documented control that lets a named person turn off AI steps in a workflow quickly, without a code deployment, while the business process keeps running through a manual or queue-based path.
According to a June 2026 Reuters report citing sources familiar with the matter, OCC and Federal Reserve examiners have asked banks about AI controls such as kill switches and contingency plans in case of failures. Reuters said the questions were aimed at understanding practices rather than setting new rules.
Not by default. Every AI vendor and CRM platform has incidents. Designing workflows that degrade safely usually matters more than switching providers, though multi-model fallback can make sense for some high-volume use cases.
Test before each go-live, after any major workflow change, and on a regular schedule your risk team sets. Document each drill so you can show auditors and examiners the plan works.