Skip to content

Claude Outage Sept 29: Build AI Workflows for Downtime

Claude's September 29 outage hit Claude.ai, Code, Cowork and the API. Learn how to design AI workflows that fall back safely when a model vendor goes down.

Claude Outage Sept 29: Build AI Workflows for Downtime
Claude Outage Sept 29: Build AI Workflows for Downtime

Claude was down for part of Tuesday morning. As a Claude Partner Network Member, we build for exactly that morning.

On September 29, 2026, Anthropic reported elevated errors across Claude.ai, Claude Code, Claude Cowork, and the Claude API. The status page puts the impact at about an hour. Teams that had built Claude into client-facing workflows found out quickly whether those workflows had a fallback.

This isn't a post about one vendor's bad day. Every model provider and every CRM platform has incidents. Resilience comes from how you design the workflow, not from which vendor you pick. Below: what happened, why regulated firms should care, and the four design rules we use so an AI outage turns into a slower queue instead of a failed transaction.

Quick Answer

The Claude outage on September 29, 2026 caused elevated errors across Claude.ai, Claude Code, Claude Cowork, and the Claude API from 14:00 to 14:59 UTC, according to Anthropic's status page. For COOs and CTOs at banks, RIAs, and insurers putting Claude or Agentforce into production, the lesson is to design for vendor downtime: keep the CRM as the system of record, route failed AI steps to a human queue, retry with backoff, and keep a documented kill switch. Vantage Point builds these controls into Claude and CRM implementations.

Key Takeaways (TL;DR)

  • What happened: Anthropic reported about an hour of elevated errors across Claude's apps and API on September 29, followed by a sign-in issue during recovery.
  • Why it matters: if an AI step is the only path through a workflow, a vendor outage becomes your outage.
  • The design rule: AI is a layer on top of the CRM, never the only path. Failed AI steps should land in a queue, not fail the transaction.
  • The control examiners ask about: bank supervisors have asked about kill switches and contingency plans for AI, according to Reuters.
  • Next step: run a one-hour "AI is off" drill before your next go-live.

What Happened in the Claude Outage on September 29, 2026?

Status as of September 30, 2026. Anthropic's incident report gives this timeline (all times UTC):

  • 14:21: Anthropic began investigating elevated error rates on Claude.ai (including desktop and mobile apps), Claude Code, and Claude Cowork. Users might see failed requests, errors loading or sending conversations, or be asked to sign in again. The notice said "retrying may succeed."
  • 14:28: Anthropic added that some requests to the Claude API were also returning errors.
  • 14:36: the original errors were mitigated. A second issue then blocked many users from signing in, including through single sign-on, and from starting new chats or Claude Code and Cowork sessions. Anthropic asked signed-in users not to sign out.
  • 14:59: most services recovered. Anthropic noted that some messages sent between 14:00 and 14:59 UTC may not have been saved.
  • 16:27: the incident was marked resolved, with impact from 14:00 to 14:59 UTC (7:00 to 7:59 a.m. Pacific).

In U.S. terms, that was roughly 10:00 to 11:00 a.m. Eastern on a Tuesday. TechRadar's live coverage reported more than 11,800 Downdetector reports as of 10:17 a.m. ET. The Next Web noted it came less than a week after an earlier incident. Anthropic's September 22 incident report lists elevated errors on several Claude models from 00:50 to 02:10 UTC.

Two details matter most for workflow design. First, the sign-in path failed during recovery, so staff who were logged out stayed locked out after the errors cleared. Second, some messages may not have been saved. Any work that existed only inside an AI conversation was at risk.

Why Does an AI Vendor Outage Matter for Regulated Firms?

For a consumer, a model outage is an inconvenience. For a bank, RIA, or insurer, it can mean a client request that never got logged, a service case with no owner, or an intake step that silently failed.

Regulators have noticed. In June, Reuters reported, citing sources familiar with the matter, that OCC and Federal Reserve examiners had begun asking banks detailed questions about AI use in routine examinations, including whether they have controls such as "kill switches" and contingency plans in case of failures. PYMNTS' summary described the concern as whether banks can shut the systems down if necessary. Reuters said supervisors were gathering information rather than writing new rules. The questions still tell you what a good answer looks like.

This isn't unique to Claude. Salesforce had its own disruption this month (see our September 2026 Salesforce outage guide), and every AI vendor publishes incident history. The right question isn't "which vendor never goes down?" It's "what does our workflow do when one does?"

What Does a Resilient AI Workflow Look Like?

We use four rules when we put Claude or Agentforce into production workflows.

1. The CRM stays the system of record

Salesforce or HubSpot holds the client record, the case, the task, and the audit trail. AI reads from and writes to the CRM, but nothing important lives only in an AI conversation. If the model is unavailable, the record still exists and a person can pick it up. That also covers the "messages may not have been saved" problem.

2. AI is a layer, never the only path

Every AI step should have a manual route that produces the same business outcome. An AI-drafted client reply has a human-drafted fallback. An AI-summarized intake has a standard form. If the manual route doesn't exist, the AI step isn't ready for client-facing work.

3. Failed AI steps fall back to a queue, not a failed transaction

When an AI call fails in a client-facing workflow, the transaction should continue: create the case, log the request, and route it to a human queue with a clear flag such as "AI step skipped." The client gets a normal acknowledgment. Your team works the queue until service returns.

4. Retries with backoff, plus a documented kill switch

Transient errors are normal. Anthropic's API error documentation advises retrying internal server errors with exponential backoff, and says its official SDKs retry transient failures twice by default. Cap retries so a long outage doesn't pile up duplicate requests, then fail over to the queue. Separately, document a kill switch: who can turn off AI steps, where the setting lives, and how the workflow behaves when it's off.

AI Workflow Resilience Checklist: Failure Mode, Control, Owner

Failure modeDesign controlOwner
Model API returns errors or times outCapped retries with exponential backoff, then route to a human queueIntegration / engineering lead
AI step fails mid-transactionTransaction completes in the CRM; AI step flagged "skipped" for follow-upCRM platform owner
Staff can't sign in to the AI app (SSO issue)Documented manual procedure that runs entirely in the CRMOperations manager
AI conversation content lostOutputs written back to the CRM record as each step completesCRM platform owner
Extended outage or bad model behaviorKill switch with a named decision-maker and a tested "AI off" modeCOO or CTO
Clients notice delaysPre-approved client messaging and service-level expectationsClient service lead + compliance
No record of what happenedIncident log: vendor status updates, affected workflows, queue volume, recovery timeRisk / compliance

How Should You Test AI Fallback Before the Next Outage?

Don't wait for a vendor to test it for you. Run a short drill in a sandbox or during low volume:

  • Flip the kill switch. Confirm the named owner can turn AI steps off without a code deployment.
  • Force failures. Simulate API errors and timeouts. Confirm retries stop at the cap and work lands in the queue.
  • Run a real transaction end to end. A client request should still create a record, get an acknowledgment, and reach a person.
  • Check the audit trail. Can you show which records skipped the AI step and who handled them?
  • Time the recovery. When AI comes back on, confirm the queue clears and nothing gets processed twice.

Write the results down. An examiner, auditor, or board member asking about AI contingency plans wants evidence that you tested the plan, not just that you have one.

What Should Teams Do This Week?

  1. Inventory AI steps in production. List every workflow where Claude, Agentforce, or another model sits in a client-facing path.
  2. Mark single points of failure. Flag any step where an AI error stops the transaction.
  3. Subscribe to vendor status pages. Route incident alerts to the people who own the affected workflows.
  4. Review your contracts. Check what your AI vendor commits to on availability and incident notification. Our guide to AI vendor incident notification clauses covers what to ask for. This isn't legal advice; review contract terms with counsel.
  5. Add the resilience checklist to vendor review. Pair it with the questions in our AI vendor review guide for regulated firms.

How Vantage Point Helps

Vantage Point is a Claude Partner Network Member and a Salesforce and HubSpot partner, so we design the AI layer and the CRM underneath it together. Our Claude AI consulting team builds fallback queues, capped retries, kill switches, and CRM write-back into Claude implementations from the start. Our managed services team helps keep those controls tested after go-live. We use the same approach with Agentforce, including for a premier insurance and financial services brokerage now running Agentforce Service Agents in production. We've completed 400+ engagements for 150+ clients, with a 4.71/5.0 average engagement rating and 95% client retention.

Is Your AI Workflow Ready for the Next Outage?

Every model will have a bad morning. Vantage Point can review your Claude or Agentforce workflows, find the single points of failure, and help you build and test a fallback plan. Talk to Vantage Point about AI workflow resilience.

Frequently Asked Questions

What happened in the Claude outage on September 29, 2026?

Anthropic reported elevated errors across Claude.ai, Claude Code, Claude Cowork, and the Claude API. Its status page lists impact from 14:00 to 14:59 UTC, followed by a sign-in issue during recovery, and the incident was marked resolved at 16:27 UTC.

Was the Claude API affected by the September 29 outage?

Yes. Anthropic's status page said at 14:28 UTC that some requests to the Claude API were returning errors, and the incident report lists the Claude API and Claude Console among the affected components.

How do you design an AI workflow to survive a vendor outage?

Keep the CRM as the system of record, give every AI step a manual path, route failed AI steps to a human queue instead of failing the transaction, retry with capped exponential backoff, and keep a documented kill switch with a named owner.

What is an AI kill switch?

An AI kill switch is a documented control that lets a named person turn off AI steps in a workflow quickly, without a code deployment, while the business process keeps running through a manual or queue-based path.

Are bank examiners asking about AI contingency plans?

According to a June 2026 Reuters report citing sources familiar with the matter, OCC and Federal Reserve examiners have asked banks about AI controls such as kill switches and contingency plans in case of failures. Reuters said the questions were aimed at understanding practices rather than setting new rules.

Should we switch AI vendors after an outage?

Not by default. Every AI vendor and CRM platform has incidents. Designing workflows that degrade safely usually matters more than switching providers, though multi-model fallback can make sense for some high-volume use cases.

How often should we test AI fallback procedures?

Test before each go-live, after any major workflow change, and on a regular schedule your risk team sets. Document each drill so you can show auditors and examiners the plan works.

Sources

David Cockrum

David Cockrum

David Cockrum is the founder and CEO of Vantage Point, a specialized Salesforce consultancy exclusively serving financial services organizations. As a former Chief Operating Officer in the financial services industry with over 13 years as a Salesforce user, David recognized the unique technology challenges facing banks, wealth management firms, insurers, and fintech companies—and created Vantage Point to bridge the gap between powerful CRM platforms and industry-specific needs. Under David’s leadership, Vantage Point has achieved over 150 clients, 400+ completed engagements, a 4.71/5 client satisfaction rating, and 95% client retention. His commitment to Ownership Mentality, Collaborative Partnership, Tenacious Execution, and Humble Confidence drives the company’s high-touch, results-oriented approach, delivering measurable improvements in operational efficiency, compliance, and client relationships. David’s previous experience includes founder and CEO of Cockrum Consulting, LLC, and consulting roles at Hitachi Consulting. He holds a B.B.A. from Southern Methodist University’s Cox School of Business.

Elements Image

Subscribe to our Blog

Get the latest articles and exclusive content delivered straight to your inbox. Join our community today—simply enter your email below!

Need help applying this to your CRM roadmap?

Talk to Vantage Point

Vantage Point helps regulated and growth-focused teams implement Salesforce, HubSpot, integrations, data migration, and managed services with practical, senior-led guidance.

Latest Articles

Claude Outage Sept 29: Build AI Workflows for Downtime

Claude Outage Sept 29: Build AI Workflows for Downtime

Claude's September 29 outage hit Claude.ai, Code, Cowork and the API. Learn how to design AI workflows that fall back safely when a model v...

Claude Sonnet 5.5: What Changed and How Regulated Teams Migrate

Claude Sonnet 5.5: What Changed and How Regulated Teams Migrate

Claude Sonnet 5.5 is here. Compare it with Opus, check pricing and safety changes, and plan a governed migration for sensitive CRM workflow...

AI Oversight Metrics: Anthropic Just Published the Board Template

AI Oversight Metrics: Anthropic Just Published the Board Template

Anthropic just published the AI oversight metrics boards will start asking AI vendors for. Here is how banks, RIAs, and insurers should use...