8 days until Shopify stops accepting app script tags. Is your store running one? Check your store free
Get started free
← Blog

How to run an outage fire drill for your store

An outage fire drill is a scheduled, 60-minute rehearsal where you simulate one vendor failure, confirm your alerts actually reach a human, and time how long it takes your team to post a customer-facing message and switch to a workaround. Run it once a quarter with one scenario, one timer, and one written debrief, and you will find the broken alert channel or the missing vendor login while nothing is on fire.

Most stores discover their alerting gaps during a real incident: the Slack channel was muted, the SMS went to a phone on Do Not Disturb, or the only person with the shipping vendor's admin password was on a flight. A drill surfaces all of that in an hour you chose.

What does a fire drill actually include?

Four parts, in this order:

  1. Alert path test. Prove that a real alert arrives on every channel, on every device, for every person on call.
  2. Scenario injection. Simulate one dependency failing, either technically or as a tabletop exercise.
  3. Timed response. Run your runbook with a stopwatch and write down the clock times.
  4. Debrief and three fixes. Fifteen minutes, three named owners, three dates.

Anything bigger than this gets skipped next quarter. Keep it to an hour.

Which scenario should you drill first?

Pick the dependency that has actually cost you minutes recently, not the scariest hypothetical. Shipping and payments infrastructure produced the heaviest concentrations of downtime in the current window:

Service90-day uptimeIncidentsTotal downtimeAvg incident
Plaid98.8%41,548 min387 min
Shippo98.96%101,344 min134 min
BigCommerce99.86%2180 min90 min
Shopify99.9%2124 min62 min
Zapier99.91%2118 min59 min

These figures come from StatusBird's independent 2-minute monitoring over the 90 days ending September 22, 2026. Two things follow for your drill design. First, frequency and length are different problems: Shippo logged 10 incidents averaging 134 minutes, so a label-printing drill should rehearse a fast, repeatable workaround. Plaid logged 4 incidents averaging 387 minutes, so a bank-linking drill should rehearse a multi-hour posture: customer messaging, support macros, and a decision about whether to keep the flow visible at all.

Second, set the drill clock to a realistic length. A 15-minute rehearsal teaches nothing about an incident that historically runs over six hours. Run the timed portion for 45 minutes and make the team state what they would do at the 2-hour and 6-hour marks.

How do you test the alert path before the drill?

Do this the day before, not during the drill, so the drill measures response rather than plumbing. Send a test alert to each channel and check each of these settings on each device:

  • SMS: on iOS, open the alert sender in Messages, tap the contact, and turn on Emergency Bypass so alerts ring through a Focus or Do Not Disturb schedule. On Android, mark the contact as a starred contact and allow starred contacts under Do Not Disturb exceptions.
  • Email: confirm the alert did not land in Gmail's Promotions or Updates tab. Create a filter on the sender address with "Never send it to Spam" and a label that triggers a mobile notification.
  • Slack: open the alerts channel, choose Notification preferences, and set it to All new messages for both desktop and mobile. Check that mobile push is not set to "Nothing" outside working hours.
  • Microsoft Teams: set the channel to Banner and feed, and verify the connector still posts. Teams connectors can silently stop after a URL rotation.
  • Discord: set the channel notification override to All Messages and confirm the server is not muted on the phone app.

Write down who confirmed receipt on which channel. If anyone does not confirm within 10 minutes, that is your first fix, before you drill anything.

How do you inject the failure without breaking your live store?

Match the injection method to the blast radius. Never simulate a payment failure on live checkout.

Safe technical injection

  • Shipping or label API: in the vendor's dashboard, rotate the API key used by a sandbox or test account, or point your ops workstation's hosts file at a dead address for the vendor domain. The team then experiences real error screens while customers see nothing.
  • Storefront app block: duplicate your live theme, remove or break the app block in the copy, and share the preview link. The team practices the "hide the broken section" edit on a theme that is not published.
  • Automation platform: turn off one Zapier or Make scenario and let the order backlog build for 30 minutes. This is the closest thing to a real partial outage, and it rehearses the replay-and-reconcile step honestly.

Tabletop injection

For payments, tax calculation, and your own storefront, use a tabletop. One person reads the scenario aloud with a fake timestamp ("at 10:14 your checkout returns a gateway error on 100% of card attempts"), and the team performs every action that does not touch customers: opening the vendor status page, drafting the banner copy, writing the support macro, and pulling up the ad platform to pause campaigns. Everything is real except the fault.

Two housekeeping rules. Announce the drill internally and prefix every Slack or Teams message with DRILL so nobody opens a real vendor ticket. And pick a low-traffic hour from your own analytics, not a generic "early morning," because your traffic curve is yours.

What should you measure?

Record six clock times on a shared doc. These are the numbers that improve between drills:

  • Alert to acknowledgement: from alert delivery to a human typing "I've got it."
  • Acknowledgement to attribution: how long to prove it is the vendor and not you, using a live status check and the vendor's own status page.
  • Attribution to customer message: banner published, or support macro enabled.
  • Attribution to workaround: manual labels being printed, alternate gateway enabled, ads paused.
  • Number of blocked steps: every point where someone lacked a login, a permission, or a phone number.
  • Recovery start: who begins the reconciliation list, and where it lives.

Set your own targets before you start rather than judging the result afterward. A common shape is: acknowledge fast, attribute within a few minutes using a second source, and publish something customer-facing before you finish diagnosing. Because StatusBird checks every 2 minutes, detection is a small part of your total clock, which means most of your improvement will come from the attribution and messaging steps. Our approach to check intervals and incident classification is described on the methodology page.

How do you debrief so the drill produces change?

Fifteen minutes, three questions, three fixes. Ask: what did we do from memory instead of from the runbook, what credential or permission was missing, and what did the customer see that we would not want them to see. Then pick exactly three fixes with an owner and a date. Typical outputs look like this: "add the shipping vendor's support phone number and account ID to the runbook, owner Maya, Friday," or "give the support lead Theme editor permission so she can publish the banner without waiting, owner Dan, Thursday."

Store the doc next to your runbook and open it at the start of the next drill. Reading last quarter's three fixes aloud takes two minutes and tells you immediately whether the process is working.

How often should you run one, and what should you rotate?

Quarterly is enough for most stores, with one extra drill four to six weeks before peak season. Rotate the scenario each time so you cover at least four categories per year: payments, shipping or fulfillment, marketing and email, and your own storefront or platform. Vendor risk moves around: the monthly totals in the outage index show 16 major or critical incidents and 40.4 combined downtime hours across 9 services in August 2026, against 20 incidents and 122.1 hours across 11 services in April 2026, so the category that hurt you last quarter may not be the one that hurts you next.

Two timing rules. Do not drill during a change freeze, since half your fixes will be things you are not allowed to deploy. And do not drill the week of a launch or a major sale, because the drill will be cancelled and you will not reschedule it.

If you need a second, independent signal to make the attribution step fast during a drill or a real incident, you can check a service's live status or set up alerts on the vendors your store depends on at statusbird.io/register.

Never find out about an outage from your customers

StatusBird monitors Stripe, Klaviyo, Google Ads, Shopify, and 80+ other services your store depends on. Get an SMS alert within minutes of any outage.

Start monitoring free