Contingency checklist for WhatsApp API outages or failures (what to do, in what order, and how to communicate with the team)
Contingency checklist for WhatsApp API outages or failures
When the WhatsApp API stops responding, there is no time to improvise. Every minute that passes without a defined plan translates into customers without responses, accumulated messages, and a team that doesn't know if the problem is with Meta, the connection, or the platform they use. This checklist is designed so you can print it, copy it, or keep it handy in Slack, and so that each team member knows exactly what to do and in what order.
The golden rule: first you verify, then you communicate, and only then do you try to fix. Most errors in a contingency come from skipping steps: someone restarts a server when the problem is with Meta, or writes to support without having checked the service status. This checklist keeps you organized so that doesn't happen.
Before the outage: preparation (do it once, review it every 3 months)
The contingency starts before the outage. Without these three elements, the checklist below is useless: you won't know who to call, how to communicate, or whether the problem is yours or Meta's.
- Define who is the person responsible for the contingency (the "incident owner"). It must be someone with access to the platform, support, and the internal communication channel. If it's not defined, during an outage everyone steps on each other.
- Have the official WhatsApp API status at hand: https://metastatus.com/whatsapp-business-api. It's not just another link: it's the source that tells you if the problem is Meta's or yours. Save it in a fixed Slack channel or in browser favorites.
- Define the internal communication channel for emergencies (WhatsApp group, Slack channel, etc.) and agree that during an outage you write THERE and not in the general chat. If the channel isn't agreed upon, information gets scattered.
If you don't have these three points defined, stop here and resolve them before continuing. The checklist below assumes they exist.
Contingency checklist: what to do, in what order
- 1Confirm that the problem exists: send a test message to your own number (a colleague's or a test phone). If the message goes through, the problem may be with a specific contact or a specific number, not with the API in general. If it does not go through, proceed to step 2.
- 2Check the official status of the WhatsApp API at https://metastatus.com/whatsapp-business-api. If an outage or incident is listed, the problem is with Meta: do not restart anything, do not change settings. Go directly to step 4 (communicate).
- 3If the official status says everything is operational, the problem is yours or your platform's. Check the internet connection where the platform runs (not your phone's), and verify that the webhook is configured and responding. If you use a platform like Wando, check the inbox status: if messages come in but do not go out, the problem may be with the number configuration.
- 4Communicate to the team: as soon as you confirm the problem is real and global (not just one contact), notify the emergency channel. The notice must state three things: what is happening, since when (exact time), and what is being done. No "it's down" without context.
- 5Communicate to customers: if the outage affects service, decide whether to send a mass message (if the API responds for sending but not receiving) or post a notice on social media. If the API does not respond at all, do not send anything: a message that does not go out is worse than sending nothing.
- 6Escalate: if the problem is with Meta, open a ticket with Meta support (if you have access). If the problem is with your platform, contact their support. Have on hand: affected phone number, exact start time, error messages if any, and a screenshot of the official status if an incident is listed.
- 7Monitor and document: every 30 minutes, someone on the team checks the official status and updates the emergency channel. Note the time it was resolved and what happened, for the post-review.
Summary table: what to check, how to know it's fine
| Stage | What to check | Sign that it's fine |
|---|---|---|
| 1. Confirm | Test message to your own number | The message goes out and arrives. If it does not, continue. |
| 2. Official status | https://metastatus.com/whatsapp-business-api | Shows "Operational" or a reported incident. If nothing shows and the message does not go out, the problem is yours. |
| 3. Own infrastructure | Internet connection, webhook, number configuration | The connection responds, the webhook is active, the number is connected to the platform. |
| 4. Communicate to the team | Emergency channel | The notice states what is happening, since when, and what is being done. |
| 5. Communicate to customers | Mass message or social media notice | The message goes out (if the API responds for sending) or the notice is posted. |
| 6. Escalate | Ticket to Meta or the platform | The ticket is open with number, time, and error. |
| 7. Monitor | Official status + internal channel | Someone updates every 30 minutes. The resolution time is noted. |
What to do when something fails: edge cases
The checklist assumes everything goes well. But real contingencies have edge cases. Here are the most common ones and how to resolve them.
- The official status says "Operational" but messages do not go out: the problem is with your configuration or your platform. Do not keep trying the test message: check the webhook, the number connection, and the platform configuration. If you use a platform like Wando, contact their support with the exact time and the affected number.
- The API responds for sending but not receiving: customers write to you but you receive nothing. In this case, communicate with customers via a mass message (if the API allows sending) and let them know you are having trouble receiving. Do not restart anything: the problem is with receiving, not sending.
- The problem is with a single number: if one specific number does not respond but the rest do, it is not a global outage. Check if the number is registered, if it has an active session, and if it has not been banned. If you cannot verify, escalate to the platform or Meta.
- There is no internet where the platform runs: if the connection went down, it is not an API problem. Let the team know the issue is connectivity, and if you have a mobile data plan as backup, activate it. Do not restart the platform until the connection is stable.
- The team does not know who is responsible: if it is not defined, the first step of the contingency is to assign it. Do not continue with the checklist until someone has the role. Without a responsible person, each step is done twice or not at all.
How to communicate with the team: the message that works
Communication during a contingency is not a luxury: it is part of the plan. A poorly written notice generates more questions than answers and delays resolution. The format that works has three fixed parts, and they are non-negotiable.
- What is happening: "The WhatsApp API is not responding. Messages are not going out or coming in since 14:32."
- Since when: the exact time, not "a while ago". The exact time helps to know if the problem is getting worse or better.
- What is being done: "We checked the official status, there is an incident listed. We opened a ticket to Meta. Next update in 30 minutes."
If the notice does not have these three parts, do not send it. An incomplete notice generates noise and questions that no one has time to answer during a contingency.
After the outage: review and adjustment
When the API starts responding again, the work is not over. The post-mortem review is what turns an outage into a lesson. Gather the team (even if it's 15 minutes) and answer three questions:
- Which step of the checklist worked and which didn't? If any step was confusing or skipped, adjust it.
- How long did it take us to detect the outage? If it took more than 10 minutes, the monitoring needs improvement.
- Was the communication clear? If there were crossed messages or people who didn't know what to do, the channel or the format of the notice needs adjustment.
Update the checklist with what you learned and store it in an accessible place. The next outage will happen: the question is whether the team will be ready.
This checklist is a starting point. Each team has its own infrastructure and its own channels: adapt it to your reality. What is not negotiable is the order: verify, communicate, and only then try to fix.
Template (HSM) errors during an outage: how to live with the 24-hour rule
When the WhatsApp API goes down, a secondary problem that appears is that of templates (HSM). If you have a scheduled campaign and the outage prevents sending, remember that templates approved by Meta have a 24-hour window from when the customer starts the conversation. If the outage left you outside that window, don't try to send the template as if it were a normal message: it will bounce. The option is to wait for the customer to write again or reschedule the campaign. There is no shortcut: Meta does not relax this rule for outages.
What NOT to do during an outage (the three mistakes that prolong the problem)
Besides knowing what to do, it's worth knowing what to avoid. These three mistakes are the ones that most prolong a contingency, and all of them have a "well-intentioned" version behind them.
- Disconnect and reconnect the WhatsApp Business number: if the outage is from Meta, this fixes nothing and could cause a temporary block for suspicious activity.
- Changing the webhook or platform configuration without logging it: if it turns out the problem was from Meta, you won't know which change was yours and which wasn't.
- Deleting messages or conversations "to clean up": it doesn't help resolution and makes you lose evidence that support might request.
If any of this already happened, it's not the end of the world: write it down in the post-review and add it to the checklist as "what not to do".
Frequently asked questions
How do I know if the problem is with Meta or my platform?+
Check the official status at https://metastatus.com/whatsapp-business-api. If there is an incident, the problem is with Meta. If it says "Operational" and messages are not going out, the problem is with your configuration or your platform.
What do I do if the API responds for sending but not for receiving?+
Communicate with your customers via a mass message (if the API allows sending) letting them know you are having trouble receiving. Do not restart anything: the problem is with receiving, not sending.
When should I notify the team?+
As soon as you confirm the problem is real and global (not just one contact). The notice should state what is happening, since when, and what is being done.
What do I do if the problem is with a single number and not the rest?+
It is not a global outage. Check if the number is registered, if it has an active session, and if it has been banned. If you cannot verify, escalate to the platform or to Meta.
Can I send an approved template during an API outage?+
If the outage left you outside the 24-hour window since the customer's last message, no. The template will bounce. Wait for the customer to write or reschedule the campaign. Meta does not relax the rule for outages.
Does disconnecting and reconnecting the number fix anything?+
No, and it can make things worse. If the outage is with Meta, reconnecting changes nothing and may cause a temporary block due to suspicious activity. If the problem is with your platform, support will ask you not to touch the number.