How do you recover data an automation missed while it was down?
Recover missed data by first measuring the gap, then rebuilding it from the source system, not from the automation's run history alone. Find the exact outage window, list every record created or changed in that window, compare against the destination, and push only the missing or stale records through a controlled backfill run.
Every automation breaks eventually. A token expires, an API changes a field name, a rate limit kicks in on a busy day, or someone renames a property in the CRM. The fix is usually quick. The hard part is the days of data that never moved while the automation was broken.
This is the playbook I use for backfills in Zapier, Make, n8n, and custom scripts. It is boring on purpose. Backfills done in a hurry are how one outage turns into two.
Why is a backfill riskier than the original outage?
A backfill is riskier because it runs many records at once, often through logic that already failed, into systems that kept changing during the outage. A careless backfill can create duplicates, overwrite newer data with older values, and trigger downstream emails or alerts that should never fire twice. The outage paused work. A bad backfill corrupts it.
Think about a form-to-CRM automation that broke for three days. During those days, sales reps may have added some of those leads by hand after seeing them in the inbox. If you replay all three days blindly, you get duplicate contacts in HubSpot, duplicate deals, and maybe a welcome email sent twice to the same buyer.
Now picture a sync between Airtable and a website CMS. Editors kept updating records during the outage. If the backfill pushes the old snapshot from the failed runs, it overwrites their newer edits. Nobody notices until a client points out a wrong page.
I wrote about the duplicate side of this problem in what to do when an automation runs twice. A backfill is that same risk, multiplied by every record in the gap.
How do you measure the exact gap?
Measure the gap by finding the last successful run and the first successful run after the fix, then listing every source record created or updated between those two timestamps. Use the source system's own timestamps, not the automation tool's, because the tool only knows about events it actually received. Missed events never show up in its history.
This is the step most people skip. They open the run history in Zapier or Make, see a list of errors, and replay those. But if the trigger itself was broken, there may be no errored runs at all. A webhook that never fired leaves no trace in the tool that was waiting for it.
So start at the source. Export the form submissions, CRM records, or database rows from the outage window. Then export the matching records from the destination. A simple comparison on a stable key, like email address or record ID, tells you exactly what is missing and what is stale.
Write the count down before you touch anything. If the gap is forty records, you will know a backfill that creates four hundred went wrong.
When is replaying runs in Zapier or Make the right move?
Replaying runs is the right move when the trigger worked, the failure happened in a later step, and the data has not changed since. Both Zapier and Make let you rerun failed work, but each replay follows specific rules about what reruns, what counts toward usage, and what must be true about the workflow first.
In Zapier, the help center separates replaying errored steps from replaying an entire run. Zapier says that replaying an entire Zap counts as a new Zap run, and that successful steps count toward your task usage again even if they already counted before. That matters for a large backfill on a tight plan.
Zapier also says you cannot replay a Zap if you changed its trigger app or event, that an edited Zap must be published before a replay uses the new version, and that the Zap must be turned on. On the Free plan, Zapier limits manual replay to runs with an errored status. Check Zapier's current docs for your plan before you count on replay.
In Make, failed runs can be stored as incomplete executions. Make's help center says you can retry incomplete executions only when the scenario is active, and that if a retry fails on a different module, Make creates a new incomplete execution starting from that module. For more on how failures route inside a scenario, see my post on error routes in a Make scenario.
When should you rebuild from the source instead?
Rebuild from the source when the trigger itself failed, when records changed during the outage, or when the fix changed the workflow's logic. In those cases, run history is either missing or wrong. A separate backfill run that reads current source data and writes through the fixed logic is safer than replaying stale payloads.
I usually build a small, separate backfill workflow rather than abusing the live one. It reads the list of missing record IDs from a sheet or a table, looks up each record fresh from the source, and writes it to the destination using the same field mapping as production. It has its own name, so nobody confuses its runs with normal traffic.
Two rules make this safe. First, every write is an upsert keyed on a stable ID, never a blind create. Second, the backfill skips any destination record that was updated after the outage started. That one check protects the edits people made by hand while the automation was down.
How do you stop a backfill from firing side effects?
Stop side effects by switching off or filtering every downstream action that should only happen once, like welcome emails, Slack alerts, owner assignment, and sequence enrollment. Mark backfilled records with a flag or a source value, and make downstream workflows ignore that flag. Then turn everything back on after the backfill finishes and you check the counts.
This is where most painful backfills go wrong. The data lands correctly, but HubSpot workflows see a burst of new contacts and fire every enrollment trigger at once. Reps get forty assignment notifications in a minute. Old leads get a "thanks for reaching out" email five days late.
A simple property such as "Backfill batch" with today's date solves most of it. Downstream workflows add one filter: skip records where that property is set. After the backfill, you decide case by case which records deserve a human follow-up instead of an automated one. A late lead deserves a personal note, not a template.
How do you check the backfill worked?
Check the backfill by comparing counts and spot-checking records. The number of records written should match the gap you measured, give or take the ones you deliberately skipped. Then open ten records at random in the destination and confirm every mapped field matches the source. Only then turn the downstream automations back on.
Log the result somewhere permanent. I keep a short incident note for every client automation: what broke, the outage window, the gap count, how it was backfilled, and what was skipped and why. If you already log every run, as I describe in how to log every run of an AI agent workflow, the incident note links straight to the relevant runs.
That note pays off later. The next time something breaks, you have a tested recipe instead of a blank page. It also gives the client a clear, honest record of what happened, which builds more trust than pretending nothing broke.
How do you make the next outage easier to recover from?
Make the next outage easier by designing for replay before anything breaks. Use stable IDs and upserts, store the raw trigger payload, alert on silence as well as on errors, and keep a backfill workflow ready to run. An automation that can be safely rerun is worth more than one that never fails on paper.
Alerting on silence is the one I push hardest. Most tools alert you when a step errors. Very few alert you when a trigger that usually fires twenty times a day has fired zero times since lunch. A daily count check catches broken webhooks and expired connections that error alerts miss.
This is how I think about the automations I run in production, including the Airtable and WhaleSync setup behind Ajust and the HubSpot and Zapier workflows behind Kismet Health. The goal is not zero failures. The goal is that every failure has a short, known path back to clean data.
What should you do next?
Pick your most important automation today and answer three questions: what is its stable record ID, where would you get the source data for a backfill, and which downstream actions must never fire twice. If you cannot answer all three in five minutes, fix that before the next outage, not during it.
Then add one silence alert for that automation and write a one-paragraph backfill plan in its documentation. That small investment turns a stressful week into a calm afternoon when something breaks.
If you are staring at a gap right now and are not sure whether to replay or rebuild, reach out. I am happy to help you plan the backfill before anything gets overwritten.
Get found, cited and the back office automated
Let's make your site the source AI engines quote and wire up the systems behind it.
Read more blogs
Let's get your website found and cited by AI
Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.