Should an automation retry, or stop and ask you?
It depends on whether the failure is temporary or wrong. A temporary failure, like a timeout or a rate limit, should be retried, because the same request will probably work in a minute. A wrong failure, like a missing required field or a rejected record, should stop and wait for a human. Retrying wrong never makes it right.
That single distinction has saved me more trouble than any monitoring tool. Most automation disasters I have seen were not caused by something failing. They were caused by something failing and then trying again, and again, on a request that was never going to succeed, or on a request that succeeded quietly the first time and got repeated.
The trouble is that the builder makes this decision once, usually while tired, usually at the end of the build, and usually by accepting whatever the platform does by default. Then it sits there for a year. So it is worth spending twenty minutes deciding it on purpose.
What actually makes a failure retryable?
A failure is retryable when nothing needs to change for the next attempt to work. The API was busy. The connection dropped. The rate limit was hit. The service was briefly down. In each case the request you sent was valid, the receiver simply could not deal with it then, and waiting is a real fix.
A failure is not retryable when the request itself is the problem. A contact record missing an email address will be missing it forever. A currency the receiving system does not accept will still be rejected on the fifth attempt. A permission your token does not have will not appear because you asked twice. These are not errors of timing, they are errors of content, and only a person can resolve them.
Most platforms hand you the signal for free. A 429 or a 500 from an API is almost always a timing problem. A 400 or a 422 is almost always a content problem. A 401 or a 403 is a credential problem, which is its own category, because retrying an expired token is pointless but refreshing it is not. If you do nothing else after reading this, sort your automation's failure handling into those three buckets.
The bucket that catches people out is the one I would call ambiguous. A webhook that returns nothing at all could mean the receiver crashed, or it could mean the receiver did the work and failed to reply. Those two outcomes look identical from your side and need completely different responses. That is the case where retrying blind is genuinely dangerous.
Why does retrying the wrong thing cost more than failing?
Because a stopped automation is visible and a repeating automation is not. When something halts, somebody notices within a day and fixes it. When something retries a write that already succeeded, you get duplicate records, duplicate emails and duplicate charges, and nobody notices until a customer complains or a report looks strange weeks later.
I have cleaned up both kinds of mess and the repeating kind is much worse. A halt costs you a delay. A silent duplicate costs you trust, because the person on the receiving end got two invoices or three welcome emails from a company that is meant to be organised. You cannot apologise your way back to where you were.
This is why retry policy and idempotency are the same conversation. If a write is safely repeatable, retrying it costs nothing but time. If it is not, every retry is a coin flip. I went through the mechanics of making writes safely repeatable in the piece on idempotency keys, and I would read that before turning aggressive retries on anywhere that creates records.
What does a platform give you by default?
Less than you think, and it varies. Zapier's Autoreplay is documented as available on Professional, Team and Enterprise plans, and not on Free, while manual replay of errored runs is documented as available on Free as well. n8n takes a different shape: you set an error workflow in Workflow Settings, and it runs if an execution fails.
Those are Zapier's and n8n's own words, from their own documentation, and they are worth reading rather than guessing at. Zapier's help also notes that Autoreplay can be switched on by an account super admin or owner to automatically replay any errored Zap runs across the entire account, which is a much bigger decision than it sounds like when you are looking at one Zap.
There is a billing detail on that page too that deserves attention. Zapier documents that when you replay an entire Zap, any successful steps will count towards your task usage even if they were already counted in a previous run. Retries are not free, and on a busy automation that difference shows up on an invoice rather than in a dashboard.
n8n gives you the opposite lever as well. Its documentation describes a Stop And Error node you can add to force executions to fail under your chosen circumstances and trigger the error workflow. That is the stop-and-ask behaviour made explicit. Being able to deliberately fail is underrated. An automation that knows how to give up cleanly is safer than one that keeps going into territory nobody designed for.
Make, Airtable automations and custom scripts each have their own arrangement, and the honest advice is to open the vendor's documentation for the platform you are actually on and read what it says today rather than trusting a blog post from anyone, including this one.
How do you decide the rule before you build?
Write one sentence for each step: what happens if this fails. Do it before you connect anything. For most steps the answer is retry a few times and then give up quietly. For a small number of steps the answer is stop immediately and tell a named person. Knowing which is which takes ten minutes and is the whole design.
The question I ask is simple. If this step fails and nobody ever looks at it, what is the worst outcome? If the worst outcome is that a report is a day late, the automation can retry and then go quiet. If the worst outcome is that a customer is billed twice or a deal is marked won when it is not, the automation stops, immediately, and somebody hears about it.
The automations I run for Ajust move case data through Airtable and WhaleSync, and that is the question I put to every step in them. Ajust has delivered more than 25,000 cases and helped more than 400,000 people, and at that volume a step you were casual about does not stay a small problem. Sorting steps by worst outcome is slow work that quietly pays for itself.
Where should a stop-and-ask automation send the person?
Somewhere a named human will actually see it, with enough context to act without opening the platform. That means the record involved, the step that failed, the error the system returned, and a link straight to the thing that needs fixing. A notification that says an automation failed and nothing else is not an escalation, it is a guilt trip.
The mistake I see most often is routing errors to a shared inbox or a busy channel where they become wallpaper. After the second week nobody reads them. If a failure is important enough to stop the automation, it is important enough to have an owner, and the owner should change when that person is on leave. I went into this properly in the article on who gets the alert overnight.
The other half of a good escalation is a path back. Once the human fixes the record, what restarts the work? If the answer is that they have to remember to go and press replay, you have built a system that depends on memory. Better to have the automation pick the record up again on its next scheduled run, so a fix is enough on its own.
How many retries is the right number?
Few, and spaced out. Three attempts over a widening gap handles almost every genuine outage that is going to resolve itself. Anything beyond that is usually hope rather than engineering, and it delays the moment a person finds out something is properly broken. The purpose of a retry is to survive a blip, not to outlast an incident.
Spacing matters more than the count. Three attempts in nine seconds all land inside the same outage and all fail. Three attempts spread across an hour give the other system a genuine chance to recover. If your platform lets you widen the gap between attempts, widen it, because tight loops also tend to make the rate limit problem you were retrying against worse.
There is also a point at which you should stop retrying entirely and switch to a scheduled catch-up. If the same job runs every hour anyway, a failure can simply be picked up by the next run rather than hammered at now. That pattern is easier to reason about and produces far less noise. The failed record just stays in the queue and disappears when it succeeds.
What does a retry policy cost you in money and trust?
Money, because most platforms count re-executed steps as usage, as Zapier documents for a full Zap replay. Trust, because every retry that lands on a system somebody else owns is a request they did not ask for. An aggressive automation is an inconsiderate guest on someone else's API, and eventually they start blocking guests.
The trust part gets skipped in most automation advice and I think it matters. When I connect HubSpot through Zapier for a client like Kismet Health, I am borrowing capacity from systems that other people depend on. A retry loop that fires a thousand times because a field name changed is not clever, it is rude, and it is the kind of thing that ends with an integration partner throttling you.
There is a quieter cost too. Teams learn to ignore systems that cry wolf. Retry quietly, escalate rarely, and the rare escalation gets taken seriously.
How do you test that the rule actually works?
Break it on purpose. Point a step at a bad credential and watch what happens. Send a record with a missing required field. Disconnect an account mid-run. If you have never watched your own automation fail, you do not know what it does when it fails, you only know what you intended it to do.
Do this in a place where the failure is harmless, which usually means a copy of the workflow writing into a sandbox table rather than the live one. This is the same instinct behind building a dry run mode into everything, which I argued for in a separate piece. You want a way to make the system misbehave without consequences.
Write down what you saw. Not in your head, in the same document that holds the automation's description, because the person who inherits this from you will need to know whether a silent failure is expected behaviour or a bug. That document is the difference between an automation somebody maintains and an automation somebody eventually deletes because nobody understands it.
What should you do next?
Pick your most important automation and answer two questions about every step that writes data. Is this failure temporary or wrong, and is this write safe to repeat? Fix the steps where the answers disagree with what the platform currently does. That is usually one or two steps, and it is an afternoon of work.
Then go and read your platform's own documentation on error handling rather than relying on what you remember. These settings change, plan tiers change, and the default you picked eighteen months ago may not be the default today. Ten minutes in the vendor's docs is cheaper than a week of cleaning up duplicates.
If you are staring at an automation that fails in ways you cannot predict, or you are not sure which of your steps are safe to retry, reach out. I spend a lot of my time on exactly this kind of unglamorous work, and it is usually a shorter conversation than people expect.
Get found, cited and the back office automated
Let's make your site the source AI engines quote and wire up the systems behind it.
Read more blogs
Let's get your website found and cited by AI
Tell me what you're working on, whether AI search is skipping your product, your back office is buried in manual work, or you need a build that does both.