Retrying a failed email safely means backoff on 4xx, stop on 5xx, and never cloning a message that already returned 250. One event key per business event. A second button click is a no-op. A for-loop is bulk. Leftover MX is not a retry. History is the only receipt that matters.
Persist the key before the socket opens. On crash, read history before you AUTH again. Disable HTTP-library retries that do not know your invoice number. Isolate keys and SMTP users per client zone so Client A’s 250 cannot mark Client B sent. Inbound recovery replay is a different button with a 14-day Free store, 90 days on Solo through Agency, or 180 days on Unlimited. Do not mix those tickets. Free has no send-as, so a retry storm on Free is a 550 storm. Confirm MailerZ pricing. Caps exist so a bug cannot become a campaign. MailerZ is not a bulk ESP. We will not help design a list loop.
Quick answer for retry failed email safely
Read history. Classify. Replay stored inbound failures once from delivery recovery, or submit outbound through your app’s idempotent path. Also: troubleshooting, docs, and features. Transport still follows IETF RFC 5321 — Simple Mail Transfer Protocol. Competitor patterns only: ImprovMX blog.
Hourly caps punish loops. Free send-as, SMTP, and API are off. Solo is 2,500 outgoing per month and 20 per hour. Starter is 5,000 and 40. Business is 12,000 and 60. Agency is 20,000 and 60. Unlimited is 100,000 and 300. Those are operational ceilings, not a campaign budget. No inbox SLA. Not SOC 2.
People-first retry docs name the event key. A page that says “just click send again” is how customers get three invoices.
The user problem and the decision criteria
Operators hit retry until the customer stops yelling. The customer then has three copies. The useful system stores “this invoice already 250’d.” The useless system treats SMTP like a fidget spinner.
| Question | If yes | If no |
|---|---|---|
| History already 250? | Do not send again. | Read the error class. |
| 4xx? | Backoff. Same event key. | — |
| 5xx? | Stop. Fix store or From. | — |
| Loop over users? | Bulk. Stop. | Good. |
| Adding leftover MX? | Wrong ticket. | Good. |
Tape the decision map next to the send function. History 250 stops the worker. 4xx waits. 5xx fixes the cause. Empty history checks env and leftover MX. Smash-click is not a state. One key. One owner. One 250. Isolate keys per client. Delete probe scripts. Do not log passwords. Search the destination before you clone a receipt. Filters and the wrong Gmail look like failure after a finished hop. They are not a license to AUTH again.
Password-reset mail that already 250’d must not go out a second time with the same link unless you minted a new event on purpose. Receipts that must arrive still get one 250. Staging workers must not share production keys. Default retries in HTTP clients are how a finished submit becomes two invoices. Teach the library your key or disable those retries.
Technical mail flow
Your app or the recovery UI submits one envelope. MailerZ returns 250, 4xx, or 550. A 250 means the next hop accepted responsibility. Cloning that event is a new message with a new Message-ID. Recipients see duplicates. IETF RFC 5321 — Simple Mail Transfer Protocol does not dedupe for you.
Destination Gmail can also 4xx. That is their store. Backoff. Do not open port 25. Do not switch MX to “retry locally.” MailerZ is not an open relay. Unhosted or unauthorized From gets 550 / 550 5.7.1.
Inbound hops are a different conversation. Envelope SRS may rewrite MAIL FROM so destination SPF can survive. Header From, Subject, Date, Message-ID, body, and MIME stay as received. Replaying a stored inbound hop uses the recovery window. It is not your app’s outbound event key.
Step-by-step setup and decision path
Give every send an event key
Invoice-1044, signup-user-9, probe-2026-04-11. Persist it before AUTH.
Read history for that key
250 ends the story. Missing history is not permission to spray.
4xx: schedule backoff
Minutes, not milliseconds. Same key. Same body. Same From.
5xx: stop
Fix 550 5.7.1 (plan or From) or destination 5xx. Then a new decision, still one key if the business event is the same.
Recovery replay once
Mark the key replayed. A second operator must see it.
Never retry a list
US Federal Trade Commission — CAN-SPAM compliance guide still applies. We will not help design that loop.
Failure modes and proof
| What you see | Likely cause | Proof |
|---|---|---|
| Two copies in Gmail | Retried after 250 | Two history 250s |
| Hourly cap | Loop | Stop the worker |
| 550 5.7.1 storm | Free or unauthorized From | Plan and map |
| Empty history, keep retrying | Never reached MailerZ | MX, host, env |
| Operator double-click | No idempotency | Add the key |
Proof is one 250 per key. Two 250s is a bug you can name without poetry.
MailerZ workflow and product boundary
MailerZ is a custom-domain delivery layer operated by Secuno LLC. Authenticated SMTP plus inbound MX. Envelope SRS only. Header From never rewritten. Not a bulk ESP. Not IMAP. Not an open relay. Unhosted 550 5.7.1. Site: mailerz.net. App: mail.mailerz.net.
- Free $0: 1 domain, 10 aliases, 14-day store, send-as Off, SMTP Off, API Off.
- Solo $40/year: 5 domains, 25 aliases, 90-day store, 2,500 outgoing, 20/hour.
- Starter $8/$80: 8 domains, 50 aliases, 5,000 outgoing, 40/hour.
- Business $19/$190: 25 / 200, 12,000 outgoing, 60/hour.
- Agency $39/$390: 100 / 500, 20,000 outgoing, 60/hour.
- Unlimited $99/$990: uncapped domains and aliases, 180-day store, 100,000 outgoing, 300/hour.
Confirm numbers on pricing. MailerZ may retry internally on a transient destination 4xx inside its hop. That is not permission for your app to fire ten submits. Your event key still owns the business event. No SOC 2, ISO 27001, HIPAA, or inbox SLA.
Cost, alternatives, and trade-offs
| Choice | What you get | What you give up |
|---|---|---|
| Event key plus history | One invoice email | The comfort of smash-retry |
| ESP with native idempotency | Their AUP. Quote live. | This hop if you needed a mapped From here |
| Human-only retry | Slow and visible | Automation without keys |
Time is a line item. One key costs less than a week of “sorry for the extra emails.” Hitting Solo’s 20/hour cap is a signal to stop the worker, not to buy Agency so the loop can run longer.
4xx versus 5xx versus 250
4xx: try again later, same event. 5xx: do not try the same way. 250: finished. 421, 450, 451, and 452 are temporary classes. 550 5.7.1 is authorization. 552 is often quota. Read the full enhanced line. Do not map every 5 to “retry tonight.”
Event keys
Store the key before you connect. If AUTH fails, the key is unused. If 250, the key is terminal. If the process crashes after 250 but before you write the flag, history is the source of truth on next boot—not “send again to be sure.”
Unique subjects help humans search. They are not keys. Two retries can share “Your invoice” and still be two messages with two Message-IDs. Recipients search feelings. You search keys and history.
Useful formats: invoice:acme:1044, signup:user:9:welcome, form:contact:uuid, probe:zone:2026-04-11. Include the zone if you send for many clients. Do not use “test” or the current minute as a key. Two tests in one minute will collide.
Human retries
Recovery UI: one replay button, then disabled. A Slack “can you send it again” needs the key in the thread. Night operators who resend from Gmail as themselves create a different message and a different From. That is not a hop retry.
If a second operator needs the order, send this page plus delivery recovery. Classify. One key. No loop. No leftover MX. The artifacts that close a retry ticket are a terminal 250 per key and a disabled replay.
Crashes after 250
The dangerous sequence is AUTH succeeds, DATA succeeds, MailerZ returns 250, then your process dies before you write “sent” in your database. The next boot looks like failure. If you send again, the destination has two messages. On restart, look up the event key in MailerZ history. If a 250 exists, mark the key terminal.
If history is empty after a crash, the submit may never have reached us. Check host, port, TLS, and credentials. A timeout is not a 250. A timeout can still have been accepted if the connection died after they queued—this is why unique event keys and Message-ID logging matter. If you cannot see a 250, wait a short backoff and check history again before a second AUTH. Do not tight-loop.
Persist the event key before you open the socket. If you generate the key after DATA, a crash loses the only handle you had. Invoice-1044 must exist in your store first. SMTP is the side effect.
Workers, queues, and hourly caps
A queue that retries every second is a loop. Backoff on 4xx should be minutes, with a ceiling, and a human after N attempts. Hourly send caps exist so a bug cannot become a campaign. Hitting the cap is a signal to stop the worker, not to add more workers.
Shared SMTP across clients couples their caps. One broken site spends the fleet’s hour. Isolate credentials per zone. Isolation also isolates retry storms.
Log the event key, the SMTP code, and the history ID. Do not log the password. Do not paste AUTH traces into Slack. Staging workers pointed at production SMTP with production keys will retry real customers. Staging should use a staging zone and a sink mailbox, or it should not send. A “test” that mails production users is bulk with a typo. US Federal Trade Commission — CAN-SPAM compliance guide still applies.
Inbound replay is not outbound retry
Recovery replay of a stored inbound hop is a different button. It uses the store window. It is not your app’s event key. Mixing the two on one ticket creates two copies plus leftover-MX theories. Classify first. Inbound empty history is MX. Inbound 5xx is destination. Outbound 550 5.7.1 is plan or From. Outbound 250 is finished.
Do not add leftover MX because a retry failed. Dual MX does not fix AUTH. It splits inbound and leaves the outbound bug in place. Do not open port 25 to “retry locally.” Submit to dashboard SMTP only.
What a duplicate looks like to a human
Two messages, two timestamps, same invoice, slightly different Message-IDs. The customer forwards both to support. Support retries again. You now have three. The fix is a terminal key and an apology that names the bug—not another send.
Gmail conversation view can hide a duplicate under the same thread. Search by Message-ID if you have it. History 250 count for the key should be one. Two 250s is the bug you can show a founder.
Header From stays the mapped alias. MailerZ does not rewrite it. A duplicate is not “Gmail showing it twice because of SRS.” Envelope SRS is inbound hop mechanics. Do not blame SRS for a second AUTH you performed.
When you must not retry at all
5xx user unknown, 5.7.1 unhosted, policy reject, and “this recipient asked you to stop” are stop conditions. Changing the body and retrying the same unwilling recipient is still a retry. We will not help design that.
Purchased lists, scraped lists, and warm-up services are out of scope. A for-loop over users is bulk. Delete the script. Free has no send-as. Paid send-as is operational mail, not a campaign ESP. No inbox SLA. Not SOC 2.
Night operators who smash retry until Slack goes quiet create the duplicate the customer will remember. Write “one key, one 250” on the ticket. If a second operator needs the order, send this page plus delivery recovery. Classify. History first. No loop. No leftover MX.
History 250 for this event key: stop. History 4xx: backoff, same key, same body, same From. History 5.7.1: fix plan or map, then one new decision. History 5xx destination: fix the store, then one replay if inbound stored, or one outbound submit if the business event never 250’d. History empty: fix host, port, TLS, leftover MX if this was inbound, then look again before a second AUTH. Human smash-click: disable the button.
Three retries that went wrong and one that did not
Story one: a worker crashed after 250 and sent again on boot. Two invoices. Fix: history first on boot. Story two: on-call hit replay and the app also retried the 4xx. Two copies. Fix: one owner for the event key. Story three: 550 5.7.1 on Free, so someone opened port 25 and mailed from Postfix. That is an open-relay path, not a retry. Fix: pay for send-as and map From. Story four: 452 on the destination, space freed, one replay, one arrival, button disabled. That is the job.
Write those stories in the runbook with the key format you actually use. On-call at 2 a.m. will not read a philosophy. They will follow a map. History 250 stop. 4xx wait. 5xx fix. Empty history check MX and env. No leftover MX. No loop. No list.
Agencies: one key namespace per client zone so logs do not collide. Shared keys across clients are how you mark Client B sent when Client A 250’d. Isolate SMTP and isolate keys. Confirm pricing so hourly caps are expected. Solo is 20 per hour. Agency is 60 per hour and 20,000 per month. Those are operational mail, not a campaign budget. No lists. No warm-up. No leftover MX on a retry ticket.
Password-reset tokens should be single-use in the app. A retry after 250 that includes a new token is a new event. A retry after 250 with the same token is a duplicate email with the same link. Decide which you mean before AUTH. Receipts that “must arrive” still get one 250. If the customer says they cannot find it, search the destination and history before a second send. 250 plus empty inbox is a filter. Leftover MX is inbound. Do not clone the receipt.
Put the decision map in the mailer module as comments next to the send function. Future you will thank present you when a queue library retries by default. Default retries in HTTP clients are a common way to clone a 250. Disable them or make them key-aware. The network library does not know your invoice number unless you teach it. Staging keys never hit production SMTP. Secrets stay in the store. Scripts get deleted after a probe.
Free has no send-as. A retry storm on Free is a 550 storm. Upgrade first. Then one key. Then one 250. Then stop. That order is the whole article in four sentences. People-first incident notes name the key and the code. A vibe that says we kept trying is how you get three copies and a founder in the thread.
FAQ
What is the safest way to handle retry failed email safely?
Retry only after a 4xx with backoff, using one event key so a second click does not create a second message. After a 5xx, stop and fix the destination or From. After a 250, never resend the same event. Do not loop users. Do not add leftover MX. History is the judge.
Does this require a new mailbox?
No. Retries use the same destination. A new mailbox does not make a 250 safe to clone. MailerZ is not IMAP.
Will it work with Gmail or Outlook?
Destination 4xx and 5xx still apply. Self-send can hide inbound issues. Send-as retries need a mapped From and a paid plan. Free has no send-as. Confirm /pricing.
What DNS records are involved?
None for a clean 4xx retry. If history is empty, leftover MX is a different ticket. See RFC 5321. Dual MX is not a retry strategy.
What should I test before production?
One unique message, one event key, confirm 250, confirm a second submit is a no-op. Delete the test script. Do not retry in a for-loop.
Key takeaways
- Never resend after a history 250 for that event key. A second AUTH is a second message.
- 4xx backoff with the same event key, same body, and same From. Minutes, not milliseconds.
- 5xx stop and fix the plan, the From map, or the destination store. Then one new decision.
- No user loops. Hourly and monthly caps are not a campaign budget. Confirm pricing.
- History beats a crash flag. Persist the key before you open the socket.
- Leftover MX is not a retry. Opening port 25 is not a retry. Personal Gmail resend is a different message.
- Free has no send-as. Isolate keys and SMTP per client. Not SOC 2. Not an inbox SLA.
- Inbound replay uses the store window and is a different button. One owner. Then stop.
Conclusion and next action
If you retry mail, retry the event, not the panic. MailerZ can 250 a mapped From. It will not collapse two 250s into one inbox. Start free for inbound, paid when the app sends with a key. Persist the key before AUTH. Read history on boot. Disable default HTTP retries that clone a 250. Isolate keys and SMTP per client.
4xx waits with the same body. 5xx stops until the cause changes. Empty history is env or leftover MX, not a smash-click. Inbound replay is a different button with a store window. Humans who resend from personal Gmail create a different message. Write the key in the thread. One owner. One 250. Then stop. Two-factor on destinations. HOLD of an unknown local-part is another class. Exclusive MX is another class. Caps are a backstop, not a design. If you need more than the card, you may be looping or you may need a real ESP for campaigns. MailerZ is not that ESP. Confirm pricing so the hourly number is known before the worker starts. Delete the probe script. Do not log SMTP passwords. Search the destination before you clone a receipt that already 250’d. If they still cannot see it, you have a filter or the wrong Gmail, not a license to send again. One 250 is enough proof the hop worked. Write the key in Slack so a second operator does not fire the same event. Then disable the replay button.
Ready to send once
Start free for inbound, paid for idempotent send-as.
One key. One 250. Sign in if the domain is already there.
Persist the event key before AUTH so a crash cannot invent a second send. On boot, history is the source of truth. If a 250 already exists for that key, mark it terminal and stop.
Review when send workers or recovery UI change. Author: MailerZ editorial, Secuno LLC.