Troubleshooting & SMTP Codes

Email incident runbook: missing custom-domain messages

Evidence before DNS. You cannot recover mail that hit another host.

MailerZ editorial · Secuno LLC17 min read

A missing email incident runbook starts with time, local-part, and a third-party sender—not with a DNS panic. Look up exclusive MX, open hop history, and classify leftover hosts, unknown HOLD, destination reject, or a folder the user did not search. Write the class in one sentence before anyone publishes a record. MailerZ keeps hops in the store window so the runbook has evidence.

Missing email incident runbook: evidence before DNS edits
Time, alias, history, then one change. Not a Friday MX flap.

Quick answer for missing email incident runbook

Incidents feel like outages. Many are a missing alias or a leftover Google MX that still owns half the planet’s lookups.

The runbook exists so two people do not edit DNS at once. One incident commander. One editor.

Transport facts live in RFC 5321. Your status page should quote a hop, not a vibe.

If the store window still covers the claimed send time, you either have a hop or you do not. Empty plus leftover MX is the usual pair.

Communicate in classes: “mail is at the old host,” “mail is held,” “destination refused,” “accepted, search junk.” Customers can act on those. They cannot act on “we are looking at servers.”

Close only after a new unique probe succeeds. Old missing messages may stay missing if they never hit you.

Authoritative mail transport is defined in IETF RFC 5321 — Simple Mail Transfer Protocol. Product path: troubleshooting, delivery recovery, and docs.

User problem and decision criteria

Decision criteria: blast radius (one alias or all), time window, whether MX is exclusive, whether history exists, whether the destination owner searched.

Criteria that do not belong: adding backup MX, rewriting From, promising recovery of mail that hit another host, inboxing percentages.

Agencies should keep the old MX screenshot in the ticket from day one of any prior migration. Incidents without that screenshot take longer.

If leadership wants a public status update, use the class sentence. Do not invent an SLA.

If legal asks for SOC 2 evidence, point at the security page. Do not mint a badge during an incident.

If the only missing mail is from one portal, isolate that sender. Do not declare the domain down.

If send-as failed, split the incident. Inbound and outbound are different commanders.

Night pages should wait on 4xx. Flapping MX overnight creates Monday’s incident.

Technical mail flow

Incident mail flow
If history is empty, the message may never have reached this MX.

Claimed send → public MX path → MailerZ or leftover host → HOLD or forward → destination.

You only see what reached you. Message traces at Google or Microsoft help when the destination is theirs.

SRS envelope and intact Header From still apply during incidents. Do not “fix” an incident by rewriting From.

Tools: leftover MX checker, MX lookup, history. Use them in that spirit.

Recovery inside the store window is for hops you actually have. It is not time travel to another provider.

Step-by-step setup / decision path

Runbook order
Scope, MX, history, classify, one change, communicate.
  1. Open the incident: time window, alias, destination, sender domain, who already edited DNS.
  2. Lookup NS and MX from two resolvers. Screenshot.
  3. If leftover owners exist, treat that as the primary hypothesis.
  4. Open history for the window. Copy codes.
  5. Classify: leftover, HOLD, 5xx, 4xx, 250-plus-folder, or no evidence / expired store.
  6. Communicate the class. Assign one DNS editor if MX must change.
  7. Change one fact. Re-probe from a third mailbox.
  8. Close with the probe Message-ID and the class. Schedule a postmortem if MX was dual.

If evidence expired, say so. Request a resend. Do not fabricate a hop.

If leftover MX is confirmed, deleting it is the fix. Demoting priority is not.

Keep a written rollback: the old MX set you screenshotted. Incidents nested inside incidents come from memory edits.

Failure modes and proof

DNS flap without class: longer outage. Proof: change log versus empty history.

Two editors: ping-pong MX. Proof: serial increments.

Promised recovery of old-host mail: you cannot. Proof: that host’s MX answered.

Self-send declared as all-clear. Proof: customers still missing.

SPF edit during inbound incident: confusion. Proof: MX still dual.

Status page silence: angry customers. Proof: no class sentence.

Catch-all FORWARD as incident response: junk later. Proof: the setting change.

Expired store plus guessing: mythology. Proof: timestamps.

Merged inbound and outbound: wrong owner. Proof: AUTH errors in an inbound ticket.

Invented SLA: legal problem. Do not.

Header rewrite “to help”: product breach and user distrust.

Closing without a probe: reopen tomorrow.

MailerZ workflow and product boundary

MailerZ is custom-domain aliasing and forwarding with optional paid send-as. Secuno LLC operates mailerz.net. The app is mail.mailerz.net. Not Workspace, not IMAP, not an open relay, not a campaign ESP.

Envelope SRS only. Header From, Subject, Date, Message-ID, body, and MIME stay intact. Exclusive MX. Hold unknown on Free. Copy SMTP host, port, and TLS or STARTTLS from the dashboard when you send. Do not invent 587 or 465 as MailerZ facts.

Free: one domain, ten aliases, one seat, fourteen-day store, send-as disabled, SMTP and API disabled. Solo forty dollars a year, twenty-five aliases, ninety-day store, 2,500 outgoing, 20 send-as per hour. Starter eight monthly or eighty yearly. Business nineteen or one hundred ninety. Agency thirty-nine or three hundred ninety. Quote the pricing page. No SOC 2, ISO, HIPAA, SLA, or inboxing percentage.

Use delivery recovery and troubleshooting as the working set. The runbook is operational, not a new product module.

Cost, alternatives, and trade-offs

A one-page runbook is cheaper than a brand-damaging weekend.

Workspace seats do not classify leftover MX.

A second vendor mid-incident is usually panic spend.

Agencies should include this runbook in retainers. It reduces emergency hours.

Honesty about unrecoverable mail costs less than a fake restore date.

HOLD review during incidents can close “missing” as “never mapped.” Cheap.

Store windows are why you do not wait a month to open a ticket.

Postmortems prevent paying for the same leftover MX twice.

Incident communications and ownership

Name an incident commander who does not edit DNS and an editor who does not write the status sentence. Mixing those jobs is how you publish “we are investigating” while also flapping MX.

The first status sentence is allowed to be “we are classifying leftover MX versus HOLD versus destination reject.” That is better than silence and better than “servers are down.”

Put the time window in UTC in the channel topic. People will quote local time. Store expiry does not care about their laptop clock.

If leftover MX is the class, the honest customer sentence is “some senders still reach the old host.” Give them the old host name if you still control it so they can look. Do not promise MailerZ can pull those messages back.

If HOLD is the class, the sentence is “the address was not on the map; we are adding it.” That is an apology, not an MX outage.

If destination 5xx is the class, the sentence names the receiver. You cannot vote Microsoft into 250. Offer a temporary alternate alias only if you can support two maps.

If 250-plus-folder is the class, the sentence is “search junk and rules.” Send a screenshot guide. Do not open a second incident for DNS.

Split outbound AUTH into another channel. Mixed incidents create two wrong owners and zero fixes.

After close, schedule a fifteen-minute postmortem if MX was dual or if two people edited. The artifact is the screenshot you should have had. The action is a calendar rule: no cut without a picture.

Do not invent an SLA in the close note. Do not invent an inboxing percentage as a make-good. Offer a probe Message-ID and the class. That is the professional close.

Night on-call should have permission to wait on 4xx. Write that permission down. Heroes who flap MX at 3 a.m. create the morning incident.

Practice the runbook on a test alias quarterly. Unused runbooks rot. Used ones stay short.

Freeze, classify, then change one fact

An email incident runbook for missing custom-domain messages starts with freeze. Do not rotate MX, SPF, and aliases in the same hour. Collect the time window, the public local-part, the dest mailbox, and whether anyone already “fixed DNS.” Print public NS and public MX from two resolvers. Open MailerZ history for that window. Name one class: leftover MX, unnamed or held alias, dest 4xx, dest 5xx, dest folder, or self-send lie. Then change one fact. MailerZ history is the artifact for hops this layer saw. Empty history is not a dest reject.

Leftover Google, Microsoft, Cloudflare routing, or registrar MX is a hard stop. Priority is order, not load balancing. Delete leftovers. Wait TTL. New unique subject from a mailbox that is not the destination. Gmail to the same Gmail can short-circuit and close a ticket that is still split. Two views must show exclusive MailerZ MX before you talk about spam.

Held unknowns look like vanishing mail. Review the store. Promote only when a real person used the leftover. Do not enable catch-all forward to make the incident green. Fan-out trains two spam buttons. Plus addressing on Gmail is not a custom-domain unknown policy. MailerZ will not strip plus tags on your domain.

Dest 250 plus spam is delivered-to-the-store. Search all mail and junk first. Classify is hop four. We do not promise Primary. We do not sell an inbox SLA. Dest 4xx is try later — quota and greylist. Dest 5xx is a final no for that attempt. Free send-as 550 is a plan boundary, not inbound missing. Unauthorized send is 550 5.7.1. MailerZ is not an open relay. Confirm /pricing.

What not to collect

Do not mail SMTP secrets or Gmail passwords to support. Send a timestamp, unique subject, public MX text, history row or the word empty, and a dest folder screenshot. Envelope SRS only. Header From stays. If another hop rewrote Header From, say so. We will not rewrite Subject, Date, Message-ID, body, or MIME.

Do not open fifty recovery tickets for one leftover. The store window is fourteen days on Free and ninety on paid. It is not legal hold. Not SOC 2. Not HIPAA. Point questionnaires at the published security and privacy pages.

Four tickets, four classes

Ticket A: plugin said sent, history empty, MX still Workspace. Class leftover. Delete, wait, probe. Ticket B: history empty, MX exclusive, local-part never named. Class hold. Name it or leave it. Ticket C: history forwarded, dest 4xx quota. Class dest. Empty Deleted Items, new subject. Ticket D: dest 250, user looked only in Primary. Class folder. Open spam. No MX change.

A loop can look like a flood incident. Dest forwards back to the public alias. Disable the dest rule. Hold unknowns. One dest you control. Repeating subjects in history are the tell.

Related: delivery recovery, troubleshooting, docs, features. RFC 5321. Google and Microsoft dest search docs. Commercial pages nofollow. After restore, still probe from a third mailbox before you close.

Close line

Class named, one change made, new subject, neighbor aliases probed, leftovers still gone on two views. If you cannot print exclusive MX, the incident is still hop one. Start free on a domain you can break. Sign in if the zone already lives here. That is the email incident runbook for missing custom-domain messages.

Agencies run the classify per client zone. A template that restores leftovers will reopen ticket A tomorrow. Offboard deletes MX you own and revokes SMTP. Review after a wizard or a staff departure. Do not “fix” inbound by editing SPF.

If the window is older than the store, history may be gone. You still have public MX and dest search. Say the window expired. Do not invent a hop. Confirm /pricing if the miss was Free send-as all along.

The classify table on the wall

Empty history plus leftover MX: delete leftover, wait two views, new subject. Empty history plus exclusive MX plus unnamed local-part: hold or name it. Empty history plus exclusive MX plus named alias: dest never mapped or dest dead — fix dest, not DNS. History accepted not forwarded: dest session failed — read the dest code. History forwarded dest 4xx: quota or greylist. History forwarded dest 5xx: dest no — read enhanced status. History forwarded dest 250, user empty: search junk and all mail. History repeating subjects: dest loop. Plugin sent, all of the above empty: leftover answered the plugin.

Freeze means no SPF edits for inbound. No second MX. No catch-all flip. No SMTP paste. No shared secret in the ticket. Collect window, alias, dest, MX print, history word. Change one fact. Neighbor aliases are the regression probe after the change. Unique subjects.

Self-send from Gmail to dest Gmail is not an incident test. A green self-send with empty customer history is leftover or short-circuit. Use Outlook.com or another provider. Confirm /pricing if the “miss” is Free send-as 550. Fourteen-day Free store, ninety-day paid. Older than the window: say expired. Do not invent a hop.

Related: delivery recovery, troubleshooting, docs, features. RFC 5321. Google dest search. MailerZ is not IMAP. Envelope SRS only. Header From stays. Not an open relay. Not SOC 2. Not HIPAA.

Escalation packet

Timestamp, subject, MX text, history or empty, dest folder, enhanced SMTP line if any. No passwords. Dest-side escalate only with forwarded plus dest 4xx/5xx. MailerZ escalate only for hops this layer stored. Leftover MX is not an escalation. Start free to practice the table on a breakable domain. Sign in if the zone already lives here. That is the email incident runbook for missing custom-domain messages.

Agencies: one classify per client. A wizard tomorrow reopens leftover. Offboard deletes MX you own. Review after staff departure. Do not merge three clients into one red Slack thread.

Classify before you touch the zone

Missing custom-domain mail is four different tickets that share a Slack title. Leftover MX: hop history empty, public resolvers show two owners. HOLD: history exists, local-part was never created, Free held it. Destination 5xx: alias existed, Gmail or Outlook rejected, SMTP line is in the hop. Folder: hop delivered, human is searching the wrong Gmail or a filter ate it. If you restore aspmx before you classify, you will create a fifth ticket called leftover MX that you invented during the incident.

Freeze DNS until classification is written. The urge to “just add Google back” is how random delivery starts. Collect: UTC window, alias as the sender typed it, destination mailbox, whether the sender is the destination (invalid), and a stranger probe to the same alias with a unique subject. Self-send from Gmail to Gmail is not an incident test. It never was.

Read two public resolvers before you open the registrar. If they disagree, you may be inside TTL, not inside a product outage. If they agree on leftover Google plus MailerZ, you have your class. Delete leftovers. Do not add more hosts. Dual MX is not a diagnostic tool. It is the failure mode.

Open hop history next. Empty after a stranger probe that two resolvers say should hit MailerZ means the name never arrived or you are looking at the wrong domain. Confirm the zone. Confirm NS. A migrated domain whose panel is pretty and whose public NS is old will produce empty history forever. That is not missing Vault. That is the wrong post office.

HOLD is success of exclusive MX plus a missing map. Tell the reporter: we have the message if it is inside the store window; the name was not an alias. Create the alias if it is a printed name. Ask for a resend if you cannot replay. Fourteen days on Free, ninety on paid. Opening Vault does not create the alias. Creating the alias does not make MailerZ IMAP.

Destination 5xx is a mailbox ticket. Copy the SMTP line into the incident. Full mailbox, blocked sender, and policy are restores on Gmail or Outlook. Do not delete the alias. Do not republish leftover MX. Do not file it as HOLD. Wrong class is how the next on-call repeats the wrong fix at 2 a.m.

Folder class needs All Mail, junk, and filters. Delegates and shared inboxes hide mail. The hop will look green. The human will swear nothing arrived. Screenshot search, then hop. Close as destination search, not as MailerZ down. There is no inbox SLA to invoke. Do not invent one on the call.

Communications that do not make leftover MX worse

Customer language: we are checking which host answered, then whether the name exists, then whether Gmail accepted. We are not “migrating again.” We are not adding a backup MX. If we ask you to resend, it is because HOLD never forwarded or the destination rejected. It is not because DNS “needs 48 hours” when two resolvers already agree.

Internal language: one writer on the zone. Registrar phone apps stay closed. Vendor tickets to Google that say “add aspmx” are their runbook, not yours. File cancel-seat tickets after exclusive MX is green. Do not file them as the fix for missing mail.

After restore, probe again from a stranger mailbox before you close. A close without a probe is a hope. Hope returns at Monday invoices. Write the class, the one change you made, and the probe subject in the ticket. That is the runbook artifact. Slack huddles are not.

Night watch if the class was leftover MX. Leftovers return when someone opens a registrar upsell. Morning MX paste from two resolvers is part of close. Unowned morning checks are how leftovers live a quarter.

Facts: envelope SRS, Header From intact, not IMAP, not SOC 2, store windows as priced. Start free if you need HOLD to teach the map. Sign in when the domain already lives on the hop. /troubleshooting for leftover MX. The incident does not grow Vault into a mailbox host.

Close criteria you can put on a card

Close leftover MX when two resolvers are exclusive, leftovers gone, stranger probe landed. Close HOLD when the alias exists or the name is refused on purpose, and the reporter knows to resend if needed. Close 5xx when the destination accepts a new probe. Close folder when the message is visible in the destination search. If you cannot meet a close, the incident is still open. Adding a second MX is never a close action.

If two classes apply, sequence them. Leftover first, then HOLD, then 5xx, then folder. Fixing folder while leftovers exist trains you to think Gmail search is the product. It is not.

FAQ

What is the safest way to handle a missing email incident runbook?
Freeze DNS until you classify. Collect window, alias, destination, and a probe. Read public MX and hop history. Name leftover MX, HOLD, 5xx, or folder. Then change one fact. MailerZ history is the artifact.
Does this require a new mailbox?
No. Missing mail is usually routing or search. MailerZ is not IMAP.
Will it work with Gmail or Outlook?
Search all mail and junk first. Self-send is not an incident test.
What DNS records are involved?
NS and exclusive MX first. Verification TXT is ownership. SPF is outbound. Do not touch SPF to find inbound mail.
What should I test before production?
This runbook is for incidents. After restore, still probe from a third mailbox before you close.

Key takeaways

  • Evidence before DNS.
  • One commander, one editor.
  • Classify leftover, HOLD, reject, or folder.
  • Empty history is a clue, not a void.
  • Do not promise mail that hit another host.
  • Re-probe to close.
  • No invented SLA or inbox rate.
  • Write the class where customers can read it.

Conclusion

Missing custom-domain messages yield to a boring runbook. Time, MX, history, one sentence, one change.

Start free and practice the runbook with a fake outage on a test alias before you need it on a Friday.

Start free on MailerZ