At 10:10 am, support receives a request for an invoice. The order is paid and the document exists in the software. The customer has not received it. The dashboard still says the website is working.
To follow the incident, this article uses a fictional batch of 120 orders paid between 9 and 10 am. An invoice was created for each. The sending provider accepted 118 messages and rejected the last two, at 9:58 and 9:59, because the software’s access credentials had expired.
The availability check opens the homepage every minute. It responds every time. There is no contradiction in these results: that check does not send an invoice.
Comparing records reveals the two missing sends
At 10:10 am, the example batch contains the following information.
| Stage | Count | What it establishes |
|---|---|---|
| Paid orders | 120 | 120 invoices expected |
| Invoices created | 120 | No documents missing at creation |
| Sends accepted by the provider | 118 | Two invoices without an accepted send |
This tracking uses invoice identifiers, rather than just an email counter. Sending the same invoice twice could bring a counter to 120 while still leaving a customer without a document.
The invoices associated with the 9:58 and 9:59 rejections have no accepted send. At 10:10, they have been waiting for twelve and eleven minutes. This fictional service expects invoices to be sent within five minutes. Both have exceeded that interval.
The check can raise an alert containing those two identifiers and their age. An overall drop in traffic does not affect this calculation: the orders are paid and the expected invoices are known.
The test passes, then access expires
Before release, a test creates an order and checks that its invoice is passed to the sending system. The test service accepts the message. That result covers the behaviour verified at that point in time.
In this incident, rejection begins later, when access to the provider expires. Rerunning the test against a simulated service that always accepts messages will not detect that expiry. The actual rejections and the two waiting invoices exist in production.
Monitoring unsent documents covers this second case. It still does not guarantee final receipt: a provider can accept a message that the destination mail server later rejects. “Accepted for sending” therefore remains the accurate label for the measured result.
Establishing that invoices reached their recipients would also require examining the available delivery information. Renaming the indicator “customers served” would not add that evidence.
At Gridky, repeated errors had to be grouped first
At Gridky, hundreds of exceptions arrived in Slack every day. I manually deduplicated them each morning to identify those recurring most often. Frequent errors often corresponded to quick fixes.
A small numerical example, separate from that experience, explains the work. If a defective operation is retried twenty times, it can produce twenty messages for one defect. Fixing the defect can remove all twenty messages. That does not mean twenty bugs were fixed.
Deduplication converted messages into problems to address. The twenty messages in this example are not a count of fixes at Gridky. In our daily work, manual triage disappeared after the stabilisation described in Drynuary.
After the repair, two invoices still need attention
In the example, the team restores access to the provider. A new invoice is accepted for sending. That establishes that new messages can pass, but the two earlier rejections are still recorded as rejections.
Recovery targets those two identifiers. If the first is accepted and the second fails again, the original batch has 119 invoices with an accepted send and one still waiting. Tracking reveals that result without confusing it with the successful new test message.
Before retrying a document, the team checks its previous attempts. An explicit rejection and a missing response carry different information. In the latter case, the provider may have accepted the message without the software receiving confirmation.
The batch is accounted for when each invoice has an explained state. For any that remain unsuccessful, support knows which documents need attention instead of waiting for another complaint.
Two possible reports of the same incident
“The website is available and a test email was sent” describes two successful checks. It leaves the status of the two rejected invoices unresolved.
A report following the batch provides different information: 120 invoices expected, 120 created, 118 initially accepted for sending, followed by the outcome of recovering the remaining two. The conclusion then depends on work actually completed.
That is why the choice of checks changes what is known about the incident. The homepage check was not broken. It answered a question that could not locate the missing documents.
This tooling case complements the five faces of technical debt.
