Every morning at Gridky, I manually deduplicated the exceptions arriving in Slack. There were hundreds a day. Several messages could correspond to one defect. The repetitions had to be identified before choosing fixes.

The team was split in two: one part handled incidents and fixes, the run work, while the other developed features, the build work. Requests from a priority client continued to arrive with considerable delivery pressure.

Drynuary paused new features so the whole team could focus on stabilisation. Planned for one month, it lasted a month and a half.

Part of the team was assigned to incidents

Separating run and build let us respond to bugs while continuing feature work. Its cost was visible in work allocation: developers assigned to incidents were unavailable for the rest of the product.

Exception triage added a recurring task. Monitoring sent the messages, but we lacked the time and tools to use that volume easily. My deduplication helped choose the day’s fixes without preventing errors from returning while their causes remained.

The target was to reunite the team and return to the product roadmap. A lower message count alone would not have established that result.

No features, including for the priority client

The decision was negotiated before the work began. Our CTO discussed reserving a month of the roadmap with the CEO for stabilisation, tests and reworking existing code.

Developers, product and sales shared the need for a stable product. The agreed scope admitted no new features during the pause, including requests from the priority client.

That rule changed how work was assigned. Developers who would have handled those requests joined the fixes. Stabilisation no longer had to fit only between features or remain with the part of the team already handling incidents.

All changed code had to be tested

We began by listing bugs and examining test coverage, which was still low. That gave the work its initial issues and identified behaviours with few checks.

The rule applied to changed code, including legacy code. A fix had to come with tests. A function written before the rule was introduced received no exemption once it was changed.

Compliance could be checked in code review. A change that did not meet the requirements could be blocked from merging.

Reporting helped us choose fixes

We had a senior team committed to the work. Each developer could choose an issue within the stabilisation scope. Severity and exception frequency provided grounds for prioritisation.

Errors that recurred frequently were also, in several cases, quick to fix. Addressing them reduced repetition in the reports. Other problems then became easier to see.

I contributed fixes, monitoring and refactoring alongside the other developers. The CTO and CEO owned the scheduling decision, while the team handled the issues. Those contributions were distinct.

Release made the effects observable

Before Drynuary, the move to Kubernetes had already enabled several deployments a day. Fixes could therefore reach production within hours, sometimes minutes.

Reporting then showed whether the relevant exceptions returned. Merged code and its production effects could be followed over a short interval instead of waiting for a bundled release weeks later.

The extension followed the first results

At the end of the planned month, the issue list was not exhausted. Management had, however, seen more tests, more stable critical features and fewer exceptions in the reports.

Those results supported an agreed extension. The work ultimately lasted a month and a half. The revised deadline was discussed with management rather than drifting as tickets remained open.

The team reunited and morning triage disappeared

The most directly observable change in my work was the end of manual deduplication every morning. The team had also reunited. Remaining bugs could be fixed between development tickets.

Before stabilisation After stabilisation
Team split between run and build Team reunited
Manual exception triage every morning That daily triage was no longer needed
Part of the team assigned to fixes Remaining bugs handled between development tickets

I estimate that production bugs fell by around 80%. The Gridky case study gives another reference point: roughly 10 bugs a week, then 1 or 2 a month over six months. These estimates do not form a measurement series from which one can be calculated using the other. The six months extend beyond the pause itself.

The hundreds of daily exceptions are another count again, since one defect can generate several messages. I therefore do not use them to calculate a bug reduction rate. The table describes organisational changes I can report directly.

The rule still blocked insufficiently tested changes

After feature work resumed, testing requirements remained in place. Code reviews could still block a change that did not meet them. The team had seen the effects of that discipline and maintained it.

That distinguishes the work from closing a stock of bugs before returning to the previous process. New development resumed, but changed code remained subject to the verification rule adopted during stabilisation.

At Gridky, we stopped incoming feature work, brought the team together on fixes and maintained testing requirements afterwards. The result was a return to the product roadmap without retaining the run/build split.

The five faces of technical debt separates the code, tooling and organisational difficulties that can come together in such a decision.

Resources for this article