The hour nobody budgets for: what a data incident really costs
A failed job costs more than the rerun. Most of the bill is the time one engineer spends finding out what happened, and the people who wait while they do.
- What failed
- The revenue number on the morning dashboard looks low.
- Who waited
- Analysts, dashboards, downstream jobs, a customer-facing number
- First sign
- 08:50
- Cause found
- 10:25, 1 h 15 min after the engineer started
- Fixed
- 10:40
- Wrong for
- 1 h 50 min
- Published
- Reading time
- 5 min in full
An example morning
A data incident rarely starts with an alarm. It starts with a message.
| 08:50 | Failed: The revenue number on the morning dashboard looks low. | |
| 09:05 | Someone in finance asks in chat whether the number is right. | |
| 09:10 | The engineer on call puts their own work down and starts looking. | |
| 10:25 | The cause is found, 75 minutes after the engineer started. | |
| 10:40 | Recovered: The fix goes out and the dashboard is right again. |
The full account, 2 more paragraphsFold the account
Someone in finance opens the morning dashboard and the revenue number looks low. Or a scheduled job went red overnight and nobody saw it until the first coffee. Either way, the first sign many teams get is a person asking a question in chat: "Is this number right?"
This post is about what happens next, and what it costs. It is the first in a series. Each later post takes one real failure and walks through it step by step. This one sets out the problem they all share.
What it looks like from the inside
08:50 to 09:10
Incidents come in two common shapes.
The full account, 3 more paragraphsFold the account
In the first, something fails loudly. A job stops with an error, the scheduler marks it failed, and an alert lands in a channel. You know that something broke. You do not yet know why, or what else it touched.
In the second, nothing fails at all. Every job finishes and every status is green, but a number is wrong. A column changed type upstream, or a source sent half its usual rows. This shape is worse, because the first person to notice is often someone who uses the data, not someone who runs it.
Then the questions arrive. Is the dashboard safe to use? Should the weekly report wait? Did last night's export go out with the bad figures? One engineer, often whoever is on call, stops what they were doing and starts to dig.
Where the time goes
09:10 to 10:25
The fix itself is often small: a cast, a filter, a rerun. What takes the time is everything before the fix, which is working out what happened.
| Step | Takes | By |
|---|---|---|
| Find which run failed, and when it last worked | 10 min | 09:20 |
| Read the error and the logs around it | 10 min | 09:30 |
| Work out whether the data changed or the code did | 15 min | 09:45 |
| Trace upstream through lineage to where it started | 15 min | 10:00 |
| Check recent commits and deploys | 10 min | 10:10 |
| Ask the people who know the business (Part of this is waiting for a reply) | 15 min | 10:25 |
| Total, estimated | 1 h 15 min | 10:25 |
The full account, 3 more paragraphsFold the account
Here is that work as a list, with a time against each step.
The times in that list are our judgement for a mid-sized team, and nobody timed them. Yours will differ.
Two things stand out. Almost none of the work is typing a fix. And the steps are not hard, they are scattered. The run history is in one tool, the logs in another, the lineage in a third if it exists at all, and the commits in a fourth. The reason a column means what it means is in somebody's head.
Who is waiting
08:50 to 10:40
While one person digs, other people wait.
- Analysts
- either stop work, or carry on and risk building on a wrong number.
- Dashboards
- keep showing stale or wrong figures to everyone who opens them. Most of those people never see the chat thread.
- Downstream jobs
- fail in turn, or run on bad input and pass it along.
- A customer-facing number
- is sometimes involved: a usage figure, an invoice line, a report that goes out on a schedule.
The full account, 1 more paragraphFold the account
The engineer's time is the part you can see. The waiting is wider, and it grows for as long as the cause stays unknown.
What reduces it without buying anything
Most of the time in the list above goes on finding context that was never written down or kept. A good part of that context can be put in place with habits, and with tools you most likely have already.
The full account, 1 more paragraphFold the account
None of this stops incidents from happening. It shortens the time each one takes.
Where Convalesce fits
Convalesce does the digging. When a run fails, it investigates the failure: the run, the error, what changed, and what sits upstream and downstream. It shows the evidence it found, so you can check the reasoning yourself. Then it proposes a fix for a person to approve, and nothing ships without that approval. The steps are set out in how it works.
The full account, 1 more paragraphFold the account
The next posts in this series each take one failure and follow it from the first red mark to the fix.
Stop reconstructing failures.
Sign in with GitHub or Google, connect a tool, and see your next failed run explained.