/blog/what-a-data-incident-really-costsBlog

The hour nobody budgets for: what a data incident really costs

A failed job costs more than the rerun. Most of the bill is the time one engineer spends finding out what happened, and the people who wait while they do.

What failed
The revenue number on the morning dashboard looks low.
Who waited
Analysts, dashboards, downstream jobs, a customer-facing number
First sign
08:50
Cause found
10:25, 1 h 15 min after the engineer started
Fixed
10:40
Wrong for
1 h 50 min
Published
Reading time
5 min in full

An example morning

A data incident rarely starts with an alarm. It starts with a message.

An example morning
08:50Failed: The revenue number on the morning dashboard looks low.
09:05Someone in finance asks in chat whether the number is right.
09:10The engineer on call puts their own work down and starts looking.
10:25The cause is found, 75 minutes after the engineer started.
10:40Recovered: The fix goes out and the dashboard is right again.
The full account, 2 more paragraphs

Someone in finance opens the morning dashboard and the revenue number looks low. Or a scheduled job went red overnight and nobody saw it until the first coffee. Either way, the first sign many teams get is a person asking a question in chat: "Is this number right?"

This post is about what happens next, and what it costs. It is the first in a series. Each later post takes one real failure and walks through it step by step. This one sets out the problem they all share.

What it looks like from the inside

08:50 to 09:10

Incidents come in two common shapes.

The full account, 3 more paragraphs

In the first, something fails loudly. A job stops with an error, the scheduler marks it failed, and an alert lands in a channel. You know that something broke. You do not yet know why, or what else it touched.

In the second, nothing fails at all. Every job finishes and every status is green, but a number is wrong. A column changed type upstream, or a source sent half its usual rows. This shape is worse, because the first person to notice is often someone who uses the data, not someone who runs it.

Then the questions arrive. Is the dashboard safe to use? Should the weekly report wait? Did last night's export go out with the bad figures? One engineer, often whoever is on call, stops what they were doing and starts to dig.

Where the time goes

09:10 to 10:25

The fix itself is often small: a cast, a filter, a rerun. What takes the time is everything before the fix, which is working out what happened.

One failure, traced by hand. Our estimate for a mid-sized team
StepTakesBy
Find which run failed, and when it last worked10 min09:20
Read the error and the logs around it10 min09:30
Work out whether the data changed or the code did15 min09:45
Trace upstream through lineage to where it started15 min10:00
Check recent commits and deploys10 min10:10
Ask the people who know the business (Part of this is waiting for a reply)15 min10:25
Total, estimated1 h 15 min10:25
The full account, 3 more paragraphs

Here is that work as a list, with a time against each step.

The times in that list are our judgement for a mid-sized team, and nobody timed them. Yours will differ.

Two things stand out. Almost none of the work is typing a fix. And the steps are not hard, they are scattered. The run history is in one tool, the logs in another, the lineage in a third if it exists at all, and the commits in a fourth. The reason a column means what it means is in somebody's head.

Who is waiting

08:50 to 10:40

While one person digs, other people wait.

Analysts
either stop work, or carry on and risk building on a wrong number.
Dashboards
keep showing stale or wrong figures to everyone who opens them. Most of those people never see the chat thread.
Downstream jobs
fail in turn, or run on bad input and pass it along.
A customer-facing number
is sometimes involved: a usage figure, an invoice line, a report that goes out on a schedule.
The full account, 1 more paragraph

The engineer's time is the part you can see. The waiting is wider, and it grows for as long as the cause stays unknown.

Why the cost stays hidden

after 10:40

Nobody sends an invoice for a data incident. The cost is paid in engineer hours and in decisions that were made late, or made on a wrong number. Neither appears as a line in a budget.

The full account, 2 more paragraphs

It hides because each incident looks small. No single one seems worth a meeting. The work that was planned for that morning slips quietly, and the slip is blamed on the plan.

It also hides because it is rarely written down. Once the job is green again, everyone goes back to work. The cause lives in a chat thread that nobody will find next time, so next time starts from nothing.

What reduces it without buying anything

Most of the time in the list above goes on finding context that was never written down or kept. A good part of that context can be put in place with habits, and with tools you most likely have already.

The full account, 1 more paragraph

None of this stops incidents from happening. It shortens the time each one takes.

Where Convalesce fits

Convalesce does the digging. When a run fails, it investigates the failure: the run, the error, what changed, and what sits upstream and downstream. It shows the evidence it found, so you can check the reasoning yourself. Then it proposes a fix for a person to approve, and nothing ships without that approval. The steps are set out in how it works.

The full account, 1 more paragraph

The next posts in this series each take one failure and follow it from the first red mark to the fix.

Stop reconstructing failures.

Sign in with GitHub or Google, connect a tool, and see your next failed run explained.