Engineering

We deleted 40% of our tests and shipped faster

Most of them were asserting that the mock returned what we told the mock to return. Here is how we worked out which ones were load-bearing.

Divya Patel · 12 min read · 4 March 2026

Design

Why our empty states get the most review time

Arun Kulkarni · 6 min

Data

Column lineage without a metadata platform

Sara Bhatt · 9 min

Ops

The incident that taught us to page less

Tolu Okafor · 7 min

Engineering

We deleted 40% of our tests and shipped faster

Divya Patel · 12 min read · 4 March 2026

Our suite took nineteen minutes. Nobody ran it locally, which meant every branch discovered its failures in CI, which meant the feedback loop was measured in coffee breaks.

What we actually measured

Before deleting anything we instrumented the suite for six weeks: which tests had ever failed on a commit that a human then fixed. That is the only definition of a useful test we could defend.

1,847 tests. 312 had ever caught a real bug. The rest were, in the strictest sense, decoration.

Failures per test over six weeks. The long tail on the right never fired once.

The pattern in what we removed

Almost everything we deleted shared one shape: it mocked a dependency, called the unit, and asserted the mock was called. That test passes forever, including when the real dependency changes its contract.

// Deleted — asserts our own mock, not our behaviour
expect(mockRepo.save).toHaveBeenCalledWith(order);
“A test that cannot fail is not a test. It is a comment with a runtime cost.”

What we kept, and added

We kept every test that touched a real database, and we added more of them. Integration tests are slower per test and enormously cheaper per bug caught.

The suite now runs in four minutes and forty seconds. People run it locally again, which was the entire point.

What we got wrong

We deleted a set of date-handling tests that looked redundant. Three weeks later a timezone bug reached production. They had never failed because the code had always been right — which is not the same as the test being useless.

We put them back and refined the rule: never delete a test guarding a category of bug you know you are prone to, however quiet it has been.