The misleading comfort of test count

Imagine a regression suite that has doubled in size. The dashboard shows more automated checks and the pipeline is green more often than it is red. Yet engineers routinely rerun failures, tests share a mutable account, and a failed build rarely tells anyone what broke. When a release decision approaches, nobody can clearly explain whether the most consequential business risks are protected.

Did confidence actually improve?

The answer is not necessarily no. Additional tests may have found defects and shortened feedback loops. But the count alone cannot tell us. It records how much automation exists, not whether the system produces dependable evidence for a decision. I prefer to evaluate an automation program by the quality of the signal it gives engineers about risks that matter.

Automation should produce an engineering signal

A script that drives an application is an implementation. Its useful output is information: which behavior was exercised, what result was expected, what was observed, and how strongly that observation bears on a product risk. A passing check should narrow uncertainty; a failing check should direct investigation.

That makes automation an engineering system with inputs, dependencies, oracles, and operating characteristics. Its value depends on more than whether a runner can execute it. If the preconditions are unstable, the assertion is shallow, or the failure has no diagnostic context, execution can succeed without giving the team much information.

This is also why confidence is always scoped. A reliable test can establish that a particular behavior held under particular conditions. It cannot certify every adjacent behavior. Google's Software Engineering at Google describes a test in terms of behavior, input, observable output, and a controlled environment; that framing is more useful than treating a test as a line item in a repository. [1]

Test count is descriptive, not sufficient

Test count is a legitimate inventory measure. It can help plan runtime, ownership, maintenance, and migration work. A sudden change in count may also reveal that a suite was split, removed, or expanded. The mistake is using it as a direct proxy for confidence.

Activity measure More decision-oriented question
Automated tests added Which important failure mode is now observable?
Executions completed How often did the same code produce a stable result?
Checks passing What risks were exercised, with what oracle and data?

Coverage percentages have a similar limitation. They can reveal unexercised code, but a numeric coverage value does not establish that assertions detect a consequential failure. Martin Fowler makes this distinction in his discussion of test coverage. [2] I would rather have a small, dependable set that checks a critical retry invariant than many checks that only confirm that a page loaded. In practice, we need a portfolio: risk coverage, reliability, feedback time, and diagnostic quality together.

Reliability before scale

An actionable regression signal requires repeatability. For the same relevant code and controlled starting state, a test should generally reach the same result. Tests should not depend on execution order, an account another worker can modify, an unbounded sleep, or an external service whose state is outside the test's control.

This does not mean every test can be perfectly hermetic. A contract test may need a real database; an end-to-end journey may need multiple services. The design question is where to isolate dependencies and where production fidelity is essential. Google's discussion of larger tests explicitly describes the tension between isolation and fidelity. [3] The right boundary follows the risk being tested.

Nondeterminism erodes the meaning of red. If engineers learn that rerunning often makes a failure disappear, they stop treating failures as evidence. A real regression can then be dismissed alongside noise. Fowler identifies isolation, asynchronous behavior, remote services, time, and resource leaks as common causes of non-deterministic tests. [4] Google's testing guidance likewise describes test data, setup, cleanup, and environmental assumptions as sources of flakiness. [5]

Reliability also requires observability. A failure should preserve the request or action, relevant identifiers, expected and actual state, and enough timing or trace information to locate the boundary that failed. A screenshot can help with a visual issue; it cannot replace an API response and state transition when the risk is financial integrity. Debuggability is part of the signal, not a convenience added after the suite grows.

Test data is part of automation architecture

One pattern I have encountered in automation programs is a shared account used as an invisible fixture. Test A changes its limit or balance. Test B assumes yesterday's value. Both pass in isolation, but parallel execution or a different order produces a random failure. The suite appears flaky even when the runner itself is working correctly.

The dependency chain is simple:

Shared mutable account → test changes state → another test assumes old state → intermittent failure

A stronger design makes the starting state explicit:

Test → deterministic fixture → known initial state → execution → verification → cleanup or disposable state

For a financial workflow, a fixture might create a uniquely identified account with a known available credit, valid eligibility, and no pending purchase. The test records those conditions before acting. It uses an idempotency key unique to that run, then verifies the intended state transition. At the end, it removes disposable data or lets an isolated test environment expire. If cleanup is impossible, the test should still own unique data rather than compete over a global record.

The fixture is not merely a helper that saves typing. It is part of the test's contract. Its setup can fail; its assumptions can drift; its cleanup can leak state. That is why fixture failures should be distinguishable from product failures. A pipeline that labels every setup error as a failed product check gives reviewers the wrong evidence.

There are trade-offs. Generating fresh data costs time. Full isolation may be expensive for systems with many real dependencies. A carefully partitioned pool can be appropriate if ownership, reset behavior, and concurrency are explicit. The important property is that state is controlled enough to reproduce and diagnose the outcome. Google's testing overview recommends tests that set up and tear down their own environment and avoid reliance on a shared database. [1]

Cover risks, not screens

The most useful layer for a test is the one that can expose the target failure with sufficient fidelity and a stable oracle. Consider the risk of a duplicate financial effect after a client retries a purchase. A browser journey can prove that a user can reach and submit the action. It may be a poor place to enumerate retry timing, duplicate callbacks, or idempotency-key behavior. API and integration tests can exercise those paths more directly, with controlled requests and observable state.

That is not an argument against UI tests. Critical user journeys still need end-to-end coverage, and UI-specific risks—accessibility, rendering, navigation, client validation, and browser behavior—must be checked at the appropriate layer. A UI test also has value when the integration between interface and backend is itself the risk. The point is to avoid forcing every scenario through the broadest layer merely because the user eventually sees its outcome.

Ham Vocke's Practical Test Pyramid argues for tests at different granularities and for avoiding redundant high-level checks when narrower tests already give the needed evidence. [6] I treat the pyramid as a cost and feedback model, not a fixed quota. For each risk, ask which boundary could fail, which layer can expose it, and what remains unverified after that test passes.

A green pipeline can still lie

Three plausible pipeline outcomes illustrate why color alone is insufficient:

  • Case A: Several checks fail because the product violates a business rule. Red is useful evidence and should block or inform the decision.
  • Case B: A few checks fail because fixture creation or a shared environment is unstable. Red deserves attention, but it does not yet demonstrate a product defect.
  • Case C: Every check passes, but no test exercises retry idempotency. Green is accurate for the executed suite and silent about that important risk.

The result must be read with its scope. Green is meaningful when the covered risks are relevant and the checks are trustworthy. Red is meaningful when a team can classify and investigate it. A pipeline can be operationally healthy and still have a coverage gap; it can also be noisy while a product change is correct.

Failure classification improves decisions

I use a simple first-pass classification: Product, Test, Fixture, Environment, Infrastructure, Unknown. Product means observed behavior contradicts the expected contract. Test means the test logic or oracle is wrong. Fixture means initial data or setup failed. Environment covers application configuration or dependent service conditions in the test environment. Infrastructure covers the runner, network, or CI substrate. Unknown remains explicit until evidence supports a category.

The distinction changes the next action. “52 failed tests” invites a rerun or a blanket release delay. The following hypothetical report is more useful:

Product:         3
Test:            0
Fixture:         1
Environment:    47
Infrastructure:  0
Unknown:         1

It suggests three product investigations, one fixture repair, a common environment incident, and one unresolved failure. The totals do not make the product defects less serious. They prevent a shared environment failure from being mistaken for 47 independent product regressions. Classification should retain links to evidence and allow revision; it should never be used to relabel inconvenient defects as “flaky.”

What I would measure instead

My preferred indicators are engineering aids, not universal KPIs. Their definitions should be visible, and trends should be interpreted with release risk and suite changes in mind:

  • Critical flow coverage: which named business invariants and failure modes have a test at an appropriate layer, and which remain untested.
  • Automation reliability: how consistently a check yields the same verdict for unchanged relevant code and controlled inputs.
  • Flaky failure rate: how often a check changes verdict without an explained product change; track affected tests and causes, not just a global percentage.
  • Fixture failure rate: how often setup, data allocation, or cleanup prevents a meaningful product assertion.
  • Unknown failure rate: how often investigation cannot yet identify the failing boundary.
  • Time to diagnose: elapsed time from failed run to a defensible classification and owner.
  • Regression execution time: feedback latency for a change, including queue and setup time when those dominate.

None of these numbers should become a target detached from behavior. A lower failure rate achieved by disabling critical tests is not improvement. A longer suite may be justified if it covers a previously invisible, high-impact risk. Review the underlying failures and uncovered risks alongside the trend.

Automation Enablement Before Automation Volume

My operating principle is Automation Enablement Before Automation Volume. Before asking a team for more scenarios, I look at what it costs to create one dependable scenario and to explain one failure.

Can tests authenticate without sharing mutable sessions? Can they establish controlled data and clean it up? Is there API support for setup and verification? Can the system expose state through a suitable public or test interface, and use database verification where architecture and risk justify it? Are UI selectors stable? Does CI preserve logs, traces, requests, and relevant artifacts? Is there a failure taxonomy and an owner for fixture problems?

If those foundations are missing, the constraint may be automation infrastructure rather than implementation capacity. Adding scripts will increase volume and maintenance before it increases confidence. Building reusable setup and observability may produce fewer visible tests in the short term, but it raises the value of every subsequent test. This is an engineering investment, not an excuse to postpone coverage indefinitely: a small set of high-risk checks can be built while the foundation improves.

A concrete financial example

Consider a fictional, generalized purchase flow. An account begins with available credit of 10,000,000 units. A purchase of 3,000,000 units is submitted. The API responds:

HTTP/1.1 200 OK
Content-Type: application/json

{"purchaseId":"P-1024","status":"SUCCESS"}

A weak test asserts only HTTP 200 and status == "SUCCESS". It passes. But the account's available credit remains 10,000,000. The expected value, assuming this operation should reserve or deduct the purchase amount immediately, is 7,000,000. The automation executed successfully; its oracle was incomplete. It observed an acknowledgement, not the full business effect.

A stronger check might combine:

Transport result
+ business operation state
+ financial transaction state
+ account state
= stronger evidence about the purchase invariant

The exact oracle depends on the product's contract. If settlement is asynchronous, the test should wait on a bounded, observable transition and distinguish pending from failed. If the architecture exposes a transaction query API, use it. If database state is the only authoritative signal in an isolated integration environment, a scoped database assertion may be appropriate. This does not mean every API test needs database access. It means the assertion must observe the effect that defines success for the risk being tested.

The same example also shows why test data and failure classification matter. A shared credit account could make the expected balance wrong before the test starts. A callback delay could make an immediate assertion premature. A missing transaction record could be a product defect, while an unavailable fixture service could prevent the test from reaching the product at all. A trustworthy check separates those possibilities instead of returning a generic red build.

The useful question

Reliable automation is not measured by one count or one green badge. It is built by choosing consequential risks, placing checks at useful layers, controlling their data, defining meaningful oracles, and making failures reproducible and diagnosable. Test count still helps us manage a suite; it cannot tell us whether the suite protects a decision.

The mature question is less “How many tests have we automated?” and more: How much trustworthy information does our automation give us about the risks we care about?

References

  1. Adam Bender, Software Engineering at Google — Testing Overview.
  2. Martin Fowler, Test Coverage.
  3. Software Engineering at Google — Larger Testing.
  4. Martin Fowler, Eradicating Non-Determinism in Tests.
  5. George Pirocanac, Google Testing Blog — Test Flakiness: One of the Main Challenges of Automated Testing (Part II).
  6. Ham Vocke, The Practical Test Pyramid.