Skip to content
Recorded

All notes  /  Billing

Estimates Against Actuals

The most valuable thing time data produces, available free once actuals exist, and almost never used because the comparison is uncomfortable.

Procedure

An organisation that records time and never compares it with what was estimated has collected the data and skipped the finding.

Why it is the best use of the data

It is checkable. The estimate was written down; the actual is recorded.

It compounds. Each completed piece of work improves the next estimate.

It is about the work, not the person, which keeps it usable.

It changes decisions: pricing, staffing, whether to take the work at all.

And it costs nothing beyond looking, since both numbers already exist.

Doing it properly

Record the estimate at the point of commitment, with a date, and do not revise it silently.

Where it is revised, keep both and record why.

Compare at completion, not at the end of a period.

Express as a ratio, not a difference: a project at 1.4 is comparable across sizes in a way that "twelve days over" is not.

Segment by work type, which is where the pattern is.

What the pattern usually looks like

Systematic underestimation, consistently, across most organisations and most fields.

Worse for unfamiliar work, which is unsurprising and worth quantifying rather than assuming.

Worse for work with many dependencies, where the waiting is not estimated at all.

Better for repeated work, which is the argument for categorising rather than treating every piece as unique.

Large variance rather than a stable bias in some categories, which is a different problem: a stable factor can be applied, a wide spread cannot.

Using it

Apply the measured factor to new estimates for the same category. This is unpopular and it works.

Report the spread as well as the factor, because a category with a factor of 1.3 and a range from 0.8 to 2.5 does not support a confident price.

Where the spread is wide, change how the work is set up rather than adding contingency indefinitely.

Feed it into pricing, which is where the value is realised.

What makes it fail

Estimates not recorded, or recorded in a document nobody can find at completion.

Silent revision, which makes every comparison flattering.

Comparing at the wrong boundary: an estimate covering design compared against an actual including delivery.

Treating it as accountability. The moment the comparison is used to assess the estimator, estimates inflate and the calibration data is lost.

Which is the central discipline here: the purpose is to learn the factor, not to find out whose estimates were worst.

Reporting it

Ratio by work category, with the spread, updated as work completes.

The number of completed instances behind each category, since three is not a basis.

The trend, which should move toward one as the factor gets applied.

And the estimates that were right, which people remember less and which matter for knowing where the process already works.

Never use it for accountability

The discipline that determines whether calibration data exists at all.

The moment estimate-to-actual is used to assess the estimator, estimates inflate.

And the factor becomes meaningless, because it now measures padding rather than difficulty.

Say explicitly that it is not used that way, and hold it.

Report the factor by category, never by person.

Where an individual genuinely estimates badly, it shows in the category data with their work in it — and it is a coaching conversation, not a metric.

Reporting the spread

A factor without a range is a false precision.

A category at 1.3 with a range from 1.1 to 1.5 supports a confident price.

The same factor with a range from 0.8 to 2.5 does not, and applying it produces a price that is wrong half the time in each direction.

Report both, always.

Where the spread is wide, the fix is in how the work is set up — scoping, dependencies, definition — rather than in adding contingency indefinitely.

And report the count behind each category, since three instances support nothing.

Check the difficult case

Use the design-team scenario to frame one representative test for this issue. The useful evidence is not the marketing description itself but the record created when a worker corrects an entry, a manager reviews it and an administrator exports it.