The Measures Worth Reporting
A short set that shows whether the system is working and whether the data can be trusted, and the popular ones that measure compliance instead.
Reference
A time tracking programme generates a great deal of countable activity. Only some of it says anything about whether the arrangement is working.
The eight
Recording delay: proportion of entries made within one working day.
Time spent recording, minutes per person per week.
Catch-all category usage, as a proportion.
Tidiness: proportion of daily totals landing exactly on a whole or half hour.
Unbilled effort as a proportion of total, trended.
Estimate against actual, for completed work, by category.
Realisation, where billing applies: recorded, billed, collected.
Hours beyond contracted, in aggregate, which is the welfare measure.
Each names a specific failure when it moves the wrong way.
What they tell you
Delay rising: friction has increased somewhere, or the visible use has stopped.
Recording time rising: the category list or the tool has become harder.
Catch-all rising: the structure no longer fits the work.
Tidiness rising: entries are being reconciled rather than recorded.
Unbilled effort falling: usually it stopped being recorded, not stopped happening. Check before celebrating.
Hours beyond contracted rising: a welfare and legal matter that the other seven will not surface.
The measures that mislead
Timesheet compliance or submission rate, which measures whether forms arrived.
Individual hours, utilisation or billable percentage, compared between people, on data with this error structure.
Total hours recorded, which rises with headcount and says nothing.
Number of projects tracked.
Any figure reported to a precision the recording method does not support.
Segmentation
By team, by project, by work type.
Never by individual as a headline, for the reasons throughout this collection: the data is not accurate enough, and the comparison measures the work assigned.
Where an individual pattern genuinely matters — someone consistently working excessive hours — that is a welfare conversation prompted by the aggregate, not a ranking.
The report
One page, monthly.
The eight numbers with their trend.
One sentence of interpretation on each, written by a person.
The data quality measures on the same page as the analysis, so a reader knows what weight to apply.
Sent to the people who record the time as well as to the people who read the analysis, which is the audience most often omitted and the one that determines whether any of it is accurate.
The annual view
What the data showed, what changed because of it, and what did not.
Whether the recording method changed, which affects comparability.
Whether the purpose limitation was tested and held.
Whether anyone would design it the same way again, which is the question that produces the useful decisions and which nobody asks unless the process asks it.
The interpretation line
One sentence per number, written by a person, that makes a report readable.
What moved, whether it is inside its normal range, and why if known.
What is being done, and by whom.
Updated when it changes rather than rewritten monthly.
Written by whoever owns the measure, which is also how you find out whether it has an owner.
Numbers without interpretation are read as noise, and readers supply their own explanation, which is usually wrong and frequently about people.
Segmenting without ranking
The distinction that keeps segmentation usable.
By team, by project, by work type: yes.
By individual as a headline: no, because the data is not accurate enough and the comparison measures the work assigned.
Where an individual pattern genuinely matters — someone consistently at sixty hours — that is a welfare conversation prompted by the aggregate.
The practical test: does the segmentation prompt a change to how work is allocated, or a judgement about a person? The first is the purpose; the second is the drift.
A concrete product reference
When translating this principle into a buying test, view a practical reference provides a concrete feature and workflow reference. Verify the relevant behaviour in a trial, retain the exported evidence and judge it against the purpose and limits described above.