All posts
MeasurementAutomatic timesheetsAccuracyOn-device matchingProduct engineering

Automatic time tracking accuracy: cutting wrong bookings by two thirds

We published the thresholds our automatic timesheet used, then measured them. Eight days later they are gone: the app now calculates them from your workspace.

TO
The TimerOS Team
Vezoft
9 min read

A week ago we published the thresholds behind automatic task matching: a score of 0.45, a margin of 0.15 and sixty seconds of agreement. Publishing them felt like the open thing to do. Then we measured what those numbers did on real workspaces and found that a threshold we picked once, in our office, cannot be right for your task list.

Since then we have rebuilt the matcher by asking one question of every number in it: how do you know that? Most of the fixed numbers are gone. The bar a booking has to clear is now calculated from your own projects, every few minutes, on your own machine. This post covers the measurement: what we count, what changed, and the one rule that survived every rewrite.

What this covers
Task matching is still marked Beta in the desktop app. It runs only in the Windows desktop app (the web dashboard has no matcher). It is on by default and can be switched off. Since v1.15.0, time-and-materials invoices bill the tracked hours on each task, so review matched time before you invoice. How the original matcher works is described separately; this post assumes you have read it.

What “accurate” has to mean before you can measure it

“Accurate” means nothing here until you say which mistake you are counting. An automatic timesheet can make two, and they cost very different amounts. It can miss work it should have recognised, and you notice at once because the row is empty. Or it can file time against a task that had nothing to do with it. That row looks exactly like a correct one, so you will not notice it at all.

A wrong booking is more expensive than a missing one, because you cannot see it.
— The same constraint the first version was built around, now with a number attached.

So the number we report is for the second kind. The test feeds the matcher work from a project the workspace does not have and counts how often it books anyway. Every booking in this test is wrong by definition, so the scoring needs no judgement.

Figure 1 · Leave one project out
1 · The window
Ship the iOS beta to TestFlight — Google Chrome
A real task title, with the extra text a real window title carries.
2 · The workspace
iOS App DevelopmentWebsite Design & BuildCloud Migration+ 17 more
The project that owns the title is removed. Everything else stays.
3 · The right answer
Say nothing
The work is not in this workspace. Anything booked here counts as a false positive.
The test runs once for each task title in the twenty project plans that ship with TimerOS, 481 windows in all. It works because the right answer is known in advance and is always the same: book nothing.

A second run does the opposite and matters just as much: each of the 481 titles is fed to a workspace that does contain it. Here, saying nothing is the failure. A matcher could get a perfect false-positive rate by never speaking at all, so the two runs only mean something together. On today’s build the second run recognises 453 of the 481, says nothing on 28, and names the wrong project zero times. The test suite fails if that zero ever changes.

The number, and how it fell

Figure 2 · Wrong bookings, as a share of 481 windows describing absent work
False bookings% of windows
16.8
11.4
5.6
v1.6.7v1.6.8v1.8.0
Lower is better. Each step is a named change with a test that holds it in place. The counts are written into the code as limits that may only go down.

Two thirds of the wrong bookings are gone. The first step cost no correct match, and after the second the matcher still recognises 98.7% of the work it recognised before. Figure 3 breaks the last run down. Note the middle slice: windows where the app was sure enough to ask, but not sure enough to file an hour.

Figure 3 · All 481 windows describing work the workspace does not have
481windows
Outcome per window
Said nothing (correct) · 43791%
Offered a wrong project · 174%
Booked the wrong task · 276%
Measured on the released build. Each of the 17 wrong suggestions costs a glance and a click to dismiss. The 27 wrong bookings are the number that matters, because nobody sees them.

Here are the three changes. The first two are the steps in Figure 2: v1.6.8 and v1.8.0. The third came in between, in v1.7.4, and made the matcher better at recognising real work.

1. Refuse to book a window the task does not explain

The old check asked whether a task scored highly enough. It never asked the reverse: does this task explain the window? On a page with nine words, a task that matched two of them could clear every threshold on those two alone. The other seven, which actually said what the page was about, were never checked.

A booking now has to pass two checks. At least half the meaningful words on screen must be words the task recognises, and at least 30% of the task’s own words must appear in the window. If either check fails, nothing is booked. The task is not discarded, though: it is offered as a one-click suggestion instead. Of the wrong bookings this removed, most disappeared entirely and a few came back as questions.

2. Let the workspace set its own bar

This is the change that makes the earlier post out of date. A threshold of 0.45 is an assumption about how easily your projects can be mixed up, and we have no way of knowing that. Take two agencies: one runs twenty unrelated client projects, the other four near-identical rebuilds. They need different numbers, and no amount of care in our office would produce the right one for the second.

So the desktop app works it out. Every few minutes it re-scores a fixed sample of your own task titles against your own projects and asks two questions: what does a right answer look like here, and what is the best wrong answer this workspace can produce? The booking bar is set above that best wrong answer, with a wider safety margin when there are few examples to learn from.

Figure 4 · Where a threshold comes from
Picked once, by usDerived from your workspace
What sets itA judgement call, made on our task namesYour task titles, re-scored against your projects
When projects overlapSame bar; more wrong bookingsBar rises above the best wrong answer available
When it cannot be worked outn/a: the number is always thereFalls back to the default number and records why
If words cannot separate the workBooks something anywayBooking switches off; everything becomes a suggestion
The limits on how much of a window a task must explain come from the workspace’s weakest right answers, not from its wrong ones. The rule: never demand more explanation than your own correct matches actually give.

That last row is the part we care about most. If the tenth percentile of a workspace’s right answers sits below the ninetieth percentile of its best wrong ones, words simply cannot tell that workspace’s projects apart. The app then says so and stops booking. The timer keeps running. Everything becomes a one-click suggestion, and if you have switched on the session record described below, the reason is written there.

3. Weigh a shared word instead of deleting it

An earlier fix simply removed every word that appeared in two projects, on the grounds that a shared word cannot tell projects apart. That was too blunt, and it took a while to see why. In a workspace with projects called Website Design and Android App, the fix rightly removed the company’s own name. It also removed design, which turned up in the Android project too, and so Website Design lost a word from its own name.

Words are now scored by how much better than chance they point to one project, and chance depends on how many projects you have. With two projects, a word split evenly between them is a coin toss and tells you nothing, yet the old scoring treated it as half-useful. The new scoring raised recognition of real work.

0
windows in the false-positive run
0.0%
of them still book (was 16.8%)
0
wrong bookings on work the workspace has
0 days
of local session record, if you turn it on

The rule that survived every rewrite

The original matcher read only the window title. It now reads more: the browser’s domain and the first part of the path, the application name, and any domains and keywords you have added to a project. It looks like nothing more than extra evidence, yet it is the riskiest change in the whole series, because those signals identify a place, not a piece of work. Being on your client’s domain tells you which project you are in. It tells you nothing about which task.

Location can produce a question. Only the words in the window can produce a booking.
— Written into the code as a rule that must always hold, and tested.
Figure 5 · What each channel is allowed to do
You pinned a task
A deliberate human statement
may book hours
Words in the window title
“Fix the invoice sync” vs your task list
may book hours
A task open inside TimerOS
The dashboard says what is on screen
may book hours
The address bar, or the app
play.google.com /console · figma.exe
may only ask
Resemblance to the project
On-device similarity, no shared word
may only ask
None of the above
Nothing recognised this window
stays unassigned
Read from top to bottom: the first channel that answers wins. A channel in the lower half can put a one-click suggestion on the timer bar and choose what it offers first, but it can never file an hour by itself.

We made this a hard rule because we broke it twice, and both times it cost real hours. The first time, a project was named after a product that puts its name in every one of its tab titles. The name only marked a place, yet it counted as the second “independent” word that a confident booking needs. The second time, an association learned from someone’s clicks (two confirmations on github.com) started booking their personal repository browsing. Learned signals no longer book anything. They only set the order of a question you still have to answer.

One word is not proof

Nineteen hours of real use made the case better than any argument could. In that time the timer asked 379 questions. Of those, 264 were triggered by a single ordinary word: monitor on a status-monitoring site, console on a scheduling service, model on GitHub, links in a search console. All of them offered the same unrelated project.

The logic behind them was not obviously wrong. A word that appears in only one of your tasks is rare, and rare looks like evidence; that is why ticket numbers work. But a word can be rare in your task list and common everywhere else. Since v1.7.3, a question that names a project needs two different words. The cases that already worked had two all along; the 264 did not. A booking can still rest on a single word that appears in just one of your tasks. That word has to be in the window title, and the booking must still clear your workspace’s bar and both checks from the first change. This kind of booking has a known gap, listed at the end of this post.

Why the app has a session record
None of this could be learned from a bug report. The desktop app can keep a local record of what it saw, what it decided and why, and what you did next. Since v1.19.2 the record is off until you switch it on yourself, under Settings → Activity Detection → Diagnostic recording, and your employer cannot switch it on for you. It covers at most 30 days, is cleaned up automatically and is never sent anywhere. It turns “it feels wrong sometimes” into 264 of 379, which is a problem we can fix.

When no word matches a task

Many windows share no word with any task. A page can clearly be about your Android work and still contain nothing your task list mentions. Every mechanism above will then stay silent, which is correct. The last release in this series added a channel for exactly that case.

The desktop app now includes a small open-source language model, MIT-licensed and bundled in the installer, that turns text into a vector. Each of your projects gets a vector built from its task titles, the page on screen gets one too, and the app compares them. When a page is close to one project and clearly farther from the rest, the app suggests that project. It is roughly 17 MB of tables and a dot product: no server, no GPU, and no request leaves the machine.

The limits around it are what make it usable, and every one was decided before release. It never books. It never overrides a project the words already named. It speaks at most once an hour about any one website or application. It works exactly where nothing else can, so nothing else can check it; the hourly limit is the check. Across all 190 pairings of the built-in project plans, 99.9% of the projects it suggested this way were right, and it showed a wrong suggestion on 0.4% of pages about work the workspace does not contain.

A note on terms, since we are strict about this elsewhere: the token matcher is not AI. It is a fixed scoring function that always gives the same result for the same input, in both Automatic and Rule-based tracking modes. The resemblance channel is a model, and it runs only in Automatic mode, so in Rule-based mode no model takes part in matching.

What it still gets wrong

  • 5.6% is not zero, and no single threshold will close the gap. In the test with the project removed, the windows that still book describe work that genuinely resembles work you have, and no limit can separate them without also refusing correct answers. The next step is a bigger vocabulary for explaining windows.
  • A booking still ignores location. On a page inside your client’s console, the one-word booking described above can file hours against a task from another project, because the part of the matcher that books was deliberately never taught to read the address bar. It is a known gap, and a failing test already tracks it.
  • Nothing is booked in a window’s first minute. A booking needs sixty seconds of consistent evidence; a question needs ten. That is deliberate, and it does feel slow.
  • Declared signals are the cheapest fix, and nothing asks you for them. Telling a project which domains and applications belong to it is the most reliable way for the timer to know which project you are in, yet the setting is a project field most people never find. That is a product failure rather than a matcher one, and it is the item on this list we find hardest to defend.
Where to find it
Update the Windows desktop app to v1.9.0 or later. Matching is on by default and can be switched off in Settings. The timer bar shows what the app thinks you are working on, and the Task time panel on the Performance page is where you review, reassign or unassign anything. Included on every plan, from Freelancer up. Download TimerOS for Windows or start a trial.
Send us the window that fooled it
Our test plans cannot contain your task names, so your real screens catch what we miss. If the timer books to the wrong task, or misses work it should obviously have caught, the window title and the task name are enough for us to reproduce it: [email protected].

If we could keep only one lesson from the week, it would be this: a threshold nobody can explain is a guess with a decimal point. Ours can now be explained. It is calculated and written down, and when a workspace cannot support one, the app says so.

See it on your own machine

Try it free for 14 days: add a card at signup, and nothing is charged if you cancel before the trial ends. Install the desktop app and watch it classify your day.