A week ago we published the thresholds behind automatic task matching: a score of 0.45, a margin of 0.15 and sixty seconds of agreement. Publishing them felt like the open thing to do. Then we measured what those numbers did on real workspaces and found that a threshold we picked once, in our office, cannot be right for your task list.
Since then we have rebuilt the matcher by asking one question of every number in it: how do you know that? Most of the fixed numbers are gone. The bar a booking has to clear is now calculated from your own projects, every few minutes, on your own machine. This post covers the measurement: what we count, what changed, and the one rule that survived every rewrite.
What “accurate” has to mean before you can measure it
“Accurate” means nothing here until you say which mistake you are counting. An automatic timesheet can make two, and they cost very different amounts. It can miss work it should have recognised, and you notice at once because the row is empty. Or it can file time against a task that had nothing to do with it. That row looks exactly like a correct one, so you will not notice it at all.
A wrong booking is more expensive than a missing one, because you cannot see it.
So the number we report is for the second kind. The test feeds the matcher work from a project the workspace does not have and counts how often it books anyway. Every booking in this test is wrong by definition, so the scoring needs no judgement.
A second run does the opposite and matters just as much: each of the 481 titles is fed to a workspace that does contain it. Here, saying nothing is the failure. A matcher could get a perfect false-positive rate by never speaking at all, so the two runs only mean something together. On today’s build the second run recognises 453 of the 481, says nothing on 28, and names the wrong project zero times. The test suite fails if that zero ever changes.
The number, and how it fell
Two thirds of the wrong bookings are gone. The first step cost no correct match, and after the second the matcher still recognises 98.7% of the work it recognised before. Figure 3 breaks the last run down. Note the middle slice: windows where the app was sure enough to ask, but not sure enough to file an hour.
Here are the three changes. The first two are the steps in Figure 2: v1.6.8 and v1.8.0. The third came in between, in v1.7.4, and made the matcher better at recognising real work.
1. Refuse to book a window the task does not explain
The old check asked whether a task scored highly enough. It never asked the reverse: does this task explain the window? On a page with nine words, a task that matched two of them could clear every threshold on those two alone. The other seven, which actually said what the page was about, were never checked.
A booking now has to pass two checks. At least half the meaningful words on screen must be words the task recognises, and at least 30% of the task’s own words must appear in the window. If either check fails, nothing is booked. The task is not discarded, though: it is offered as a one-click suggestion instead. Of the wrong bookings this removed, most disappeared entirely and a few came back as questions.
2. Let the workspace set its own bar
This is the change that makes the earlier post out of date. A threshold of 0.45 is an assumption about how easily your projects can be mixed up, and we have no way of knowing that. Take two agencies: one runs twenty unrelated client projects, the other four near-identical rebuilds. They need different numbers, and no amount of care in our office would produce the right one for the second.
So the desktop app works it out. Every few minutes it re-scores a fixed sample of your own task titles against your own projects and asks two questions: what does a right answer look like here, and what is the best wrong answer this workspace can produce? The booking bar is set above that best wrong answer, with a wider safety margin when there are few examples to learn from.
That last row is the part we care about most. If the tenth percentile of a workspace’s right answers sits below the ninetieth percentile of its best wrong ones, words simply cannot tell that workspace’s projects apart. The app then says so and stops booking. The timer keeps running. Everything becomes a one-click suggestion, and if you have switched on the session record described below, the reason is written there.
3. Weigh a shared word instead of deleting it
An earlier fix simply removed every word that appeared in two projects, on the grounds that a shared word cannot tell projects apart. That was too blunt, and it took a while to see why. In a workspace with projects called Website Design and Android App, the fix rightly removed the company’s own name. It also removed design, which turned up in the Android project too, and so Website Design lost a word from its own name.
Words are now scored by how much better than chance they point to one project, and chance depends on how many projects you have. With two projects, a word split evenly between them is a coin toss and tells you nothing, yet the old scoring treated it as half-useful. The new scoring raised recognition of real work.
The rule that survived every rewrite
The original matcher read only the window title. It now reads more: the browser’s domain and the first part of the path, the application name, and any domains and keywords you have added to a project. It looks like nothing more than extra evidence, yet it is the riskiest change in the whole series, because those signals identify a place, not a piece of work. Being on your client’s domain tells you which project you are in. It tells you nothing about which task.
Location can produce a question. Only the words in the window can produce a booking.
We made this a hard rule because we broke it twice, and both times it cost real hours. The first time, a project was named after a product that puts its name in every one of its tab titles. The name only marked a place, yet it counted as the second “independent” word that a confident booking needs. The second time, an association learned from someone’s clicks (two confirmations on github.com) started booking their personal repository browsing. Learned signals no longer book anything. They only set the order of a question you still have to answer.
One word is not proof
Nineteen hours of real use made the case better than any argument could. In that time the timer asked 379 questions. Of those, 264 were triggered by a single ordinary word: monitor on a status-monitoring site, console on a scheduling service, model on GitHub, links in a search console. All of them offered the same unrelated project.
The logic behind them was not obviously wrong. A word that appears in only one of your tasks is rare, and rare looks like evidence; that is why ticket numbers work. But a word can be rare in your task list and common everywhere else. Since v1.7.3, a question that names a project needs two different words. The cases that already worked had two all along; the 264 did not. A booking can still rest on a single word that appears in just one of your tasks. That word has to be in the window title, and the booking must still clear your workspace’s bar and both checks from the first change. This kind of booking has a known gap, listed at the end of this post.
When no word matches a task
Many windows share no word with any task. A page can clearly be about your Android work and still contain nothing your task list mentions. Every mechanism above will then stay silent, which is correct. The last release in this series added a channel for exactly that case.
The desktop app now includes a small open-source language model, MIT-licensed and bundled in the installer, that turns text into a vector. Each of your projects gets a vector built from its task titles, the page on screen gets one too, and the app compares them. When a page is close to one project and clearly farther from the rest, the app suggests that project. It is roughly 17 MB of tables and a dot product: no server, no GPU, and no request leaves the machine.
The limits around it are what make it usable, and every one was decided before release. It never books. It never overrides a project the words already named. It speaks at most once an hour about any one website or application. It works exactly where nothing else can, so nothing else can check it; the hourly limit is the check. Across all 190 pairings of the built-in project plans, 99.9% of the projects it suggested this way were right, and it showed a wrong suggestion on 0.4% of pages about work the workspace does not contain.
A note on terms, since we are strict about this elsewhere: the token matcher is not AI. It is a fixed scoring function that always gives the same result for the same input, in both Automatic and Rule-based tracking modes. The resemblance channel is a model, and it runs only in Automatic mode, so in Rule-based mode no model takes part in matching.
What it still gets wrong
- 5.6% is not zero, and no single threshold will close the gap. In the test with the project removed, the windows that still book describe work that genuinely resembles work you have, and no limit can separate them without also refusing correct answers. The next step is a bigger vocabulary for explaining windows.
- A booking still ignores location. On a page inside your client’s console, the one-word booking described above can file hours against a task from another project, because the part of the matcher that books was deliberately never taught to read the address bar. It is a known gap, and a failing test already tracks it.
- Nothing is booked in a window’s first minute. A booking needs sixty seconds of consistent evidence; a question needs ten. That is deliberate, and it does feel slow.
- Declared signals are the cheapest fix, and nothing asks you for them. Telling a project which domains and applications belong to it is the most reliable way for the timer to know which project you are in, yet the setting is a project field most people never find. That is a product failure rather than a matcher one, and it is the item on this list we find hardest to defend.
If we could keep only one lesson from the week, it would be this: a threshold nobody can explain is a guess with a decimal point. Ours can now be explained. It is calculated and written down, and when a workspace cannot support one, the app says so.