James BoothSep 25, 20263 min read
Every sentence the model wrote was true
On 21 August a person working in our own workspace asked the assistant to take their team to the brain page. The assistant replied that it had directed them there, and printed the link. Nobody moved. They had asked four times over five days, in four different phrasings, and got four fluent accounts of something that never happened.
The easy story is that the model made it up. It did not. The tool it needed was not there.
A tool cut by 38 tokens
On every turn, an assistant is handed a set of tools, and each tool is described in a few hundred tokens. There is a ceiling on how much of that description fits. When too many tools are connected, the most expensive ones are dropped until the rest fit.
That turn, 168 tools were measured against a 24,000-token ceiling and 84 were dropped. The navigation tool, at 242 tokens, was one of them. It missed the cut by about 38 tokens.
The prompt still told the model to use it. So the model did the closest thing it could, and printed the link the prompt had shown it.
It was never one tool
The next day the plan tool went missing on our Interact page. No plan tool means no plan card, and no plan card means no way to start background work. So we measured it properly. On our template workspace since 1 August, 29 of 191 chat turns had dropped a tool, and all 29 had lost one of our own.
The ceiling protected each connected app's main read and write actions by name. It did not protect ours. We had defended the vendors' tools and let our own fall.
Our client workspaces were fine. The template, with six connected apps, was the outlier.
The step that said it was done
Four days later the same workspace drew a step marked delivered, "Opened a page for you", over a turn that had opened nothing.
Every sentence the model wrote was true. It had told the person the builder was unavailable. The screen around those sentences was not true. When one of our tools refuses a request, it returns an error rather than throwing one. The code that drew each step only checked whether the tool had returned something, so every refusal on every chat surface had been drawn as a completed step.
That turn also showed the cut from a new side. 169 tools came to 69,237 tokens against the raised ceiling, and 54 were dropped, including the builder. The tools we meant to protect would have fit the whole time.
What we changed
A prompt may not name a tool the turn does not hold. We now build one list of every tool a turn has promised, from the page it is on, the plan mode and the agent's own instructions, and that list is protected from the ceiling. When a promised tool still goes missing, the system logs an error instead of going quiet.
We raised the ceiling from 24,000 to 32,000 tokens. Protecting more tools without more room would have cut every other tool at once.
A refusal now draws as a refusal. The live view and the saved view of a turn read the same rule, and a test pins that they agree.
We proved it by forcing the ceiling low and sending the same request twice. With the fix off, the builder was dropped and the reply apologised about the budget. With it on, the builder stayed and the card appeared.
What to ask any vendor
When your assistant says it did something, what checks that it did? What happens when a tool it needs is missing on a given turn? Can a refused action ever show as done?
Good answers come with a demo. The insights are free. If you want this level of engineering pointed at your operation, start with the free audit. The plan is yours to keep either way.