James BoothOct 6, 20263 min read
One ask, one run
On 12 September someone in our template workspace asked the team for a deck. They got four runs across two agents, three approval cards for actions inside the app, a promise instead of a deck, and a build that died four minutes after it said done. The runs page showed four rows that did not look related.
Every piece of that was the product following its own rules. Twelve changes later, the same kind of ask finished in about 11 minutes with no approvals.
Approvals for things that could not hurt anyone
The approval rule sent anything without a known target to a person. A saved note, a design build and a hand-off to a teammate address nobody, so their target was never known and every one waited for a yes. Over seven days in that workspace, 39 hand-offs, 9 notes, 13 builds and 3 exports were each proposed rather than done.
The setting was also stricter than anyone had chosen. Every client workspace sat on "propose only" because of an old default, not a decision.
The rule now reads: a person is asked before an action that is externally consequential or irreversible, and nothing else.
A run that said it was done
The Chief of Staff handed step 1 of a seven-step plan to a teammate, got the report back, ended with a promise about the deck, and marked itself completed with six steps still open.
A run that runs out while it still owes steps now ends as failed, and names the first step it did not do.
A build that never stopped arriving
Writing the deck was one model call with no way to stop it. On the builds that finished it took between 79 and 193 seconds, against a 240-second deadline. When a build died, the conversation kept saying the deck was being made and would appear when it landed. It never did. 4 of the 22 builds in the previous two weeks ended that way.
The call now stops itself 15 seconds before the deadline, so a slow build fails inside its deadline instead of dying after it.
Four rows for one conversation
The only link from each run back to its chat was an optional field in its metadata. So the runs page showed four unrelated rows, each done by one agent, while the chat only ever named the Chief of Staff.
Every hand-off now records the conversation it came from. The runs page groups a conversation's runs into one entry that names everyone who worked on it.
Limits sized for a different problem
The limits on a run and its hand-offs were set in July to stop runaway loops, and the second hand-off of any real plan hit them. One of them was written for a run nobody is watching and was stopping a conversation somebody was watching. We raised the limits for a plan's hand-offs and gave a run started from a conversation a ceiling of its own.
Proving it on production
We then sent the same shape of ask through the live app and checked it end to end.
The third attempt delivered the deck with no approvals, and the test still failed, because the workspace's database was unreachable for eleven minutes along the way. The fourth attempt closed in ninety seconds by pointing at the deck from the earlier conversation, which was not in this one. So the plan now says a deliverable counts only when it is a file, a link or a card the person can open in this conversation.
The fifth attempt started at 02:28 UTC on 13 September. It produced three runs in one conversation, and all three completed. None asked for approval. The deck build was ready about three minutes after it started, and the whole ask took about 11 minutes from the first run to the last.
What to ask any vendor
When you ask for one thing, how many runs does it start, and can you see them as one? Which actions ask for your approval, and why those? When a run says it is done, what did it actually deliver?
Good answers come with a demo. The insights are free. If you want this level of engineering pointed at your operation, start with the free audit. The plan is yours to keep either way.