Skip to content
All insights

Sep 29, 20263 min read

A skip on a critical alarm is a silent failure

On 25 August, three questions to the brain in our template workspace failed within 23 seconds. The person asking got an error and a Retry button. About a minute after the third failure, our health sweep ran.

The check that watches the brain reported "No questions were asked recently, so there is nothing to check." The check that watches our AI provider account reported that the provider's billing API had refused it. Both checks are marked critical. Both returned skip. Nothing on the board was red.

1,791
skips in a row from one critical check, over 12 days

A skip raises no alarm

A health check can pass, fail, error or skip. Skip is meant for a check that had nothing to judge, so a skip raises no alarm.

The brain check had skipped 380 times in a row, back to 23 August. The credit check had skipped 1,791 times in a row, back to 13 August. A check that never decides looks exactly like a system that never breaks.

Two checks, two blind spots

The brain check sampled recent answers and counted only the ones that recorded whether their search step worked. A question that failed outright recorded nothing. The failures were the rows it threw away, and with nothing left to count it concluded nobody had asked.

The credit check asked the provider for the account balance. That endpoint only answers an administrator's key, and the application does not carry one. It was refused on every run.

So the credit check was already blind on 24 August, when our key reached its spend limit and every chat turn in the template workspace failed. The account still had money. The limit sat on the key, and a balance check would have called the account healthy even if it had worked.

Why the questions failed

The failed questions had a quiet cause of their own. One line of code discarded the provider's error. The answer stream then waited until a 20-second timeout fired and filed the failure as slow. Every failure of that kind was recorded as a timeout, and none as what it was. With the fix, the same failure is named in 1.2 seconds.

What we changed

The brain check now counts every question asked in its window, so it cannot report silence when there was traffic. The credit check reads the key's own remaining limit, which the application can see and which is the limit that actually stops work. A new check looks for the sentence a person sees when an answer fails, so it works even on a workspace that has not taken the new error handling yet.

The first real verdicts came the same night. The brain check passed at 22:20 UTC. The credit check passed at 23:09. The new check failed, because 5 of the last 9 questions had come back empty with no reason recorded. That was true.

Which zero

Skips still happen, and some are honest. In the last day both brain checks skipped 111 times, and the record agrees: nobody asked the brain a question in that time. What changed is that "nobody asked" and "everyone failed" are now two different answers.

Nothing in our system yet raises an alarm when a critical check keeps skipping for days. That is the next change. The rule behind it is short: a critical check that cannot decide has failed.

What to ask any vendor

When one of your health checks cannot run, what does your dashboard show? When did each critical check last return a real verdict? Which of your alarms would have stayed quiet through your last outage?

Good answers come with a demo. The insights are free. If you want this level of engineering pointed at your operation, start with the free audit. The plan is yours to keep either way.

5x ROI in 30 days. Or we work for free.