Research
We publish what we measure, including the parts that do not flatter us
Most writing about enterprise AI is opinion with a statistic attached. We run AI systems in production for real firms, so we can publish observations instead. Every study here states its method, its sample and what it cannot tell you.
Published
The Agent Failure Index
Version 1.0
Every production failure Advizr has recorded, with its root cause, the signal it gave at the time, and whether it was ever closed. Published work on agent failure studies coding assistants in an IDE or frameworks in a lab. This is business automation running against real client data.
359 failures · 5 systems · 68 days
The AI Evidence Index
Version 1.0
The enterprise AI debate runs on about a dozen numbers, and almost nobody quotes them correctly. What each study measured, its sample, and how it gets restated wrong.
93 studies · 77 publishers · 2024 to 2026
Being counted now
Listed before publication on purpose. Announcing what is being measured, before the result is known, is the part that makes the result worth reading.
Model performance by task area
Which models clear a written rubric on real business tasks, at what cost per passing case, ranked by the lower bound of the confidence interval rather than the headline rate.
Retrieval at scale
The corpus size at which a production retrieval system stopped returning correct results without raising an error, what caused it, and what the fix cost.
How we publish
Every number states its sample. Where a rate comes from a small sample we publish the interval, not the point estimate, because a rate from 30 cases and a rate from 3,000 are different claims that look identical once rounded.
Client data is anonymized unless the client has approved being named in writing. An anonymized descriptor is a real engagement, never a composite of several.
Corrections are versioned rather than silently edited. If a number changes, the old one stays visible with the date it was replaced.
Advizr sells AI systems, which is a commercial interest in every subject here. The limits sections are written to be usable against us.