Research
We publish what we measure, including the parts that do not flatter us
Most writing about enterprise AI is opinion with a statistic attached. We run AI systems in production for real firms, so we can publish observations instead. Every study here states its method, its sample and what it cannot tell you.
Being counted now
Listed before publication on purpose. Announcing what is being measured, before the result is known, is the part that makes the result worth reading.
The agent failure index
Every production failure Advizr has recorded, with its root cause, its severity and how often it came back. Published work on agent failure studies coding assistants in an IDE or frameworks in a lab. This is business automation running against real client data.
Model performance by task area
Which models clear a written rubric on real business tasks, at what cost per passing case, ranked by the lower bound of the confidence interval rather than the headline rate.
Retrieval at scale
The corpus size at which a production retrieval system stopped returning correct results without raising an error, what caused it, and what the fix cost.
How we publish
Every number states its sample. Where a rate comes from a small sample we publish the interval, not the point estimate, because a rate from 30 cases and a rate from 3,000 are different claims that look identical once rounded.
Client data is anonymized unless the client has approved being named in writing. An anonymized descriptor is a real engagement, never a composite of several.
Corrections are versioned rather than silently edited. If a number changes, the old one stays visible with the date it was replaced.
Advizr sells AI systems, which is a commercial interest in every subject here. The limits sections are written to be usable against us.