custom AI agents

Custom agents for the jobs only your business has

Research agents, reporting agents, knowledge brains. Scoped to one workflow, wired into where the work already happens. 5x ROI in 30 days. Or we work for free.

Teams we build for

  • Hoyes Michalos
  • Nurse Next Door
  • Fedi
  • UBC Sauder
  • Merchant House Capital
  • Picton Investments
  • Campbell Froh May & Rice LLP
  • Barnakl
  • Hungerford
  • Breez

The problems this solves

01

You evaluated six tools. None fit.

Your workflow is specific to you. Off-the-shelf software keeps asking you to change it.

02

Your knowledge lives in inboxes.

Every answer the firm has ever produced exists somewhere. Finding it takes longer than redoing it.

03

The work is judgment plus assembly, and assembly wins.

Your best people spend their hours collecting, formatting and re-keying before they get to think.

04

Vendor AI is a black box.

You can't see why it answered, can't change how it behaves, can't take it with you.

05

Pilots demo well and die quietly.

The proof of concept worked. Nobody wired it into the real workflow, so nobody used it.

What we build

01 · Scoping

One workflow, chosen because senior hours meet assembly work there. Not a platform. A job.

02 · DOE architecture

Directives your team can read and edit in plain text. Orchestration that decides what runs when. Execution tools that do the work. You can see all three layers.

03 · Integration

The agent lives inside the tools your team already uses. No new destination, no new login, no adoption cliff.

04 · Evaluation

Tested against real cases with your people judging output before anything deploys.

05 · Training

Your team learns to run it, tune it and extend it. The directives are theirs.

What we won't build: agents that make final calls on legal positions, investment decisions, or anything your regulator expects a human to sign. Agents assemble, retrieve, draft and monitor. People decide.

What's included

  • One scoped agent, built on DOE, deployed in your stack
  • An evaluation harness with your real cases
  • Plain-text directives your team owns and edits
  • Training, drift monitoring and iteration. Cancel anytime

Weeks, not quarters.

How the full engagement works

01 Discovery

WEEK 0

02 Prototype

WEEKS 1-3

03 Deploy & train

WEEKS 4-8

04 Run & improve

WEEK 9+

How we build it

Step 01

One workflow, binary success criteria.

The agents that survive contact with production are scoped to a single workflow where success is unambiguous. Broad assistants are the ones that quietly get abandoned.

Step 02

Gate on evals before autonomy.

An agent earns autonomy by clearing a graded set, not by seeming impressive in a demo. Published pilot-to-production rates are the reason this gate exists.

Step 03

Widen scope only from measured ground.

New skills are added one at a time, each with its own evals. Scope creep without measurement is how a working agent becomes an unreliable one.

What we will not automate

Building an agent where a script would do, and shipping anything customer-facing without a human gate and an eval suite. If a vendor cannot offer no-training terms, it does not get into the stack.

How it is built

See the full stack

Retrieval

Agents work from your documents and data rather than general knowledge, and every answer carries its source. Retrieval quality is measured separately from generation, and before it.

  • pgvector
  • Full-text BM25
  • Reciprocal rank fusion

Agents and orchestration

Each agent has a defined skill set, an autonomy level, a spend ceiling and a promotion gate it must clear before acting without review. Tool access goes through one protocol rather than one-off integrations.

  • agent-worker
  • Autonomy guard
  • Promotion gate

Evaluation

An agent that has not been evaluated is a demo. Graded sets run before a change ships and again after, and the harness resolves which evals a change actually touches.

  • Eval graders
  • Regression targeting
  • quality-worker

Models

Model per job, not one model everywhere. We default to frontier reasoning for tool use and long documents, and reach for something smaller or a plain classifier when it benchmarks better on the specific task.

  • Frontier models, one gateway, routed per task
  • Embedding models via the same gateway
  • Task-specific classifiers

Data boundary

Memory and access control at the database layer, including row-level scoping where data must not cross teams. Agents inherit the permissions of the person they act for.

  • Supabase row-level security
  • Scoped agent memory
  • No-training API terms

Named tools in this build: Frontier models, routed per task · MCP · Agent runtime · Supabase

FIG. 01
Custom agents: how the system fits togetherYour documents, Your systems, Team knowledge feed a hybrid index. The agent runtime works from that index, and every consequential action passes a human approval before it reaches Your AI team, Ask your team, Approval queue.Your documentsYour systemsTeam knowledgeHybridindexAgentruntimeYourapprovalYour AI teamAsk your teamApproval queueYour systemsWhere your team works
Custom agents: how the system fits together.REV 2026.08

What that means in practice

Agents and orchestration

AI that does the work instead of just answering: looks things up, calls your systems, completes multi-step tasks, and knows when to hand off to a human.

Where we stop. Multi-agent swarms are oversold; most jobs need one well-guarded loop. If a cron job and a script solve it, that is what we build, because 90 percent per-step accuracy compounds to 59 percent over five chained steps and no framework changes that arithmetic.

How we use it

Retrieval and RAG

Connecting AI to your actual documents so it answers from your knowledge, accurately and with citations, instead of making things up.

Where we stop. A dedicated vector database is justified by scale, not by default. Most of RAG quality is won or lost in chunking and indexing strategy, not in the model choice, and we have walked clients back from RAG to plain search when that was the honest answer.

How we use it

Evaluation and observability

How we prove the AI actually works: measured, monitored and regression-tested like real software, not vibes.

Where we stop. There is no engagement where we skip this. The honest variable is depth: a document pipeline gets faithfulness and extraction suites, an outbound agent gets human review sampling, a classifier gets a held-out test set. We size the harness to the risk, never to zero.

How we use it

Questions? Straight answers.

Three ways in. No long discovery.

Free · 3-5 days · No obligation

AI Opportunity Audit

We map your operations, find the highest-ROI automations, and hand you a ranked plan with payback math. Yours to keep, whoever builds it.

No obligation. No follow-up sequence.

Paid · Fixed scope

First Build

One high-ROI system, built on your real data and deployed in your stack, with your team trained to run it. Fixed scope. Quoted after the audit. Covered by the 5x ROI guarantee.

Scope my first build

Paid up front. Cancel anytime after.

Free · 15 min

Intro Call

Fifteen minutes with James, not a sales rep. Bring your worst bottleneck, leave with a straight answer.

Book 15 minutes

No pitch deck.

Every first build is covered: 5x ROI in 30 days. Or we work for free. Read the full terms

Start free. Know your number in five days.

A 3 to 5 day audit of your operations, ending in a plan with the ROI math attached. No obligation.

5x ROI in 30 days. Or we work for free.