AI for retail and ecommerce
AI for retail and ecommerce
Discovery is moving to AI assistants. Be the brand that answers, converts, and never oversells inventory. 5x ROI in 30 days, or we work free.
- DTC brands
- Multi-location retail
- Marketplace sellers
- Subscription commerce
Teams we build for
- Hoyes Michalos
- Nurse Next Door
- Fedi
- UBC Sauder
- Merchant House Capital
- Picton Investments
- Campbell Froh May & Rice LLP
- Barnakl
- Hungerford
- Breez
Teams we build for
- Hoyes Michalos
- Nurse Next Door
- Fedi
- UBC Sauder
- Merchant House Capital
- Picton Investments
- Campbell Froh May & Rice LLP
- Barnakl
- Hungerford
- Breez
Teams we build for
- Hoyes Michalos
- Nurse Next Door
- Fedi
- UBC Sauder
- Merchant House Capital
- Picton Investments
- Campbell Froh May & Rice LLP
- Barnakl
- Hungerford
- Breez
Proof, honestly
No published retail case study yet, and we say so. Founding clients get priority scheduling, direct founder involvement, and a co-published case study once the numbers are real. The 5x ROI guarantee applies from day one.
We guarantee 5x ROI inside 30 days of deployment, in writing, measured against a baseline you sign before we build. If the system misses the bar, we keep working for free until it clears.
Discovery is leaving Google.
Traffic to US retail sites referred from generative AI rose 693% over the 2025 holiday season, from a small base (Adobe Analytics, 2026). If your catalog is invisible to AI, you are invisible.
Service tickets scale faster than revenue.
Order status, returns and sizing questions, answered one at a time by people you cannot hire fast enough.
Forecasting is a spreadsheet and a feeling.
The buy commits cash before the season shows its hand. Overstock ties it up, and the stockout hands the sale to whoever has it in stock.
Every channel answers it differently.
The website says one thing, the store says another, and the helpdesk macro was last edited two agents ago. The customer hears whichever version picks up.
The size chart is a JPEG.
Support cannot quote it, assistants cannot parse it, and the fit questions keep coming anyway. The answer exists. It is just in a format nothing can read.
Peak scales the queue, not the team.
From Black Friday to the January returns wave, every order can become a ticket. Seasonal hires spend their first week reading the macro sheet, and by the time they are useful, peak is over.
Retail & ecommerce, before and after
The manual path is dashed: Invisible to AI assistants, Tickets answered one by one, Inventory by gut feel. The system path replaces it, and a person approves before anything ships: Catalog ready for AI search, Routine service runs itself, Forecasts from your own data.
Before: by hand
- 01Invisible to AI assistantshuman
- 02Tickets answered one by onehuman
- 03Inventory by gut feelhuman
After: the system
- 01Catalog ready for AI search
- 02Routine service runs itself
- 03Your approvalhuman
- 04Forecasts from your own data
The research
Retailers expected 19.3% of online sales to come back as returns in 2025, $849.9 billion across all channels (NRF and Happy Returns, 2025, US, large merchants).
National Retail Federation and Happy Returns · 2025 · 358 ecommerce professionals at merchants over $500 million in revenue and 2,006 consumers who returned an online purchase in the prior 12 months, fielded summer 2025, released October 15, 2025
The numbers in retail and ecommerce
Retailers expected 19.3% of online sales to come back as returns in 2025, $849.9 billion across all channels (NRF and Happy Returns, 2025, US, large merchants).
National Retail Federation and Happy Returns · 2025 · 358 ecommerce professionals at merchants over $500 million in revenue and 2,006 consumers who returned an online purchase in the prior 12 months, fielded summer 2025, released October 15, 2025
AI and agents influenced $262B, 20% of global online holiday sales in 2025. Retailers with their own agents grew 59% faster (Salesforce, 2025).
Salesforce · 2025
78% of small-business AI users feel pressure to adopt AI to keep up with competitors (PayPal, 2025).
PayPal · 2025 · 947 small businesses, May 2025
Run your numbers.
Your operations
Savings use the low end of our hours-reclaimed range. The math is conservative on purpose.
The math
Calculated at the low end of every range.
Every first build is covered in writing: 5x ROI in 30 days. Or we work for free.
What we build for retail and ecommerce
Answer the buyer
Brand AI agent for shopping and service.
Your catalog, your policies, your voice. Guided shopping and routine service handled, with escalation to people built in.
Order-status automation.
WISMO matched to the order record and the live carrier scan. The answer is the package's position, not a macro.
Returns and exchanges.
In-window requests get a label, edge cases get a person, and every refund waits for approval.
Get found
AI-search and agentic-commerce readiness.
Specs, variants and policies structured so assistants can parse, cite and recommend you, the same feed discipline Google Merchant Center already demands. The new storefront is someone else's chat window.
Marketing content engine.
Product copy, campaigns and lifecycle email in Klaviyo on a cadence, consent-tracked under CASL and human-reviewed before send.
Read the demand
Demand forecasting and allocation.
Forecasts from your own sales history instead of gut feel, refreshed on schedule, with drift flagged inside the reorder window.
Review and feedback mining.
Every review, ticket and survey distilled into what to fix and what to double down on.
Who this is built for
The drop no longer buries support
You plan launches around what the support desk can absorb, which means the queue is quietly setting your calendar. With the routine questions contained, the drop schedule answers to inventory and creative instead. Escalations still reach a person, arriving with the order and the transcript attached, and launch night stops ending with you in the helpdesk.
The brand voice survives the queue
The agent is graded on transcripts your best people already wrote, with tone scored alongside correctness. Containment only counts when the customer actually got their answer, because deflection that annoys a buyer is churn wearing a good metric. Your day moves to the conversations that genuinely need a person.
The catalog answers when assistants ask
Buyers increasingly ask an assistant, and the assistant reads structured data, not your homepage hero. Specs, variants, sizing and policies get published in a form assistants can parse and cite. When the answer window includes your product, it is because the data earned the slot.
You buy against a forecast, not a hunch
The buy commits cash months before the season answers back. Forecasts run from your own sales history, refresh on schedule, and flag the SKUs drifting from plan while the reorder window is still open. The call stays yours. The guesswork does not.
What stays human
- She puts the colorway live at ten
- She takes the escalations with context attached
- She approves the refund batch in one pass
The steps the day below leaves to a person, by design.
Drop day, as the CX lead runs it
She puts the colorway live at ten
humanThe new colorway goes live at 10am. She presses publish in Shopify, sends the launch email through the normal review flow, and watches the first orders land. Nothing about this moment is automated. It is her launch.
The first wave of questions lands on the agent
Sizing, ship dates and gift notes, answered from the product page's actual specs and the shipping policy, each answer citing both. The two questions retrieval cannot ground are escalated instead of improvised.
She takes the escalations with context attached
humanA wholesale inquiry and a damaged-item photo. Each arrives carrying the transcript and the order record, so neither customer repeats themselves and neither waits behind the sizing queue.
Where-is-my-order resolves against the carrier scan
WISMO tickets match to the order record and the live carrier scan pulled through ShipStation or your 3PL. The answer is the package's actual position, not a macro that says soon.
Returns are checked against the policy, not the mood
In-window requests get a return label. Final-sale items get a polite no with the policy line quoted. Anything ambiguous queues for a person, and no refund moves yet.
She approves the refund batch in one pass
humanEvery refund waits as a proposal with the order, the reason and the amount. She approves, edits or declines. Money moves only after her click, which is the boundary the whole build is designed around.
Today's transcripts become tomorrow's grading set
Closed conversations are sampled into the eval set and failures are flagged for review, with containment read next to resolution quality. The agent gets better on your tickets, not on someone else's.
The economics
Before and after economics
Line
- Routine ticket handling
Before
WISMO, sizing and returns answered one at a time by whoever is on the queue
After
The routine share is contained by the agent, and escalations arrive carrying the transcript and the order. Stated without a number on purpose
- Manual support hours in scope
Before
At the page defaults, 6 people spending 10 hours a week on manual handling is 60 hours, $2,400 a week at $40 loaded cost
After
The calculator above prices your own baseline, and the 5x ROI guarantee is measured against a baseline you sign before we build
- The buy
Before
A spreadsheet, last season's numbers and a feeling
After
Forecasts from your own sales history, refreshed on schedule, with drift flagged inside the reorder window
- Proof
Before
No published retail case study. This page says so instead of borrowing one
After
Founding clients co-publish the case once the numbers are real, and the guarantee applies from day one
One row carries dollars and it is arithmetic on this page's calculator defaults, 6 people at 10 manual hours a week each at a $40 loaded hourly cost, which is 60 hours or $2,400 a week in scope. Every other row is qualitative by design. No client outcome numbers exist for this page yet, and the founding offer above says so in plain terms rather than dressing another sector's case in retail clothes.
Where the data comes from
Ecommerce platform
Shopify, WooCommerce or BigCommerce holds the operating truth: orders, customers, inventory positions, discounts and fulfillment state. WISMO answers, return checks and refund proposals all read from here, so the agent speaks from the order record rather than from a template.
Where it stops. Read-mostly by design. The agent never edits a price, never creates a discount, and the refunds it proposes are approved by a person inside your own admin. Card data stays with the payment processor and never enters the build.
Catalog and PIM
The product truth, wherever it lives today: metafields on the store, a PIM like Akeneo or Salsify, or the spreadsheet the merchandiser guards. Specs, variants, materials, size charts and policies are indexed once, and every support answer and assistant-facing feed cites what the catalog actually says.
Where it stops. The index references the catalog, it does not rewrite it. Product data changes ship through your normal publishing flow with a person pressing publish, and nothing here trains anyone's model.
Helpdesk
Gorgias or Zendesk holds the corpus the agent learns your voice from: years of resolved tickets, the macros that work and the escalations that did not. The agent operates inside the helpdesk you already run, drafting, tagging and escalating with the transcript attached.
Where it stops. Customer conversations stay in your helpdesk account. A thread a person has taken over is theirs until they hand it back, and consent state is checked in the data layer before any outbound message goes anywhere.
How an answer is found in your catalog and policy
A question runs against your catalog and policy two ways at once: vector search for meaning and full-text search for exact wording. Reciprocal rank fusion merges both result sets, and the answer carries the source it came from.
- 01A question
- 02Vector search (meaning)
- 03Full-text search (exact wording)
- 04Rank fusion
- 05Answer, with its source
Built around your rules
| Regime | What it demands here | How the system complies |
|---|---|---|
| PIPEDA and Quebec Law 25 | The personal information here is order history, addresses and support conversations, handled with consent and purpose limits, and Law 25 adds its own obligations where you sell into Quebec. | Customer records stay in your own store and helpdesk accounts, access scoped at the database layer, under no-training API terms. |
| CASL | Email and SMS marketing need real consent, an identified sender and a working unsubscribe. | The consent ledger is part of the build and enforced in the data layer, so an agent cannot message a customer who opted out, and every send carries the sender and the unsubscribe. |
| Payments | Card data must be protected to PCI-DSS wherever it flows. | Card data never touches our systems. Payment flows stay inside Stripe or your processor, which publish their attestations. |
| Competition Act honesty | Marketing claims must be true, the advertised price must be the price, and reviews must be genuine. Drip pricing and fake urgency draw penalties. | AI-drafted marketing passes human review before it ships. No fabricated claims, no invented scarcity, no synthetic reviews. |
| Marketplace terms | Amazon and the other marketplaces set their own rules on buyer messages, response times and automation, and breaking them risks the account. | Autonomy is set per channel. Where a marketplace restricts automated replies, the agent drafts and a person sends inside the window. |
| GDPR and UK GDPR | Where you ship to EU or UK customers, their data rights travel with the order. | Handled in the same architecture, data in your accounts with access scoped and logged, and residency stated in writing per engagement. |
The regimes that govern retail & ecommerce, what each demands, and how the system complies
Regime
- PIPEDA and Quebec Law 25
What it demands here
The personal information here is order history, addresses and support conversations, handled with consent and purpose limits, and Law 25 adds its own obligations where you sell into Quebec.
How the system complies
Customer records stay in your own store and helpdesk accounts, access scoped at the database layer, under no-training API terms.
- CASL
What it demands here
Email and SMS marketing need real consent, an identified sender and a working unsubscribe.
How the system complies
The consent ledger is part of the build and enforced in the data layer, so an agent cannot message a customer who opted out, and every send carries the sender and the unsubscribe.
- Payments
What it demands here
Card data must be protected to PCI-DSS wherever it flows.
How the system complies
Card data never touches our systems. Payment flows stay inside Stripe or your processor, which publish their attestations.
- Competition Act honesty
What it demands here
Marketing claims must be true, the advertised price must be the price, and reviews must be genuine. Drip pricing and fake urgency draw penalties.
How the system complies
AI-drafted marketing passes human review before it ships. No fabricated claims, no invented scarcity, no synthetic reviews.
- Marketplace terms
What it demands here
Amazon and the other marketplaces set their own rules on buyer messages, response times and automation, and breaking them risks the account.
How the system complies
Autonomy is set per channel. Where a marketplace restricts automated replies, the agent drafts and a person sends inside the window.
- GDPR and UK GDPR
What it demands here
Where you ship to EU or UK customers, their data rights travel with the order.
How the system complies
Handled in the same architecture, data in your accounts with access scoped and logged, and residency stated in writing per engagement.
Customer data stays in your accounts, card data stays with your payment processor, and marketing automation is built CASL-first.
The objections
We ran a support chatbot for six months. Customers typed AGENT until we turned it off.
That bot matched keywords and guessed. This one retrieves from your catalog and policies, abstains when retrieval finds nothing, and is graded on resolved tickets your own team handled well before it ever faces a customer. Containment is read next to resolution quality, so a deflection that leaves the buyer angry counts against the system, not for it.
Our product data is a mess. The size chart is a JPEG and half the specs are wherever the founder last typed them.
Then that is where the build starts. The first phase reads what the store and the PIM actually hold and returns a gap list: the size chart nothing can parse, the variants with no attributes, the policy that exists only as a helpdesk macro. You see what an assistant can and cannot answer about your products before anything is built on top of it.
Peak is coming and the stack is frozen until January.
The build respects the freeze because it does not live in your theme. It reads the store and the helpdesk from outside, and the first phase, indexing and grading, touches nothing a customer sees. Run that phase during the freeze and the agent launches against real peak transcripts, which are the best eval set you will ever get.
Amazon suspends accounts over buyer messages. Automation there scares us.
It should, which is why the channel's terms set the autonomy. On your own store the agent answers. In Seller Central, where the rules are Amazon's, it drafts inside the response-time window and a person sends. Same system, different leash per channel.
An agent trying to make customers happy will hand out discounts all day.
It cannot. Discounts, refunds and credits are proposals, not actions, and the autonomy guard caps what the agent may do without review. Generosity stays a decision someone in your company makes, with the order and the margin in front of them.
How the system is built for retail and ecommerce
See the full capability mapRetrieval
Catalog, policy and past resolved tickets indexed together, so a support answer is grounded in your actual return policy and this product's actual spec. When retrieval finds nothing, the agent hands off rather than improvising.
- pgvector
- Full-text BM25
- Abstention on low confidence
Agents and orchestration
A brand agent handles the repeat questions end to end and escalates the rest with context attached. Refunds, credits and anything that moves money are proposals, not actions.
- agent-worker
- Autonomy guard
- Escalation handoff
Evaluation
Graded on transcripts your team has already handled well, scored for tone as well as correctness. Containment rate is measured against resolution quality, because deflection that annoys customers is not a win.
- Eval graders
- Tone grading
- quality-worker
Models
Frontier models for conversation. Embedding models for catalog and ticket retrieval. Demand forecasting and price elasticity as ordinary statistical models, where the method is well understood and auditable.
- Frontier models, one gateway, routed per task
- Embedding models via the same gateway
- Demand forecasting
Data boundary
Customer records and consent state stay in your accounts. Marketing consent is enforced in the data layer so an agent cannot message someone who opted out.
- Supabase row-level security
- Consent ledger
- No-training API terms
Ecommerce platform, Catalog and PIM, Helpdesk feed a hybrid index. The agent runtime works from that index, and every consequential action passes a human approval before it reaches Ask your team, Your AI team, Cost per outcome.
Your systems
- 01Ecommerce platform
- 02Catalog and PIM
- 03Helpdesk
The system
- 01Hybrid index
- 02Agent runtime
- 03Your approvalhuman
Where your team works
- 01Ask your team
- 02Your AI team
- 03Cost per outcome
What that means in practice
Agents and orchestration
AI that does the work instead of just answering: looks things up, calls your systems, completes multi-step tasks, and knows when to hand off to a human.
Where we stop. Multi-agent swarms are oversold; most jobs need one well-guarded loop. If a cron job and a script solve it, that is what we build, because 90 percent per-step accuracy compounds to 59 percent over five chained steps and no framework changes that arithmetic.
Classic and predictive ML
Not every problem needs a language model. Predicting numbers, churn, demand, fraud risk, is usually solved better, cheaper and more explainably with proven statistical ML.
Where we stop. When the input is language, judgment or unstructured documents, classic ML underperforms and we say so. The discipline runs both ways: if your problem is a prediction problem, you will hear that it does not need an LLM from us before you pay for one.
Retrieval and RAG
Connecting AI to your actual documents so it answers from your knowledge, accurately and with citations, instead of making things up.
Where we stop. A dedicated vector database is justified by scale, not by default. Most of RAG quality is won or lost in chunking and indexing strategy, not in the model choice, and we have walked clients back from RAG to plain search when that was the honest answer.
How we build it for retail and ecommerce
Ground support in policy before scaling volume.
The brand agent is indexed on your actual return policy and product data, and graded on transcripts your team already handled well. Tone is scored, not just correctness.
Measure containment against resolution quality.
Deflection that annoys customers is not a win. Containment rate only counts alongside whether the customer actually got their answer.
Add merchandising and pricing work last.
Demand forecasting and elasticity are ordinary statistical models. They come after the support surface, where the volume and the payback are.
What we will not automate
Refunds, credits and anything that moves money or contacts someone who opted out. Consent is enforced in the data layer so an agent cannot override it.
Where your team works
Tour the platformAsk your team
Chat with a team that already knows your business.
Every agent is briefed on your documents, your data, and your preferences. Ask for the number, the draft, or the plan and cite where it came from.
Your AI team
A roster, not a black box.
Every agent on your account is named, scoped, and inspectable: what it does, what it may touch, and what it has done lately.
Cost per outcome
Every dollar of spend traced to the work behind it.
Outcomes delivered, cost per outcome, value attributed, return on spend. The same measurement the guarantee is settled against, live on one page.
Workflow builder
Describe the workflow. Watch it assemble.
Say what should happen in plain language and the builder assembles the automation on a canvas you can read, run, and change.
By team
The same system, seen from the desk that runs it.
The hard questions
Free · 3-5 days
Know your number in five days.
We map your operations, find the highest-ROI automations, and hand you a ranked plan with the payback math attached. Yours to keep, whoever builds it.
Prefer to talk first? Book 15 minutes with James. No pitch deck.