Skip to main content
  • AI & Retail
  • Research

The state of AI in retail commerce

AI is everywhere in retail roadmaps. Real production impact is still concentrated in a handful of places: search, fraud, inventory and content.

Invenzo editorialThe applied-AI team
Sep 11, 20268 min read
An AI-native retail robot at a store counter

Every retail deck for the last two years has an AI slide. Fewer of those slides survive contact with a P&L. The gap is not model quality, it is where the model sits: bolted onto a dashboard nobody opens twice, or wired into the exact screen an operator already works from every day.

That distinction is the whole story. A recommendation engine running inside a data-science notebook and a recommendation engine running inside the storefront’s own search bar can be the identical model, and only one of them changes a conversion rate. The teams getting real value out of AI in commerce right now are not the ones with the newest model, they are the ones who solved the boring integration problem first: one signal store, one place the model writes back to, one team accountable for the metric it is meant to move.

The gap between the slide and the P&L

Ask any retail technology buyer what they are piloting and the list is long: a chatbot for support, a forecasting model for buying, a personalisation layer for the storefront, an agent for merchandising. Ask the same buyer which of those pilots is now load-bearing, something that would be missed operationally if it were switched off today, and the list gets short fast.

Most AI pilots never reach a line item

The pattern we keep seeing is not a failed model. It is a pilot that was never wired into the workflow it was meant to change, so nobody can point to the day it started mattering.

That is not a comment on any one vendor or team. It is a structural problem: a pilot proves a model can predict something and stops there, while making it matter operationally means owning the write path back into an order, a stock ledger, a support queue or a storefront, which is a much less glamorous project than the pilot itself.

Four places we actually see the needle move

Across the products we build and run for retailers, the AI that survives past the pilot stage keeps showing up in the same four places. None of them are exotic; all four are places where a wrong prediction has an immediate, visible cost, which is exactly what forces the integration work to get done properly rather than left half-finished.

  • Search & discovery

    Conversational and voice search, inside the storefront itself

    Replacing keyword search with an assistant that understands intent (and, for a multilingual shopper base, the language they actually speak) shows up immediately in a "did they find it" metric, so it gets fixed fast when it is wrong.

  • Fraud & checkout risk

    Scoring that retrains on a daily cadence

    A fraud model that is a week stale is worse than a rules engine. The ones that hold up retrain against yesterday’s attempted checkouts, not last quarter’s.

  • Inventory & fulfilment

    Forecasts that recompute when reality changes

    A demand forecast is only useful if it reacts the moment a supplier ETA slips or a store sells through early across channels, not on a weekly batch job that already assumes yesterday’s stock position.

  • Content & personalisation

    Recommendations wired into the page that gets clicked

    The same ranking model produces very different outcomes depending on whether it writes to a live storefront slot or a report a merchandiser reads once a week.

The pattern behind the ones that work

One signal store, not five dashboards

Every deployment that has stuck runs on a single event stream feeding every surface. The storefront, the warehouse, support and the buying team see the same signal at the same time, rather than five teams each reconciling their own export of it a day later. That is unglamorous plumbing, and it is the actual precondition for any of the four areas above to work at all.

Judged by the metric it changes, not by the model

The deployments that survive a budget review are the ones where someone could name the metric before the model shipped, and can show the same metric measured the same way afterward. "Improved relevance" is not a metric. "Search-to-add-to-cart went from X to Y, measured identically before and after" is.

Shipped where the operator already works

A model that requires a new tab, a new login and a new habit competes with everything else on someone’s desk, and mostly loses. The ones that stick are wired into the screen the picker, the support agent or the merchandiser already has open, so adopting the AI costs nothing more than a click they were already going to make.

“The question we ask before building any AI feature is not "can the model predict this" but "what screen does this need to live inside, and whose day does it need to change."”

Invenzo, applied-AI team

How to tell a pilot from a deployment

None of this requires reading a model card. A buyer evaluating any AI claim in a retail platform can ask five plain questions and learn more than a benchmark slide would tell them:

  1. If this were switched off tomorrow, who would notice, and how fast?
  2. What is the one metric this is supposed to move, and how was it measured before the model shipped?
  3. Does the output write back into a live workflow, or into a report someone reads separately?
  4. How often does it retrain, and does that cadence match how fast the underlying behaviour actually changes?
  5. Is there a fallback to the manual process it replaces, or was this a hard cutover?

A "yes" to the first four and a real answer to the fifth is what a deployment looks like. Vague answers to more than one of them is what a pilot still in search of a home looks like, however good the demo was.

What this looks like in Indian retail specifically

Two things make this market a harder, more interesting test than most. Distribution here is still genuinely fragmented, so a stock forecast has to hold across a general-trade counter, a modern-trade store and a quick-commerce dark store carrying the same SKU at three different velocities, not one clean DTC funnel. And omnichannel stopped being a differentiator here well before it did elsewhere; it is the baseline shoppers already expect, which raises the bar for what "personalisation" or "search" even means.

That context is also why voice and language matter more here than the average AI feature list suggests. A search assistant that only understands English keyword queries is solving last decade’s problem; one built to understand a shopper speaking their own language, mid-sentence, on a phone, is solving this one.

What actually correlates with AI showing up on the P&L

  • The model writes back into a live workflow, not a report someone reads separately.
  • One event stream feeds every surface that touches the same customer or the same order.
  • A named metric was agreed before the model shipped, and is measured the same way after.
  • The rollout keeps a fallback to the process it replaces, rather than a hard cutover.
  • Retraining is scheduled against how fast the underlying behaviour actually changes, not against a calendar.

The next 12 months

The platforms pulling ahead are the ones that stopped treating "AI" as a feature bolted onto a roadmap and started treating it as a property of the data layer underneath every feature: one signal store, one set of models, and channel-specific surfaces that all read from the same place. That is the architecture behind every one of the four areas above, and it is also, not coincidentally, the same architecture a retailer needs for a unified storefront, warehouse and last mile regardless of whether AI is involved at all.

Key takeaways

  • Real AI impact in retail today concentrates in search, fraud, inventory and content: the places where a wrong prediction has an immediate operational cost.
  • The difference between a pilot and a line item is almost always integration, not model quality.
  • A model judged against a named, consistently measured metric survives budget reviews; one described only in adjectives does not.
  • The precondition for any of this is a single signal store feeding every surface, not a model per team.
  • Opinion · free to read

Subscribe

Get the best of the blog in your inbox.

Product notes, applied-AI write-ups and retail benchmarks roughly twice a month. No spam.