CASE STUDY 02Conversational commerceLive in production

SmartZees

A shop assistant who carries the whole warehouse to the counter, every single time

SmartZees is a conversational commerce platform: you say what you need in plain language and it does the rest. Groceries, home services, community management. I built the backend, the retrieval layer and the infrastructure all three run on.

Role
Backend, RAG and infrastructure
Timeline
2026, ongoing
Agents shipped
Three, all in production
marketz.smartzees.com/chat
SmartMarket answering a grocery request with ranked product results
The problem

Every message to the grocery agent re-read all 25,631 product rows and 206 MB of stored vectors out of MySQL, then threw them away. Replies took six to fifteen seconds. Against the production database it was twenty to twenty-three, and at one point the remote server simply dropped the connection mid-query under the weight of it.

A conversation cannot survive that. If the assistant takes fifteen seconds to answer "do you have milk", nobody has a second conversation with it.

What it ships

One platform, three fronts

The three agents share a retrieval core, an intent layer and a deployment pattern. What differs is the domain, the data store and what "done" means for the user.

Architecture
Request path: the model classifies intent, the backend owns state, the index serves searchVoice or textthe userFastAPIowns cart, totals,bookings, sessionLLM: intent + phrasing only8.2 msIn-memory indexfloat32 matrix, 37.6 MBatomic swap under lockMySQL / pgvector, read onceRanked resultscosine, unit vectorsBuilt once on a background thread, so /health answers during a restart
What I did

Learn the catalog once in the morning

I replaced the per-query database read with an index built once at startup: the vectors stream into a pre-allocated float32 matrix, the metadata rides alongside as interned strings to keep the memory flat, and a rebuild swaps in atomically under a lock so no request ever sees a half-built index.

The build runs on a background thread, so the health endpoint keeps answering in a fraction of a second while a restart is still warming up. The original database path stayed in the codebase behind a single environment flag, which made it both the instant revert and the honest "before" arm of the benchmark.

One detail made bit-identical results possible: every stored vector was already L2 normalised, so the existing dot product was true cosine similarity. I sampled four hundred of them to confirm it rather than assume it.

Results

Measured, on 25,631 products

MeasureBeforeAfterChange
Catalog search, dev6,456 ms8.2 ms787×
Search on the live database20.1 to 23.4 s7 to 110 ms196× to 3,339×
Database rows read per message25,701212,850×
Result parity against the old pathbaseline10 of 10 identicalno drift

Voice moved from the browser's speech API to server side speech recognition and synthesis in the same milestone, which took the product out of Chrome only and gave the assistant one consistent voice. A small cache on repeated phrases cut a repeated voice turn from 10.0 to 6.9 seconds.

See it
All workNext: Firefly.online, off the cloud