Case study · RAG and AI chatbots

StoreFilter: an AI analyst for any Shopify store, built on 2.5 million records

An AI analytics chat for e-commerce, built with Next.js, Claude Code, Supabase, pgvector and GPT-4. Ask about any Shopify store, get answers with charts.

StoreFilter chat answering a question about a Shopify store's conversion rate, with the calculation, insights and follow-up questions
Role
Full-stack developer
Timeline
– (4 months)
Stack
  • Next.js
  • Claude Code
  • Supabase
  • PostgreSQL + pgvector
  • Edge Functions
  • OpenAI GPT-4
  • Stripe
  • Vercel

StoreFilter is an AI analytics assistant for e-commerce. You type a question about any Shopify store, mentioning its domain, and get an answer back: traffic, conversion rate, revenue, social media presence, and how the store compares. No dashboards to learn, no reports to build. Visitors get two free analyses, then sign up.

From March to June 2025 I was the full-stack developer on it: the Next.js app, the Supabase database and its import pipeline, the store-matching search, the AI answers and the Stripe subscriptions.

The problem

The data to answer these questions existed, but in a form only an analyst could use: millions of product, order and marketing records. Two things stood in the way of a simple chat box:

  • Finding the right store. People type the same store a dozen ways: https://www.eightsleep.com, www.eightsleep.com, eightsleep.com, or something close to it. Every variation has to land on the right merchant, every time.
  • Turning a question into an answer. A founder asking “what is the conversion rate for this store?” needs the right numbers pulled, the calculation shown, and an explanation in plain words, quickly.
product and order records in the database
2.5M+
Stripe subscription events handled by webhooks
5
months, March to June 2025
4

What I built

Store matching: search by meaning, then score

Every domain a user types is first normalised (protocol and www. stripped), so the common variations collapse into one form. For everything else, the store index lives in pgvector inside Supabase: the input is compared against stored stores by similarity, the nearest candidates come back with a score, and an algorithm in a Supabase edge function picks the highest-scoring match. It’s the retrieval half of a RAG system, used to make sure the answer is about the right store.

  1. Question

    Plain English, with a store domain in it

  2. Normalise

    URL variations reduced to one domain

  3. Vector match

    pgvector finds the nearest stores and scores them

  4. Pick and fetch

    Edge function picks the best match, pulls its data

  5. GPT-4 answer

    Calculation, insight, chart and follow-up questions

Answers that show their working

Once the store is identified, the relevant figures are pulled from the database and handed to GPT-4 with prompts routed by the kind of question: revenue, inventory, product performance, customer behaviour or conversion. The answer shows the formula and the numbers behind it, explains what they mean, and suggests follow-up questions, so the next question is one click away.

StoreFilter answering 'What is the conversion rate for eightsleep.com?' with the formula, estimated monthly sales and visits, the result, revenue insights and follow-up questions
A conversion rate answer: the formula, the inputs, the result and suggested follow-up questions.

Where a comparison reads better as a picture, the answer includes a chart.

StoreFilter answering which social media platform drives the most engagement for a store, with a bar chart comparing followers, and a sign-up prompt below the chat
An answer with a chart, and the sign-up prompt once the free analyses are used.

A database built for millions of records

The Supabase PostgreSQL database holds more than 2.5 million product and order records. To get them in reliably I built an internal import tool with batch uploads, retries, progress tracking, validation and error logging, so a failed batch could be resumed instead of starting over. Indexes, query tuning and aggregation pipelines keep the common questions fast: revenue, conversion, inventory and returns.

Subscriptions with Stripe

Free visitors get two analyses, then sign up. Paid plans run on Stripe Checkout and Subscriptions, with webhook handlers for new subscriptions, renewals, failed payments, upgrades and cancellations, so access always matches what the customer is paying for.

The app itself

The front end is a Next.js app, written with Claude Code and deployed on Vercel: a chat interface with conversation history, sign-up and login, and access rules that keep each account’s data its own.

Build log

March 2025

Started

Joined as the full-stack developer, responsible for the app, the database and the AI layer.

Along the way

Data and imports

Set up the Supabase database and the batch import tool, and brought more than 2.5 million product and order records in.

Along the way

Store matching and AI answers

Built domain normalisation and the pgvector similarity search with a scoring step in an edge function, then the GPT-4 answer layer with routed prompts, charts and follow-up questions.

June 2025

Wrapped up

Handed over with the Next.js app, store matching, AI answers, the import tool and Stripe subscriptions in place.

What this project shows

  • Retrieval that has to be right. Vector search plus a scoring step, so the AI answers about the store you meant, not the one with a similar name.
  • AI answers people can check. The numbers and the formula are in every answer, not just a confident sentence.
  • A complete SaaS, not a demo. Millions of records, a reliable import pipeline, accounts and Stripe subscriptions, in four months.

Thinking about an AI assistant over your own data, whether it’s documents, products or orders? See RAG development and custom AI chatbots.

Building something like this?

A multi-tenant SaaS, a scan-to-data pipeline, dashboards for two kinds of users. A 30-minute call is enough to scope your version.

Book a 30-minute call