A support agent that investigates customer problems, searches company policies, checks real order data, and takes safe actions only with human approval.
Status: Lesson 1 of 14 is done: a normal Rails app with real and planted data, and no AI yet. Lesson 2 (the first LLM call) is next. This is a learning project, built step by step, not a finished product. The plan below describes where it is going, not what works today.
Most AI support demos are chatbots that sound confident and guess. A real support case needs more than a fluent answer:
"My order #1042 has not arrived and it has been two weeks. If it is lost, please send a replacement."
To answer this correctly, a system has to look up the order, check the shipment, find the lost-package rule in the shipping policy, decide whether a replacement is allowed, and get a person to approve it before anything is changed. Then it has to cite where each fact came from.
This project builds that system, and measures whether it gets it right.
An agent that reasons and uses tools instead of answering from memory:
Text version of the diagram
- Route: decide what the message needs (order data, policy, or both).
- Act: call read tools such as
get_orderandget_shipment. - Retrieve: search policy documents (RAG) for the relevant rule.
- Decide: a replacement or refund is a write action, so it needs approval.
- Ask for approval: a support person sees the proposed action with the evidence.
- Execute and audit: the action runs and is written to an audit log.
- Answer: the customer gets a reply that cites the order status and the policy section.
Steps 1 to 4 only read data, and if the agent does not have enough facts after deciding, it goes back to step 2 and calls more tools. Steps 5 and 6 change something, so a human has to say yes first.
Every step is traced, and each scenario is also an eval case with a known correct answer.
Tools are plain Ruby classes. The agent and the MCP server only wrap them, so the same tools work from the chat, the approval screen and Claude.
| Tool | Type | Rule |
|---|---|---|
get_order, get_shipment, get_customer, search_orders, search_policies |
Read | Automatic |
update_address |
Write, low risk | Automatic after a policy check |
create_return, create_replacement_order |
Write, high risk | Human approval |
issue_refund, cancel_order |
Write, high risk | Human approval and confirmation |
About five minutes, three of them the import. You need Ruby 4.0.7, PostgreSQL and a free Kaggle account.
- Download the Olist dataset and unzip it into
data/olist/. The importer reads six of its files: customers, products, orders, order items, order payments and the category translation file. The files are not in the repo because of the dataset license. - Set up the app, load the data and plant the test scenarios:
bundle install
bin/rails db:create db:migrate
bin/rails data:import_olist
bin/rails data:plant_scenariosThe import prints how many rows each step kept and skipped (nothing is skipped on the full dataset):
15:40:47 importing customers
15:41:43 customers: 96096, lookup entries: 99441
15:41:43 importing products
15:41:57 products: 32951
15:41:57 importing orders
15:43:03 orders: 99441, skipped: 0
15:43:03 importing order items
15:43:22 order items: 112650, skipped: 0
15:43:22 importing payments
15:43:39 payments: 103886, skipped: 0
data:plant_scenarios generates a shipment for every order that reached a carrier, then edits five fixed orders so each tells one story with a known answer:
shipments: 97658
delayed shipment: order 1
lost package: order 2
return window closed: order 3
final sale item: order 4
duplicate charge: order 5
- Check that everything works:
bundle exec rspec # 20 examples, 0 failuresFor a quick try, LIMIT=200 bin/rails data:import_olist loads only 200 customers and their orders in under a minute. There is no chat or web screen yet, so there is nothing to open in a browser. The agent arrives from lesson 2.
| Area | Choice |
|---|---|
| App | Rails 8, Hotwire, Tailwind CSS |
| Database | PostgreSQL, pgvector (added in the RAG lesson) |
| LLM library and provider | RubyLLM, Groq (free tier) |
| Embeddings | Ollama (local) |
| MCP | Official MCP Ruby SDK |
| Observability | OpenTelemetry, Langfuse |
| Evals | eval-ruby, rspec-agents |
Gems are added in the lesson that needs them, not all at once. All LLM calls will go through one small LlmClient class, so changing provider is a config change.
Three milestones, 14 lessons. Each lesson changes the same application and gets a Git tag (lesson-01, lesson-02, ...).
Milestone 1: a working support agent
- A normal Rails app with real and planted data (no AI yet) (done)
- First LLM call and structured output
- Tool calling with read tools
- The agent loop, written by hand
- MCP server and Claude connector
- RAG over policy documents
Milestone 2: quality and safety
- Routing: tools, RAG or both
- Hybrid search and reranking
- Write tools, permissions and human approval
- Evals
- Observability and cost tracking
Milestone 3: advanced
- Memory and context engineering
- Multi-agent orchestration
- Feedback loop and deployment
Real public data makes the app believable, and planted scenarios give every eval a known answer.
- Base data: the Brazilian E-Commerce Public Dataset by Olist (Kaggle): customers, orders, items, products and payments.
- Planted scenarios: added by a seed script with a fixed random seed, so everyone gets the same data. First set: delayed shipment, lost package, return window closed, final sale item, duplicate charge.
- Policies: written by hand, with deliberate exceptions, because exceptions are where weak retrieval fails.
Six tables hold the customers, orders, payments and shipments the agent will investigate. The full notes are in docs/DATA_MODEL.md.
Click a table to open its model:
Dataset files are not committed. They live in data/ (git-ignored) and are downloaded separately. Check each dataset's license on its own page before using it.
Attribution: the base data is the Brazilian E-Commerce Public Dataset by Olist, licensed under CC BY-NC-SA 4.0. It is used here for non-commercial learning only. This repository contains the import and seed code, not the data or a database built from it.
Customer messages for routing and evals (from Lesson 7) come from the Bitext customer support dataset, licensed under CDLA-Sharing-1.0. The data file is not included here either.
This is a learning project about building a reliable AI agent, not a complete support product. Real Shopify or BigCommerce integration, multi-modal input and new eval tooling are out of scope for now.
- The full plan, with architecture, gems by lesson, data mapping and the evaluation plan: docs/PROJECT_PLAN.md
- The database tables and how they relate, as a diagram: docs/DATA_MODEL.md
- Lesson progress, decisions and what broke: docs/lessons/01-data-model-and-import.md
- Index of all documentation: docs/README.md
- Where the code lives: this GitHub repo is the source of truth, and pull requests and issues are here. A read-only mirror is kept on GitLab.