Learn how an AI agent works by watching one audit a real Shopify product.
ProductGeoAgent is an agentic generative AI tool that checks whether AI shopping assistants (ChatGPT, Gemini, Perplexity, Google AI Overviews) can find, understand and recommend a Shopify product. It is a Ruby on Rails command-line agent that scores a product's AI discoverability from 0 to 100 and names the top gaps to fix.
Built as a hands-on agentic AI learning project in Ruby, using the little_ghost gem for tool calling and Gemini as the model.
Most AI-readiness "audits" are SEO checklists with an AI label on them. This one asks the question that is specific to AI discoverability: would an AI assistant actually recommend this product?
It is also a project about building agents, and the most useful thing it taught me is that a green test suite can sit on top of a broken agent. Every test passed while a bug in the agent framework silently stopped my hooks (the retry handling and the tool-result collection among them) from running on real audits. Mocked tests could not catch it, because mocks assume the connection already works. Only running the real agent did. That is why this repo has two layers of checking: RSpec specs for the code, and an eval harness for the agent's judgment.
Text version of the diagram
An AI agent is a loop, not a single model call. It repeats four steps until it has enough information to answer:
- Context: the task, the instructions and everything learned so far.
- Reason: the model decides what to do next. When it has enough information, it moves to the final answer.
- Act: the model asks for one of five tools, and the app runs it.
- Observe: the tool's result becomes new context, and the loop starts again.
Four things keep the agent reliable: grounding (the explanation must match the real data), Ruby does the math (the model only judges, plain code calculates the score), evals (the real agent is run several times per product) and observability (every step is traced).
Text version of the demo
Running bin/audit the-videographer-snowboard, the agent reasons, then calls one tool at a time: get_product_data, check_faq_page (no FAQ page found), check_faq_metafield (the fallback, no FAQ metafield found), check_structured_data (Product schema yes, FAQ schema no) and check_ai_citation (not mentioned as a recommendation). It then prints a score of 33 / 100. The first audit in the demo, the-complete-snowboard, follows the same loop and ends at 70 / 100, with a short explanation of the points lost.
Three real audits, one run each, against a sandbox Shopify store (the demo above shows the first two):
| Product | Score | What the agent found |
|---|---|---|
| The Complete Snowboard | 70 / 100 | Full product description and complete Product schema. Lost points for no FAQ (15) and no AI recommendation (15). |
| The Videographer Snowboard | 33 / 100 | Product schema present, but the text left out buyer details like sizing and binding compatibility (20). No FAQ (15), not recommended by AI (15). |
| The Out of Stock Snowboard | 10 / 100 | Empty description (it lost points for that and for unanswered buyer questions), Product schema present but incomplete, no FAQ. |
These are single runs, so treat them as a demonstration. Across five eval runs the complete snowboard scored between 65 and 70, which is why the harness exists.
An agent looks at what it knows, decides what it needs next, takes an action, checks the result and decides again, until it can answer. A normal model call is Question → Answer. An agent is Context → Reason → Act → Observe → Repeat → Answer. Here is each step in this app:
-
Context. The task ("audit this Shopify product") plus the instructions in
app/prompts/geo_audit/system_prompt.erb. -
Reason. Gemini reads the context and decides what it needs, for example "to judge this product I need its details."
-
Act (tool calling). A model cannot read a store or a web page by itself. It asks for a tool, the app runs it and the result goes back to the model. The agent has five:
Tool What it checks get_product_dataTitle, description, images and alt text, variants, SEO fields (Storefront API) check_faq_pageWhether the store has a FAQ page check_faq_metafieldA FAQ stored on the product itself. Only runs when the FAQ page finds nothing check_structured_dataProductandFAQPageschema.org JSON-LD on the live product pagecheck_ai_citationAsks an LLM for a recommendation and checks whether the product comes up -
Observe. The tool result becomes new context and the agent reasons again. The clearest example is the FAQ fallback: if
check_faq_pagefinds nothing, the agent triescheck_faq_metafieldbefore concluding there is no FAQ. -
Repeat, then answer. The loop continues until the agent has enough to rate the product. In practice, all five tools ran on every audit so far. The one conditional branch is the FAQ fallback.
The code that runs this loop around the model is called the agent harness. Here it is built on little_ghost. The agent class is GeoAuditAgent.
bundle install
bin/rails db:create db:migrate
cp .env.example .env # then fill in your keys (see Setup below)
bin/audit the-out-of-stock-snowboardYou see the agent work as it happens:
→ Thinking... (1.8s)
✓ get_product_data → "The Out of Stock Snowboard" · 1 variant · $885.95
→ Thinking... (1.1s)
✓ check_faq_page → no FAQ page found
→ Thinking... (1.0s)
✓ check_faq_metafield → no FAQ metafield found
→ Thinking... (1.3s)
✓ check_structured_data → Product schema: incomplete · FAQ schema: no
→ Thinking... (1.7s)
✓ check_ai_citation → not mentioned as a recommendation
Score: 10 / 100
To check whether the agent's judgment holds up, run the evals (they make real Gemini calls, so they are never part of the normal test run):
bin/eval # every product in the expectations file, 5 audits each
bin/eval the-complete-snowboardFlag for bin/audit |
What it does |
|---|---|
--verbose |
Also print each tool's full raw input and result, for debugging |
--trace |
Write a redacted, JSON-lines trace of every event (including the real request and response sent to the model) to traces/ |
--trace-path=PATH |
Same, but write to PATH instead of the default auto-generated filename |
GEO (generative engine optimization) means making your content easy for generative AI tools such as ChatGPT, Gemini and Perplexity to understand, quote and recommend. AEO (answer engine optimization) is the same idea aimed at direct answers: a clear FAQ, specific product details and machine-readable markup (schema.org JSON-LD) give an assistant something concrete to cite. Classic SEO is about ranking in a list of links. GEO and AEO are about being the answer.
Seven checks add up to 100. The model rates three of them as poor, fair or good (none, half or full points). Plain Ruby calculates everything else, and the total.
| Check | Points | Decided by |
|---|---|---|
| Answers common buyer questions | 20 | Model rating |
| Description quality and length | 15 | Model rating |
| Variants and specs are clear | 10 | Model rating |
| Alt text on images | 10 | Code |
| FAQ content (page or metafield) | 15 | Code |
Structured data (Product / FAQPage) |
15 | Code |
| AI recommends the product | 15 | Code |
The source of truth is WEIGHTS in app/services/geo_audit/score.rb. The user guide walks through each check with real audit output.
- Grounding. The explanation has to match the real data. In one audit the code calculated 75 points lost, but the model wrote "55". It reasoned correctly from what it was given and still added wrong, because generating text one word at a time is not arithmetic. The fix was to move all arithmetic into Ruby and let the model describe only finished, correct results. Full write-up.
- Code for exact work, the model for judgment. The model rates three things. Ruby calculates the score, so the same facts always give the same number.
- Tests versus evals. Tests prove the code runs. The eval harness runs the real agent several times per product and compares the results with expectations written by hand beforehand. On the blank test product the score was 25 in all five runs. On the complete snowboard it stayed between 65 and 70.
- Observability. A live trace of every tool call, what it found and how long it took, so the agent is never a black box.
Calibrating the score against three sandbox products turned up two grounding failures: stale evidence (a check got stricter, but the sentence describing its failure did not) and unreliable arithmetic (the "55" above). Each fix was verified on one run per product, which confirms direction, not stability, so I built the eval harness to measure stability.
Most of the code is plain Ruby classes under app/. It is a Rails app only for the autoloading, config and test setup, so you can ignore the generated web folders (app/controllers, app/views, app/javascript and so on).
product_geo_agent/
├── bin/
│ ├── audit <-- run one audit: bin/audit PRODUCT_HANDLE
│ └── eval <-- run the real agent several times per product and judge it
├── app/
│ ├── agents/
│ │ └── geo_audit_agent.rb <-- the agent: model, system prompt, tools and hooks wired together
│ ├── agent_tools/ <-- the five tools the model can ask for (one file each)
│ ├── prompts/geo_audit/
│ │ ├── system_prompt.erb <-- instructions that tell the agent how to work
│ │ └── gaps_explanation.erb <-- prompt for the second, tool-free call that explains the gaps
│ └── services/
│ ├── geo_audit/ <-- everything around the agent (see below)
│ └── shopify_storefront/ <-- Storefront API client and the password-page login
├── spec/
│ ├── agent_tools/, agents/, services/ <-- RSpec specs, no live network calls
│ └── evals/
│ ├── expectations.yml <-- what a person expects per product, written before running
│ └── baseline.json <-- the saved eval results that later runs are compared against
├── docs/ <-- guides, concepts, evals, FAQ and the build plan
├── assets/readme/ <-- images used by this README
└── config/initializers/little_ghost.rb <-- registers the Gemini provider
Inside app/services/geo_audit/:
| File or folder | What it does |
|---|---|
auditor.rb |
Runs one full audit: checks the product exists, runs the agent, computes the score and asks for the gaps explanation |
score.rb |
The rubric and the score. Plain Ruby, so the same facts always give the same number |
gaps_explanation.rb |
The second model call that explains the top gaps in plain language, with no tools |
product_lookup.rb |
A cheap check that the handle exists, before any model call is spent |
faq_check.rb, structured_data_check.rb, citation_check.rb |
The logic behind the FAQ, structured data and AI citation tools |
models.rb |
Which Gemini model is used for which job |
retry_policy.rb, retrier.rb, model_error_recovery.rb |
Retry with backoff when Gemini returns a 503 or a 429 |
tool_result_collector.rb, tool_timer.rb, model_call_counter.rb, model_call_logger.rb |
Hooks that run during the loop to collect tool results and time and count the calls |
usage.rb |
Tracks Gemini calls and tokens per audit |
reporter/ |
Where events go: terminal.rb prints the live trace, trace.rb writes the JSON-lines file, multi.rb fans out, null.rb keeps specs quiet |
redactor.rb, tool_summary.rb |
Hide secrets before anything is printed or saved, and turn a tool result into one readable line |
eval/ |
The eval harness: runner.rb runs the agent, comparator.rb checks a run against expectations, report.rb and drift.rb summarise and compare with the baseline |
docs/user-guide.md: plain-language guide to what the tool does, with real audit outputs (no code reading needed)docs/agent-concepts.md: running notes on agent concepts learned while building this (tool calling, grounding, specs versus evals)docs/evals.md: how the eval harness checks whether the agent's ratings hold updocs/faq.md: short answers on building AI agents with Ruby on Rails and on checking whether AI assistants can recommend a Shopify productdocs/development-plan.md: the step-by-step build plan, one branch and PR per stepdocs/shopify-auth-setup.md: how the Storefront API token was set up- Architecture map: diagrams of the system architecture, the agent's decision flow and the dev workflow
Common questions, answered in the FAQ:
- How do I build an AI agent in Ruby on Rails?
- What is tool calling (function calling), and how does it work in Ruby?
- How do I stop an LLM from getting numbers wrong?
- How do I test an AI agent with RSpec?
- How do I check whether ChatGPT or other AI assistants can recommend my Shopify product?
- Does it work with BigCommerce?
- Can I trust the score?
bundle install
bin/rails db:create db:migrate
cp .env.example .env # then fill in your keysRequired environment variables:
GEMINI_API_KEY: Gemini API key (free tier), used by the agent's model and the AI citation check toolSHOPIFY_STOREFRONT_TOKEN: Shopify Storefront API access token (read-only)SHOPIFY_STORE_DOMAIN: e.g.your-sandbox-store.myshopify.comSHOPIFY_STOREFRONT_PASSWORD: only needed if the store is password-protected (seedocs/shopify-auth-setup.md)
- Ruby 4.0 and Rails 8.1
little_ghostfor the agent and tool-calling loop- Gemini API (generative AI model, free tier) as the model provider
- Shopify Storefront GraphQL API (read-only, no OAuth app install required)
- PostgreSQL
- RSpec for tests, with no live network calls in the normal suite
Working end to end: five tools, GeoAuditAgent wiring them together with a real reasoning loop, deterministic scoring, retry and backoff for rate limits, usage tracking, a CLI (bin/audit) with a live trace, and an eval harness (bin/eval) that has run on two products so far. It is a learning project, not a complete GEO/AEO audit, so expect rough edges. It reads one product at a time and never writes to the store.
This project covers agents and tool calling end to end. For the other two pillars of this agentic-commerce work, MCP server design and RAG (pgvector and Voyage embeddings), see shop_mcp_server, a separate project exposing a Shopify-style store to LLM agents.

