Skip to content

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

This file is 196 lines long; read all of them.

Zyte

Zyte Agentic Web Data for Codex CLI

From a plain-English prompt to a working Scrapy spider.

Version 0.4.0 Zyte EULA GitHub stars


Not using exclusively Codex CLI? See Zyte Coding Agent Add-Ons for alternatives.

Install

codex plugin marketplace add zytedata/codex-skills
codex plugin add zyte-web-data@zyte-ai

If Codex CLI is already running, restart the active session to load the plugin.

Codex in ChatGPT

A workspace admin, or anyone with plugin import permissions, can install the plugin for the whole ChatGPT workspace:

  1. Go to Workspace settings → Plugins → Add → Import marketplace.
  2. Enter zytedata/codex-skills as the GitHub repository. Leave Path and Branch empty, so that the workspace keeps getting plugin updates.
  3. Select Import marketplace, and authorize GitHub access if prompted.
  4. Search for Zyte in the plugin directory, open it, and select Install.

To use it, select Try it from the plugin page, @-mention it, or pick it from Plugins in the composer.

The plugin declares the Zyte MCP server itself, so do not add it with codex mcp add as well. Sign in to it once, in the browser:

codex mcp login zyte

What it does

This is Zyte's official Codex CLI plugin that generates production-ready Scrapy spiders with web-poet page objects from a plain-English prompt. Give it a URL and describe what you want to extract. It handles site exploration, schema discovery, code generation, and smoke testing: no boilerplate, no manual selector hunting.

The plugin explores the target site, discovers available fields, and presents a schema for your approval before generating a single line of code. After you confirm the schema, it creates a Scrapy project with all dependencies configured, generates web-poet page objects and test fixtures, wires up the spider, and runs a smoke test to verify that extraction is working before handing the project back to you.

Optionally, use /zyte to deploy directly to Scrapy Cloud for scheduled runs, job history, and monitoring, which the Zyte MCP handles once the spider is deployed. A free tier is available.


Use cases

The /scrape skill works on any website with repeating structured content: detail pages linked from a listing or category page. Examples from the skill:

  • Product catalogs
  • Job listings
  • Recipes

How does it work?

The /scrape skill orchestrates two stages automatically:

1. Plan and validate the scrape     →  /scrape-plan
2. Build the project and spider     →  /scrapy-extra

Each stage feeds directly into the next. When the pipeline completes, you have a runnable spider and a passing test suite:

uv run scrapy crawl <spider_name>
uv run pytest fixtures/

Skills

Orchestration

Skill Description
scrape End-to-end web scraping workflow — from URL to working spider with web-poet page objects

Pipeline stages (called automatically by /scrape)

Skill Description
scrape-plan Plan the scrape and author a validated extraction spec: discover fields, download diverse pages, compare HTML variants, optional browser review
scrape-analyze-page Extract all available fields with values from a detail page
scrapy-extra Hands-on Scrapy coding: write/debug spiders, web-poet page objects, and projects; configure scrapy-poet and scrapy-zyte-api

Zyte APIs

Skill Description
zyte Zyte work the Zyte MCP does not cover: set up your Zyte account and credentials; deploy projects to Scrapy Cloud and export all items, logs or requests of a job; look up Zyte API plan pricing; and answer how-to and documentation questions about Zyte from the official docs

Prerequisites

  • Codex CLI
  • uv — used to create and manage the Scrapy project
  • A Zyte account, signed in once through the Zyte MCP (see Install) — used for Scrapy Cloud jobs, Zyte API usage stats and per-website prices

Project dependencies (scrapy, scrapy-poet, scrapy-zyte-api, web-poet, extruct, price-parser, pytest) are installed automatically by the skills.


Quickstart

Any scraping prompt triggers the skill automatically. For example:

Scrape books.toscrape.com

The plugin walks you through schema approval interactively, then generates a complete, tested Scrapy project.


Update

To update manually, refresh the marketplace snapshot and reinstall the plugin:

codex plugin marketplace upgrade
codex plugin add zyte-web-data@zyte-ai

Then restart the Codex CLI session.


Evaluation

We automatically evaluate skills and track both wall time and cost. We measure and aim to improve these metrics over time.


Feedback

If you find any issue — such as prompts that did not work as expected, or that caused excessive wall time or cost — please open a GitHub issue.

Provide as much detail as possible to help us reproduce the issue. You are welcome to anonymize target websites or other data.


Frequently asked questions

Is a Zyte account required?

No. The generated spider is a standard Scrapy project that runs locally with uv. A Zyte account is required only if you want to deploy to Scrapy Cloud or use Zyte API to access sites that block standard scrapers. If you want to use Zyte API, you'll need an account to generate an API key.

Does it handle JavaScript-rendered pages?

The generated project includes scrapy-zyte-api as a dependency. Enabling headless browser rendering requires a Zyte API key. The /zyte skill guides you through setting up your credentials.

What Python libraries does the generated project use?

The project template includes scrapy, scrapy-poet, scrapy-zyte-api, web-poet, extruct, price-parser, and pytest. All dependencies are installed automatically via uv sync.

Can the generated spider run without Codex CLI?

Yes. The plugin generates a standard Scrapy project. Run it directly with:

uv run scrapy crawl <spider_name>

You can extend, modify, and deploy it independently of Codex CLI.


License

See LICENSE.md for the Zyte End User License Agreement.

About

No description, website, or topics provided.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages