~500 LOC. One-click deployment. Full control. Your custom skills.
xarnes-agent is a small, self-hosted pull-request review runner. It watches your GitHub repository
and uses OpenAI Codex to run your configured skill on each eligible
PR. Your skill chooses the checks, verdict, and complete review body. The runner posts the result:
pass approves, comment comments, and block requests changes.
- Easy to verify. The whole runner is one ~500-line file using only Node.js built-ins. Read it end to end to see how credentials, reviews, retries, and delivery work.
- Host it yourself. Deploy to your Render account with one click, or run Docker on your own infrastructure. Your persistent volume holds the sign-ins, queue, and results.
- Full control. Choose the repository, branches, model, concurrency, and reviewer identity. Inspect and change the source, configuration, and saved state.
- Your custom skills. Set
SKILLto a skill in your repository. It defines the review policy and the Markdown body to post; the runner handles execution, retries, and GitHub delivery.
Click Deploy to Render and provide the five required environment values:
| Field | Value |
|---|---|
GITHUB_REPO |
owner/name of the repository to watch |
SKILL |
Directory in that repository holding SKILL.md, for example skills/review |
GITHUB_TOKEN |
Fine-grained token with Contents: read, Pull requests: write, Commit statuses: write |
STATUS_CONTEXT |
Review/check name, for example security-review. GitHub displays security-review/pr-123 for PR #123 |
MAX_CONCURRENCY |
1 to start. Each extra slot needs its own Codex sign-in on this service's disk, so it is set per service and never overwritten by a Blueprint sync |
All optional watcher settings are declared in render.yaml and applied automatically. Open your service's Environment page to see all 18 settings, including these defaults:
| Optional setting | Render default |
|---|---|
TARGET_BRANCHES |
main,dev |
MAX_ATTEMPTS |
3 |
POLL_SECONDS |
15 |
RUN_EXISTING |
false |
WATCH_UPDATES |
true |
MODEL |
gpt-6-sol |
FAST_MODE |
false |
TASK_TIMEOUT_SECONDS |
1800 |
SANDBOX |
container |
ALLOW_NETWORK |
false |
DATA_DIR |
/data |
CODEX_HOME |
/data/codex |
CODEX_BIN |
codex |
Render's creation form prompts only for the five required values. To change optional settings
permanently, edit your fork's render.yaml: a later Blueprint sync can overwrite changes made on
the Environment page. Settings used only by local task mode are listed under Configuration below.
Codex uses your ChatGPT account. On first start, follow the sign-in link and code in the logs for each configured slot. Sign-ins run one at a time; polling starts after all slots are authenticated. No shell access is needed.
The Render blueprint builds the Dockerfile as a background worker with a persistent
1 GB disk at /data. The GitHub token and saved Codex sign-in live on your Render service.
If you fork this repository, update the button above to point to your fork. Render builds directly from the repository; no published image is required.
every 15 s ─► list open PRs ─► keep non-draft PRs into TARGET_BRANCHES ─► queue new heads
│
┌─────────────────────────────────────────────────┘
▼ (up to MAX_CONCURRENCY at once, one Codex sign-in each)
fetch HEAD + target ─► merge-base = BASE ─► checkout HEAD, worktree BASE ─► load skill from BASE
│
▼
codex exec ◄── skill path, repo, PR number, BASE, HEAD, PR metadata file
│ returns JSON: verdict, body
▼
validate ─► save result + marker ─► recheck PR ─► POST review ─► commit status
- Discover. Every
POLL_SECONDSthe agent lists open pull requests and keeps the non-draft ones targetingTARGET_BRANCHES. A PR is queued when it is new, has a new head commit, or was retargeted. Discovery keeps running while reviews are in progress, so new work is never blocked behind a long review, and a finished review wakes the next scan immediately. - Prepare. For each queued PR the agent fetches the exact head and target commits, computes the merge base, checks out the head, and creates a second worktree at the merge base. The skill is loaded from the merge base, never from the PR, so a PR that edits the skill still gets reviewed by the skill the target branch had.
- Review. Codex runs non-interactively in its own process group and Codex home, starting
outside both checkouts with user configuration, rules, and
AGENTS.mdloading disabled. It receives the skill path, repository, PR number,BASE,HEAD, and a metadata file with the PR title and description, which the prompt marks as untrusted data. It never receives the GitHub token. - Validate. The response must be a JSON object with a verdict of
pass,comment, orblock, and a non-emptybodystring. The skill decides the verdict and writes the complete review. Nothing is posted for a malformed result. - Deliver. The validated result is saved to the volume with a unique marker before anything is posted. The agent then re-reads the PR (still open, same head, same base, not a draft), submits the review bound to the reviewed commit, and updates the commit status. If the response is lost, the next scan finds the marker on GitHub and does not post again.
- Retry. Failed review execution uses backoff (5, 10, 20, 40 minutes, capped at 60) until
MAX_ATTEMPTSis spent. Interrupted runs retry on restart when attempts remain; a graceful stop does not consume an attempt. GitHub delivery failures retry the saved result without rerunning Codex or consuming review attempts.
A status check named <STATUS_CONTEXT>/pr-<number> appears in the PR's checks section. The agent
publishes it as a GitHub commit status and updates it automatically as the review progresses:
| Stage | Status check | GitHub review action |
|---|---|---|
| Waiting for a worker | 🟡 Pending: Queued for review | No review posted yet |
| Picked up by a worker | 🟡 Pending: Review running (attempt n/N) | No review posted yet |
| Review finished, delivery pending | 🟡 Pending: Review finished; posting the result | Posting or confirming delivery of the review |
Result pass |
✅ Success: Review passed | Approves the PR and posts the review report |
Result comment |
✅ Success: Review complete with non-blocking comments | Posts the review report with comments, without approval |
Result block |
❌ Failure: Review blocked; changes requested | Requests changes and posts the review report with blocking findings |
| Retry scheduled | 🟡 Pending: Review failed; retry scheduled | No final result yet |
| Codex usage limit reached | 🟡 Pending: Codex usage limit reached; review will retry automatically | Not an attempt; reviews pause 15 minutes, then resume |
| Retry budget exhausted | Delivery has not completed | |
| Superseded revision or closed PR | No new review posted for the ineligible revision |
For a passing review, the sequence is queued → running → posting the review → approved + check passed. The final success or failure status is published only after GitHub confirms the review was posted (or the agent finds an already-posted review when reconciling a retry). The report is the body of that approval, comment, or change request; it is not a separate duplicate PR comment.
The skill writes the complete review body, including any headings, icons, tables, evidence, and commit references. The runner posts it with credential redaction and a hidden delivery marker, without adding visible formatting.
Running. The check is pending and shows the current attempt. Required approval remains outstanding while the agent reviews the PR.
Changes requested. A blocking verdict posts a change request and marks the check as failed. In this example, the repository's review rules also block merging.
Passed with comments. A non-blocking comment verdict marks the check as passed and posts the
review comments. It does not approve the PR, so GitHub can still show Review required and
Merging is blocked, as in this example. A pass verdict submits an approval instead.
These statuses are informational: their names include the PR number, so they cannot serve as one reusable required status check for the branch. To gate merging on reviews, configure required approvals and dismiss stale approvals in your branch rules. GitHub does not let a token approve its owner's own PRs, so use a dedicated reviewer identity.
docker build -t xarnes-agent:local .
cp agent.env.example agent.env # set GITHUB_REPO, SKILL, GITHUB_TOKEN, STATUS_CONTEXT
chmod 600 agent.envSign in once, or skip this and follow the link the watcher prints on first start:
docker run --rm -it --no-healthcheck --env-file ./agent.env -v my-agent:/data xarnes-agent:local loginWatch:
docker run -d --name pr-agent --stop-timeout 15 --restart unless-stopped \
--log-driver local --log-opt max-size=10m --log-opt max-file=3 \
--env-file ./agent.env -v my-agent:/data xarnes-agent:local watch
docker logs -f pr-agentReuse the same volume on every run: it holds the sign-in, the state file, and the saved results.
Run exactly one watcher per volume; the container holds a lock on it and refuses a second owner.
With MAX_CONCURRENCY=N, each slot has its own Codex home (/data/codex, /data/codex-2, …) and
its own sign-in; login signs in every slot that lacks one, login 2 re-signs one slot. Slots
signed in with the same ChatGPT account share its rate limit, so concurrency buys wall-clock time,
not quota.
One-off run against a local checkout, without GitHub (no review is posted):
docker run --rm -v my-agent:/data -v "$PWD:/workspace" -e SKILL=skills/review xarnes-agent:local taskTask mode runs the explicitly configured SKILL. For a free-form task, set INSTRUCTIONS or
INSTRUCTIONS_FILE (see instructions.example.md) and leave SKILL unset.
Any edits remain in the mounted workspace.
Everything is an environment variable. agent.env.example is a commented starting point.
The table lists the runner's defaults when variables are absent. The Render blueprint and local
example explicitly select gpt-6-astra with FAST_MODE=true.
| Variable | Default | Purpose |
|---|---|---|
GITHUB_REPO |
required | Repository to watch, owner/name. Its presence selects watch when no command is given |
SKILL |
required | Skill directory in that repository; a bare name means skills/<name> |
GITHUB_TOKEN |
required | GitHub credential, supplied as an environment variable |
TARGET_BRANCHES |
main |
Comma-separated base branches; only non-draft PRs into these are reviewed |
STATUS_CONTEXT |
required | Commit-status check name; the agent appends /pr-<number> |
MAX_CONCURRENCY |
1 |
Reviews run in parallel, one Codex sign-in each |
MAX_ATTEMPTS |
3 |
Attempts per head commit before the failure becomes terminal |
POLL_SECONDS |
15 |
Discovery interval and minimum delivery-retry delay; integer ≥ 15 |
RUN_EXISTING |
false |
Also review PRs already open on the first scan |
WATCH_UPDATES |
true |
Review new head commits on already-reviewed PRs |
MODEL |
Codex default | Codex model override |
FAST_MODE |
false |
true requests Codex's Fast tier (more ChatGPT credits, if available for the model) |
TASK_TIMEOUT_SECONDS |
1800 |
Limit for one Codex run; the whole process group is killed on expiry |
SANDBOX |
container in the image |
container trusts the container boundary; read-only / workspace-write use Codex's inner sandbox |
ALLOW_NETWORK |
false |
Network access for skill commands in workspace-write mode |
DATA_DIR |
/data |
Volume root: sign-ins, state, results, checkouts |
CODEX_HOME |
$DATA_DIR/codex |
Slot 1's Codex home; slot N uses <CODEX_HOME>-N |
CODEX_BIN |
codex |
Codex executable |
WORKSPACE |
/workspace |
Task mode: the checkout to run against |
SKILLS_DIR |
$WORKSPACE/skills |
Task mode: where bare skill names are resolved |
INSTRUCTIONS / INSTRUCTIONS_FILE |
none | Task mode: free-form prompt instead of a skill |
Credentials are read at startup; restart after rotating them. The image runs as the non-root node
user and pins its Node base image by digest and the Codex CLI by version.
Configure SKILL with the directory of the skill you want the runner to execute, for example
SKILL=skills/your-review. That directory must contain SKILL.md and any references it needs,
committed to the repository being reviewed. Your configured skill owns the review policy, the
verdict, and the complete review body.
Its final response must be a JSON object with two required fields, without Markdown fences or
surrounding text:
{
"verdict": "block",
"body": "## Changes requested\n\nThe endpoint allows anonymous writes. Restore the authorization check."
}verdict: exactlypass,comment, orblock.body: a non-empty string containing the complete GitHub review in Markdown.
| Verdict | GitHub review | Commit status |
|---|---|---|
pass |
APPROVE |
success |
comment |
COMMENT |
success |
block |
REQUEST_CHANGES |
failure |
These are the only required fields. Extra fields are ignored. The runner validates the verdict and body, then posts the skill's Markdown using the matching GitHub review action.
The runner posts body with credential redaction and a hidden marker for duplicate prevention.
The JSON response is saved under /data/runs/*.json. The runner's 60,000-byte publishing limit
includes the marker; oversized reviews fail rather than truncate. Skills should keep their bodies
below this limit.
- State lives in
/data/prs-OWNER--REPO.json: per PR the reviewed head, base, attempt count, result path, marker, and delivered review id. Writes are fsynced and atomically renamed. Invalid state stops startup without changing the file. Closed PRs retain completed and pending review records so reopening can reuse a saved result or recognize an already-posted review. Other closed-PR records are removed. - Concurrency is one process. Workers share the process; the state file and GitHub status writes are serialized; a PR never has two reviews in flight, so a new head on an active PR waits.
- Status delivery posts directly to GitHub. Retrying after a lost response may add an identical entry to the commit's status history. Review delivery still checks its marker to avoid duplicate reviews.
- Codex usage limit. When Codex exits with "You've hit your usage limit", the run is not counted
against
MAX_ATTEMPTS: the PR's status says the limit was reached, every slot pauses for 15 minutes, and the same attempt is retried when the pause ends. Nothing is escalated to an operator for a limit that resets on its own. - Healthcheck (
node /app/agent.mjs healthcheck) is a liveness check: the discovery loop ticked withinmax(2 min, 3 × POLL_SECONDS). A failing GitHub scan logsPoll failedbut is not "unhealthy", because a restart would not fix it. - Shutdown on
SIGTERMaborts GitHub calls, signals every Codex process group, escalates toSIGKILLafter five seconds, and marks interrupted runs for immediate retry without spending an attempt. Allow at least 15 seconds for graceful shutdown. - Storage:
codex[-N]/holds sign-ins,runs/saved results,prs-*.jsonwatcher state, andagent.lockthe single-owner lock. Each review uses temporary checkouts underworkspaces/; both snapshots and their Git history are removed after success or failure. Startup clears any checkouts left by a crash. Saved results have no automatic expiry. Allow space forMAX_CONCURRENCY × (repository history + two working trees), plus saved results and Codex data.
The skill is loaded from the merge base. Skill files and any symlink targets are trusted; the runner
does not scan or restrict symlinks. Codex workers do not receive the GitHub token; only the wrapper
fetches and publishes. Review bodies, saved results, and logs redact the exact GITHUB_TOKEN value;
other credentials and encoded tokens are not redacted. The default SANDBOX=container trusts the
container as the boundary: skill commands run as the same user as the wrapper, with the volume and
network reachable. Use this for
repositories you control. Reviewing hostile or public pull requests needs workers isolated from the
publisher and its credentials.
| File | What it is |
|---|---|
agent.mjs |
The runner. Everything described above |
Dockerfile |
Node + Git + Codex CLI, non-root, tini + flock entrypoint |
agent.env.example |
Commented configuration template |
render.yaml |
Deploy to Render blueprint |
.github/workflows/ci.yml |
Runs tests and builds the Docker image locally in CI; does not publish or deploy |
*.test.mjs |
Tests: local Git fixtures, a fake Codex, a mocked GitHub API. No network, no credentials |
node --test --test-timeout=120000 agent.test.mjs reviews.test.mjs lifecycle.test.mjs status.test.mjsTests cover the verdict mapping, skill-written bodies, delivery reconciliation across restarts and lost responses, retries and backoff, concurrency, queue persistence, status transitions, shutdown, crash recovery, state validation, and first-start sign-in from the logs. They use temporary directories and mocked services; no network or credentials are needed.
References: Codex documentation, skills, GitHub reviews API, GitHub commit statuses.


