docs: start historical retrieval shadow benchmark - #356
Conversation
|
Warning Review limit reached
Next review available in: 11 minutes Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (3)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 761c81f41d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| The decision requires at least ten organic questions covering dev-backlog, | ||
| dev-relay, and consumer repositories, at least two weeks between first and last | ||
| real reuse, and no decision before 2026-08-28. If live sources remain sufficient, | ||
| close #350 with no compiler. If Arm C wins every gate, propose a separate | ||
| `project-memory` skill and human-gated charter amendment; do not productize it |
There was a problem hiding this comment.
Require organic coverage for every compared arm
When the organic observations are imbalanced—for example, nine Arm B questions and one easy Arm C question—the ten-question and two-week conditions are satisfied even though Arm C's recall, error rate, and medians are based on a single sample. Because each log record contains only one arm and this decision gate specifies neither paired B/C observations nor a per-arm minimum, the Day-0 result could support proposing project-memory without demonstrating Arm C across the required organic question set. Require paired observations for each question, or equivalent per-arm coverage across all repository classes.
Useful? React with 👍 / 👎.
Summary
Supports #350. This PR intentionally does not close the issue. Earliest go/no-go remains 2026-08-28.
Day-0 evidence
These are synthetic Day-0 results, not organic reuse evidence.
Verification
node --test skills/backlog-triage/scripts/*.test.js skills/dev-backlog/scripts/*.test.js— 454 pass, 1 skip, 0 failnode --test skills/dev-backlog/scripts/contract-prose.test.js— 9/9 passgit diff --check