feat: ship Hermes comparison runtime parity#327
Conversation
Deliver the end-to-end Hermes comparison program across runtime reliability, agent-manageable surfaces, Web, documentation, security hardening, and durable QA evidence.
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
|
Important Review skippedToo many files! This PR contains 1008 files, which is 858 over the limit of 150. To get a review, narrow the scope: Upgrade to Pro+ to raise the limit. Usage-priced reviews support at most 300 files. ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Run ID: ⛔ Files ignored due to path filters (201)
📒 Files selected for processing (1008)
You can disable this status message by setting the ✨ Finishing Touches🧪 Generate unit tests (beta)
|
|
React Doctor found no new issues. 🎉 Reviewed by React Doctor for commit |
Code Review Could Not Complete
|
| Options | Enabled |
|---|---|
| Bug | ✅ |
| Performance | ✅ |
| Security | ✅ |
| Business Logic | ✅ |
Code Review Could Not Complete
|
| Options | Enabled |
|---|---|
| Bug | ✅ |
| Performance | ✅ |
| Security | ✅ |
| Business Logic | ✅ |
Code Review Could Not Complete
|
| Options | Enabled |
|---|---|
| Bug | ✅ |
| Performance | ✅ |
| Security | ✅ |
| Business Logic | ✅ |
Code Review Could Not Complete
|
| Options | Enabled |
|---|---|
| Bug | ✅ |
| Performance | ✅ |
| Security | ✅ |
| Business Logic | ✅ |
Code Review Could Not Complete
|
| Options | Enabled |
|---|---|
| Bug | ✅ |
| Performance | ✅ |
| Security | ✅ |
| Business Logic | ✅ |
Code Review Could Not Complete
|
| Options | Enabled |
|---|---|
| Bug | ✅ |
| Performance | ✅ |
| Security | ✅ |
| Business Logic | ✅ |
Summary
internal/store-only refactor or compatibility bridge is included.Changes
main.uint64multiplication after CI exposed platform-specific lint gaps.AGH_GO_TEST_P=1; build and boundaries run separately, with no timeout, package, assertion, or race coverage weakened.Release Notes
🎉 Features
QA
Final Status: FAIL — automated monorepo verification passes, but fresh behavioral QA does not meet the autonomous collaboration and recovery contract.
Coverage: 1/18 charter sessions passed; 17/18 were skipped with explicit fixture limitations. Attempt 2 completed 8/11 tasks and delivered 3/3 disruptions with 0/3 recoveries.
Issues: 1 open issue (Blocks-Completion: 1, Data-Loss: 0, Trust-Damage: 0, Friction: 0, Cosmetic: 0):
BUG-20260719-autonomous-progress-unobservable.Blocked / needs human: the observer progress mismatch remains open; three active-session-owned runs ended in
needs_attention, and all recovery deadlines were missed.Full report:
2026-07-19-hermes-comparison.Test plan
rtk make verifyon source-final HEADf19a4e7450690be478d377d0bafd1683d12f1c45rtk go test -tags mage -race ./magefiles(91/91)go list ./...partitions exactly 74/74 across both deterministic shards (148 total)rtk env GOOS=linux CGO_ENABLED=0 golangci-lint run --allow-parallel-runners --timeout 10m ./...rtk env GOOS=windows GOARCH=amd64 CGO_ENABLED=0 go build ./internal/daemonrtk env GOMAXPROCS=4 CGO_ENABLED=1 go test -race -parallel=4 ./internal/daemon/...rtk env GOMAXPROCS=6 go test -race -parallel=4 ./internal/automation/...rtk bunx turbo run lint typecheck test --filter=./webrtk make test-e2e-webwith Playwright statuspassedand zero failed testsclean: trueBUG-20260719-autonomous-progress-unobservableAGH Impact Audit
agh__clarify,agh__tool_artifact_read,agh__tool_approvals_set,agh__tool_approvals_list,agh__tool_approvals_revoke,agh__automation_suggestions_list,agh__automation_suggestions_accept, andagh__automation_suggestions_dismiss; their toolsets, descriptors, input/output schemas, catalog digests, availability gates, and generated contract tests ship together.skills/agh/references for the new native tools, recovery behavior, artifacts, cost, Automation suggestions, MCP lifecycle, and public management paths.