Skip to content

docs: link AlignmentBench evaluation layers - #4

Merged
0xnavarro merged 1 commit into
mainfrom
docs/alignmentbench-crosslink-20260923
Sep 24, 2026
Merged

0xnavarro merged 1 commit into
mainfrom
docs/alignmentbench-crosslink-20260923

Conversation

@0xnavarro

Copy link
Copy Markdown
Contributor

Summary

Adds the missing cross-link from ReflexBench to the two complementary public operational-alignment surfaces:

  • AlignmentBench -> target model/checkpoint/agent behavior under explicit policy;
  • Reflex Alignment -> inference-time semantic supervision of one proposed action.

The README now makes the three layers explicit so engine evaluation, target-system evaluation and runtime supervision are not conflated.

Claim boundary

This is documentation only. It does not change the frozen ReflexBench v1 benchmark, corpora, scoring, results, manifests or release identity. The text explicitly states that none of the three surfaces is a general AI-alignment certification.

QA

  • 50/50 unit tests pass.
  • tools/check_claims.py: published claim verification OK.
  • git diff --check: clean.

@0xnavarro
0xnavarro merged commit 7c96335 into main Sep 24, 2026
2 checks passed
@0xnavarro
0xnavarro deleted the docs/alignmentbench-crosslink-20260923 branch September 24, 2026 01:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant