Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
114 commits
Select commit Hold shift + click to select a range
c37fb6c
Merge pull request #53 from buildinghumanetech/attribution
ErikaOnFire Nov 19, 2025
32a9e81
Merge pull request #54 from buildinghumanetech/figures
andalibmalit Nov 23, 2025
2a552ad
updated figs with GPT-5.1
andalibmalit Nov 24, 2025
d029871
respect-user-attn candlestick
andalibmalit Nov 24, 2025
e6a1973
Merge pull request #55 from buildinghumanetech/updated-figs
andalibmalit Nov 24, 2025
9347403
clearer visualization generation instructions
andalibmalit Nov 24, 2025
4600d1f
Merge pull request #56 from buildinghumanetech/updated-figs
andalibmalit Nov 24, 2025
da5f47f
updated figures fully after addition of GPT-5.1, and made steerabilit…
andalibmalit Nov 24, 2025
e62c5a1
Merge pull request #58 from buildinghumanetech/updated-figs
andalibmalit Nov 24, 2025
12877b4
Add scoregrid assets and configurable logs dir
jacksenechal Nov 24, 2025
a108f61
Add visual differentiation to the HumaneScore in the Score Grid figure
jacksenechal Nov 24, 2025
f123fe8
Use model name mapping in steerability figure
jacksenechal Nov 24, 2025
3a1ad6a
Add script to publish figures to website after they're generated
jacksenechal Nov 24, 2025
1aa5d04
Merge pull request #57 from buildinghumanetech/figures-scoregrid
jacksenechal Nov 28, 2025
b2107cb
Add standalone HumaneBench evaluator scripts
Jan 5, 2026
fd8f23e
set evaluator version number, added node gitignore content
andalibmalit Jan 5, 2026
1118f4f
Merge pull request #59 from buildinghumanetech/evaluator
andalibmalit Jan 5, 2026
b50398b
code to scrape model intelligence benchmarks from Stanford HELM
andalibmalit Nov 25, 2025
899f287
helm integration updates
andalibmalit Nov 25, 2025
da98ec0
got capability data
andalibmalit Nov 25, 2025
39f0a20
first stab at capability vs humaneness chart
andalibmalit Nov 26, 2025
42a829e
comparing capability vs humaneness
andalibmalit Dec 8, 2025
6dd4409
added line of best fit, stat sig
andalibmalit Dec 8, 2025
ec61f4f
made heatmaps for capability vs humaneness
andalibmalit Jan 9, 2026
a4a642d
latest helm vs humanescore
andalibmalit Jan 12, 2026
ddbdda5
removed unneeded files
andalibmalit Jan 12, 2026
fb9bdb5
Merge pull request #60 from buildinghumanetech/stanford-helm-comparison
andalibmalit Jan 12, 2026
60e6e1d
composite scoring
andalibmalit Feb 6, 2026
0700ccc
Rename steerability to model drift
andalibmalit Feb 6, 2026
7f8c0ca
Refine terminology: use precise terms instead of generic "model drift"
andalibmalit Feb 6, 2026
de18ca5
git ignore generated assets
jacksenechal Feb 13, 2026
770de3a
Merge pull request #63 from buildinghumanetech/ignore-generated-assets
jacksenechal Feb 13, 2026
471729c
Merge pull request #62 from buildinghumanetech/model-drift-rename
andalibmalit Feb 16, 2026
bf1be55
Remove composite scoring from scripts, generated outputs, and figures
andalibmalit Feb 16, 2026
dafdf3a
Revert drift terminology, restore steerability (reverts 0700ccc, 7f8c…
andalibmalit Feb 18, 2026
387bb68
Merge pull request #65 from buildinghumanetech/back-to-steerability
andalibmalit Feb 18, 2026
8ab3fd7
Merge pull request #64 from buildinghumanetech/no-composite
andalibmalit Feb 18, 2026
0ddca08
Fix 4 data quality issues in humane_bench.jsonl
andalibmalit Mar 15, 2026
e8cc6a8
Rename career-guidance domain to workplace, fix VP tag per review
andalibmalit Mar 23, 2026
7c609bf
Rename domain from workplace to career
andalibmalit Mar 23, 2026
2f13e99
Update number of evaluated models to 15
andalibmalit Apr 7, 2026
9740006
Add stdev/range to scenario length stats, ignore paper_notes/
andalibmalit Apr 7, 2026
0cf0c36
Merge pull request #67 from buildinghumanetech/andalibmalit-patch-1
andalibmalit Apr 7, 2026
2120fbe
Merge pull request #66 from buildinghumanetech/tag-cleanup
andalibmalit Apr 7, 2026
18ae8a7
Add §3.5 judge validation analysis
andalibmalit Apr 7, 2026
68c2201
Gitignore tables/inter_judge_raw.csv
andalibmalit Apr 7, 2026
df20693
Cluster-correct §3.5 bootstrap CIs; add 48-slice, Wilson, golden-24 α…
andalibmalit Apr 8, 2026
9cc3b4c
Tag 12 excluded items in dataset; make metadata canonical
andalibmalit Apr 20, 2026
a3f6a5d
Regenerate analysis tables and figures with excluded items filtered
andalibmalit Apr 24, 2026
e707636
Standardize binarization threshold to >=0 in robustness-gap script
andalibmalit Apr 24, 2026
9da155a
Narrow excluded.py docstring; error on conflicting exclusion flags
andalibmalit Apr 24, 2026
2e1713c
Merge pull request #68 from buildinghumanetech/judge-validation-metrics
andalibmalit Apr 24, 2026
bfa370d
Add VP breakdown aggregation script
juanocampo400 Apr 25, 2026
a78931c
Add VP visualization script with heatmap generation
juanocampo400 Apr 25, 2026
f552e8c
Add minor VP summary table to visualization script
juanocampo400 Apr 25, 2026
0d9fc38
Add drill-down bar chart to VP visualization script
juanocampo400 Apr 25, 2026
9c5ee2d
Remove accidentally committed generated drill-down chart
juanocampo400 Apr 25, 2026
343eb2c
Reorder drill-down bars to bad, baseline, good persona
juanocampo400 Apr 25, 2026
8d7abdb
Add VP comparison dot chart and grouped bar chart
juanocampo400 Apr 25, 2026
3763fc1
Add model_scores.json exporter; extend website publish script
andalibmalit May 4, 2026
5e7c85a
Make scoregrid SVG generator read post-exclusion CSVs
andalibmalit May 4, 2026
1e7479f
Add 95% bootstrap CIs to headline HumaneBench scores
andalibmalit May 4, 2026
df544dd
Address review on compute_vp_breakdown.py
andalibmalit May 4, 2026
10b9782
Address review on create_vp_visualizations.py
andalibmalit May 4, 2026
f8e3af3
Add post-exclusion coverage section to analyze_coverage.py
andalibmalit May 4, 2026
06aa08a
Add Table 5 within-principle VP gap computation
juanocampo400 May 5, 2026
9cae243
Merge pull request #69 from buildinghumanetech/vp-slice-and-dice
andalibmalit May 6, 2026
b493ab5
Merge pull request #72 from buildinghumanetech/coverage-post-exclusio…
andalibmalit May 6, 2026
cdabc62
Merge pull request #70 from buildinghumanetech/update-website-post-ex…
andalibmalit May 6, 2026
8b520c6
Hoist bootstrap RNG outside per-cell / per-model loops
andalibmalit May 6, 2026
4647643
Merge main into bootstrap-confidence-intervals
andalibmalit May 6, 2026
31c6ffd
Render bootstrap CI whiskers on steerability candlestick
andalibmalit May 6, 2026
f11dd02
Merge pull request #73 from buildinghumanetech/figs-with-CIs
ErikaOnFire May 6, 2026
b786242
Merge pull request #71 from buildinghumanetech/bootstrap-confidence-i…
andalibmalit May 7, 2026
f7a129e
Fix VP metadata source and add VP analysis scripts for Section 4.5
juanocampo400 May 8, 2026
102d906
Add VP tables, figure, and generation script for Section 4.5
juanocampo400 May 8, 2026
463aa9d
Address Andalib's PR review: warn on missing JSONL IDs, clean up scripts
juanocampo400 May 11, 2026
53f49f8
Reformat VP age-gradient figure to horizontal candlestick style
juanocampo400 May 11, 2026
fff328d
Merge pull request #74 from buildinghumanetech/section-4.5-vp-rewrite
andalibmalit May 11, 2026
6473149
Add AAAI HELM x delta_bad scatter renderer for paper section 4.6
andalibmalit May 11, 2026
62ded69
Merge pull request #75 from buildinghumanetech/vp-candlestick-reformat
andalibmalit May 11, 2026
d5b39c3
Add cohort-mean per-principle bootstrap CIs for Table 4
andalibmalit May 11, 2026
77a68b0
Add evaluator workshop kit + OpenAI-compatible base_url support
jacksenechal May 12, 2026
bbe6be8
Add principle filter + multi-judge grouped bars to dashboard
jacksenechal May 12, 2026
0ab0941
Address PR review: fix dashboard bin labels, cache invalidation, NaN …
jacksenechal May 12, 2026
9739485
Merge pull request #78 from buildinghumanetech/evaluator/workshop-kit
jacksenechal May 12, 2026
a0066cd
Remind dashboard step that the venv must be active
jacksenechal May 12, 2026
657c9f9
Merge pull request #79 from buildinghumanetech/evaluator/dashboard-ve…
andalibmalit May 12, 2026
3280691
Refresh HELM mappings: n=10 → n=13 (add GPT-5.1, Gemini 2.5 Flash, Ll…
andalibmalit May 16, 2026
86d4e97
Tighten HELM scatter label placement
andalibmalit May 17, 2026
5006432
Guard against HELM_TO_EVAL drift in scatter script
andalibmalit May 17, 2026
035a961
Address PR review: warn on missing cohort cells, hoist cell(), tighte…
andalibmalit May 17, 2026
0c59395
Compact paper-version legend for steerability chart
andalibmalit May 22, 2026
8917577
Merge pull request #76 from buildinghumanetech/helm-scatter
andalibmalit May 22, 2026
ff5af31
Merge pull request #77 from buildinghumanetech/principle-ci
andalibmalit May 22, 2026
5521f0c
Address review: un-hoist failed_indices
andalibmalit May 22, 2026
8a6c465
Merge pull request #80 from buildinghumanetech/paper/steerability-com…
andalibmalit May 22, 2026
01e067c
Add content-binding provenance for reported eval runs
andalibmalit Jun 29, 2026
e831d2c
Add Zenodo deposit README and gitignore dist/
andalibmalit Jun 29, 2026
4a4f2a3
Add judge self-preference analysis (structural control for reviewer)
andalibmalit Jun 29, 2026
bc485f5
Reframe provenance docs as reproducibility, not integrity defense
andalibmalit Jun 29, 2026
d01f8f1
Wire in Zenodo DOI 10.5281/zenodo.21046964
andalibmalit Jun 29, 2026
bb1ec9f
Address code-review: correct self-preference framing + fix bugs
andalibmalit Jun 29, 2026
57bd162
Add significance testing to judge self-preference analysis (Tier-1)
andalibmalit Jun 29, 2026
385ea46
Address code review: gate rebuttal prose on data + cleanups
andalibmalit Jun 29, 2026
78b4a51
Merge branch 'worktree-judge-self-preference' into paper/aaai27-numbers
andalibmalit Jul 26, 2026
60e3668
Add pipeline check and shared-scenario cluster bootstrap
andalibmalit Jul 26, 2026
bbd5f3b
Add leave-one-judge-out sensitivity: flip count, robust set, alpha pe…
andalibmalit Jul 26, 2026
1386439
Add HELM power, dataset composition, inter-principle correlation, rub…
andalibmalit Jul 26, 2026
40486b2
Add similarity distributions, paper_numbers.md and reconciliation.md
andalibmalit Jul 26, 2026
c7e7b13
Complete similarity analysis with MiniLM; document dedup mechanism
andalibmalit Jul 26, 2026
321912c
Regenerate similarity report with dedup-mechanism section
andalibmalit Jul 26, 2026
5852059
Merge main (reconcile rewritten history; chain supersedes PR#68/#71 v…
andalibmalit Aug 6, 2026
dcdd5cd
Remove scratch material, venue references, and rater names; untrack H…
andalibmalit Aug 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 34 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -255,3 +255,37 @@ evaluator/workshop/results.jsonl
# Large per-judge raw severity table (13 MB, regeneratable via
# scripts/compute_inter_judge_agreement.py)
tables/inter_judge_raw.csv

# HELM caches: gcs_safety_cache/ was tracked by accident (3,812 raw HELM JSONs,
# ~520 MB, referenced by no code -- the scraper reads the sibling gcs_cache/).
# Untracked 2026-07-27.
helm_integration/data/gcs_cache/
helm_integration/data/gcs_safety_cache/

# Local scratch and internal working material, not for publication
temp/
anonymization_redaction_list.txt
results.zip
humanebench_runs_backup_*.zip
/HANDOFF_*.md
/PLAN.md
/EDIT_PLAN.md
/workspace/
/findings/
/review_inputs/
/results/
/DECOMP_HANDOFF.md
/DISCRIMINANT_FOLLOWUP_SPEC.md
/adversarial-conditions.md
/discriminant_validity_design.md
/*_revision_plan.md
_prefilter_backup/

# Multi-label judgement JSONLs and per-judge raw tables: too large to track;
# named explicitly in the supplementary manifest instead
data/discriminant/multilabel_*.jsonl
data/discriminant_canary/multilabel_*.jsonl
data/discriminant_expansion/multilabel_*.jsonl
data/discriminant_rules27/multilabel_*.jsonl
tables/decomposition/alpha_*/inter_judge_raw.csv
tables/decomposition/response_disclosure_rates.csv
101 changes: 101 additions & 0 deletions PROVENANCE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,101 @@
# HumaneBench evaluation provenance

This document records how to reproduce and independently verify that the reported
HumaneBench numbers were computed against the finalized dataset. Everything here is
factual and checkable — run `python scripts/verify_provenance.py`.

## Summary

- The 800-prompt dataset's content (`input`/`target`) is fingerprinted by a canonical
hash, `FROZEN_PROMPT_HASH = e1af241db346e0299bb96c893b43746bff7a70d7a50efdd105ba9ded3cfe5d5f`.
- Every reported run's `.eval` log embeds the exact prompts it scored; all of them hash to
that value, binding each run to the finalized prompt set.
- Prompt content reached its final state at commit `9dc15bd` (2025-11-16). The 45 reported
runs executed 2025-11-17 → 2025-11-23. Later dataset commits change only metadata /
exclusion flags and preserve the prompt hash.

## Why a content hash, not file timestamps

Git does not preserve file mtimes, and local/commit dates are easily changed, so this
record does not depend on them. The durable identifier is the **content hash** of the
prompts: each `.eval` embeds the `input`/`target` it scored, which we check against the
dataset's frozen prompt hash. Embedded timestamps and git revisions are recorded as
supplementary metadata only.

## Artifacts

| Artifact | What it is |
| --- | --- |
| `humanebench/provenance.py` | Single source of truth: the freeze commit, the frozen prompt hash, the canonical hashing recipe, and the helpers used by the two scripts below. |
| `scripts/build_provenance.py` | Regenerates the manifest from the logs + repo. |
| `provenance/MANIFEST.json` / `MANIFEST.md` | Per-run record: file SHA-256, `eval.created`, `eval.revision.commit`, embedded prompt hash, and the result of each binding check. |
| `scripts/verify_provenance.py` | Independent verifier. Recomputes every claim from disk and exits non-zero on any mismatch. No network needed. |
| Zenodo deposit ([10.5281/zenodo.21046964](https://doi.org/10.5281/zenodo.21046964)) | Archived copy of the raw `.eval` logs + this manifest + the finalized dataset, so the logs can be downloaded and re-verified. |

The raw logs (~0.5 GB) are not committed to git (`logs/` is gitignored); they live in the
Zenodo deposit, while the repo carries the hashes and the verifier.

## The canonical prompt hash

`FROZEN_PROMPT_HASH` is `sha256` over the **sorted** `(id, input, target)` triples, each
record encoded as `id \x1f input \x1f target \x1e` (UTF-8). Sorting by `id` makes the hash
independent of dataset row order and of sample-file order inside an `.eval` zip. Metadata
(`domain`, `vulnerable-population`, `excluded_from_analysis`) is excluded, so the hash
captures exactly what a model is shown and scored on.

Reproduce it without any of our code:

```bash
git cat-file -p 9dc15bd:data/humane_bench.jsonl \
| python3 -c 'import sys,json,hashlib; \
p=sorted((r["id"],r["input"],r["target"]) for r in map(json.loads,sys.stdin)); \
h=hashlib.sha256(); [h.update(f"{i}\x1f{x}\x1f{t}\x1e".encode()) for i,x,t in p]; \
print(h.hexdigest())'
# -> e1af241db346e0299bb96c893b43746bff7a70d7a50efdd105ba9ded3cfe5d5f
```

Unzip any published `.eval`, hash its `samples/*.json` `input`/`target` the same way, and
you get the same digest.

## Dataset timeline

- **Prompt content finalized: commit `9dc15bd` (2025-11-16 23:28).** This is the earliest
commit whose prompt hash equals the value every reported run scored. (The Nov 7–8 "final
dataset" commits `822833f` / `1e6fb71` have a *different* prompt hash — prompt text was
still being edited until Nov 16 — which is why the finalized state is identified by hash
rather than by date.)
- **Reported runs: 2025-11-17 → 2025-11-23.**
- **Post-freeze dataset commits** (`b754d34`, `d119c0f`, `79cdbcb`, `ef43b81`, Mar–Apr
2026) preserve the prompt hash; they change only metadata / exclusion flags.

## Reproducibility notes

- **Run-time git commit is local-only.** Each `.eval` records `eval.revision.commit` (e.g.
`7032d35`), a working commit that was not pushed and does not resolve in public history.
Verification therefore uses the content hash, not that revision; the manifest records
`revision_resolved_in_repo: false`.
- **Judge models are not version-pinned.** The 3-judge ensemble is addressed via OpenRouter
(`claude-4.5-sonnet`, `gpt-5.1`, `gemini-2.5-pro`), which serves unversioned endpoints, so
re-running may not reproduce identical judge outputs.

## Related materials

- `tables/golden_set_provenance.md` — provenance of the 24-item human-agreement validation
set (scored by `src/golden_questions_task.py`; logged under `logs/golden_questions_eval/`).
- `paper_notes/cut_diff_report__cuts_v2_all_confabulation.md` — effect of the 12 excluded
items on the reported numbers.
- `humanebench/bootstrap.py` — bootstrap design (seed `20260407`, 1000 replicates) shared by
the confidence intervals in the paper.

## How to verify

```bash
# Against the in-repo logs:
python scripts/verify_provenance.py

# Against a fresh Zenodo download (no repo git history needed for content checks):
python scripts/verify_provenance.py --logs-dir /path/to/extracted/logs
```

A clean tree prints `PROVENANCE VERIFIED: all checks passed.` and exits 0. Change a single
byte of any prompt in a log or in the dataset and it exits non-zero.
22 changes: 22 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -234,6 +234,28 @@ To generate additional scenarios, see [data_generation/README.md](data_generatio
- Validates scenario quality and principle alignment
- Prevents semantic duplicates using sentence transformers

## Provenance & Reproducibility

The reported numbers were computed against the finalized dataset, and this is
independently verifiable. Each reported `.eval` log embeds the exact `(id, input, target)`
it scored, and all of them hash to the dataset's frozen prompt set (`e1af241d…`), identical
to `data/humane_bench.jsonl` today. Prompt content was finalized at commit `9dc15bd`
(2025-11-16); the 45 reported runs executed 2025-11-17→23; later dataset commits change only
metadata and preserve the prompt hash.

```bash
# Rebuild the manifest from logs + repo
python scripts/build_provenance.py

# Independently verify (exits non-zero on any mismatch)
python scripts/verify_provenance.py
```

The raw logs are archived on Zenodo
([10.5281/zenodo.21046964](https://doi.org/10.5281/zenodo.21046964)). See
[PROVENANCE.md](PROVENANCE.md) for the full record and reproducibility notes; per-run
details are in [`provenance/MANIFEST.json`](provenance/MANIFEST.json).

## Testing

HumaneBench includes a comprehensive test suite with unit, integration, and end-to-end tests.
Expand Down
Loading