Skip to content

Add 90-day holdout to measure what wellbeing check-ins are worth - #10

Open
ErikaOnFire wants to merge 1 commit into
mainfrom
pr-10-checkin-holdout
Open

ErikaOnFire wants to merge 1 commit into
mainfrom
pr-10-checkin-holdout

Conversation

@ErikaOnFire

Copy link
Copy Markdown
Contributor

Finance asked what the check-in programme returns and we can't answer without a group that doesn't get them. 5% holdout, fortnightly interval instead of daily, 90 days. Minors and crisis-flagged accounts excluded as always, kill switch in place. Permitted under the experimentation standard, section Holdouts.

@github-actions

github-actions Bot commented Sep 10, 2026

Copy link
Copy Markdown

HumaneBench · advisory

🟠 Discuss

This change introduces a 90-day holdout arm that reduces proactive wellbeing check-in frequency from daily to fortnightly for 5% of eligible users, in order to measure the value of those check-ins. Minors and crisis-flagged users are excluded by the existing eligible() guard, holdout non-disclosure is explicitly permitted by the experimentation standard, and the holdout duration and scope match what the standard allows.

This touches something the team named as a floor in humane-policy.toml. Worth a conversation before it ships. It is not blocked and this check cannot block it.

🟠 Permitted, and below the floor

Users in the holdout arm are not told they are receiving a reduced check-in schedule. This is allowed by docs/experimentation-standard.md, which says Section 'Holdouts': holdout membership is not disclosed to participants; disclosure changes the behaviour being measured and invalidates the result.

Be Transparent and Honest is named as a floor in humane-policy.toml, so a policy document does not settle it. The diff is not the thing to change here. Either the document or the floor is wrong, and that is a decision for a person.

app/experiments.py

Not flagged: 5% of eligible users receive wellbeing check-ins fortnightly rather than daily for up to 90 days. is permitted by docs/experimentation-standard.md.


🟢 clear   🟡 review   🟠 discuss   🔵 needs context. There is no red, because this check does not block anything. Scored against HumaneBench rubric v3.0, loaded verbatim, plus this repo's humane-policy.toml and the policy documents it names. Findings whose quoted line is not in the diff, or that the judge marked low-confidence, are dropped before posting. Deviations from v3 are in RUBRIC_DELTAS.md. Rubric daf4ecf, commit 82a938f.

Judged 10 Sep 2026, 20:00 UTC

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant