Code and source manifests for collecting, cleaning, normalizing, generating and reproducing Najd Research datasets. Consumers download data from Hugging Face; this repository owns how it is built.
| Dataset | Rows | Purpose |
|---|---|---|
| System One | 5,184 | Public decision-model research draft; 11 separately licensed packs, independent review pending |
| Najd Benchmark | 6,089 | Arabic/Saudi evaluation collection: all 6,089 cases in one default configuration |
| Najd Legacy 31 | 31 | Selected original Najd cases, required fixtures and historical row mapping |
| Arabic Riddles and Questions | 333 | Freshly collected question–answer pairs, with source attribution on every row and the dataset card |
The standalone datasets overlap with Najd Benchmark; do not add their counts as independent cases. Current public versions omit review annotations. Questions, answers, IDs, splits and scoring were not changed by that metadata migration. Dataset availability is not a claim of semantic correctness.
The current benchmark pin is in releases/current.json. Earlier metadata-only revisions are in the migration ledger. See standalone dataset builds and metadata migration.
The published Hugging Face dataset card is also tracked here for documentation updates.
| Repository | Responsibility |
|---|---|
| datasets | Sources, collection, normalization, synthetic generation, provenance and dataset releases |
| benchmark | Task contracts, shared evaluation/scoring, local reports and pinned execution package |
| najd-arena | Website, organizations, managed jobs, private reports, publication and public results |
The current Hugging Face version, 2026.09.27, exposes all 6,089 cases in default/test without audit_status labels. Prompts, answers, IDs, provenance and specific issue notes are preserved. To produce it after the historical public build:
uv run python scripts/build_unified_release.py build/public-release/release build/currentEarlier pinned revisions remain unchanged. Historical commands below reproduce those original artifacts.
See the complete source dependency audit for all 39 sources, the public-only execution command, and evidence required to close each remaining gap.
uv sync --extra dev --extra parquet --extra excel
uv run pytest
uv run najd-datasets audit-public-inputs releases/2026.09.14/sources.json --sources sourcesThe offline audit inventories original manifests. To rebuild all 6,089 cases from public sources, use the public release builder:
uv run python scripts/reconstruct_public_release.py --output build/public-release
uv run python scripts/verify_public_release.py build/public-releaseBoth JSONL files reproduce the published bytes. Parquet exports reproduce values and schema. Historical audit documents are preserved separately. Original authoring/extract history and some redistribution permission evidence remain unavailable; see the documented reconstruction boundary and source attribution.
| Job | Entry point |
|---|---|
| Download every historical published file | Pinned release inventory |
| Reproduce historical bytes | Historical replay |
| Rebuild an individual source | Source catalog and najd-datasets reproduce-source |
| Collect a new dataset | Source standard and najd-datasets collect |
| Clean or generate | najd-datasets clean / generate; small original specs in fixtures/specs/ |
| Recollect the 333 questions | Standalone publications |
| Package/publish a new version | package, check-approval, publish; release protocol |
| Propose a new task family | Task roadmap |
| Contribute | CONTRIBUTING.md |
For example:
uv run najd-datasets reproduce-source sources/arabic-agent-eval.json --output build/arabic-agent-eval
uv run najd-datasets generate fixtures/specs/reminder-variants.json --output build/generated.jsonl
uv run najd-datasets validate build/generated.jsonlSource adapters verify the historical expected hashes in a compatibility context, then emit current rows without review annotations. Reports distinguish the historical checksum from the new output checksum. Historical replay commands intentionally retain the original artifact representation.
Every source needs its origin, pinned revision/hash, redistribution basis, transformation and split. Keep credentials, private cases and large artifacts out of Git. Builds go in ignored build/; private permission records remain private. Attribution is preserved on dataset cards and rows.
The generic publication command requires a hash-bound approval record and refuses to overwrite an existing version. Omission of review metadata does not disable rights/privacy controls or record a completed review. See the public reconstruction gates.
Build the internal Arabic customer-support version-comparison task pack using original fictional policies and scenario-level splits. Other task families are documented placeholders, not released benchmark packs.
The original bilingual pilot can now be packaged locally with hashes and development-only provenance. See fixture instructions. It is not a reviewed public release or a fresh holdout.
Dataset v1 documentation describes the completed local candidate: 6,884 cases across controlled policies, natural development drafts, and public references/diagnostics. Build with python -m najd_datasets.system_one_release --include-references --output <fresh-directory> after restoring the pinned inputs. Cases are not automatically eligible for publication.
benchmark and Arena use the immutable identity in releases/current.json: 6,089 cases across 25 tracks. Task-specific scorers may select a documented subset. Saved results keep their original revisions; changing the default release never rewrites prior scores. New rows do not become scorable merely by removing metadata: missing references and tasks requiring an execution harness must remain explicitly ungraded until supported.
Version 2026.09.27.1 adds verified public fixture files and four source-based answer
corrections. Run scripts/build_executable_release.py after the historical and unified
builders; see the correction guide. Case count remains
6,089. Old revisions remain immutable; these corrected cases require a fresh evaluation.
Dataset manifests and contribution checks now share pinned schemas with benchmark and Arena. The original Arabic support routing example rebuilds deterministically and demonstrates the end-to-end development contract. It is separate from the existing Hugging Face benchmark release.