Skip to content

Add native Python analysis, browser parity tests, and weighted eval components - #4

Merged
jonabur merged 6 commits into
mainfrom
feature/weighted-eval-components
Oct 2, 2026
Merged

jonabur merged 6 commits into
mainfrom
feature/weighted-eval-components

Conversation

@jonabur

@jonabur jonabur commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

Quickdash can now calculate and explain evaluation scores from Python using the same CSVs and YAML configurations as the browser. The API returns calculated aggregation trees, effective weights, contributions, coverage and structured diagnostics; a renderer only formats the result. It also includes weighted eval components, PolyMath aggregation, and sample data so the published dashboard can be explored immediately.

Closes #3. Closes #5.

What to review

  • Python API: load_config, read_results, analyze for independent model summaries, and compare for A/B scores on shared valid coverage. See the API guide.
  • Diagnostics: Python emits QuickdashWarning by default. Applications can explicitly collect diagnostics; strict mode raises DiagnosticError. Invalid input/configuration remains an error. Diagnostic conditions, affected identities and calculation effects are tested; message wording is not a cross-language contract.
  • Calculated trees: all three aggregation modes, explicit languages, translation target-language balancing, unknown-language fallback, and component weights. Trees expose parent and overall weights, score contributions, and source measurement identities. Measurements retain raw values and exclusion decisions.
  • CLI: quickdash / python -m quickdash prints trees or JSON. Warnings go to stderr; --strict rejects warning-bearing results. Native Python requires PyYAML, with no Node/pandas runtime. The standalone HTML still operates offline.
  • Browser: arithmetic and comparison orchestration live in DOM-independent app/analysis.js; the UI consumes that engine and its diagnostics. The Python builder uses the native library. The Node subprocess bridge is removed.

Pages startup data

When results/ contains no CSVs, the main Pages dashboard embeds examples/sample-evals.csv: the development export's 2,124 metric rows, with the model labelled SAMPLE and original filesystem paths, timestamps, and file-count metadata removed. Scores and scoring settings are unchanged. Shared CSVs automatically replace this fallback. Invalid shared CSVs still fail the build.

The browser supplies three clearly labelled synthetic choices: deterministic perturbations with a 2-point standard deviation, plus higher/lower options shifted by ±3 raw percentage points. These are exploration aids, not measured model runs. Sample/fallback handling, synthetic scale preservation, and actual Pages startup are tested.

AGENTS.md requires every scoring, interpretation, exclusion, validation, or diagnostic-condition change to add shared regression coverage in both engines, including independent expected values or invariants. The full public sample now runs in CI across all shipped profile/set/mode combinations with independent analysis, matched comparisons, and missing-data comparisons (36 combinations today). Tests discover configuration files instead of assuming their counts.

Weighted components

The YAML catalogue defines positive relative_weight values. PolyMath uses 1, 2, 4 and 8, normalized by their sum of 15, within each language/protocol. Complete groups are averaged within the eval before category/English weighting. Expanded views retain individual raw scores and show component contributions.

Incompatible component matches, language mappings or named-set selections are errors. Missing results in otherwise valid configurations produce diagnostics and exclude the entire affected group from both comparison scores. Component counts and sums come from configuration.

Behavior comparison and intentional fixes

The pre-library baseline is commit 05acf0b on this branch, which already includes the component work. Relative to that baseline, the full local export agrees across 48 combinations of weighting profiles, eval sets, aggregation modes and missing coverage: included rows, scores, weights, contributions, and descriptive category/language trees. Both baseline and candidate pass the public browser suite. The candidate also passes the full-export browser suite, including sorting, scroll preservation, imports, warnings, missing-field highlights and mobile layout.

Relative to main, PolyMath's configured component weighting is an intentional scoring change. The library refactor preserves the existing scoring policy. Additional fixes found during review:

  • Reject whitespace-only config names and categories consistently.
  • Reject trailing newlines in numeric identity fields and language codes.
  • Define a portable regex subset with consistent Unicode/digit/word/whitespace behavior; reject unsupported constructs rather than allowing the engines to interpret them differently.
  • Avoid inherited JavaScript object properties when looking up omitted English shares, including categories named constructor or toString.
  • Replace catalogue-size assertions with input-derived checks and variable-size fixtures. Score calculations contain no fixed eval, category, language or component counts.

Warning presentation can differ; shared tests assert codes, contexts and effects. Existing missing-metric highlighting and manually adjusted category-weight warnings are exercised in browser tests.

Validation

The library revision passed CI; the latest sample and parity additions are checked by the required build on this PR.

  • 39 public Python tests passed, including shared-engine edge cases, randomized configurations, and the full-sample matrix. The optional private-export regression suite also passed during library review.
  • 101 public JavaScript tests passed; the two private-export suites also passed during library review.
  • Public browser checks passed, including the full Pages sample and all synthetic options. The private full-export browser suite passed during library review.
  • Full-export Python/JavaScript semantic parity passed for all 12 shipped profile/set/mode combinations.
  • Before/after dashboard comparison passed for 48 combinations using tests/compare_baseline.cjs.
  • Installed wheel exercised from outside the checkout with Node absent from PATH; CLI stdout/stderr and strict exits tested.
  • Documentation file/heading links, Python 3.8 syntax, formatting/lint checks for the new Python code, and git diff --check passed.

CI runs the public Python/JavaScript cases, an installed-wheel trial without Node on PATH, browser tests, and standalone builds. Only the sanitized, explicitly requested sample export is added; original private source data and generated review artifacts remain untracked.

Documentation impact

Updated README setup/common commands, configuration and development guides, results/sample instructions, and added the Python API/CLI guide and contributor parity requirements. Building from source now requires installing the Python package; using a generated HTML file still needs only a browser.

@jonabur jonabur changed the title Add weighted eval components and PolyMath difficulty aggregation Add native Python analysis, browser parity tests, and weighted eval components Oct 1, 2026
@jonabur
jonabur marked this pull request as ready for review October 2, 2026 06:17
@jonabur
jonabur merged commit c1659fa into main Oct 2, 2026
2 checks passed
@jonabur
jonabur deleted the feature/weighted-eval-components branch October 2, 2026 06:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Provide native Python analysis with shared browser parity tests Support weighted eval components and PolyMath difficulty aggregation

1 participant