Skip to content

Add static prompt quality evaluations - #1441

Merged
Cedric Vidal (cedricvidal) merged 39 commits into
microsoft:mainfrom
cedricvidal:cedricvidal-plan-static-prompt-evals
Sep 30, 2026
Merged

Cedric Vidal (cedricvidal) merged 39 commits into
microsoft:mainfrom
cedricvidal:cedricvidal-plan-static-prompt-evals

Conversation

@cedricvidal

@cedricvidal Cedric Vidal (cedricvidal) commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Replaces closed #1398 with a new PR from the restored fork branch.
The original PR remains closed. The CI foundation #1442 is merged;
its published history is incorporated without changing evaluation content.

  • Add production-backed quality evaluations for 10 static prompt families and 15 adapter variants.
  • Commit 222 curated input cases; keep generated responses and evaluation artifacts ignored.
  • Add rubric-driven grading, gate-based reports, historical rubric snapshots, immutable selective regrading, and case-level evidence.
  • Retain the previously implemented native grader parsing, task context, SDK tool-call schemas, incomplete-sample handling, and replay source protection.
  • Extract unused cloud red-team orchestration into a separate dependent contribution. The quality runner defaults to quality and retains explicit --mode quality; removed red-team modes fail before execution.
  • Preserve normal evaluation logic, prompts, rubrics, datasets, dependency versions, and OneRAI-shared evidence. The extraction applies none of the review fixes from Add static prompt evaluation and red-team framework #1398.

Extraction validation

  • TypeScript typecheck passed; 32 remaining TypeScript tests and 94 Python tests passed.
  • Dataset validation: 222 approved cases, 10 quality families, 15 quality targets.
  • All 1,106 protected quality/dependency/evidence file hashes match the preservation snapshot.
  • Pre/post fake-transport comparison: all 666 rows match across 222 cases and three samples, excluding only latencyMs.
  • No live model generation, paid Azure evaluation, or assessment regrading was run.
  • No generated results or unrelated local editor settings are included.
  • CI for the newly published head must be assessed separately.

Evaluation evidence

Quality orchestration runs locally using Azure AI Evaluation SDK graders.
The retained 222-case assessment reuses recorded responses: 221 successful generation rows and one error.
Execution completed, acceptance failed, and assessment integrity remains incomplete.
Blocking gates: 65 passed, 14 failed, 8 unresolved; 2 advisory violations.
No new model generations or paid Azure evaluations were run for this replacement or extraction.

Before this extraction, aggregate pass-rate floors above 80% were lowered in the rubric only.
The engine continues to honor configured rates through 100%; mean-score requirements and lower floors are unchanged.
The extraction makes no additional rubric or evidence changes.
The interactive review extension remains session-scoped, outside the repository.

Known limitations and follow-up

  • Suggested-feature novelty receives existing-feature detection output instead of generated suggestions in 28 of 30 current cases; these scores are not reliable novelty assessments.
  • Feature-authoring candidate checking can omit child references because it selects the first configured field instead of combining parent and child IDs.
  • Variation duplicate checking can omit the original existingPrompt and does not combine all relevant prompt collections.
  • Dependency references need validation and explicit authority/completeness rules before adopting deterministic set-based grading.
  • Feature-extraction mean requirements of 1.0 still require perfection despite separate 80% pass-rate floors. Custom 1-5 rubrics need behavioral score anchors.
  • Missing original judge tool traces, one report-generation error, and incomplete generator provenance remain visible limitations; no evidence was fabricated.
  • The eight quality review findings from Add static prompt evaluation and red-team framework #1398 remain deferred; extraction is not remediation.
  • Quality results are not a completed safety/red-team assessment or a clean pass. Red teaming was not used for the OneRAI evaluation; its previous live baseline was blocked by missing Foundry Microsoft.CognitiveServices/accounts/AIServices/evaluations/write permission.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Retain lower floors and exact mean/case requirements; preserve hash-verified historical policy.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Validate nested labels and finite scores; normalize SDK conversations and recorded tool history without fabricating evidence.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Share versioned decisions between Markdown and review clients, retain policy snapshots and hashes, and explicitly select affected graders while reusing source responses. Document integrity limits and regression coverage.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Emit decision-summary.json for new runs and provide report --decision-output --decision-only outside the source run, retaining hash-verified historical policy identities.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Remove runtime and baseline caps, accept configured pass-rate floors through 100%, and retain exact historical snapshot semantics and policy-change provenance.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Strip standalone forwarded separators before argparse and cover decision-only exports through both direct and pnpm-style arguments.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… early

Require expected sample coverage before case verdicts or means, retaining case-level diversity semantics. Resolve source/results paths before creating artifacts and reject equal, nested, or symlink-aliased source destinations.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ostics

Use flat name/arguments/tool_call_id fields and validate actual SDK converter inputs before evaluation. Capture SDK run-summary errors for input-only native failures and preserve diagnostic sidecars through replay.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Resolve prompt composition conflicts by retaining shared builders and routing production calls through adaptive chat completion. Preserve both evaluation guidance and upstream contribution instructions.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Declare shared and telemetry workspace dependencies for recursive build ordering. Allow only the verified OpenAPI content fingerprint in the dataset manifest under the generic API key rule.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve redesigned onboarding and evaluation documentation; reconcile the evaluation workspace importer with upstream dependency resolutions.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@cedricvidal
Cedric Vidal (cedricvidal) marked this pull request as ready for review September 29, 2026 05:55
Keep the canonical microsoft/scope execution gate while distinguishing upstream fork-head PRs from fork repository workflows. Run secret-free queue checks on upstream PRs and keep credentials and OIDC out of fork code.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Restore all unrelated workflow, documentation, dependency, and test changes. The aggregate PR diff now only replaces the retired repository name with microsoft/scope in the four existing gates.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Select integration checks for workflow edits, keep ACP tool checks credential-free on upstream fork PRs, make Docker Hub login optional, and repair demonstrated video, shell, and Windows path failures. Remove only obsolete matrix rows with no worker implementation; retain PR reporting and artifact uploads.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Remove the obsolete VS Code assignments rather than retaining dead case arms. Current ACP tag behavior is unchanged.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Remove internal ACR and cross-repository publication/status automation from OSS while preserving local ACP builds, CLI validation, videos, reporting, Pages and maintenance workflows. Document legacy CLI distribution separately from OSS release publishing.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Validate the Windows base, pinned dependencies and worker on a hosted Windows runner using local images only. Propagate native Dockerfile failures and require Windows validation in CI Summary.

Restore manual main-only CLI releases to microsoft/scope with the repository token, serialized version selection, tested artifacts and a public installer/updater destination. Preserve Linux integration jobs and public automation without internal infrastructure.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Bring microsoft#1442 into microsoft#1441 without rewriting either branch. Merge the CI foundation PR first.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use the canonical public installer throughout onboarding and delegate the website compatibility entry point to it. Preserve API access requirements and verify anonymous installs, failure safety, partial-download rejection and rendered public website links.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Prevent shared server imports from leaking into the bundle and isolate bundle subprocess tests from workspace module resolution. Remove built-in and port-derived API destinations while keeping help, version, and updates usable without configuration.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Prevent shared server imports from leaking into the bundle and isolate bundle subprocess tests from workspace module resolution. Remove built-in and port-derived API destinations while keeping help, version, and updates usable without configuration.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve upstream shared ACP test-utils extraction and auto-labeling. Retain additive test-utils path selection alongside OSS validation filters, and build test-utils before Windows workers with native-command failure checks.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve both the static-prompt evaluation workspace and upstream test-utils package, plus the CI branch's integrated workflow and Windows build changes.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
History-only synchronization after microsoft#1442 was squash-merged; evaluation files unchanged.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve the validated quality path, curated cases, rubrics, dependencies, and historical evidence. Keep red-team work recoverable for a dependent contribution; apply no evaluation review fixes.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep OneRAI evidence and known evaluation limitations unchanged; move cloud red-team instructions to the dependent contribution.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@cedricvidal Cedric Vidal (cedricvidal) changed the title Add static prompt evaluation and red-team framework Add static prompt quality evaluations Sep 30, 2026
@cedricvidal

Cedric Vidal (cedricvidal) commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor Author

Wassim Chegham (@manekinekko) I've addressed the eight static-prompt quality-evaluation review comments carried over from the now-closed #1398:

  • Separate generator configuration from grader configuration.
  • Treat generation infrastructure failures as unresolved rather than prompt-quality failures, including sample diversity.
  • Retain judge completion metadata and associate tool traces with their requests.
  • Include generated feature suggestions in novelty-grader inputs.
  • Allow explicitly selected regrading of missing or unusable retained assessments.
  • Accept empty historical observation streams during replay.
  • Reject nonexistent explicit dataset paths instead of silently loading the full manifest.
  • Accept valid empty feature-catalog outputs while retaining feature-coverage checks.

Publication status: these fixes are implemented in local commit a1c23a94 but have not yet been pushed, so they are not in the current PR head (781b4c57).

Production prompt text, curated cases, and score thresholds are unchanged. Existing OneRAI evidence has not been modified or regraded, and no live generation or paid grading was run. The three red-team-specific review comments remain deferred to the separate red-team work.

@cedricvidal
Cedric Vidal (cedricvidal) merged commit 395969f into microsoft:main Sep 30, 2026
23 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant