Add static prompt quality evaluations - #1441
Merged
Cedric Vidal (cedricvidal) merged 39 commits intoSep 30, 2026
Merged
Cedric Vidal (cedricvidal) merged 39 commits into
Cedric Vidal (cedricvidal) merged 39 commits into
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Retain lower floors and exact mean/case requirements; preserve hash-verified historical policy. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Validate nested labels and finite scores; normalize SDK conversations and recorded tool history without fabricating evidence. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Share versioned decisions between Markdown and review clients, retain policy snapshots and hashes, and explicitly select affected graders while reusing source responses. Document integrity limits and regression coverage. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Emit decision-summary.json for new runs and provide report --decision-output --decision-only outside the source run, retaining hash-verified historical policy identities. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Remove runtime and baseline caps, accept configured pass-rate floors through 100%, and retain exact historical snapshot semantics and policy-change provenance. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Strip standalone forwarded separators before argparse and cover decision-only exports through both direct and pnpm-style arguments. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… early Require expected sample coverage before case verdicts or means, retaining case-level diversity semantics. Resolve source/results paths before creating artifacts and reject equal, nested, or symlink-aliased source destinations. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ostics Use flat name/arguments/tool_call_id fields and validate actual SDK converter inputs before evaluation. Capture SDK run-summary errors for input-only native failures and preserve diagnostic sidecars through replay. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Resolve prompt composition conflicts by retaining shared builders and routing production calls through adaptive chat completion. Preserve both evaluation guidance and upstream contribution instructions. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Declare shared and telemetry workspace dependencies for recursive build ordering. Allow only the verified OpenAPI content fingerprint in the dataset manifest under the generic API key rule. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve redesigned onboarding and evaluation documentation; reconcile the evaluation workspace importer with upstream dependency resolutions. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Cedric Vidal (cedricvidal)
marked this pull request as ready for review
September 29, 2026 05:55
Keep the canonical microsoft/scope execution gate while distinguishing upstream fork-head PRs from fork repository workflows. Run secret-free queue checks on upstream PRs and keep credentials and OIDC out of fork code. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
6 tasks
Restore all unrelated workflow, documentation, dependency, and test changes. The aggregate PR diff now only replaces the retired repository name with microsoft/scope in the four existing gates. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Select integration checks for workflow edits, keep ACP tool checks credential-free on upstream fork PRs, make Docker Hub login optional, and repair demonstrated video, shell, and Windows path failures. Remove only obsolete matrix rows with no worker implementation; retain PR reporting and artifact uploads. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Remove the obsolete VS Code assignments rather than retaining dead case arms. Current ACP tag behavior is unchanged. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Remove internal ACR and cross-repository publication/status automation from OSS while preserving local ACP builds, CLI validation, videos, reporting, Pages and maintenance workflows. Document legacy CLI distribution separately from OSS release publishing. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Validate the Windows base, pinned dependencies and worker on a hosted Windows runner using local images only. Propagate native Dockerfile failures and require Windows validation in CI Summary. Restore manual main-only CLI releases to microsoft/scope with the repository token, serialized version selection, tested artifacts and a public installer/updater destination. Preserve Linux integration jobs and public automation without internal infrastructure. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Bring microsoft#1442 into microsoft#1441 without rewriting either branch. Merge the CI foundation PR first. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Use the canonical public installer throughout onboarding and delegate the website compatibility entry point to it. Preserve API access requirements and verify anonymous installs, failure safety, partial-download rejection and rendered public website links. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Prevent shared server imports from leaking into the bundle and isolate bundle subprocess tests from workspace module resolution. Remove built-in and port-derived API destinations while keeping help, version, and updates usable without configuration. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Prevent shared server imports from leaking into the bundle and isolate bundle subprocess tests from workspace module resolution. Remove built-in and port-derived API destinations while keeping help, version, and updates usable without configuration. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve upstream shared ACP test-utils extraction and auto-labeling. Retain additive test-utils path selection alongside OSS validation filters, and build test-utils before Windows workers with native-command failure checks. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve both the static-prompt evaluation workspace and upstream test-utils package, plus the CI branch's integrated workflow and Windows build changes. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
History-only synchronization after microsoft#1442 was squash-merged; evaluation files unchanged. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Preserve the validated quality path, curated cases, rubrics, dependencies, and historical evidence. Keep red-team work recoverable for a dependent contribution; apply no evaluation review fixes. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Keep OneRAI evidence and known evaluation limitations unchanged; move cloud red-team instructions to the dependent contribution. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Contributor
Author
|
Wassim Chegham (@manekinekko) I've addressed the eight static-prompt quality-evaluation review comments carried over from the now-closed #1398:
Publication status: these fixes are implemented in local commit Production prompt text, curated cases, and score thresholds are unchanged. Existing OneRAI evidence has not been modified or regraded, and no live generation or paid grading was run. The three red-team-specific review comments remain deferred to the separate red-team work. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Replaces closed #1398 with a new PR from the restored fork branch.
The original PR remains closed. The CI foundation #1442 is merged;
its published history is incorporated without changing evaluation content.
--mode quality; removed red-team modes fail before execution.Extraction validation
latencyMs.Evaluation evidence
Quality orchestration runs locally using Azure AI Evaluation SDK graders.
The retained 222-case assessment reuses recorded responses: 221 successful generation rows and one error.
Execution completed, acceptance failed, and assessment integrity remains incomplete.
Blocking gates: 65 passed, 14 failed, 8 unresolved; 2 advisory violations.
No new model generations or paid Azure evaluations were run for this replacement or extraction.
Before this extraction, aggregate pass-rate floors above 80% were lowered in the rubric only.
The engine continues to honor configured rates through 100%; mean-score requirements and lower floors are unchanged.
The extraction makes no additional rubric or evidence changes.
The interactive review extension remains session-scoped, outside the repository.
Known limitations and follow-up
Microsoft.CognitiveServices/accounts/AIServices/evaluations/writepermission.