Skip to content

Make the evals more realistic and improve the skills accordingly - #39217

Draft
AndriySvyryd wants to merge 1 commit into
mainfrom
ImproveSkills
Draft

AndriySvyryd wants to merge 1 commit into
mainfrom
ImproveSkills

Conversation

@AndriySvyryd

Copy link
Copy Markdown
Member

Update the models used for evals
Allow to specify the judge effort level for cases where judgement is not deterministic

Update the models used for evals
Allow to specify the judge effort level for cases where judgement is not deterministic
@AndriySvyryd
AndriySvyryd requested a balanced review from Copilot October 10, 2026 00:22
@github-actions

Copy link
Copy Markdown
Contributor

Agent harness evaluation required

This PR affects: change-tracking, code-review, copilot-instructions, cosmos-provider, make-custom-agent, make-github-actions-workflow, make-instructions, make-skill, migrations, model-building, query-pipeline, run-apichief, scaffolding, servicing-pr, sqlite-adonet, tooling, triage, update-pipeline. Its author cannot use repository secrets in an automatic run. A contributor with write access can comment /eval or /eval component-name, or run the workflow manually for pull request #39217.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Required comparison evidence is omitted and fixture provenance hashes are incorrect.

4 open findings
What changed in this PR

Updates the evaluation harness for more realistic fixtures, configurable judge effort, and richer comparison evidence.

Changes:

  • Adds comparison-input preparation and completed-trial handling.
  • Upgrades evaluation models and refreshes scenarios/fixtures.
  • Expands repository skills and associated harness coverage.
File Description
eng/​harness-evaluation/​test/​harness.test.mjs Tests validation, judge effort, and completion gates.
eng/​harness-evaluation/​test/​eval-text-graders.test.mjs Empty test file in the reviewed snapshot.
eng/​harness-evaluation/​test/​comparison.test.mjs Tests comparison evidence preparation and safety.
eng/​harness-evaluation/​src/​harness.mjs Adds completed-variant detection and fixture validation.
eng/​harness-evaluation/​src/​comparison.mjs Builds bounded comparison evidence from workspace patches.
eng/​harness-evaluation/​src/​cli.mjs Forwards judge effort and prepares comparisons.
eng/​harness-evaluation/​skills/​update-pipeline/​fixtures/​shared-identity/​source-excerpts.md Documents fixture context and provenance.
eng/​harness-evaluation/​skills/​update-pipeline/​fixtures/​shared-identity/​proposed/​NonSharedModelUpdatesTestBase.cs Adds proposed update-pipeline test snapshot.
eng/​harness-evaluation/​skills/​update-pipeline/​fixtures/​shared-identity/​proposed/​ColumnAccessorsFactory.cs Adds proposed accessor snapshot.
eng/​harness-evaluation/​skills/​update-pipeline/​fixtures/​shared-identity/​before/​NonSharedModelUpdatesTestBase.cs Adds baseline test snapshot.
eng/​harness-evaluation/​skills/​update-pipeline/​fixtures/​shared-identity/​before/​ColumnAccessorsFactory.cs Adds baseline accessor snapshot.
eng/​harness-evaluation/​skills/​update-pipeline/​eval.yaml Reworks shared-identity evaluation.
eng/​harness-evaluation/​skills/​triage/​eval.yaml Updates model defaults.
eng/​harness-evaluation/​skills/​tooling/​eval.yaml Updates model defaults.
eng/​harness-evaluation/​skills/​sqlite-adonet/​eval.yaml Updates model defaults.
eng/​harness-evaluation/​skills/​servicing-pr/​eval.yaml Updates model defaults.
eng/​harness-evaluation/​skills/​scaffolding/​eval.yaml Corrects schema-filter evaluation semantics.
eng/​harness-evaluation/​skills/​run-apichief/​eval.yaml Updates model defaults.
eng/​harness-evaluation/​skills/​query-pipeline/​fixtures/​query-pipeline/​split-enumerables.cs Adds split-query review fixture.
eng/​harness-evaluation/​skills/​query-pipeline/​fixtures/​query-pipeline/​projection-retry.cs Adds projection retry fixture.
eng/​harness-evaluation/​skills/​query-pipeline/​fixtures/​query-pipeline/​fallback-and-tests.cs Adds fallback and concurrency fixture.
eng/​harness-evaluation/​skills/​query-pipeline/​fixtures/​provider-test-placement/​source-context.md Supplies provider-test hierarchy context.
eng/​harness-evaluation/​skills/​query-pipeline/​fixtures/​provider-test-placement/​proposed/​NorthwindMiscellaneousQuerySqlServerTest.cs Adds proposed provider-only test snapshot.
eng/​harness-evaluation/​skills/​query-pipeline/​fixtures/​provider-test-placement/​before/​NorthwindMiscellaneousQueryTestBase.cs Adds specification baseline snapshot.
eng/​harness-evaluation/​skills/​query-pipeline/​fixtures/​provider-test-placement/​before/​NorthwindMiscellaneousQuerySqlServerTest.cs Adds SQL Server baseline snapshot.
eng/​harness-evaluation/​skills/​query-pipeline/​fixtures/​provider-test-placement/​before/​NorthwindMiscellaneousQueryRelationalTestBase.cs Adds relational baseline snapshot.
eng/​harness-evaluation/​skills/​query-pipeline/​eval.yaml Introduces query-pipeline skill evaluation.
eng/​harness-evaluation/​skills/​model-building/​eval.yaml Updates model defaults.
eng/​harness-evaluation/​skills/​migrations/​fixtures/​upgrade-report.txt Expands the upgrade regression report.
eng/​harness-evaluation/​skills/​migrations/​fixtures/​upgrade-only/​source-context.md Documents migration-diff code paths.
eng/​harness-evaluation/​skills/​migrations/​fixtures/​upgrade-only/​SnapshotModelProcessor.cs Adds reduced processor snapshot.
eng/​harness-evaluation/​skills/​migrations/​fixtures/​upgrade-only/​CustomerContextModelSnapshot.cs Adds historical model snapshot.
eng/​harness-evaluation/​skills/​migrations/​fixtures/​upgrade-only/​CustomerContext.cs Adds current model fixture.
eng/​harness-evaluation/​skills/​migrations/​eval.yaml Reworks upgrade compatibility evaluation.
eng/​harness-evaluation/​skills/​make-skill/​eval.yaml Updates model defaults.
eng/​harness-evaluation/​skills/​make-instructions/​eval.yaml Focuses evaluation on proposal review.
eng/​harness-evaluation/​skills/​make-github-actions-workflow/​eval.yaml Focuses evaluation on external PR policy.
eng/​harness-evaluation/​skills/​make-custom-agent/​eval.yaml Focuses evaluation on agent boundaries.
eng/​harness-evaluation/​skills/​cosmos-provider/​fixtures/​CosmosTransactionalBatchTest.cs Adds Cosmos transactional-batch fixture.
eng/​harness-evaluation/​skills/​cosmos-provider/​eval.yaml Replaces analysis with offline test generation.
eng/​harness-evaluation/​skills/​code-review/​fixtures/​dbset-sequence/​DbSetConsumer.cs Adds binary compatibility consumer fixture.
eng/​harness-evaluation/​skills/​code-review/​fixtures/​concurrency-detector/​proposed/​ConcurrencyDetectorTest.cs Adds proposed concurrency tests.
eng/​harness-evaluation/​skills/​code-review/​fixtures/​concurrency-detector/​proposed/​ConcurrencyDetector.cs Adds proposed detector implementation.
eng/​harness-evaluation/​skills/​code-review/​fixtures/​concurrency-detector/​before/​ConcurrencyDetectorTest.cs Adds baseline concurrency tests.
eng/​harness-evaluation/​skills/​code-review/​fixtures/​concurrency-detector/​before/​ConcurrencyDetector.cs Adds baseline detector implementation.
eng/​harness-evaluation/​skills/​code-review/​eval.yaml Reworks compatibility and concurrency reviews.
eng/​harness-evaluation/​skills/​change-tracking/​eval.yaml Updates model defaults.
eng/​harness-evaluation/​README.md Documents snapshots, comparison evidence, and judge controls.
eng/​harness-evaluation/​instructions/​copilot-instructions/​fixtures/​query-pipeline/​projection-retry.cs Expands instruction fixture behavior.
eng/​harness-evaluation/​instructions/​copilot-instructions/​fixtures/​activation-proposal.cs Adds API implementation proposal.
eng/​harness-evaluation/​instructions/​copilot-instructions/​eval.yaml Replaces planning with implementation evaluation.
.github/​copilot-instructions.md Refines shared test-infrastructure guidance.
.agents/​skills/​update-pipeline/​SKILL.md Updates shared-row value-access guidance.
.agents/​skills/​scaffolding/​SKILL.md Corrects anti-semi-join guidance.
.agents/​skills/​query-pipeline/​SKILL.md Adds query-pipeline domain guidance.
.agents/​skills/​migrations/​SKILL.md Clarifies historical metadata normalization.
.agents/​skills/​make-instructions/​SKILL.md Adds initial-authoring validation order.
.agents/​skills/​make-github-actions-workflow/​SKILL.md Tightens bot-comment ownership checks.
.agents/​skills/​make-custom-agent/​SKILL.md Adds proposal-review boundary guidance.

🧠 Review effort: Balanced

Comment thread .agents/skills/query-pipeline/SKILL.md
Comment thread eng/harness-evaluation/instructions/copilot-instructions/eval.yaml
Comment thread .agents/skills/update-pipeline/SKILL.md
@github-code-quality

Copy link
Copy Markdown

Code Coverage Overview

Languages: C#

C# / code-coverage

The overall line coverage in commit c2b520d in the ImproveSkills branch remains at 76%, unchanged from commit 2b99285 in the main branch.

@AndriySvyryd

Copy link
Copy Markdown
Member Author

Addressed the four substantive review findings in the local worktree (not committed or pushed):

  • Added review-split-query-completion-change, staging the formerly unused split-enumerable fixture plus a self-contained proposal. Its rubric covers both consumers and sync/async paths, grouped lookahead versus outer exhaustion, empty results, early disposal, caller/internal buffering, and consistency across separate commands.
  • Added an exact-path file-matches grader for src/EFCore/Properties/CoreStrings.resx. A regression test loads the actual async-API eval and verifies resource edits reach comparison evidence in both answer orders.
  • Corrected “original valued” to “original value.”
  • Removed the stale fixture hash table and unstaged-before provenance references. Fixture context now uses source-path provenance without commit IDs, commit URLs, or revision-fetch instructions; documented this authoring rule.

The eight Where/Select suggestions on copied CommandBatchPreparer and InternalDbSet snapshots are intentionally not applied. These are existing hot-path loops, not the behavior proposed by these fixtures. LINQ wrappers add iterator/delegate overhead without correcting a defect; changing both snapshot versions would also introduce unrelated refactoring into the eval inputs. The direct loops preserve the existing per-entry state transitions and batching boundaries.

Validation: all 47 harness tests pass, strict Vally lint passes for all 18 components, and git diff --check is clean. No model-based behavioral evals were rerun for this review-fix pass.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants