Repository navigation
Make the evals more realistic and improve the skills accordingly - #39217
AndriySvyryd wants to merge 1 commit into
Conversation
Update the models used for evals Allow to specify the judge effort level for cases where judgement is not deterministic
Agent harness evaluation requiredThis PR affects: |
There was a problem hiding this comment.
🟡 Changes recommended
Required comparison evidence is omitted and fixture provenance hashes are incorrect.
4 open findings
What changed in this PR
Updates the evaluation harness for more realistic fixtures, configurable judge effort, and richer comparison evidence.
Changes:
- Adds comparison-input preparation and completed-trial handling.
- Upgrades evaluation models and refreshes scenarios/fixtures.
- Expands repository skills and associated harness coverage.
| File | Description |
|---|---|
eng/harness-evaluation/test/harness.test.mjs |
Tests validation, judge effort, and completion gates. |
eng/harness-evaluation/test/eval-text-graders.test.mjs |
Empty test file in the reviewed snapshot. |
eng/harness-evaluation/test/comparison.test.mjs |
Tests comparison evidence preparation and safety. |
eng/harness-evaluation/src/harness.mjs |
Adds completed-variant detection and fixture validation. |
eng/harness-evaluation/src/comparison.mjs |
Builds bounded comparison evidence from workspace patches. |
eng/harness-evaluation/src/cli.mjs |
Forwards judge effort and prepares comparisons. |
eng/harness-evaluation/skills/update-pipeline/fixtures/shared-identity/source-excerpts.md |
Documents fixture context and provenance. |
eng/harness-evaluation/skills/update-pipeline/fixtures/shared-identity/proposed/NonSharedModelUpdatesTestBase.cs |
Adds proposed update-pipeline test snapshot. |
eng/harness-evaluation/skills/update-pipeline/fixtures/shared-identity/proposed/ColumnAccessorsFactory.cs |
Adds proposed accessor snapshot. |
eng/harness-evaluation/skills/update-pipeline/fixtures/shared-identity/before/NonSharedModelUpdatesTestBase.cs |
Adds baseline test snapshot. |
eng/harness-evaluation/skills/update-pipeline/fixtures/shared-identity/before/ColumnAccessorsFactory.cs |
Adds baseline accessor snapshot. |
eng/harness-evaluation/skills/update-pipeline/eval.yaml |
Reworks shared-identity evaluation. |
eng/harness-evaluation/skills/triage/eval.yaml |
Updates model defaults. |
eng/harness-evaluation/skills/tooling/eval.yaml |
Updates model defaults. |
eng/harness-evaluation/skills/sqlite-adonet/eval.yaml |
Updates model defaults. |
eng/harness-evaluation/skills/servicing-pr/eval.yaml |
Updates model defaults. |
eng/harness-evaluation/skills/scaffolding/eval.yaml |
Corrects schema-filter evaluation semantics. |
eng/harness-evaluation/skills/run-apichief/eval.yaml |
Updates model defaults. |
eng/harness-evaluation/skills/query-pipeline/fixtures/query-pipeline/split-enumerables.cs |
Adds split-query review fixture. |
eng/harness-evaluation/skills/query-pipeline/fixtures/query-pipeline/projection-retry.cs |
Adds projection retry fixture. |
eng/harness-evaluation/skills/query-pipeline/fixtures/query-pipeline/fallback-and-tests.cs |
Adds fallback and concurrency fixture. |
eng/harness-evaluation/skills/query-pipeline/fixtures/provider-test-placement/source-context.md |
Supplies provider-test hierarchy context. |
eng/harness-evaluation/skills/query-pipeline/fixtures/provider-test-placement/proposed/NorthwindMiscellaneousQuerySqlServerTest.cs |
Adds proposed provider-only test snapshot. |
eng/harness-evaluation/skills/query-pipeline/fixtures/provider-test-placement/before/NorthwindMiscellaneousQueryTestBase.cs |
Adds specification baseline snapshot. |
eng/harness-evaluation/skills/query-pipeline/fixtures/provider-test-placement/before/NorthwindMiscellaneousQuerySqlServerTest.cs |
Adds SQL Server baseline snapshot. |
eng/harness-evaluation/skills/query-pipeline/fixtures/provider-test-placement/before/NorthwindMiscellaneousQueryRelationalTestBase.cs |
Adds relational baseline snapshot. |
eng/harness-evaluation/skills/query-pipeline/eval.yaml |
Introduces query-pipeline skill evaluation. |
eng/harness-evaluation/skills/model-building/eval.yaml |
Updates model defaults. |
eng/harness-evaluation/skills/migrations/fixtures/upgrade-report.txt |
Expands the upgrade regression report. |
eng/harness-evaluation/skills/migrations/fixtures/upgrade-only/source-context.md |
Documents migration-diff code paths. |
eng/harness-evaluation/skills/migrations/fixtures/upgrade-only/SnapshotModelProcessor.cs |
Adds reduced processor snapshot. |
eng/harness-evaluation/skills/migrations/fixtures/upgrade-only/CustomerContextModelSnapshot.cs |
Adds historical model snapshot. |
eng/harness-evaluation/skills/migrations/fixtures/upgrade-only/CustomerContext.cs |
Adds current model fixture. |
eng/harness-evaluation/skills/migrations/eval.yaml |
Reworks upgrade compatibility evaluation. |
eng/harness-evaluation/skills/make-skill/eval.yaml |
Updates model defaults. |
eng/harness-evaluation/skills/make-instructions/eval.yaml |
Focuses evaluation on proposal review. |
eng/harness-evaluation/skills/make-github-actions-workflow/eval.yaml |
Focuses evaluation on external PR policy. |
eng/harness-evaluation/skills/make-custom-agent/eval.yaml |
Focuses evaluation on agent boundaries. |
eng/harness-evaluation/skills/cosmos-provider/fixtures/CosmosTransactionalBatchTest.cs |
Adds Cosmos transactional-batch fixture. |
eng/harness-evaluation/skills/cosmos-provider/eval.yaml |
Replaces analysis with offline test generation. |
eng/harness-evaluation/skills/code-review/fixtures/dbset-sequence/DbSetConsumer.cs |
Adds binary compatibility consumer fixture. |
eng/harness-evaluation/skills/code-review/fixtures/concurrency-detector/proposed/ConcurrencyDetectorTest.cs |
Adds proposed concurrency tests. |
eng/harness-evaluation/skills/code-review/fixtures/concurrency-detector/proposed/ConcurrencyDetector.cs |
Adds proposed detector implementation. |
eng/harness-evaluation/skills/code-review/fixtures/concurrency-detector/before/ConcurrencyDetectorTest.cs |
Adds baseline concurrency tests. |
eng/harness-evaluation/skills/code-review/fixtures/concurrency-detector/before/ConcurrencyDetector.cs |
Adds baseline detector implementation. |
eng/harness-evaluation/skills/code-review/eval.yaml |
Reworks compatibility and concurrency reviews. |
eng/harness-evaluation/skills/change-tracking/eval.yaml |
Updates model defaults. |
eng/harness-evaluation/README.md |
Documents snapshots, comparison evidence, and judge controls. |
eng/harness-evaluation/instructions/copilot-instructions/fixtures/query-pipeline/projection-retry.cs |
Expands instruction fixture behavior. |
eng/harness-evaluation/instructions/copilot-instructions/fixtures/activation-proposal.cs |
Adds API implementation proposal. |
eng/harness-evaluation/instructions/copilot-instructions/eval.yaml |
Replaces planning with implementation evaluation. |
.github/copilot-instructions.md |
Refines shared test-infrastructure guidance. |
.agents/skills/update-pipeline/SKILL.md |
Updates shared-row value-access guidance. |
.agents/skills/scaffolding/SKILL.md |
Corrects anti-semi-join guidance. |
.agents/skills/query-pipeline/SKILL.md |
Adds query-pipeline domain guidance. |
.agents/skills/migrations/SKILL.md |
Clarifies historical metadata normalization. |
.agents/skills/make-instructions/SKILL.md |
Adds initial-authoring validation order. |
.agents/skills/make-github-actions-workflow/SKILL.md |
Tightens bot-comment ownership checks. |
.agents/skills/make-custom-agent/SKILL.md |
Adds proposal-review boundary guidance. |
🧠 Review effort: Balanced
|
Addressed the four substantive review findings in the local worktree (not committed or pushed):
The eight Validation: all 47 harness tests pass, strict Vally lint passes for all 18 components, and |


Update the models used for evals
Allow to specify the judge effort level for cases where judgement is not deterministic