Improve template and JSON skill guidance - #1097
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Skill Coverage Report
Uncovered:
|
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review tier: Lite
Findings: 2
New issues introduced by this change (2)
| Severity | Finding |
|---|---|
plugins/dotnet-template-engine/skills/template-validation/SKILL.md — Step 2 says to report a malformed JSON parse error with only a line number, but the new parse… |
|
plugins/dotnet-template-engine/skills/template-smart-defaults/SKILL.md — The validation checklist still says advice-only requests should be flagged as "to-confirm" rather… |
What changed in this PR
This PR updates skill guidance for the .NET 11 System.Text.Json skill and several dotnet new template-engine skills to reduce common failure modes observed in cross-model evaluations (e.g., continuing semantic validation after JSON parse errors, guessing template flags, and emitting unsafe commands for no-write advice).
Changes:
- Strengthens the .NET 11 System.Text.Json skill with explicit
DictionaryKeyPolicyguidance and a non-admin, localdotnet-installfallback when execution is requested. - Adds a “parse gate” to template validation (stop after syntax errors) and tightens constraint/schema guidance (notably
hostconstraints). - Requires grounding option claims in observed
dotnet new <template> --helpoutput, and makes “no-write” advice commands safe via--dry-run; improves CPM-aware instantiation guidance by avoiding restore before centralization.
| File | Description |
|---|---|
| plugins/dotnet11/skills/system-text-json-net11/SKILL.md | Adds net11 execution enablement guidance (local SDK install) and clarifies PascalCase for dictionary keys via DictionaryKeyPolicy. |
| plugins/dotnet-template-engine/skills/template-validation/SKILL.md | Introduces a parse-first gate and expands rules for constraints/schema-sensitive validation. |
| plugins/dotnet-template-engine/skills/template-smart-defaults/SKILL.md | Requires --help grounding for exact commands and enforces --dry-run for no-write advice commands. |
| plugins/dotnet-template-engine/skills/template-instantiation/SKILL.md | Adds a concise situation→action table and improves CPM flow by preferring --no-restore when available. |
| plugins/dotnet-template-engine/skills/template-discovery/SKILL.md | Tightens “inspection requires inspection” and reinforces --dry-run for no-write command output. |
| plugins/dotnet-template-engine/skills/template-comparison/SKILL.md | Establishes an evidence contract for side-by-side comparisons and avoids unapproved installs. |
| plugins/dotnet-template-engine/skills/template-authoring/SKILL.md | Adds schema-accurate authoring patterns for XML conditionals, restore actions, and host constraints. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
Review tier: Lite
Findings: None
Issues resolved since last review (2)
| Severity | Finding |
|---|---|
plugins/dotnet-template-engine/skills/template-smart-defaults/SKILL.md — The validation checklist still says advice-only requests should be flagged as "to-confirm" rather… View resolved comment |
|
plugins/dotnet-template-engine/skills/template-validation/SKILL.md — Step 2 says to report a malformed JSON parse error with only a line number, but the new parse… View resolved comment |
📊 Skill Evaluation Results14 model/skill results across 7 skills and 2 models — ✅ 2 improved, ➖ 11 not proven improved, Measurement identity: evaluated commit Measurement health: 14 expected / 14 observed / 14 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — system-text-json-net11 (claude-sonnet-4.6)Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +75.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Moderate (score 0.36) Repeated-run reliability (not used by the gate): 8 paired runs (6W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — system-text-json-net11 (gpt-5.6-luna)Why: Net win +66.7% (4W/2T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +32.5% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 4W/2T/0L; d=4; p=0.063; net +66.7%; 2 dormancy excluded Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-authoring (claude-sonnet-4.6)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +40.0% across 7 paired run(s) — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1% Overfit: Moderate (score 0.30) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-authoring (gpt-5.6-luna)Why: Net win -14.3% (2W/2T/3L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s) — no improvement Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 2W/2T/3L; d=5; p=0.500; net -14.3% Overfit: Low (score 0.08) Repeated-run reliability (not used by the gate): 7 paired runs (2W/2T/3L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-comparison (claude-sonnet-4.6)Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +17.1% across 7 paired run(s) — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9% Overfit: Moderate (score 0.26) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-comparison (gpt-5.6-luna)Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +17.1% across 7 paired run(s) — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9% Overfit: Low (score 0.12) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-discovery (gpt-5.6-luna)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +22.9% across 7 paired run(s) — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1% Overfit: Low (score 0.16) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-instantiation (claude-sonnet-4.6)Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3% Overfit: Moderate (score 0.27) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-instantiation (gpt-5.6-luna)Why: Net win -14.3% (0W/6T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference -5.7% across 7 paired run(s) — no improvement Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 0W/6T/1L; d=1; p=0.500; net -14.3% Overfit: Low (score 0.15) Repeated-run reliability (not used by the gate): 7 paired runs (0W/6T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-smart-defaults (gpt-5.6-luna)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +31.4% across 7 paired run(s) — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1% Overfit: Moderate (score 0.24) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-validation (claude-sonnet-4.6)Why: Net win +71.4% (6W/0T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +54.3% across 7 paired run(s) — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 6W/0T/1L; d=7; p=0.063; net +71.4% Overfit: Moderate (score 0.49) Repeated-run reliability (not used by the gate): 7 paired runs (6W/0T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-validation (gpt-5.6-luna)Why: Net win +71.4% (6W/0T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +54.3% across 7 paired run(s) — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 6W/0T/1L; d=7; p=0.063; net +71.4% Overfit: Moderate (score 0.20) Repeated-run reliability (not used by the gate): 7 paired runs (6W/0T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — template-discovery (claude-sonnet-4.6)Why: Net win +71.4% (5W/2T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +45.7% across 7 paired run(s) — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 5W/2T/0L; d=5; p=0.031; net +71.4% Overfit: Moderate (score 0.37) Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — template-smart-defaults (claude-sonnet-4.6)Why: Net win +85.7% (6W/1T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +68.6% across 7 paired run(s) — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 6W/1T/0L; d=6; p=0.016; net +85.7% Overfit: Moderate (score 0.43) Repeated-run reliability (not used by the gate): 7 paired runs (6W/1T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. 🔍 Full Results - all metrics and investigation details
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review tier: Lite
Findings: 2
New issues introduced by this change (2)
| Severity | Finding |
|---|---|
plugins/dotnet11/skills/system-text-json-net11/SKILL.md — Step 0 says to install the “current .NET 11 preview” when a net11-capable SDK is missing. For… |
|
plugins/dotnet-template-engine/skills/template-validation/SKILL.md — The parse gate requires malformed-JSON responses to have exactly two parts (a one-line verdict + a… |
Suppressed comments (2)
Previously missed (2) — in code that hasn't changed since the last review.
plugins/dotnet11/skills/system-text-json-net11/SKILL.md:51
- The PowerShell example assumes
./dotnet-install.ps1already exists in the working directory. That’s not a safe assumption; download the official script from dot.net before invoking it so the steps are runnable as written.
```powershell
./dotnet-install.ps1 -Channel 11.0 -Quality preview -InstallDir ./.dotnet
./.dotnet/dotnet run --project ./Sample
**plugins/dotnet11/skills/system-text-json-net11/SKILL.md:56**
* The bash example assumes `./dotnet-install.sh` already exists. Download the official script from dot.net (or otherwise explain where it comes from) so the instructions are self-contained and runnable.
./dotnet-install.sh --channel 11.0 --quality preview --install-dir ./.dotnet
./.dotnet/dotnet run --project ./Sample</details>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Once you've addressed the issues Copilot identified, you can request another Copilot review.
Review tier: Lite
Findings: 1
New issues introduced by this change (1)
| Severity | Finding |
|---|---|
plugins/dotnet11/skills/system-text-json-net11/SKILL.md — The PowerShell snippet assumes dotnet-install.ps1 is already present in the working directory and… |
Issues resolved since last review (2)
| Severity | Finding |
|---|---|
plugins/dotnet-template-engine/skills/template-validation/SKILL.md — The parse gate requires malformed-JSON responses to have exactly two parts (a one-line verdict + a… View resolved comment |
|
plugins/dotnet11/skills/system-text-json-net11/SKILL.md — Step 0 says to install the “current .NET 11 preview” when a net11-capable SDK is missing. For… View resolved comment |
Suppressed comments (1)
plugins/dotnet11/skills/system-text-json-net11/SKILL.md:56
- The bash snippet assumes
dotnet-install.shis already present and also references./Samplewithout defining it anywhere in this skill. Download the script from the official URL and clarify the project path placeholder so the instructions are actually runnable.
./dotnet-install.sh --channel 11.0 --install-dir ./.dotnet
./.dotnet/dotnet run --project ./Sample
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Review tier: Lite
Findings: None
Issues resolved since last review (2)
| Severity | Finding |
|---|---|
tests/dotnet-template-engine/template-smart-defaults/eval.yaml — This framework negative check is anchored to lines that start with dotnet new, but the same… View resolved comment |
|
tests/dotnet-template-engine/template-smart-defaults/eval.yaml — The eval’s --dry-run matcher explicitly allows backtick-wrapped dotnet new commands, but this… View resolved comment |
Suppressed comments (7)
Previously missed (3) — in code that hasn't changed since the last review.
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:19
- The first stimulus asks for a worker service, but the graders don’t assert that the generated command uses the
workertemplate. As written, an answer could emit adotnet newcommand for a different template and still pass as long as it includes--dry-runand mentions AOT/framework.
This issue also appears in the following locations of the same file:
- line 38
- line 67
- type: output-matches
config:
pattern: (?i)net(8|9|10|11)\.0
- type: output-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s--dry-run\b|`dotnet\s+new\s+\S+[^`\r\n#]*\s--dry-run\b[^`\r\n#]*`)
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:22
- The “no unsafe
dotnet newcommand” negative regex only detects commands that start the line (or follow;/&/|) and therefore can be bypassed by prefixing an unsafe command with common markdown/shell markers (e.g.- dotnet new ...,1. dotnet new ...,$ dotnet new ...). That can allow a response to include an additional unsafe creation command without failing this grader.
This issue also appears in the following locations of the same file:
- line 147
- line 174
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?!(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s(?:--dry-run|--help)\b)(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*|`dotnet\s+new\s+\S+(?![^`\r\n#]*\s(?:--dry-run|--help)\b)[^`\r\n#]*`)
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:103
- This negative grader is intended to reject commands that override the user’s explicit net8.0 choice, but it only matches when the command starts the line with
dotnet new. If the command is formatted as a list item (e.g.- dotnet new ... --framework net9.0), it won’t be detected and the scenario could false-pass.
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*(?:--framework|-f)[ =]+net(9|10|11)\.0|`dotnet\s+new[^`\r\n#]*(?:--framework|-f)[ =]+net(9|10|11)\.0[^`\r\n#]*`)
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:51
- The “Auth implies HTTPS stays enabled” stimulus asks for a
webapicommand, but the graders never assert the template short name. A response that produces a non-webapi command (still containingSingleOrgand--dry-run) could incorrectly pass this scenario.
graders:
- type: exit-success
- type: output-matches
config:
pattern: (?i)SingleOrg
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*--no-https\b|`dotnet\s+new[^`\r\n#]*--no-https\b[^`\r\n#]*`)
- type: output-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s--dry-run\b|`dotnet\s+new\s+\S+[^`\r\n#]*\s--dry-run\b[^`\r\n#]*`)
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?!(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s(?:--dry-run|--help)\b)(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*|`dotnet\s+new\s+\S+(?![^`\r\n#]*\s(?:--dry-run|--help)\b)[^`\r\n#]*`)
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:80
- The “Controllers exclude the minimal-API flag” stimulus asks for a
webapicommand, but the graders don’t currently assert the template short name. Adding an explicitdotnet new webapimatch would prevent false passes where the response discusses controllers but uses a different template.
graders:
- type: exit-success
- type: output-matches
config:
pattern: (?i)(use-controllers|controllers)
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*--[a-z-]*minimal|`dotnet\s+new[^`\r\n#]*--[a-z-]*minimal[^`\r\n#]*`)
- type: output-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s--dry-run\b|`dotnet\s+new\s+\S+[^`\r\n#]*\s--dry-run\b[^`\r\n#]*`)
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?!(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s(?:--dry-run|--help)\b)(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*|`dotnet\s+new\s+\S+(?![^`\r\n#]*\s(?:--dry-run|--help)\b)[^`\r\n#]*`)
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:185
- In the minimal-API stimulus, both the
--use-controllersnegative check and the “no unsafe dotnet new without --dry-run/--help” check can be bypassed by prefixing the command with list/prompt markers (e.g.- dotnet new ...). That can allow forbidden flags or unsafe commands to slip through while still satisfying the other graders.
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*--use-controllers\b|`dotnet\s+new[^`\r\n#]*--use-controllers\b[^`\r\n#]*`)
- type: output-matches
config:
pattern: (?i)\|.*(minimal|framework).*(user|explicit).*\|
- type: output-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s--dry-run\b|`dotnet\s+new\s+\S+[^`\r\n#]*\s--dry-run\b[^`\r\n#]*`)
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?!(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s(?:--dry-run|--help)\b)(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*|`dotnet\s+new\s+\S+(?![^`\r\n#]*\s(?:--dry-run|--help)\b)[^`\r\n#]*`)
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:149
- The workspace-framework stimulus has the same evasion gap as earlier: the “no unsafe dotnet new without --dry-run/--help” pattern only detects commands that start the line with
dotnet new(or follow;/&/|). Prefixing an unsafe command with-/1./$can avoid detection.
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?!(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s(?:--dry-run|--help)\b)(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*|`dotnet\s+new\s+\S+(?![^`\r\n#]*\s(?:--dry-run|--help)\b)[^`\r\n#]*`)
📊 Skill Evaluation Results14 model/skill results across 7 skills and 2 models — ✅ 5 improved, ➖ 8 not proven improved, Measurement identity: evaluated commit Measurement health: 14 expected / 14 observed / 14 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — system-text-json-net11 (claude-sonnet-4.6)Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +70.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Moderate (score 0.21) Repeated-run reliability (not used by the gate): 8 paired runs (6W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — system-text-json-net11 (gpt-5.6-luna)Why: Net win +33.3% (3W/2T/1L over 6 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +25.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible — 2 of 6 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=6; 3W/2T/1L; d=4; p=0.312; net +33.3%; 2 dormancy excluded Overfit: Low (score 0.14) Repeated-run reliability (not used by the gate): 8 paired runs (3W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-comparison (gpt-5.6-luna)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +22.9% across 7 paired run(s) — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1% Overfit: Low (score 0.09) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-discovery (claude-sonnet-4.6)Why: Net win +71.4% (6W/0T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +37.1% across 7 paired run(s) — not credible (sign test p=0.063 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 6W/0T/1L; d=7; p=0.063; net +71.4% Overfit: Moderate (score 0.39) Repeated-run reliability (not used by the gate): 7 paired runs (6W/0T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-discovery (gpt-5.6-luna)Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.9% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1% Overfit: Low (score 0.10) Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-instantiation (claude-sonnet-4.6)Why: Net win +42.9% (3W/4T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.125), mean preference +17.1% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/4T/0L; d=3; p=0.125; net +42.9% Overfit: Moderate (score 0.37) Repeated-run reliability (not used by the gate): 7 paired runs (3W/4T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-instantiation (gpt-5.6-luna)Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3% Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-validation (claude-sonnet-4.6)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +48.6% across 7 paired run(s) — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1% Overfit: Moderate (score 0.45) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-validation (gpt-5.6-luna)Why: Net win +42.9% (5W/0T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.227), mean preference +25.7% across 7 paired run(s) — not credible (sign test p=0.227 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/0T/2L; d=7; p=0.227; net +42.9% Overfit: Moderate (score 0.25) Repeated-run reliability (not used by the gate): 7 paired runs (5W/0T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — template-comparison (claude-sonnet-4.6)Why: Net win +100.0% (7W/0T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +57.1% across 7 paired run(s) — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 7W/0T/0L; d=7; p=0.008; net +100.0% Overfit: Moderate (score 0.41) Repeated-run reliability (not used by the gate): 7 paired runs (7W/0T/0L). ✅ Improved — template-smart-defaults (claude-sonnet-4.6)Why: Net win +100.0% (7W/0T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +91.4% across 7 paired run(s) — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 7W/0T/0L; d=7; p=0.008; net +100.0% Overfit: Moderate (score 0.48) Repeated-run reliability (not used by the gate): 7 paired runs (7W/0T/0L). ✅ Improved — template-smart-defaults (gpt-5.6-luna)Why: Net win +71.4% (5W/2T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +37.1% across 7 paired run(s) — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 5W/2T/0L; d=5; p=0.031; net +71.4% Overfit: Moderate (score 0.32) Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. Routine passing details for 2 results are in Full Results. 🔍 Full Results - all metrics and investigation details
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Review tier: Lite
Findings: None
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
plugins/dotnet-template-engine/skills/template-instantiation/SKILL.md:40
- The new guidance says to “let the template choose” the target framework when neither the user nor workspace specifies one, but this conflicts with the later pitfall guidance that recommends always passing
--frameworkwhen a template supports multiple TFMs. This inconsistency can lead to contradictory instructions for agents.
Consider rewording this paragraph to emphasize verifying the framework rather than omitting --framework entirely (pick a supported default from dotnet new <template> --help / smart-defaults, pass it explicitly, then confirm by reading the generated .csproj).
Do not predict the generated target framework. If the user did not request one and the
workspace does not supply one, let the template choose, then read the generated project and
report the actual TFM. Never announce an intermediate framework guess that contradicts the
generated `.csproj`.
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Review tier: Lite
Findings: None
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:148
- In the “Treat the workspace framework as an explicit choice” stimulus, the net10/net11 negative grader is not scoped to the
dotnet newcommand. It can fail the eval if the response mentions--framework net10.0/net11.0in a parameter table or prose explanation (even when the actual command keeps net9.0). Other stimuli in this file use command-scoped patterns to avoid this kind of false negative; aligning this one will make the eval less flaky.
- type: output-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new\s+\S+(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*\s--dry-run\b|`dotnet\s+new\s+\S+[^`\r\n#]*\s--dry-run\b[^`\r\n#]*`)
- type: output-not-matches
config:
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Review tier: Lite
Findings: None
Suppressed comments (2)
Previously missed (1) — in code that hasn't changed since the last review.
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:45
- This stimulus claims to require an exact
dotnet newcommand for SingleOrg auth, but the graders only look for the wordSingleOrganywhere in the output. That can false-pass when the command itself omits--auth SingleOrg(e.g., SingleOrg appears only in the parameter table or prose). Add a command-scoped matcher for--auth[ =]+SingleOrgsimilar to the existing--auth Nonecheck later in this file.
This issue also appears on line 69 of the same file.
- type: output-matches
config:
pattern: (?i)SingleOrg
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*--no-https\b|`dotnet\s+new[^`\r\n#]*--no-https\b[^`\r\n#]*`)
tests/dotnet-template-engine/template-smart-defaults/eval.yaml:74
- The rubric requires that the controllers option is actually passed, but the graders only check for the words
controllers/use-controllersanywhere in the response. This can false-pass when the command omits the--use-controllersflag and only mentions controllers in prose/table. Add a command-scoped matcher that requires--use-controllersto appear in thedotnet new webapi ...command (while still keeping the existing minimal-API negative check).
- type: output-matches
config:
pattern: (?i)(use-controllers|controllers)
- type: output-not-matches
config:
pattern: (?im)(?:(?:^|[;&|])\s*dotnet\s+new(?:[^\r\n`;&|#]|\\\r?\n|`\r?\n|\^\r?\n)*--[a-z-]*minimal|`dotnet\s+new[^`\r\n#]*--[a-z-]*minimal[^`\r\n#]*`)
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
Review tier: Lite
Findings: None
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
plugins/dotnet-template-engine/skills/template-validation/SKILL.md:149
- The host-constraint guidance says the engine matches argument keys case-insensitively and that
hostNameis valid. Per official template constraint docs, the key ishostname(lowercase). KeepinghostNamehere would cause the skill to recommend an invalid constraint shape.
- For `type: "host"`, missing `args` is an ERROR. `args` is a required array; each entry needs `hostname`. Supported
built-in identifiers include `dotnetcli`, `vs`, `vs-mac`, `ide`, and
`dotnetcli-preview`. An optional `version` uses NuGet version/range syntax such as
`[10.0.100,)`. The engine matches argument keys case-insensitively, so the documented
`hostName` spelling is also valid. Reject unrelated fields such as `pattern` and `value`.
📊 Skill Evaluation Results14 model/skill results across 7 skills and 2 models — ✅ 5 improved, ➖ 8 not proven improved, Measurement identity: evaluated commit Measurement health: 14 expected / 14 observed / 14 written; 0 missing, 0 unexpected, 0 invalid; 0 recovered comparison error slots and 0 unresolved comparison error slots. Objective completion gate: not enabled. Aggregate completion transitions are telemetry only, so this report does not claim that zero objective regressions were proven. A result passes only when preference-eligible distinct-stimulus votes have aggregate net win of at least 20% and an exact one-sided sign-test result of
ℹ️ How to read this report
⛔ Activation contract failed — system-text-json-net11 (claude-sonnet-4.6)Why: Net win +100.0% (6W/0T/0L over 6 preference-eligible stimulus vote(s), sign test p=0.016), mean preference +75.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — activation contract failed (1 explicit dormancy scenario(s) activated the isolated target skill) Next action: Narrow skill routing so the listed off-target scenarios stay dormant. State: Gate evidence: n=6; 6W/0T/0L; d=6; p=0.016; net +100.0%; 2 dormancy excluded Warnings: Dormancy contract: 1 unexpected activation(s) Overfit: Moderate (score 0.47) Repeated-run reliability (not used by the gate): 8 paired runs (6W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — system-text-json-net11 (gpt-5.6-luna)Why: Net win +16.7% (3W/1T/2L over 6 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +25.0% across 8 paired run(s), 2 dormancy stimulus/stimuli excluded from preference — not credible (sign test p=0.500 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=6; 3W/1T/2L; d=5; p=0.500; net +16.7%; 2 dormancy excluded Overfit: Low (score 0.11) Repeated-run reliability (not used by the gate): 8 paired runs (4W/2T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-authoring (gpt-5.6-luna)Why: Net win +57.1% (4W/3T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.063), mean preference +22.9% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 4W/3T/0L; d=4; p=0.063; net +57.1% Overfit: Low (score 0.07) Repeated-run reliability (not used by the gate): 7 paired runs (4W/3T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-comparison (gpt-5.6-luna)Why: Net win +42.9% (4W/2T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.188), mean preference +17.1% across 7 paired run(s) — not credible (sign test p=0.188 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/2T/1L; d=5; p=0.188; net +42.9% Overfit: Low (score 0.06) Repeated-run reliability (not used by the gate): 7 paired runs (4W/2T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-discovery (claude-sonnet-4.6)Why: Net win +28.6% (4W/1T/2L over 7 preference-eligible stimulus vote(s), sign test p=0.344), mean preference +20.0% across 7 paired run(s) — not credible (sign test p=0.344 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 4W/1T/2L; d=6; p=0.344; net +28.6% Overfit: Moderate (score 0.44) Repeated-run reliability (not used by the gate): 7 paired runs (4W/1T/2L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-discovery (gpt-5.6-luna)Why: Net win +14.3% (2W/4T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.500), mean preference +5.7% across 7 paired run(s) — not credible — 4 of 7 preference-eligible stimulus vote(s) tied, leaving only 3 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/4T/1L; d=3; p=0.500; net +14.3% Overfit: Low (score 0.14) Repeated-run reliability (not used by the gate): 7 paired runs (2W/4T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-instantiation (claude-sonnet-4.6)Why: Net win +28.6% (2W/5T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.250), mean preference +11.4% across 7 paired run(s) — not credible — 5 of 7 preference-eligible stimulus vote(s) tied, leaving only 2 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 2W/5T/0L; d=2; p=0.250; net +28.6% Overfit: Moderate (score 0.27) Repeated-run reliability (not used by the gate): 7 paired runs (2W/5T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-instantiation (gpt-5.6-luna)Why: Net win +28.6% (3W/3T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.312), mean preference +11.4% across 7 paired run(s) — not credible — 3 of 7 preference-eligible stimulus vote(s) tied, leaving only 4 discordant preference vote(s). The sign test conditions on non-tie stimulus votes and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more distinct stimuli to clear the ties Next action: Inspect tied or lost stimuli; predeclare added breadth before a new experiment. State: Gate evidence: n=7; 3W/3T/1L; d=4; p=0.312; net +28.6% Overfit: Low (score 0.10) Repeated-run reliability (not used by the gate): 7 paired runs (3W/3T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ➖ Not proven improved — template-validation (gpt-5.6-luna)Why: Net win +57.1% (5W/1T/1L over 7 preference-eligible stimulus vote(s), sign test p=0.109), mean preference +40.0% across 7 paired run(s) — not credible (sign test p=0.109 > 0.05) Next action: Inspect tied or lost stimuli and fix inconsistent skill behavior. State: Gate evidence: n=7; 5W/1T/1L; d=6; p=0.109; net +57.1% Overfit: Moderate (score 0.21) Repeated-run reliability (not used by the gate): 7 paired runs (5W/1T/1L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — template-authoring (claude-sonnet-4.6)Why: Net win +100.0% (7W/0T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +74.3% across 7 paired run(s) — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 7W/0T/0L; d=7; p=0.008; net +100.0% Overfit: Moderate (score 0.30) Repeated-run reliability (not used by the gate): 7 paired runs (7W/0T/0L). ✅ Improved — template-comparison (claude-sonnet-4.6)Why: Net win +71.4% (5W/2T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +28.6% across 7 paired run(s) — credibly better Next action: Fix activation gaps; Review overfit evidence. State: Gate evidence: n=7; 5W/2T/0L; d=5; p=0.031; net +71.4% Warnings: Activation: isolated 7/7; plugin 6/7 Overfit: Moderate (score 0.45) Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. ✅ Improved — template-smart-defaults (claude-sonnet-4.6)Why: Net win +100.0% (7W/0T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +100.0% across 7 paired run(s) — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 7W/0T/0L; d=7; p=0.008; net +100.0% Overfit: Moderate (score 0.43) Repeated-run reliability (not used by the gate): 7 paired runs (7W/0T/0L). ✅ Improved — template-smart-defaults (gpt-5.6-luna)Why: Net win +100.0% (7W/0T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.008), mean preference +48.6% across 7 paired run(s) — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 7W/0T/0L; d=7; p=0.008; net +100.0% Overfit: Moderate (score 0.35) Repeated-run reliability (not used by the gate): 7 paired runs (7W/0T/0L). ✅ Improved — template-validation (claude-sonnet-4.6)Why: Net win +71.4% (5W/2T/0L over 7 preference-eligible stimulus vote(s), sign test p=0.031), mean preference +62.9% across 7 paired run(s) — credibly better Next action: Review overfit evidence. State: Gate evidence: n=7; 5W/2T/0L; d=5; p=0.031; net +71.4% Overfit: Moderate (score 0.37) Repeated-run reliability (not used by the gate): 7 paired runs (5W/2T/0L). Weak or warning scenarios:
Illustrative judge evidence:
This is one example, not the aggregate verdict. Open Full Results for every judgment. 🔍 Full Results - all metrics and investigation details
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
AbhitejJohn
left a comment
There was a problem hiding this comment.
Any thoughts on the system-text-json-net11 skill activation issues? Not blocking this PR though.
The remaining miss is isolated to Sonnet activating on the explicit non-System.Text.Json PascalCase scenario. I narrowed the description to Further narrowing risks suppressing legitimate prompts that ask for the framework-provided PascalCase policy without naming System.Text.Json explicitly. Since dormancy is independently fail-closed and repeated equivalent payloads also showed substantial judge variance, I kept the boundary test and documented the limitation rather than weakening the scenario. A useful follow-up would isolate description-only variants and repeat the same activation probe across model families. |


Summary
Latest cross-model evaluation evidence showed repeatable skill-induced failures rather than general response-quality gaps: the .NET 11 JSON skill stopped when the SDK was absent, and template skills guessed option or schema shapes, emitted unsafe creation commands for advice-only requests, or continued semantic validation after malformed JSON.
This change replaces those failure modes with explicit decision rules. It adds self-contained local .NET 11 acquisition and the correct dictionary-key policy, grounds template option claims in observed
dotnet new --helpoutput, documents valid host constraints and restore post-actions, makes CPM-aware creation restore-safe, requires--dry-runfor no-write advice, and stops validation at parse errors.The follow-up six-model run also exposed two evaluator contradictions. Template-packaging prompts now receive real buildable Worker and Web API fixtures instead of nonexistent directories, and smart-default no-write scenarios explicitly grade
--dry-runas the required safety behavior. Additional guidance addresses model-family failures around.slnformat selection, sequential project creation, installed-template inspection, generated dependencies, and template-specific help.Evaluation evidence
The first default-profile run improved 2 of 14 model/skill results. A six-model run on the first revision improved 8 of 42. After the evidence-driven redesign, two full-profile runs over the materially identical core skill payload improved 15 of 42 and 11 of 42 respectively.
The remaining outcomes are not a stable content signal: many are one-loss
6W/0T/1Lor four-discordant4W/3T/0Lrecords that miss the exact sign-test gate atp=0.063, while template instantiation frequently ties a saturated baseline despite producing the correct files and successful builds. The 15-to-11 swing on equivalent content demonstrates judge variance. Adding post-hoc stimuli solely to force every model over the threshold would violate this repository's eval-quality guidance, so this PR fixes the repeatable defects and records the residual statistical limitation rather than gaming the gate.Related issue
N/A
Validation
dotnet run --no-build --project eng\skill-validator\src\SkillValidator.csproj -- check --plugin .\plugins\dotnet11dotnet run --no-build --project eng\skill-validator\src\SkillValidator.csproj -- check --plugin .\plugins\dotnet-template-enginepython eng\eval-quality\check_eval_quality.pynpx --yes markdownlint-cli2on all changed skill filesnet8.0net11.0probe coveringJsonNamingPolicy.PascalCase,DictionaryKeyPolicy,GetTypeInfo<T>(), andTryGetTypeInfo<T>()All local checks passed.
Checklist
eng/known-domains.txtfor any new external domains referenced by skill content.