Harden eval verdict correctness and failure accounting - #1027
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
…ectness-infrastructure
Add Vally 0.13 all-error compatibility, a practical net-win floor, objective completion criteria, and CI fault-injection coverage. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
Assert combined stdout and stderr because GitHub warning annotations are emitted on stdout in Actions. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
📊 Skill Evaluation Results70 skill(s) evaluated — ✅ 6 improved, ❌ 22 no credible change, 🔻 0 objective regressions, 📉 0 preference losses (report only).
A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
|
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Determine required Android SDK packages for specific .NET version | +0.0% | +0.0% | 0/0/0 |
| ▲ Diagnose non-Microsoft JDK causing build failure | +100.0% | +100.0% | 1/0/0 |
| ▲ Fix stale MAUI workloads after SDK update | +100.0% | +100.0% | 1/0/0 |
| = Guardrail against workload update and repair | +0.0% | +0.0% | 0/1/0 |
| ▲ Plan Linux MAUI environment for Android | +100.0% | +40.0% | 1/0/0 |
| ▲ Plan complete MAUI setup on Windows | +100.0% | +40.0% | 1/0/0 |
| ▲ Plan macOS MAUI setup with Xcode | +100.0% | +100.0% | 1/0/0 |
| ▲ Prevent incorrect JAVA_HOME configuration | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +25.0% (2W/1T/1L over 4 trial(s), sign test p=0.500), mean preference +10.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (2W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Avoid legacy Xamarin.Forms lifecycle methods | -100.0% | -40.0% | 0/0/1 |
| = Platform-specific lifecycle mapping | +0.0% | +0.0% | 0/1/0 |
| ▲ Save and restore state on background | +100.0% | +40.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +70.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid legacy Xamarin.Forms lifecycle methods | +100.0% | +40.0% | 1/0/0 |
| ▲ Platform-specific lifecycle mapping | +100.0% | +40.0% | 1/0/0 |
| ▲ Save and restore state on background | +100.0% | +100.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +55.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid legacy Xamarin.Forms lifecycle methods | +100.0% | +40.0% | 1/0/0 |
| ▲ Platform-specific lifecycle mapping | +100.0% | +40.0% | 1/0/0 |
| ▲ Save and restore state on background | +100.0% | +100.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +40.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid legacy Xamarin.Forms lifecycle methods | +100.0% | +40.0% | 1/0/0 |
| ▲ Platform-specific lifecycle mapping | +100.0% | +40.0% | 1/0/0 |
| ▲ Save and restore state on background | +100.0% | +40.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +40.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid legacy Xamarin.Forms lifecycle methods | +100.0% | +40.0% | 1/0/0 |
| ▲ Platform-specific lifecycle mapping | +100.0% | +40.0% | 1/0/0 |
| ▲ Save and restore state on background | +100.0% | +40.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +40.0% | 1/0/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (1W/4T/0L over 5 trial(s), sign test p=0.500), mean preference +8.0% — not credible — 4 of 5 trial(s) tied, leaving only 1 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (1W/4T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid ListView and ViewCell mistakes | +100.0% | +40.0% | 1/0/0 |
| = Basic CollectionView with data binding and DataTemplate | +0.0% | +0.0% | 0/1/0 |
| = Grid layout with CollectionView | +0.0% | +0.0% | 0/1/0 |
| = ItemSizingStrategy placement for uniform items | +0.0% | +0.0% | 0/1/0 |
| = Selection and pull-to-refresh with CollectionView | +0.0% | +0.0% | 0/1/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (2W/2T/1L over 5 trial(s), sign test p=0.500), mean preference +8.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (2W/2T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid ListView and ViewCell mistakes | +0.0% | +0.0% | 0/1/0 |
| ▼ Basic CollectionView with data binding and DataTemplate | -100.0% | -40.0% | 0/0/1 |
| = Grid layout with CollectionView | +0.0% | +0.0% | 0/1/0 |
| ▲ ItemSizingStrategy placement for uniform items | +100.0% | +40.0% | 1/0/0 |
| ▲ Selection and pull-to-refresh with CollectionView | +100.0% | +40.0% | 1/0/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (3W/2T/0L over 5 trial(s), sign test p=0.125), mean preference +24.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/2T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid ListView and ViewCell mistakes | +0.0% | +0.0% | 0/1/0 |
| = Basic CollectionView with data binding and DataTemplate | +0.0% | +0.0% | 0/1/0 |
| ▲ Grid layout with CollectionView | +100.0% | +40.0% | 1/0/0 |
| ▲ ItemSizingStrategy placement for uniform items | +100.0% | +40.0% | 1/0/0 |
| ▲ Selection and pull-to-refresh with CollectionView | +100.0% | +40.0% | 1/0/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +32.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid ListView and ViewCell mistakes | +0.0% | +0.0% | 0/1/0 |
| ▲ Basic CollectionView with data binding and DataTemplate | +100.0% | +40.0% | 1/0/0 |
| ▲ Grid layout with CollectionView | +100.0% | +40.0% | 1/0/0 |
| ▲ ItemSizingStrategy placement for uniform items | +100.0% | +40.0% | 1/0/0 |
| ▲ Selection and pull-to-refresh with CollectionView | +100.0% | +40.0% | 1/0/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (2W/2T/1L over 5 trial(s), sign test p=0.500), mean preference +8.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (2W/2T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid ListView and ViewCell mistakes | +100.0% | +40.0% | 1/0/0 |
| ▼ Basic CollectionView with data binding and DataTemplate | -100.0% | -40.0% | 0/0/1 |
| = Grid layout with CollectionView | +0.0% | +0.0% | 0/1/0 |
| ▲ ItemSizingStrategy placement for uniform items | +100.0% | +40.0% | 1/0/0 |
| = Selection and pull-to-refresh with CollectionView | +0.0% | +0.0% | 0/1/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +25.0% (2W/1T/1L over 4 trial(s), sign test p=0.500), mean preference +10.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (2W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Create and use an IValueConverter | -100.0% | -40.0% | 0/0/1 |
| ▲ Diagnose missing BindingContext and non-compiled bindings | +100.0% | +40.0% | 1/0/0 |
| = Implement MVVM ViewModel with ObservableObject | +0.0% | +0.0% | 0/1/0 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +75.0% (3W/1T/0L over 4 trial(s), sign test p=0.125), mean preference +30.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Create and use an IValueConverter | +0.0% | +0.0% | 0/1/0 |
| ▲ Diagnose missing BindingContext and non-compiled bindings | +100.0% | +40.0% | 1/0/0 |
| ▲ Implement MVVM ViewModel with ObservableObject | +100.0% | +40.0% | 1/0/0 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +0.0% (2W/0T/2L over 4 trial(s), sign test p=0.687), mean preference +0.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (2W/0T/2L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Create and use an IValueConverter | -100.0% | -40.0% | 0/0/1 |
| ▲ Diagnose missing BindingContext and non-compiled bindings | +100.0% | +40.0% | 1/0/0 |
| ▼ Implement MVVM ViewModel with ObservableObject | -100.0% | -40.0% | 0/0/1 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +50.0% (3W/0T/1L over 4 trial(s), sign test p=0.312), mean preference +20.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create and use an IValueConverter | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose missing BindingContext and non-compiled bindings | +100.0% | +40.0% | 1/0/0 |
| ▼ Implement MVVM ViewModel with ObservableObject | -100.0% | -40.0% | 0/0/1 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +25.0% (1W/3T/0L over 4 trial(s), sign test p=0.500), mean preference +10.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (1W/3T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Create and use an IValueConverter | +0.0% | +0.0% | 0/1/0 |
| = Diagnose missing BindingContext and non-compiled bindings | +0.0% | +0.0% | 0/1/0 |
| = Implement MVVM ViewModel with ObservableObject | +0.0% | +0.0% | 0/1/0 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
❌ maui-dependency-injection — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +56.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid AddScoped pitfall in MAUI | +0.0% | +0.0% | 0/1/0 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +100.0% | 1/0/0 |
| ▲ Platform-specific service registration with fallback | +100.0% | +100.0% | 1/0/0 |
| ▲ Register services with correct lifetimes in MauiProgram.cs | +100.0% | +40.0% | 1/0/0 |
| ▲ Shell navigation auto-resolves DI-registered pages | +100.0% | +40.0% | 1/0/0 |
❌ maui-dependency-injection — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +36.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid AddScoped pitfall in MAUI | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +100.0% | 1/0/0 |
| ▲ Platform-specific service registration with fallback | +100.0% | +40.0% | 1/0/0 |
| ▼ Register services with correct lifetimes in MauiProgram.cs | -100.0% | -40.0% | 0/0/1 |
| ▲ Shell navigation auto-resolves DI-registered pages | +100.0% | +40.0% | 1/0/0 |
❌ maui-dependency-injection — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +36.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Avoid AddScoped pitfall in MAUI | -100.0% | -40.0% | 0/0/1 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +100.0% | 1/0/0 |
| ▲ Platform-specific service registration with fallback | +100.0% | +40.0% | 1/0/0 |
| ▲ Register services with correct lifetimes in MauiProgram.cs | +100.0% | +40.0% | 1/0/0 |
| ▲ Shell navigation auto-resolves DI-registered pages | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-dependency-injection — details
State: INVALID_INCONCLUSIVE (unmatched_trajectories)
Reason: Net win +50.0% (3W/0T/1L over 4 trial(s), sign test p=0.312), mean preference +20.0%, 1 unmatched — inconclusive (unmatched trajectories)
Effective scenarios (report only): 4 (3W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid AddScoped pitfall in MAUI | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +40.0% | 1/0/0 |
| = Platform-specific service registration with fallback | +0.0% | +0.0% | 0/0/0 |
| ▼ Register services with correct lifetimes in MauiProgram.cs | -100.0% | -40.0% | 0/0/1 |
| ▲ Shell navigation auto-resolves DI-registered pages | +100.0% | +40.0% | 1/0/0 |
❌ maui-dependency-injection — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (2W/2T/1L over 5 trial(s), sign test p=0.500), mean preference +20.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (2W/2T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid AddScoped pitfall in MAUI | +0.0% | +0.0% | 0/1/0 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +100.0% | 1/0/0 |
| = Platform-specific service registration with fallback | +0.0% | +0.0% | 0/1/0 |
| ▲ Register services with correct lifetimes in MauiProgram.cs | +100.0% | +40.0% | 1/0/0 |
| ▼ Shell navigation auto-resolves DI-registered pages | -100.0% | -40.0% | 0/0/1 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +85.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +40.0% | 1/0/0 |
| ▲ Edge-to-edge layout with SafeAreaEdges | +100.0% | +100.0% | 1/0/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +100.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +100.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +100.0% | 1/0/0 |
| ▲ Edge-to-edge layout with SafeAreaEdges | +100.0% | +100.0% | 1/0/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +100.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +85.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +40.0% | 1/0/0 |
| ▲ Edge-to-edge layout with SafeAreaEdges | +100.0% | +100.0% | 1/0/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +100.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +75.0% (3W/1T/0L over 4 trial(s), sign test p=0.125), mean preference +45.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +40.0% | 1/0/0 |
| = Edge-to-edge layout with SafeAreaEdges | +0.0% | +0.0% | 0/1/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +40.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +70.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +100.0% | 1/0/0 |
| ▲ Edge-to-edge layout with SafeAreaEdges | +100.0% | +40.0% | 1/0/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +100.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +40.0% | 1/0/0 |
❌ maui-shell-navigation — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +44.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose common Shell navigation mistakes | +0.0% | +0.0% | 0/1/0 |
| ▲ Handle back navigation and unsaved changes guard | +100.0% | +100.0% | 1/0/0 |
| ▲ Navigate with parameters using GoToAsync | +100.0% | +40.0% | 1/0/0 |
| ▲ Set up Shell navigation with tabs and flyout | +100.0% | +40.0% | 1/0/0 |
| ▲ Stable routes for deep linking into tabs | +100.0% | +40.0% | 1/0/0 |
❌ maui-shell-navigation — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (3W/2T/0L over 5 trial(s), sign test p=0.125), mean preference +36.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/2T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose common Shell navigation mistakes | +100.0% | +40.0% | 1/0/0 |
| ▲ Handle back navigation and unsaved changes guard | +100.0% | +100.0% | 1/0/0 |
| ▲ Navigate with parameters using GoToAsync | +100.0% | +40.0% | 1/0/0 |
| = Set up Shell navigation with tabs and flyout | +0.0% | +0.0% | 0/1/0 |
| = Stable routes for deep linking into tabs | +0.0% | +0.0% | 0/1/0 |
❌ maui-shell-navigation — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +32.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose common Shell navigation mistakes | +0.0% | +0.0% | 0/1/0 |
| ▲ Handle back navigation and unsaved changes guard | +100.0% | +40.0% | 1/0/0 |
| ▲ Navigate with parameters using GoToAsync | +100.0% | +40.0% | 1/0/0 |
| ▲ Set up Shell navigation with tabs and flyout | +100.0% | +40.0% | 1/0/0 |
| ▲ Stable routes for deep linking into tabs | +100.0% | +40.0% | 1/0/0 |
❌ maui-shell-navigation — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +40.0% (3W/1T/1L over 5 trial(s), sign test p=0.312), mean preference +52.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose common Shell navigation mistakes | +100.0% | +100.0% | 1/0/0 |
| ▼ Handle back navigation and unsaved changes guard | -100.0% | -40.0% | 0/0/1 |
| ▲ Navigate with parameters using GoToAsync | +100.0% | +100.0% | 1/0/0 |
| = Set up Shell navigation with tabs and flyout | +0.0% | +0.0% | 0/1/0 |
| ▲ Stable routes for deep linking into tabs | +100.0% | +100.0% | 1/0/0 |
❌ maui-theming — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +24.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Add light/dark mode support using AppThemeBinding | +100.0% | +40.0% | 1/0/0 |
| ▲ Avoid common theming mistakes | +100.0% | +40.0% | 1/0/0 |
| ▼ Create custom themes with ResourceDictionary switching | -100.0% | -40.0% | 0/0/1 |
| ▲ Detect and respond to system theme changes | +100.0% | +40.0% | 1/0/0 |
| ▲ Swap theme dictionaries without destroying app styles | +100.0% | +40.0% | 1/0/0 |
❌ maui-theming — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +36.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Add light/dark mode support using AppThemeBinding | +100.0% | +40.0% | 1/0/0 |
| ▼ Avoid common theming mistakes | -100.0% | -40.0% | 0/0/1 |
| ▲ Create custom themes with ResourceDictionary switching | +100.0% | +100.0% | 1/0/0 |
| ▲ Detect and respond to system theme changes | +100.0% | +40.0% | 1/0/0 |
| ▲ Swap theme dictionaries without destroying app styles | +100.0% | +40.0% | 1/0/0 |
❌ maui-theming — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +40.0% (3W/1T/1L over 5 trial(s), sign test p=0.312), mean preference +28.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Add light/dark mode support using AppThemeBinding | +100.0% | +40.0% | 1/0/0 |
| ▲ Avoid common theming mistakes | +100.0% | +100.0% | 1/0/0 |
| ▲ Create custom themes with ResourceDictionary switching | +100.0% | +40.0% | 1/0/0 |
| = Detect and respond to system theme changes | +0.0% | +0.0% | 0/1/0 |
| ▼ Swap theme dictionaries without destroying app styles | -100.0% | -40.0% | 0/0/1 |
❌ maui-theming — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +40.0% (3W/1T/1L over 5 trial(s), sign test p=0.312), mean preference +16.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Add light/dark mode support using AppThemeBinding | +100.0% | +40.0% | 1/0/0 |
| = Avoid common theming mistakes | +0.0% | +0.0% | 0/1/0 |
| ▲ Create custom themes with ResourceDictionary switching | +100.0% | +40.0% | 1/0/0 |
| ▼ Detect and respond to system theme changes | -100.0% | -40.0% | 0/0/1 |
| ▲ Swap theme dictionaries without destroying app styles | +100.0% | +40.0% | 1/0/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +40.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| ▲ Validate a template.json file | +100.0% | +40.0% | 1/0/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +70.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| ▲ Validate a template.json file | +100.0% | +100.0% | 1/0/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +50.0% (1W/1T/0L over 2 trial(s), sign test p=0.500), mean preference +20.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (1W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| = Validate a template.json file | +0.0% | +0.0% | 0/1/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +40.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| ▲ Validate a template.json file | +100.0% | +40.0% | 1/0/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +40.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| ▲ Validate a template.json file | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +33.3% (2W/0T/1L over 3 trial(s), sign test p=0.500), mean preference +13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (2W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Choose between blazor and blazorwasm | +100.0% | +40.0% | 1/0/0 |
| ▼ Compare webapi vs webapp side by side | -100.0% | -40.0% | 0/0/1 |
| ▲ Decide which template fits a background processing scenario | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +66.7% (2W/1T/0L over 3 trial(s), sign test p=0.250), mean preference +26.7% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (2W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Choose between blazor and blazorwasm | +0.0% | +0.0% | 0/1/0 |
| ▲ Compare webapi vs webapp side by side | +100.0% | +40.0% | 1/0/0 |
| ▲ Decide which template fits a background processing scenario | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +66.7% (2W/1T/0L over 3 trial(s), sign test p=0.250), mean preference +46.7% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (2W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Choose between blazor and blazorwasm | +0.0% | +0.0% | 0/1/0 |
| ▲ Compare webapi vs webapp side by side | +100.0% | +100.0% | 1/0/0 |
| ▲ Decide which template fits a background processing scenario | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +40.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (3W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Choose between blazor and blazorwasm | +100.0% | +40.0% | 1/0/0 |
| ▲ Compare webapi vs webapp side by side | +100.0% | +40.0% | 1/0/0 |
| ▲ Decide which template fits a background processing scenario | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +33.3% (1W/2T/0L over 3 trial(s), sign test p=0.500), mean preference +13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (1W/2T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Choose between blazor and blazorwasm | +0.0% | +0.0% | 0/1/0 |
| ▲ Compare webapi vs webapp side by side | +100.0% | +40.0% | 1/0/0 |
| = Decide which template fits a background processing scenario | +0.0% | +0.0% | 0/1/0 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (3W/0T/2L over 5 trial(s), sign test p=0.500), mean preference +20.0% — not credible (sign test p=0.500 > 0.05)
Effective scenarios (report only): 5 (3W/0T/2L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +100.0% | 1/0/0 |
| ▼ Inspect template parameters and compare choices | -100.0% | -40.0% | 0/0/1 |
| ▲ Preview project creation with dry run | +100.0% | +40.0% | 1/0/0 |
| ▼ Resolve ambiguous project intent to multiple candidates | -100.0% | -40.0% | 0/0/1 |
| ▲ Search NuGet for specialized template | +100.0% | +40.0% | 1/0/0 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +40.0% (3W/1T/1L over 5 trial(s), sign test p=0.312), mean preference +40.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +100.0% | 1/0/0 |
| ▲ Inspect template parameters and compare choices | +100.0% | +40.0% | 1/0/0 |
| = Preview project creation with dry run | +0.0% | +0.0% | 0/1/0 |
| ▲ Resolve ambiguous project intent to multiple candidates | +100.0% | +100.0% | 1/0/0 |
| ▼ Search NuGet for specialized template | -100.0% | -40.0% | 0/0/1 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +44.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +40.0% | 1/0/0 |
| ▲ Inspect template parameters and compare choices | +100.0% | +40.0% | 1/0/0 |
| ▲ Preview project creation with dry run | +100.0% | +100.0% | 1/0/0 |
| ▲ Resolve ambiguous project intent to multiple candidates | +100.0% | +40.0% | 1/0/0 |
| = Search NuGet for specialized template | +0.0% | +0.0% | 0/1/0 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +32.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +40.0% | 1/0/0 |
| = Inspect template parameters and compare choices | +0.0% | +0.0% | 0/1/0 |
| ▲ Preview project creation with dry run | +100.0% | +40.0% | 1/0/0 |
| ▲ Resolve ambiguous project intent to multiple candidates | +100.0% | +40.0% | 1/0/0 |
| ▲ Search NuGet for specialized template | +100.0% | +40.0% | 1/0/0 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +12.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +40.0% | 1/0/0 |
| ▲ Inspect template parameters and compare choices | +100.0% | +40.0% | 1/0/0 |
| ▲ Preview project creation with dry run | +100.0% | +40.0% | 1/0/0 |
| ▲ Resolve ambiguous project intent to multiple candidates | +100.0% | +40.0% | 1/0/0 |
| ▼ Search NuGet for specialized template | -100.0% | -100.0% | 0/0/1 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win -100.0% (0W/0T/1L over 1 trial(s), sign test p=0.500), mean preference -40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (0W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Create a console application | -100.0% | -40.0% | 0/0/1 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (1W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create a console application | +100.0% | +40.0% | 1/0/0 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +100.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (1W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create a console application | +100.0% | +100.0% | 1/0/0 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (1W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create a console application | +100.0% | +40.0% | 1/0/0 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (0W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Create a console application | +0.0% | +0.0% | 0/1/0 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win -50.0% (1W/0T/3L over 4 trial(s), sign test p=0.312), mean preference -20.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (1W/0T/3L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ AOT implies a compatible framework | -100.0% | -40.0% | 0/0/1 |
| ▼ Auth implies HTTPS stays enabled | -100.0% | -40.0% | 0/0/1 |
| ▼ Controllers exclude the minimal-API flag | -100.0% | -40.0% | 0/0/1 |
| ▲ Never override an explicit user value | +100.0% | +40.0% | 1/0/0 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +75.0% (3W/1T/0L over 4 trial(s), sign test p=0.125), mean preference +45.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ AOT implies a compatible framework | +100.0% | +40.0% | 1/0/0 |
| ▲ Auth implies HTTPS stays enabled | +100.0% | +40.0% | 1/0/0 |
| = Controllers exclude the minimal-API flag | +0.0% | +0.0% | 0/1/0 |
| ▲ Never override an explicit user value | +100.0% | +100.0% | 1/0/0 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win -50.0% (0W/2T/2L over 4 trial(s), sign test p=0.250), mean preference -20.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (0W/2T/2L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = AOT implies a compatible framework | +0.0% | +0.0% | 0/1/0 |
| ▼ Auth implies HTTPS stays enabled | -100.0% | -40.0% | 0/0/1 |
| = Controllers exclude the minimal-API flag | +0.0% | +0.0% | 0/1/0 |
| ▼ Never override an explicit user value | -100.0% | -40.0% | 0/0/1 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +50.0% (3W/0T/1L over 4 trial(s), sign test p=0.312), mean preference +20.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ AOT implies a compatible framework | +100.0% | +40.0% | 1/0/0 |
| ▲ Auth implies HTTPS stays enabled | +100.0% | +40.0% | 1/0/0 |
| ▼ Controllers exclude the minimal-API flag | -100.0% | -40.0% | 0/0/1 |
| ▲ Never override an explicit user value | +100.0% | +40.0% | 1/0/0 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +0.0% (1W/2T/1L over 4 trial(s), sign test p=0.750), mean preference +0.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (1W/2T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ AOT implies a compatible framework | +100.0% | +40.0% | 1/0/0 |
| ▼ Auth implies HTTPS stays enabled | -100.0% | -40.0% | 0/0/1 |
| = Controllers exclude the minimal-API flag | +0.0% | +0.0% | 0/1/0 |
| = Never override an explicit user value | +0.0% | +0.0% | 0/1/0 |
⚠️ template-validation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +70.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Validate correct template and suggest improvements | +100.0% | +100.0% | 1/0/0 |
| ▲ Validate template with multiple errors | +100.0% | +40.0% | 1/0/0 |
⚠️ template-validation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +70.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Validate correct template and suggest improvements | +100.0% | +100.0% | 1/0/0 |
| ▲ Validate template with multiple errors | +100.0% | +40.0% | 1/0/0 |
⚠️ template-validation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win -50.0% (0W/1T/1L over 2 trial(s), sign test p=0.500), mean preference -20.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (0W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Validate correct template and suggest improvements | -100.0% | -40.0% | 0/0/1 |
| = Validate template with multiple errors | +0.0% | +0.0% | 0/1/0 |
⚠️ template-validation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +100.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Validate correct template and suggest improvements | +100.0% | +100.0% | 1/0/0 |
| ▲ Validate template with multiple errors | +100.0% | +100.0% | 1/0/0 |
Per-scenario details for 7 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown.
🔍 Full Results - additional metrics and failure investigation steps
To investigate failures, paste this to your AI coding agent:
For PR 1027 in dotnet/skills, download eval artifacts with
gh run download 32085934523 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/24d0a0697fe5c868928d452ca11856a9268a24e2/eng/vally-adapter/InvestigatingResults.md and follow it to analyze the results.json files. Diagnose each failure, suggest fixes to the eval.yaml and skill content, and tell me what to fix first.
▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions
📊 Skill Evaluation Results70 skill(s) evaluated — ✅ 6 improved, ❌ 22 no credible change, 🔻 0 objective regressions, 📉 0 preference losses (report only).
A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
|
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Determine required Android SDK packages for specific .NET version | +0.0% | +0.0% | 0/0/0 |
| ▲ Diagnose non-Microsoft JDK causing build failure | +100.0% | +100.0% | 1/0/0 |
| ▲ Fix stale MAUI workloads after SDK update | +100.0% | +100.0% | 1/0/0 |
| = Guardrail against workload update and repair | +0.0% | +0.0% | 0/1/0 |
| ▲ Plan Linux MAUI environment for Android | +100.0% | +40.0% | 1/0/0 |
| ▲ Plan complete MAUI setup on Windows | +100.0% | +40.0% | 1/0/0 |
| ▲ Plan macOS MAUI setup with Xcode | +100.0% | +100.0% | 1/0/0 |
| ▲ Prevent incorrect JAVA_HOME configuration | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +25.0% (2W/1T/1L over 4 trial(s), sign test p=0.500), mean preference +10.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (2W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Avoid legacy Xamarin.Forms lifecycle methods | -100.0% | -40.0% | 0/0/1 |
| = Platform-specific lifecycle mapping | +0.0% | +0.0% | 0/1/0 |
| ▲ Save and restore state on background | +100.0% | +40.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +70.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid legacy Xamarin.Forms lifecycle methods | +100.0% | +40.0% | 1/0/0 |
| ▲ Platform-specific lifecycle mapping | +100.0% | +40.0% | 1/0/0 |
| ▲ Save and restore state on background | +100.0% | +100.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +55.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid legacy Xamarin.Forms lifecycle methods | +100.0% | +40.0% | 1/0/0 |
| ▲ Platform-specific lifecycle mapping | +100.0% | +40.0% | 1/0/0 |
| ▲ Save and restore state on background | +100.0% | +100.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +40.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid legacy Xamarin.Forms lifecycle methods | +100.0% | +40.0% | 1/0/0 |
| ▲ Platform-specific lifecycle mapping | +100.0% | +40.0% | 1/0/0 |
| ▲ Save and restore state on background | +100.0% | +40.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-app-lifecycle — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +40.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid legacy Xamarin.Forms lifecycle methods | +100.0% | +40.0% | 1/0/0 |
| ▲ Platform-specific lifecycle mapping | +100.0% | +40.0% | 1/0/0 |
| ▲ Save and restore state on background | +100.0% | +40.0% | 1/0/0 |
| ▲ Window lifecycle event subscription | +100.0% | +40.0% | 1/0/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (1W/4T/0L over 5 trial(s), sign test p=0.500), mean preference +8.0% — not credible — 4 of 5 trial(s) tied, leaving only 1 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (1W/4T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid ListView and ViewCell mistakes | +100.0% | +40.0% | 1/0/0 |
| = Basic CollectionView with data binding and DataTemplate | +0.0% | +0.0% | 0/1/0 |
| = Grid layout with CollectionView | +0.0% | +0.0% | 0/1/0 |
| = ItemSizingStrategy placement for uniform items | +0.0% | +0.0% | 0/1/0 |
| = Selection and pull-to-refresh with CollectionView | +0.0% | +0.0% | 0/1/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (2W/2T/1L over 5 trial(s), sign test p=0.500), mean preference +8.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (2W/2T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid ListView and ViewCell mistakes | +0.0% | +0.0% | 0/1/0 |
| ▼ Basic CollectionView with data binding and DataTemplate | -100.0% | -40.0% | 0/0/1 |
| = Grid layout with CollectionView | +0.0% | +0.0% | 0/1/0 |
| ▲ ItemSizingStrategy placement for uniform items | +100.0% | +40.0% | 1/0/0 |
| ▲ Selection and pull-to-refresh with CollectionView | +100.0% | +40.0% | 1/0/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (3W/2T/0L over 5 trial(s), sign test p=0.125), mean preference +24.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/2T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid ListView and ViewCell mistakes | +0.0% | +0.0% | 0/1/0 |
| = Basic CollectionView with data binding and DataTemplate | +0.0% | +0.0% | 0/1/0 |
| ▲ Grid layout with CollectionView | +100.0% | +40.0% | 1/0/0 |
| ▲ ItemSizingStrategy placement for uniform items | +100.0% | +40.0% | 1/0/0 |
| ▲ Selection and pull-to-refresh with CollectionView | +100.0% | +40.0% | 1/0/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +32.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid ListView and ViewCell mistakes | +0.0% | +0.0% | 0/1/0 |
| ▲ Basic CollectionView with data binding and DataTemplate | +100.0% | +40.0% | 1/0/0 |
| ▲ Grid layout with CollectionView | +100.0% | +40.0% | 1/0/0 |
| ▲ ItemSizingStrategy placement for uniform items | +100.0% | +40.0% | 1/0/0 |
| ▲ Selection and pull-to-refresh with CollectionView | +100.0% | +40.0% | 1/0/0 |
❌ maui-collectionview — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (2W/2T/1L over 5 trial(s), sign test p=0.500), mean preference +8.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (2W/2T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid ListView and ViewCell mistakes | +100.0% | +40.0% | 1/0/0 |
| ▼ Basic CollectionView with data binding and DataTemplate | -100.0% | -40.0% | 0/0/1 |
| = Grid layout with CollectionView | +0.0% | +0.0% | 0/1/0 |
| ▲ ItemSizingStrategy placement for uniform items | +100.0% | +40.0% | 1/0/0 |
| = Selection and pull-to-refresh with CollectionView | +0.0% | +0.0% | 0/1/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +25.0% (2W/1T/1L over 4 trial(s), sign test p=0.500), mean preference +10.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (2W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Create and use an IValueConverter | -100.0% | -40.0% | 0/0/1 |
| ▲ Diagnose missing BindingContext and non-compiled bindings | +100.0% | +40.0% | 1/0/0 |
| = Implement MVVM ViewModel with ObservableObject | +0.0% | +0.0% | 0/1/0 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +75.0% (3W/1T/0L over 4 trial(s), sign test p=0.125), mean preference +30.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Create and use an IValueConverter | +0.0% | +0.0% | 0/1/0 |
| ▲ Diagnose missing BindingContext and non-compiled bindings | +100.0% | +40.0% | 1/0/0 |
| ▲ Implement MVVM ViewModel with ObservableObject | +100.0% | +40.0% | 1/0/0 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +0.0% (2W/0T/2L over 4 trial(s), sign test p=0.687), mean preference +0.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (2W/0T/2L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Create and use an IValueConverter | -100.0% | -40.0% | 0/0/1 |
| ▲ Diagnose missing BindingContext and non-compiled bindings | +100.0% | +40.0% | 1/0/0 |
| ▼ Implement MVVM ViewModel with ObservableObject | -100.0% | -40.0% | 0/0/1 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +50.0% (3W/0T/1L over 4 trial(s), sign test p=0.312), mean preference +20.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create and use an IValueConverter | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose missing BindingContext and non-compiled bindings | +100.0% | +40.0% | 1/0/0 |
| ▼ Implement MVVM ViewModel with ObservableObject | -100.0% | -40.0% | 0/0/1 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-data-binding — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +25.0% (1W/3T/0L over 4 trial(s), sign test p=0.500), mean preference +10.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (1W/3T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Create and use an IValueConverter | +0.0% | +0.0% | 0/1/0 |
| = Diagnose missing BindingContext and non-compiled bindings | +0.0% | +0.0% | 0/1/0 |
| = Implement MVVM ViewModel with ObservableObject | +0.0% | +0.0% | 0/1/0 |
| ▲ Set up compiled bindings with x:DataType on a page | +100.0% | +40.0% | 1/0/0 |
❌ maui-dependency-injection — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +56.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid AddScoped pitfall in MAUI | +0.0% | +0.0% | 0/1/0 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +100.0% | 1/0/0 |
| ▲ Platform-specific service registration with fallback | +100.0% | +100.0% | 1/0/0 |
| ▲ Register services with correct lifetimes in MauiProgram.cs | +100.0% | +40.0% | 1/0/0 |
| ▲ Shell navigation auto-resolves DI-registered pages | +100.0% | +40.0% | 1/0/0 |
❌ maui-dependency-injection — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +36.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid AddScoped pitfall in MAUI | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +100.0% | 1/0/0 |
| ▲ Platform-specific service registration with fallback | +100.0% | +40.0% | 1/0/0 |
| ▼ Register services with correct lifetimes in MauiProgram.cs | -100.0% | -40.0% | 0/0/1 |
| ▲ Shell navigation auto-resolves DI-registered pages | +100.0% | +40.0% | 1/0/0 |
❌ maui-dependency-injection — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +36.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Avoid AddScoped pitfall in MAUI | -100.0% | -40.0% | 0/0/1 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +100.0% | 1/0/0 |
| ▲ Platform-specific service registration with fallback | +100.0% | +40.0% | 1/0/0 |
| ▲ Register services with correct lifetimes in MauiProgram.cs | +100.0% | +40.0% | 1/0/0 |
| ▲ Shell navigation auto-resolves DI-registered pages | +100.0% | +40.0% | 1/0/0 |
⚠️ maui-dependency-injection — details
State: INVALID_INCONCLUSIVE (unmatched_trajectories)
Reason: Net win +50.0% (3W/0T/1L over 4 trial(s), sign test p=0.312), mean preference +20.0%, 1 unmatched — inconclusive (unmatched trajectories)
Effective scenarios (report only): 4 (3W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid AddScoped pitfall in MAUI | +100.0% | +40.0% | 1/0/0 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +40.0% | 1/0/0 |
| = Platform-specific service registration with fallback | +0.0% | +0.0% | 0/0/0 |
| ▼ Register services with correct lifetimes in MauiProgram.cs | -100.0% | -40.0% | 0/0/1 |
| ▲ Shell navigation auto-resolves DI-registered pages | +100.0% | +40.0% | 1/0/0 |
❌ maui-dependency-injection — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (2W/2T/1L over 5 trial(s), sign test p=0.500), mean preference +20.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (2W/2T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Avoid AddScoped pitfall in MAUI | +0.0% | +0.0% | 0/1/0 |
| ▲ Diagnose a page whose injected dependencies are missing | +100.0% | +100.0% | 1/0/0 |
| = Platform-specific service registration with fallback | +0.0% | +0.0% | 0/1/0 |
| ▲ Register services with correct lifetimes in MauiProgram.cs | +100.0% | +40.0% | 1/0/0 |
| ▼ Shell navigation auto-resolves DI-registered pages | -100.0% | -40.0% | 0/0/1 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +85.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +40.0% | 1/0/0 |
| ▲ Edge-to-edge layout with SafeAreaEdges | +100.0% | +100.0% | 1/0/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +100.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +100.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +100.0% | 1/0/0 |
| ▲ Edge-to-edge layout with SafeAreaEdges | +100.0% | +100.0% | 1/0/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +100.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +85.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +40.0% | 1/0/0 |
| ▲ Edge-to-edge layout with SafeAreaEdges | +100.0% | +100.0% | 1/0/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +100.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +75.0% (3W/1T/0L over 4 trial(s), sign test p=0.125), mean preference +45.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +40.0% | 1/0/0 |
| = Edge-to-edge layout with SafeAreaEdges | +0.0% | +0.0% | 0/1/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +40.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +100.0% | 1/0/0 |
⚠️ maui-safe-area — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (4W/0T/0L over 4 trial(s), sign test p=0.063), mean preference +70.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (4W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Avoid deprecated UseSafeArea in new .NET 10 project | +100.0% | +100.0% | 1/0/0 |
| ▲ Edge-to-edge layout with SafeAreaEdges | +100.0% | +40.0% | 1/0/0 |
| ▲ Handle notch and status bar safe areas on iOS | +100.0% | +100.0% | 1/0/0 |
| ▲ Keyboard avoidance with safe area for chat UI | +100.0% | +40.0% | 1/0/0 |
❌ maui-shell-navigation — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +44.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose common Shell navigation mistakes | +0.0% | +0.0% | 0/1/0 |
| ▲ Handle back navigation and unsaved changes guard | +100.0% | +100.0% | 1/0/0 |
| ▲ Navigate with parameters using GoToAsync | +100.0% | +40.0% | 1/0/0 |
| ▲ Set up Shell navigation with tabs and flyout | +100.0% | +40.0% | 1/0/0 |
| ▲ Stable routes for deep linking into tabs | +100.0% | +40.0% | 1/0/0 |
❌ maui-shell-navigation — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (3W/2T/0L over 5 trial(s), sign test p=0.125), mean preference +36.0% — not credible — 2 of 5 trial(s) tied, leaving only 3 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/2T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose common Shell navigation mistakes | +100.0% | +40.0% | 1/0/0 |
| ▲ Handle back navigation and unsaved changes guard | +100.0% | +100.0% | 1/0/0 |
| ▲ Navigate with parameters using GoToAsync | +100.0% | +40.0% | 1/0/0 |
| = Set up Shell navigation with tabs and flyout | +0.0% | +0.0% | 0/1/0 |
| = Stable routes for deep linking into tabs | +0.0% | +0.0% | 0/1/0 |
❌ maui-shell-navigation — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +32.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Diagnose common Shell navigation mistakes | +0.0% | +0.0% | 0/1/0 |
| ▲ Handle back navigation and unsaved changes guard | +100.0% | +40.0% | 1/0/0 |
| ▲ Navigate with parameters using GoToAsync | +100.0% | +40.0% | 1/0/0 |
| ▲ Set up Shell navigation with tabs and flyout | +100.0% | +40.0% | 1/0/0 |
| ▲ Stable routes for deep linking into tabs | +100.0% | +40.0% | 1/0/0 |
❌ maui-shell-navigation — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +40.0% (3W/1T/1L over 5 trial(s), sign test p=0.312), mean preference +52.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Diagnose common Shell navigation mistakes | +100.0% | +100.0% | 1/0/0 |
| ▼ Handle back navigation and unsaved changes guard | -100.0% | -40.0% | 0/0/1 |
| ▲ Navigate with parameters using GoToAsync | +100.0% | +100.0% | 1/0/0 |
| = Set up Shell navigation with tabs and flyout | +0.0% | +0.0% | 0/1/0 |
| ▲ Stable routes for deep linking into tabs | +100.0% | +100.0% | 1/0/0 |
❌ maui-theming — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +24.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Add light/dark mode support using AppThemeBinding | +100.0% | +40.0% | 1/0/0 |
| ▲ Avoid common theming mistakes | +100.0% | +40.0% | 1/0/0 |
| ▼ Create custom themes with ResourceDictionary switching | -100.0% | -40.0% | 0/0/1 |
| ▲ Detect and respond to system theme changes | +100.0% | +40.0% | 1/0/0 |
| ▲ Swap theme dictionaries without destroying app styles | +100.0% | +40.0% | 1/0/0 |
❌ maui-theming — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +36.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Add light/dark mode support using AppThemeBinding | +100.0% | +40.0% | 1/0/0 |
| ▼ Avoid common theming mistakes | -100.0% | -40.0% | 0/0/1 |
| ▲ Create custom themes with ResourceDictionary switching | +100.0% | +100.0% | 1/0/0 |
| ▲ Detect and respond to system theme changes | +100.0% | +40.0% | 1/0/0 |
| ▲ Swap theme dictionaries without destroying app styles | +100.0% | +40.0% | 1/0/0 |
❌ maui-theming — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +40.0% (3W/1T/1L over 5 trial(s), sign test p=0.312), mean preference +28.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Add light/dark mode support using AppThemeBinding | +100.0% | +40.0% | 1/0/0 |
| ▲ Avoid common theming mistakes | +100.0% | +100.0% | 1/0/0 |
| ▲ Create custom themes with ResourceDictionary switching | +100.0% | +40.0% | 1/0/0 |
| = Detect and respond to system theme changes | +0.0% | +0.0% | 0/1/0 |
| ▼ Swap theme dictionaries without destroying app styles | -100.0% | -40.0% | 0/0/1 |
❌ maui-theming — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +40.0% (3W/1T/1L over 5 trial(s), sign test p=0.312), mean preference +16.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Add light/dark mode support using AppThemeBinding | +100.0% | +40.0% | 1/0/0 |
| = Avoid common theming mistakes | +0.0% | +0.0% | 0/1/0 |
| ▲ Create custom themes with ResourceDictionary switching | +100.0% | +40.0% | 1/0/0 |
| ▼ Detect and respond to system theme changes | -100.0% | -40.0% | 0/0/1 |
| ▲ Swap theme dictionaries without destroying app styles | +100.0% | +40.0% | 1/0/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +40.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| ▲ Validate a template.json file | +100.0% | +40.0% | 1/0/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +70.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| ▲ Validate a template.json file | +100.0% | +100.0% | 1/0/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +50.0% (1W/1T/0L over 2 trial(s), sign test p=0.500), mean preference +20.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (1W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| = Validate a template.json file | +0.0% | +0.0% | 0/1/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +40.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| ▲ Validate a template.json file | +100.0% | +40.0% | 1/0/0 |
⚠️ template-authoring — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +40.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create template from existing project | +100.0% | +40.0% | 1/0/0 |
| ▲ Validate a template.json file | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +33.3% (2W/0T/1L over 3 trial(s), sign test p=0.500), mean preference +13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (2W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Choose between blazor and blazorwasm | +100.0% | +40.0% | 1/0/0 |
| ▼ Compare webapi vs webapp side by side | -100.0% | -40.0% | 0/0/1 |
| ▲ Decide which template fits a background processing scenario | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +66.7% (2W/1T/0L over 3 trial(s), sign test p=0.250), mean preference +26.7% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (2W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Choose between blazor and blazorwasm | +0.0% | +0.0% | 0/1/0 |
| ▲ Compare webapi vs webapp side by side | +100.0% | +40.0% | 1/0/0 |
| ▲ Decide which template fits a background processing scenario | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +66.7% (2W/1T/0L over 3 trial(s), sign test p=0.250), mean preference +46.7% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (2W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Choose between blazor and blazorwasm | +0.0% | +0.0% | 0/1/0 |
| ▲ Compare webapi vs webapp side by side | +100.0% | +100.0% | 1/0/0 |
| ▲ Decide which template fits a background processing scenario | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (3W/0T/0L over 3 trial(s), sign test p=0.125), mean preference +40.0% — underpowered (3 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (3W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Choose between blazor and blazorwasm | +100.0% | +40.0% | 1/0/0 |
| ▲ Compare webapi vs webapp side by side | +100.0% | +40.0% | 1/0/0 |
| ▲ Decide which template fits a background processing scenario | +100.0% | +40.0% | 1/0/0 |
⚠️ template-comparison — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +33.3% (1W/2T/0L over 3 trial(s), sign test p=0.500), mean preference +13.3% — underpowered (3 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 3 (1W/2T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Choose between blazor and blazorwasm | +0.0% | +0.0% | 0/1/0 |
| ▲ Compare webapi vs webapp side by side | +100.0% | +40.0% | 1/0/0 |
| = Decide which template fits a background processing scenario | +0.0% | +0.0% | 0/1/0 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +20.0% (3W/0T/2L over 5 trial(s), sign test p=0.500), mean preference +20.0% — not credible (sign test p=0.500 > 0.05)
Effective scenarios (report only): 5 (3W/0T/2L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +100.0% | 1/0/0 |
| ▼ Inspect template parameters and compare choices | -100.0% | -40.0% | 0/0/1 |
| ▲ Preview project creation with dry run | +100.0% | +40.0% | 1/0/0 |
| ▼ Resolve ambiguous project intent to multiple candidates | -100.0% | -40.0% | 0/0/1 |
| ▲ Search NuGet for specialized template | +100.0% | +40.0% | 1/0/0 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +40.0% (3W/1T/1L over 5 trial(s), sign test p=0.312), mean preference +40.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (3W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +100.0% | 1/0/0 |
| ▲ Inspect template parameters and compare choices | +100.0% | +40.0% | 1/0/0 |
| = Preview project creation with dry run | +0.0% | +0.0% | 0/1/0 |
| ▲ Resolve ambiguous project intent to multiple candidates | +100.0% | +100.0% | 1/0/0 |
| ▼ Search NuGet for specialized template | -100.0% | -40.0% | 0/0/1 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +44.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +40.0% | 1/0/0 |
| ▲ Inspect template parameters and compare choices | +100.0% | +40.0% | 1/0/0 |
| ▲ Preview project creation with dry run | +100.0% | +100.0% | 1/0/0 |
| ▲ Resolve ambiguous project intent to multiple candidates | +100.0% | +40.0% | 1/0/0 |
| = Search NuGet for specialized template | +0.0% | +0.0% | 0/1/0 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +80.0% (4W/1T/0L over 5 trial(s), sign test p=0.063), mean preference +32.0% — not credible — 1 of 5 trial(s) tied, leaving only 4 discordant trial(s). The sign test conditions on non-tie trials and cannot reach 0.05 below 5, so no record could have passed here — this is not a measured null. Either the skill is inert on these scenarios (make them discriminate) or the eval needs more trials to clear the ties (more scenarios or defaults.runs)
Effective scenarios (report only): 5 (4W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +40.0% | 1/0/0 |
| = Inspect template parameters and compare choices | +0.0% | +0.0% | 0/1/0 |
| ▲ Preview project creation with dry run | +100.0% | +40.0% | 1/0/0 |
| ▲ Resolve ambiguous project intent to multiple candidates | +100.0% | +40.0% | 1/0/0 |
| ▲ Search NuGet for specialized template | +100.0% | +40.0% | 1/0/0 |
❌ template-discovery — details
State: VALID_NO_CHANGE (no_credible_preference_change)
Reason: Net win +60.0% (4W/0T/1L over 5 trial(s), sign test p=0.188), mean preference +12.0% — not credible (sign test p=0.188 > 0.05)
Effective scenarios (report only): 5 (4W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Find template for web API project | +100.0% | +40.0% | 1/0/0 |
| ▲ Inspect template parameters and compare choices | +100.0% | +40.0% | 1/0/0 |
| ▲ Preview project creation with dry run | +100.0% | +40.0% | 1/0/0 |
| ▲ Resolve ambiguous project intent to multiple candidates | +100.0% | +40.0% | 1/0/0 |
| ▼ Search NuGet for specialized template | -100.0% | -100.0% | 0/0/1 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win -100.0% (0W/0T/1L over 1 trial(s), sign test p=0.500), mean preference -40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (0W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Create a console application | -100.0% | -40.0% | 0/0/1 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (1W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create a console application | +100.0% | +40.0% | 1/0/0 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +100.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (1W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create a console application | +100.0% | +100.0% | 1/0/0 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (1W/0T/0L over 1 trial(s), sign test p=0.500), mean preference +40.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (1W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Create a console application | +100.0% | +40.0% | 1/0/0 |
⚠️ template-instantiation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +0.0% (0W/1T/0L over 1 trial(s), sign test p=1.000), mean preference +0.0% — underpowered (1 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 1 (0W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = Create a console application | +0.0% | +0.0% | 0/1/0 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win -50.0% (1W/0T/3L over 4 trial(s), sign test p=0.312), mean preference -20.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (1W/0T/3L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ AOT implies a compatible framework | -100.0% | -40.0% | 0/0/1 |
| ▼ Auth implies HTTPS stays enabled | -100.0% | -40.0% | 0/0/1 |
| ▼ Controllers exclude the minimal-API flag | -100.0% | -40.0% | 0/0/1 |
| ▲ Never override an explicit user value | +100.0% | +40.0% | 1/0/0 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +75.0% (3W/1T/0L over 4 trial(s), sign test p=0.125), mean preference +45.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/1T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ AOT implies a compatible framework | +100.0% | +40.0% | 1/0/0 |
| ▲ Auth implies HTTPS stays enabled | +100.0% | +40.0% | 1/0/0 |
| = Controllers exclude the minimal-API flag | +0.0% | +0.0% | 0/1/0 |
| ▲ Never override an explicit user value | +100.0% | +100.0% | 1/0/0 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win -50.0% (0W/2T/2L over 4 trial(s), sign test p=0.250), mean preference -20.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (0W/2T/2L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| = AOT implies a compatible framework | +0.0% | +0.0% | 0/1/0 |
| ▼ Auth implies HTTPS stays enabled | -100.0% | -40.0% | 0/0/1 |
| = Controllers exclude the minimal-API flag | +0.0% | +0.0% | 0/1/0 |
| ▼ Never override an explicit user value | -100.0% | -40.0% | 0/0/1 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +50.0% (3W/0T/1L over 4 trial(s), sign test p=0.312), mean preference +20.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (3W/0T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ AOT implies a compatible framework | +100.0% | +40.0% | 1/0/0 |
| ▲ Auth implies HTTPS stays enabled | +100.0% | +40.0% | 1/0/0 |
| ▼ Controllers exclude the minimal-API flag | -100.0% | -40.0% | 0/0/1 |
| ▲ Never override an explicit user value | +100.0% | +40.0% | 1/0/0 |
⚠️ template-smart-defaults — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +0.0% (1W/2T/1L over 4 trial(s), sign test p=0.750), mean preference +0.0% — underpowered (4 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 4 (1W/2T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ AOT implies a compatible framework | +100.0% | +40.0% | 1/0/0 |
| ▼ Auth implies HTTPS stays enabled | -100.0% | -40.0% | 0/0/1 |
| = Controllers exclude the minimal-API flag | +0.0% | +0.0% | 0/1/0 |
| = Never override an explicit user value | +0.0% | +0.0% | 0/1/0 |
⚠️ template-validation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +70.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Validate correct template and suggest improvements | +100.0% | +100.0% | 1/0/0 |
| ▲ Validate template with multiple errors | +100.0% | +40.0% | 1/0/0 |
⚠️ template-validation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +70.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Validate correct template and suggest improvements | +100.0% | +100.0% | 1/0/0 |
| ▲ Validate template with multiple errors | +100.0% | +40.0% | 1/0/0 |
⚠️ template-validation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win -50.0% (0W/1T/1L over 2 trial(s), sign test p=0.500), mean preference -20.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (0W/1T/1L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▼ Validate correct template and suggest improvements | -100.0% | -40.0% | 0/0/1 |
| = Validate template with multiple errors | +0.0% | +0.0% | 0/1/0 |
⚠️ template-validation — details
State: INVALID_INCONCLUSIVE (underpowered)
Reason: Net win +100.0% (2W/0T/0L over 2 trial(s), sign test p=0.250), mean preference +100.0% — underpowered (2 counted trial(s); a credible verdict needs at least 5, and this eval won every one of them) — raise the eval's trial count with more scenarios or defaults.runs
Effective scenarios (report only): 2 (2W/0T/0L).
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Validate correct template and suggest improvements | +100.0% | +100.0% | 1/0/0 |
| ▲ Validate template with multiple errors | +100.0% | +100.0% | 1/0/0 |
Per-scenario details for 7 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown.
🔍 Full Results - additional metrics and failure investigation steps
To investigate failures, paste this to your AI coding agent:
For PR 1027 in dotnet/skills, download eval artifacts with
gh run download 32085934523 --repo dotnet/skills --pattern "vally-results-*" --dir ./eval-results, then fetch https://raw.githubusercontent.com/dotnet/skills/24d0a0697fe5c868928d452ca11856a9268a24e2/eng/vally-adapter/InvestigatingResults.md and follow it to analyze the results.json files. Diagnose each failure, suggest fixes to the eval.yaml and skill content, and tell me what to fix first.
▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
📊 Skill Evaluation Results5 skill(s) evaluated — ✅ 2 improved, ❌ 3 no credible change, 🔻 0 regressed. A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
❌ coverage-analysis — detailsReason: Net win +38.9% (11W/3T/4L over 18 trial(s), sign test p=0.059), mean preference +38.9% — not credible (sign test p=0.059 > 0.05)
❌ coverage-analysis — detailsReason: Net win +27.8% (10W/3T/5L over 18 trial(s), sign test p=0.151), mean preference +11.1% — not credible (sign test p=0.151 > 0.05)
❌ generate-testability-wrappers — detailsReason: Net win +40.0% (9W/3T/3L over 15 trial(s), sign test p=0.073), mean preference +40.0% — not credible (sign test p=0.073 > 0.05)
Per-scenario details for 2 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown. 🔍 Full Results - additional metrics and failure investigation steps
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
Might this happen to be too strict for some niche skills? Should there be option to override it? |
JanKrivanek
left a comment
There was a problem hiding this comment.
Overall - looks good to go
📊 Skill Evaluation Results7 skill(s) evaluated — ✅ 4 improved, ❌ 3 no credible change, 🔻 0 regressed. A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
❌ coverage-analysis — detailsReason: Net win +38.9% (11W/3T/4L over 18 trial(s), sign test p=0.059), mean preference +38.9% — not credible (sign test p=0.059 > 0.05)
❌ coverage-analysis — detailsReason: Net win +27.8% (10W/3T/5L over 18 trial(s), sign test p=0.151), mean preference +11.1% — not credible (sign test p=0.151 > 0.05)
❌ generate-testability-wrappers — detailsReason: Net win +40.0% (9W/3T/3L over 15 trial(s), sign test p=0.073), mean preference +40.0% — not credible (sign test p=0.073 > 0.05)
Per-scenario details for 4 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown. 🔍 Full Results - additional metrics and failure investigation steps
▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
|
@JanKrivanek The 20% floor is repository policy rather than a Vally default, and I did not add a per-eval override. It does not make small or niche evals harder to pass: for every instrument from 5 through 25 stimuli, any W/T/L record that can pass the exact one-sided sign test already has net win >=20%. The existing exhaustive boundary test proves this. The first record newly rejected by the floor is |
There was a problem hiding this comment.
Review details
Suppressed comments (2)
Previously missed (1) — in code that hasn't changed since the last review.
.github/workflows/evaluation-run.yml:168
- The pinned SHA for
actions/upload-artifactis annotated asv4.6.2here, but the same SHA is used elsewhere in this workflow with av7comment. This inconsistency can confuse audits and future upgrades; align the version comment with the actual action version you intend to pin.
This issue also appears on line 433 of the same file.
- name: Upload trusted skill-validator archive
uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v4.6.2
.github/workflows/evaluation-run.yml:436
- The pinned SHA for
actions/download-artifactis annotated asv4.3.0here, but the same SHA is referenced elsewhere in the repo asv8.x. Keeping the inline version comment accurate avoids confusion during dependency reviews.
- name: Download trusted skill-validator archive
if: steps.find-evals.outputs.has_evals == 'true'
uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v4.3.0
with:
- Files reviewed: 28/28 changed files
- Comments generated: 0 new
- Review effort level: Lite
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
Resolve the PAT probe overlap by retaining capability-based effort fallback, which covers the incoming Haiku fix without a model allowlist. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
There was a problem hiding this comment.
Review details
Suppressed comments (1)
Previously missed (1) — in code that hasn't changed since the last review.
eng/vally-adapter/consolidate.mjs:498
- In the main markdown table,
SkillandModelcells are only passed throughtd()(pipe/newline escaping) and are not HTML-escaped. Because these values originate from results produced against evaluated content, a crafted name containing</&could inject HTML and break or spoof the PR comment rendering. EscapeskillNameandmodelwithhtml()before table rendering (while keepingtd()for table-safety).
const common = [
verdict.skillName,
verdict.model,
resultLabel(verdict),
gateEvidence(verdict),
];
- Files reviewed: 28/28 changed files
- Comments generated: 0 new
- Review effort level: Lite
📊 Skill Evaluation Results5 skill(s) evaluated — ✅ 4 improved, ❌ 0 no credible change, 🔻 0 regressed. A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
|
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Ask for a bounded list instead of grading the workspace | +100.0% | +100.0% | 3/0/0 |
| ▲ Grade C# tests against available production code | +100.0% | +100.0% | 3/0/0 |
| ▲ Grade Go table-driven tests without misreading the loop as branching | +100.0% | +100.0% | 3/0/0 |
| ▲ Grade pytest test methods using the same rubric | +100.0% | +60.0% | 3/0/0 |
| ▲ Grade tests when the production code under test is unavailable | +100.0% | +100.0% | 3/0/0 |
| ▲ Keep a 62-test grading report readable as a PR comment | +100.0% | +70.0% | 2/0/0 |
Per-scenario details for 4 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown.
🔍 Full Results - additional metrics and failure investigation steps
▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
📊 Skill Evaluation Results5 skill(s) evaluated — ✅ 4 improved, ❌ 0 no credible change, 🔻 0 regressed. A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
|
| Scenario | Net win | Δ Pref | Trials (W/T/L) |
|---|---|---|---|
| ▲ Decline wrapper generation for already-abstracted code | +100.0% | +100.0% | 2/0/0 |
| ▲ Generate TimeProvider adoption for DateTime.UtcNow | +33.3% | +13.3% | 2/0/1 |
| ▲ Generate custom Environment wrapper | +33.3% | +13.3% | 1/2/0 |
| ▲ Make time controllable in a library that has no DI container | +100.0% | +40.0% | 3/0/0 |
| ▲ Recommend System.IO.Abstractions for file system calls | +100.0% | +100.0% | 3/0/0 |
Per-scenario details for 4 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown.
🔍 Full Results - additional metrics and failure investigation steps
▶ Sessions Visualisation -- interactive replay of all evaluation sessions
📊 Session Analytics (preview) -- aggregated metrics across evaluation sessions
📊 Skill Evaluation Results6 skill(s) evaluated — ✅ 6 improved, ❌ 0 no credible change, 🔻 0 regressed. A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
Per-scenario details for 6 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown. 🔍 Full Results - additional metrics and failure investigation steps ▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
📊 Skill Evaluation Results7 skill(s) evaluated — ✅ 7 improved, ❌ 0 no credible change, 🔻 0 regressed. A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
Per-scenario details for 7 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown. 🔍 Full Results - additional metrics and failure investigation steps ▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
📊 Skill Evaluation Results8 skill(s) evaluated — ✅ 8 improved, ❌ 0 no credible change, 🔻 0 regressed. A skill passes only on a credible net win over baseline: more wins than losses, by an exact one-sided sign test at
ℹ️ Column legend
Per-scenario details for 8 skill(s) were omitted to keep this comment under GitHub's 65,536-character limit — open the job's step summary or Full Results for the complete breakdown. 🔍 Full Results - additional metrics and failure investigation steps ▶ Sessions Visualisation -- interactive replay of all evaluation sessions |
Why this is needed
The old gate could mistake three different things for skill quality:
This PR makes the inference unit, failure accounting, and verdict contract explicit.
Example for #986: repeated runs manufactured significance
Assume an eval has 4 distinct stimuli and
runs: 3. The treatment wins all 12 paired runs.12W/0L, exact one-sided sign-testp = 1/4096, so the eval could pass.4W/0L,p = 1/16 = 0.0625. It cannot pass atalpha = 0.05.The 12 runs show repeatability on four tasks. They do not create 12 independent task samples. This PR therefore makes four stimulus votes authoritative and keeps the 12 paired runs as reliability evidence.
Example for #909: one judge timeout is not a regression
Assume five comparison slots. Four judge calls succeed, but one returns
Timeout after 120000ms waiting for session.idle.This PR:
(stimulusName, trialIndex).INVALID_INCONCLUSIVEwith a classified cause if recovery fails.A disabled organization, rate limit, service failure, missing result, malformed report, or empty run now fails closed. None can become
VALID_PASS,VALID_NO_CHANGE, orVALID_REGRESSION.Example for #970: significance alone is not enough
Assume 100 distinct stimuli produce
5W/95T/0L.p = 1/32 = 0.03125.(5 - 0) / 100 = 5%.Statistical significance says the direction is unlikely under a 50/50 null among non-ties. It does not say the effect is large enough to matter. This PR also requires an absolute task-level net win of at least 20%, so this sparse result does not pass.
The completion hard gate is not enabled. Vally's aggregate
passedvalue can combine weighted deterministic and LLM graders. Treating that mixed value as objective completion could manufacture a P0 regression.VALID_REGRESSIONstays reserved until the harness can prove deterministic grader provenance.Decision flow
flowchart LR A[Distinct stimulus] --> B[Repeated Vally runs] B --> C[Pair by stimulusName and trialIndex] C --> D{Judge slot failed?} D -- Yes --> E[Retry once; freeze successes] D -- No --> F[Run-level W/T/L] E --> F F --> G[One majority-direction vote per stimulus] F --> H[Reliability evidence only] G --> I[Exact sign test at alpha 0.05] G --> J[Absolute net-win floor at 20%] I --> K[Verdict state] J --> K H --> L[Report pass rate, retries, and flakiness]What is implemented
alpha = 0.05, based on distinct stimulus votes.abs((wins - losses) / stimulusVotes) >= 0.20.VALID_REGRESSIONstays reserved until the harness supplies trusted spec-to-result grader identity.Vally guidance versus repository policy
Vally documents that:
pass@k,pass^k, and flakiness;static/complex-staticgraders from non-deterministicllmgraders.Vally does not prescribe a distinct-stimulus minimum, a sign-test alpha, or a practical net-win floor. Those are repository policies:
alpha = 0.05;Five is not a general power target. Under a no-tie model, exact calculations show that the sign test alone reaches 80% power at about 158 discordant votes for a 60% conditional win rate, 37 for 70%, 18 for 80%, and 8 for 90%. The 20% practical floor is intentionally binding at a 60% conditional win rate: at 158 votes the combined gate passes about 52% of records and approaches 50% as the sample grows. The gate is designed to certify effects above its practical threshold. Eval authors must choose breadth from the smallest effect they need to detect; they must not treat five as "well powered."
Official references:
vally compareObjective completion contract
A future
VALID_REGRESSIONhard gate must have all of these properties:staticandcomplex-staticgrader types is eligible.pass,fail, orunknown; missing, duplicate, ambiguous, or LLM-backed evidence becomesunknown.Current Vally comparison output exposes aggregate pass transitions, but not a trusted unique mapping from each eval-spec grader declaration to each result. The PR reports those transitions but does not overstate them as objective proof.
Output and compatibility
scenarioEvidencefrom non-gatingcomparisonTrialEvidence.stimulusVoteCountandminCredibleStimuliare the canonical fields.trialCountandminCredibleTrialsremain compatibility aliases for current consumers.main.(stimulusName, trialIndex)keys, with no missing index and no duplicate slot.PR result UI and trusted validator handoff
The old PR comment could show the same skill twice without naming the model, label a positive but unproven result with a red cross, say "zero objective regressions" while that gate was disabled, and call all non-passes failures. It also hid exact result accounting and judge retry health.
The new comment:
n, W/T/L, discordant votes, one-resultp, and aggregate net win;For example,
7W/0T/2Lover nine stimuli has a strong+55.6%net win butp=0.090. It now appears as ➖ Not proven improved, with the two losing stimuli and the next action, rather than as a generic red failure.The validator transport also no longer depends on a writable cache. The trusted
prepare-validatorjob builds or restores one archive before PR content is checked out, uploads it as a same-run artifact, and every matrix job downloads that exact archive. The cache remains an optimization only. This removes the issue-comment path where cache save failed and every matrix job then tried to rebuild against the restricted evaluation package sources.Validation
actionlint1.7.7.session.idletimeout;780c999546a4007bde9742f511c29adb65c67cd8.8 expected / 8 observed / 8 written, with zero missing, unexpected, or invalid evals.trialIndexwas present and non-negative, each(stimulusName, trialIndex)key was unique, and paired key sets had zero differences.scenarioEvidence.gateEligible=true,comparisonTrialEvidence.gateEligible=false, zero unmatched trials, and zero errors.0434d9e58437e6beda3f27814cf9bed996eb57c9.8 expected / 8 observed / 8 writtenaccounting, zero missing, unexpected, invalid, underpowered, or unresolved-error results.trialIndexwas present and non-negative, every arm key was unique, and paired key sets had zero differences.Related issues