Skip to content

Complete the eval gate: completion regression, rerun, and multiple-comparison policies #970

Description

@AbhitejJohn

Current scope

PR #1027 completed the practical-significance work from this issue:

  • Exact one-sided sign test
  • 20% task-level net-win floor
  • Failed-slot judge retry and frozen successful judgments
  • Exact fail-closed result accounting

The remaining work is intentionally narrower:

  • Add a trusted objective-completion regression hard gate when deterministic grader identity is available.
  • Define the rerun policy so repeated runs cannot weaken the nominal significance guarantee.
  • Define the multiple-comparison policy for suite-wide evaluation.

The completion gate remains deferred because the current infrastructure cannot identify a trusted deterministic completion grader. Scenario-level aggregation and any threshold changes must be back-tested before they become merge gates.

Related implementation: #1027.

(Copilot, updating this issue on Abhitej's behalf.)

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions