Skip to content

Reject misaligned observations in Model QA discrimination reports - #1038

Merged
msitarzewski merged 1 commit into
msitarzewski:mainfrom
rudycelekli:campaign/agency-oct5-model-alignment
Oct 6, 2026
Merged

msitarzewski merged 1 commit into
msitarzewski:mainfrom
rudycelekli:campaign/agency-oct5-model-alignment

Conversation

@rudycelekli

Copy link
Copy Markdown
Contributor

Agent Information

Agent file: specialized/specialized-model-qa.md

Motivation

AUC consumes the arrays positionally while pandas class masks align by index: reordered scores produce two inconsistent metrics in one report.

Testing

  • Reproduced the original example failure; the corrected example passes the same focused fixture.
  • Changed-agent lint and git diff --check pass.
  • Full scripts/test-convert-outputs.sh --drift=advisory passes all 32 checks over 282 agents and 15 tool formats. Expected content drift is advisory; the generated-output manifest is left to maintainers.

Native extracted example with controlled local inputs; no production system or model-effectiveness claim. No paid API call, real invoice, production data or actual model-effectiveness evaluation was performed.

Candidate listed before publication in #917: #917 (comment). This single-profile correction adds no tooling or generated outputs. AI assistance was used to investigate, implement and test it; the commit is signed and includes DCO sign-off.

Signed-off-by: Rudy Celekli <47457359+rudycelekli@users.noreply.github.com>
@msitarzewski
msitarzewski merged commit ac58638 into msitarzewski:main Oct 6, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants