Repository navigation
RFC: Agent quality scoring with promptfoo eval harness #434
Replies: 2 comments
|
I’d like to help turn this into evidence people can use when selecting a specialist. I read #371 and the revert in #433; I’m treating this RFC as the coordination point, and will preserve @jonesrussell’s attribution if adapting that work. A concrete first slice I’m preparing locally is a zero-API-call scorer for user-supplied outputs, with three checkable engineering/testing task packs: seeded defect review, API boundary cases and a documented failure scenario. It pairs baseline/specialist runs under the same recorded model/prompt revision/tools/input/limits, reports correct findings, misses and false positives, and produces evidence cards with sample size, setup, failures, measured or unknown latency/cost. Subjective quality stays a separate blinded human rubric; a model judge is not the sole evidence, and synthetic scorer fixtures will not be presented as evidence that an agent improved a model. I can contribute the protocol and card examples through the existing Tool Evaluator agent content first, which CONTRIBUTING welcomes, while keeping the executable prototype outside the upstream tree. Before any new harness directories/scripts/CI PR: would you prefer a companion project, or an optional local-only tool here? Does this objective-first supplied-output approach fit the RFC’s scope? No mandatory paid CI gate is proposed. |
|
sure hel
…On Sat, 3 Oct 2026 at 7:42 am, Rudy Mizrahi Celekli < ***@***.***> wrote:
I’d like to help turn this into evidence people can use when selecting a
specialist. I read #371
<#371> and the revert
in #433 <#433>; I’m
treating this RFC as the coordination point, and will preserve
@jonesrussell <https://github.com/jonesrussell>’s attribution if adapting
that work.
A concrete first slice I’m preparing locally is a zero-API-call scorer for
user-supplied outputs, with three checkable engineering/testing task packs:
seeded defect review, API boundary cases and a documented failure scenario.
It pairs baseline/specialist runs under the same recorded model/prompt
revision/tools/input/limits, reports correct findings, misses and false
positives, and produces evidence cards with sample size, setup, failures,
measured or unknown latency/cost. Subjective quality stays a separate
blinded human rubric; a model judge is not the sole evidence, and synthetic
scorer fixtures will not be presented as evidence that an agent improved a
model.
I can contribute the protocol and card examples through the existing Tool
Evaluator agent content first, which CONTRIBUTING welcomes, while keeping
the executable prototype outside the upstream tree. Before any new harness
directories/scripts/CI PR: would you prefer a companion project, or an
optional local-only tool here? Does this objective-first supplied-output
approach fit the RFC’s scope? No mandatory paid CI gate is proposed.
—
Reply to this email directly, view it on GitHub
<#434?email_source=notifications&email_token=CEDVMPJHUGC3D2RDTAAADND5SAONDA5CNFSNUABIM5UWIORPF5TWS5BNNB2WEL2ENFZWG5LTONUW63SDN5WW2ZLOOQXTCOBXGIZDMNBUUZZGKYLTN5XKU43VMJZWG4TJMJSWJJLFOZSW45FMMZXW65DFOJPWG3DJMNVQ#discussioncomment-18722644>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/CEDVMPKC5R7QRN3GD2ZJMFD5SAONDAVCNFSNUABIKJSXA33TNF2G64TZHMYTANZVGM3TENJUGU5UI2LTMN2XG43JN5XDWOJYGYZDGMZSUF3AE>
.
You are receiving this because you are subscribed to this thread.Message
ID: <msitarzewski/agency-agents/repo-discussions/434/comments/18722644@
github.com>
|
Uh oh!
There was an error while loading. Please reload this page.
Context
PR #371 by @jonesrussell introduced a promptfoo eval harness for agent quality scoring — an LLM-as-judge system that scores agents on task completion, instruction adherence, identity consistency, deliverable quality, and safety.
The PR was merged and then reverted in #433 because per CONTRIBUTING.md, new tooling and CI infrastructure should go through a Discussion first to get community alignment.
The Proposal
An
evals/directory with:Questions for the Community
agency-agents-evalsrepo?The original PR (#371) is ready to re-submit once we align on approach. @jonesrussell — thank you for the excellent work, and apologies for the revert. We want to make sure this lands with full community buy-in.
Ref: PR #371, reverted in #433
All reactions