Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 47 additions & 0 deletions HELDOUT-REPORT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
# Held-out retrieval eval

Independent eval for issue #24. Questions were written from the 17 verbatim raw sources under:

`/Users/ved/Developer/grimoire/.growth/truth/workspaces/agent-sdk-claude/raw/uncategorized/`

Wiki contents were not read; only wiki slugs were listed. Excluded files from `SPEC.md` were not opened or modified. Ranking code was not tuned.

## Commands

```sh
npm run build
node scripts/retrieval-eval.mjs /Users/ved/Developer/grimoire/.growth/truth/workspaces/agent-sdk-claude test/fixtures/retrieval-questions.heldout.json --no-fail
npm test
```

## Results

| Set | Present top-1 | Present top-3 | Absent abstain | All top-1-or-abstain | All top-3-or-abstain | Gate |
| --- | ---: | ---: | ---: | ---: | ---: | --- |
| Tuned set from PR #21 | 10/10 | 10/10 | n/a | 10/10 | 10/10 | PASS |
| Held-out set | 9/10 | 10/10 | 0/2 | 9/12 | 10/12 | FAIL |

## Per-question output

| ID | Coverage | Top-1 | Top-3 | No-hit | Expected | Returned |
| ---: | --- | --- | --- | --- | --- | --- |
| 1 | present | yes | yes | - | `permissions-and-approval-control` | `permissions-and-approval-control`, `hooks`, `production-security-and-sandboxing` |
| 2 | present | yes | yes | - | `custom-tools-in-process-mcp` | `custom-tools-in-process-mcp`, `agent-loop-and-turns`, `hooks` |
| 3 | present | yes | yes | - | `file-checkpointing` | `file-checkpointing`, `permissions-and-approval-control`, `what-is-the-agent-sdk` |
| 4 | present | yes | yes | - | `external-mcp-integration` | `external-mcp-integration`, `custom-tools-in-process-mcp`, `tool-search` |
| 5 | present | - | yes | - | `hosting-and-deployment-patterns` | `sessions-continue-resume-fork`, `runtime-and-subprocess-architecture`, `hosting-and-deployment-patterns` |
| 6 | present | yes | yes | - | `observability-opentelemetry` | `observability-opentelemetry`, `hooks`, `authentication-and-providers` |
| 7 | present | yes | yes | - | `migration-from-claude-code-sdk` | `migration-from-claude-code-sdk`, `what-is-the-agent-sdk`, `runtime-and-subprocess-architecture` |
| 8 | present | yes | yes | - | `sessions-continue-resume-fork` | `sessions-continue-resume-fork`, `runtime-and-subprocess-architecture`, `file-checkpointing` |
| 9 | present | yes | yes | - | `subagents` | `subagents`, `observability-opentelemetry`, `permissions-and-approval-control` |
| 10 | present | yes | yes | - | `tool-search` | `tool-search`, `external-mcp-integration`, `custom-tools-in-process-mcp` |
| 11 | absent | - | - | - | absent | `what-is-the-agent-sdk`, `migration-from-claude-code-sdk`, `runtime-and-subprocess-architecture` |
| 12 | absent | - | - | - | absent | `what-is-the-agent-sdk`, `custom-tools-in-process-mcp`, `hosting-and-deployment-patterns` |

## Finding

The independent set does not reproduce the tuned set's 10/10 gate. Retrieval is strong for covered topics at top-3, but the current behavior does not abstain on absent topics: both absent questions returned wiki articles instead of no hit.

## Test status

`npm test` passed: 23 test files, 368 tests.
50 changes: 50 additions & 0 deletions test/fixtures/retrieval-questions.heldout.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
[
{
"question": "If I set a PreToolUse hook to allow a command, can a deny rule still block it later in the permission flow?",
"expectedSlug": "permissions-and-approval-control"
},
{
"question": "For a custom in-process tool, what happens to the agent loop if my handler throws instead of returning an error result?",
"expectedSlug": "custom-tools-in-process-mcp"
},
{
"question": "Can file rewind undo changes made by a shell command like sed, or only SDK edit operations?",
"expectedSlug": "file-checkpointing"
},
{
"question": "Where should I check for failed MCP server connections before letting an agent start using those tools?",
"expectedSlug": "external-mcp-integration"
},
{
"question": "When hosting many active agent sessions, why do I need separate working directories or per-session cwd values?",
"expectedSlug": "hosting-and-deployment-patterns"
},
{
"question": "Which telemetry switch records full tool inputs and outputs, and why should it stay off unless the pipeline is approved?",
"expectedSlug": "observability-opentelemetry"
},
{
"question": "What changed in the rename from the old Claude Code SDK packages to the Agent SDK packages?",
"expectedSlug": "migration-from-claude-code-sdk"
},
{
"question": "How can I fork a conversation to try another direction without changing the original session history?",
"expectedSlug": "sessions-continue-resume-fork"
},
{
"question": "If a subagent needs to run without seeing the parent conversation history, what information must the parent provide?",
"expectedSlug": "subagents"
},
{
"question": "Why might enabling deferred tool discovery improve both prompt size and tool selection once a server exposes hundreds of tools?",
"expectedSlug": "tool-search"
},
{
"question": "How do I configure vector embedding dimensions and chunk overlap for the Agent SDK retriever?",
"expectedSlug": null
},
{
"question": "Which managed Postgres schema does the Agent SDK create for storing durable task queues?",
"expectedSlug": null
}
]