Skip to content

Latest commit

 

History

History
19 lines (11 loc) · 3.79 KB

File metadata and controls

19 lines (11 loc) · 3.79 KB

Bounded native generation qualification

Native failures now return fixed diagnostic codes and next actions. Documented native error items are classified as rejected host-error observations, including coarse authentication, usage, context-limit and connectivity categories when present. Even a nonfatal error followed by a final response remains rejected under this qualification policy. Unsupported item kinds retain rejection and a type hash; provider error text, source and credentials are never copied into the diagnostic. This improves diagnosis without accepting tools, unknown items or late completions. Historical ITEM_UNKNOWN outcomes retain their original codes; a later diagnostic cannot retrospectively identify or qualify them.

scripts/codex-worker-proof.js is an explicit, local native exercise. Importing it and running the ordinary test suite never invokes a model. It requires clean committed source, an installed Codex CLI and an existing CLI-reported ChatGPT login; it does not install tools, initiate login, alter permissions or fall back to API credentials.

node scripts/codex-worker-proof.js --execute --output .tddswarm/worker-proof/new-run --timeout-ms 60000

Use a new output directory every time. Private receipts retain inputs, outputs, native audits, rejected attempts and independently executed outcomes. The fixed budget allows at most three controller role invocations—architect, author, reviewer—with no controller retries, a 60-second default per-role deadline and a 32 KiB response/event limit. CLI version and login-status probes consume the same role deadline. The model identity and provider billing remain unknown. Reported token counters are not a verified invoice or provider request count.

The exercise freezes one unsigned big-endian integer contract and three separately authored faults before generation. Reference tests demonstrate each fault in two native runs. The worker receives the correct source and behavior requirements; reference tests and fault payloads are excluded from its input and request directory. Generated tests must pass twice and catch each identical fault twice with stable named outcomes. A reviewer approval, more cases or a process exit code cannot substitute for those executions. Failed qualification has zero qualified detections; partial observations remain separately labeled.

The native adapter rejects reported tool attempts, malformed or incomplete events, missing terminal success, invalid role schemas, event/file disagreement, late results and launcher/schema/adapter drift. It preserves structured stdout compatibility and rejects preexisting response evidence. Worker protocol explains the event and process-group guarantees. Auditing does not establish OS read confinement, complete nested-runtime identity or a verified provider model.

This is one constructed contract, not a production corpus or an independent leaderboard. It does not test a learning effect or fully qualify an agent's MCP execution permissions. Frozen learning results and native host outcomes remain separate evidence.

The September 30 aggregate retains two attempts. The first completed two audited roles but was rejected before review because the harness had omitted its required file-layout rule from the prompt. The corrected retest completed three audited roles, produced 43 stable passing cases per baseline and detected all three identical faults twice. Five role invocations and 115.9 seconds across both attempts remain counted; qualified detections are 3/6 attempted fault opportunities. The same-contract retest is protocol and retention evidence, not an untouched quality or learning experiment. Native CLI-reported version was 0.155.1; the default model and billing remain unverified.