Skip to content

v0.2 Roadmap: CodeRabbit-Class Reviewer Pilot #3

Description

@100yenadmin

Goal

Turn evaos-code-review-bot from the v0.1 safe ZCode inline reviewer into a CodeRabbit-class EVAOS reviewer with calibrated confidence, richer PR context, and low-noise repo policy controls.

Non-goals

  • No auto-merge.
  • No APPROVE reviews.
  • No all-org rollout.
  • No repo mutation by default review runs.
  • No public 95% calibrated confidence until measured reliability supports it.

Current state

  • v0.1 merged in PR Build ZCode PR review bot pilot #2 at merge commit 2dfad82600719c359e40df10f4ce0e92b5b7efbb.
  • Launchd live canary is running from main with dryRun=false and canary PRs electricsheephq/WorldOS#1161, 100yenadmin/evaOS-GUI#497.
  • Current bot deliberately disables native ZCode skills/MCP/memory and permits only read-only inspection tools.
  • GitNexus is available locally for indexed repos but is not integrated into the bot.

Ordered work

  1. Add CodeRabbit-style walkthrough and pre-merge summary.
  2. Add repo policy config for review profile, path filters/instructions, labels/reviewers, checks, and finishing-touch commands.
  3. Add read-only GitNexus context provider with stale/missing-index degraded mode.
  4. Add curated read-only skill-pack prompt injection before any native skill: true enablement.
  5. Add confidence calibration and public confidence display rules.
  6. Add CodeRabbit/human comparison eval harness.
  7. Add issue/PR enrichment lane.
  8. Add opt-in finishing touches commands.

Acceptance criteria

  • Each child issue has tests/evidence and no expansion of GitHub App permissions unless explicitly justified.
  • The default review path remains read-only and App-authored.
  • Duplicate suppression remains per {repo, pr, head_sha}.
  • No secrets are written to comments, logs, or shareable evidence.
  • P0/P1 REQUEST_CHANGES remains gated by validated findings.
  • Public confidence labels say uncalibrated until calibration thresholds pass.

Validation / eval gates

  • Eval required: yes
  • Eval claim class: advisory
  • Required eval suites: canary_shadow, historical_pr_replay, seeded_defect_recall, safety_redaction, duplicate_suppression, coderabbit_comparison
  • Eval name/version: evaos-code-review-bot-v0.2-roadmap
  • Dataset/scenario refs: WorldOS#1161, evaOS-GUI#497, historical CodeRabbit-reviewed PRs, seeded P0/P1 fixtures, negative-control docs-only PRs
  • Baseline/comparison: CodeRabbit comments, human review outcomes, CI failures, merged-fix diffs
  • Metrics and thresholds: 100% current-line validity; 100% secret redaction on seeded cases; zero mutation; no duplicate posting; P0/P1 calibrated precision Wilson lower bound >=95% before displaying 95% calibrated
  • Runner/CI location: local focused eval packets under Lexar, with later GitHub Actions shadow mode if useful
  • Failure owner: bot maintainer
  • Eval evidence path: /Volumes/LEXAR/Codex/evals/zcode-glm-pr-review/<date>/<run-id>/
  • Trace feedback target: this tracker, child issues, JSONL eval artifacts
  • Eval proof boundary: proves pilot reviewer quality on scoped repos/datasets only; does not prove production all-org readiness or human-review replacement.

Evidence / resume links

Current status

Created after v0.1 merge. Exact next action is to implement the walkthrough/pre-merge summary issue first while keeping live canary allowlist in place for 24h.

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions