Skip to content

Commit 30c84c1

Browse files
IMNMVclaude
andcommitted
Referee Mode v2: configurable lenses, reviewer stances, model tiers, cross-vendor
ClaudeR 0.8.0 (R-side only; bridge unchanged at 0.10.0). Mode-collapse hardening for the referee lenses. Same-model parallel reviewers share blind spots; lens mandates decorrelate attention but not priors. New controls: - referee_prompt(lenses, reviewers_per_lens, model, cross_vendor): run any lens subset; 2 reviewers per lens = adversarial prosecutor + verifier stances, 3 adds a backwards reader; model takes one tier for all subagents ("haiku" quick / "opus" submission-grade) or a named per-lens vector, passed through to the host CLI subagent dispatch (Claude Code Task tool model parameter); cross_vendor = TRUE dispatches logic/methods reviewers to another vendor via codex/agy/ qwen one-shots -- the strongest decorrelation available - Anti-collapse rules now mandatory in the protocol: reviewer prompts written from scratch per lens and stance (template-with-one-word- swapped forbidden), consistency lens reads back-to-front, findings carry n_independent corroboration counts through dedup and into the report, and the run configuration is reported honestly - reviewer_zero_prompt(referee = TRUE) uses the configured defaults CI: referee v2 configuration test. R CMD check: Status OK. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1 parent 0443129 commit 30c84c1

6 files changed

Lines changed: 230 additions & 33 deletions

File tree

‎.github/scripts/checks.R‎

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -271,5 +271,21 @@ r <- tryCatch({
271271
}, error = function(e) conditionMessage(e))
272272
if (isTRUE(r)) pass("cross-reference checker: dangling + orphans, ranges resolve") else fail("crossref:", r)
273273

274+
# --- 12. referee mode v2 configuration ---
275+
r <- tryCatch({
276+
assign("system.file", function(..., package = NULL) file.path("inst", ...), envir = env)
277+
t1 <- capture.output(env$referee_prompt(lenses = c("logic", "methods"),
278+
reviewers_per_lens = 2, model = "haiku"))
279+
t2 <- capture.output(env$referee_prompt(model = c(logic = "opus"), cross_vendor = TRUE))
280+
any(grepl("Lenses: logic, methods", t1, fixed = TRUE)) &&
281+
any(grepl('model = "haiku"', t1, fixed = TRUE)) &&
282+
any(grepl("PROSECUTOR", t1, fixed = TRUE)) &&
283+
any(grepl('logic -> "opus"', t2, fixed = TRUE)) &&
284+
any(grepl("codex exec", t2, fixed = TRUE)) &&
285+
!any(grepl("{{", c(t1, t2), fixed = TRUE)) &&
286+
inherits(tryCatch(env$referee_prompt(lenses = "vibes"), error = function(e) e), "error")
287+
}, error = function(e) conditionMessage(e))
288+
if (isTRUE(r)) pass("referee v2: lenses, models, stances, cross-vendor, validation") else fail("referee v2:", r)
289+
274290
if (!ok) quit(status = 1)
275291
cat("\nAll checks passed.\n")

‎DESCRIPTION‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,6 @@
11
Package: ClaudeR
22
Title: R Integration for Claude AI
3-
Version: 0.7.0
3+
Version: 0.8.0
44
Authors@R: person("Nykko", "Vitali", email = "nykvt@icloud.com", role = c("aut", "cre"))
55
Description: Connects RStudio with Claude AI to enable interactive coding sessions.
66
License: MIT + file LICENSE

‎R/ui.R‎

Lines changed: 106 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -2457,15 +2457,92 @@ reviewer_zero_prompt <- function(prereg_path = NULL, robustness = FALSE,
24572457
}
24582458

24592459
if (isTRUE(referee)) {
2460-
ext_path <- system.file("prompts", "reviewer_zero_referee.md", package = "ClaudeR")
2461-
ext <- paste(readLines(ext_path, warn = FALSE), collapse = "\n")
2462-
txt <- paste0(txt, "\n", ext)
2460+
txt <- paste0(txt, "\n", build_referee_text())
24632461
}
24642462

24652463
cat(txt, "\n")
24662464
invisible(txt)
24672465
}
24682466

2467+
# Build the Referee Mode protocol text with a concrete run configuration.
2468+
# Shared by referee_prompt() and reviewer_zero_prompt(referee = TRUE).
2469+
build_referee_text <- function(lenses = c("logic", "methods", "consistency",
2470+
"evidence", "framing"),
2471+
reviewers_per_lens = 1L,
2472+
model = NULL,
2473+
cross_vendor = FALSE) {
2474+
valid <- c("logic", "methods", "consistency", "evidence", "framing")
2475+
bad <- setdiff(lenses, valid)
2476+
if (length(bad) > 0) {
2477+
stop(sprintf("Unknown lens(es): %s. Valid: %s.",
2478+
paste(bad, collapse = ", "), paste(valid, collapse = ", ")),
2479+
call. = FALSE)
2480+
}
2481+
if (!is.numeric(reviewers_per_lens) || reviewers_per_lens < 1 ||
2482+
reviewers_per_lens > 3) {
2483+
stop("`reviewers_per_lens` must be 1, 2, or 3 (stances: balanced; prosecutor+verifier; +backwards reader).",
2484+
call. = FALSE)
2485+
}
2486+
reviewers_per_lens <- as.integer(reviewers_per_lens)
2487+
2488+
model_directive <- if (is.null(model)) {
2489+
paste0("inherit the session's model for every reviewer subagent. If the ",
2490+
"user asked for a quick pass, prefer a fast tier (e.g. haiku or ",
2491+
"sonnet); for a submission-grade review, prefer the strongest ",
2492+
"tier available (e.g. opus or fable).")
2493+
} else if (is.null(names(model)) && length(model) == 1) {
2494+
sprintf(paste0("pass model = \"%s\" when dispatching EVERY reviewer ",
2495+
"subagent (the Task tool's model parameter on Claude ",
2496+
"Code, or your host's equivalent)."), model)
2497+
} else {
2498+
if (is.null(names(model)) || any(!nzchar(names(model)))) {
2499+
stop("`model` must be a single tier or a fully named vector, e.g. c(logic = \"opus\", consistency = \"haiku\").",
2500+
call. = FALSE)
2501+
}
2502+
bad_names <- setdiff(names(model), valid)
2503+
if (length(bad_names) > 0) {
2504+
stop(sprintf("`model` names must be lenses. Unknown: %s.",
2505+
paste(bad_names, collapse = ", ")), call. = FALSE)
2506+
}
2507+
paste0("per-lens models: ",
2508+
paste(sprintf("%s -> \"%s\"", names(model), model), collapse = "; "),
2509+
"; lenses not listed inherit the session's model.")
2510+
}
2511+
2512+
vendor_directive <- if (isTRUE(cross_vendor)) {
2513+
paste0(
2514+
"ENABLED. Where another vendor's CLI is installed (check with ",
2515+
"`which codex agy qwen` from Bash), dispatch at least one reviewer of ",
2516+
"the logic and methods lenses to a DIFFERENT model vendor as a ",
2517+
"one-shot subprocess: `codex exec` (pipe the reviewer prompt via ",
2518+
"stdin; flags: --skip-git-repo-check -c mcp_servers={}), ",
2519+
"`agy -p \"<prompt>\"`, or `qwen --prompt \"<prompt>\"`. First write ",
2520+
"the extracted manuscript to a plain-text file ",
2521+
"(writeLines(doc_lines, \"ms_extract.txt\")) and reference that path ",
2522+
"in the prompt, along with the lens mandate, stance, and the finding ",
2523+
"format. Cross-vendor findings enter the same registry and the same ",
2524+
"adjudication. Agreement across vendors is strong corroboration; ",
2525+
"vendor-unique findings deserve scrutiny in both directions. If no ",
2526+
"other vendor CLI is available, record that in the report and proceed ",
2527+
"single-vendor. Same-model reviewers share blind spots; a second ",
2528+
"vendor is the strongest decorrelation available."
2529+
)
2530+
} else {
2531+
"disabled for this run: all reviewers run as host-native subagents."
2532+
}
2533+
2534+
prompt_path <- system.file("prompts", "reviewer_zero_referee.md", package = "ClaudeR")
2535+
if (!nzchar(prompt_path) || !file.exists(prompt_path)) {
2536+
stop("Referee prompt template not found. Is ClaudeR installed correctly?")
2537+
}
2538+
txt <- paste(readLines(prompt_path, warn = FALSE), collapse = "\n")
2539+
txt <- gsub("{{LENSES}}", paste(lenses, collapse = ", "), txt, fixed = TRUE)
2540+
txt <- gsub("{{REVIEWERS_PER_LENS}}", as.character(reviewers_per_lens), txt, fixed = TRUE)
2541+
txt <- gsub("{{MODEL_DIRECTIVE}}", model_directive, txt, fixed = TRUE)
2542+
txt <- gsub("{{VENDOR_DIRECTIVE}}", vendor_directive, txt, fixed = TRUE)
2543+
txt
2544+
}
2545+
24692546
#' Print the Referee Mode prompt (standalone)
24702547
#'
24712548
#' Referee Mode is a substantive review of a manuscript's reasoning:
@@ -2476,14 +2553,34 @@ reviewer_zero_prompt <- function(prereg_path = NULL, robustness = FALSE,
24762553
#' numeric audit passes; to run it after a full audit, use
24772554
#' `reviewer_zero_prompt(referee = TRUE)`.
24782555
#'
2556+
#' @param lenses Which review lenses to run. Any subset of
2557+
#' `c("logic", "methods", "consistency", "evidence", "framing")`.
2558+
#' @param reviewers_per_lens 1, 2, or 3 independent reviewers per lens.
2559+
#' With 2, each lens gets a prosecutor (hunts flaws) and a verifier
2560+
#' (confirms each step); with 3, a backwards reader is added. Opposed
2561+
#' stances decorrelate reviewers built on the same model.
2562+
#' @param model Optional model directive for reviewer subagents. A single
2563+
#' tier applies to all lenses (e.g. `"haiku"` for a quick pass,
2564+
#' `"opus"` for a submission-grade review); a named vector sets tiers
2565+
#' per lens, e.g. `c(logic = "opus", consistency = "haiku")`. The
2566+
#' orchestrating agent passes this to its subagent dispatch (Claude
2567+
#' Code's Task tool `model` parameter). Default: inherit the session's
2568+
#' model.
2569+
#' @param cross_vendor If TRUE, the protocol instructs the orchestrator to
2570+
#' dispatch at least one logic and one methods reviewer to a different
2571+
#' model vendor (codex/agy/qwen one-shot CLI calls) where installed.
2572+
#' Cross-vendor agreement is the strongest available guard against
2573+
#' same-model blind spots.
24792574
#' @return The prompt text (invisibly), printed to the console.
24802575
#' @export
2481-
referee_prompt <- function() {
2482-
prompt_path <- system.file("prompts", "reviewer_zero_referee.md", package = "ClaudeR")
2483-
if (!nzchar(prompt_path) || !file.exists(prompt_path)) {
2484-
stop("Referee prompt template not found. Is ClaudeR installed correctly?")
2485-
}
2486-
txt <- paste(readLines(prompt_path, warn = FALSE), collapse = "\n")
2576+
referee_prompt <- function(lenses = c("logic", "methods", "consistency",
2577+
"evidence", "framing"),
2578+
reviewers_per_lens = 1L,
2579+
model = NULL,
2580+
cross_vendor = FALSE) {
2581+
txt <- build_referee_text(lenses = lenses,
2582+
reviewers_per_lens = reviewers_per_lens,
2583+
model = model, cross_vendor = cross_vendor)
24872584
cat(txt, "\n")
24882585
invisible(txt)
24892586
}

‎README.md‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -44,6 +44,7 @@ claudeAddin()
4444
<details>
4545
<summary><b>Recent Updates</b> (click to expand)</summary>
4646

47+
- **Referee Mode v2: configurable reviewers (R 0.8.0).** `referee_prompt()` now takes `lenses` (run any subset), `reviewers_per_lens` (2 = adversarial prosecutor + verifier pairs, 3 adds a backwards reader), `model` (a tier for all subagents, e.g. `"haiku"` for quick passes or `"opus"` for submission-grade, or a named per-lens vector), and `cross_vendor = TRUE` to dispatch logic/methods reviewers to a different model vendor via codex/agy/qwen one-shots. Anti-collapse rules are now mandatory in the protocol: reviewer prompts must be written from scratch per lens and stance, the consistency lens reads back-to-front, and every finding carries a corroboration count across independent reviewers.
4748
- **Referee Mode (R 0.7.0 / clauder-mcp 0.10.0).** The substantive manuscript review that paid services charge ~$50 a pass for, running free on the subscription you already have, and delivered where it belongs: as Word comments in your manuscript. `reviewer_zero_prompt(referee = TRUE)` (or standalone `referee_prompt()`) runs five content-only review lenses — argument logic, methods, internal consistency, evidence presentation, framing — as parallel subagents where the CLI supports them. Every finding must anchor to a verbatim quote and survive an independent verification pass before it lands in the document, severities are kept honest, and a clean report on a sound paper is a valid outcome. Alongside it, the new `check_cross_references` tool deterministically catches dangling references ("see Table 4" when Table 4 no longer exists) and tables or figures the text never mentions.
4849

4950
- **Value-first auditing (R 0.6.0 / clauder-mcp 0.9.0).** Built from field feedback after a full manuscript+supplement audit. `read_file` now transparently extracts `.docx`/`.pdf` (previously returned raw bytes), and the extractor preserves structure: headings marked, table cells emitted row-wise with separators (they were previously dropped entirely). New `reconcile_values` tool: enumerates every number in a manuscript and reconciles each against the corpus your code produced, respecting displayed precision (5038.5 matches 5038.46), commas, percents, scientific notation, and `< .001` thresholds; the per-value `values_registry` makes numeric completeness a construction, not a diligence hope. Reviewer Zero now sets audit-clean print options (no more tibble 3-sig-fig false alarms), gates on the value sweep, and takes final verdicts from clean-room runs via `probe_scripts(capture_output = TRUE)`. CrossRef lookups retry with backoff instead of silently truncating the reference check on 429s.
@@ -189,6 +190,17 @@ reviewer_zero_prompt(writeback = TRUE)
189190

190191
# All extensions together
191192
reviewer_zero_prompt(prereg_path = "prereg.docx", robustness = TRUE, writeback = TRUE)
193+
194+
# Referee Mode, configured. Quick-and-dirty pass on a fast model tier:
195+
referee_prompt(model = "haiku")
196+
197+
# Submission-grade: strongest tier, adversarial reviewer pairs per lens
198+
# (prosecutor + verifier), and lenses dispatched across model vendors
199+
# (codex/agy/qwen) so same-model blind spots can't hide the same flaw twice:
200+
referee_prompt(model = "opus", reviewers_per_lens = 2, cross_vendor = TRUE)
201+
202+
# Mix tiers per lens: deep model where the reasoning is hard, fast elsewhere
203+
referee_prompt(model = c(logic = "opus", methods = "opus", consistency = "haiku"))
192204
```
193205

194206
The protocol works with `.docx`, `.pdf`, `.qmd`, `.Rmd`, `.tex`, or plain text manuscripts and supports multi-script R projects.

‎inst/prompts/reviewer_zero_referee.md‎

Lines changed: 66 additions & 22 deletions
Original file line numberDiff line numberDiff line change
@@ -27,12 +27,14 @@ table/figure becomes a finding. If a class is reported unverifiable
2727

2828
### Step R2: Review lenses
2929

30-
Five lenses, each with exactly one lane. If your environment supports
31-
subagents (Claude Code's Task tool, Codex agents), dispatch them in
32-
parallel, one lens per subagent, each instructed to read the ENTIRE
33-
manuscript via paginated `read_file` and report findings only in its lane.
34-
Without subagents, run the lenses sequentially yourself, completing one
35-
before starting the next.
30+
Configuration for this run:
31+
32+
- Lenses: {{LENSES}}
33+
- Independent reviewers per lens: {{REVIEWERS_PER_LENS}}
34+
- Subagent model directive: {{MODEL_DIRECTIVE}}
35+
- Cross-vendor dispatch: {{VENDOR_DIRECTIVE}}
36+
37+
The lens mandates:
3638

3739
1. **logic** -- Does each conclusion follow from what is established?
3840
Unsupported leaps, circular arguments, quantifier slips, claims proved
@@ -55,46 +57,84 @@ before starting the next.
5557
explanations left unaddressed, positioning gaps (flag, do not demand
5658
specific citations).
5759

60+
If your environment supports subagents (Claude Code's Task tool, Codex
61+
agents), dispatch reviewers in parallel, one subagent per reviewer, each
62+
instructed to read the ENTIRE manuscript via paginated `read_file` and
63+
report findings only in its lane. Without subagents, run the reviewers
64+
sequentially yourself, completing one before starting the next.
65+
66+
#### Anti-collapse rules (mandatory)
67+
68+
Parallel reviewers built on the same model share blind spots. These rules
69+
exist to break that correlation; do not skip them:
70+
71+
1. Write each reviewer's prompt from scratch for its lens and stance. The
72+
prompts may share only the output format. If two of your subagent
73+
prompts differ by one sentence, rewrite them.
74+
2. Reviewer stances per lens, by reviewer count:
75+
- 1 reviewer: balanced examiner.
76+
- 2 reviewers: a PROSECUTOR ("find where this lane's reasoning fails;
77+
assume there is a flaw and hunt it") and a VERIFIER ("attempt to
78+
confirm each relevant step actually holds; report every step you
79+
could not confirm").
80+
- 3 reviewers: prosecutor, verifier, and a BACKWARDS reader.
81+
3. Regardless of count, the consistency lens's first reviewer reads the
82+
manuscript BACK TO FRONT: conclusions first, then checking whether the
83+
premises for each claim were ever established earlier. This traversal
84+
catches promissory abstracts and unproven dependencies that
85+
front-to-back reading normalizes.
86+
4. Convergence is information, not waste: findings surfaced independently
87+
by multiple reviewers are corroborated. Track how many independent
88+
reviewers surfaced each finding.
89+
5890
Every finding uses this exact structure:
5991

6092
- `lens`: which lens produced it
93+
- `reviewer`: which reviewer (e.g. "logic/prosecutor", "logic/verifier")
6194
- `severity`: `major` (undermines a conclusion) / `moderate` (weakens or
6295
confuses an argument) / `minor` (substantive but small)
6396
- `anchor`: a VERBATIM quote of 10-30 characters from the manuscript at the
6497
location of the issue (copy-paste; it will be existence-checked)
6598
- `comment`: the referee comment, written to the author: specific, concrete,
6699
and stating why it matters. One issue per finding.
67100

68-
A lens that finds nothing in its lane must say so explicitly rather than
69-
inventing filler findings.
101+
A reviewer that finds nothing in its lane must say so explicitly rather
102+
than inventing filler findings.
70103

71104
### Step R3: Adjudication
72105

73106
Merge all findings (including Step R1's) into a registry:
74107

75108
```r
76109
referee_registry <- data.frame(
77-
finding_id = integer(0), lens = character(0), severity = character(0),
78-
anchor = character(0), comment = character(0),
79-
status = character(0), # "confirmed" or "dropped"
80-
reason = character(0), # required when dropped
110+
finding_id = integer(0), lens = character(0), reviewer = character(0),
111+
severity = character(0), anchor = character(0), comment = character(0),
112+
n_independent = integer(0), # reviewers that surfaced this finding
113+
status = character(0), # "confirmed" or "dropped"
114+
reason = character(0), # required when dropped
81115
stringsAsFactors = FALSE
82116
)
83117
```
84118

85-
Then verify EVERY finding yourself; subagent output is a draft, not a
119+
Then verify EVERY finding yourself; reviewer output is a draft, not a
86120
verdict:
87121

88-
1. Prove the anchor exists:
122+
1. Deduplicate first: findings from different reviewers describing the
123+
same issue merge into one row with `n_independent` set to the count of
124+
distinct reviewers that surfaced it.
125+
2. Prove the anchor exists:
89126
`stopifnot(any(grepl(anchor, doc_lines, fixed = TRUE)))` --
90127
a finding whose anchor is not in the document is dropped as
91128
hallucinated, whatever its content.
92-
2. Re-read the anchored passage with surrounding context. Drop findings
93-
that misread the text, duplicate another finding, or are stylistic
94-
despite the rules. Record the reason.
95-
3. Sanity-check severity: downgrade anything a careful author would
129+
3. Re-read the anchored passage with surrounding context. Drop findings
130+
that misread the text or are stylistic despite the rules. Record the
131+
reason.
132+
4. Sanity-check severity: downgrade anything a careful author would
96133
shrug at. `major` must be reserved for issues that change what a
97-
reader should believe.
134+
reader should believe. A finding with `n_independent >= 2` (and
135+
especially one surfaced across model vendors) warrants extra care
136+
before dropping; a single-reviewer finding warrants extra care before
137+
confirming as major.
98138

99139
#### Referee gate
100140

@@ -118,8 +158,10 @@ ClaudeR::annotate_manuscript(
118158
)
119159
```
120160

121-
Add a "Referee Review" section to the Final Report: counts by severity and
122-
lens, the confirmed findings ranked most severe first, dropped findings
161+
Add a "Referee Review" section to the Final Report: the reviewer
162+
configuration that ran (lenses, reviewers per lens, models/vendors used),
163+
counts by severity and lens, the confirmed findings ranked most severe
164+
first with their `n_independent` corroboration counts, dropped findings
123165
with reasons (so the user can audit your filtering), the cross-reference
124166
results, and the annotated file path if one was written.
125167

@@ -129,10 +171,12 @@ results, and the annotated file path if one was written.
129171
belong here.
130172
2. Every finding is anchored to a verbatim quote and the anchor is
131173
grepl-proven before the finding survives.
132-
3. Verify subagent findings independently; you own every comment that
174+
3. Verify reviewer findings independently; you own every comment that
133175
lands in the author's document.
134176
4. Severity honesty: three real major findings help the author more than
135177
thirty inflated ones.
136178
5. Absence is reportable: if a lens legitimately finds nothing, the report
137179
says so. A clean referee report on a sound paper is a success, not a
138180
failure to find things.
181+
6. Report the configuration honestly, including any lens you could not
182+
dispatch as requested (e.g. a vendor CLI not installed).

0 commit comments

Comments
 (0)