Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 68 additions & 0 deletions qa/evidence/plate-sprint/outdoor-lora/SMOKE-TEST-VERDICT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
# OUTDOOR-LORA smoke test verdict (2026-07-10)

Model: `model_RsWEcQL2NWXwoyEodWVE2vWG` ("WorldOS Painterly Exterior (FLUX)"), trained on the
18-image quality-passed set in `training_manifest.md` (job `job_U3HDsSj7T4aPETy7MRCuX6oK`,
1080 CU, ~172 min).

Pipeline per room: mint anchor via schema-correct `model_run` (`model_bfl-flux-1-dev` + depth
ControlNet str 0.65/end 0.6 + the new LoRA, `lorasScale` 0.7) -> Gemini instruction-edit
re-registration (`model_google-gemini-3-1-flash`, structure-lock + dimetric-lock prompt,
`referenceImages=[minted anchor]`) over the room's own greybox -> overlay + advisory edge-recall
(`qa/plate_overlays.py` / `advisory_eval.py`, content-blind per #1491) -> 5-scorer blind panel
(control + reference/incumbent + candidate, shuffled, `qa/scores_db.py` surface=visual).

## Results

| Room | Candidate median | Reference | Baseline cap | Success bar | Result |
|---|---|---|---|---|---|
| forest_road | **2.0 / 10** | cross-lane anchor (camp_clearing_night_v2) 7.0, control 9.0 | 6.0 | >6.5 | **SEVERE REGRESSION** |
| camp_clearing_night | **6.5 / 10** | incumbent (camp_clearing_night_v2, actual adopted plate) 6.0, control 9.0 | 6.0 | >6.5 | **BORDERLINE** (right at the line, not clearly above) |

Full scores, mapping, and defect notes: `forest_road/panel_verdict.json`, `camp/panel_verdict.json`.
scores_db rows: `outdoor-lora-smoke-forest_road-2026-07-10`, `outdoor-lora-smoke-camp-2026-07-10`.

## Headline finding

The new exterior-only, quality-passed LoRA **does not uniformly fix the outdoor generalization
gap** the effort was built to close (issue #1481). It generalizes cleanly to **camp**
(no invented architecture, no characters, modest +0.5 median lift over the incumbent — though
short of a clean pass) but **regresses severely on forest_road**: all 5 blind scorers
independently flagged broken/wireframe/"melted" architecture — the greybox's dense tree-line box
volumes got repainted as disconnected wooden shrine/hut structures with carved doors and
curtains, not trees. This is the SAME invented-architecture failure mode ARM C (the interior
LoRA) produced on this room (CB-FOREST precedent, median 6.0 there) — just worse (median 2.0)
and with a different visual signature (broken wireframe vs solid stone).

**Root-cause hypothesis (new, beyond the original training-data-domain framing):**
forest_road's greybox represents its tree line as adjacent axis-aligned rectangular box volumes
(`forest_road_greybox.png` / `forest_road_greybox_depth.png`) — a coarser, more literally
"architectural-looking" placeholder shape than camp's greybox. The model appears to read blocky
rectangular volumes as buildings regardless of whether its LoRA was trained on interior or
exterior images — a **geometry-shape bias**, not purely a training-domain bias. Edge-recall
(the automated advisory metric) is misleading here: forest_road scored a deceptively HIGH 0.9678
specifically because the candidate preserved the box edges too literally (as walls, not trees),
while camp's genuinely better result scored a LOWER 0.6645 because it correctly reinterpreted
the boxes as rounded organic forms. The overlay images are the binding evidence, not the number

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Compress the committed evidence frames

This verdict now cites the overlay images as the binding visual evidence, but the newly committed PNG frames in this evidence directory are all over the repo’s visual-evidence cap from qa/evidence/README.md/docs/OPERATIONS.md (find shows ~1.19MB to ~5.75MB per frame, versus ≤400KB/frame). Merging these oversized frames bloats the repo and violates the evidence contract reviewers rely on; please downsample/compress the stills to reviewer-visible JPEG/PNG files under the cap.

Useful? React with 👍 / 👎.

(confirmed both visually and by the 5-scorer panel, independent of the metric).

## What this means for issue #1481

- The OUTDOOR-LORA is a genuine, validated improvement for **camp-class** rooms (open clearings
with sparse, discrete tree/prop placeholders) but is **not ready to replace** ARM C or be
declared the fix for forest_road-class rooms (dense linear tree corridors with adjacent box
placeholders) without either (a) a differently-authored greybox for that structure class
(more organic/clustered tree footprints instead of axis-aligned boxes), or (b) further LoRA/
prompt iteration specifically targeting box-shaped placeholder reinterpretation.
- Do not adopt either candidate as a new canonical_plate off this single smoke test. camp's
result is a borderline-promising lead worth one more iteration (fix the firepit ring artifact);
forest_road's result should be treated as a confirmed non-fix.

## Cost actuals (record on #1481)

- Training: 1080 CU (job_U3HDsSj7T4aPETy7MRCuX6oK, ~172 min, model_RsWEcQL2NWXwoyEodWVE2vWG)
- Smoke test: 4 anchor mints @ 9 CU = 36 CU + 2 Gemini re-registration passes @ 20 CU = 40 CU
-> 76 CU
- **Total: 1156 CU (~$22.13 at the ~$0.0191/CU rate implied by the 1080 CU dry-run estimate)**
- Account model-count limit could not be read via API (403, role scope) — noted for any future
training on this account; if creation/training ever fails on a count-limit error, stop and
report, do not delete existing models.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
14 changes: 14 additions & 0 deletions qa/evidence/plate-sprint/outdoor-lora/camp/gemini_meta.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
{
"image_ref": "asset_JacJoMavdMzU4VMGj3K13Bkp",
"requested_reference_images": true,
"reference_asset_id": "asset_MYs9LVhFtvEv7wJUfDmu1oCW",
"job_id": "job_pPSriYXXqnm6sB4Xet3msWNW",
"assets": [
{
"asset_id": "asset_zxqvrguC2iSXztMdrtsi4Zvu",
"path": "/Users/lume/worldos-session-notes/plate-sprint/camp/outdoor_lora/camp_outdoorlora_registered.png",
"bytes": 5319843
}
],
"prompt": "Repaint this image in the painterly hand-painted CRPG environment-background style of the reference image: visible oil-brush strokes, muted painterly palette, warm firelit key light from the central campfire falling off into cool blue-violet shadow at the clearing's edge, dappled moonlight through canopy gaps. The blocky grey placeholder trees, crates, and rocks must be repainted as natural organic tree trunks with bark texture and foliage, rounded moss-covered boulders, and weathered wooden crates/bedrolls with visible grain -- never flat-faced geometric boxes or glassy/translucent prisms.\n\nCRITICAL STRUCTURE-LOCK (non-negotiable, checked pixel-by-pixel against the original): every wall, pillar, archway, doorway, staircase, tree, boulder, road edge, and prop must stay in EXACTLY its current position, size, and shape -- do NOT move, add, remove, resize, re-project, or recompose ANYTHING; only the paint and lighting treatment changes. Do NOT introduce any repeated, cloned, or tiled decorative motifs. Do NOT add any text, letters, numbers, labels, legends, map insets, UI elements, frames, or borders. Do NOT add any characters, people, figures, or creatures.\n\nThis is a TOP-DOWN DIMETRIC (2:1 isometric) game BACKGROUND PLATE viewed at a fixed ~30-degree downward camera pitch -- you are looking DOWN onto the scene from high above and to the side, exactly like a Pillars of Eternity II pre-rendered isometric level. Keep that EXACT high dimetric bird's-eye angle: there is NO horizon and NO sky band across the top of the frame -- the top of the frame is MORE ground/canopy seen from above. Do NOT tilt the camera to an eye-level, first-person, or landscape-horizon view."
}
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
5 changes: 5 additions & 0 deletions qa/evidence/plate-sprint/outdoor-lora/camp/panel_mapping.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
{
"img_1.jpg": "CONTROL",
"img_2.jpg": "INCUMBENT_CAMP_V2",
"img_3.jpg": "CANDIDATE_CAMP_OUTDOORLORA"
Comment on lines +2 to +4

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Emit the panel-mapping contract consumed by the verifier.

This file omits slot_to_label and label_to_slot, and its labels do not match the A/B/C contract generated by qa/plate_loop.py at Line 356-430. As written, downstream ingestion cannot reliably identify candidate, incumbent, or control medians.

Normalize this file and the forest mapping to one exact slot namespace:

Proposed contract
-{
-  "img_1.jpg": "CONTROL",
-  "img_2.jpg": "INCUMBENT_CAMP_V2",
-  "img_3.jpg": "CANDIDATE_CAMP_OUTDOORLORA"
-}
+{
+  "slot_to_label": {
+    "image_1": "C",
+    "image_2": "B",
+    "image_3": "A"
+  },
+  "label_to_slot": {
+    "A": "image_3",
+    "B": "image_2",
+    "C": "image_1"
+  },
+  "labels": {
+    "A": "candidate",
+    "B": "incumbent-canonical",
+    "C": "disguised-real-art-control"
+  }
+}

If img_*.jpg is the canonical slot naming, use those exact keys consistently in both the mapping and verdict scores.

Confidence: 99%.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"img_1.jpg": "CONTROL",
"img_2.jpg": "INCUMBENT_CAMP_V2",
"img_3.jpg": "CANDIDATE_CAMP_OUTDOORLORA"
{
"slot_to_label": {
"image_1": "C",
"image_2": "B",
"image_3": "A"
},
"label_to_slot": {
"A": "image_3",
"B": "image_2",
"C": "image_1"
},
"labels": {
"A": "candidate",
"B": "incumbent-canonical",
"C": "disguised-real-art-control"
}
}
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@qa/evidence/plate-sprint/outdoor-lora/camp/panel_mapping.json` around lines 2
- 4, Normalize the outdoor LoRA panel mapping and corresponding forest mapping
to the A/B/C contract produced by qa/plate_loop.py, including both slot_to_label
and label_to_slot fields. Use the canonical img_*.jpg slot names consistently in
mappings and verdict scores, with labels matching the verifier’s candidate,
incumbent, and control assignments; update all related mapping data accordingly.

}
28 changes: 28 additions & 0 deletions qa/evidence/plate-sprint/outdoor-lora/camp/panel_verdict.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
{
"run": "camp-outdoor-lora-smoke-2026-07-10",
"arm": "OUTDOOR-LORA smoke test (#1481): mint anchor via schema-correct model_run (model_bfl-flux-1-dev + depth ControlNet str0.65/end0.6 + model_RsWEcQL2NWXwoyEodWVE2vWG, the new 18-image exterior-only LoRA, lorasScale 0.7, seed 42) -> Gemini instruction-edit re-registration (model_google-gemini-3-1-flash, structure-lock + dimetric-lock prompt, referenceImages=[minted anchor]) over camp_clearing_night_greybox.jpg.",
"panel_mapping": {"img_1.jpg": "CONTROL", "img_2.jpg": "INCUMBENT_CAMP_V2", "img_3.jpg": "CANDIDATE_CAMP_OUTDOORLORA"},
"scores": {
"img_1_CONTROL_poe2": [9, 9, 9, 9, 9],
"img_2_INCUMBENT_CAMP_V2": [6, 6, 7, 6, 6],
"img_3_CANDIDATE": [6, 7, 6, 8, 6.5]
},
"medians": {"CONTROL": 9.0, "INCUMBENT_CAMP_V2": 6.0, "CANDIDATE": 6.5},
Comment on lines +4 to +10

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Keep panel keys and medians in one canonical namespace.

panel_mapping declares img_1.jpg, but scores uses img_1_CONTROL_poe2, img_2_INCUMBENT_CAMP_V2, and img_3_CANDIDATE. These keys cannot be joined. Additionally, ingest_verdict expects A/B/C after relabeling, while medians uses CONTROL, INCUMBENT_CAMP_V2, and CANDIDATE; downstream cand, ctrl, and delta_vs_control therefore become null.

Normalize the embedded mapping, score keys, and medians together.

Confidence: 99%.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@qa/evidence/plate-sprint/outdoor-lora/camp/panel_verdict.json` around lines 4
- 10, Normalize panel_mapping, scores, and medians to one canonical key
namespace so every panel identifier joins consistently. Update the embedded
mapping and score keys together, and use the same normalized labels expected by
ingest_verdict’s A/B/C relabeling, ensuring downstream cand, ctrl, and
delta_vs_control are populated rather than null.

"control_valid": false,
"control_integrity_note": "Control (poe2_control.jpg) scored median 9.0 with human figures visible and flagged by all 5 scorers -- same imperfect-control caveat as the forest_road panel and the CB-FOREST precedent; scorers judged craft only as instructed.",
"incumbent_cross_validation": "Incumbent (camp_clearing_night_v2.jpg, the room's actual adopted canonical_plate) scored median 6.0 blind, independently reproducing its already-registered backdrop_score=6.0 in room_recipes.json -- this cross-validates the blind-panel methodology (scorers were not told which image was which, yet landed on the known number).",
"candidate_defect_notes": [
"img_3: 6 | odd concentric-ring ground texture radiating from firepit reads as a repeated/tiled artifact, hanging bags on tree at right lack a visible attachment point",
"img_3: 7 | clean brushwork and lighting, faint repetitive concentric rings in the dirt around the firepit read slightly artificial",
"img_3: 6 | concentric ring/banding pattern in dirt around firepit looks like an artificial glow-radius artifact rather than natural ground texture",
"img_3: 8 | hanging satchels upper-right look faintly disconnected/floating from the branch, no characters or structures, otherwise clean",
"img_3: 6.5 | decent fire/smoke and moss-rock brushwork, but crate stack has a mildly ambiguous overlap/perspective and hanging satchels read slightly flat/pasted-on"
],
"advisory_metrics": {
"ncc_candidate_vs_base": 0.8749,
"ncc_candidate_vs_greybox": 0.4558,
"edge_recall_vs_greybox": 0.6645,
"edge_recall_caveat": "Content-blind per #1491 -- lower than forest_road's number DESPITE being the visually cleaner, better result, because the candidate correctly reinterpreted the greybox's rectangular placeholder volumes as rounded organic forms (crates, boulders) rather than preserving their blocky edges literally. Confirms edge-recall alone is not a reliable quality signal here; the overlay + panel are the binding evidence."
},
"verdict": "BORDERLINE PASS, NOT a clean win. Candidate median 6.5 vs incumbent median 6.0 -- a +0.5 delta over the existing adopted plate, sitting exactly AT (not clearly above) the >6.5 success threshold, and well short of the 7.0+ target. No invented architecture, no characters, no melted/broken geometry -- the LoRA generalizes CLEANLY to camp, in sharp contrast to forest_road's severe regression. One consistent, real defect: 3/5 scorers independently flagged a repeated concentric-ring ground-texture artifact around the campfire (likely a training-data or diffusion-radial-falloff tell) and 2/5 flagged hanging-satchel attachment ambiguity. Given the borderline number and this recurring artifact, this is NOT a strong enough result to recommend replacing the adopted camp_clearing_night_v2 canonical_plate on its own -- it would need at least one more iteration (prompt-level fix for the ring artifact) before being a genuine promote.py-grade candidate."

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Do not record a strict-threshold miss as a pass.

The verdict says BORDERLINE PASS while also stating that 6.5 is exactly at, not above, the >6.5 success threshold. The score ledger uses that same strict success bar, so this artifact currently records a passing outcome for a numeric failure and could mislead promotion logic. Emit a failing/inconclusive verdict or change the threshold contract explicitly.

Confidence: 99%.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@qa/evidence/plate-sprint/outdoor-lora/camp/panel_verdict.json` at line 27,
Change the verdict in the recorded evidence from “BORDERLINE PASS” to a failing
or inconclusive classification because the candidate median of 6.5 does not
satisfy the strict >6.5 success threshold. Update the surrounding wording in the
verdict field to consistently describe the result as below threshold and not
promotable, unless the scoring contract is explicitly changed everywhere to
allow 6.5.

}
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
14 changes: 14 additions & 0 deletions qa/evidence/plate-sprint/outdoor-lora/forest_road/gemini_meta.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
{
"image_ref": "asset_rbKpGkes716PxfvTa9pmfee3",
"requested_reference_images": true,
"reference_asset_id": "asset_QwasHA86jHyJmRib1q5s3bUK",
"job_id": "job_mLNVWjrNLCRDfyRHwEhUUyJV",
"assets": [
{
"asset_id": "asset_1n3pD4HasyJ6Lwx1TyS9WYYd",
"path": "/Users/lume/worldos-session-notes/plate-sprint/forest_road/outdoor_lora/fr_outdoorlora_registered.png",
"bytes": 5753219
Comment on lines +9 to +10

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
set -euo pipefail

pwd
printf '\n--- file ---\n'
nl -ba qa/evidence/plate-sprint/outdoor-lora/forest_road/gemini_meta.json | sed -n '1,40p'

printf '\n--- surrounding artifacts with path fields ---\n'
rg -n '"path"\s*:\s*"/Users/lume/worldos-session-notes|"/Users/lume/WorldOS|/Volumes/LEXAR/Codex' qa/evidence -g '*.json' -g '*.md' -g '*.txt' | sed -n '1,120p'

Repository: electricsheephq/WorldOS

Length of output: 234


🏁 Script executed:

#!/bin/bash
set -euo pipefail

printf 'PWD: '; pwd

printf '\n--- gemini_meta.json (first 40 lines) ---\n'
sed -n '1,40p' qa/evidence/plate-sprint/outdoor-lora/forest_road/gemini_meta.json

printf '\n--- matching path references in qa/evidence ---\n'
rg -n '"/Users/lume/worldos-session-notes|"/Users/lume/WorldOS|/Volumes/LEXAR/Codex' qa/evidence -g '*.json' -g '*.md' -g '*.txt' || true

printf '\n--- file type / size ---\n'
stat -c '%n %s bytes' qa/evidence/plate-sprint/outdoor-lora/forest_road/gemini_meta.json

Repository: electricsheephq/WorldOS

Length of output: 7265


Replace the local absolute asset path

qa/evidence/plate-sprint/outdoor-lora/forest_road/gemini_meta.json still hardcodes /Users/lume/worldos-session-notes/..., so replay only works on one workstation. Use a repo-relative path, /Users/lume/WorldOS/..., or an approved artifact root like /Volumes/LEXAR/Codex so the metadata stays portable. Confidence: 93%.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@qa/evidence/plate-sprint/outdoor-lora/forest_road/gemini_meta.json` around
lines 9 - 10, The metadata contains a workstation-specific absolute asset path
in the gemini_meta.json path field. Replace it with a portable repo-relative
path or an approved shared artifact root, preserving the referenced asset
location so replay works across environments.

Source: Coding guidelines

}
],
"prompt": "Repaint this image in the painterly hand-painted CRPG environment-background style of the reference image: visible oil-brush strokes, muted painterly palette, warm-amber dusk key light raking down the dirt road falling off into cool blue-violet shadow toward the dense forest flanks. The blocky grey placeholder trees and boulders must be repainted as natural, organic tree trunks with bark texture and foliage, and rounded moss-covered boulders -- never flat-faced geometric boxes or glassy/translucent prisms.\n\nCRITICAL STRUCTURE-LOCK (non-negotiable, checked pixel-by-pixel against the original): every wall, pillar, archway, doorway, staircase, tree, boulder, road edge, and prop must stay in EXACTLY its current position, size, and shape -- do NOT move, add, remove, resize, re-project, or recompose ANYTHING; only the paint and lighting treatment changes. Do NOT introduce any repeated, cloned, or tiled decorative motifs. Do NOT add any text, letters, numbers, labels, legends, map insets, UI elements, frames, or borders. Do NOT add any characters, people, figures, or creatures.\n\nThis is a TOP-DOWN DIMETRIC (2:1 isometric) game BACKGROUND PLATE viewed at a fixed ~30-degree downward camera pitch -- you are looking DOWN onto the scene from high above and to the side, exactly like a Pillars of Eternity II pre-rendered isometric level. Keep that EXACT high dimetric bird's-eye angle: there is NO horizon and NO sky band across the top of the frame -- the top of the frame is MORE ground/canopy seen from above. Do NOT tilt the camera to an eye-level, first-person, or landscape-horizon view."

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Relax the exact-shape lock for organic placeholders.

The prompt simultaneously requires every tree and boulder to retain its exact pixel shape and requires those same blocky volumes to become organic trunks and rounded rocks. The forest result recorded in qa/scores_ledger.md Line 13 is the direct failure mode: all five scorers saw wireframe-box/open-cubicle architecture.

Lock camera, road topology, boundaries, and object anchors, but allow silhouette and volume reinterpretation for placeholder trees and boulders before rerunning the forest-road smoke test.

Confidence: 97%.

Proposed prompt correction
- every ... tree, boulder ... must stay in EXACTLY its current position, size, and shape
+ keep the camera, road boundaries, topology, and object anchor positions fixed;
+ placeholder trees and boulders may change silhouette, volume, occlusion, and materials
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
"prompt": "Repaint this image in the painterly hand-painted CRPG environment-background style of the reference image: visible oil-brush strokes, muted painterly palette, warm-amber dusk key light raking down the dirt road falling off into cool blue-violet shadow toward the dense forest flanks. The blocky grey placeholder trees and boulders must be repainted as natural, organic tree trunks with bark texture and foliage, and rounded moss-covered boulders -- never flat-faced geometric boxes or glassy/translucent prisms.\n\nCRITICAL STRUCTURE-LOCK (non-negotiable, checked pixel-by-pixel against the original): every wall, pillar, archway, doorway, staircase, tree, boulder, road edge, and prop must stay in EXACTLY its current position, size, and shape -- do NOT move, add, remove, resize, re-project, or recompose ANYTHING; only the paint and lighting treatment changes. Do NOT introduce any repeated, cloned, or tiled decorative motifs. Do NOT add any text, letters, numbers, labels, legends, map insets, UI elements, frames, or borders. Do NOT add any characters, people, figures, or creatures.\n\nThis is a TOP-DOWN DIMETRIC (2:1 isometric) game BACKGROUND PLATE viewed at a fixed ~30-degree downward camera pitch -- you are looking DOWN onto the scene from high above and to the side, exactly like a Pillars of Eternity II pre-rendered isometric level. Keep that EXACT high dimetric bird's-eye angle: there is NO horizon and NO sky band across the top of the frame -- the top of the frame is MORE ground/canopy seen from above. Do NOT tilt the camera to an eye-level, first-person, or landscape-horizon view."
"prompt": "Repaint this image in the painterly hand-painted CRPG environment-background style of the reference image: visible oil-brush strokes, muted painterly palette, warm-amber dusk key light raking down the dirt road falling off into cool blue-violet shadow toward the dense forest flanks. The blocky grey placeholder trees and boulders must be repainted as natural, organic tree trunks with bark texture and foliage, and rounded moss-covered boulders -- never flat-faced geometric boxes or glassy/translucent prisms.\n\nCRITICAL STRUCTURE-LOCK (non-negotiable, checked pixel-by-pixel against the original): keep the camera, road boundaries, topology, and object anchor positions fixed; placeholder trees and boulders may change silhouette, volume, occlusion, and materials -- do NOT move, add, remove, resize, re-project, or recompose ANYTHING; only the paint and lighting treatment changes. Do NOT introduce any repeated, cloned, or tiled decorative motifs. Do NOT add any text, letters, numbers, labels, legends, map insets, UI elements, frames, or borders. Do NOT add any characters, people, figures, or creatures.\n\nThis is a TOP-DOWN DIMETRIC (2:1 isometric) game BACKGROUND PLATE viewed at a fixed ~30-degree downward camera pitch -- you are looking DOWN onto the scene from high above and to the side, exactly like a Pillars of Eternity II pre-rendered isometric level. Keep that EXACT high dimetric bird's-eye angle: there is NO horizon and NO sky band across the top of the frame -- the top of the frame is MORE ground/canopy seen from above. Do NOT tilt the camera to an eye-level, first-person, or landscape-horizon view."
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@qa/evidence/plate-sprint/outdoor-lora/forest_road/gemini_meta.json` at line
13, Revise the prompt’s CRITICAL STRUCTURE-LOCK to preserve the camera, road
topology, boundaries, and object anchors while explicitly allowing placeholder
trees and boulders to be reshaped into organic trunks, foliage, and rounded
rocks. Remove the requirement that those organic placeholders retain their exact
pixel silhouettes, sizes, and volumes, then rerun the forest-road smoke test.

}
Loading
Sorry, something went wrong. Reload?
Sorry, we cannot display this file.
Sorry, this file is invalid so it cannot be displayed.
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
{
"img_1.jpg": "ANCHOR_CAMP_V2",
"img_2.jpg": "CONTROL",
"img_3.jpg": "CANDIDATE_FR_OUTDOORLORA"
Comment on lines +2 to +4

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Use the verifier’s A/B/C slot contract here as well.

This mapping has the same integration defect as the camp mapping: it omits slot_to_label/label_to_slot and uses ad-hoc labels such as ANCHOR_CAMP_V2 and CANDIDATE_FR_OUTDOORLORA. qa/plate_loop.py at Line 465-497 cannot relabel these slots or compute the candidate/control delta.

Normalize this file and its forest verdict artifact to the exact harness schema and slot names.

Confidence: 99%.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@qa/evidence/plate-sprint/outdoor-lora/forest_road/panel_mapping.json` around
lines 2 - 4, The forest-road panel mapping uses ad-hoc labels and lacks the
verifier’s A/B/C slot contract. Update the mapping and corresponding forest
verdict artifact to the exact harness schema, including slot_to_label and
label_to_slot with the expected A, B, and C slot names, so qa/plate_loop.py can
relabel slots and compute the candidate/control delta.

}
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
{
"run": "forest_road-outdoor-lora-smoke-2026-07-10",
"arm": "OUTDOOR-LORA smoke test (#1481): mint anchor via schema-correct model_run (model_bfl-flux-1-dev + depth ControlNet str0.65/end0.6 + model_RsWEcQL2NWXwoyEodWVE2vWG, the new 18-image exterior-only LoRA, lorasScale 0.7, seed 7) -> Gemini instruction-edit re-registration (model_google-gemini-3-1-flash, structure-lock + dimetric-lock prompt, referenceImages=[minted anchor]) over forest_road_greybox.png.",
"panel_mapping": {"img_1.jpg": "ANCHOR_CAMP_V2", "img_2.jpg": "CONTROL", "img_3.jpg": "CANDIDATE_FR_OUTDOORLORA"},
"scores": {
"img_1_ANCHOR_CAMP_V2": [7, 7, 7, 8, 7],
"img_2_CONTROL_poe2": [9, 9, 9, 9, 9],
"img_3_CANDIDATE": [2, 2, 4, 3, 2]
},
"medians": {"ANCHOR_CAMP_V2": 7.0, "CONTROL": 9.0, "CANDIDATE": 2.0},
"control_valid": false,
"control_integrity_note": "Control (poe2_control.jpg) scored median 9.0 with human figures visible and flagged by all 5 scorers -- consistent with the CB-FOREST precedent's finding that this control is imperfect for genre-fit judgment, but scorers here were asked to judge craft only, and did.",
"candidate_defect_notes": [
"img_3: 2 | broken/nonsensical geometry -- three structures render as transparent wireframe frames with wood-carved panels floating disconnected from walls/roofs, no coherent solid form, reads as melted/AI-artifacted architecture",
"img_3: 2 | broken/melted geometry -- the two forest structures read as disconnected wireframe boxes with doors/curtains floating in impossible non-volumetric frames, repeated carved-door and curtain motifs tiled across separate huts",
"img_3: 4 | wooden hut modules read as flat disconnected box-frames with unnatural thin wireframe-like edge seams, nonsensical open-cubicle architecture with no coherent walls/roofline",
"img_3: 3 | nonsensical wireframe/open-frame box structures with disconnected beams and textures applied to only some faces -- reads as broken/melted geometry",
"img_3: 2 | broken/melted geometry -- 'buildings' are disconnected wireframe-edge fragments with walls that don't close or support each other, nonsensical structure placement in the forest, incoherent construction logic"
],
"advisory_metrics": {
"ncc_candidate_vs_base": 0.6974,
"ncc_candidate_vs_greybox": -0.0800,
"edge_recall_vs_greybox": 0.9678,
"edge_recall_caveat": "MISLEADINGLY HIGH -- content-blind per #1491. The candidate preserved the greybox's blocky rectangular edges too literally, reading them as walls/doors instead of trees, so edges 'recall' well while content is badly wrong. The overlay image (overlay_fr_outdoorlora.png) is the binding evidence, not this number."
},
"verdict": "NOT CONVERGED -- SEVERE REGRESSION. Candidate median 2.0 is far below both the 6.0 baseline cap and the >6.5 success bar, and is WORSE than ARM C's own (interior LoRA) forest_road attempt (CB-FOREST precedent, median 6.0 tie, qa/evidence/... /cb_forest/panel_verdict.json -- kept in session-notes, not yet migrated to this evidence tree). All 5 scorers independently and unprompted flagged broken/wireframe/melted architecture -- the SAME invented-architecture failure mode the OUTDOOR-LORA effort was built to fix, still present despite training on 18 char-free, quality-passed, exterior-only images (with fortress/temple content deliberately trimmed in the pre-train quality pass). ROOT CAUSE HYPOTHESIS: forest_road's greybox represents its dense tree line as adjacent rectangular box volumes (see forest_road_greybox.png / forest_road_greybox_depth.png) -- a coarser, more literally 'architectural-looking' placeholder geometry than camp's greybox. The model appears to read blocky rectangular volumes as buildings/structures regardless of whether its training data was interior or exterior -- a GEOMETRY-SHAPE bias, not purely a training-domain bias. This is a genuinely new finding beyond the original hypothesis (that exterior-only training data would fix outdoor generalization) and suggests forest_road-class rooms may need a differently-shaped greybox (more organic/clustered tree footprints, not axis-aligned boxes) rather than a different LoRA, if this failure mode is to be fixed."
}
Loading
Loading