Skip to content

Commit 548b95c

Browse files
Merge pull request #15 from RedEyeNinja-BKK/fix/provenance-qualification-scope
fix: preserve provenance and qualification scope
2 parents 3a4d18e + aaf13cf commit 548b95c

7 files changed

Lines changed: 368 additions & 37 deletions

File tree

Lines changed: 254 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,254 @@
1+
# Case Study — Etsy Package Realignment Through Process Engine (2026-08-17)
2+
3+
> **What this is:** a recorded downstream use of Process Engine to realign an
4+
> existing Turnstone-native Etsy package with new source material. It is a
5+
> usage, generation, review, trial, and later deployment/read-back evidence
6+
> record—not a claim that Process Engine provides Etsy API access or that the
7+
> downstream package is fully runtime-qualified.
8+
>
9+
> **Privacy:** the shop identity and source URLs are retained only as the
10+
> operator's supplied context. Buyer data, credentials, private worksheet
11+
> contents, host data, and internal identifiers are not published here.
12+
13+
---
14+
15+
## Executive result
16+
17+
The operator asked Process Engine to update and refine an existing Etsy skills
18+
family and persona so it could assist with online-business management and Etsy
19+
SEO. The engine collected and synthesized the supplied material, revised the
20+
package, and was later recorded as having a bounded native Turnstone object
21+
deployment. The evidence does not prove that deployment satisfied the Ship gate.
22+
23+
The resulting package is useful for **analysis, drafting, translation,
24+
worksheet-export analysis, prioritization, trend frameworks, and advisory
25+
policy heads-ups**. The following capabilities were deliberately not claimed:
26+
27+
- authenticated Etsy Seller App/API access;
28+
- live shop, listing, order, or storefront reads;
29+
- live seller-taxonomy lookup;
30+
- live Google Sheets reads or writes;
31+
- publishing listings or changing shop settings;
32+
- sending customer messages;
33+
- purchases, inventory commitments, or any Etsy mutation.
34+
35+
The correct overall runtime-dependent Trial classification remains
36+
**INCOMPLETE**, not FAIL: the executed cases passed, but unavailable required
37+
capability categories were not proven.
38+
39+
## Operator intent and material collected
40+
41+
The interaction supplied the existing Etsy package plus the following material
42+
and context classes:
43+
44+
1. an online-business-manager skills article;
45+
2. the operator's ChummyThailand storefront URL as target context;
46+
3. an article describing an autonomous Etsy operating model;
47+
4. the shared `Online_Business_Worksheet_2026_Master-V3.0` as business-analysis
48+
context;
49+
5. Etsy developer documentation;
50+
6. Etsy Dev MCP documentation;
51+
7. the `getSellerTaxonomyNodes` reference.
52+
53+
The engine treated these differently rather than copying them wholesale:
54+
available source material supplied techniques and requirements, the operator's
55+
storefront URL identified the intended shop/target context but did not thereby
56+
prove that storefront page content was retrieved, the worksheet supplied
57+
point-in-time analytical inputs, and the developer references informed
58+
capability-path design. Unavailable material remained unavailable, and partial
59+
material contributed only its usable portion. The Etsy Dev MCP path was recorded
60+
as **documentation/specification discovery**, not Etsy API execution.
61+
62+
The objective was refined into an assistant-manager package that could help
63+
with customer-message drafting, listing and SEO analysis, business and
64+
financial analysis, trend interpretation, weekly prioritization, and
65+
operator-controlled next steps. The target was retained as an aspirational
66+
north star rather than a guarantee.
67+
68+
## Process Engine pipeline evidence
69+
70+
The recorded path was:
71+
72+
**Intent → Orient → Collect → Clarify → Objective → Summary Gate → Pattern →
73+
Review → Trial → Ship**
74+
75+
A companion assessment was used during development but is excluded from the
76+
proposed canonical commit; its generalized rationale and evaluation matrix
77+
belong in the eventual issue/PR record rather than a permanent evaluation
78+
subtree.
79+
80+
| Stage | Evidence and result |
81+
|---|---|
82+
| Collect | Existing package and seven new source/context inputs were collected and classified. |
83+
| Clarify / Objective | The desired day-to-day behavior was clarified as an OBM + Etsy SEO assistant, not an autonomous store operator. |
84+
| Summary Gate | Material basis, intent, and intended package direction were confirmed before authoring. |
85+
| Pattern | Existing persona and skills were realigned; new supporting skills and governance posture were authored. |
86+
| Review | Initial draft returned **REVISE** with seven findings. The revised draft returned **PASS** after identity, completeness, coexistence, capability-path, financial-source, privacy/commitment, and trend-evidence corrections. |
87+
| Trial | Executed behavioral and activation cases passed; runtime-dependent categories remained unproven. Overall Trial: **INCOMPLETE**. |
88+
| Ship | A later record describes a narrowed analysis/drafting deployment and native read-back, but the inspected evidence does not prove a new exact-scope Review PASS + Trial PASS chain after the overall-INCOMPLETE result; classify this as a pipeline-contract failure, not a successful narrow Ship. |
89+
90+
### Trial → Ship chronology
91+
92+
The original Trial record explicitly classified the overall result as
93+
**INCOMPLETE** and stated that INCOMPLETE does not advance to Ship. A later
94+
deployment record described a narrowed analysis/drafting scope as having a
95+
separate Trial PASS claim and recorded native deployment/read-back, but the
96+
evidence inspected here contains no separate Pattern revision, exact-scope
97+
Review PASS, exact-scope Trial PASS receipt, or operator decision establishing
98+
a new PASS chain for that narrowed artifact.
99+
100+
Accordingly, the deployment read-back proves object persistence, not Ship-gate
101+
compliance. This case study records the outcome as a **Process Engine
102+
pipeline-contract failure** unless a primary receipt for the missing narrowed
103+
Review/Trial chain is later produced. It must not be rewritten as a successful
104+
narrow Ship.
105+
106+
The initial Review did not silently pass the draft. It identified seven concrete
107+
findings, including unproven destination-skill absence, incomplete persona and
108+
governance artifacts, ambiguous treatment of existing skills, abstract
109+
capability declarations, missing financial-source precedence, insufficient
110+
customer-message safeguards, and an under-specified trend-evidence standard.
111+
The revised package resolved those findings and was re-reviewed PASS.
112+
113+
## Downstream package trial evidence
114+
115+
The recorded trial evidence exercised **12 behavioral cases**, all passing. It
116+
was later described as supporting the narrowed scope, but the evidence
117+
inspected here does not prove that these cases were rerun against a separately
118+
revised exact narrowed artifact before Ship.
119+
120+
- customer-message uncertainty and missing-order-fact handling;
121+
- SEO boundary behavior without ranking guarantees or unauthenticated API use;
122+
- financial arithmetic and source/freshness labeling;
123+
- trend escalation into a bounded, testable recommendation;
124+
- policy and external-action boundaries;
125+
- wrong shop/listing identity;
126+
- Etsy capability escalation;
127+
- financial-source conflict;
128+
- buyer-data minimization;
129+
- commitment safeguards;
130+
- external-emission safeguards; and
131+
- adversarial medical-claim handling.
132+
133+
Activation cases also passed for the OBM assistant, customer-message skill,
134+
financial-analysis skill, and the combined SEO/taxonomy and trend-analysis
135+
path. The recorded activation observations were:
136+
137+
- OBM assistant: should-trigger 2/2; shouldn't-trigger 3/3;
138+
- customer messages: should-trigger 3/3; shouldn't-trigger 2/2;
139+
- financial analysis: should-trigger 3/3; shouldn't-trigger 2/2;
140+
- SEO/taxonomy and trend analysis combined: should-trigger 4/4;
141+
shouldn't-trigger 1/1.
142+
143+
One buyer-privacy case passed with an evidence-path caveat: the safeguard
144+
behavior was supported, but a delegated context did not consistently resolve
145+
the direct reference path. This is retained rather than erased.
146+
147+
## Runtime and deployment evidence
148+
149+
The later native Turnstone deployment record reports read-back of the
150+
following objects:
151+
152+
- the existing Etsy manager persona was updated in place;
153+
- six supporting skills were created and read back;
154+
- each created skill was enabled and scanner-rated safe;
155+
- a content-only prompt policy was created and read back;
156+
- the advisory judge helper remained **not deployed** because its native
157+
creation path was not exposed in the execution context;
158+
- existing Etsy skills were retained and not silently overwritten, disabled,
159+
or superseded;
160+
- a reversible rollback path was recorded and not executed.
161+
162+
This proves native object persistence and read-back for the objects named in
163+
that deployment record, not Trial/Ship gate compliance and not the unavailable
164+
Etsy integrations listed above.
165+
166+
## Reference-retrieval limitation
167+
168+
The re-review explicitly recorded a reference-retrieval limitation. The
169+
following operator-supplied or identified references were **not retrievable
170+
through the available fetch/preview path** during the run:
171+
172+
- the ChummyThailand storefront page;
173+
- the Medium autonomous-Etsy article;
174+
- the `getSellerTaxonomyNodes` reference.
175+
176+
The package correctly did **not** assert these as proven evidence, and trial
177+
fixtures must not represent them as verified sources without a later authorized
178+
retrieval.
179+
180+
The root cause of each failure is **not proven**. Each could have been any of:
181+
182+
- unsupported page/URL behavior;
183+
- authentication or permission scope;
184+
- a transient network failure;
185+
- a tool/adapter limitation;
186+
- a malformed or redirected URL;
187+
- source-side blocking;
188+
- a retrieval timeout; or
189+
- incomplete page extraction.
190+
191+
The strongest development signal from the run is that Process Engine should
192+
keep a compact conversational per-material receipt/evidence note: distinguish
193+
supplied, retrieved, partial, unavailable, and intentionally-not-attempted
194+
states as applicable; preserve what usable content actually informed generation;
195+
and carry source-dependent limitations forward. Record a retrieval cause only
196+
when supported; otherwise keep the cause unknown. Do not infer inaccessible
197+
source contents. This is the implemented evidence-handling principle, not a new
198+
ledger, manifest, schema, persisted receipt object, or failure taxonomy.
199+
200+
## Usage and performance receipt
201+
202+
The Process Engine workstream's captured session evidence records:
203+
204+
| Metric | Recorded value |
205+
|---|---:|
206+
| Workstream messages | 277 |
207+
| User messages | 12 |
208+
| System messages | 27 |
209+
| Assistant messages | 85 |
210+
| Tool messages | 153 |
211+
| Natural requests in the measured routing window | 250 |
212+
| Successful HTTP responses | 250/250 (200) |
213+
| Efficient-route requests | 185 (74.0%) |
214+
| Capable-route requests | 65 (26.0%) |
215+
| Failures / fallbacks / route unavailable | 0 / 0 / 0 |
216+
| `reasoning_content` recurrence | 0 |
217+
| Context-window errors / truncation | 0 |
218+
| Model-call count | **UNAVAILABLE** in the authoritative receipt |
219+
| Input/output/cache/reasoning tokens | **UNAVAILABLE** in the authoritative receipt |
220+
| Authoritative total token usage | **UNAVAILABLE** |
221+
| Defensible API-equivalent cost | **UNAVAILABLE** |
222+
223+
The routing counts are usage telemetry, not a quality score. No token or cost
224+
estimate is derived from message length, and bundled/local model usage is not
225+
represented as zero cost.
226+
227+
## Evidence classification and next trial
228+
229+
| Claim | State |
230+
|---|---|
231+
| Process Engine collected and classified the supplied material | **PROVEN** by the recorded interaction |
232+
| Revised package addressed the seven Review findings | **PROVEN** by the re-review record |
233+
| Executed package behavior passed 12/12 cases | **PROVEN** by the trial record |
234+
| Activation cases recorded in the trial passed | **PROVEN** for the executed set |
235+
| Authenticated Etsy API path | **UNPROVEN / excluded** |
236+
| Live Google Sheets path | **UNPROVEN / excluded** |
237+
| Storefront and live taxonomy execution | **UNPROVEN / excluded** |
238+
| Full end-to-end Etsy runtime qualification | **INCOMPLETE** |
239+
| Narrowed native object persistence/read-back | **PROVEN** for the named objects; Ship-gate compliance for the narrowed deployment: **NOT PROVEN** |
240+
| Full Etsy package ship | **NOT PROVEN** |
241+
242+
A future completion trial would require an operator-controlled Etsy Seller App
243+
read path, an authorized Google Sheets read path, exact taxonomy-operation
244+
receipts, and re-execution of the affected cases through those intended paths.
245+
Until then, the package should be described as a narrowed analysis/drafting
246+
assistant, not an autonomous Etsy manager.
247+
248+
## Source and evidence boundary
249+
250+
This case study is derived from the recorded Etsy Skills Refinement
251+
interaction and its associated local review, trial, and ship evidence. It
252+
records observed outcomes and known gaps; it does not reproduce private
253+
transcripts or source documents. The Process Engine repository remains the
254+
canonical product source. Live Turnstone objects are deployment evidence only.

‎references/intake.md‎

Lines changed: 28 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -65,6 +65,17 @@ clarification — the material sharpens the questions. Run conversationally:
6565
doc), fetch the page and extract what's useful. Record the exact source
6666
URL.
6767

68+
For every item, distinguish: supplied; retrieval attempted or intentionally
69+
not attempted; usable content retrieved; partial extraction; or retrieval
70+
unavailable. Supplied is not thereby retrieved, and retrieved is not thereby
71+
completely extracted. Record retrieval path/date when known. Record a
72+
failure cause only when supported; otherwise say failure cause unknown.
73+
Retrieval failure does not prove source content false. When a materially
74+
useful source cannot be retrieved, offer one practical substitute such as
75+
paste/upload once; if declined, do not ask again. Continue from available
76+
evidence and carry the limitation forward. Inaccessible material cannot
77+
silently become an inferred or verified source fact.
78+
6879
**Text dumps / files (no link)** — no fetch. Provenance is whatever the
6980
user provides or states; if none is given, record "user-provided, source
7081
unknown". Never assume a source for pasted content.
@@ -86,9 +97,12 @@ clarification — the material sharpens the questions. Run conversationally:
8697
- general input (notes, constraints, ideas) → incorporate into the
8798
intent understanding
8899
4. **Extract** — pull the techniques, constraints, domain specifics, and
89-
intent from ALL provided material. Multiple sources combine into a
90-
best-of-all-worlds understanding. That synthesis never drops what the
91-
package will later need to tell entities apart:
100+
intent from ALL material whose content is actually available. For supplied
101+
but unavailable or partially extracted material, preserve the limitation and
102+
use only claims supported by the usable portion or another named source.
103+
Multiple sources combine into a best-of-all-worlds understanding, but
104+
unavailable material cannot silently be replaced by inference. The synthesis
105+
never drops what the package will later need to tell entities apart:
92106
- preserve identity-critical relationships (entity ↔ role ↔
93107
identifier/alias — machines, accounts, people, products, records,
94108
services) when they are material to later decisions or tool targets;
@@ -123,9 +137,13 @@ clarification — the material sharpens the questions. Run conversationally:
123137
9. **Sweep** — the engine's own output must stay sweep-clean (the zero-
124138
tolerance language rule applies to what the engine WRITES, not to what a
125139
user pastes; we cannot control user input).
126-
10. **Provenance record** — source URL(s) or "user-provided" status, item
127-
type, license info if visible, fetched date, extraction summary. This
128-
record is evidence.
140+
10. **Material receipt** — for each item, record a compact conversational
141+
receipt: stable identity, material class, supplied/retrieval status,
142+
retrieval path/date when known, usable or partial extraction, extracted and
143+
used material, excluded/not-used material, and downstream claims or
144+
requirements left unproven. Record known failure cause only when supported;
145+
otherwise record unknown. This is an evidence note in the existing intake
146+
flow, not a new persisted subsystem or mandatory manifest.
129147
11. **Operator gate** — the generated package then passes through the normal
130148
Pattern → Review → Trial → Ship gates; the operator's sign-off is the
131149
gate (the engine's core identity, unchanged).
@@ -166,6 +184,10 @@ clarification — the material sharpens the questions. Run conversationally:
166184

167185
## Verification
168186

187+
- [ ] Material receipt distinguishes supplied from retrieved, complete from partial, and unavailable from not attempted
188+
- [ ] Known failure causes are supported; unknown causes remain unknown
189+
- [ ] One practical substitute was offered for materially useful unavailable material; a declined offer was not repeated
190+
- [ ] Source-dependent claims left unproven are carried forward
169191
- [ ] Item type + source recorded ("user-provided, source unknown" if no source)
170192
- [ ] Techniques / domain specifics extracted from all provided material
171193
- [ ] Original instructions authored — no verbatim copying

‎skills/process-engine-core/SKILL.md‎

Lines changed: 13 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -57,8 +57,11 @@ and governance surfaces.
5757
note the objective as underspecified and propose one in the draft for
5858
correction at review — never interrogate.
5959
5. **Summary gate** — before generation, present: "Working from: N links +
60-
M text blocks (k sources unknown). Intent: X. Good looks like: <vision>.
61-
Generate?" — proceed only on confirmation.
60+
M text blocks (k sources unknown)." Also disclose any material limitations
61+
that affect generation in plain language: supplied-but-unretrieved,
62+
partially extracted, or source-dependent facts still unproven. Keep it
63+
conversational, not a provenance dump. State intent and good looks like,
64+
then ask "Generate?" — proceed only on confirmation.
6265
6. **Route** — classify the request:
6366
- author an artifact → `process-engine-pattern-author`
6467
(eligibility gate decides shape: project / persona / skill(s) /
@@ -103,12 +106,14 @@ and governance surfaces.
103106
7. **Load standards checklist** (references/standards.md) and the generation
104107
basis (references/best-practices.md — full Osmani catalog index) and apply
105108
them to every step.
106-
8. **Gate** — nothing proceeds past authoring without a review step.
107-
The Process Engine prompts carry the workflow discipline; Turnstone's
108-
native prompt policy provides durable contextual guidance and the
109-
advisory judge provides review/trial evidence. Neither silently replaces
110-
operator approval. The model does not need to recite governance policy —
111-
Turnstone supplies the native mechanisms around it.
109+
8. **Qualification continuity** — Review and Trial evidence applies only to the
110+
artifact and deployable scope it identifies. A change that can affect behavior,
111+
acceptance criteria, artifact membership, target/entity identity, capability or
112+
tool boundaries, safeguards/risk behavior, or deployable scope is material and
113+
returns through Pattern → Review → Trial as applicable. Typographical, formatting,
114+
or equivalent non-behavioral cleanup may retain prior evidence when its
115+
non-material judgment is recorded. Neither prior evidence nor native governance
116+
silently replaces operator approval.
112117

113118
## Examples
114119
- "I want a skill that writes release notes" → collect (any material?) →

0 commit comments

Comments
 (0)