Skip to content

Commit 317f97b

Browse files
docs: add evidence-bound architecture walkthrough
1 parent 27be5c1 commit 317f97b

1 file changed

Lines changed: 91 additions & 0 deletions

File tree

‎docs/architecture-walkthrough.md‎

Lines changed: 91 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,91 @@
1+
# Architecture Walkthrough: Two-Stage Serving
2+
3+
This is an 8-minute walkthrough script and diagram set for the current implementation. It describes
4+
the code in this repository; it does not claim that an ANN index, external feature store, or runtime
5+
reasoning service exists.
6+
7+
## 0:00–1:00 — Problem and decision
8+
9+
The original serving path applied the sequence model's dense output projection to every catalogue
10+
item. That is simple and correct for the current small bundle, but its work grows with catalogue
11+
size and provides no replaceable candidate boundary. The two-stage design separates *recall*
12+
from *precision*: retrieval produces a bounded set of plausible item IDs, then ranking spends the
13+
sequence model's richer scoring capacity only on that set.
14+
15+
```mermaid
16+
flowchart LR
17+
H["Known interaction history"] --> Q["Embedding query"]
18+
Q --> R["Stage 1: exact embedding retrieval"]
19+
R --> C["Bounded candidate IDs"]
20+
C --> K["Stage 2: BiLSTM + attention ranking"]
21+
H --> K
22+
K --> O["Top-K recommendations"]
23+
```
24+
25+
## 1:00–3:00 — Retrieval is not ranking
26+
27+
`app/core/retrieval.py` implements `ExactEmbeddingRetriever`. It averages embeddings of known
28+
history items, normalizes that query, compares it with normalized catalogue embeddings, excludes
29+
items already seen, and returns a bounded candidate list. Its default is an exact in-memory vector
30+
scan. That makes behavior reproducible and keeps model bundles self-contained, but it is **not**
31+
FAISS, HNSW, or another approximate-nearest-neighbor index.
32+
33+
Retrieval optimizes candidate recall and cost. It can return items that are semantically or
34+
behaviorally near the history, but it has no access to the full sequence-ordering signal used by
35+
the ranker. The stable `RetrievalResult` contract is the seam where an ANN implementation can
36+
later be added after index lifecycle, recall, freshness, and operational evidence are available.
37+
38+
## 3:00–5:00 — Candidate-only ranking
39+
40+
`DeepSequenceModel.rank_candidates` runs the existing padding-aware bidirectional LSTM and
41+
attention model, then gathers scores only for the retrieved IDs. This preserves the existing model
42+
and training compatibility while making the ranking stage explicit. The API uses retrieval after
43+
admission control and before decoding recommendations; cache, authentication, rate limits, and
44+
fallback behavior remain unchanged.
45+
46+
```mermaid
47+
sequenceDiagram
48+
participant Client
49+
participant API as FastAPI route
50+
participant Retriever as ExactEmbeddingRetriever
51+
participant Ranker as DeepSequenceModel
52+
53+
Client->>API: history + top_k
54+
API->>API: validate, authorize, rate-limit, cache check
55+
API->>Retriever: retrieve(history, exclusions, top_k)
56+
Retriever-->>API: candidate_ids
57+
API->>Ranker: rank_candidates(history, candidate_ids, top_k)
58+
Ranker-->>API: ordered item IDs
59+
API-->>Client: recommendations + model version + latency
60+
```
61+
62+
The tradeoff is intentional: the first stage is logically separated but still scans all embeddings,
63+
so it does not yet deliver the latency or memory profile of a production ANN system. Candidate-pool
64+
size is configurable with `RETRIEVAL_CANDIDATE_POOL_SIZE`; it should be measured against
65+
Recall@K and latency before changing it in production.
66+
67+
## 5:00–6:30 — Reasoning and explanations
68+
69+
No runtime `reasoning/` package or per-recommendation explanation endpoint exists in the current
70+
repository. That is deliberate in this walkthrough: an LSTM score is not a causal explanation, and
71+
the service should not invent a user-facing reason from hidden states. Issue #16 records the
72+
separate evidence-backed explanation contract, including privacy review and insufficient-history
73+
handling. Until that work is implemented and tested, the only trustworthy response-level evidence
74+
is model version, fallback state, cache state, and bounded request latency.
75+
76+
## 6:30–8:00 — Engineering tradeoffs and next decisions
77+
78+
- **Exact retrieval now:** easy to test and bundle; unsuitable for large catalogues without an ANN
79+
index and index-refresh lifecycle.
80+
- **Sequence ranker retained:** preserves current training artifacts; a future ranker change needs
81+
temporal evaluation against the popularity baseline.
82+
- **No fabricated confidence:** offline ranking quality and per-recommendation confidence are
83+
different measurements.
84+
- **No hidden reasoning:** user-facing explanations must be constrained to permitted evidence,
85+
not chain-of-thought or unvalidated causal language.
86+
- **Safety preserved:** existing authentication, credential-derived rate limiting, admission
87+
control, cache keying, fallback, and model-bundle checks remain the API boundary.
88+
89+
Before a large-catalogue deployment, add an evaluated ANN backend, version and validate its index
90+
with the model bundle, measure candidate recall and end-to-end p95/p99 latency, and keep the
91+
candidate-ranker contract stable during rollout.

0 commit comments

Comments
 (0)