|
| 1 | +# Architecture Walkthrough: Two-Stage Serving |
| 2 | + |
| 3 | +This is an 8-minute walkthrough script and diagram set for the current implementation. It describes |
| 4 | +the code in this repository; it does not claim that an ANN index, external feature store, or runtime |
| 5 | +reasoning service exists. |
| 6 | + |
| 7 | +## 0:00–1:00 — Problem and decision |
| 8 | + |
| 9 | +The original serving path applied the sequence model's dense output projection to every catalogue |
| 10 | +item. That is simple and correct for the current small bundle, but its work grows with catalogue |
| 11 | +size and provides no replaceable candidate boundary. The two-stage design separates *recall* |
| 12 | +from *precision*: retrieval produces a bounded set of plausible item IDs, then ranking spends the |
| 13 | +sequence model's richer scoring capacity only on that set. |
| 14 | + |
| 15 | +```mermaid |
| 16 | +flowchart LR |
| 17 | + H["Known interaction history"] --> Q["Embedding query"] |
| 18 | + Q --> R["Stage 1: exact embedding retrieval"] |
| 19 | + R --> C["Bounded candidate IDs"] |
| 20 | + C --> K["Stage 2: BiLSTM + attention ranking"] |
| 21 | + H --> K |
| 22 | + K --> O["Top-K recommendations"] |
| 23 | +``` |
| 24 | + |
| 25 | +## 1:00–3:00 — Retrieval is not ranking |
| 26 | + |
| 27 | +`app/core/retrieval.py` implements `ExactEmbeddingRetriever`. It averages embeddings of known |
| 28 | +history items, normalizes that query, compares it with normalized catalogue embeddings, excludes |
| 29 | +items already seen, and returns a bounded candidate list. Its default is an exact in-memory vector |
| 30 | +scan. That makes behavior reproducible and keeps model bundles self-contained, but it is **not** |
| 31 | +FAISS, HNSW, or another approximate-nearest-neighbor index. |
| 32 | + |
| 33 | +Retrieval optimizes candidate recall and cost. It can return items that are semantically or |
| 34 | +behaviorally near the history, but it has no access to the full sequence-ordering signal used by |
| 35 | +the ranker. The stable `RetrievalResult` contract is the seam where an ANN implementation can |
| 36 | +later be added after index lifecycle, recall, freshness, and operational evidence are available. |
| 37 | + |
| 38 | +## 3:00–5:00 — Candidate-only ranking |
| 39 | + |
| 40 | +`DeepSequenceModel.rank_candidates` runs the existing padding-aware bidirectional LSTM and |
| 41 | +attention model, then gathers scores only for the retrieved IDs. This preserves the existing model |
| 42 | +and training compatibility while making the ranking stage explicit. The API uses retrieval after |
| 43 | +admission control and before decoding recommendations; cache, authentication, rate limits, and |
| 44 | +fallback behavior remain unchanged. |
| 45 | + |
| 46 | +```mermaid |
| 47 | +sequenceDiagram |
| 48 | + participant Client |
| 49 | + participant API as FastAPI route |
| 50 | + participant Retriever as ExactEmbeddingRetriever |
| 51 | + participant Ranker as DeepSequenceModel |
| 52 | +
|
| 53 | + Client->>API: history + top_k |
| 54 | + API->>API: validate, authorize, rate-limit, cache check |
| 55 | + API->>Retriever: retrieve(history, exclusions, top_k) |
| 56 | + Retriever-->>API: candidate_ids |
| 57 | + API->>Ranker: rank_candidates(history, candidate_ids, top_k) |
| 58 | + Ranker-->>API: ordered item IDs |
| 59 | + API-->>Client: recommendations + model version + latency |
| 60 | +``` |
| 61 | + |
| 62 | +The tradeoff is intentional: the first stage is logically separated but still scans all embeddings, |
| 63 | +so it does not yet deliver the latency or memory profile of a production ANN system. Candidate-pool |
| 64 | +size is configurable with `RETRIEVAL_CANDIDATE_POOL_SIZE`; it should be measured against |
| 65 | +Recall@K and latency before changing it in production. |
| 66 | + |
| 67 | +## 5:00–6:30 — Reasoning and explanations |
| 68 | + |
| 69 | +No runtime `reasoning/` package or per-recommendation explanation endpoint exists in the current |
| 70 | +repository. That is deliberate in this walkthrough: an LSTM score is not a causal explanation, and |
| 71 | +the service should not invent a user-facing reason from hidden states. Issue #16 records the |
| 72 | +separate evidence-backed explanation contract, including privacy review and insufficient-history |
| 73 | +handling. Until that work is implemented and tested, the only trustworthy response-level evidence |
| 74 | +is model version, fallback state, cache state, and bounded request latency. |
| 75 | + |
| 76 | +## 6:30–8:00 — Engineering tradeoffs and next decisions |
| 77 | + |
| 78 | +- **Exact retrieval now:** easy to test and bundle; unsuitable for large catalogues without an ANN |
| 79 | + index and index-refresh lifecycle. |
| 80 | +- **Sequence ranker retained:** preserves current training artifacts; a future ranker change needs |
| 81 | + temporal evaluation against the popularity baseline. |
| 82 | +- **No fabricated confidence:** offline ranking quality and per-recommendation confidence are |
| 83 | + different measurements. |
| 84 | +- **No hidden reasoning:** user-facing explanations must be constrained to permitted evidence, |
| 85 | + not chain-of-thought or unvalidated causal language. |
| 86 | +- **Safety preserved:** existing authentication, credential-derived rate limiting, admission |
| 87 | + control, cache keying, fallback, and model-bundle checks remain the API boundary. |
| 88 | + |
| 89 | +Before a large-catalogue deployment, add an evaluated ANN backend, version and validate its index |
| 90 | +with the model bundle, measure candidate recall and end-to-end p95/p99 latency, and keep the |
| 91 | +candidate-ranker contract stable during rollout. |
0 commit comments