Skip to content

Latest commit

 

History

History
575 lines (490 loc) · 24.1 KB

File metadata and controls

575 lines (490 loc) · 24.1 KB

Observation plane

Status: target design contract. Last updated: 2026-07-07.

The observation plane owns trace, audit, introspect, reflect, status, diagnostics, and supportability. It describes what the system knows about itself and what the agent/user may inspect.

This is a normative design contract, not an implementation status report.

Principles

  • Trace is structured runtime fact, not log text.
  • Audit is durable accountability, not debug noise.
  • Introspect reports current state and capability boundaries.
  • Reflect reasons about strategy and quality within policy.
  • Diagnostics must be specific enough to support recovery.

Fact layers

Layer Examples
C0 control session, run, task, checkpoint, leases.
C1 transcript user-visible messages.
C2 audit run events, permission decisions, provider decisions.
C3 trace LLM rounds, tool lifecycle, retry/cache/sync decisions.
C4 debug raw captures, support bundles, manifests.
C5 learning redacted examples, labels, eval artifacts.

Default C3 events

An assistant answer is C1 conversation content, not an automatic decision fact. Explicit decision records link to the event that establishes the decision and its owner; copying answer text or tool names into an audit row cannot establish authorization or the admitted tool surface. Canonical answer and usage producers remain independent of that explicit decisions API.

The default trace schema should include:

  • llm_round_completed;
  • tool_call_started;
  • tool_call_completed;
  • tool_call_failed;
  • provider_decision;
  • step_verdict;
  • retry_decision;
  • cache_decision;
  • sync_status.

Common causal fields:

session_id
run_id
parent_run_id
turn_id
round_index
tool_call_id
provider_id
capability
cause_event_id

Introspect

Introspect should answer:

  • what state am I in;
  • what providers are available;
  • what tools are visible and why;
  • what is blocked and how to unblock it;
  • what context was loaded;
  • what sync state is safe or degraded;
  • what recent failures matter.

Reflect

Reflect is agent reasoning over observation facts. It should not mutate state by itself. It can propose strategy, identify uncertainty, and request action.

Debug bundles

Raw debug bundles are C4 and off by default.

Requirements:

  • explicit user or policy enablement;
  • short TTL;
  • manifest;
  • redaction boundary;
  • export and delete operations;
  • audit event for creation/access/deletion;
  • exclusion from default learning pipeline.

Diagnostics quality

Bad diagnostic examples:

  • unknown tool reported as missing runtime binding;
  • provider offline reported as malformed call;
  • plan-mode policy denial reported as tool absence;
  • sync poison hidden as generic failure.

Good diagnostics include cause, scope, affected capability, resumability, and next action.

Trace payload contracts

provider_decision

event_type = provider_decision
session_id
run_id
turn_id
capability
tool_name
provider_type
provider_id
route
admission_status
runtime_binding_status
fallback_policy
fallback_from
degraded_reason
offline_reason

retry_decision

event_type = retry_decision
session_id
run_id
turn_id
round_index
tool_call_id
retry_reason
retryable
attempt
max_attempts
next_action

cache_decision

event_type = cache_decision
session_id
run_id
turn_id
prompt_contract_version
stable_prefix_hash
dynamic_block_hash
cache_expected
cache_hit
miss_reason

step_verdict

event_type = step_verdict
session_id
run_id
turn_id
step_id
verdict
confidence
reasons
next_action

tool_call_started

event_type = tool_call_started
session_id
run_id
turn_id
round_index
tool_call_id
tool_name
provider_id
route
arguments_hash
started_at

tool_call_completed

event_type = tool_call_completed
session_id
run_id
turn_id
round_index
tool_call_id
tool_name
provider_id
status
duration_ms
quality_status
result_artifact_ref
completed_at

tool_call_failed

event_type = tool_call_failed
session_id
run_id
turn_id
round_index
tool_call_id
tool_name
provider_id
error_kind
retryable
quality_status
fallback_available
failed_at

Model request attribution and usage

For local CLI child runs, the existing agent_spawned journal and Trace event carry a credential-free model_configuration snapshot: effective thinking and first-output cap at spawn, the requested Offering when known, and a prepared Offering/model identity when available. catalog_resolved means directory resolution only; admission_validated means batch admission passed before launch. The requested Offering may be inherited, not explicitly chosen by the user. These are spawn-time configuration facts, never proof that a provider accepted the request or consumed a particular reasoning level. Actual request and usage attribution remains with inference evidence; absent evidence is unknown, not zero.

The inference ledger stores content-free accepted/terminal request diagnostics in model_request_context_events. These records complement agent_events and are exposed through the existing owner-scoped request queries and session observability projection. They do not require full prompt capture.

ModelRequestContextEvent.route projects non-secret facts from the admitted inference plan:

route_id
invocation_id
upstream_model
execution_placement
access_kind

The adjacent request identity owns the Offering, configured model, purpose, session/run/turn/round, and physical request ID. Accepted and terminal events must agree on route attribution, including recovery after process failure. Older records without this projection deserialize with an unknown route; readers must not infer selection policy from a model name or missing field. Credentials, owner/admission tokens, endpoint URLs, and raw prompt content do not belong in this projection.

Request diagnostics use strict typed readers. Deployments adding route fields must upgrade readers and recovery workers before new producers, or drain old instances during a coordinated upgrade. Backward reading of stored records without route fields does not imply older binaries can read the new projection.

Usage coverage is independent of request outcome:

  • provider_exact: observed usage is available, including a measured zero. Full-input budget errors and cache-read share may be computed.
  • provider_partial: retain observed token lanes, but do not report them as a complete input measurement, budget estimate error, or cache-read share.
  • unavailable: usage and measured diagnostic fields are absent. Placeholder zeroes in legacy records are not evidence of zero-token billing.

Session Audit request usage reads retained physical provider attempts for the authenticated owner and session. It counts failures, cancellation and unknown delivery as attempts; each token lane reports a known sum and the number of attempts that observed that lane. A missing lane is unknown, not zero. This snapshot is not proof that retained records cover the session's lifetime. Inference routes retain optional admission-time price snapshots. Session Audit prices each retained physical attempt once, including priced retries and failures, using its retained snapshot and separately observed fresh-input, output, cache-read and cache-creation lanes. The USD estimate is not a provider invoice or proof of complete lifetime coverage. Empty ledgers, nonterminal attempts or missing price or usage evidence keep cost unavailable with historical_attempt_coverage_incomplete. Current catalog prices or a display model name never fill those gaps or reprice history.

Live runtime and feedback cache-read percentages require a positive observed input-token denominator. With no input measurement they display unknown, not 0%; an observed positive input with zero cache reads may display 0%. This live snapshot is not a durable session or task-tree usage aggregate. It does not establish whether a provider omitted cache buckets; request-ledger usage coverage remains authoritative for that distinction.

Accepted events have no provider measurements. Terminal diagnostics with unavailable usage store nullable token columns as NULL. Aggregate request counters still count those attempts; token sums contain only observed usage and do not establish complete billing coverage. Consumers estimating cost or building learning examples must retain per-request coverage and treat missing or expired diagnostics as unknown. Foreground settlement and recovery share the same coverage projection.

Persisted token evidence and audit projections

Canonical persisted token lanes are disjoint fresh input, cache read, cache creation, and output. Missing or null is unknown, not measured zero. Derived input and total counts require all contributing lanes. Malformed supplied counts, contradictory totals, and overflowing known subtotals are rejected at the shared parsing boundary. An observed but unavailable accounting sample is distinct from absence of an accounting sample.

Journal numeric evidence stays in the existing nullable token fields. Only an all-unknown sample carries metadata.qualified_usage: {} as a presence witness; this metadata field must not duplicate numeric lanes. Invalid witnesses or witnesses conflicting with numeric fields quarantine the sample as unavailable. Run summaries retain total_tokens: null when accounting scope metadata would otherwise erase the canonical empty-sample shape. Recovery requires explicit nullable qualified and latest-physical usage fields; checkpoints missing them are rejected rather than inferred from observed subtotals.

Audit turn/detail token fields and session request-usage lanes serialize unknown values as explicit nulls. Session lane sums use checked addition; missing evidence or overflow makes that lane unknown without discarding other lanes. An explicitly examined empty sample set has zero totals, whereas a missing summary defaults to unknown. A newer cumulative response replaces the older response; it must not borrow older values to fill unknown fields.

Cost estimates require all four lanes and their applicable prices. Partial samples count as unpriced, not as free requests. An all-unpriced sample set omits the cost estimate and per-model map. These counts describe selected accounting rows, not proof of coverage of every physical provider attempt. Top-level session SQL token sums remain observed subtotals, not certified complete totals. Primary and auxiliary attribution must be established before interpreting a sample as one model's cost.

Workspace/session restoration counters remain observed subtotals. Combining local and cloud counters does not establish common request coverage or model attribution. /ask must carry that limitation instead of deriving a session cache percentage. The current-session /cost and TUI /stats cost view is a current-rate scenario on those counters, not session billing; missing prices remain unavailable, and no actual savings or per-turn average is inferred. Qualified per-turn metrics remain independent of these historical counters.

Live cumulative monetary estimates are nullable through session state, rollback, TUI and dump projections. A known-empty local state may start at zero, but attaching a session does not establish historical cost coverage. Aggregate token counters plus current or fallback model prices cannot qualify the complete cost of primary, auxiliary, delegated and retried work. Until that attribution is available, successful and failed turns retain an unknown cost, not a zero amount or a priced subset. Journal-only exports likewise have no complete monetary evidence. Current-rate scenarios remain available separately.

Semantic judgment trace

Server lifecycle observability flushes prepare one annotated journal batch for local persistence and the existing bounded ingestion sender. Only generic TraceSpan records use this handoff; other event families retain their existing durable owners. Trusted owner/session and the captured execution generation bind the process-local sink once; this is not per-fact generation verification. Generic trace IDs need not be run IDs. Typed semantic readers validate their own run correlation. Historical facts may flush after waiting, cancellation or owner transfer without granting the old execution new authority. Local IO failure does not suppress enqueue or consume the retained buffer; missing/closed/full ingestion does not suppress local persistence. Exact batch replay retains the content-addressed storage ID. Changed interruption/eviction annotations change that ID; semantic readers reconcile equivalent observation identities, while generic trace readers must not assume annotation-changing retries are unique. The shared ingestion queue is bounded and prioritizes critical audit traffic; it does not promise per-owner telemetry fairness or complete trace capture. One flush partitions accepted facts by authenticated (owner, session) and commits each partition in its own transaction. Per worker, active partition transactions are capped by the configured ceiling, one half of the pool, and the pool maximum minus a two-connection reserve. Each term has a minimum of one so one-connection test or emergency pools can still make progress; that small-pool exception cannot reserve foreground capacity. Event rows, causal edges, the exact session counter delta, inserted session-end effects, and config-version projections share that partition transaction. A blocked or retryable partition therefore cannot retain locks for, roll back, or replay an unrelated session. Successful and permanently rejected partitions leave the retry set immediately; only unresolved partitions retain their already- accounted queue payloads. This per-worker cap limits connection amplification but is not a shared-pool reservation or a fairness guarantee across workers or server processes.

The worker coordinator keeps database attempts and retry timers outside its receive path. Accepted facts enter a per-owner/per-session FIFO; only one attempt for that key may be active, and a failed head remains ahead of later facts for the same session. Runnable owners rotate first, then their runnable sessions, so one owner with many hot sessions cannot consume every dispatch turn. batch_size bounds one session transaction. A sparse session becomes runnable at an absolute deadline established by its oldest buffered fact; later arrivals do not reset that deadline. Retry backoff consumes neither a database slot nor the receive loop. This is process-local dispatch fairness, not equal SQL execution time or cluster-wide fairness.

Every accepted fact owns one admission lease from before channel entry through queued, retrying, and in-flight states. Global, per-owner, and per-session limits apply to both event count and compact-JSON bytes; one blocked session or owner therefore cannot consume all process-local headroom. A maximum event size rejects oversized payloads before acceptance. Lease release is tied to terminal drop rather than scheduler bookkeeping, so cancellation and closed-channel paths return capacity as well.

Database attempts use a process-local limiter that can be shared by all workers over one pool. A whole-attempt client deadline includes connection acquisition, transaction work, and commit. If an exchange times out, the physical connection is detached and closed rather than returned to the idle pool; acknowledged commits remain terminal even if later cleanup fails. The limiter and deadline do not establish cross-process fairness or database-cluster capacity. The client deadline also does not establish a server-side rollback deadline: discarding a timed-out connection can restore logical pool capacity before the server releases transaction locks. Consequently, unrelated session writes must not share a transaction on the assumption that socket cancellation bounds their lock coupling. Independent session transactions preserve that isolation boundary.

Enqueue-to-terminal latency uses a fixed-size process-local histogram rather than retaining per-event samples. A terminal outcome is commit, durable admission rejection, or explicit shutdown abandonment; retryable attempts keep their original enqueue timestamp. Shutdown seals the receiver before draining, so deferred sends that never entered the channel remain pre-acceptance drops. If the runtime deadline expires, it aborts and awaits the worker instead of detaching a task that may still own pool resources. Delivery accounting belongs to the shared admission lease: channel acceptance precedes dispatch, and a known commit or rejection settles the lease before control returns to the scheduler. Explicit shutdown failure or final-owner drop counts an accepted, unsettled delivery as unresolved exactly once across retry clones. Residency snapshots are not terminal evidence and never add to that count.

Request-classification observations use the existing trace_span envelope with name semantic_judgment and a bounded typed JSON string in attrs["semantic_judgment.v1"]. They do not create another usage or execution ledger. Initial classification and clarification are separate semantic stages, not a count of physical provider attempts. A classification that succeeds before planning fails remains a successful classification, not a failed judgment.

The payload contains closed reasons, normalized answer values and provenance, run/turn/round correlation and a preflight evaluation identity. It excludes raw provider responses, prompts, tool output, parser error text and credentials. Trace shape alone is not producer authentication; these facts cannot authorize execution or settlement. Readers use authenticated owner/session storage scope, bound bytes before decoding, and exclude conflicting observation or terminal evaluation identities. Optional database reads use the existing cancellation-safe connection boundary.

Known preparation failures are not-dispatched facts. Transport failure, cancellation and timeout do not by themselves establish whether a request was dispatched; delivery remains unresolved unless response receipt is known. Classification results do not authorize execution or prove model adoption or quality improvement. Success does not invent scores absent from the classifier's result. Unsupported old semantic payloads are rejected rather than migrated.

Trace delivery remains lossy. Successful empty reads do not prove inactivity; query truncation, display omission and potential upstream loss remain distinct. Physical attempts and token totals continue to come from the inference ledger.

Agent event field requirements

Agent event storage should support the following logical fields, whether physically normalized or stored with indexed metadata:

event_id
user_id
session_id
run_id
parent_run_id
turn_id
turn_seq
round_index
tool_call_id
event_type
trace_kind
provider_id
capability
cause_event_id
parent_event_id
created_at
server_received_at
payload_hash
redaction_status
retention_class
metadata

event_id must be stable and collision-resistant. If the same event id arrives with different payload hash, ingestion must treat it as a collision, not idempotent success.

Durable capture distinguishes Inserted, Replayed, and Collision within the owning database transaction. Only Inserted may apply event counters, parent edges, terminal-session effects, configuration projections, or manifest artifact references. An exact retry after a lost commit acknowledgement therefore does not repeat those effects. Each database attempt uses a fresh write marker; retries must not reuse a marker from an attempt whose commit outcome is unknown.

Capture outcomes remain available to downstream projections. A collision must not create a transcript row or snapshot link from the rejected payload, nor be reported as a successfully captured response. Exact replay may repair a missing projection from the accepted payload under the existing transaction and session fences. Valid sibling events in the same batch continue to be captured.

Atomic run-terminal settlement is stricter than ordinary batch capture: every canonical event in the settlement (including user intents, rounds, and tool events) must be inserted or exactly replayed. Any collision rolls back the settlement before terminal status, usage, or transcript changes can commit. Initial commit and lost-acknowledgement recovery therefore require the same complete evidence; recovery must not relax hash verification to accept a subset.

Manifest identity and item identity are tenant-scoped: (user_id, manifest_id) and (user_id, manifest_id, item_order). The manifest digest covers its header and the complete item set ordered by item_order. Every item lookup, join, and delete must carry the owner. Unknown-reason diagnostics commit with the first manifest capture and are not repeated by replay or collision.

Collision receipts retain identities and hashes, never observation payloads. They aggregate into one row per (user_id, identity_kind, identity_id), with a count and latest conflicting hash. Their fixed seven-day expiry is not extended by repeated conflicts. The bounded runtime-maintenance sweep removes expired receipts, and explicit session deletion removes its owner-scoped receipts. Batch producers submit receipts in input order through one shared writer, in chunks of at most 128 receipts. Each multi-receipt chunk uses one database grouping query followed by one weighted upsert (at most 1,024 binds); a single receipt needs only the upsert. Grouping inherits the target columns' types and collation and returns first/last input ordinals and counts. This avoids passing duplicate fresh identities to MatrixOne's multi-row upsert. Callers do not normalize identities, and every receipt is checked against schema field bounds before reduction. Each input increments the count once. Duplicate updates preserve the first stored hash, session, source, first-seen time, and fixed expiry, while the last input supplies the attempted hash. Timestamps have statement granularity; expiry remains seven days after first observation. Empty batches issue no SQL, and all chunks share the caller's transaction and rollback boundary. This saves statement round trips and repeated row updates within a batch, not contention between independent transactions on the same identity.

The v83 core schema requires capture hashes and attempt markers on event and manifest writes. Deployments using an earlier table shape require a fresh-schema cutover; startup rejects missing capture columns rather than assigning empty hashes to old rows. Hashes include producer occurrence timestamps. Root execution freezes the turn-start timestamp before its first durable write, then freezes the terminal offset after execution. Delayed retries and lost-acknowledgement resolution reuse both bounds, not a later event-buffer start or the current clock. Content and lineage still undergo full hash comparison. Configuration version pushes are the exception: their envelope has a generated delivery time, so that field is hashed as JSON null and the first stored delivery time is kept.

Event ingestion unhappy paths

Path Required behavior
Invalid payload shape Reject or quarantine without poisoning unrelated records.
Missing causal fields Accept only if event type permits; otherwise degraded/quarantine.
Oversized metadata Store summary/artifact ref or reject according to policy.
Redaction failure Fail closed.
Duplicate same hash Idempotent.
Duplicate different hash Collision/poison.
Unknown event type Store only if policy allows extension; otherwise quarantine.
Retention class missing Apply safe default, not infinite retention.

Observation metrics

agent_events_ingested_total
agent_events_rejected_total
agent_events_quarantined_total
agent_event_collision_total
trace_events_by_type_total
introspection_requests_total
reflection_requests_total
debug_bundle_created_total
debug_bundle_access_total
debug_bundle_expired_total

The retired process-local decision-explainer recorder had no production writer; its empty /telemetry decisions view and counters are removed. Context assembly explanations, provider decisions, durable journal evidence and Explain Analyze remain with their existing producers and readers. Retiring unused adaptive/drift journal construction helpers does not remove decoding of durable event records.