Skip to content

skills: add mz-debug-replica-latency triage skill - #37292

Draft
antiguru wants to merge 1 commit into
MaterializeInc:mainfrom
antiguru:mz-debug-replica-latency-skill
Draft

skills: add mz-debug-replica-latency triage skill#37292
antiguru wants to merge 1 commit into
MaterializeInc:mainfrom
antiguru:mz-debug-replica-latency-skill

Conversation

@antiguru

@antiguru antiguru commented Jun 25, 2026

Copy link
Copy Markdown
Member

Adds an agent skill for triaging a production replica that stalls or runs slow.

The skill distinguishes the four common root causes of a multi-second internal latency warning, which usually names a downstream symptom rather than the cause: backing-store (CRDB) slowness, a stalled process runtime, shard contention, and local memory pressure.
It walks the decision tree with Grafana (Prometheus) and PolarSignals, using metric names as the stable contract and discovering datasources and identifiers at runtime, so it carries no environment- or customer-specific values.

Reference files cover the swap-thrash and CPU-starvation branches in detail.

Derived from a recent investigation where a persist reader heartbeat call took 60s warning turned out to be swap thrashing on a single replica, not a consensus problem. See DB-153.

Triage procedure for production replica stalls and slow-persist-op
warnings. Distinguishes backing-store slowness, runtime stalls, shard
contention, and memory pressure using Grafana metrics and PolarSignals
profiles. References cover the swap-thrash and CPU-starvation branches.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant