Skip to content

feat(#263): implement alert aggregation to prevent alert fatigue - #323

Open
Ebuka042-pixel wants to merge 1 commit into
Pidoko257:mainfrom
Ebuka042-pixel:feat/issue-263-alert-aggregation
Open

feat(#263): implement alert aggregation to prevent alert fatigue#323
Ebuka042-pixel wants to merge 1 commit into
Pidoko257:mainfrom
Ebuka042-pixel:feat/issue-263-alert-aggregation

Conversation

@Ebuka042-pixel

Copy link
Copy Markdown
  • Add AlertAggregator class that groups alerts by service + alertType
  • Suppress transient alerts until threshold is reached within a time window
  • Fire a single aggregated notification once threshold or window elapses
  • Configurable grouping rules: service, alertType (wildcards supported), threshold count, and windowMs — all replaceable at runtime via PUT endpoint
  • Custom groupBy key for sub-grouping within a service+alertType pair
  • Track maxSeverity across grouped alerts (info → warning → error → critical)
  • Redis persistence of group state for cross-replica awareness (TTL=600s)
  • Prometheus metrics: alerts_ingested_total, alerts_fired_total, alerts_deduplicated_total, active_alert_groups
  • Default rules: provider_timeout(5/2min), high_error_rate(1), aml(1), queue_depth(10/5min), catch-all(5/5min)
  • Add admin REST endpoints: GET/PUT/POST /api/admin/alerts/rules, GET /api/admin/alerts/groups, POST /api/admin/alerts/flush, POST /api/admin/alerts/ingest (non-prod only)
  • Wire route into index.ts under requireAuth guard
  • Tests verify: threshold firing, window flushing, deduplication ≥70% reduction, rule resolution priority, severity tracking, flushAll

Description

Brief description of changes.

Related Issue

Fixes #(issue number)

Type of Change

  • Bug fix
  • New feature
  • Documentation update
  • Code refactoring
  • Performance improvement

Changes Made

Testing

How did you test these changes?

Checklist

  • Code follows project style
  • Self-reviewed my code
  • Commented complex code
  • Updated documentation
  • No new warnings
  • Added tests (if applicable)

Screenshots (if applicable)

Additional Notes

closes #263

…igue

- Add AlertAggregator class that groups alerts by service + alertType
- Suppress transient alerts until threshold is reached within a time window
- Fire a single aggregated notification once threshold or window elapses
- Configurable grouping rules: service, alertType (wildcards supported),
  threshold count, and windowMs — all replaceable at runtime via PUT endpoint
- Custom groupBy key for sub-grouping within a service+alertType pair
- Track maxSeverity across grouped alerts (info → warning → error → critical)
- Redis persistence of group state for cross-replica awareness (TTL=600s)
- Prometheus metrics: alerts_ingested_total, alerts_fired_total,
  alerts_deduplicated_total, active_alert_groups
- Default rules: provider_timeout(5/2min), high_error_rate(1), aml(1),
  queue_depth(10/5min), catch-all(5/5min)
- Add admin REST endpoints: GET/PUT/POST /api/admin/alerts/rules,
  GET /api/admin/alerts/groups, POST /api/admin/alerts/flush,
  POST /api/admin/alerts/ingest (non-prod only)
- Wire route into index.ts under requireAuth guard
- Tests verify: threshold firing, window flushing, deduplication ≥70%
  reduction, rule resolution priority, severity tracking, flushAll
@drips-wave

drips-wave Bot commented Jul 29, 2026

Copy link
Copy Markdown

@Ebuka042-pixel Great news! 🎉 Based on an automated assessment of this PR, the linked Wave issue(s) no longer count against your application limits.

You can now already apply to more issues while waiting for a review of this PR. Keep up the great work! 🚀

Learn more about application limits

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Implement Alert Aggregation to Prevent Alert Fatigue

1 participant