|
| 1 | +# Amazon Data Firehose + S3 telemetry path |
| 2 | + |
| 3 | +SentinelAI can optionally mirror accepted inference telemetry from the Go ingestion service to Amazon Data Firehose. Firehose buffers the records and delivers GZIP-compressed newline-delimited JSON (NDJSON) objects to a private S3 telemetry bucket. |
| 4 | + |
| 5 | +This AWS path is **optional**. Local Docker Compose keeps `FIREHOSE_ENABLED=false`, so PostgreSQL remains the default local persistence path and no AWS account is required for development. |
| 6 | + |
| 7 | +## Architecture |
| 8 | + |
| 9 | +```mermaid |
| 10 | +flowchart LR |
| 11 | + Client[Model / application] --> Gateway[NGINX ingestion gateway] |
| 12 | + Gateway --> Go[Go ingestion replicas] |
| 13 | + Go --> DB[(PostgreSQL / Snowflake path)] |
| 14 | + Go -. bounded fail-open mirror .-> Queue[In-memory Firehose queue] |
| 15 | + Queue --> Firehose[Amazon Data Firehose] |
| 16 | + Firehose --> S3[(Amazon S3 telemetry lake)] |
| 17 | + Firehose --> CW[CloudWatch delivery logs] |
| 18 | +``` |
| 19 | + |
| 20 | +The Firehose mirror intentionally does not participate in `/ready`. A temporary AWS failure increments `ingestion_firehose_records_total{status="error"}` and is logged, while the primary ingestion response continues to reflect the primary warehouse write. If the bounded queue fills, records are dropped from the mirror and counted with `status="dropped"` rather than allowing telemetry backpressure to take down ingestion. |
| 21 | + |
| 22 | +## Provision the AWS resources |
| 23 | + |
| 24 | +Prerequisites: |
| 25 | + |
| 26 | +- Terraform 1.6+ |
| 27 | +- AWS credentials available through the standard AWS credential chain |
| 28 | +- Permission to create S3, Firehose, CloudWatch Logs, IAM policy/role, and ECR resources |
| 29 | + |
| 30 | +```bash |
| 31 | +cd terraform |
| 32 | +terraform init |
| 33 | +terraform fmt -check |
| 34 | +terraform validate |
| 35 | +terraform plan |
| 36 | +terraform apply |
| 37 | +``` |
| 38 | + |
| 39 | +The default configuration creates: |
| 40 | + |
| 41 | +- a private, versioned S3 bucket with SSE-S3 encryption; |
| 42 | +- a 30-day telemetry lifecycle policy; |
| 43 | +- an Amazon Data Firehose delivery stream named `sentinelai-telemetry`; |
| 44 | +- 60-second / 5-MiB buffering with GZIP compression; |
| 45 | +- time-partitioned S3 keys under `inference/year=.../month=.../day=.../hour=.../`; |
| 46 | +- a CloudWatch log group for Firehose delivery errors; |
| 47 | +- a Firehose service role with only the S3 and CloudWatch permissions it needs; |
| 48 | +- a separate `sentinelai-firehose-writer` IAM policy granting `PutRecord` and `PutRecordBatch` to the SentinelAI workload; |
| 49 | +- the existing SentinelAI ECR repository, now defined in a valid Terraform `.tf` file. |
| 50 | + |
| 51 | +Use `terraform output` after apply to retrieve the generated S3 bucket name, stream ARN, stream name, and writer-policy ARN. |
| 52 | + |
| 53 | +## Give the ingestion workload permission |
| 54 | + |
| 55 | +Do not put long-lived AWS access keys in the repository. Attach the Terraform output `firehose_writer_policy_arn` to the workload identity used by SentinelAI. On EKS, the intended production pattern is an IAM role associated with the ingestion service account (IRSA / EKS workload identity). |
| 56 | + |
| 57 | +For local development against a real AWS account, the AWS SDK for Go v2 uses its normal credential provider chain. Keep credentials outside the repo. |
| 58 | + |
| 59 | +## Enable the mirror |
| 60 | + |
| 61 | +```bash |
| 62 | +export FIREHOSE_ENABLED=true |
| 63 | +export FIREHOSE_DELIVERY_STREAM=sentinelai-telemetry |
| 64 | +export FIREHOSE_QUEUE_SIZE=1000 |
| 65 | +export AWS_REGION=us-east-1 |
| 66 | +``` |
| 67 | + |
| 68 | +Then start SentinelAI and send a normal inference log: |
| 69 | + |
| 70 | +```bash |
| 71 | +curl -X POST http://localhost:8080/log \ |
| 72 | + -H "Content-Type: application/json" \ |
| 73 | + -d '{"model_id":"demo","model_version":"v1","latency_ms":120,"tokens_in":32,"tokens_out":64,"status":"ok"}' |
| 74 | +``` |
| 75 | + |
| 76 | +The producer appends a newline to every JSON record before `PutRecord`. That keeps individual events parseable after Firehose concatenates buffered records into S3 objects. |
| 77 | + |
| 78 | +## Observability |
| 79 | + |
| 80 | +The ingestion service exports: |
| 81 | + |
| 82 | +```text |
| 83 | +ingestion_firehose_records_total{status="queued"} |
| 84 | +ingestion_firehose_records_total{status="delivered"} |
| 85 | +ingestion_firehose_records_total{status="error"} |
| 86 | +ingestion_firehose_records_total{status="dropped"} |
| 87 | +``` |
| 88 | + |
| 89 | +CloudWatch delivery logs cover the managed Firehose-to-S3 leg. Application metrics cover the producer-side queue and `PutRecord` result. |
| 90 | + |
| 91 | +## Failure semantics |
| 92 | + |
| 93 | +This feature is a telemetry mirror, not a transactional dual-write guarantee. The in-memory queue is intentionally bounded and is not persisted across process termination. AWS SDK retries may also produce duplicate Firehose records in some failure scenarios. Consumers should therefore treat the S3 telemetry dataset as at-least-once/best-effort observability data and use stable event identifiers if strict de-duplication is later required. |
| 94 | + |
| 95 | +## Cost control |
| 96 | + |
| 97 | +Firehose, S3, CloudWatch Logs, and related data transfer can incur AWS charges. The default 30-day S3 expiration and 14-day CloudWatch log retention are intended to keep a portfolio/dev deployment bounded. Review the Terraform plan and AWS pricing before leaving the stack running. |
0 commit comments