Skip to content

Commit 8045a84

Browse files
docs: document Firehose to S3 telemetry architecture
1 parent cd3dfe5 commit 8045a84

1 file changed

Lines changed: 97 additions & 0 deletions

File tree

‎docs/aws-firehose-s3.md‎

Lines changed: 97 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,97 @@
1+
# Amazon Data Firehose + S3 telemetry path
2+
3+
SentinelAI can optionally mirror accepted inference telemetry from the Go ingestion service to Amazon Data Firehose. Firehose buffers the records and delivers GZIP-compressed newline-delimited JSON (NDJSON) objects to a private S3 telemetry bucket.
4+
5+
This AWS path is **optional**. Local Docker Compose keeps `FIREHOSE_ENABLED=false`, so PostgreSQL remains the default local persistence path and no AWS account is required for development.
6+
7+
## Architecture
8+
9+
```mermaid
10+
flowchart LR
11+
Client[Model / application] --> Gateway[NGINX ingestion gateway]
12+
Gateway --> Go[Go ingestion replicas]
13+
Go --> DB[(PostgreSQL / Snowflake path)]
14+
Go -. bounded fail-open mirror .-> Queue[In-memory Firehose queue]
15+
Queue --> Firehose[Amazon Data Firehose]
16+
Firehose --> S3[(Amazon S3 telemetry lake)]
17+
Firehose --> CW[CloudWatch delivery logs]
18+
```
19+
20+
The Firehose mirror intentionally does not participate in `/ready`. A temporary AWS failure increments `ingestion_firehose_records_total{status="error"}` and is logged, while the primary ingestion response continues to reflect the primary warehouse write. If the bounded queue fills, records are dropped from the mirror and counted with `status="dropped"` rather than allowing telemetry backpressure to take down ingestion.
21+
22+
## Provision the AWS resources
23+
24+
Prerequisites:
25+
26+
- Terraform 1.6+
27+
- AWS credentials available through the standard AWS credential chain
28+
- Permission to create S3, Firehose, CloudWatch Logs, IAM policy/role, and ECR resources
29+
30+
```bash
31+
cd terraform
32+
terraform init
33+
terraform fmt -check
34+
terraform validate
35+
terraform plan
36+
terraform apply
37+
```
38+
39+
The default configuration creates:
40+
41+
- a private, versioned S3 bucket with SSE-S3 encryption;
42+
- a 30-day telemetry lifecycle policy;
43+
- an Amazon Data Firehose delivery stream named `sentinelai-telemetry`;
44+
- 60-second / 5-MiB buffering with GZIP compression;
45+
- time-partitioned S3 keys under `inference/year=.../month=.../day=.../hour=.../`;
46+
- a CloudWatch log group for Firehose delivery errors;
47+
- a Firehose service role with only the S3 and CloudWatch permissions it needs;
48+
- a separate `sentinelai-firehose-writer` IAM policy granting `PutRecord` and `PutRecordBatch` to the SentinelAI workload;
49+
- the existing SentinelAI ECR repository, now defined in a valid Terraform `.tf` file.
50+
51+
Use `terraform output` after apply to retrieve the generated S3 bucket name, stream ARN, stream name, and writer-policy ARN.
52+
53+
## Give the ingestion workload permission
54+
55+
Do not put long-lived AWS access keys in the repository. Attach the Terraform output `firehose_writer_policy_arn` to the workload identity used by SentinelAI. On EKS, the intended production pattern is an IAM role associated with the ingestion service account (IRSA / EKS workload identity).
56+
57+
For local development against a real AWS account, the AWS SDK for Go v2 uses its normal credential provider chain. Keep credentials outside the repo.
58+
59+
## Enable the mirror
60+
61+
```bash
62+
export FIREHOSE_ENABLED=true
63+
export FIREHOSE_DELIVERY_STREAM=sentinelai-telemetry
64+
export FIREHOSE_QUEUE_SIZE=1000
65+
export AWS_REGION=us-east-1
66+
```
67+
68+
Then start SentinelAI and send a normal inference log:
69+
70+
```bash
71+
curl -X POST http://localhost:8080/log \
72+
-H "Content-Type: application/json" \
73+
-d '{"model_id":"demo","model_version":"v1","latency_ms":120,"tokens_in":32,"tokens_out":64,"status":"ok"}'
74+
```
75+
76+
The producer appends a newline to every JSON record before `PutRecord`. That keeps individual events parseable after Firehose concatenates buffered records into S3 objects.
77+
78+
## Observability
79+
80+
The ingestion service exports:
81+
82+
```text
83+
ingestion_firehose_records_total{status="queued"}
84+
ingestion_firehose_records_total{status="delivered"}
85+
ingestion_firehose_records_total{status="error"}
86+
ingestion_firehose_records_total{status="dropped"}
87+
```
88+
89+
CloudWatch delivery logs cover the managed Firehose-to-S3 leg. Application metrics cover the producer-side queue and `PutRecord` result.
90+
91+
## Failure semantics
92+
93+
This feature is a telemetry mirror, not a transactional dual-write guarantee. The in-memory queue is intentionally bounded and is not persisted across process termination. AWS SDK retries may also produce duplicate Firehose records in some failure scenarios. Consumers should therefore treat the S3 telemetry dataset as at-least-once/best-effort observability data and use stable event identifiers if strict de-duplication is later required.
94+
95+
## Cost control
96+
97+
Firehose, S3, CloudWatch Logs, and related data transfer can incur AWS charges. The default 30-day S3 expiration and 14-day CloudWatch log retention are intended to keep a portfolio/dev deployment bounded. Review the Terraform plan and AWS pricing before leaving the stack running.

0 commit comments

Comments
 (0)