drop-observability-plan

Drop — System Support & Observability Plan

Client: Drop (Digital Banking) Executing Company: FlowForge (DevOps & Infrastructure) Support Company: HelixSupport (Production Support & SLA) Status: DRAFT — čeka CEO approval Created: 2026-02-20


Current State (AS-IS)

Drop već ima:

Drop nema:


Target State (TO-BE)

Tier 1: Essential (Week 1)

Cost: ~$0 — free tiers

Component Tool Why
Uptime monitoring BetterStack (free: 10 monitors) Independent od AWS — znaš kad padne PRIJE korisnika
Error tracking Sentry (free: 5K events/mo) Stack traces, user context, release tracking
Log shipping AWS CloudWatch Logs (App Runner native) Searchable logs, retention, metric filters

Tier 2: Visibility (Week 2)

Cost: ~$0-20/mo

Component Tool Why
DB monitoring RDS Performance Insights (free tier) Slow queries, connection pool, wait events
CDN analytics Cloudflare Analytics (free) Traffic patterns, threats, cache hit rate
Alerting escalation BetterStack On-call (free: 1 team) Slack → Email → SMS escalation chain

Tier 3: Intelligence (Week 3-4)

Cost: ~$0-50/mo

Component Tool Why
Business metrics Custom endpoint + Grafana Cloud (free: 10K metrics) Tx/hour, success rate, revenue
Application metrics Prometheus client in app Request latency, error rate, saturation
Dashboards Grafana Cloud (free tier) SLO tracking, operational dashboards

Tier 4: Advanced (Future)

Kad bude potreba (production scale)

Component Tool Why
Distributed tracing OpenTelemetry → Grafana Tempo Cross-service request flow
Log aggregation Grafana Loki Centralized searchable logs
Chaos engineering Manual game days Resilience validation

FlowForge Execution Plan

Phase 1: Plan (ovaj dokument)

Gate: Alem kaže GO

Phase 2: Provision (Week 1)

FlowForge SRE → implementacija:

2a. BetterStack Uptime (30 min)

2b. Sentry Re-integracija (1-2h)

2c. CloudWatch Logs (1h)

Phase 3: Deploy (Week 2)

3a. RDS Performance Insights (15 min)

3b. Cloudflare Analytics (15 min)

3c. Alerting Escalation (30 min)

Phase 4: Monitor (Week 3-4)

4a. Business Metrics Endpoint (2-3h)

4b. Application Metrics (2-3h)

4c. SLO Dashboard (1-2h)

Phase 5: Optimize (Ongoing)

HelixSupport preuzima:


Cost Summary

Item Monthly Cost
BetterStack Uptime (free tier) $0
Sentry (free tier, 5K events) $0
CloudWatch Logs (App Runner) ~$5
RDS Performance Insights (free) $0
Grafana Cloud (free tier) $0
Total ~$5/mo

Kad Drop skalira → upgrade na paid tiers (~$50-100/mo).


Success Metrics

Metric Target
MTTD (Mean Time to Detect) < 5 min
MTTR (Mean Time to Recover) < 1h (P1)
Uptime SLO 99.9%
Undetected outages 0
Alert noise (false positives) < 10%

Dependencies


Approval

CEO Decision Required:


Revision #4
Created 2026-02-20 09:59:12 UTC by John
Updated 2026-06-28 20:01:19 UTC by John