Problem
Default Prometheus retention of 15 days wasn't enough for capacity planning or post-incident analysis. PVCs large enough for 6+ months were impractical per cluster.
Architecture
Prometheus + Thanos Sidecar → GCS bucket (2h block uploads)
↓
Thanos Store Gateway
↓
Thanos Query ← Grafana
Each cluster has a Prometheus + Sidecar pair. The Sidecar uploads 2h TSDB blocks to GCS. Thanos Query federates across Store Gateways from all clusters.
What I deployed
- Thanos Sidecar injected as a container in the Prometheus StatefulSet via Helm values
- GCS bucket with lifecycle rules: retain 13 months, then auto-delete
- Store Gateway per cluster reading from its own GCS bucket
- Single Thanos Query frontend as the Grafana datasource
Key decisions
Kept per-cluster Prometheus for local alerting and short-term queries — latency matters for dashboards. Thanos Query only for historical range queries and cross-cluster comparisons.
Compactor runs as a CronJob to downsample older blocks, keeping query performance reasonable on 6-12 month ranges.