← worklog

2024-09-15

Long-term metrics storage with Thanos Sidecar

PrometheusThanosGCSGrafanaKubernetes

Problem

Default Prometheus retention of 15 days wasn't enough for capacity planning or post-incident analysis. PVCs large enough for 6+ months were impractical per cluster.

Architecture

Prometheus + Thanos Sidecar → GCS bucket (2h block uploads)
                                   ↓
                            Thanos Store Gateway
                                   ↓
                            Thanos Query ← Grafana

Each cluster has a Prometheus + Sidecar pair. The Sidecar uploads 2h TSDB blocks to GCS. Thanos Query federates across Store Gateways from all clusters.

What I deployed

  • Thanos Sidecar injected as a container in the Prometheus StatefulSet via Helm values
  • GCS bucket with lifecycle rules: retain 13 months, then auto-delete
  • Store Gateway per cluster reading from its own GCS bucket
  • Single Thanos Query frontend as the Grafana datasource

Key decisions

Kept per-cluster Prometheus for local alerting and short-term queries — latency matters for dashboards. Thanos Query only for historical range queries and cross-cluster comparisons.

Compactor runs as a CronJob to downsample older blocks, keeping query performance reasonable on 6-12 month ranges.