Skip to main content

Future / To Explore

Things I have not done yet, and want to. One line each, no setup steps.

Table of Contents


Why​

This module covers the base: metrics, logs, one dashboard and one alert. A real team goes further. This page keeps a list so the next steps are not forgotten.


To explore​

  • Tempo (traces). Shows which call inside a slow request took the time. Needs OpenTelemetry in the app (see What Needs the Dev Team).
  • Mimir (long-term metrics). Prometheus here has no volume and loses history on removal. Mimir keeps metrics for months, so you can compare this release with last quarter's.
  • Grafana Cloud. The same stack, hosted. No containers to run, with a free tier to start.
  • Alloy as a service on real Linux servers. Here Alloy runs in a container on a laptop. On a real server it runs as a system service and reads that server's own CPU and disk.
  • On-call and alert routing. Send an alert to the right person, at the right time, through a real channel. Add escalation if nobody answers.
  • Provisioning dashboards as code. Load the dashboard and alert JSON from files when Grafana starts, so a new container is ready with nothing to import.