Future / To Explore
Things I have not done yet, and want to. One line each, no setup steps.
Table of Contents
Why
This module covers the base: metrics, logs, one dashboard and one alert. A real team goes further. This page keeps a list so the next steps are not forgotten.
To explore
- Tempo (traces). Shows which call inside a slow request took the time. Needs OpenTelemetry in the app (see What Needs the Dev Team).
- Mimir (long-term metrics). Prometheus here has no volume and loses history on removal. Mimir keeps metrics for months, so you can compare this release with last quarter's.
- Grafana Cloud. The same stack, hosted. No containers to run, with a free tier to start.
- Alloy as a service on real Linux servers. Here Alloy runs in a container on a laptop. On a real server it runs as a system service and reads that server's own CPU and disk.
- On-call and alert routing. Send an alert to the right person, at the right time, through a real channel. Add escalation if nobody answers.
- Provisioning dashboards as code. Load the dashboard and alert JSON from files when Grafana starts, so a new container is ready with nothing to import.