Dashboards and Alerts
Grafana can now read Prometheus and Loki. This page turns that into the everyday view from Everyday Monitoring: one dashboard, and one alert that tells you when the api starts failing.
Table of Contents
- Why
- How: build the dashboard
- How: or import the dashboard
- How: create the alert
- How: trigger the alert on purpose
- Tips
Why
A dashboard is what you look at. An alert is what looks for you. A team needs both. Nobody stares at a dashboard all day, so the alert watches the one thing that must never go unnoticed: errors.
The dashboard has four rows, one for each golden signal. The alert uses the same ERROR log query as the Errors row. So what the alert says matches what the dashboard shows.
How: build the dashboard
Open http://localhost:3001. In the left menu go to Dashboards, then New, then New dashboard. Click Add visualization and pick a data source. For every panel below:
- Pick the data source in the table.
- Switch the query editor to Code, paste the query and run it.
- Set the panel Title.
Add a Row with the name of each signal, and put its panels under it. Name the dashboard Pizza Playground — Everyday.
| Row | Panel | Data source | Query |
|---|---|---|---|
| Traffic | DB transactions per second | Prometheus | sum(rate(pg_stat_database_xact_commit{datname="playground"}[1m])) |
| Traffic | DB rows returned per second | Prometheus | sum(rate(pg_stat_database_tup_returned{datname="playground"}[1m])) |
| Errors | api ERROR lines per minute | Loki | sum(count_over_time({service="api"} |= "ERROR" [1m])) |
| Errors | Error log lines (visualization Logs) | Loki | {service=~"api|mock-gateway"} |= "ERROR" |
| Latency | Text panel (visualization Text) | none | No latency metric without app changes. During a test, k6 provides it (page 10). For every day, see page 11. |
| Saturation | CPU per service | Prometheus | sum by (service) (rate(container_cpu_usage_seconds_total{service!=""}[1m])) |
| Saturation | Memory per service | Prometheus | sum by (service) (container_memory_working_set_bytes{service!=""}) |
| Saturation | DB connections | Prometheus | sum(pg_stat_database_numbackends{datname="playground"}) and, as a second query, pg_settings_max_connections |
In the table, \| is how a pipe is written inside a table cell. In Grafana you type a plain |.
Save with Save dashboard. Open the playground in the browser and click around.
Expected: the Traffic and Saturation panels show lines, with api, db, web and mock-gateway in the CPU and memory panels. The api ERROR lines per minute panel and the Error log lines panel are empty. That is the good result. The playground has no errors right now.
How: or import the dashboard
Skip the building and load the finished one from monitoring/sample/grafana/everyday-dashboard.json.
- Go to Dashboards, then New, then Import.
- Click Upload dashboard JSON file and pick
monitoring/sample/grafana/everyday-dashboard.json. - Under the data source choices, select
PrometheusandLoki. - Click Import.
Expected: the same dashboard opens, and every panel except the error ones shows data.
How: create the alert
The alert counts api ERROR lines over the last 5 minutes. It fires when there are more than 10 for a full minute.
-
Go to Alerting, then Alert rules, then New alert rule.
-
Rule name:
api ERROR lines. -
Pick the
Lokidata source. In the Code editor, paste:sum(count_over_time({service="api"} |= "ERROR" [5m])) -
The 5 minute look-back is the
[5m]inside the query. Use the Instant query type. -
Add a Reduce expression on that query using the function
Last. Add a Threshold expression on the reduced value: with the condition above10. Make the threshold the alert condition. -
Under Add folder and labels, create the folder
Playground. Add the labelseverity=warning. -
Under Set evaluation behavior, create the New evaluation group
everydaywith interval1m. Set the Pending period to1m. -
Set what happens with no data to Normal. Quiet logs give no data, and that is not a problem.
-
Add the annotation summary
More than 10 ERROR lines from the api in 5 minutes, then save the rule.
monitoring/sample/grafana/api-error-alert.yaml is a reference export of this rule, in Grafana's file-provisioning format (loaded from /etc/grafana/provisioning/alerting/, see provisioning as code). The steps above are how you create it. The UI import form does not read this file. To use it by provisioning, first replace datasourceUid: fg0ht4plq27eoe with your own Loki uid. Find it under Connections, then Data sources, then Loki: the uid is the last part of the page address, .../connections/datasources/edit/<uid>.
Where the message goes. An alert that nobody hears is useless. Grafana sends messages through a contact point, picked by the notification policy. Look at both:
- Alerting, then Contact points. Each contact point is one channel. A team creates one with Create contact point and picks the integration: Email, Slack, Microsoft Teams, Webhook and more.
- Alerting, then Notification policies. The default policy chooses which contact point gets the alert.
This playground has none yet, so the alert only changes state on screen. Do not add a real channel for this exercise. A team would add one here for the people on call.
Expected: Alerting, then Alert rules lists api ERROR lines in the folder Playground, in state Normal.
How: trigger the alert on purpose
Break the database for about 45 seconds. The api then fails every request and writes ERROR lines.
From the playground folder:
docker compose stop db
for i in $(seq 12); do curl -s -m 5 -o /dev/null http://127.0.0.1:8080/api/v1/menu & done; wait
sleep 40
docker compose start db
Open Alerting, then Alert rules, and watch api ERROR lines. Also keep the dashboard open on the Errors row.
What I saw, counting from the moment the db was stopped:
| After | State | What happened |
|---|---|---|
| 0 s | Normal | db stopped, 12 requests sent |
| about 45 s | Normal | ERROR lines appear in Loki: 36 in all (3 per request), logged 30 s after each request, plus ingest delay |
| about 50 s | Pending | The next evaluation sees 36, which is above 10 |
| about 1 min 55 s | Firing | The condition held for the 1 minute pending period |
| about 5 min 50 s | Normal | The 5 minute window moved past the ERROR lines |
Expected: the rule goes Normal, then Pending, then Firing, then back to Normal. The db is running again from second 45, so the playground works while the alert is still firing.
Tips
- The Latency row is a text box on purpose. Prometheus has no request timing for the api, because the app does not publish any. Without changes to the app there is nothing to draw. During a test, k6 measures latency itself (page 10). For every day, page 11 lists what the dev team would add.
- There is no network panel.
container_network_receive_bytes_totaldoes not appear on OrbStack, so the Traffic row uses database activity instead. On plain Linux or Docker Desktop the network metric may exist. Check in Explore before you add it. - Keep the db down for the whole 45 seconds. The api logs the errors only after its 30 second wait for a connection. If the db returns sooner, the waiting requests succeed and no ERROR lines appear.
curl -m 5is there for speed, not for the alert. Without it, each request hangs for 30 seconds. The api still logs its errors either way.- A contact point needs a real channel in team use. On screen the alert is fine for learning. In a team, an alert with no contact point is a silent alarm.
- The import expects data sources called
PrometheusandLoki. If yours have other names, pick them in the import form. The names from Set Up Grafana work with no choices. - Keep the files as a backup. Grafana here has no volume, so removing the container erases the dashboard and the rule. The dashboard JSON imports as is. Rebuild the alert from the steps above, or provision the YAML after you replace the uid.
Next: Watch a Load Test.