Watch a Load Test
You have a dashboard for the server. k6 can write its own numbers into Prometheus. This page puts both on one screen and shows how to read them during a load test.
Table of Contents
- Why
- How: take a baseline
- How: run k6 with results in Prometheus
- How: open the load test dashboard
- How: what to read, in order
- How: the same view with JMeter
- Tips
Why
k6 tells you what the user felt: response time, requests per second, failures. The server dashboard tells you why. Apart, you have two graphs and a guess. Together, with the same time axis, you can say "p95 rose at 14:03, and api CPU rose at the same moment".
The setup is from Monitoring During a Performance Test: baseline first, then the test, then compare. The load test itself is explained in the k6 load model and execute and analyze pages.
How: take a baseline
Stop everything else that talks to the playground. Leave the stack idle for 5 minutes. Then read the idle values from Prometheus, in the Explore view on the Prometheus data source, or on the dashboard:
sum by (service) (rate(container_cpu_usage_seconds_total{service!=""}[5m]))
sum(pg_stat_database_numbackends{datname="playground"})
Write them down. These are your "normal".
What I saw at idle:
| Measure | Idle value |
|---|---|
| api CPU | 0.013 cores |
| db CPU | 0.015 cores |
| DB connections | 12 |
Expected: every number is small and steady.
How: run k6 with results in Prometheus
Prometheus must be started with --web.enable-remote-write-receiver. That is in the start command from Set Up Prometheus. Run from the repo root, or from the folder where you unzipped both samples (get the k6 sample if you do not have it):
cd k6/sample
K6_PROMETHEUS_RW_SERVER_URL=http://127.0.0.1:9090/api/v1/write \
K6_PROMETHEUS_RW_TREND_STATS="p(95),avg,max" \
k6 run -o experimental-prometheus-rw -e PROFILE=load -e SCALE=0.1 -e USERS=30 05-load-model.js
SCALE=0.1 shrinks the 30 minute load profile to about 3 minutes. USERS=30 is the peak number of virtual users. I used 30 because 10 barely moved the api CPU.
Check that the data arrived:
curl -s localhost:9090/api/v1/query --data-urlencode 'query=k6_http_reqs_total' | python3 -c 'import sys,json;print(len(json.load(sys.stdin)["data"]["result"]),"series")'
Expected: a number above 0. k6 also prints its normal summary at the end. In my run all 16230 checks passed, with 0% failed requests.
How: open the load test dashboard
Import monitoring/sample/grafana/load-test-dashboard.json the same way as the everyday one:
- Go to Dashboards, then New, then Import.
- Click Upload dashboard JSON file and pick
monitoring/sample/grafana/load-test-dashboard.json. - Select
PrometheusandLokias the data sources. - Click Import.
The dashboard is named Pizza Playground — Load Test. It has three rows:
| Row | Panel | Query |
|---|---|---|
| k6 | Virtual users | k6_vus |
| k6 | Requests per second | sum(rate(k6_http_reqs_total[1m])) |
| k6 | p95 response time | max by (name) (k6_http_req_duration_p95{scenario!=""}) |
| Saturation | CPU per service (cores) | sum by (service) (rate(container_cpu_usage_seconds_total{service!=""}[1m])) |
| Saturation | Memory per service | sum by (service) (container_memory_working_set_bytes{service!=""}) |
| Saturation | DB connections | sum(pg_stat_database_numbackends{datname="playground"}) and pg_settings_max_connections |
| Errors | api ERROR lines per minute | sum(count_over_time({service="api"} |= "ERROR" [1m])) |
| Errors | Error log lines | {service=~"api|mock-gateway"} |= "ERROR" |
The k6 metric names are the ones Prometheus shows after a run. The p95 value is in seconds, and there is one line for each request name.
To see the test live, click the time picker at the top right and choose Last 15 minutes. The dashboard refreshes every 5 seconds. Start the k6 command, then watch.
Expected: during the run the virtual users line climbs to 30 and stays there. After the run all lines fall back.
How: what to read, in order
Read the dashboard top to bottom, with four questions.
- Did p95 rise? Look at the p95 panel. A flat line means users felt nothing.
- When? Note the time. Did it rise with the users, or suddenly later?
- Which resource rose at the same moment? Look straight down at CPU, memory and DB connections. Match the time.
- Any ERROR lines? The Errors row is empty when all is well.
My result, with USERS=30 and SCALE=0.1. Baseline is the idle value from above. Peak is the highest value in the 3 minute run:
| Measure | Baseline | Peak |
|---|---|---|
| api CPU | 0.013 cores | 0.78 cores |
| db CPU | 0.015 cores | 0.088 cores |
| DB connections | 12 | 12 |
p95 GET /menu | 11 ms | 23 ms |
p95 POST /orders | 17 ms | 78 ms |
The baseline p95 is the first value k6 reported, when only a few users were running. k6 had no idle value. I measured the CPU peak with a 30 second rate, so the 1 minute panel shows a slightly lower peak.
How I read it: p95 rose with the users, and the api CPU rose at the same moment, about 60 times higher. The database did far less work, and its connections never moved. So in this run the api was the busy part. No ERROR lines appeared, and 0% of requests failed.
How: the same view with JMeter
JMeter does not write to Prometheus. Its Backend Listener sends results to InfluxDB, as in Backend Listener. Grafana can have both data sources at once. A dashboard can show JMeter panels from InfluxDB beside the server panels from Prometheus. The reading steps above stay the same.
Tips
- Forgetting the two
K6_PROMETHEUS_RW_variables gives empty k6 panels. k6 runs fine and prints its summary, but it sends nothing. Check with thecurlquery above before you open the dashboard. - The laptop is both the load generator and the server here. k6 and the playground share the same CPU. Use the numbers to learn how to read cause and effect, not to size a real system.
- Set the Grafana time range to the test window. The default may show hours of idle line and flatten the part you care about.
- DB connections did not move in my run. That is a real result. The api keeps a fixed pool of connections, so the count stayed at 12. Raise
USERSif you want to see it change, and watch the line againstpg_settings_max_connections. - Database locks are not on the dashboard. Check them in Explore on
Prometheuswithsum(pg_locks_count{datname="playground"}). - Setup requests are left out of the p95 panel. The login and reset calls that k6 makes before the test have no
scenariolabel. The{scenario!=""}filter hides them. - Raise
USERSuntil something moves. With 10 users the api CPU barely changed. Record the value you used, so the next run is comparable. - Keep the dashboard JSON as a backup. Grafana here has no volume, so removing the container erases it.
Next: What Needs the Dev Team.