Everyday Monitoring
A production system runs 24 hours a day, not only when someone is testing it. The team needs to know something is wrong before customers start to complain. This page shows how teams do that day to day.
Table of Contents
Why
Customers rarely report a problem quickly. Most just leave. If the first you hear of an issue is an angry message, the problem has already cost you orders.
Monitoring turns that around. The system tells the team first. Then the team can fix it, or at least say "we know, we are on it".
How: the four golden signals
With hundreds of possible measurements, where do you start? Google's SRE book suggests four signals that cover most problems (Monitoring Distributed Systems). Here they are for a pizza shop.
| Signal | Question it answers | Pizza shop example |
|---|---|---|
| Traffic | How busy are we? | Orders per minute |
| Errors | How many requests fail? | Failed payments |
| Latency | How slow are we? | How long the menu takes to load |
| Saturation | How full is the system? | Database connections in use, out of the maximum |
If you can answer these four for a system, you know most of what matters about its health.
How: dashboards
One always-on dashboard per system. Put the four golden signals at the top. Put the details below, for when someone needs to dig in.
| Who looks | When |
|---|---|
| Ops or on-call engineers | Every day, and when an alert fires |
| Testers | Before and after a release, to compare against normal |
How: alerts
An alert has three parts:
| Part | Example |
|---|---|
| A rule | Error log lines above 10 |
| A duration | In 5 minutes |
| Who to tell | The team channel |
Read together: "error logs above 10 in 5 minutes, tell the team channel". Alert on symptoms customers feel, such as failed orders or slow pages. Do not alert on every metric that moves.
Tips
- Too many alerts become ignored alerts. Keep only the ones that need a person to act.
- Give every alert an owner and a "what to do" note. An alert nobody owns gets muted.
- The Pizza Playground has no latency metric without help from the dev team. See What Needs the Dev Team.
Next: Monitoring During a Performance Test shows what to watch on the server while a load test runs.