Skip to main content

Monitoring in General

Monitoring means watching something continuously, so you notice a problem early and can explain it. A person decides what "healthy" looks like. The system does the watching.

This page is about monitoring in general, not any tool. Get this mental model first. The tools come later and are the same idea in software.

Table of Contents


Monitoring you already know​

ExampleWhat it watchesWhat it does when something is wrong
Car dashboardFuel level, engine temperatureLights a warning light
Hospital patient monitorHeart rate, oxygen levelSounds an alarm at the nurse station
Home smoke detectorSmoke in the airBeeps loudly
Phone battery widgetBattery percentageWarns you at 20%
Bank SMS alertEvery transaction on your cardSends an SMS when a large amount leaves

The shape of every monitoring setup​

Every example above does the same four steps. Software monitoring is no different.

StepQuestion it answersPatient monitor
CollectWhat do we measure?Sensors on the finger and chest read heart rate and oxygen
StoreWhere do the readings go?The monitor keeps the last hours of readings
ShowHow does a person see them?A screen draws the heart rate as a line
AlertWho is told when it goes wrong?An alarm sounds when oxygen drops below a set level

Without the store step, you cannot look back at what happened at 3am. Without the alert step, someone has to stare at the screen all day.


Two kinds of signal: metrics and logs​

MetricLog
What it isA number measured over timeA written line about one event
Patient monitor exampleHeart rate: 72"14:02 nurse gave medicine"
Good forSeeing trends and spotting "something changed"Finding out what exactly happened
Software exampleOrders per minute"ERROR payment failed for order 123"

Metrics tell you that something is wrong. Logs help you find out why. Traces are a third kind of signal. They are covered in Future / To Explore.


Why it matters to the business​

  • Find the real cause, not the symptom. A slow checkout may really be a database running out of connections. Monitoring shows where the time went.
  • Avoid outages after go-live. A memory line that climbs every hour warns you days before the server falls over.
  • Right-size servers. If the CPU never goes above 20%, the team is paying for a machine it does not need.
  • Evidence for go / no-go sign-off. "Response time stayed steady and errors stayed at zero during the test" is stronger than "it felt fine".
  • Less guessing and fewer re-runs. When a test fails, the recorded data tells you why. You do not repeat the run and hope to catch it.

Tips​

  • Monitoring does not fix anything. It shows you what is happening, so a person can act.
  • A dashboard nobody looks at is not monitoring. Decide who looks, and when.
  • Start with a few signals you understand, not hundreds you do not.

Next: Everyday Monitoring shows how a team uses this every day. Later pages set up Grafana, Prometheus, Loki and Alloy.