Skip to main content

Set Up Alloy

Alloy is the collector. It reads the playground containers and sends what it finds to Prometheus and Loki.

Table of Contents


Why​

Alloy is the "sensors" in The Grafana Stack. Without it, Prometheus and Loki are empty.

Alloy needs the most setup of the four pieces, because it has to look inside Docker. The config is one file, monitoring/sample/alloy/config.alloy. It is split into blocks. Each block does one small job, and the blocks pass data to each other.


How: the config in four blocks​

Block 1: container metrics. The cadvisor block reads CPU and memory for every container from Docker. Network metrics are missing on OrbStack. The scrape block collects those numbers every 5 seconds and passes them to the next block.

// ---- Container metrics: cAdvisor reads CPU and memory per container ----
prometheus.exporter.cadvisor "containers" {
docker_host = "unix:///var/run/docker.sock"
docker_only = true
storage_duration = "5m"
// cAdvisor also needs the containerd socket. OrbStack keeps it here. Plain Linux usually
// uses the default /run/containerd/containerd.sock, so delete this line and mount that path.
containerd_host = "/run/docker/containerd/containerd.sock"
}

prometheus.scrape "containers" {
targets = prometheus.exporter.cadvisor.containers.targets
forward_to = [prometheus.relabel.service.receiver]
scrape_interval = "5s"
scrape_timeout = "5s"
}

Block 2: the service relabel. This block keeps only the containers of the playground project. Then it copies the compose service name (api, db, web, mock-gateway) into a short label called service. Every dashboard uses that label.

// Keep only the playground's containers (PLAYGROUND_PROJECT is passed in by docker run),
// then copy the compose service name (api, db, web, mock-gateway) into a short "service" label.
prometheus.relabel "service" {
forward_to = [prometheus.remote_write.local.receiver]
rule {
source_labels = ["container_label_com_docker_compose_project"]
regex = sys.env("PLAYGROUND_PROJECT")
action = "keep"
}
rule {
source_labels = ["container_label_com_docker_compose_service"]
target_label = "service"
}
}

Block 3: Postgres. This block asks the database for its own stats: connections, locks and transactions. It also holds prometheus.remote_write, the one place where all metrics leave Alloy for Prometheus.

// ---- Postgres: connections, locks, transactions ----
prometheus.exporter.postgres "db" {
data_source_names = ["postgresql://playground:playground@db:5432/playground?sslmode=disable"]
}

prometheus.scrape "db" {
targets = prometheus.exporter.postgres.db.targets
forward_to = [prometheus.remote_write.local.receiver]
scrape_interval = "5s"
scrape_timeout = "5s"
}

prometheus.remote_write "local" {
endpoint {
url = "http://prometheus:9090/api/v1/write"
}
}

Block 4: logs. This block lists the containers, keeps the playground ones, adds the same service label, and pushes every log line to Loki.

// ---- Container logs ----
discovery.docker "containers" {
host = "unix:///var/run/docker.sock"
}

discovery.relabel "logs" {
targets = discovery.docker.containers.targets
// Keep only the playground's containers, so other apps and the monitoring
// containers themselves are not collected.
rule {
source_labels = ["__meta_docker_container_label_com_docker_compose_project"]
regex = sys.env("PLAYGROUND_PROJECT")
action = "keep"
}
rule {
source_labels = ["__meta_docker_container_label_com_docker_compose_service"]
target_label = "service"
}
}

loki.source.docker "containers" {
host = "unix:///var/run/docker.sock"
targets = discovery.relabel.logs.output
forward_to = [loki.write.local.receiver]
}

loki.write "local" {
endpoint {
url = "http://loki:3100/loki/api/v1/push"
}
}

How: start Alloy​

Use the same terminal as before, so NET and PROJECT are still set. Prometheus and Loki must already be running.

# Mounts, smallest set that worked on OrbStack (macOS):
# docker.sock lets Alloy list containers and read their logs
# /sys cgroup files where CPU and memory are counted
# /var/lib/docker cAdvisor reads container disk usage here
# /run/docker/containerd containerd socket, needed by cAdvisor to name containers
# Docker Desktop and plain Linux were not tested: the containerd socket is usually
# /run/containerd/containerd.sock there, so mount that instead and remove the
# containerd_host line in config.alloy. No --privileged and no "-v /:/rootfs" were needed.
docker run -d --name alloy --network "$NET" -p 12345:12345 \
-e PLAYGROUND_PROJECT="$PROJECT" \
-v "$PWD/monitoring/sample/alloy/config.alloy:/etc/alloy/config.alloy:ro" \
-v /var/run/docker.sock:/var/run/docker.sock:ro \
-v /sys:/sys:ro \
-v /var/lib/docker/:/var/lib/docker:ro \
-v /run/docker/containerd:/run/docker/containerd:ro \
grafana/alloy:v1.20.1 \
run --server.http.listen-addr=0.0.0.0:12345 /etc/alloy/config.alloy

What each part is for:

PartWhy Alloy needs it
-e PLAYGROUND_PROJECTThe project name from page 05. The config uses it to keep only playground containers.
config.alloyThe config file from the last section.
docker.sockLets Alloy see the containers and read their logs.
/sysThe cgroup files where Linux counts CPU and memory.
/var/lib/dockercAdvisor reads container disk usage here.
/run/docker/containerdThe containerd socket. cAdvisor needs it to name the containers.

How: check it​

First, is Alloy up?

curl -s localhost:12345/-/ready

Expected:

Alloy is ready.

Open http://localhost:12345. Every component in the graph should be healthy.

Next, open http://localhost:9090/targets. The alloy target, which was DOWN on page 05, should now be UP.

Then check that per-container CPU arrives. Wait about 30 seconds after the start, then run:

curl -s localhost:9090/api/v1/query --data-urlencode 'query=rate(container_cpu_usage_seconds_total{service!=""}[1m])' | python3 -c 'import sys,json;print(sorted(x["metric"]["service"] for x in json.load(sys.stdin)["data"]["result"]))'

Expected:

['api', 'db', 'mock-gateway', 'web']

Last, check that Loki receives the logs:

curl -s localhost:3100/loki/api/v1/label/service/values

Expected:

{"status":"success","data":["api","db","mock-gateway","web"]}

You can also paste the query rate(container_cpu_usage_seconds_total{service!=""}[1m]) into the Prometheus UI at http://localhost:9090. You get one series per service.


Tips​

  • If curl does not say Alloy is ready., run docker logs alloy. The error names the block of config.alloy that failed.
  • If the CPU list is empty, look at the cAdvisor lines in docker logs alloy. Most often it is the containerd mount.
  • I tested this on OrbStack (macOS) only. On Docker Desktop and plain Linux the containerd socket is usually /run/containerd/containerd.sock. Mount that instead, and delete the containerd_host line in the config.
  • After you edit config.alloy, recreate the container: docker rm -f alloy, then the docker run again. A plain docker restart alloy can lose the mounted file when an editor replaces it instead of changing it in place.
  • After you recreate Alloy, you may see duplicate series for about 5 minutes. They fade away by themselves.
  • On the first start, Alloy also sends old log lines it finds. Loki refuses the ones older than its limit, so docker logs alloy shows a few entry too far behind errors. They are harmless.
  • On a real server, Alloy runs as a service, not as a container like in this lab. See Future / To Explore.

Next: Set Up Grafana.