Counting systemd service restarts: NRestarts and the D-Bus API behind it

A while ago I debugged a resource leak caused by a service stuck in a restart loop. Each restart left a process behind, and after a few minutes the host had over a thousand copies of it. Nothing alerted on the loop itself. We noticed CPU and memory climbing, and only found the restart loop after reading the journal. Following up on the incident, we decided to emit a metric for how many times each service has restarted since boot, and alert when the number climbs. The metric has since helped us identify six bugs that caused unexpected service restarts, including the port conflict described in Apply sysctl before systemd. This post covers how systemd (v235+) tracks that restart count, how it is read, what D-Bus has to do with it, and how to export the value in Prometheus format. Everything below was run on Amazon Linux 2023, which ships systemd 252 as of today.

systemd NRestarts

The v235 release notes (October 2017) introduced NRestarts:

1* For each service unit a restart counter is now kept: it is increased
2  each time the service is restarted due to Restart=, and may be
3  queried using "systemctl show -p NRestarts …".

How the counter behaves

The implementation is a single unsigned integer, n_restarts, in the Service struct. It increments when systemd schedules an automatic restart. service_enter_restart() is called when a service has exited and Restart= says to bring it back. It enqueues a restart job, bumps the counter, and logs the line you see in the journal:

 1/* Count the jobs we enqueue for restarting. This counter is maintained as long as the unit isn't fully
 2 * stopped, i.e. as long as it remains up or remains in auto-start states. The user can reset the counter
 3 * explicitly however via the usual "systemctl reset-failure" logic. */
 4s->n_restarts ++;
 5s->flush_n_restarts = false;
 6
 7log_unit_struct(UNIT(s), LOG_INFO,
 8                "MESSAGE_ID=" SD_MESSAGE_UNIT_RESTART_SCHEDULED_STR,
 9                LOG_UNIT_INVOCATION_ID(UNIT(s)),
10                LOG_UNIT_MESSAGE(UNIT(s),
11                                 "Scheduled restart job, restart counter is at %u.", s->n_restarts),
12                "N_RESTARTS=%u", s->n_restarts);
13
14/* Notify clients about changed restart counter */
15unit_add_to_dbus_queue(UNIT(s));

The journal entry carries the count as a structured field, N_RESTARTS, and a fixed MESSAGE_ID, so you can filter for restart events by the MESSAGE_ID efficiently. And the last line queues a D-Bus PropertiesChanged signal, so anyone subscribed to the systemd unit finds out immediately. A few more rules about the counter:

  1. Manual restarts do not count. systemctl restart goes through a different path. The service is stopped, and because the stop was explicit, systemd does not enter service_enter_restart(). The counter is not incremented. In fact it gets reset, which is the next rule.

  2. NRestarts resets lazily on the next non-automatic start. When a service goes fully inactive without a pending automatic restart, service_enter_dead() does not zero the counter right away:

    1} else
    2        /* If we shan't restart, then flush out the restart counter. But don't do that immediately, so that the
    3         * user can still introspect the counter. Do so on the next start. */
    4        s->flush_n_restarts = true;
    

    The zeroing happens at the top of the next service_start(). So after a service dies for good, or after systemctl stop, you can still read how many times it had restarted. The next systemctl start wipes it.

  3. systemctl reset-failed zeroes it. service_reset_failed() sets n_restarts to 0 along with the failure result.

  4. It survives daemon-reexec. The value is serialized as "n-restarts" and restored when PID 1 re-executes itself, for example during a systemd package upgrade. It does not survive a reboot, since it lives only in PID 1's memory. "Restarts since boot" is the right way to think about it.

Seeing it on a flapping service

You can follow along on an AL2023 EC2 instance (recommended), or in a privileged Docker container if you don't want to start a new instance.

1docker run -d --name sysd --privileged --cgroupns=host \
2  -v /sys/fs/cgroup:/sys/fs/cgroup:rw --tmpfs /run --tmpfs /tmp \
3  public.ecr.aws/amazonlinux/amazonlinux:2023 \
4  bash -c 'dnf install -y -q systemd dbus-daemon >/dev/null 2>&1; exec /sbin/init'
5sleep 20
6docker exec -it sysd bash

Create a service that exits two seconds after it starts and asks systemd to restart it forever:

 1# /etc/systemd/system/flappy.service
 2[Unit]
 3Description=A service that exits shortly after start
 4
 5[Service]
 6ExecStart=/bin/bash -c "sleep 2; exit 1"
 7Restart=always
 8RestartSec=1
 9StartLimitIntervalSec=0
10
11[Install]
12WantedBy=multi-user.target

StartLimitIntervalSec=0 disables the start rate limit so the loop does not get cut off after five attempts. Start it and let it flap for a bit:

1# systemctl daemon-reload && systemctl enable --now flappy.service
2# sleep 12
3# systemctl show flappy.service -p NRestarts -p ActiveState -p SubState -p Result
4Result=success
5NRestarts=3
6ActiveState=activating
7SubState=auto-restart

The journal shows the same counter:

1# journalctl -u flappy.service -o short-monotonic | tail -7
2[ 1296.756035] ip-172-31-0-127.us-west-2.compute.internal systemd[1]: Stopped flappy.service - A service that exits shortly after start.
3[ 1296.771829] ip-172-31-0-127.us-west-2.compute.internal systemd[1]: Started flappy.service - A service that exits shortly after start.
4[ 1298.775225] ip-172-31-0-127.us-west-2.compute.internal systemd[1]: flappy.service: Main process exited, code=exited, status=1/FAILURE
5[ 1298.775517] ip-172-31-0-127.us-west-2.compute.internal systemd[1]: flappy.service: Failed with result 'exit-code'.
6[ 1300.005497] ip-172-31-0-127.us-west-2.compute.internal systemd[1]: flappy.service: Scheduled restart job, restart counter is at 19.
7[ 1300.006005] ip-172-31-0-127.us-west-2.compute.internal systemd[1]: Stopped flappy.service - A service that exits shortly after start.
8[ 1300.031798] ip-172-31-0-127.us-west-2.compute.internal systemd[1]: Started flappy.service - A service that exits shortly after start.

And the structured fields are there if you ask for JSON output and filter by the message ID:

1# journalctl -u flappy.service -o json MESSAGE_ID=5eb03494b6584870a536b337290809b3 | tail -1 | jq '{MESSAGE, N_RESTARTS, UNIT, INVOCATION_ID}'
2{
3  "MESSAGE": "flappy.service: Scheduled restart job, restart counter is at 45.",
4  "N_RESTARTS": "45",
5  "UNIT": "flappy.service",
6  "INVOCATION_ID": "cdefd54fee4549bdb372a1fc7b23a7b2"
7}

Now try the reset:

 1# systemctl show flappy -p NRestarts
 2NRestarts=13
 3
 4# systemctl stop flappy && systemctl show flappy -p NRestarts -p ActiveState
 5NRestarts=13
 6ActiveState=inactive
 7
 8# systemctl start flappy && systemctl show flappy -p NRestarts
 9NRestarts=0
10
11# sleep 7 && systemctl show flappy -p NRestarts
12NRestarts=2
13
14# systemctl restart flappy && systemctl show flappy -p NRestarts
15NRestarts=0
16
17# sleep 10 && systemctl show flappy -p NRestarts && systemctl reset-failed flappy && systemctl show flappy -p NRestarts
18NRestarts=3
19NRestarts=0

Stop keeps the value for inspection, the next start flushes it, a manual restart flushes it, and reset-failed zeroes it.

There is a gotcha with the "kept after stop" behavior. If the unit has no [Install] section and nothing else references it, systemd garbage-collects the unit from memory as soon as it goes inactive. The next systemctl show loads a fresh copy from disk, and NRestarts reads 0. I hit this on the first attempt and spent a few minutes confused. With WantedBy=multi-user.target, the active target holds a reference and the unit stays loaded, so the counter is visible after stop as shown above.

How systemctl reads it: D-Bus

systemctl show is not reading a file. There is no /proc or /sys entry for unit state. systemctl is a D-Bus client, and every value it prints came out of a D-Bus property read against PID 1.

What is D-Bus?

D-Bus is an inter-process communication system from freedesktop.org, released in 2003, seven years before systemd. It came out of the desktop world, where GNOME and KDE each had incompatible IPC (CORBA and DCOP), and D-Bus was the neutral replacement. It is now how most system daemons on Linux expose their control interfaces: NetworkManager, logind, udisks, and systemd itself. The pieces:

A message bus. A daemon sits in the middle: historically dbus-daemon, and on AL2023 and most current distros the faster reimplementation dbus-broker. Processes connect over a Unix socket and the daemon routes messages between them. There is one system bus per machine at /run/dbus/system_bus_socket, and one session bus per logged-in user for desktop applications.

Names. A connection gets a unique name like :1.9 when it connects, and can additionally claim well-known names like org.freedesktop.systemd1. Clients address services by well-known name so they never need to know a PID.

Objects, interfaces, members. A service exposes objects at paths like /org/freedesktop/systemd1/unit/flappy_2eservice. Each object implements interfaces, such as org.freedesktop.systemd1.Unit and org.freedesktop.systemd1.Service. An interface has methods you can call, properties you can read (and sometimes write), and signals you can subscribe to. It is a remote object model.

Typed messages. Values on the wire carry type signatures: s is a string, u is uint32, o is an object path, a{sv} is a dictionary of string to variant. The protocol is binary.

Introspection. Every object answers org.freedesktop.DBus.Introspectable.Introspect with an XML description of its interfaces, so generic tools can explore any service without a schema in hand.

Standard interfaces. org.freedesktop.DBus.Properties gives you Get, GetAll, and the PropertiesChanged signal for any object. This is the interface systemctl show uses.

D-Bus is a specification with many implementations: libdbus, GDBus (GLib), QtDBus, sd-bus (systemd's own, part of libsystemd), godbus for Go, zbus for Rust, dbus-python.

systemd's D-Bus API

systemd uses D-Bus as its control protocol. PID 1 claims org.freedesktop.systemd1 on the system bus and exposes a Manager object at /org/freedesktop/systemd1 plus one object per loaded unit and per job. When you run systemctl restart flappy, systemctl calls Manager.RestartUnit("flappy.service", "replace") and gets back a job object path. Everything systemctl does is a method call or property read you can make yourself.

There is one exception. PID 1 starts before the bus daemon does, and it has to stay controllable during early boot and late shutdown when there is no bus. So systemd also listens on a private socket, /run/systemd/private, which speaks the D-Bus wire protocol point-to-point with no daemon in between. systemctl, when running as root, talks to this socket rather than the bus. That is why node_exporter's documentation tells you to mount /run/systemd/private into its container rather than the system bus socket. The exporter needs one of the two, and the private socket works even when the container has no access to the bus.

Hands-on with busctl

busctl ships with systemd and is the fastest way to see all of this. Run these as root on the host with flappy.service running.

Who is on the system bus:

 1# busctl list
 2NAME                       PID PROCESS         USER            CONNECTION    UNIT                     SESSION DESCR>
 3:1.0                      1647 systemd-resolve systemd-resolve :1.0          systemd-resolved.service -       -
 4:1.1                      1773 systemd-inhibit root            :1.1          acpid.service            -       -
 5:1.2                      1781 systemd-homed   root            :1.2          systemd-homed.service    -       -
 6:1.3                      1782 systemd-logind  root            :1.3          systemd-logind.service   -       -
 7:1.4                      1787 systemd-network systemd-network :1.4          systemd-networkd.service -       -
 8:1.43                     2640 systemd         ec2-user        :1.43         [email protected]        -       -
 9:1.49                     3518 busctl          root            :1.49         session-10.scope         10      -
10:1.5                         1 systemd         root            :1.5          init.scope               -       -
11:1.6                      1785 dbus-broker-lau root            :1.6          dbus-broker.service      -       -
12org.freedesktop.DBus         1 systemd         root            -             init.scope               -       -
13org.freedesktop.home1     1781 systemd-homed   root            :1.2          systemd-homed.service    -       -
14org.freedesktop.hostname1    - -               -               (activatable) -                        -       -
15org.freedesktop.locale1      - -               -               (activatable) -                        -       -
16org.freedesktop.login1    1782 systemd-logind  root            :1.3          systemd-logind.service   -       -
17org.freedesktop.network1  1787 systemd-network systemd-network :1.4          systemd-networkd.service -       -
18org.freedesktop.oom1         - -               -               (activatable) -                        -       -
19org.freedesktop.portable1    - -               -               (activatable) -                        -       -
20org.freedesktop.resolve1  1647 systemd-resolve systemd-resolve :1.0          systemd-resolved.service -       -
21org.freedesktop.systemd1     1 systemd         root            :1.5          init.scope               -       -
22org.freedesktop.timedate1    - -               -               (activatable) -                        -       -
23org.freedesktop.timesync1    - -               -               (activatable) -                        -       -

PID 1 is a regular bus client alongside resolved, networkd, and logind. Ask the Manager for the object path of our unit. Note the escaping: a dot in a unit name becomes _2e in the path.

1# busctl call org.freedesktop.systemd1 /org/freedesktop/systemd1 \
2    org.freedesktop.systemd1.Manager GetUnit s flappy.service
3o "/org/freedesktop/systemd1/unit/flappy_2eservice"

Introspect the unit object. The full output is a few hundred lines; the interesting subset:

 1# busctl introspect org.freedesktop.systemd1 /org/freedesktop/systemd1/unit/flappy_2eservice
 2NAME                                TYPE      SIGNATURE  RESULT/VALUE  FLAGS
 3org.freedesktop.DBus.Introspectable interface -          -             -
 4.Introspect                         method    -          s             -
 5org.freedesktop.DBus.Properties     interface -          -             -
 6.Get                                method    ss         v             -
 7.GetAll                             method    s          a{sv}         -
 8.PropertiesChanged                  signal    sa{sv}as   -             -
 9org.freedesktop.systemd1.Service    interface -          -             -
10.ExecMainPID                        property  u          249           emits-change
11.NRestarts                          property  u          7             emits-change
12.Restart                            property  s          "always"      const
13.Result                             property  s          "success"     emits-change
14org.freedesktop.systemd1.Unit       interface -          -             -
15.ResetFailed                        method    -          -             -
16.Restart                            method    s          o             -
17.Start                              method    s          o             -
18.Stop                               method    s          o             -
19.ActiveState                        property  s          "active"      emits-change
20.SubState                           property  s          "running"     emits-change

NRestarts is a uint32 property on the Service interface with the emits-change flag, which is the SD_BUS_VTABLE_PROPERTY_EMITS_CHANGE flag from the vtable in src/core/dbus-service.c:

1SD_BUS_PROPERTY("NRestarts", "u", bus_property_get_unsigned, offsetof(Service, n_restarts), SD_BUS_VTABLE_PROPERTY_EMITS_CHANGE),

Read it directly. This is exactly what systemctl show -p NRestarts does:

1# busctl get-property org.freedesktop.systemd1 /org/freedesktop/systemd1/unit/flappy_2eservice \
2    org.freedesktop.systemd1.Service NRestarts
3u 7

Restart it the way systemctl does, by calling a method on the Manager:

1# busctl call org.freedesktop.systemd1 /org/freedesktop/systemd1 \
2    org.freedesktop.systemd1.Manager RestartUnit ss flappy.service replace
3o "/org/freedesktop/systemd1/job/2052"

Now the part that polling exporters leave on the table. Because NRestarts emits change signals, you can watch restarts happen instead of asking. Subscribe to signals from the unit's object path and let it flap for one cycle:

 1# busctl monitor org.freedesktop.systemd1 \
 2    --match "type=signal,path=/org/freedesktop/systemd1/unit/flappy_2eservice"
 3...
 4‣ Type=signal  Sender=:1.9  Path=/org/freedesktop/systemd1/unit/flappy_2eservice
 5  Interface=org.freedesktop.DBus.Properties  Member=PropertiesChanged
 6  MESSAGE "sa{sv}as" {
 7          STRING "org.freedesktop.systemd1.Service";
 8          ARRAY "{sv}" {
 9                  DICT_ENTRY "sv" { STRING "MainPID";   VARIANT "u" { UINT32 0; }; };
10                  DICT_ENTRY "sv" { STRING "Result";    VARIANT "s" { STRING "exit-code"; }; };
11                  DICT_ENTRY "sv" { STRING "NRestarts"; VARIANT "u" { UINT32 4; }; };
12                  ...
13‣ Type=signal  Sender=:1.9  Path=/org/freedesktop/systemd1/unit/flappy_2eservice
14  Interface=org.freedesktop.DBus.Properties  Member=PropertiesChanged
15  MESSAGE "sa{sv}as" {
16          STRING "org.freedesktop.systemd1.Unit";
17          ARRAY "{sv}" {
18                  DICT_ENTRY "sv" { STRING "ActiveState"; VARIANT "s" { STRING "activating"; }; };
19                  DICT_ENTRY "sv" { STRING "SubState";    VARIANT "s" { STRING "auto-restart"; }; };
20                  ...

One restart cycle produces a burst of PropertiesChanged signals: Service properties (Result flips to exit-code, NRestarts ticks to 4, MainPID goes to 0), then Unit properties (ActiveState to activating, SubState to auto-restart), then JobNew on the Manager, then another round when the new process is up. A daemon that wanted to react to a restart in real time would subscribe to this and never poll. Neither node_exporter nor the OpenTelemetry systemd receiver does; both read the property on each scrape, which is fine for a metric.

The private socket: /run/systemd/private

To see why /run/systemd/private matters, stop the bus daemon:

 1# systemctl stop dbus-broker.service dbus.socket
 2
 3# busctl get-property org.freedesktop.systemd1 /org/freedesktop/systemd1/unit/flappy_2eservice \
 4    org.freedesktop.systemd1.Service NRestarts
 5Failed to connect to bus: Connection refused
 6
 7# systemctl show flappy -p NRestarts
 8NRestarts=10
 9
10# systemctl start dbus.socket dbus-broker.service

busctl needs the bus. systemctl kept working because, as root, it never used the bus in the first place. bus_connect_system_systemd() in src/shared/bus-util.c dials /run/systemd/private directly whenever euid is 0, with the comment "If we are root then let's talk directly to the system instance, instead of going via the bus", and only falls back to the system bus if that fails. The go-systemd library goes the other way around: dbus.NewWithContext() tries the system bus first and, if that fails and the caller is root, dials unix:path=/run/systemd/private. Its source notes that it skips the Hello handshake on the private connection, because there is no bus daemon on the other end to say hello to. This is also why pointing busctl at the private socket with --address fails with "Invalid request descriptor": busctl always sends Hello, PID 1 answers with an unknown-method error, and sd-bus maps that to EBADR.

Reading it from Go

This is the core of what node_exporter does for the restart metric, stripped to a dozen lines:

 1package main
 2
 3import (
 4	"context"
 5	"fmt"
 6	"os"
 7
 8	"github.com/coreos/go-systemd/v22/dbus"
 9)
10
11func main() {
12	ctx := context.Background()
13	// System bus first; falls back to /run/systemd/private when root.
14	conn, err := dbus.NewWithContext(ctx)
15	if err != nil {
16		panic(err)
17	}
18	defer conn.Close()
19
20	for _, unit := range os.Args[1:] {
21		p, err := conn.GetUnitTypePropertyContext(ctx, unit, "Service", "NRestarts")
22		if err != nil {
23			fmt.Printf("%s: %v\n", unit, err)
24			continue
25		}
26		fmt.Printf("%s NRestarts=%d\n", unit, p.Value.Value().(uint32))
27	}
28}
1# ./nrestarts flappy.service systemd-logind.service nonexistent.service
2flappy.service NRestarts=18
3systemd-logind.service NRestarts=0
4nonexistent.service NRestarts=0

GetUnitTypePropertyContext computes the object path from the unit name and calls org.freedesktop.DBus.Properties.Get on it. Note the last line: asking about a unit that does not exist does not error. systemd lazily loads a not-found placeholder unit for the path and returns default values. Check LoadState if you care.

Exporting to Prometheus: node_exporter or otel

node_exporter. The systemd collector is off by default, and the restart metric is a further opt-in because it costs one extra property read per unit on every scrape (the rest of the collector's data comes from a single ListUnits call):

1--collector.systemd
2--collector.systemd.unit-include="(kubelet|containerd|flappy)\.service"
3--collector.systemd.enable-restarts-metrics

This gives you node_systemd_service_restart_total{name="flappy.service"}. When running as a container or DaemonSet, mount /run/systemd/private read-write (it is a socket, so the exporter needs to connect to it) and run as root. An alert on increase(node_systemd_service_restart_total[15m]) > 2 would have caught the restart loop from the incident in the first couple of minutes.

OpenTelemetry Collector. The contrib distribution has a systemd receiver that reads the same properties over the same go-systemd library. The restart metric is optional and, at the time of writing, marked development stability:

1receivers:
2  systemd:
3    units: ["kubelet.service", "containerd.service"]
4    metrics:
5      systemd.service.restarts:
6        enabled: true
7exporters:
8  prometheus:
9    endpoint: 0.0.0.0:9464

It comes out as systemd_service_restarts_total{systemd_unit_name="kubelet.service"}.

Whichever you use, remember the reset rules. The counter drops to zero on reboot, on systemctl start after a stop, on systemctl restart, and on reset-failed. So alert on rate or increase over a window, never on the raw value.