Logs and System Monitoring

Advanced
12 min

Logs and System Monitoring

When something is wrong on a server the question is always the same: what is it doing right now, and what did it do just before it broke? The answers live in the load average, memory and disk counters, and in the log files under /var/log. This lesson gives you a repeatable first-five-minutes routine, explains how to read the numbers those commands print, and shows how logs are organised and rotated so they do not fill the disk.

The Log Directory

Most services either write to the journal (previous lesson) or to a file under /var/log, and many do both.

| File or directory | Contents | |---|---| | /var/log/syslog (Debian) / /var/log/messages (RHEL) | general system messages | | /var/log/auth.log / /var/log/secure | logins, sudo, SSH attempts | | /var/log/kern.log | kernel messages | | /var/log/dpkg.log, /var/log/apt/ | package installs and upgrades | | /var/log/nginx/access.log, error.log | web server | | /var/log/journal/ | binary journal read with journalctl |

bash
sudo tail -f /var/log/syslog # follow sudo tail -f /var/log/nginx/*.log # several files at once sudo grep -i "failed password" /var/log/auth.log | tail sudo grep -h "Failed password" /var/log/auth.log | awk '{print $(NF-3)}' | sort | uniq -c | sort -rn | head sudo zgrep "error" /var/log/syslog.*.gz # search rotated, compressed logs last -n 10 # recent logins sudo lastb | head # recent failed logins

Log Rotation

Without rotation a busy access log grows until the disk is full. logrotate runs daily from a timer or cron job and rotates files according to /etc/logrotate.conf and drop-in files in /etc/logrotate.d/.

bash
# /etc/logrotate.d/api /srv/api/logs/*.log { daily rotate 14 compress delaycompress missingok notifempty copytruncate }

rotate 14 keeps two weeks, compress gzips old files, delaycompress leaves the most recent rotation uncompressed for easy reading, and copytruncate copies then empties the live file so an application that keeps it open need not be restarted. Test a config with sudo logrotate -d /etc/logrotate.d/api (dry run) and force it with -f.

Load, CPU and Processes

bash
uptime # 14:02:11 up 12 days, 3:41, 2 users, load average: 3.85, 2.10, 1.42 nproc # 4

Load average is the number of processes running or waiting to run (or blocked on disk I/O), averaged over 1, 5 and 15 minutes. Compare it with the core count: 3.85 on 4 cores is busy but coping; 12 on 4 cores means work is queueing. A rising 1-minute figure with a lower 15-minute one means the problem just started.

bash
top -bn1 | head -15 # one snapshot; %us user, %sy system, %wa waiting on I/O htop # interactive ps aux --sort=-%cpu | head -5 # top CPU consumers ps aux --sort=-%mem | head -5 # top memory consumers mpstat -P ALL 1 5 # per-core usage, 5 samples (sysstat package)

High %wa (I/O wait) means the CPU is idle waiting for disk — look at storage, not code.

Memory

bash
free -h # total used free shared buff/cache available # Mem: 7.8Gi 3.1Gi 0.4Gi 120Mi 4.3Gi 4.4Gi # Swap: 2.0Gi 0.3Gi 1.7Gi

The column that matters is available, not free. Linux uses spare RAM as disk cache (buff/cache) and gives it back instantly when a program needs it. Low available memory plus growing swap usage means real pressure. The kernel's OOM killer records its victims in the kernel log:

bash
sudo dmesg -T | grep -i "out of memory" journalctl -k | grep -i oom vmstat 1 5 # si/so columns: swap in/out per second; non-zero is bad

Disk I/O and Network

bash
iostat -xz 1 3 # per-device utilisation; %util near 100 is saturated iotop -o # processes doing I/O right now (needs root) df -h; df -i # space and inodes sudo ss -s # socket summary sudo ss -tulpn # listening ports and owning processes ip -s link # per-interface packet and error counters sar -n DEV 1 3 # network throughput (sysstat)

Install sysstat for iostat, mpstat and sar; sar also keeps a history so you can ask what the load was at 03:00 last night with sar -q -s 03:00 -e 04:00.

A Triage Routine

  1. uptime and nproc — is the machine overloaded?
  2. free -h and dmesg -T | tail — memory pressure or OOM kills?
  3. df -h — any filesystem full?
  4. top or ps --sort=-%cpu — which process?
  5. journalctl -p err --since "1 hour ago" and the service's own log — what did it say?
  6. ss -tulpn — is the service actually listening?

Working through these in order finds the cause of most incidents in a few minutes, and the same numbers are what you would put on a dashboard.

Common Mistakes

  • Reading free (the column) instead of available and concluding memory is exhausted.
  • Ignoring %wa and blaming the application for slowness caused by a saturated disk.
  • Deleting a rotated log with rm while a process still holds the live file open; use copytruncate or restart the service.
  • Grepping only the current log and missing yesterday's .1 or .gz file.
Quick Quiz
Question 1 of 3

A 2-core server shows a load average of `8.20, 7.90, 7.50`. What does this indicate?

Key Takeaways

  • System logs live in /var/log and the journal; tail -f, grep, zgrep, last and lastb cover most investigations.
  • logrotate keeps logs bounded; copytruncate and delaycompress are the options that matter for applications.
  • Judge load average against nproc; watch %wa for disk-bound systems.
  • Use the available column of free, and check dmesg for OOM kills when processes vanish.
  • iostat -xz, vmstat, ss -tulpn and sar complete the picture; follow the six-step triage routine.

Next lesson: Networking Commands: ip, ss, ping, curl and wget — inspect interfaces, ports and connectivity, and talk to web services from the shell.

Logs and System Monitoring - Linux & Command Line | CodeYourCraft | CodeYourCraft