When something is wrong on a server the question is always the same: what is it doing right now, and what did it do just before it broke? The answers live in the load average, memory and disk counters, and in the log files under /var/log. This lesson gives you a repeatable first-five-minutes routine, explains how to read the numbers those commands print, and shows how logs are organised and rotated so they do not fill the disk.
Most services either write to the journal (previous lesson) or to a file under /var/log, and many do both.
| File or directory | Contents |
|---|---|
| /var/log/syslog (Debian) / /var/log/messages (RHEL) | general system messages |
| /var/log/auth.log / /var/log/secure | logins, sudo, SSH attempts |
| /var/log/kern.log | kernel messages |
| /var/log/dpkg.log, /var/log/apt/ | package installs and upgrades |
| /var/log/nginx/access.log, error.log | web server |
| /var/log/journal/ | binary journal read with journalctl |
sudo tail -f /var/log/syslog # follow
sudo tail -f /var/log/nginx/*.log # several files at once
sudo grep -i "failed password" /var/log/auth.log | tail
sudo grep -h "Failed password" /var/log/auth.log | awk '{print $(NF-3)}' | sort | uniq -c | sort -rn | head
sudo zgrep "error" /var/log/syslog.*.gz # search rotated, compressed logs
last -n 10 # recent logins
sudo lastb | head # recent failed loginsWithout rotation a busy access log grows until the disk is full. logrotate runs daily from a timer or cron job and rotates files according to /etc/logrotate.conf and drop-in files in /etc/logrotate.d/.
# /etc/logrotate.d/api
/srv/api/logs/*.log {
daily
rotate 14
compress
delaycompress
missingok
notifempty
copytruncate
}rotate 14 keeps two weeks, compress gzips old files, delaycompress leaves the most recent rotation uncompressed for easy reading, and copytruncate copies then empties the live file so an application that keeps it open need not be restarted. Test a config with sudo logrotate -d /etc/logrotate.d/api (dry run) and force it with -f.
uptime
# 14:02:11 up 12 days, 3:41, 2 users, load average: 3.85, 2.10, 1.42
nproc # 4Load average is the number of processes running or waiting to run (or blocked on disk I/O), averaged over 1, 5 and 15 minutes. Compare it with the core count: 3.85 on 4 cores is busy but coping; 12 on 4 cores means work is queueing. A rising 1-minute figure with a lower 15-minute one means the problem just started.
top -bn1 | head -15 # one snapshot; %us user, %sy system, %wa waiting on I/O
htop # interactive
ps aux --sort=-%cpu | head -5 # top CPU consumers
ps aux --sort=-%mem | head -5 # top memory consumers
mpstat -P ALL 1 5 # per-core usage, 5 samples (sysstat package)High %wa (I/O wait) means the CPU is idle waiting for disk — look at storage, not code.
free -h
# total used free shared buff/cache available
# Mem: 7.8Gi 3.1Gi 0.4Gi 120Mi 4.3Gi 4.4Gi
# Swap: 2.0Gi 0.3Gi 1.7GiThe column that matters is available, not free. Linux uses spare RAM as disk cache (buff/cache) and gives it back instantly when a program needs it. Low available memory plus growing swap usage means real pressure. The kernel's OOM killer records its victims in the kernel log:
sudo dmesg -T | grep -i "out of memory"
journalctl -k | grep -i oom
vmstat 1 5 # si/so columns: swap in/out per second; non-zero is badiostat -xz 1 3 # per-device utilisation; %util near 100 is saturated
iotop -o # processes doing I/O right now (needs root)
df -h; df -i # space and inodes
sudo ss -s # socket summary
sudo ss -tulpn # listening ports and owning processes
ip -s link # per-interface packet and error counters
sar -n DEV 1 3 # network throughput (sysstat)Install sysstat for iostat, mpstat and sar; sar also keeps a history so you can ask what the load was at 03:00 last night with sar -q -s 03:00 -e 04:00.
uptime and nproc — is the machine overloaded?free -h and dmesg -T | tail — memory pressure or OOM kills?df -h — any filesystem full?top or ps --sort=-%cpu — which process?journalctl -p err --since "1 hour ago" and the service's own log — what did it say?ss -tulpn — is the service actually listening?Working through these in order finds the cause of most incidents in a few minutes, and the same numbers are what you would put on a dashboard.
free (the column) instead of available and concluding memory is exhausted.%wa and blaming the application for slowness caused by a saturated disk.rm while a process still holds the live file open; use copytruncate or restart the service..1 or .gz file.A 2-core server shows a load average of `8.20, 7.90, 7.50`. What does this indicate?
/var/log and the journal; tail -f, grep, zgrep, last and lastb cover most investigations.logrotate keeps logs bounded; copytruncate and delaycompress are the options that matter for applications.nproc; watch %wa for disk-bound systems.available column of free, and check dmesg for OOM kills when processes vanish.iostat -xz, vmstat, ss -tulpn and sar complete the picture; follow the six-step triage routine.Next lesson: Networking Commands: ip, ss, ping, curl and wget — inspect interfaces, ports and connectivity, and talk to web services from the shell.