awk: Column-Based Text Processing

Intermediate
14 min

awk: Column-Based Text Processing

Where cut extracts columns and sed rewrites lines, awk is a small programming language built for text that comes in columns. It splits every line into fields, lets you filter with conditions, keeps variables and arrays across lines, and prints reports at the end. It is installed on every Unix system, which makes it the safest choice for data crunching on a server where you cannot install Python. After this lesson you will be able to write one-liners that filter, aggregate and reformat any columnar text.

The awk Model

An awk program is a list of pattern { action } pairs. For every input line, each pattern is tested; when it matches, the action runs. Omit the pattern and the action runs on every line; omit the action and the matching line is printed.

bash
awk '{ print }' file.txt # print every line (like cat) awk '/ERROR/' app.log # print matching lines (like grep) awk '/ERROR/ { print $1, $2 }' app.log # print fields 1 and 2 of matching lines

Fields and Built-in Variables

awk splits each line on runs of whitespace by default. Fields are $1, $2, ..., and $0 is the whole line.

| Variable | Meaning | |---|---| | $0 | the entire current line | | $1, $2, $NF | first field, second field, last field | | NF | number of fields on this line | | NR | current line number across all input | | FNR | line number within the current file | | FS / OFS | input / output field separator | | FILENAME | name of the current input file |

bash
ls -l | awk '{print $NF}' # just filenames from a long listing awk '{print NR, NF}' file.txt # line number and its field count awk 'NR==5' file.txt # only line 5 awk 'NR>=10 && NR<=20' file.txt # a range of lines awk '{print $(NF-1)}' data.txt # second-to-last field

Changing the Separator

Use -F for CSV, colon-separated or tab-separated data. Set OFS to control how printed fields are joined.

bash
awk -F: '{print $1, $7}' /etc/passwd # username and shell awk -F, 'NR>1 {print $2}' users.csv # skip the header row awk -F'\t' '{print $3}' data.tsv awk -F, -v OFS='\t' '{print $1, $3}' users.csv # CSV in, TSV out awk 'BEGIN{FS=","; OFS=" | "} {$1=$1; print}' users.csv # rebuild line with new OFS

The $1=$1 trick forces awk to rebuild $0 using OFS, which is how you convert delimiters in one step.

Conditions and Comparisons

Patterns can be regular expressions, comparisons or combinations of both.

bash
awk -F: '$3 > 1000' /etc/passwd # lines where field 3 is greater than 1000 awk -F: '$3 >= 1000 && $7 != "/usr/sbin/nologin" {print $1}' /etc/passwd awk '$9 ~ /^5/' access.log # field 9 matches a regex (5xx status) awk '$9 !~ /^2/' access.log # does not match awk 'length($0) > 80' script.sh # long lines awk '$5 == "root" {print $2}' <(ps aux)

BEGIN, END and Aggregation

BEGIN runs before any input, END after the last line. Variables need no declaration and start as zero or empty, which makes sums and counts trivial. Associative arrays (dictionaries) are the key to group-by reports.

bash
# Total and average of column 2 awk '{sum += $2} END {print "total:", sum, "avg:", sum / NR}' numbers.txt # Requests per IP, then sort descending awk '{hits[$1]++} END {for (ip in hits) print hits[ip], ip}' access.log | sort -rn | head # Bytes served per status code, in MB awk '{bytes[$9] += $10} END {for (s in bytes) printf "%s %.1f MB\n", s, bytes[s]/1048576}' access.log # Largest value in column 3 and which line it came from awk 'NR==1 || $3 > max {max = $3; line = $0} END {print line}' data.txt

printf uses C-style formats: %s string, %d integer, %.2f two decimals, %-10s left-aligned width 10. Unlike print, it does not add a newline unless you include \n.

Formatting Output

bash
awk -F, 'NR>1 {printf "%-15s %-25s %s\n", $1, $2, $3}' users.csv awk '{printf "%5d %s\n", NR, $0}' file.txt # numbered lines, aligned df -h | awk 'NR>1 && $5+0 > 80 {print $6, "is", $5, "full"}' # $5+0 converts "85%" to 85

Adding +0 to a string coerces it to a number, discarding trailing characters such as %.

Scripts in Files

When a one-liner grows, put it in a file and run it with -f, or use a shebang.

bash
#!/usr/bin/awk -f BEGIN { FS = ","; print "Report" } NR > 1 && $3 == "IN" { count++; names = names $1 ", " } END { print count, "Indian users:", names }
bash
chmod +x report.awk && ./report.awk users.csv

Common Mistakes

  • Using double quotes around the program in the shell, which makes $1 a shell variable. Always single-quote the awk program.
  • Forgetting -F and getting the whole CSV line in $1.
  • Comparing numbers stored as strings: "10" < "9" is true as strings; force numeric with +0.
  • Expecting uniq-style ordering from for (k in arr); array iteration order is undefined, so pipe to sort.
Quick Quiz
Question 1 of 3

What does `awk -F: '{print $1}' /etc/passwd` print?

Key Takeaways

  • awk 'pattern { action }' runs the action on every line where the pattern matches; either part is optional.
  • Fields are $1..$NF; NR is the line number; -F sets the separator and OFS controls output joins.
  • Conditions can be comparisons ($3 > 1000) or regex matches ($9 ~ /^5/).
  • BEGIN/END with variables and associative arrays turn awk into a group-by and sum engine.
  • printf formats aligned reports; add +0 to force numeric comparison.

Next lesson: File Permissions and chmod — understand read, write and execute bits and how to set them.

awk: Column-Based Text Processing - Linux & Command Line | CodeYourCraft | CodeYourCraft