Text Processing: cut, sort, uniq, tr and wc

Intermediate
13 min

Text Processing: cut, sort, uniq, tr and wc

Most data you meet on a server is line-oriented text: logs, CSV exports, /etc/passwd, the output of other commands. The five tools in this lesson each do one small job — extract columns, order lines, remove duplicates, translate characters, count — and combine into pipelines that replace a surprising amount of scripting. You will finish able to answer questions like "which country has the most users?" in a single line.

cut: Extract Columns

cut selects fields (split by a delimiter) or character positions from each line.

bash
cut -d: -f1 /etc/passwd # first field, colon delimited: usernames cut -d, -f1,3 users.csv # fields 1 and 3 cut -d, -f2- users.csv # field 2 to the end cut -c1-10 app.log # first 10 characters of each line echo "2026-09-27T10:15:00" | cut -dT -f1 # 2026-09-27

cut uses a single-character delimiter and treats consecutive delimiters as separate empty fields, which makes it awkward for space-aligned output such as ls -l. For that, use awk (next lessons) or squeeze the spaces first with tr -s.

sort: Order Lines

By default sort orders lines alphabetically by the whole line. The flags select numeric, reverse, key-based and human-readable ordering.

| Flag | Effect | |---|---| | -n | numeric order (2 before 10) | | -r | reverse | | -k 2 | sort by field 2 (whitespace separated) | | -t , | field separator for -k | | -h | human-readable sizes (2K, 1G) | | -u | unique: drop duplicate lines | | -s | stable: keep input order for equal keys | | -o file | write output to file (safe even if it is the input) |

bash
sort -t: -k3 -n /etc/passwd | head # by UID, numeric sort -t, -k2,2 users.csv # by second field only du -sh * | sort -rh # biggest first sort -u emails.txt -o emails.txt # deduplicate in place

-k2,2 means "start and end at field 2". A bare -k2 means "from field 2 to the end of the line", which is rarely what you intend.

uniq: Collapse Duplicates

uniq removes adjacent duplicate lines, so input almost always needs to be sorted first. Its counting mode is the basis of many ad-hoc reports.

bash
sort words.txt | uniq # distinct lines sort words.txt | uniq -c # prefix each with its count sort words.txt | uniq -d # only lines that appear more than once sort words.txt | uniq -u # only lines that appear exactly once

The classic frequency-table pipeline:

bash
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -10
bash
4821 203.0.113.7 1930 198.51.100.24 412 192.0.2.99

Read it left to right: extract the field, sort so duplicates touch, count them, sort by count descending, show the top ten.

tr: Translate or Delete Characters

tr maps characters in the first set to the corresponding characters in the second. It works on characters, never on words, and reads only from standard input.

bash
echo "Hello World" | tr 'a-z' 'A-Z' # HELLO WORLD tr -d '\r' < windows.txt > unix.txt # remove carriage returns echo "a b c" | tr -s ' ' # squeeze repeats: a b c echo "user:pass:home" | tr ':' '\n' # split onto lines echo "phone: 555-0142" | tr -cd '0-9\n' # -c complement: keep only digits tr '[:upper:]' '[:lower:]' < README.md

wc: Count

bash
wc -l app.log # lines wc -w essay.txt # words wc -c image.png # bytes wc -m unicode.txt # characters ls | wc -l # entries in the current directory grep -c ERROR app.log # same as grep ERROR app.log | wc -l but faster

The output of wc -l file includes the filename; use wc -l < file to get just the number, which matters in scripts.

Putting Them Together

bash
# Distinct file extensions in a project, most common first find . -type f | grep -oE '\.[a-zA-Z0-9]+$' | sort | uniq -c | sort -rn | head # Which hours of the day had the most 500 errors grep ' 500 ' access.log | cut -d[ -f2 | cut -d: -f2 | sort | uniq -c # Lines present in new.txt but not in old.txt sort old.txt > /tmp/o; sort new.txt > /tmp/n; comm -13 /tmp/o /tmp/n

comm compares two sorted files column by column: -1 hides lines only in the first, -2 only in the second, -3 common lines. paste and join (merge columns and join on a key) round out the family when you need them.

Common Mistakes

  • Running uniq without sort first and missing duplicates that are not adjacent.
  • Using sort without -n on numbers, which puts 10 before 9.
  • Expecting cut -d' ' to handle multiple spaces; squeeze with tr -s ' ' or use awk.
  • Redirecting sort file > file, which truncates the file before reading it. Use -o.
Quick Quiz
Question 1 of 3

Which pipeline lists the distinct values in column 3 of a CSV, with counts?

Key Takeaways

  • cut -d<delim> -f<n> extracts fields; -c extracts character ranges.
  • sort -n for numbers, -h for sizes, -k for a specific field, -u to deduplicate, -o to write in place.
  • uniq -c counts adjacent duplicates, so always sort first; sort | uniq -c | sort -rn is the frequency idiom.
  • tr translates, deletes (-d) or squeezes (-s) characters, not words.
  • wc -l counts lines; redirect the file in to get a bare number.

Next lesson: sed: Stream Editing and Find-and-Replace — edit text in a pipeline or in place without opening an editor.

Text Processing: cut, sort, uniq, tr and wc - Linux & Command Line | CodeYourCraft | CodeYourCraft