Text Processing & Shell Tools

0%
20 mintextshellgrepawksed

Text Processing & Shell Tools

Learn text processing & shell tools concepts for DevOps interviews

08 — Text Processing & Shell Tools

Fluency here is what separates people who use Linux from people who are fast on Linux. Interviewers give you a log file and ask “count the top 10 IPs by request.” This file gives you grep/sed/awk, pipes/redirects, find, and the sort/uniq/cut toolkit to answer instantly.


1. Pipes and redirection (the foundation)

Linux tools are small and composable; you chain them with pipes and control their input/output with redirection.

cmd1 | cmd2            # pipe: stdout of cmd1 → stdin of cmd2
cmd > file            # redirect stdout → file (overwrite)
cmd >> file           # append stdout → file
cmd 2> err.log        # redirect stderr (fd 2)
cmd > out.log 2>&1    # stdout AND stderr → same file  (order matters!)
cmd &> all.log        # bash shorthand for the above
cmd < input.txt       # feed a file as stdin
cmd1 |& cmd2          # pipe stdout+stderr
: > file              # truncate a file to zero bytes (frees space, keeps inode)

The three standard streams: stdin (fd 0), stdout (fd 1), stderr (fd 2).

⚠️ Order matters: > out 2>&1 sends stderr to where stdout currently points (the file); 2>&1 >out sends stderr to the original stdout (terminal) then stdout to the file — a classic gotcha. 🎯 Interview signal: explaining why 2>&1 must come after > file shows you actually understand fd duplication.

tee                    # write to a file AND pass through the pipe
cmd | tee out.log | grep ERROR      # save everything, also filter live

2. grep — search text

grep "ERROR" app.log             # lines containing ERROR
grep -i error app.log            # case-insensitive
grep -r "TODO" src/              # recursive
grep -v "DEBUG" app.log          # INVERT: lines NOT matching
grep -c "ERROR" app.log          # count matching lines
grep -n "ERROR" app.log          # show line numbers
grep -E "warn|error|fatal" log   # extended regex (alternation)
grep -A3 -B2 "panic" log         # 3 lines After, 2 Before the match (context)
grep -o "[0-9]\+" log            # print only the matched part
ss -tlnp | grep :443             # filter another command's output

🎯 grep -A/-B/-C (context lines) and grep -v (invert) are the ones that impress — most people only know plain grep.


3. sed — stream editor (find/replace, edit lines)

sed 's/foo/bar/' file            # replace FIRST foo per line
sed 's/foo/bar/g' file           # replace ALL (global)
sed -i 's/foo/bar/g' file        # edit the file IN PLACE (⚠️ no undo — back up first)
sed -i.bak 's/foo/bar/g' file    # in place, but keep a .bak backup
sed -n '10,20p' file             # print only lines 10–20
sed '/^#/d' file                 # delete comment lines
sed '/^$/d' file                 # delete blank lines
sed 's/[0-9]\+/N/g' file         # replace all numbers with N

⚠️ sed -i on a production config with a wrong regex is a real outage cause — always test without -i first (it prints to stdout), or use -i.bak.


4. awk — field-based processing (the power tool)

awk splits each line into fields ($1, $2, … ; $0 = whole line) and runs a pattern { action }. Default field separator is whitespace.

awk '{print $1}' access.log            # first field of every line
awk '{print $1, $9}' access.log        # multiple fields
awk -F: '{print $1}' /etc/passwd       # -F sets the field separator (: here)
awk '$9 == 500 {print $7}' access.log  # print URL ($7) where status ($9) is 500
awk '{sum += $5} END {print sum}' file # SUM a column
awk 'NR==1{next} {print}' file         # skip the header line (NR = record number)
df -h | awk '$5+0 > 80 {print $6, $5}' # mountpoints over 80% full

🎯 Interview signal: awk’s {sum+=$N} END{print sum} (aggregation) and -F (custom separator) show you can process structured logs, not just search them.


5. sort, uniq, cut, tr, wc, head/tail

sort file                    # sort lines
sort -n / -rn                # numeric / reverse numeric
sort -k2 -t,                 # sort by 2nd field, comma-delimited
uniq                         # collapse ADJACENT duplicate lines (sort first!)
uniq -c                      # count occurrences
cut -d: -f1 /etc/passwd      # cut field 1, delimiter ':'
cut -c1-10 file              # characters 1–10
tr 'a-z' 'A-Z'               # translate/transform chars
tr -d '\r'                   # delete carriage returns (fix Windows line endings)
wc -l                        # count lines (also -w words, -c bytes)
head -n20 / tail -n20        # first / last 20 lines
tail -f app.log              # follow a growing file (live logs)
tail -f app.log | grep ERROR # live-filter

⚠️ uniq only removes adjacent duplicates — you almost always sort | uniq. Forgetting this is a common mistake interviewers watch for.


6. The classic interview one-liner: top N by frequency

“Given an Nginx access log, show the top 10 client IPs”:

awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -10
#     ↑ extract IP    ↑ group   ↑ count    ↑ rank desc   ↑ top 10

🎯 This exact pipeline (extract → sort → uniq -c → sort -rn → head) is the canonical “analyze a log” answer. Variations: top URLs ($7), count of 5xx (awk '$9~/^5/'), requests per minute, etc. Being able to build it fluently is a strong signal.


7. find — locate files by criteria (and act on them)

find /var/log -name "*.log"                 # by name
find / -type f -size +100M                  # files over 100MB
find /tmp -type f -mtime +7                  # modified > 7 days ago
find . -type f -mmin -60                     # modified in the last 60 minutes
find / -perm -4000 -type f 2>/dev/null       # setuid files (security audit, file 12)
find /data -user arpit                        # owned by a user
find /var/log -name "*.log" -mtime +30 -delete        # delete old logs
find /var/log -name "*.log" -mtime +30 -exec gzip {} \;  # act on each match
find . -type f -print0 | xargs -0 grep -l ERROR       # safe with spaces in names

⚠️ find ... -delete and -exec rm are destructive — run without the action first to preview, and quote/-print0 | xargs -0 to handle filenames with spaces/newlines safely.


8. xargs — build command lines from input

# Turn a list of items into arguments for another command
cat urls.txt | xargs -n1 curl -sO           # download each URL
find . -name "*.tmp" | xargs rm             # delete matches (⚠️ spaces! use -print0/-0)
find . -name "*.log" -print0 | xargs -0 gzip
echo "1 2 3" | xargs -n1 echo               # one arg per invocation
some_list | xargs -P4 -n1 process           # run 4 in parallel

🎯 xargs -P (parallelism) and -0 (null-delimited, space-safe) are the pro touches.


Interview Questions

Q1. Show me how to find the top 10 IPs in an access log.

awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -10. awk extracts the client IP (field 1), sort groups identical lines together, uniq -c counts each, sort -rn ranks by count descending, and head -10 takes the top 10. The same pattern handles top URLs or error counts by changing the field or adding a filter.

Q2. What’s the difference between >, >>, and 2>&1?

> redirects stdout to a file, overwriting; >> appends. 2>&1 redirects stderr (fd 2) to wherever stdout (fd 1) currently points. Order matters: > file 2>&1 sends both to the file, but 2>&1 > file sends stderr to the original terminal and only stdout to the file, because the redirections are applied left to right.

Q3. Why do you almost always pair sort with uniq?

uniq only collapses adjacent duplicate lines, so unsorted input won’t dedupe correctly. Sorting first groups identical lines together; then uniq (or uniq -c to count) works as expected.

Q4. When would you use awk over grep/sed?

When the data is field-structured and you need to select or compute on specific columns — e.g., print the URL where the status code is 500, sum a byte column, or reformat fields. grep finds lines, sed edits streams, but awk understands records and fields and can aggregate ({sum+=$5} END{print sum}), which grep/sed can’t do cleanly.

Q5. How do you safely replace a string across a config file?

Test the sed expression first without -i so it prints to stdout and I can verify the match, then apply with sed -i.bak 's/old/new/g' file to edit in place while keeping a .bak backup. A wrong in-place regex on a production config with no backup is a real outage cause.

Q6. Find all files over 100MB modified in the last week under /var and compress them.

find /var -type f -size +100M -mtime -7 -exec gzip {} \; — or safer for odd filenames, find /var -type f -size +100M -mtime -7 -print0 | xargs -0 gzip. I’d run the find without the action first to preview what it matches before compressing.

Q7. What does xargs do and why use -print0 | xargs -0?

xargs turns lines of input into arguments for another command (e.g., delete or process each match). find -print0 | xargs -0 uses null bytes as separators so filenames containing spaces or newlines don’t get split incorrectly — the standard safe pattern. xargs -P also lets you run invocations in parallel.

Q8. (Senior) Live-tail a busy log for errors while also archiving everything. How?

tail -f app.log | tee -a archive.log | grep --line-buffered -iE "error|fatal". tee -a appends the full stream to an archive while passing it through; grep --line-buffered shows matching lines immediately (line buffering matters in a pipe, or output gets stuck in a buffer). For structured logs I’d swap grep for awk to filter on a specific field/severity.


Next: 09 — Bash Scripting — turning these tools into reliable automation.