Storage & Filesystems
Storage & Filesystems
Learn storage & filesystems concepts for DevOps interviews
07 — Storage & Filesystems
Disks, partitions, LVM, mounting, and the ever-present “disk full” incident. Interviewers test whether you can grow a filesystem safely, write a correct
/etc/fstab, and systematically find what ate the disk (including the tricky non-obvious causes).
1. The storage stack (bottom to top)
Physical disk (/dev/sda, /dev/nvme0n1)
└─ Partition (/dev/sda1) ── optional
└─ LVM: PV → VG → LV ── optional, flexible layer
└─ Filesystem (ext4/xfs) ── formats the space
└─ Mount point (/data) ── where you access it
lsblk # tree of disks, partitions, LVs, mount points ← start here
lsblk -f # + filesystem type, UUID, mountpoint
fdisk -l # partition tables
blkid # UUIDs and fs types of block devices
df -h # space usage per mounted filesystem
df -i # INODE usage (file 01 — different from space!)
2. Partitions vs LVM
- Partition: a fixed slice of a disk (MBR or GPT). Resizing is awkward.
- LVM (Logical Volume Manager): an abstraction layer that makes storage flexible — grow/shrink volumes, span multiple disks, snapshot.
LVM has three layers:
Physical Volume (PV) = a disk/partition marked for LVM (pvcreate /dev/sdb)
│
Volume Group (VG) = a pool of one or more PVs (vgcreate data_vg /dev/sdb)
│
Logical Volume (LV) = a "virtual partition" from the VG (lvcreate -n data_lv -L 50G data_vg)
│
Filesystem on the LV = mkfs.ext4 /dev/data_vg/data_lv
🎯 Interview signal: “LVM decouples the filesystem from physical disks, so I can add a disk to the VG and grow the LV + filesystem online, without repartitioning or downtime.” That flexibility is the whole reason LVM exists.
3. Creating and growing storage (the common task)
# 1. New disk /dev/sdb → make it an LVM PV, add to a VG, carve an LV
sudo pvcreate /dev/sdb
sudo vgcreate data_vg /dev/sdb
sudo lvcreate -n data_lv -L 50G data_vg
sudo mkfs.ext4 /dev/data_vg/data_lv
sudo mkdir /data && sudo mount /dev/data_vg/data_lv /data
# 2. GROW it later (online): extend LV, then grow the filesystem
sudo lvextend -L +20G /dev/data_vg/data_lv # or -l +100%FREE
sudo resize2fs /dev/data_vg/data_lv # ext4: grow the filesystem
# for XFS: sudo xfs_growfs /data # XFS grows via mountpoint, ONLINE only
⚠️ Two classic gotchas:
- Growing the LV is not enough — you must also grow the filesystem (
resize2fsfor ext4,xfs_growfsfor XFS). People extend the LV, see no extra space, and are confused. - XFS can only grow, never shrink. ext4 can shrink but only offline (unmounted). If shrinking matters, that influences filesystem choice.
4. Filesystem types
| FS | Notes |
|---|---|
| ext4 | Mature, general-purpose default on many distros; can grow (online) and shrink (offline) |
| XFS | RHEL default; great for large files/high throughput; grow-only, no shrink |
| btrfs / ZFS | Copy-on-write, snapshots, checksums; more features, more complexity |
| tmpfs | RAM-backed (e.g., /run, /dev/shm) — fast, volatile |
| vfat/exFAT | Cross-OS removable media |
🎯 Interview signal: “RHEL defaults to XFS (grow-only); Ubuntu commonly ext4. If a workload needs shrinking, that argues for ext4; for huge sequential I/O, XFS.”
5. Mounting & /etc/fstab (persistent mounts)
mount /dev/data_vg/data_lv /data # mount now (not persistent)
umount /data # unmount
mount -a # mount everything in fstab (TEST after editing!)
/etc/fstab fields:
# <device> <mountpoint> <fstype> <options> <dump> <pass>
UUID=1234-abcd /data ext4 defaults,noatime 0 2
/dev/data_vg/swap_lv none swap sw 0 0
tmpfs /dev/shm tmpfs defaults 0 0
<pass>: fsck order at boot —1for root/,2for others,0to skip.- Useful options:
noatime(don’t update access times — perf),nofail(don’t block boot if missing — good for optional/network mounts),ro,noexec,nosuid.
⚠️ A bad fstab entry can break boot (drops to emergency shell). Always:
- Use
UUID=(device names like/dev/sdb1can reorder). - Run
mount -aafter editing to validate before rebooting. - Use
nofailfor non-critical/network mounts so a missing disk doesn’t hang boot.
6. Swap
Swap is disk space used as overflow when RAM is full (or for hibernation).
swapon --show # active swap
free -h # RAM + swap usage
sudo swapoff -a # disable (e.g., Kubernetes historically requires swap off)
⚠️ Kubernetes traditionally requires swap disabled on nodes (the scheduler/kubelet assume no swap). Heavy swapping (“thrashing”) also means you’re out of RAM and performance collapses (file 10). Swap masks memory pressure — don’t rely on it for real workloads.
7. The “disk full” incident (a guaranteed question)
There are three distinct causes — naming all three is the senior answer:
1. BLOCKS full → df -h shows 100%. Find big files/dirs:
sudo du -xh / --max-depth=1 | sort -rh | head (-x stays on one filesystem)
sudo du -sh /var/log/* (logs are the usual culprit)
2. INODES exhausted → df -h has space but writes fail. df -i shows 100% inodes.
Millions of tiny files (sessions, cache, mail spool). Delete them.
3. DELETED-BUT-OPEN → df full but du finds nothing (file 01).
A process holds a deleted file open. Find it: sudo lsof +L1
Fix: restart the process holding it (often a service still writing a deleted log).
Systematic approach:
df -h # which filesystem is full
df -i # is it inodes, not blocks?
sudo du -xh / --max-depth=1 | sort -rh | head # biggest top-level consumers
sudo du -sh /var/log/* | sort -rh | head # logs are #1 offender
sudo lsof +L1 # deleted-but-open files
Common fixes: rotate/compress/truncate logs (file 11), clear package caches
(apt clean/dnf clean all), remove old kernels, prune Docker (docker system prune),
and add monitoring/alerts before it fills.
🎯 Interview signal: leading with “first I check whether it’s blocks (df -h), inodes
(df -i), or a deleted-but-open file (lsof +L1)” immediately marks you as experienced —
most candidates only know du.
⚠️ On /-full emergencies, a subtle one: you sometimes can’t even log in or run commands if
/tmp or /var is full. Truncate a big log with : > /var/log/huge.log (frees space without
rm, avoiding the deleted-but-open trap) to get breathing room.
Interview Questions
Q1. Explain the LVM layers and why LVM is useful.
Physical Volumes (disks/partitions initialized for LVM) join a Volume Group (a storage pool), from which you carve Logical Volumes (virtual partitions) that you format with a filesystem. LVM decouples filesystems from physical disks, so you can add disks to the VG and grow (and, with ext4, shrink) volumes and the filesystem online, span multiple disks, and take snapshots — all without repartitioning or downtime.
Q2. You extended a Logical Volume but the filesystem still shows the old size. Why?
Extending the LV only enlarges the block device; you must then grow the filesystem on it. For ext4 that’s
resize2fs <lv>; for XFS it’sxfs_growfs <mountpoint>(online). Until you do that, the extra space isn’t usable.
Q3. Difference between ext4 and XFS that matters operationally?
Both are journaling filesystems. XFS (RHEL default) excels at large files/high throughput but can only grow, never shrink. ext4 (common on Ubuntu) can grow online and shrink offline. If a workload may need shrinking, ext4 is safer; for very large sequential I/O, XFS is a good fit.
Q4. Walk me through debugging a full disk.
First determine the type:
df -hfor block usage,df -ifor inode exhaustion (space free but writes fail from too many tiny files), andlsof +L1for deleted-but-open files (df full but du finds nothing because a process holds a deleted file). Then, for blocks,du -xh / --max-depth=1 | sort -rhto find the biggest consumers — usually/var/log. Fix by rotating/ truncating logs, clearing caches, pruning containers/old kernels, and adding alerts.
Q5. Why can deleting a large log file not free any space?
If a process still has that file open,
rmonly removes the directory entry; the inode and its blocks stay allocated until the process closes it (deleted-but-open).lsof +L1shows such files. You must restart or signal the holding process (often a service still writing the log) to actually reclaim the space, or truncate through its fd.
Q6. How do you write a safe /etc/fstab entry and avoid breaking boot?
Use
UUID=(or a label) instead of/dev/sdXsince device names can reorder; pick correct options (defaults,noatime,nofailfor optional/network mounts); set the fsck pass (1 for root, 2 for others, 0 to skip); and always runmount -aafter editing to validate before rebooting. A bad entry can drop the system into an emergency shell.
Q7. What is swap and why does Kubernetes want it off?
Swap is disk space used as overflow when RAM is exhausted. Kubernetes traditionally requires swap disabled because the kubelet’s memory accounting and QoS/eviction logic assume no swap — swapping would let pods exceed memory limits invisibly and hurt scheduling predictability. Heavy swapping also means you’re out of RAM and performance collapses.
Q8. How would you add 100 GB to a full /data filesystem with no downtime?
If it’s on LVM: attach a new disk,
pvcreateit,vgextendthe VG,lvextendthe LV (-l +100%FREE), then grow the filesystem online (resize2fsfor ext4 orxfs_growfsfor XFS). No unmount needed. If it’s a raw partition without LVM, it’s much harder — which is exactly why production volumes use LVM (or cloud elastic volumes) in the first place.
Q9. (Senior) The root filesystem is 100% full and you can barely run commands. First moves?
Get breathing room without triggering the deleted-but-open trap: truncate the biggest log in place with
: > /var/log/<huge>.log(frees space immediately while keeping the file’s inode so services keep writing). Thendf -h/df -ito classify,du -xh / --max-depth=1to find consumers, clear package caches and old kernels, and checklsof +L1for deleted-but-open files needing a service restart. Afterward, add disk-usage alerts and fix log rotation so it doesn’t recur.
Next: 08 — Text Processing & Shell Tools — the grep/sed/awk toolkit that defines a fluent Linux user.