Storage & Filesystems

0%
20 minstoragefilesystemslvmraid

Storage & Filesystems

Learn storage & filesystems concepts for DevOps interviews

07 — Storage & Filesystems

Disks, partitions, LVM, mounting, and the ever-present “disk full” incident. Interviewers test whether you can grow a filesystem safely, write a correct /etc/fstab, and systematically find what ate the disk (including the tricky non-obvious causes).


1. The storage stack (bottom to top)

Physical disk (/dev/sda, /dev/nvme0n1)
   └─ Partition (/dev/sda1)                 ── optional
        └─ LVM: PV → VG → LV                 ── optional, flexible layer
             └─ Filesystem (ext4/xfs)        ── formats the space
                  └─ Mount point (/data)     ── where you access it
lsblk                    # tree of disks, partitions, LVs, mount points  ← start here
lsblk -f                 # + filesystem type, UUID, mountpoint
fdisk -l                 # partition tables
blkid                    # UUIDs and fs types of block devices
df -h                    # space usage per mounted filesystem
df -i                    # INODE usage (file 01 — different from space!)

2. Partitions vs LVM

  • Partition: a fixed slice of a disk (MBR or GPT). Resizing is awkward.
  • LVM (Logical Volume Manager): an abstraction layer that makes storage flexible — grow/shrink volumes, span multiple disks, snapshot.

LVM has three layers:

Physical Volume (PV)   = a disk/partition marked for LVM   (pvcreate /dev/sdb)
        │
Volume Group (VG)      = a pool of one or more PVs          (vgcreate data_vg /dev/sdb)
        │
Logical Volume (LV)    = a "virtual partition" from the VG  (lvcreate -n data_lv -L 50G data_vg)
        │
Filesystem on the LV   = mkfs.ext4 /dev/data_vg/data_lv

🎯 Interview signal: “LVM decouples the filesystem from physical disks, so I can add a disk to the VG and grow the LV + filesystem online, without repartitioning or downtime.” That flexibility is the whole reason LVM exists.


3. Creating and growing storage (the common task)

# 1. New disk /dev/sdb → make it an LVM PV, add to a VG, carve an LV
sudo pvcreate /dev/sdb
sudo vgcreate data_vg /dev/sdb
sudo lvcreate -n data_lv -L 50G data_vg
sudo mkfs.ext4 /dev/data_vg/data_lv
sudo mkdir /data && sudo mount /dev/data_vg/data_lv /data

# 2. GROW it later (online): extend LV, then grow the filesystem
sudo lvextend -L +20G /dev/data_vg/data_lv      # or -l +100%FREE
sudo resize2fs /dev/data_vg/data_lv             # ext4: grow the filesystem
#   for XFS:  sudo xfs_growfs /data             # XFS grows via mountpoint, ONLINE only

⚠️ Two classic gotchas:

  • Growing the LV is not enough — you must also grow the filesystem (resize2fs for ext4, xfs_growfs for XFS). People extend the LV, see no extra space, and are confused.
  • XFS can only grow, never shrink. ext4 can shrink but only offline (unmounted). If shrinking matters, that influences filesystem choice.

4. Filesystem types

FS Notes
ext4 Mature, general-purpose default on many distros; can grow (online) and shrink (offline)
XFS RHEL default; great for large files/high throughput; grow-only, no shrink
btrfs / ZFS Copy-on-write, snapshots, checksums; more features, more complexity
tmpfs RAM-backed (e.g., /run, /dev/shm) — fast, volatile
vfat/exFAT Cross-OS removable media

🎯 Interview signal: “RHEL defaults to XFS (grow-only); Ubuntu commonly ext4. If a workload needs shrinking, that argues for ext4; for huge sequential I/O, XFS.”


5. Mounting & /etc/fstab (persistent mounts)

mount /dev/data_vg/data_lv /data     # mount now (not persistent)
umount /data                          # unmount
mount -a                              # mount everything in fstab (TEST after editing!)

/etc/fstab fields:

# <device>              <mountpoint> <fstype> <options>              <dump> <pass>
UUID=1234-abcd           /data        ext4     defaults,noatime       0      2
/dev/data_vg/swap_lv     none         swap     sw                     0      0
tmpfs                    /dev/shm     tmpfs    defaults               0      0
  • <pass>: fsck order at boot — 1 for root /, 2 for others, 0 to skip.
  • Useful options: noatime (don’t update access times — perf), nofail (don’t block boot if missing — good for optional/network mounts), ro, noexec, nosuid.

⚠️ A bad fstab entry can break boot (drops to emergency shell). Always:

  1. Use UUID= (device names like /dev/sdb1 can reorder).
  2. Run mount -a after editing to validate before rebooting.
  3. Use nofail for non-critical/network mounts so a missing disk doesn’t hang boot.

6. Swap

Swap is disk space used as overflow when RAM is full (or for hibernation).

swapon --show        # active swap
free -h              # RAM + swap usage
sudo swapoff -a      # disable (e.g., Kubernetes historically requires swap off)

⚠️ Kubernetes traditionally requires swap disabled on nodes (the scheduler/kubelet assume no swap). Heavy swapping (“thrashing”) also means you’re out of RAM and performance collapses (file 10). Swap masks memory pressure — don’t rely on it for real workloads.


7. The “disk full” incident (a guaranteed question)

There are three distinct causes — naming all three is the senior answer:

1. BLOCKS full        → df -h shows 100%.  Find big files/dirs:
     sudo du -xh / --max-depth=1 | sort -rh | head        (-x stays on one filesystem)
     sudo du -sh /var/log/*      (logs are the usual culprit)
2. INODES exhausted   → df -h has space but writes fail. df -i shows 100% inodes.
     Millions of tiny files (sessions, cache, mail spool). Delete them.
3. DELETED-BUT-OPEN   → df full but du finds nothing (file 01).
     A process holds a deleted file open. Find it:  sudo lsof +L1
     Fix: restart the process holding it (often a service still writing a deleted log).

Systematic approach:

df -h                                   # which filesystem is full
df -i                                   # is it inodes, not blocks?
sudo du -xh / --max-depth=1 | sort -rh | head   # biggest top-level consumers
sudo du -sh /var/log/* | sort -rh | head        # logs are #1 offender
sudo lsof +L1                            # deleted-but-open files

Common fixes: rotate/compress/truncate logs (file 11), clear package caches (apt clean/dnf clean all), remove old kernels, prune Docker (docker system prune), and add monitoring/alerts before it fills.

🎯 Interview signal: leading with “first I check whether it’s blocks (df -h), inodes (df -i), or a deleted-but-open file (lsof +L1)” immediately marks you as experienced — most candidates only know du.

⚠️ On /-full emergencies, a subtle one: you sometimes can’t even log in or run commands if /tmp or /var is full. Truncate a big log with : > /var/log/huge.log (frees space without rm, avoiding the deleted-but-open trap) to get breathing room.


Interview Questions

Q1. Explain the LVM layers and why LVM is useful.

Physical Volumes (disks/partitions initialized for LVM) join a Volume Group (a storage pool), from which you carve Logical Volumes (virtual partitions) that you format with a filesystem. LVM decouples filesystems from physical disks, so you can add disks to the VG and grow (and, with ext4, shrink) volumes and the filesystem online, span multiple disks, and take snapshots — all without repartitioning or downtime.

Q2. You extended a Logical Volume but the filesystem still shows the old size. Why?

Extending the LV only enlarges the block device; you must then grow the filesystem on it. For ext4 that’s resize2fs <lv>; for XFS it’s xfs_growfs <mountpoint> (online). Until you do that, the extra space isn’t usable.

Q3. Difference between ext4 and XFS that matters operationally?

Both are journaling filesystems. XFS (RHEL default) excels at large files/high throughput but can only grow, never shrink. ext4 (common on Ubuntu) can grow online and shrink offline. If a workload may need shrinking, ext4 is safer; for very large sequential I/O, XFS is a good fit.

Q4. Walk me through debugging a full disk.

First determine the type: df -h for block usage, df -i for inode exhaustion (space free but writes fail from too many tiny files), and lsof +L1 for deleted-but-open files (df full but du finds nothing because a process holds a deleted file). Then, for blocks, du -xh / --max-depth=1 | sort -rh to find the biggest consumers — usually /var/log. Fix by rotating/ truncating logs, clearing caches, pruning containers/old kernels, and adding alerts.

Q5. Why can deleting a large log file not free any space?

If a process still has that file open, rm only removes the directory entry; the inode and its blocks stay allocated until the process closes it (deleted-but-open). lsof +L1 shows such files. You must restart or signal the holding process (often a service still writing the log) to actually reclaim the space, or truncate through its fd.

Q6. How do you write a safe /etc/fstab entry and avoid breaking boot?

Use UUID= (or a label) instead of /dev/sdX since device names can reorder; pick correct options (defaults, noatime, nofail for optional/network mounts); set the fsck pass (1 for root, 2 for others, 0 to skip); and always run mount -a after editing to validate before rebooting. A bad entry can drop the system into an emergency shell.

Q7. What is swap and why does Kubernetes want it off?

Swap is disk space used as overflow when RAM is exhausted. Kubernetes traditionally requires swap disabled because the kubelet’s memory accounting and QoS/eviction logic assume no swap — swapping would let pods exceed memory limits invisibly and hurt scheduling predictability. Heavy swapping also means you’re out of RAM and performance collapses.

Q8. How would you add 100 GB to a full /data filesystem with no downtime?

If it’s on LVM: attach a new disk, pvcreate it, vgextend the VG, lvextend the LV (-l +100%FREE), then grow the filesystem online (resize2fs for ext4 or xfs_growfs for XFS). No unmount needed. If it’s a raw partition without LVM, it’s much harder — which is exactly why production volumes use LVM (or cloud elastic volumes) in the first place.

Q9. (Senior) The root filesystem is 100% full and you can barely run commands. First moves?

Get breathing room without triggering the deleted-but-open trap: truncate the biggest log in place with : > /var/log/<huge>.log (frees space immediately while keeping the file’s inode so services keep writing). Then df -h/df -i to classify, du -xh / --max-depth=1 to find consumers, clear package caches and old kernels, and check lsof +L1 for deleted-but-open files needing a service restart. Afterward, add disk-usage alerts and fix log rotation so it doesn’t recur.


Next: 08 — Text Processing & Shell Tools — the grep/sed/awk toolkit that defines a fluent Linux user.