Networking Fundamentals

0%
25 minnetworkingtcpipdns

Networking Fundamentals

Learn networking fundamentals concepts for DevOps interviews

06 — Networking

Half of production incidents are “it can’t connect.” This file makes you fluent in the modern ip tooling, routing, DNS resolution, socket/port inspection with ss, firewalls, and — most importantly — a layered troubleshooting method interviewers want to hear.


1. Interfaces & addresses (the modern ip command)

⚠️ ifconfig, netstat, and route are deprecated — use the ip suite and ss. Interviewers notice when you reach for the old tools.

ip addr show            # (ip a) interfaces + IP addresses
ip link show            # (ip l) link state: UP/DOWN, MAC
ip route show           # (ip r) routing table
ip neigh                # ARP/neighbor table (MAC ↔ IP)
ip -s link              # stats (errors, drops)
ip addr add 10.0.0.5/24 dev eth0     # add an IP (runtime)
ip link set eth0 up                   # bring interface up

Persistent config depends on the distro:

  • Ubuntu: netplan (/etc/netplan/*.yaml) → renders to systemd-networkd/NetworkManager.
  • RHEL: NetworkManager (nmcli), configs in /etc/NetworkManager/.
nmcli device status          # NetworkManager view
nmcli connection show

2. Routing — how a packet chooses its exit

The routing table maps destination networks → gateway + interface. The default route (0.0.0.0/0) is where everything without a more specific match goes.

ip route
# default via 10.0.0.1 dev eth0        ← default gateway
# 10.0.0.0/24 dev eth0 proto kernel scope link src 10.0.0.5
ip route get 8.8.8.8      # show which route/interface a packet to 8.8.8.8 would take
  • Longest-prefix match: the most specific matching route wins.
  • No route to a destination → “Network is unreachable.”

🎯 Interview signal: “ip route get <ip> tells me exactly which interface and gateway a packet will use — the fastest way to answer ‘why can’t this host reach that network?’”


3. DNS resolution — name → IP

Order and config:

  • /etc/nsswitch.conf (hosts: line) decides the lookup order — usually files dns (check /etc/hosts first, then DNS).
  • /etc/hosts — static name→IP overrides (checked first).
  • /etc/resolv.conf — the DNS servers (nameserver) and search domains. ⚠️ On systemd-resolved systems this is often a symlink to a stub (127.0.0.53); edit via resolvectl/netplan, not by hand.
dig example.com                 # full DNS query (preferred)
dig +short example.com
dig @8.8.8.8 example.com        # query a specific server (bypass local resolver)
nslookup example.com            # older, still common
resolvectl status               # systemd-resolved: servers per link
getent hosts example.com        # resolve THE WAY the system will (respects nsswitch)

⚠️ Classic trap: ping works but the app fails, or vice versa — because /etc/hosts, search domains, or the app’s own resolver differ. Use getent hosts to resolve the way the OS libc will, and dig to test DNS directly. “It’s always DNS” is a running ops joke for a reason.


4. Ports & sockets — ss (replaces netstat)

ss -tlnp                # TCP, listening, numeric, with process (who owns the port)
ss -tulnp               # TCP + UDP listening
ss -tan                 # all TCP sockets + states
ss -s                   # summary counts
ss -tp state established '( dport = :443 )'   # filter by state/port

Flags: t=TCP, u=UDP, l=listening, n=numeric (no name resolution), p=process.

lsof -i :8080           # what process is using port 8080
fuser 8080/tcp          # PID holding the port

⚠️ “Address already in use” on start = something already listens on that port (ss -tlnp finds it) — often a previous instance that didn’t exit, or a TIME_WAIT/SO_REUSEADDR issue.

🎯 Interview signal: know the TCP handshake (SYN → SYN-ACK → ACK) and states like LISTEN, ESTABLISHED, TIME_WAIT, CLOSE_WAIT. Lots of CLOSE_WAIT = the local app isn’t closing sockets (a bug); lots of TIME_WAIT = normal for a busy client making many short connections.


5. Firewalls

Under the hood: netfilter in the kernel, historically via iptables, now nftables. Front-ends make it manageable:

Front-end Family Basics
ufw Ubuntu ufw allow 22/tcp, ufw enable, ufw status
firewalld RHEL zones + services: firewall-cmd --add-service=https --permanent
raw nft/iptables any fine-grained rules
# Ubuntu
sudo ufw allow 22/tcp && sudo ufw allow 80,443/tcp && sudo ufw enable
sudo ufw status verbose
# RHEL
sudo firewall-cmd --permanent --add-service=https
sudo firewall-cmd --reload
sudo firewall-cmd --list-all

⚠️ Lock-yourself-out risk: enabling a default-deny firewall without first allowing SSH (22) drops your session on the next connection. Always allow SSH before enabling, and on cloud VMs remember the security group / cloud firewall is a separate layer from the host firewall — a blocked port could be either.


6. Connectivity testing tools

ping -c4 8.8.8.8         # ICMP reachability (may be blocked by firewalls)
traceroute example.com   # hop-by-hop path (mtr for continuous)
mtr example.com          # live traceroute + loss %
curl -v https://api:443  # test HTTP(S) incl. TLS handshake and headers
curl -I https://site     # headers only
nc -zv host 5432         # test if a TCP port is open (no app protocol)
telnet host 25           # legacy port test
tcpdump -ni eth0 port 443   # capture packets (see what's actually on the wire)

🎯 Interview signal: distinguish the layers each tool probes — ping (L3 reachability), nc/telnet (L4 TCP port open), curl (L7 app/TLS). “Port open but app fails” vs “can’t even reach the port” points you to different layers.


7. The layered troubleshooting method (say this in interviews)

“Service X can’t reach service Y” — work the layers:

1. Is it a NAME or a NETWORK problem?
   getent hosts Y   /  dig Y      → DNS resolves? to the right IP?
2. L3 reachability:  ping Y (or the gateway)   → routing OK? ip route get Y
3. L4 port open:     nc -zv Y <port>           → is the port listening/allowed?
4. Firewall:         host firewall (ufw/firewalld) AND cloud SG/NACL
5. L7 / app:         curl -v ; check the service is actually LISTENING (ss -tlnp on Y)
6. Logs both sides:  journalctl -u serviceY ; look for refused/timeouts

⚠️ Connection refused vs timeout is a key tell:

  • Refused = reached the host, but nothing is listening on that port (or it actively rejected) → app down / wrong port.
  • Timeout = packets went into a black hole → firewall/security group/routing dropping them silently.

🎯 This single distinction (refused = app layer, timeout = network/firewall layer) is one of the highest-value things you can say in a networking interview.


Interview Questions

Q1. Which commands do you use for interfaces, routing, and sockets, and what replaced the old ones?

The iproute2 suite and ss: ip addr/ip link for interfaces, ip route for routing, ss -tulnp for sockets/ports, ip neigh for ARP. These replace the deprecated ifconfig, route, and netstat. Reaching for ip/ss signals you’re current.

Q2. Walk me through how a hostname becomes an IP on Linux.

The resolver library follows /etc/nsswitch.conf’s hosts: order — typically files dns: check /etc/hosts first, then DNS servers in /etc/resolv.conf (often a systemd-resolved stub at 127.0.0.53). getent hosts <name> resolves the way the OS libc will; dig tests DNS directly and can target a specific server with @.

Q3. Connection refused vs connection timeout — what does each tell you?

Refused means the packet reached the host but nothing was listening on that port (or it was actively rejected) — so it’s an application/port issue: service down, crashed, or wrong port. Timeout means packets got no response at all — a firewall, security group, NACL, or routing problem silently dropping traffic. They point at completely different layers.

Q4. How do you find which process is listening on a port, and what’s using a port that won’t free?

ss -tlnp shows listening TCP sockets with the owning process; lsof -i :<port> or fuser <port>/tcp also identify the PID. “Address already in use” usually means a previous instance still holds the port or it’s in TIME_WAIT — find and stop the holder or fix the app’s socket reuse.

Q5. ping to a server works but your application can’t connect. How do you debug?

ping only proves L3 reachability, not that the app port is open or serving. I’d check the port at L4 (nc -zv host port), confirm the service is actually listening on that host (ss -tlnp), test L7 with curl -v, then check both host firewall and cloud security group, and finally the service logs. Also verify DNS resolves to the IP I think it does with getent hosts.

Q6. How do firewalls work on Linux and what’s the risk when enabling one?

They’re built on the kernel’s netfilter framework, managed via nftables/iptables, with friendlier front-ends: ufw (Ubuntu) and firewalld (RHEL, zone/service based). The risk is locking yourself out — enabling a default-deny policy without first allowing SSH (22) drops your session. On cloud, the host firewall is separate from the security group, so a blocked port could be at either layer.

Q7. What do lots of CLOSE_WAIT or TIME_WAIT sockets indicate?

Many CLOSE_WAIT means the remote closed but the local application isn’t calling close() — usually an app bug leaking sockets/file descriptors. Many TIME_WAIT is normal on a busy client making lots of short-lived connections (the OS holds the socket briefly to catch stray packets); it’s tuned with reuse settings, not usually a bug.

Q8. How do you check the exact route a packet will take?

ip route get <destination> shows the chosen source IP, interface, and gateway using longest-prefix match against the routing table. It’s the quickest way to confirm whether the default gateway or a specific route is being used and to diagnose “network unreachable.”

Q9. (Senior) Intermittent latency and occasional failures to one upstream only. How do you approach it?

Isolate the layer and make it continuous: mtr to that host to spot per-hop packet loss/ latency, dig repeatedly to rule out flaky/round-robin DNS, ss/curl -w to measure connect vs TLS vs first-byte times, and tcpdump on both ends to see retransmits/resets. Check interface error/drop counters (ip -s link), MTU/path-MTU issues, and whether it correlates with load or a specific backend behind a load balancer. The goal is to localize it to DNS, network path, the upstream app, or the LB before touching config.


Next: 07 — Storage & Filesystems — partitions, LVM, and the “disk full” incident.