Linux Fundamentals for DevOps
The core Linux concepts every DevOps engineer touches daily: where the system keeps its files, how users and permissions gate access, and how processes are started, inspected, and stopped.
Learning objectives
- Navigate the Filesystem Hierarchy Standard and explain what belongs in /etc, /var, /opt, and /usr
- Read and modify permission bits with chmod and ownership with chown
- Inspect running processes with ps and top, and terminate them safely with kill and signals
Concept Overview Every Linux distribution follows, more or less, the Filesystem Hierarchy Standard (FHS) — a convention that dictates what kind of data belongs under which top-level directory. As a DevOps engineer you don't memorize the FHS for trivia; you rely on it constantly, because every tool you install, every config you edit, and every log you tail assumes this layout.
/etc holds system-wide configuration — plain text files, almost always human-editable, with no binaries. Nginx's config lives at /etc/nginx/nginx.conf, SSH's server config at /etc/ssh/sshd_config, and systemd unit files at /etc/systemd/system/. If you're hunting for 'how is this service configured,' /etc is the first place to look.
/var holds variable data — files that change size and content while the system runs: logs (/var/log), package caches (/var/cache), spool files for mail or cron (/var/spool), and often application data and databases (/var/lib). This is the directory that fills up a disk over time, which is why disk-usage incidents almost always trace back to something under /var.
/opt is reserved for optional, self-contained third-party software — think of a vendor's agent or a monitoring tool that ships its own directory tree rather than scattering files across the system. You'll commonly see /opt// as a self-contained install.
/usr contains the bulk of the actual operating system — user-space binaries (/usr/bin, /usr/sbin), shared libraries (/usr/lib), and architecture-independent data (/usr/share). Despite the name, /usr has nothing to do with home directories; it stands for 'Unix System Resources' historically. Your package manager installs almost everything here.
Rounding this out: /home holds per-user personal directories, /root is the superuser's home, /tmp is wiped on reboot (don't store anything important there), /bin and /sbin (often symlinked into /usr on modern distros) hold essential binaries needed even in single-user/recovery mode, and /proc and /sys are virtual filesystems exposing kernel and process information as if they were files — they don't consume real disk space.
Understanding this map means that when something breaks, you instinctively know where to look: a misbehaving service's config is in /etc, its logs are in /var/log, and if it's a third-party agent, it's probably self-contained under /opt.
💻 Code example
# Explore the hierarchy relevant to a running service ls -la /etc/nginx/ # configuration files sudo tail -n 50 /var/log/nginx/error.log # logs du -sh /var/lib/* # which app data directories are largest ls /opt/ # any vendor-installed software file /bin/bash # on modern distros this is a symlink into /usr ls -l /bin | head -1
Concept Overview
Every file and directory on a Linux system carries an owner (a user), a group, and a set of permission bits that apply separately to the owner, the group, and everyone else ('other'). Running ls -l shows this directly: a string like -rwxr-xr-- breaks down into the file type (the leading - for a regular file, d for directory, l for symlink), then three triplets of read/write/execute for owner, group, and other respectively.
Each permission has a numeric value: read = 4, write = 2, execute = 1. Add them per triplet to get the familiar three-digit mode: 755 means owner has read+write+execute (4+2+1=7), group has read+execute (4+1=5), and other has read+execute (5). 644 is the classic 'readable by everyone, writable only by the owner' mode for config files. For a directory, the execute bit means something slightly different — it controls whether you can cd into it or access files inside by name, not whether you can 'run' it.
chmod changes these bits. You can use numeric mode (chmod 755 deploy.sh) or symbolic mode (chmod u+x deploy.sh to add execute for the owner, chmod g-w file to remove write from the group, chmod o=r file to set 'other' to read-only). Symbolic mode is often clearer when you're adjusting a single bit without touching the rest.
chown changes ownership — which user and/or group owns the file. chown deploy:deploy app.jar sets both user and group to deploy. You typically need root (or sudo) privileges to change ownership to a user you're not already, since handing files to another user is a privileged operation. chgrp changes just the group if you only need that.
Users belong to groups, which is how Linux implements shared access without giving everyone individual permissions on every file. A deploy user might belong to a docker group to run Docker without sudo, or a www-data group to write into a web server's directory. /etc/passwd lists users, /etc/group lists groups and their members, and id <username> shows a given user's primary and secondary group memberships.
Two special bits matter in production: the setuid/setgid bits (run a binary as its owner/group rather than the invoking user — rare, and a security review item when you see it) and the sticky bit (on a directory like /tmp, it means users can only delete their own files even though the directory itself is world-writable).
💻 Code example
# Inspect and fix permissions on a deploy script ls -l deploy.sh # -rw-r--r-- 1 ubuntu ubuntu 812 Oct 3 10:02 deploy.sh chmod u+x deploy.sh # make it executable for the owner chmod 750 deploy.sh # owner: rwx, group: r-x, other: none # Hand the file to a service account and its group sudo chown deploy:deploy deploy.sh # Check what groups a user belongs to before granting docker access id deploy sudo usermod -aG docker deploy # add deploy user to the docker group groups deploy # confirm membership (re-login required to take effect)
Concept Overview A process is a running instance of a program, identified by a numeric PID. Linux exposes a live view of every process on the system, and you'll spend a fair amount of time inspecting that view to diagnose stuck deploys, runaway memory, or a service that just won't die.
ps gives you a snapshot. The most commonly used invocation is ps aux, which lists every process on the system (a = all, including processes not attached to a terminal, u = user-oriented output showing %CPU and %MEM, x = include processes without a controlling terminal, like daemons). Combine it with grep to find a specific process: ps aux | grep nginx. ps -ef is the other common style, showing parent PIDs (PPID), which is handy for understanding process trees — pstree makes that tree visual.
top (and its friendlier cousin htop, if installed) gives a live, continuously refreshing view sorted by resource usage by default. It's your first stop for 'the server feels slow' — you can immediately see which process is eating CPU or memory, sort interactively (press M for memory, P for CPU in top), and kill a process directly from the UI (k, then enter the PID).
Terminating a process isn't a single operation — it's sending it a signal, and the process (if well-behaved) decides how to react. kill <PID> sends SIGTERM (signal 15) by default — a polite request to shut down, which lets the process clean up open files, close connections, and exit gracefully. Well-written services trap SIGTERM and shut down cleanly within some grace period. kill -9 <PID> sends SIGKILL (signal 9) — this is not a request, it's the kernel immediately removing the process with no chance to clean up. Reach for SIGKILL only when SIGTERM has been ignored (e.g., the process is hung), because skipping cleanup can leave corrupted state, orphaned locks, or unflushed writes.
Other signals worth knowing: SIGHUP (1) is often used by daemons as a 'reload your config without restarting' signal (kill -HUP <PID> on nginx reloads config); SIGINT (2) is what Ctrl+C sends; SIGSTOP/SIGCONT pause and resume a process. kill -l lists all signal names and numbers. pkill and pkill -f let you signal processes by name pattern instead of PID, and killall <name> kills every process matching an exact name — useful but dangerous on a shared host, since it's easy to match more than you intended.
💻 Code example
# Find and gracefully stop a stuck service, escalating only if needed ps aux | grep '[j]ava' # bracket trick avoids matching the grep process itself top -b -n 1 | head -15 # one-shot snapshot sorted by CPU/mem PID=$(pgrep -f myapp.jar) kill "$PID" # SIGTERM: ask nicely sleep 5 if kill -0 "$PID" 2>/dev/null; then # kill -0 just checks if the PID still exists echo "Still running, forcing it down" kill -9 "$PID" # SIGKILL: no cleanup, last resort fi # Reload nginx config without dropping connections sudo kill -HUP $(cat /run/nginx.pid)
Want a visual for this concept?
Generate a diagram tailored to “Linux Fundamentals for DevOps” — the AI picks whichever visual (flowchart, comparison, sequence, etc.) best fits.
Sign in to generate a visual →