Book cover: Mastering Linux/UNIX C Programming

Mastering Linux/UNIX C Programming

A hands-on course from first steps to system programming

IRays Teknology Ltd.

Copyright © 2026 IRays Teknology Ltd. All rights reserved.

Preface

What this book is. A complete course in C, taught the way working programmers meet the language on Linux and the other UNIX systems: in a terminal, with a compiler under strict flags, a manual page one keystroke away, and the whole operating system available as the standard library's energetic downstairs neighbor. The book is an original work. It honors — and frequently points at — the classic 1978 text by Brian Kernighan and Dennis Ritchie, whose two hundred and seventy pages remain the most elegant short treatment of C ever written; this book takes the longer road through the system underneath, because that is where a Linux C programmer actually lives.

The map. Two halves, promised here and delivered in nine chapters: the language first — a first tour and the toolchain (1), types and operators (2), control flow and program structure (3), pointers and arrays (4), structures and memory management (5), and strings (6) — and then the system — files and the kernel's own calls (7), processes, signals, and pipes (8), and the practices, tools, and the view from the end of the road (9). Standard C first, then the system: the two halves of a Linux C programmer's education.

The method. Every program in this book was compiled under the strict flag set it teaches (-Wall -Wextra -Werror, with -g) and — with the exceptions labeled in place — every one was run, and every output block you will see is what the program actually printed. When the machine running the book differed from a Linux system, the difference was captured and taught instead of hidden: type widths, missing functions, translation quirks, missing tool runtimes. The handful of transcripts that could not be captured on the drafting machine are labeled by kind — derived from the tools' documented behavior — and their exercises run them for real on any Linux system. Nothing in this book is meant to be taken on faith; the whole point of the book is that you do not have to.

How to read. Lines beginning with $ are shell commands; what follows, without the $, is the program's output. Programs are printed in full and live beside the chapters in examples/chNN/. Every chapter ends with exercises, and the exercises are not decoration: they are where each chapter's claims become your own evidence. The manual is assumed open — man man first.

What you need. A Linux system — a distribution, a spare machine, or WSL2 on Windows — and a compiler, gcc or clang. Chapters 1–6 speak strict standard C (-std=c17); from Chapter 7 on, the book crosses to the kernel's interface and speaks -std=gnu17, and says why.

Contents


Chapter 1 · First Steps with C on Linux

C is a small language. You can learn the whole of it in weeks; using all of it well takes longer, but everything you learn transfers, because so much of the computing world is built with it. The Linux kernel, the shell utilities in /bin and /usr/bin, git, sqlite, the runtimes of Python and Perl — all C. This book teaches C the way working programmers actually meet it on Linux and the other UNIX systems: in a terminal, with a compiler, a manual page one keystroke away, and the whole operating system available as your standard library.

C and UNIX were born together at Bell Labs in the early 1970s, and the partnership never ended: UNIX was one of the first operating systems written mostly in a high-level language, and C spread across machines precisely because a small, honest compiler for it was everywhere. Linux inherited both halves of that bargain. When you write C on Linux you are using the same tools, the same headers, and largely the same idioms as the people who built the system under you.

To follow this chapter you need three skills: running commands in a shell, editing text files (any editor — vim, nano, VS Code), and reading. Everything else is C.

Conventions used throughout this book. Lines that begin with $ are commands you type at the shell; what follows, without the $, is the program's output. Every program is printed in full, compiled with the flags recommended in §1.11, and run; the output you see under each program is what it actually printed. The source for each program sits beside this chapter in examples/ch01/, and later chapters follow the same pattern (examples/ch02/, and so on).

1.1 The compiler

C is a compiled language: you write a program in a text file, a compiler translates it into machine code, and the result is a self-contained executable you run directly. There is no interpreter to install, no runtime to ship, no intermediate step between compiling and running.

On Linux two compilers matter: GCC, the GNU Compiler Collection, the default on every distribution, and Clang, its main rival, which gives clearer diagnostics and is worth having for that alone. Both are excellent; everything in this book works with either. The commands below install the standard toolchain (compiler, make, headers) on the major distributions:

System Command
Debian, Ubuntu sudo apt install build-essential
Fedora sudo dnf install gcc make
Arch sudo pacman -S base-devel
openSUSE sudo zypper install gcc make
Windows install WSL (wsl --install -d Ubuntu), then the Debian/Ubuntu command
macOS xcode-select --install (installs clang)

Check what you have:

$ gcc --version
gcc (GCC) 15.2.0
Copyright (C) 2025 Free Software Foundation, Inc.
...
$ clang --version
clang version ...

(The exact wording and version number will differ on your machine; any release of the last few years is fine.) The name cc — "C compiler" — is the traditional neutral spelling, and on Linux it is usually a link to whichever compiler is installed. This book says gcc; if you prefer clang, substitute it everywhere, flag for flag.

One note on scope: this book is about Linux, but nearly everything applies to the whole UNIX family — the BSDs and macOS use the same standard C library concepts and the same POSIX interfaces. When a detail is Linux-specific or not portable to, say, macOS, the text says so.

1.2 Your first program

Here is the traditional first program. Type it into a file named hello.c — the .c suffix is what tells tools (and people) that the file holds C source:

/* hello.c -- the traditional first program */
#include <stdio.h>

int main(void)
{
    puts("Hello, Linux!");
    return 0;
}

Compile it, then run it:

$ gcc -o hello hello.c
$ ./hello
Hello, Linux!
$ echo $?
0

Three things happened, and each is worth a slow look.

gcc -o hello hello.c reads hello.c and writes an executable named hello. Without -o, the compiler would name the result a.out — a genuine fossil: the name stands for "assembler output" and comes from the first UNIX C compiler, which is also why some UNIXes still name their object files a.out underneath. Use -o; it costs four keystrokes and saves confusion.

./hello runs the program. The ./ (this directory) is required because the shell, by design, does not search the current directory for programs — only the directories listed in $PATH. That is a security habit, not an inconvenience: imagine cd-ing into a directory where someone left a file named ls.

echo $? prints the exit status of the last command: the number main returned. 0 means success. Every program you have ever run in a shell — ls, grep, gcc — has reported its outcome this way, and shell scripts are built on it. We return to exit statuses in §1.10.

Now the program itself, piece by piece:

Layout is free: C does not care about indentation or line breaks, only punctuation. But readers care, so this book uses one statement per line, four-space indentation for what a brace encloses, and a } on its own line. Pick a style early and stay with it; your code, and this book's, will be easier to live with.

1.3 Formatted output: printf

puts prints one string. Its workhorse colleague is printf — "print formatted" — which takes a format string containing ordinary text and conversion specifiers (introduced by %), followed by one value for each specifier:

/* hello2.c -- formatted output with printf */
#include <stdio.h>

int main(void)
{
    printf("Hello, %s!\n", "Linux");
    printf("%d + %d = %d\n", 2, 3, 2 + 3);
    printf("1024 * 1024 = %d\n", 1024 * 1024);
    return 0;
}
$ gcc -o hello2 hello2.c
$ ./hello2
Hello, Linux!
2 + 3 = 5
1024 * 1024 = 1048576

The specifiers you will use constantly:

Conversion Argument type Prints
%d or %i int signed decimal integer
%u unsigned int unsigned decimal integer
%x, %X unsigned int hexadecimal (7ff, 7FF)
%c int (a character) one character
%s string (§1.9, Chapter 6) a sequence of characters
%f double fixed-point; six decimals by default
%e, %g double scientific notation; "as short as possible"
%ld, %lu long, unsigned long decimal
%lld, %llu long long, unsigned long long decimal
%zu size_t (§1.5) decimal — the type sizeof produces
%p pointer (Chapter 4) a pointer value
%% — a single %

Between % and the conversion letter you may put a width and, for some conversions, a precision, optionally with flags. A width right-aligns the value in that many columns; a - flag left-aligns; a leading 0 pads with zeros; for %f a precision is the number of decimals:

printf("[%8.2f]\n", 3.14159);    /* [    3.14] */
printf("[%-8s]\n", "left");      /* [left    ] */
printf("[%05d]\n", 42);          /* [00042]   */

Field widths are how you line up columns of a table — you will meet them at work in §1.8.

One rule is absolute: each specifier must match the type of its argument. Mismatches are undefined behavior — §1.7 explains what that phrase costs you — and, mercifully, -Wall catches most of them at compile time, as you will see in §1.11.

1.4 What the compiler actually does

gcc -o hello hello.c looks like one step, but it is four, and knowing the seams matters later, when programs grow to many files.

$ gcc -E hello.c -o hello.i      # 1. preprocessing: expand #include, #define
$ gcc -S hello.i                # 2. compiling: C becomes assembly (hello.s)
$ gcc -c hello.s                # 3. assembling: assembly becomes an object file (hello.o)
$ gcc hello.o -o hello          # 4. linking: join object files and libraries into a program

The preprocessor (-E, output .i) is a text editor on a macro level: it inserts the contents of headers, expands macros, and strips comments. The compiler proper (-S, output .s) translates the result into assembly language for your CPU. The assembler (-c, output .o) turns that into an object file: real machine code with external references left dangling. The linker resolves them: it joins your object file with the libraries it needs — including the C standard library, libc, where the code for printf actually lives — and writes the executable.

Two consequences. First, this is why a header can be used without shipping code: stdio.h only declares printf; the linker attaches libc automatically. Some libraries are not automatic — the math functions need -lm — and §2.4's fabs note shows the habit. Second, this pipeline is what makes multi-file programs and make natural: compile each .c to an object file once, relink only what changed. §9.5's make builds on exactly that.

If you are curious, gcc -v prints every sub-command as it runs; ls -l on the intermediate files shows how size shrinks at each stage.

1.5 Variables, types, and sizeof

A variable is a named piece of memory with a type, and every variable must be declared before use:

int count = 0;               /* declaration, with initialization */
double ratio = 1.5;
char initial = 'K';

Names are case-sensitive, made of letters, digits, and underscores, and may not begin with a digit; keywords like int and return are reserved. This book uses lower-case-with-underscores, the style of the Linux kernel. Since C99 a declaration may appear anywhere in a block, just before first use; in much older code all declarations sit at the top of a function — you will recognize the habit instantly (§1.13).

C's basic types come in families:

Type x86-64 Linux Guaranteed at least
char 1 byte 8 bits
short 2 bytes 16 bits
int 4 bytes 16 bits
long 8 bytes 32 bits
long long 8 bytes 64 bits
float 4 bytes ~6 significant digits
double 8 bytes ~15 significant digits

Only the right-hand column is promised by the language; the middle column is what you will almost certainly meet. And you should not take our word for the middle column — ask the machine. sizeof is an operator (yes, an operator: the parentheses are needed only around type names, sizeof x works for expressions) that yields the size of its operand in bytes; the type it returns, size_t, is printed with %zu:

/* sizes.c -- how big are the basic types on this machine? */
#include <stdio.h>

int main(void)
{
    printf("char        %zu byte\n",  sizeof(char));
    printf("short       %zu bytes\n", sizeof(short));
    printf("int         %zu bytes\n", sizeof(int));
    printf("long        %zu bytes\n", sizeof(long));
    printf("long long   %zu bytes\n", sizeof(long long));
    printf("float       %zu bytes\n", sizeof(float));
    printf("double      %zu bytes\n", sizeof(double));
    printf("a pointer   %zu bytes\n", sizeof(void *));
    return 0;
}
$ gcc -o sizes sizes.c
$ ./sizes
char        1 byte
short       2 bytes
int         4 bytes
long        8 bytes
long long   8 bytes
float       4 bytes
double      8 bytes
a pointer   8 bytes

That is what a 64-bit x86-64 Linux system prints, and it hides a real portability lesson: on Windows — with MSYS2, MinGW, or Visual Studio — long prints 4 even in a 64-bit program. Windows uses a memory model (LLP64) in which only pointers and long long are 64 bits, while Linux and macOS use LP64, which makes long 64 bits as well. long is therefore the least trustworthy integer in C. When a size must be exact, do not guess from the table — name it:

#include <stdint.h>               /* exact-width integer types, since C99 */

uint32_t exact = 4000000000u;     /* exactly 32 bits, on every system */
int64_t  total = 0;               /* exactly 64 bits */

stdint.h also provides int8_t, int16_t, and the uint family, plus _least and _fast variants, and uintptr_t, an integer type wide enough to hold a pointer — all in Chapter 4's territory. Two more standard types matter now: size_t, an unsigned type for sizes and (as you will see) array indices, returned by sizeof and used throughout the library; and bool, with values true and false, from <stdbool.h> since C99 and a genuine keyword since C23.

And char deserves one more look, because it is quietly doing two jobs. A char stores a small integer — one byte, on every mainstream system — and character literals are just small integers with a spelling:

char c = 'A';
printf("%c %d\n", c, c);          /* prints: A 65 */

'A' is 65 in ASCII, the character set of the C basic execution environment, and every Linux locale you are likely to use day to day. Where things get subtle: whether a plain char is signed or unsigned is up to the compiler. When a variable holds a character, use char. When it holds a byte — data read from a file, a checksum, a buffer element — use unsigned char, which has a guaranteed range of 0 to 255 and no surprises. This distinction becomes important in Chapter 7, when we read real files.

1.6 Constants, literals, and macros

Numbers and strings that appear in programs are literals, and their exact spelling carries type information:

For named constants, modern C offers three tools, in increasing order of distrust:

const double PI = 3.14159265358979;
enum { BUFLEN = 256 };
#define PROGRAM_NAME "nl"

const makes an object read-only after initialization: typed, checked, scoped. One C quirk: a const int is not a compile-time constant, so it cannot size a static array. enum declares integer constants that are true compile-time constants, with scope and no text-substitution risk — the idiomatic way to name array sizes and small sets of related constants. #define names a macro, substituted textually by the preprocessor everywhere it appears; it is untyped and scopeless, and until you know enough to miss its power (§9.6), prefer enum for integers and const for the rest. We use #define sparingly, and always with parentheses, in this book.

1.7 Arithmetic and expressions

The arithmetic operators are + - * / % with the usual precedence, parentheses to override it, and one standing piece of advice: when in doubt, parenthesize. You are writing for the next reader as much as for the compiler.

The operator with the most surprises for newcomers is integer division, which truncates:

printf("%d\n", 7 / 2);        /* 3   -- integer division truncates toward zero */
printf("%d\n", 7 % 2);       /* 1   -- % is the remainder, integers only */
printf("%f\n", 7 / 2.0);     /* 3.500000 -- one double operand makes the math double */

Mixing types follows usual arithmetic conversions: narrower integer types widen to int, and if either operand is double the other is converted too. Use double for floating-point by default; float earns its keep mainly in large arrays where its size matters (Chapter 5).

Division by zero is not a defined value — it kills the process at once with a signal (on Linux, helpfully named SIGFPE, "floating-point exception", even for integers; signals are Chapter 8). More insidious is overflow. Unsigned arithmetic wraps modulo its range:

unsigned char u = 255;
u = u + 1;                   /* now 0, not 256 -- defined, predictable wrapping */

but signed overflow is undefined behavior. "Undefined" is a technical term, and it is harsher than "wrong answer": the standard permits the compiler to assume it cannot happen, and modern optimizers — gcc -O2 especially — act on that assumption in ways that can mangle unrelated code. The practical rules: use unsigned types for bit patterns and sizes; check ranges at boundaries; and let the tools police you — UBSan, §1.11, makes signed overflow stop the program at the moment it happens.

Rounding out the expression toolkit: ++ and -- increment and decrement (as a statement, n++;, where their prefix/postfix difference is moot); compound assignments += -= *= /= %= are shorthand; comparisons == != < > <= >= yield int values 0 or 1, so printf("%d\n", 2 < 3); prints 1; and &&, ||, ! combine conditions, short-circuiting — the left operand of && is fully evaluated first, and if it decides the answer, the right is not evaluated at all. That is not an optimization footnote; it is a guarantee you will build on (Chapter 3).

1.8 Control flow

C has a deliberately small set of control statements — enough to build anything, small enough to hold in your head.

if / else chooses:

if (n > 0)
    printf("positive\n");
else if (n < 0)
    printf("negative\n");
else
    printf("zero\n");

The braces are optional around single statements, but this book almost always writes them: code grows, and "forgot braces after adding a line" is a bug you have to find. Write

if (n > 0) {
    printf("positive\n");
}

while repeats while a condition holds; for bundles initialize-test-increment on one line, and (since C99) the loop variable may be declared in the statement itself, scoped to the loop:

int i = 0;
while (i < 3) {
    printf("i is %d\n", i);
    i++;
}

for (int i = 0; i < 3; i++) {
    printf("i is %d\n", i);
}

A first real use of the for loop, with the field widths of §1.3 doing honest work lining up a table:

/* pow2.c -- a table of powers of two */
#include <stdio.h>

int main(void)
{
    unsigned long long p = 1;

    for (int i = 0; i <= 63; i++) {
        printf("2^%-2d = %20llu\n", i, p);
        p *= 2;
    }
    return 0;
}
$ gcc -o pow2 pow2.c
$ ./pow2
2^0  =                    1
2^1  =                    2
2^2  =                    4
...
2^10 =                 1024
2^16 =                65536
...
2^31 =           2147483648
2^32 =           4294967296
...
2^62 =  4611686018427387904
2^63 =  9223372036854775808

Read the loop as "for i from 0 to 63, print p and double it". Note how much the loop depends on types: p is unsigned long long precisely because 2^63 does not fit in an int or a long on every system, and wrapping (§1.7) is what would happen otherwise. If you change the bound to i <= 64, watch the last row lie to you — a worthwhile minute with UBSan.

do-while tests after the body, so the body runs at least once — the shape of a "read a command, then decide" loop; it is rare, and right, exactly when that guarantee is what you mean.

break leaves the nearest enclosing loop or switch immediately; continue jumps to the next iteration's test. Both are honest tools; both are easier to misuse than to use, so this book reaches for them sparingly.

switch dispatches on an integer-valued expression, jumping to the matching constant label:

switch (c) {
case 'y':
case 'Y':
    answer = 1;
    break;
case 'n':
case 'N':
    answer = 0;
    break;
default:
    answer = -1;
    break;
}

Two labels in a row share one arm — that is how y and Y land together, and the only sanctioned use of "falling through" a label. Everywhere else, each arm ends in break: forget it and execution falls into the next arm, historically a bottomless source of bugs. (GCC can police this with -Wimplicit-fallthrough, which is why the recommended flags in §1.11 matter even for a language this small.) Labels must be integer constant expressions — case 'y': works because, as §1.5 promised, a character literal is an integer.

1.9 A first useful program: a line-numbering filter

The most durable idea the UNIX system gave software is the filter: a program that reads lines of text from its standard input, does one thing well, and writes text to its standard output. cat, grep, sort, wc are filters, and because they all speak the same protocol — text in, text out — the shell's pipe | composes them into tools no single author had to write. Time to join that economy: a program that numbers its input lines, a pocket version of cat -n:

/* nl.c -- number the lines of standard input (a tiny `cat -n`) */
#include <stdio.h>

int main(void)
{
    char line[256];
    int n = 0;

    while (fgets(line, sizeof line, stdin) != NULL) {
        n++;
        printf("%6d\t%s", n, line);
    }
    return 0;
}

Two new ideas, both worth stopping for.

char line[256]; declares an array: 256 adjacent bytes named line. So far we only use it as a buffer, a place for characters to land; arrays and the pointer arithmetic built on them are the heart of Chapter 4, and nothing here requires you to understand them yet.

fgets(line, sizeof line, stdin) reads one line of input into that buffer. Its contract, read carefully, is very good: it stores at most 255 characters plus the newline if it fits, always adds a terminating null character so the buffer is a well-formed string, and returns the buffer — or NULL when input is exhausted. The loop runs the body once per line and stops at end of input with no counter, no flag, no fuss. Passing sizeof line instead of retyping 256 means the size lives in exactly one place (§1.6's lesson, applied). stdin is the standard input stream, opened for you before main runs — three streams in §1.10.

The printf prints the count right-aligned in six columns (%6d), a tab, then the line — fgets kept its newline, so the format string adds none.

Try it — on its own source, then as a cog in a pipeline:

$ gcc -o nl nl.c
$ ./nl < nl.c
     1  /* nl.c -- number the lines of standard input (a tiny `cat -n`) */
     2  #include <stdio.h>
     3  
     4  int main(void)
     5  {
     6      char line[256];
     7      int n = 0;
     8  
     9      while (fgets(line, sizeof line, stdin) != NULL) {
    10          n++;
    11          printf("%6d\t%s", n, line);
    12      }
    13      return 0;
    14  }
$ ls | ./nl
     1  fail.c
     2  hello.c
     3  hello2.c
     4  nl.c
     5  pow2.c
     6  sizes.c
     7  wrong.c

Both are the UNIX idea in miniature: < redirects a file into standard input, and | makes one program's standard output the next one's standard input — our nl, thirty lines into your C career, interoperates with ls exactly as grep and sort do.

Honesty about limits, because a book that hides them teaches you to distrust books: a line longer than 255 characters arrives in pieces, and each piece gets its own number — real cat -n does not do that, and the honest repair waits for the string tools of Chapter 6. And fgets returns NULL both at end of input and on error; we treat both as "done", which is wrong about a real disk error — the right check is ferror(stdin), and it takes until Chapter 7 to pay that debt. Meanwhile the program is small, correct for the cases it claims, and composed of tools you now know.

1.10 Three streams, exit status, and redirection

Every process starts life with three open files, the standard streams:

Stream Number Default destination
standard input 0 the keyboard
standard output 1 the terminal
standard error 2 the terminal

The numbers are file descriptors, the currency of all UNIX input and output — Chapter 7 is about them. The distinction between the two output streams is the one that matters now: stdout is for a program's product — the thing it exists to produce, fine to filter, pipe, or redirect — and stderr is for diagnostics: errors, progress, anything a human should see even when the product is being consumed by another program. A program that writes its errors to stdout breaks the moment someone pipes it.

Separating the streams is half the discipline; the other half is exiting honestly. main's return value becomes the process's exit status, visible to the shell as $?:

/* fail.c -- demonstrate stderr and a nonzero exit status */
#include <stdio.h>

int main(void)
{
    fprintf(stderr, "fail: this program deliberately fails\n");
    return 1;
}
$ gcc -o fail fail.c
$ ./fail
fail: this program deliberately fails
$ echo $?
1
$ ./fail > /dev/null
fail: this program deliberately fails
$ ./fail > /dev/null 2>&1
$ echo $?
1

fprintf is printf's sibling, writing to a chosen stream — here stderr. The first run shows the message; the second shows the crucial property: > redirected stdout to /dev/null (the byte-shredder device), and the message still appeared, because stderr was untouched. The third run silences both: 2>&1 sends stream 2 to wherever stream 1 is going, and it must follow the > it modifies. The exit status survives all of it, because it is not a stream: it is the process's last word to whoever launched it. The shell keeps only its low 8 bits (return 256 and $? reads 0 — try it), and the convention is 0 = success, 1–255 = failure, with many tools using 2 specifically for "you used the command wrong" (check ls --not-an-option; echo $?).

Two closing notes. Deep in a function — nowhere near main's return — exit(1) from <stdlib.h> terminates the whole program with that status at once. And buffering: stderr is unbuffered (messages appear immediately, which is what you want from a diagnostic), while stdout is line-buffered on a terminal but block-buffered when piped or redirected. That is why interleaved printf and fprintf(stderr, ...) can appear out of order in a log file. It becomes a real tool in Chapter 7; for now it explains an observation you will eventually make and otherwise find maddening.

1.11 Warnings, sanitizers, and the debugger

Here is a program with exactly one wrong character — %d where %f belongs:

/* wrong.c -- deliberately wrong, to see what a warning looks like */
#include <stdio.h>

int main(void)
{
    double half = 0.5;

    printf("half is %d\n", half);   /* %d is for ints, not doubles */
    return 0;
}

Compiling it with the flags this section builds toward, GCC says (Clang's wording is similar):

$ gcc -std=c17 -Wall -Wextra -g -o wrong wrong.c
wrong.c: In function 'main':
wrong.c:8:22: warning: format '%d' expects argument of type 'int', but argument 2 has type 'double' [-Wformat=]
    8 |     printf("half is %d\n", half);   /* %d is for ints, not doubles */
      |                     ~^     ~~~~
      |                      |     |
      |                      int   double
      |                     %f

Read it like an address: file wrong.c, line 8, column 22; the specifier %d wants an int and was given a double; and the last line is GCC going further than complaint — it proposes the fix, %f. Now run it:

$ ./wrong
half is 0
$ echo $?
0

half is 0 — wrong, silently, with exit status 0, "success". This is worth internalizing early: a program that compiles and runs is not thereby correct, and most of your career's bugs will be in that category. printf with a mismatched specifier is undefined behavior (§1.7): it read an int where a double was passed, and happened to find a zero. Another compiler, another day, another number.

The defense is a compile line, and this is the one to adopt as a reflex — every program in this book passes it:

$ gcc -std=c17 -Wall -Wextra -Werror -g -o program program.c

For bugs that run rather than compile, the modern toolchain has two sanitizers, which add runtime checks that stop the program at the exact moment something goes wrong, instead of letting it corrupt state and lie about it:

$ gcc -std=c17 -g -fsanitize=address,undefined -o program program.c

AddressSanitizer (ASan) catches out-of-bounds array indexing, use-after-free, and (on Linux) leaks — memory you forgot to release. UBSan catches undefined behavior as it executes: signed overflow, bad shifts, misaligned access. A program under sanitizers that completes silently has earned real confidence. The alternative valgrind does much of the same work without recompiling, at a larger speed cost. Chapter 9 makes all of these routine.

And the debugger itself. gdb runs your program under observation; a minimal session looks like:

$ gdb -q ./nl
(gdb) break main          # stop on entry to main
(gdb) run                 # start, with arguments after run if any
(gdb) next                # step over a line
(gdb) print n             # inspect a variable
(gdb) quit

Chapter 9 turns this from a shape into a working skill; what matters on day one is that a debugger attached to a -g build shows you your source lines and your variable names — and that "just add printf" is not the only tool you own.

1.12 The manual

Linux ships its documentation with it, and the tool is man. A man page opens in a pager: move with arrows or space, search with /pattern, leave with q. Start with the manual about the manual:

$ man man

Man pages are divided into sections by what they document, and the section is part of the name:

Section Documents Examples
1 user commands man 1 ls, man 1 gcc
2 system calls man 2 open, man 2 fork
3 library functions man 3 printf, man 3 malloc
5 file formats man 5 passwd
7 conventions, tables man 7 ascii, man 7 signal

The distinction between sections 2 and 3 is the C programmer's map of the world: section 3 is the standard C library — printf, fgets, malloc, functions that exist on any C implementation — while section 2 is the kernel's interface — open, fork, write, the UNIX system itself, reached from C through functions declared in headers like <unistd.h>. This book walks both sides, and the sections tell you which side you are standing on: man 2 write and man 3 printf are different kinds of thing, even though both are called from C. (There is even a printf in section 1 — the shell-adjacent coreutils printf — so when a page is not what you expected, give the section number: man 3 printf.)

Search by keyword — same tool, different name — with man -k:

$ man -k 'exit status'
...

On a minimal installation, section 3 pages can be missing; Debian and Ubuntu install them with sudo apt install manpages-dev. You will know within your first week whether you need it.

Two habits worth forming now. First, read the SYNOPSIS at the top of a library page: it shows the #include you need and the function's shape — for fgets, #include <stdio.h> and char *fgets(char *s, int size, FILE *stream) tells you, in one line, the include, the return type, and the argument types. Second, remember that headers are just files: less /usr/include/stdio.h (search with /fgets) shows you the machine's actual declarations, the very text #include pastes into your program. The truth is not hidden; it is in /usr/include, and man is its friendly index.

1.13 C standards and old dialects

C has been standardized for nearly four decades, and you will meet code from all of it. The timeline, compressed to what matters:

Current GCC and Clang default to a recent dialect with GNU extensions (-std=gnu17 or newer — GCC 15 defaults to gnu23); this book compiles with -std=c17, strict and portable, and flags the few places where C23 changes things.

Why the history lesson on day one: you will read old code, on Linux and everywhere, and it was written for an earlier language. The two fossils to recognize:

/* Old-style function definition, pre-1989.  Do not write this;
   shown so you can recognize it in old code. */
int square(x)
int x;
{
    return x * x;
}
/* Implicit int, pre-C99: "main()" with no return type.  GCC 14 and
   newer reject this by default; older compilers only warned. */
main()
{
    ...
}

Both compile in old programs, both are gone from modern practice, and neither should appear in anything you write. The habit this book wants: write modern C — full prototypes, explicit types, (void) parameter lists — and when an old file needs touching, modernize opportunistically, not dogmatically.

1.14 Summary and the road ahead

You can now do the whole loop: write a .c file; compile it with gcc -std=c17 -Wall -Wextra -Werror -g -o program program.c; run it as ./program; read its exit status with echo $?; compose it into pipelines with <, >, and |; and look things up in man. On the language side you have: variables and the basic types, with sizeof, size_t, and stdint.h's exact-width integers; literals, const, enum, and (sparingly) #define; arithmetic with its two traps — integer truncation and undefined signed overflow; the full small set of control flow, with braces and break discipline; printf and its conversions; and a first real program, nl, built on the fgets contract and interoperating with ls through the standard streams.

The road ahead: Chapter 2 goes deeper into the machine's view of the language — every variable's scope and lifetime, the complete set of integer and floating types, conversions and casts, and the full operator set, including the bitwise tools that systems code lives on. Chapter 3 breaks programs into functions — declaration, definition, call-by-value, headers, and multiple .c files, plus main's command-line arguments. Chapters 4–6 are the deep water of C's design: pointers and arrays, then structures and memory management, then strings — the machinery that makes line[256] and "..." and %s fully yours. Chapter 7 crosses to the system's side of the map for files: the whole standard I/O library (FILE * streams, block reads and writes, getline) and, underneath it, the kernel's own interface — open, read, write, close, and stat. From Chapter 8 on come the rest of the system's moving parts — processes and signals, memory in depth, and finally the tools that build real programs (make, libraries, debugging for keeps). Standard C first, then the system: exactly the two halves of a Linux C programmer's education.

1.15 Exercises

The only way to learn a language is to write programs in it — so write programs. Suggested order: 1–4 while the chapter is fresh; the rest as comfort grows. Every exercise compiles under the recommended flags; keep them on.

  1. Personalize hello.c to greet the current user, using getenv ("man 3 getenv"; include <stdlib.h>, which declares it). The shape of it:

    char *name = getenv("USER");     /* NULL if there is no such variable */
    if (name != NULL)
        printf("Hello, %s!\n", name);

    A variable of type char * ("pointer to char") and NULL ("points at nothing") are Chapter 4's business — for now, treat them as the spell that fetches an environment variable.

  2. Add a hexadecimal column to pow2.c with %#llx (the # flag prefixes 0x; the ll matches p's type). Verify at least 2^10 = 0x400.

  3. Watch int run out of room. Change pow2.c's bound to i <= 31 and keep printing p with %20llu; then add:

    unsigned int big = 2147483648u;   /* 2^31 -- one more than INT_MAX */
    printf("%u\n", big);
    printf("%d\n", (int) big);

    Explain both lines of output. (Here int is 32 bits, INT_MAX is 2147483647, and negatives are two's complement.)

  4. Make nl.c report its total by printing n to stderr after the loop ("%d lines\n"), then run ls | ./nl and ./nl < /etc/passwd and watch the summary arrive on the terminal even when stdout is piped into something else — §1.10, experienced firsthand.

  5. Run sizes.c on another architecture. If your distribution supports it, gcc -m32 -o sizes32 sizes.c (Debian/Ubuntu: sudo apt install gcc-multilib) builds a 32-bit version; compare every row against the 64-bit one and explain the pointer row's change.

  6. Man page safari. From man 3 printf: what does %-8.3f do, and what does printf itself return? From man 2 write: which header does it require, and what does its return value mean? (You have just met sections 3 and 2 — §1.12's map.)

  7. Break nl.c, one bug at a time, and read each diagnostic before fixing it: (a) misspell stdio.h; (b) delete the semicolon after int n = 0;; (c) pass n to %s; (d) write if (fgets(line, sizeof line, stdin) = NULL). For (d), note what the compiler says about a single = where == belongs — and get in the habit of reading warnings as free bug reports rather than interruptions.


Chapter 2 · Variables, Data Types, and Operators

Chapter 1 got programs written, compiled, and running, and met each piece of the language just long enough to use it. This chapter slows down and looks hard at the three things every expression is built from: variables (where values live and how long), types (what values mean and how far they can go), and operators (how values combine). C is sometimes described as "portable assembly" — the phrase is unfair to both sides, but it holds this chapter's truth: nothing about data is hidden from you, and nothing is forgiven either. You will see exactly where a variable lives and dies, exactly what the machine promises about each type (and what it pointedly does not promise), and the one conversion trap every C programmer falls into exactly once.

The Linux context matters more here than in Chapter 1: type sizes and limits are the machine's answers, not the language's; the constants that describe them are files on your disk, readable like any other; and the bitwise operators — this chapter's second half — are the working language of file permissions, device flags, and the kernel's own data structures.

The chapter's programs sit in examples/ch02/, same conventions as Chapter 1. Two of them are traps on purpose — uninit.c reads an uninitialized variable, and cmpu.c walks into the signed/unsigned comparison trap — and are compiled without -Werror precisely so their warnings can appear in the text. Everything else compiles clean under the recommended flags.

2.1 Scope, lifetime, and initialization

A variable is four things at once: a name, a type, a location in memory, and a lifetime over which that location belongs to you. The type was Chapter 1's subject; the other three are this section's.

Scope is where in the source text a name is visible. A declaration at the top level of a file has file scope: visible from that line to the end of the file. A declaration inside braces { } has block scope: visible from that line to the closing brace. An inner block may redeclare an outer name — shadowing it: legal, occasionally useful, and a reliable source of confusion. (GCC's opt-in -Wshadow flags every shadow; many projects enable it, and this book suggests you try it for a week and see why.)

Lifetime (the standard says storage duration) is how long the location is yours. Variables declared in a block are automatic: the storage exists while the block runs and evaporates when it ends. Variables declared at file scope are static: the storage exists for the whole program. The two differ in one more way that no beginner should have to learn by accident:

And declaration versus definition: a definition creates the object (int count = 0;); a declaration announces a name and type without creating it (extern int count;). The distinction becomes a tool when programs grow to several files, and Chapter 3 takes it up properly.

All of that in one program:

/* scope.c -- where names are visible */
#include <stdio.h>

int count = 0;                 /* file scope: visible from here down */

int main(void)
{
    int inner = 1;             /* block scope: inside main's braces */

    printf("count is %d\n", count);
    {
        int inner = 2;        /* a new inner: shadows the outer one */
        count = count + 1;
        printf("inner is %d, count is %d\n", inner, count);
    }
    printf("inner is %d, count is %d\n", inner, count);
    return 0;
}
$ gcc -o scope scope.c
$ ./scope
count is 0
inner is 2, count is 1
inner is 1, count is 1

Read the output against the rules. The file-scope count is visible everywhere below its declaration, so the nested block can assign to it — and its change (count = count + 1) outlives the block, which is why the final line still sees 1: the nested block changed an outer object. The two inners never interfere: the outer inner prints 1 after the block ends, untouched by the shadow, which was a different object that died at its closing brace.

Now the uninitialized case, as a program that is wrong on purpose:

/* uninit.c -- wrong on purpose: reading an uninitialized variable */
#include <stdio.h>

int main(void)
{
    int not_set;               /* declared, never given a value */

    printf("not_set is %d\n", not_set);
    return 0;
}

Compiling with the recommended flags (but without -Werror, so the warning can be shown), GCC says:

$ gcc -std=c17 -Wall -Wextra -g -o uninit uninit.c
uninit.c: In function 'main':
uninit.c:8:5: warning: 'not_set' is used uninitialized [-Wuninitialized]
    8 |     printf("not_set is %d\n", not_set);
      |     ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
uninit.c:6:9: note: 'not_set' was declared here
    6 |     int not_set;               /* declared, never given a value */
      |         ^~~~~~~

Note that GCC did two things: named the sin, and — the note: — pointed back at the declaration that made it possible. Run it anyway:

$ ./uninit
not_set is 32759

That number is one machine's garbage, from one run; on your machine, or the next run, or after the next unrelated edit, it will be something else. The value is meaningless — the read is the bug, and the number only proves the program ran. This is the cheapest kind of undefined behavior: it is caught before the program even runs, by a warning that ships in -Wall, and the cure is free. The rules this book keeps:

2.2 The integer types, completely

The complete set of standard integer types, with the x86-64 Linux answers and the language's actual promises side by side:

Type x86-64 Linux The language promises
bool (_Bool) 1 byte stores 0 or 1
char 1 byte exactly 1 byte
signed char 1 byte holds at least −128 … 127
unsigned char 1 byte holds at least 0 … 255
short, unsigned short 2 bytes at least 16 bits
int, unsigned 4 bytes at least 16 bits
long, unsigned long 8 bytes at least 32 bits
long long, unsigned long long 8 bytes at least 64 bits

Read the right-hand column as a contract: the middle column is merely what a 64-bit Linux system happens to answer. sizeof (§1.5) is how you ask; limits.h is how the machine tells you its ranges, and it deserves a program of its own:

/* limits.c -- what the headers promise about this machine */
#include <stdio.h>
#include <limits.h>
#include <float.h>

int main(void)
{
    printf("CHAR_BIT  = %d\n",   CHAR_BIT);
    printf("INT_MIN   = %d\n",   INT_MIN);
    printf("INT_MAX   = %d\n",   INT_MAX);
    printf("UINT_MAX  = %u\n",   UINT_MAX);
    printf("LONG_MAX  = %ld\n",  LONG_MAX);
    printf("ULONG_MAX = %lu\n",  ULONG_MAX);
    printf("LLONG_MAX = %lld\n", LLONG_MAX);
    printf("DBL_DIG   = %d decimal digits\n", DBL_DIG);
    printf("DBL_MAX   = %g\n",    DBL_MAX);
    return 0;
}

On 64-bit x86-64 Linux it prints:

$ gcc -o limits limits.c
$ ./limits
CHAR_BIT  = 8
INT_MIN   = -2147483648
INT_MAX   = 2147483647
UINT_MAX  = 4294967295
LONG_MAX  = 9223372036854775807
ULONG_MAX = 18446744073709551615
LLONG_MAX = 9223372036854775807
DBL_DIG   = 15 decimal digits
DBL_MAX   = 1.79769e+308

Two rows of that table are machines talking, not the language. Run the same program on Windows (MSYS2, MinGW, or Visual Studio) and LONG_MAX prints 2147483647 and ULONG_MAX 4294967295: §1.5's LLP64 lesson again — on Windows only long long reaches 64 bits, so LONG_MAX and LLONG_MAX disagree there, while on Linux and macOS they coincide. Every one of those constants is a macro defined in a real file; open it: less /usr/include/limits.h (search with /INT_MAX) and you are reading the exact text the preprocessor pastes into this program.

Three further details complete the standard types. First, char, signed char, and unsigned char are three distinct types (§1.5): plain char has the representation of one of the other two, chosen by the compiler. Second, stdint.h (§1.5) is not just the exact-width types; it is a family:

Third — and this is the section to reread whenever you choose a type — some rules of thumb that working Linux code follows:

You need Reach for
a general-purpose integer int
a size, length, or index that came from the library size_t
a width that is part of a data format int32_t, uint64_t, …
a raw byte (file contents, buffers, checksums) unsigned char
a character char
truth bool
a signed library size (rare, POSIX) ssize_t
avoid plain long (LLP64 trap), short (nothing to gain)

One more conversion worth naming while types are on the table: assigning any nonzero value to a bool yields true (1) — bool b = 42; stores 1 — and zero yields false. bool is not "any nonzero stored faithfully"; it is a genuine 0-or-1 type.

2.3 Conversions, casts, and the signed/unsigned trap

Values move between types constantly — in assignments, in comparisons, in arithmetic — and C has one set of rules for all of it. The story that covers everything you will meet this year, in two steps:

  1. Integer promotion. In any expression, bool, char, and short operands are first promoted to int (or unsigned int, on the exotic machines where int cannot hold them — not yours).
  2. Common type. If the two operands still differ: when both have the same signedness, the narrower converts to the wider; when they have the same width but different signedness, the unsigned type wins; and an int that meets an unsigned int becomes unsigned itself.

(The complete rule — the standard's "usual arithmetic conversions" — has exactly one more wrinkle, for long and unsigned int on LLP64 systems, and after §2.2 you know precisely why that wrinkle exists.)

Step 2's boldface clause is the trap, and it deserves a program:

/* cmpu.c -- the signed/unsigned comparison trap */
#include <stdio.h>

int main(void)
{
    int i = -1;
    unsigned int u = 1;

    if (i < u)
        printf("-1 is less than 1\n");
    else
        printf("-1 is NOT less than 1: the comparison sees %u\n", (unsigned int) i);
    return 0;
}

Compiled without -Werror (so the warning survives), the compiler announces the problem before the program lies to you:

$ gcc -std=c17 -Wall -Wextra -g -o cmpu cmpu.c
cmpu.c: In function 'main':
cmpu.c:9:11: warning: comparison of integer expressions of different signedness: 'int' and 'unsigned int' [-Wsign-compare]
    9 |     if (i < u)
      |           ^

-Wsign-compare rides along with -Wextra. Run the program:

$ ./cmpu
-1 is NOT less than 1: the comparison sees 4294967295

What happened: the comparison converted i to unsigned int to match u — step 2's boldface clause — and the conversion of −1 is not "clamp to 0" but reinterpretation: the same 32 bits, 1111...1111, now read as an unsigned number: 4294967295. That is two's complement arithmetic doing exactly what it is designed to do, in a context where it silently reverses the meaning of <.

The everyday version of this bug never looks like -1 < 1u: it looks like a size check. if (count - 1 < limit) with count == 0 computes count - 1 as a huge unsigned value (because limit came from the library as a size_t), and the guard admits what it was built to refuse. This is a large fraction of all buffer-overflow bugs, which is why memory-safe code (Chapter 9) checks sizes with care bordering on paranoia, and why you should read every mixed-signedness warning as a stop sign, not a formality.

Conversions you write yourself are casts: (unsigned int) i, (char) c, (double) n. A cast is two statements in five characters — to the compiler, "convert this"; to the reader, "yes, I meant this". Use casts to say exactly what you mean at exactly the points where the implicit rules would do something surprising — as cmpu.c did in its printf, to match the %u conversion honestly. And use the opt-in -Wconversion flag when you want the compiler to question every implicit narrowing, not just the mixed-signedness ones. What a cast must never be is a way to silence a warning you have not understood: that is taping over the smoke detector.

2.4 Floating-point, honestly

C's floating types are float (single precision), double (double precision), and long double (extended precision — on x86-64 Linux, 80 bits in a 16-byte space; rarely needed, occasionally exactly right). All mainstream machines implement them as IEEE 754 binary floating point, and their one property that changes how you write code is this: they store binary fractions, and decimal fractions like 0.1 are not binary fractions. The stored value is the nearest representable double, and "nearest" is close enough for physics and far enough to break ==:

/* floatcmp.c -- floating-point arithmetic, honestly */
#include <stdio.h>
#include <math.h>

int main(void)
{
    double sum = 0.1 + 0.2;

    printf("0.1 + 0.2  = %.17g\n", sum);
    printf("0.3        = %.17g\n", 0.3);
    printf("difference = %g\n", sum - 0.3);

    if (sum == 0.3)
        printf("== says equal\n");
    else
        printf("== says not equal\n");

    if (fabs(sum - 0.3) < 1e-9)
        printf("but the difference is within epsilon: close enough\n");
    return 0;
}
$ gcc -o floatcmp floatcmp.c
$ ./floatcmp
0.1 + 0.2  = 0.30000000000000004
0.3        = 0.29999999999999999
difference = 5.55112e-17
== says not equal
but the difference is within epsilon: close enough

The %.17g conversions print 17 significant digits — one more than DBL_DIG (15, §2.2), so you are seeing the honest value the machine stored, not the 0.3 the default printing would round to. The two honest values differ in the 17th digit, the difference is about 5.5 × 10⁻¹⁷, and == therefore says no. The working rules:

A word on fabs and libraries, paying a debt from §1.4: the math functions live in the math library, and on Linux that library is separate — -lm on the link line. GCC recognizes the common ones (fabs, sqrt, ...) and usually generates them inline, which is why floatcmp.c links with no extra flag; when some day a link fails with undefined reference to a math function, add -lm and move on. Making library use routine is Chapter 3's job.

2.5 The operator set, complete

Here is every operator in the language, loosest-binding last, with the two pointer operators and two structure operators we have not met yet marked for their chapters:

Precedence Operators Associates
1 () (call) [] . -> (Chapter 5), postfix ++ -- left
2 unary ! ~ ++ -- + - * & (Chapter 4), (cast) sizeof right
3 * / % left
4 + - left
5 << >> left
6 < <= > >= left
7 == != left
8 & (bitwise) left
9 ^ left
10 | left
11 && left
12 `
13 ? : right
14 = += -= *= /= %= &= ^= |= <<= >>= right
15 , (comma operator) left

Memorize the table's shape, not its rows: unary binds tightest, then arithmetic, then shifts, then comparisons, then the bitwise trio, then the logical pair, then conditional, then assignment loosest of the meaningful. And when in doubt — §1.7's law holds — parenthesize. The one precedence fact every C programmer must know cold, because it is counterintuitive and silent: comparisons bind tighter than &, ^, and |. v & mask != 0 parses as v & (mask != 0) — bitwise-AND of v with 0 or 1 — which is why every test of a bit in this book is written (v & mask) != 0, parentheses included.

The operators Chapter 1 met briefly, in full now:

Assignment is an expression with a value: a = b = 0 assigns right to left (b = 0 yields 0, stored into a), and if ((c = getchar()) != EOF) — Chapter 4's world — is the idiom that makes it useful. The compound assignments (+=, &=, <<=, ...) all mean "operate and store": p *= 2 from Chapter 1's pow2.c. Assignment's cousins == and != are comparisons; the single-= typo inside an if is the oldest bug in the family, and GCC asks for parentheses whenever you write it.

Relational and equality operators yield int 0 or 1 (§1.7) — and do not chain the way mathematics writes them. a < b < c parses (a < b) < c: the result of the first comparison (0 or 1) is compared with c. "Between" must be spelled out: a < b && b < c.

Logical &&, ||, ! combine conditions, yielding 0 or 1. Their defining property is short-circuit: the left operand is evaluated first, and if it decides the answer (&& with 0, || with nonzero), the right operand is not evaluated at all — a guarantee, not an optimization. Its everyday shape is the guard:

if (n != 0 && total % n == 0)     /* the guard makes the % safe */

The conditional operator ? : is a two-armed expression: condition ? when-true : when-false yields exactly one of the two arms — the other is never evaluated. It is right-associative (row 13), and its natural home is inside larger expressions where an if would be clumsy:

printf("%s\n", ok ? "yes" : "no");
printf("%d%s", (v & mask) != 0, n % 4 == 3 ? " " : "");

The second line is from this chapter's bits.c, below — a conditional as an argument, choosing between two strings.

The bitwise operators are C's machinery room, and the second half of this chapter: & (and), | (or), ^ (exclusive or, either/or), ~ (complement, flip every bit), << and >> (shift left/right). They work on integers as bit patterns, and every one of them is defined and portable when written on unsigned operands. The idioms, in the form you will use them weekly:

You want to Write Why it works
test a bit (v & mask) != 0 the AND keeps only that bit
set a bit v |= mask OR with 1 forces the bit on
clear a bit v &= ~mask AND with the inverted mask forces it off
flip a bit v ^= mask XOR with 1 inverts it
build a mask for bit n 1u << n shift the 1 into position

Two shift rules, both cheap to keep and expensive to break: the shift count must be 0 through width − 1 of the (promoted) left operand — anything else is undefined behavior; and shift unsigned operands only — shifting a negative value, or shifting a 1 into a signed type's sign bit, is undefined too. (Compilers do not warn reliably about shift counts, because the count may be a runtime value; this is one place the rule must live in you.) Shifts double and halve by powers of two, but write * 8 when you mean multiply — the compiler turns it into a shift anyway, and the reader sees your intent.

All of it working together, a program that prints a number's bits:

/* bits.c -- show an integer, bit by bit */
#include <stdio.h>
#include <limits.h>

int main(void)
{
    unsigned int v = 42;
    unsigned int mask = 1u << (CHAR_BIT * sizeof v - 1);
    int n = 0;

    while (mask != 0) {
        printf("%d%s", (v & mask) != 0, n % 4 == 3 ? " " : "");
        mask = mask >> 1;
        n++;
    }
    printf("= %u (0x%x)\n", v, v);
    return 0;
}
$ gcc -o bits bits.c
$ ./bits
0000 0000 0000 0000 0000 0000 0010 1010 = 42 (0x2a)

Walk it once, slowly, because every line is this chapter in miniature:

Check the arithmetic by hand: 42 = 32 + 8 + 2, so bits 5, 3, and 1 are set — 0010 1010. Then verify the last group of the output against perm.c, next.

Because the payoff for all this bit-work is the most Linux thing in the book so far. Every file on a Linux system carries a nine-bit permission mask — who may read, write, or execute it — and ls -l prints it as rw-r--r---style letters. The bits are three groups of three (owner, group, others; read, write, execute), and they are named constants in a real header, <sys/stat.h> — our first POSIX header, arriving three chapters early because it is too good to wait:

/* perm.c -- decode a permission mask the way ls does */
#include <stdio.h>
#include <sys/stat.h>          /* S_IRUSR, S_IWUSR, ... and mode_t */

int main(void)
{
    mode_t mode = 0644;        /* rw-r--r-- */

    printf("mode %04o is ", (unsigned int) mode);
    printf("%c", (mode & S_IRUSR) ? 'r' : '-');
    printf("%c", (mode & S_IWUSR) ? 'w' : '-');
    printf("%c", (mode & S_IXUSR) ? 'x' : '-');
    printf("%c", (mode & S_IRGRP) ? 'r' : '-');
    printf("%c", (mode & S_IWGRP) ? 'w' : '-');
    printf("%c", (mode & S_IXGRP) ? 'x' : '-');
    printf("%c", (mode & S_IROTH) ? 'r' : '-');
    printf("%c", (mode & S_IWOTH) ? 'w' : '-');
    printf("%c", (mode & S_IXOTH) ? 'x' : '-');
    printf("\n");
    return 0;
}
$ gcc -o perm perm.c
$ ./perm
mode 0644 is rw-r--r--

Everything in this chapter is in those lines. 0644 is an octal literal (§1.6: the leading zero), and octal is permission spelling because each digit is three bits — one rwx triple: 6 = 110 = rw-, 4 = 100 = r--. The S_ constants are the named bits — S_IRUSR (IR = read, USR = user) is 0400, the user-read bit, and its eight colleagues complete the nine. mode_t is POSIX's type for exactly this mask (an unsigned integer; its exact width varies by system, which is why the printf passes (unsigned int) mode — a cast that says what we mean to the reader and to %o). Each of the nine lines tests one bit with & and lets the conditional choose 'r' or '-' — the letter version of bits.c's 0 and 1. The whole program is ls's middle third, in twenty lines; chmod 644 file writes these very bits, and Exercise 5 meets the umask idiom that builds masks like this one.

One honest observation before moving on: those nine printf lines are the same line nine times with one letter and one constant changed. You have felt the want of a way to name a repeated shape and reuse it. That thing is the function, it is the single most important construction in the language, and it is — at last — the next chapter.

sizeof, one last time, now with its full behavior: it yields a size_t, it applies to expressions without parentheses (sizeof v, as in bits.c) and to type names with them (sizeof(int)), it is computed at compile time, and it never evaluates its operand's side effects. The comma operator (row 15) evaluates its left operand, discards the value, and yields its right: i++, j++ in a for header is its natural habitat. It is not the comma that separates function arguments — same character, entirely different grammar — and outside for loops it is rarely the clearest way to say anything.

2.6 Summary

You now know, precisely: that a variable is a name, type, location, and lifetime, and that scope says where a name is visible (file, block, shadowing) while storage duration says how long the location is yours — with automatic objects uninitialized (and -Wuninitialized on your side) and static objects zeroed; the complete integer family with its guarantees (limits.h, stdint.h's exact/least/fast/max/pointer sizes, size_t) and the rules of thumb for choosing among them; how conversions work — promotions, the common-type rule, the unsigned-wins clause, the -1 < 1u trap and its everyday size-check disguise — and how casts say what you mean; what floating point honestly does (binary fractions, DBL_DIG, the 0.1 + 0.2 demonstration) and the epsilon discipline that replaces ==; and the complete operator set — assignment as an expression, short-circuit as a guard, ?: as a two-armed expression, the precedence table with its one mandatory fact (& binds looser than comparisons), the bitwise idioms (test, set, clear, flip, 1u << n), the shift rules, and two real programs — bits.c and perm.c — in which all of it earns its keep.

Next: Chapter 3 breaks programs into functions — prototypes, definitions, call-by-value, headers and multiple .c files, main's command-line arguments — and among its first tasks is collapsing perm.c's nine lines into the one shape they all share.

2.7 Exercises

Same rules as Chapter 1: recommended flags on, predict before you run, and read every warning to the end.

  1. Break scope.c deliberately. Move int inner = 1; from main's top into the nested block (deleting the outer inner), predict the compiler's exact complaint before compiling, then confirm. Then put it back and enable -Wshadow; explain in one sentence why projects ban shadowing.

  2. Run bits.c on a permission. Change v to 0644, predict all 32 bits on paper, run, and check yourself. Then match the nine low bits, group by group, against perm.c's rw-r--r--.

  3. Widen bits.c to print an unsigned long long (mask from 1ull, printed with %llu and %llx). Predict the number of groups before running, and explain why 1u would be wrong now.

  4. Cast arithmetic, predicted then verified. What is (unsigned char) -1 and why? What is (signed char) 200 on this machine, and why is the answer −56? Write a six-line program to check both predictions; be able to draw the eight bits of each.

  5. Meet umask. Print 0666 & ~S_IWOTH in octal (%o), decode the result with perm.c's nine lines, and then read man 2 umask enough to explain, in one sentence, what the kernel does with such a mask when a process creates a file.

  6. How many digits does one-third deserve? In a loop over p from 1 to 17, print 1.0 / 3.0 with printf("%.*g\n", p, 1.0 / 3.0); — the * takes the precision from the argument. Find the smallest p whose output you would accept as "a third", and compare it with DBL_DIG.

  7. A classic sin. Write int i = 0; i = i++ + 1; and print i. Compile with -Wall -Wextra and read the -Wsequence-point warning carefully — this is the sequencing rule of §2.5's opening paragraph in the flesh: one statement, one modification. Then write the obviously-correct i = i + 1;, confirm the output, and resolve never to write the first form again.


Chapter 3 · Control Flow and Program Structure

Chapter 1 met every control statement just long enough to use it, and Chapter 2 left a promise unpaid: perm.c's nine repeated lines, and the felt want of a way to name a repeated shape. This chapter pays both debts. Its first half is control flow in full — what a statement is, how multi-way decisions are built, what switch really does, how loops nest, and the one job goto is actually good at. Its second half is program structure — functions, call-by-value, prototypes, recursion, and the anatomy of a program split across several files that build with one command. The two halves meet in the chapter's last program: a small tool, in three files, with a usage message and honest exit statuses — the shape of every real Linux command you have ever typed.

The Linux context is everywhere here. The goto cleanup pattern in §3.5 is the Linux kernel's own house style. The read-a-line-and-dispatch loop in §3.3 is the skeleton of shells and text editors. And the interface/implementation/client file trio in §3.8 is how every C program larger than a page is organized, from git down.

Same conventions as before: programs live in examples/ch03/ (this chapter's capstone is three files, not one), every one was compiled with the recommended flags from §1.11 with zero warnings, and every output block below is what the program actually printed.

3.1 Statements

Everything a C program does at run time happens in a statement, and there are only three kinds to know:

The statements that decide — if, switch, while, for, do — all take a condition, and a condition is any scalar value with one rule: zero is false, anything else is true. Comparisons produce 0 or 1 (§2.5), so if (n != 0) and if (n) mean the same thing — but they do not read the same. This book's rule: with numbers, write the comparison out (if (n != 0) says "if there is a count"); the bare form comes into its own with pointers, in Chapter 4.

One trap closes the section, because it is invisible in exactly the way bugs love. else binds to the nearest unmatched if, regardless of indentation:

if (a > 0)
    if (b > 0)
        printf("both positive\n");
else                              /* pairs with if (b > 0), not if (a > 0)! */
    printf("a only\n");           /* runs when a > 0 and b <= 0 */

The compiler is perfectly happy; the human is fooled by the indentation. Braces settle ownership in writing:

if (a > 0) {
    if (b > 0)
        printf("both positive\n");
} else {
    printf("a not positive\n");
}

This is the braces-always rule of §1.8 doing quiet, permanent work.

3.2 Choosing among many: the else-if chain

Decisions with two outcomes are if/else. Decisions among many outcomes chain elses into a cascade, each test narrowing the field:

/* grades.c -- a multi-way decision, in a function */
#include <stdio.h>

char letter_grade(int score)
{
    if (score >= 90)
        return 'A';
    else if (score >= 80)
        return 'B';
    else if (score >= 70)
        return 'C';
    else if (score >= 60)
        return 'D';
    else
        return 'F';
}

int main(void)
{
    for (int s = 55; s <= 100; s += 5)
        printf("%3d -> %c\n", s, letter_grade(s));
    return 0;
}
$ gcc -o grades grades.c
$ ./grades
 55 -> F
 60 -> D
 65 -> D
 70 -> C
 75 -> C
 80 -> B
 85 -> B
 90 -> A
 95 -> A
100 -> A

Read the cascade as a sequence of ever-lower hurdles. The order is not decoration: 95 satisfies every test in the chain, so the tests must stand most-exclusive first; had we tested score >= 60 first, everyone would be a D. return ends the function on the spot, so only the tests before the match ever run — an else-if chain is a cascade of exits, not a sequence of assignments. And the final else, no condition, is what makes the chain exhaustive: every possible score produces a letter. A multi-way decision without a final catch-all has a crack in it, and one day a value will find it.

letter_grade is the book's first real function: defined above main, called inside main's loop. The whole of §3.6 is about that shape; here, notice only how much it buys immediately — the loop reads as data, the decision reads as policy, and each can change without disturbing the other.

3.3 switch, completely

An else-if chain tests conditions in order. A switch jumps straight to a match. Its controlling expression must be an integer value; its labels must be integer constant expressions — literals, character constants, or enum constants, known to the compiler, never two the same:

enum { CMD_HELP, CMD_QUIT, CMD_RUN };

switch (cmd) {
case CMD_HELP:
    print_usage();
    break;
case CMD_QUIT:
    return 0;
case CMD_RUN:
    run();
    break;
default:
    printf("unknown command %d\n", cmd);
    break;
}

Three properties distinguish switch from a chain of ifs. It jumps, it does not test in sequence: because the labels are compile-time constants, the compiler can build a jump table and go directly to the arm — the price of that speed is the constant-label rule, which is why case n: with a runtime n is a syntax error. Arms fall through: after a label's statements run, execution continues into the next label's statements unless a break (or a return, or the end of the switch) stops it. Falling through from case 'y': into case 'Y': with no statements between (§1.8) is the sanctioned use; every other intentional fallthrough deserves a /* fall through */ comment, and GCC's -Wimplicit-fallthrough (in -Wall) treats an unlabeled one as a warning. break binds to the switch, not to any enclosing loop — the subtlety this chapter's program exists to demonstrate.

That program is the read–dispatch loop: read a line, dispatch on what it starts with, repeat. It is the skeleton of every shell, text editor, and network protocol handler ever written:

/* cmd.c -- a tiny command loop: dispatch on the first character */
#include <stdio.h>

int main(void)
{
    char line[256];

    while (fgets(line, sizeof line, stdin) != NULL) {
        switch (line[0]) {          /* line[0]: the first character */
        case 'q':
            printf("bye\n");
            return 0;
        case '#':
            break;                  /* comments print nothing */
        case 'e':
            printf("echo: %s", line);
            break;
        default:
            printf("unknown command: %s", line);
            break;
        }
    }
    return 0;
}

line[0] is the first character of the line buffer — Chapter 4 owns array indexing properly; here it is "look at the first character", the way you read it. The '#' arm is the break-binds-to-the-switch lesson in one line: its break ends only the switch, and the while rides on to the next line. And 'q' shows the other exit: return 0 from inside loop-and-switch both, leaving main — and the program — immediately. When a command means "stop everything", no amount of break will do; you need return (or, in a deeper program, exit, §1.10).

Run it with a scripted session piped in — the testing habit worth forming on day one:

$ gcc -o cmd cmd.c
$ printf 'hello\ne demo\n# comment\nq\n' | ./cmd
unknown command: hello
echo: e demo
bye

Four lines in; three lines and one silence out. hello matched default; e demo echoed (%s with the line fgets kept, newline included); # comment hit the '#' arm, printed nothing, and — because break only ended the switch — the loop continued to q, which said bye and returned. Interactive use is identical: run ./cmd alone and type the same four lines.

3.4 Loops, completely

Chapter 1 introduced the three loops; here is how to choose among them, and what happens when they nest.

while is "as long as" — including zero times. It is the loop for reading input, because the first read may already find end-of-file: the fgets loop of §1.9 is the canonical shape, and cmd.c above is that loop wearing a dispatch.

for is the counted loop — initialization, test, and step on one line, variable scoped to the loop (§1.8). Use it whenever the number of iterations is the point. The comma operator (§2.5) can run two counters: for (i = 0, j = n - 1; i < j; i++, j--), and an empty test means always: for (;;) is the honest spelling of "loop until a break stops me".

do-while tests after the body, so the body runs at least once — exactly the shape of "act, then decide whether to go again": prompt for a password, process a line, retry a failed operation. It is rare, and right, precisely when that first-time guarantee is what you mean.

Loops nest: an inner loop runs to completion once per outer step. Three nested loops, one bit each, enumerate a whole universe — every combination of three permission bits, which is to say every rwx group a Linux file can carry:

/* table.c -- nested loops enumerate every combination */
#include <stdio.h>

int main(void)
{
    for (int r = 0; r <= 1; r++) {
        for (int w = 0; w <= 1; w++) {
            for (int x = 0; x <= 1; x++) {
                printf("%d%d%d  %c%c%c\n", r, w, x,
                       r ? 'r' : '-', w ? 'w' : '-', x ? 'x' : '-');
            }
        }
    }
    return 0;
}
$ gcc -o table table.c
$ ./table
000  ---
001  --x
010  -w-
011  -wx
100  r--
101  r-x
110  rw-
111  rwx

Read the digits column as what they are: the binary numbers 0 through 7, in order — one octal digit's worth of possibilities. That is Chapter 2's permission arithmetic returning with its hat off: each row is one value of the 6 in 0644 (§2.5's perm.c), and the conditionals turning 0/1 into -/r are the same three, evaluated by perm.c's nine lines. Nested loops and enumeration are the same idea: the outer loops choose the slowest-changing digits, the inner the fastest, and the combination count is the product — three loops of two, eight rows.

Two rules close the section. break and continue affect the innermost loop or switch only — a break in table.c's innermost loop would leave one inner pass, not the whole table; escaping two loops at once takes either a bool done flag (Chapter 2's type doing honest work) or — one way and one way only — the next section's subject. And prefer the loop whose shape matches the sentence you would say aloud: "for each score from 55 to 100" is a for; "as long as there is input" is a while; "ask, then maybe ask again" is a do-while. When the code and the sentence agree, both are easier to trust.

3.5 goto, and the kernel's cleanup pattern

goto label; jumps to a named point in the same function. It is the most maligned word in programming — the 1968 letter that named "considered harmful" was about goto — and the honest position after six decades is precise: unrestricted goto makes control flow untraceable, C's structured statements (§3.2–3.4) cover nearly all of it, and the small remainder is a real, important shape: error unwinding.

The shape: a routine acquires resources one by one — config, then a lock, then a buffer — and any acquisition can fail. On failure, everything already acquired must be released, in reverse. Written with nested ifs, each new resource indents everything after it deeper and duplicates the releases; by the third resource the code is a triangle of copies. The kernel's house style writes it flat instead, with one cleanup path that every failure enters at the right point:

/* cleanup.c -- the goto-unwind pattern: flat code, one cleanup path */
#include <stdio.h>

static int acquire(const char *what, int fail)
{
    printf("acquire %s\n", what);
    return fail;
}

static int run(int fail_at)
{
    if (acquire("config", fail_at == 1))
        goto fail_config;
    if (acquire("lock", fail_at == 2))
        goto fail_lock;
    if (acquire("buffer", fail_at == 3))
        goto fail_buffer;

    printf("work done\n");
    printf("release buffer\n");
fail_buffer:
    printf("release lock\n");
fail_lock:
    printf("release config\n");
fail_config:
    return -1;
}

int main(void)
{
    printf("--- happy path ---\n");
    run(0);
    printf("--- failure at stage 2 ---\n");
    run(2);
    return 0;
}
$ gcc -o cleanup cleanup.c
$ ./cleanup
--- happy path ---
acquire config
acquire lock
acquire buffer
work done
release buffer
release lock
release config
--- failure at stage 2 ---
acquire config
acquire lock
release config

Walk both paths. The happy path (fail_at == 0: acquire always returns 0) acquires all three, does the work, then releases the buffer and falls through the label region — fail_buffer:'s statements, then fail_lock:'s, then fail_config:'s return — releasing in reverse order down one shared path. The failure path enters the label region at the on-ramp for what failed: stage 2's goto fail_lock jumps past work done and release buffer (neither applies — no buffer was acquired) and lands on release config — because when the lock could not be acquired, the config is the one thing that must be unwound. Each label means "enter here when the thing named just failed", and what follows it is exactly the release of what that failure leaves behind.

(The acquire function takes a const char *what — a parameter that receives a string; Chapter 4 owns the type, and this is simply how a function is handed a message to print. The static on both functions makes them file-private, §3.6. And a caution for the tempted: the pattern is forward-only jumps into one cleanup region at function bottom — jumping backward into loops, or sideways past checks you dislike, is exactly what forty years of discipline is against. GCC's -Wall also flags labels nobody jumps to, via -Wunused-label, so dead on-ramps do not linger.)

The standard library has one more escape, far past this chapter's needs: setjmp/longjmp save and restore an entire call-stack position, letting an error path leap out of several functions at once. It exists, it is occasionally exactly right in deep protocol code, and it is not for us — mentioned so the word is recognizable when you meet it. (man 3 setjmp when that day comes.)

3.6 Functions

Chapter 2 ended wanting a way to name a repeated shape. Here it is:

return-type name(parameter declarations)
{
    statements
}

grades.c's letter_grade is one; cmd.c and table.c were pure main only because they had nothing to factor yet. A function packages a computation under a name, gives it inputs (parameters) and an output (return value), and — the property that makes programs buildable — can be called wherever an expression is expected.

The one semantic rule that surprises every newcomer, worth a program of its own: arguments are passed by value. The parameter is a new variable, initialized with a copy of the argument's value. Assigning to it changes the copy, never the caller's original:

/* zerofail.c -- call-by-value: the callee works on a copy */
#include <stdio.h>

void try_to_zero(int x)
{
    printf("inside: x is %d, now zeroing it\n", x);
    x = 0;
    printf("inside: x is now %d\n", x);
}

int main(void)
{
    int n = 42;

    try_to_zero(n);
    printf("back in main: n is still %d\n", n);
    return 0;
}
$ gcc -o zerofail zerofail.c
$ ./zerofail
inside: x is 42, now zeroing it
inside: x is now 0
back in main: n is still 42

x really was zeroed — its x, the copy, which then evaporated when the function returned. Consequences, in order of importance: a function communicates with its caller through its return value (and, soon, by writing through pointers — Chapter 4 exists to make "modify my n" possible, and you can now feel exactly why); call-by-value means the callee cannot corrupt the caller's variables by accident, which is a large part of why C programs remain debuggable; and parameters are ordinary local variables — they can be assigned, shadowed, and passed on, with the caller none the wiser.

The rest of the function rules, compactly:

3.7 Recursion

Functions may call themselves. The discipline is always the same two clauses: a base case that returns without recursing, and a recursive step that moves strictly closer to it. Printing a number's binary digits — Chapter 2's bits.c again, this time without any loop or mask — is a natural fit, because "print v in binary" is "print v / 2 in binary, then print v % 2":

/* digits.c -- recursion: print a number in binary */
#include <stdio.h>

void print_bits(unsigned int v)
{
    if (v < 2) {                    /* base case: one digit left */
        printf("%u", v);
        return;
    }
    print_bits(v / 2);              /* all the earlier digits... */
    printf("%u", v % 2);            /* ...then this one */
}

int main(void)
{
    print_bits(42);
    printf("\n");
    print_bits(255);
    printf("\n");
    print_bits(1);
    printf("\n");
    return 0;
}
$ gcc -o digits digits.c
$ ./digits
101010
11111111
1

Trace print_bits(42) once, slowly, because the output order is the lesson. It cannot print 42's last bit yet — the earlier bits must come first — so it calls print_bits(21) and waits; that calls print_bits(10); then 5; then 2; then 1, which is the base case and prints 1 and returns. Now the waiting calls resume, most recent first, each appending its own last digit: 0, 1, 0, 1, 0. The digits appear most-significant-first with no array, no reversal pass — the call stack itself is the array. Each call has its own v, its own printf position, its own place to resume; the recursion depth is the number of digits (six for 42); and the unwinding returns are what produce the left-to-right order.

When to reach for recursion: when the data is nested — trees, directories within directories, expressions within expressions — recursion's shape matches the problem's shape (Chapter 5 will walk a structure that has this property). When the data is flat — scan a buffer, count lines, sum a column — a loop is clearer, cheaper, and cannot run out of stack. Recursion costs a stack frame per level, and "a million-deep recursion" is a crash (Chapter 9 explains the stack properly); factorial, the traditional textbook demo, is linear data wearing a costume, and a loop does it better.

3.8 Many files, one program

Programs outgrow one file fast — not because of length, but because of roles. The shape that carries C programs from one page to a million lines is three kinds of file: an interface (a header, .h) that declares what a module offers; an implementation (a .c) that defines it; and a client (another .c) that uses it through the interface. The chapter's capstone tool, iseq — print an integer sequence, like a pocket seq — is deliberately small, but wears the full costume:

/* iseq.h -- interface of iseq: print an integer sequence */
#ifndef ISEQ_H
#define ISEQ_H

/* Print the integers start..stop, one per line, stepping by step.
   Returns 0 on success, or -1 if the arguments describe no sequence
   (step is 0, or the range runs against the step's direction). */
int iseq(int start, int stop, int step);

#endif
/* iseq.c -- implementation of iseq */
#include <stdio.h>
#include "iseq.h"

int iseq(int start, int stop, int step)
{
    if (step == 0)
        return -1;
    if (start < stop && step < 0)
        return -1;
    if (start > stop && step > 0)
        return -1;

    for (int n = start; step > 0 ? n <= stop : n >= stop; n += step)
        printf("%d\n", n);
    return 0;
}

Four things in those twenty lines deserve their own paragraphs.

The prototype. int iseq(int start, int stop, int step); — a semicolon where the body was — is a declaration: it promises the compiler "this function exists, with exactly these types" without saying what it does. It is how a call in one file is checked against a definition in another. In Chapter 1's terms (§1.4): each .c compiles alone, against declarations; the linker joins the definitions later.

The include guard. #ifndef ISEQ_H / #define ISEQ_H ... #endif makes the header safe to include twice — and it will be, directly and transitively, the day after a program grows past three files. Function declarations tolerate repetition; the typedefs and structure definitions coming in Chapters 4–6 do not. Guards cost three lines and prevent a class of error messages famous for their length, so every header gets one, always, from the first day — the macro named after the file, in caps, by convention.

One definition, many declarations. iseq.c defines iseq; the header, included by iseq.c itself and by the client below, declares it. Including your own interface is not vanity: it makes the compiler verify that your definition matches your published promise, in the file that would otherwise never see it. Functions are shared across files by default (the linker resolves the name); sharing variables across files needs the extern keyword, and waits until Chapter 5, when we have data worth sharing. And a rule with no exceptions: never #include a .c file. It compiles, once, and then the same definitions exist in two translation units and the linker rejects the duplicate — a beginner classic, and the error message names both copies.

The build. Two .c files, one command:

$ gcc -std=c17 -Wall -Wextra -Werror -g -o iseq iseq_main.c iseq.c

Every recommended flag, unchanged; gcc compiles each .c to its own object file and links them — §1.4's pipeline, exercised for real. (The two-step form — gcc -c each file, then link the .os — is exactly what §9.5's make automates.)

3.9 main's arguments, usage, and exit statuses

The client file is the smallest of the three, and the densest with convention:

/* iseq_main.c -- iseq: print integers from first to last */
#include <stdio.h>
#include <stdlib.h>              /* atoi */
#include "iseq.h"

int main(int argc, char *argv[])
{
    if (argc != 3) {
        fprintf(stderr, "usage: iseq first last\n");
        return 2;
    }
    if (iseq(atoi(argv[1]), atoi(argv[2]), 1) != 0)
        return 1;
    return 0;
}

main(int argc, char *argv[]) is the second form of main promised in §1.2. argc is the argument count, including the program's own name; argv holds the arguments as strings — argv[0] is the name the program was invoked by, argv[1] through argv[argc - 1] are the arguments, and argv[argc] is a NULL sentinel (§1.9's "points at nothing"). With ./iseq 1 5: argc is 3, argv[1] is "1", argv[2] is "5". The pointers themselves are Chapter 4's subject; today they are the spell that hands you the command line.

The shape of the body is the convention, and it is worth reading as a checklist. Check the count first. An argument you did not receive is not "", it is absent, and indexing into argv past argc is exactly the out-of-bounds class §2.3 warned about. On bad usage: a usage message on stderr — one line, the command's name, its arguments in the order it wants them — and exit status 2, §1.10's "you used the command wrong", here made real in our own tool. Then do the work, and translate its result into the exit status the shell will see: 1 for "ran, but the work failed", 0 for success.

atoi ("man 3 atoi") converts the argument string to an int — the string-handling details are Chapter 6's, and here is its honesty clause: atoi("x") returns 0 and reports nothing. Garbage in, zero out. Our tool accepts that for now (Exercise 6 hardens it) and flags the debt; a real tool must validate its input, and Chapter 6 shows how strtol does it properly.

Now the tool at work — every exit status demonstrated:

$ ./iseq 1 5
1
2
3
4
5
$ ./iseq -3 2
-3
-2
-1
0
1
2
$ ./iseq 7 7
7
$ ./iseq 5 1
$ echo $?
1
$ ./iseq 1
usage: iseq first last
$ echo $?
2

Read the transcripts against the contract in iseq.h. 1 5 and -3 2 (negative arguments are ordinary integers) run the sequence; 7 7 is the one-element case. 5 1 prints nothing and exits 1 — our documented decision that a range running against the step's direction is an error, checked by iseq before any printing (a tool that prints half its output before failing has worse manners than one that fails first). Real GNU seq decides differently — seq 5 1 is an empty sequence, exit 0 — and which choice is right is a genuine design question; we made ours in writing, in the header, where the client can read it. Exercise 6 asks you to switch sides.

3.10 Summary

Control flow, complete: statements come in three kinds (expression, null, block), conditions are "zero or not", and else binds to the nearest if — braces make ownership visible. Multi-way decisions are cascades of elses, ordered most-exclusive first and closed with an unconditional final else. switch is a jump to a constant label — fast, with the constant-label price — with fallthrough discipline and a break that binds to the switch itself, as the '#' arm of the read–dispatch loop demonstrated. Loops are chosen by the sentence you would say; nested loops enumerate combinations, break/continue touch only the innermost, and two-level escapes want a flag or the one disciplined goto: the kernel's flat, single-cleanup-path error-unwind, entered at the label named for what just failed. Functions: defined as return type, name, parameters, body; arguments passed by value — the callee works on copies, so results travel by return (until pointers, next chapter); every non-void path returns; static makes a function file-private; prototypes declare interfaces, and interface/implementation/client is how programs of every size are cut. Recursion is base case plus a step that shrinks — right for nested data, wasteful for flat. And a real tool: main(int argc, char *argv[]), count checked first, usage on stderr, exit 2 for bad usage, 1 for failed work, 0 for success — the conventions every Linux command already follows.

Next: Chapter 4 — pointers and arrays, the machinery underneath argv, line[0], %s, and the whole "back in main, n is still 42" story: how a function can reach back into its caller's variables, and what a pointer is that makes it honest.

3.11 Exercises

Recommended flags on; predict before running; read every warning and every man page to the end.

  1. Close grades' crack. letter_grade returns a letter for every int, including −20 and 140, which no real course issues. Make out-of-range scores return '?', widen the test loop to run from −10 to 110 in steps of 10, and predict the new table before running it.

  2. Count cmd's work. Add a counter of commands processed (everything except 'q' and '#'), and have cmd print the total to stderr when it exits — including on the 'q' path. Then re-run the scripted session from §3.3 and check the total against the transcript by hand.

  3. table, two groups deep. Rewrite table.c to print, in one loop over i from 0 to 63, the two-permission-group rows 0644-style — the high group's bits are i / 8, the low group's are i % 8 — with both rwx-letter renderings on each line. Predict the first three rows before running, then compare the readability of one-division-loop versus eight-nested-rows and write down which you would rather debug.

  4. Pay Chapter 2's promise. Write static void print_group(mode_t mode, mode_t r, mode_t w, mode_t x) — three bit-tests and three printfs of 'r'/'w'/'x' versus '-' — and collapse examples/ch02/perm.c's nine lines into three calls: print_group(mode, S_IRUSR, S_IWUSR, S_IXUSR); and its two colleagues. Output must be byte-identical to Chapter 2's.

  5. Generalize the recursion. Turn print_bits into print_digits(unsigned v, unsigned base) for any base 2 through 10 — the base case and the two printfs change, the shape does not — and print 42 in every base from 2 to 10, one line each, from a loop in main.

  6. Harden iseq, twice. (a) Accept an optional third argument, the step: ./iseq first last [step]; reject a step of 0 with the usage message and exit 1 (note that argc can now be 3 or 4). (b) Change the empty-range policy to GNU seq's — print nothing, exit 0 — and update the contract comment in iseq.h to match, because a changed promise in the interface is part of the change.

  7. Split, then sabotage, grades. Split grades.c into grade.h (the prototype), grade.c (the definition), and grades_main.c (the loop), and build with one gcc command naming both .c files. Then delete #include "grade.h" from grade.c, predict the diagnostic, and compile: modern GCC refuses the implicit declaration outright — §1.13's implicit-int story, function edition — and the error is the compiler insisting on exactly the interface discipline this chapter built.


Chapter 4 · Pointers and Arrays

Every chapter so far has used pointers by proxy. argv arrived as a spell in Chapter 1 and was demoted to "Chapter 4's business". line[256] was "a buffer" whose first character we peeked at in Chapter 3. const char *what was "how a function is handed a string". And Chapter 3 ended on a cliffhanger: try_to_zero worked on a copy, and n was still 42 — with the promise that this chapter is where a function learns to reach back. The proxies end here.

Pointers are C's model of memory itself — a machine address, with a type attached — and on Linux that model is the native language of the whole system: the kernel hands your process a virtual address space, the loader places your program inside it, ASLR shuffles the layout every run so an attacker cannot guess it, and the moment your program touches an address that is not its own, the kernel kills it with the crash every C programmer knows by name. Arrays are the other half of the same idea: a contiguous run of memory with one type. The two are braided together in C — most uses of an array become a pointer on contact — and the braid is the source of both the language's power and its most famous bug class.

Same conventions: programs in examples/ch04/, all compiled under the recommended flags (two trap programs — buggy.c and dangle.c — compile clean or with exactly their expected warning and crash at run time), and every output block is captured program output, with platform notes where this capture machine's answers differ from a Linux system's.

4.1 What a pointer is

A pointer is the address of an object, together with the type of whatever lives there. Two operators make the idea usable:

A pointer's declaration spells the type it points at: int *p; reads, right to left around the *, as "p is a pointer to int". The pointer itself is a variable like any other — it has a size (sizeof(int *) is 8 on x86-64 Linux; §1.5's sizes program printed it), an address of its own, and a value: the address it holds.

        p                          n
  +--------------+          +--------------+
  | 0x7f9c..a1c4 |--------->|      42      |
  +--------------+          +--------------+
  "p holds the address        "n is an int
   of n"                       living at 0x7f9c..a1c4"

A program to make it concrete:

/* addr.c -- a pointer is an address with a type attached */
#include <stdio.h>

int main(void)
{
    int n = 42;
    int *p = &n;
    int *q = p + 1;

    printf("n lives at %p\n", (void *) p);
    printf("*p is %d\n", *p);
    printf("p + 1 is %p\n", (void *) q);
    printf("q - p is %td (measured in ints, not bytes)\n", q - p);
    return 0;
}
$ gcc -o addr addr.c
$ ./addr
n lives at 0000005DE31FF74C
*p is 42
p + 1 is 0000005DE31FF750
q - p is 1 (measured in ints, not bytes)

Read the output slowly; four lessons are in it.

Addresses are values, printable like any other. %p prints a pointer; the idiom is to cast to (void *) first, because %p is only specified for that type. (Rendering varies by C library: glibc, on Linux, prints 0x7ffc8a1b2c34 — lowercase with a 0x prefix, and NULL as (nil); this capture, on a Windows toolchain, prints padded uppercase hex. Both are the same address in different typography.)

*p is n. Not "a copy of n", not "the value of p": the expression *p is the object at that address. *p = 99; would store 99 into n through p, and n — the name — would see it.

p + 1 is not "the next byte". Compare the last hex digits: ...74C and ...750 — four bytes apart, exactly one sizeof(int). Pointer arithmetic counts in elements, not bytes, because the type attached to the pointer tells it what an element is. That is the whole point of the type being attached.

q - p is 1. Pointer subtraction measures the distance between two pointers into the same object, in elements — its type, ptrdiff_t, is printed with %td. (Note what q is: a pointer one past n — legal to compute and compare, never to dereference, a rule §4.4 makes central.)

And one thing the program cannot show but you should do: run it twice. The addresses differ every run. ASLR — address space layout randomization, a Linux security feature — places the stack at a random offset each time precisely so that no one can predict where your variables live. Pointers are how your program navigates a map that is redrawn at every launch.

4.2 The payoff: a function that reaches back

Chapter 3's zerofail.c demonstrated the wall: arguments are passed by value, so the callee's x was a copy, and zeroing it left the caller's n untouched. The wall is real and it stays — but there is a door in it. Do not pass the variable; pass its address. The address is copied like any value, but a copy of an address still points at the same object:

/* swap.c -- pointers: a function that reaches into its caller */
#include <stdio.h>

void swap(int *a, int *b)
{
    int t = *a;                   /* t gets the value where a points */
    *a = *b;                      /* write through a */
    *b = t;                       /* write through b */
}

int main(void)
{
    int x = 1;
    int y = 2;

    printf("before: x = %d, y = %d\n", x, y);
    swap(&x, &y);
    printf("after:  x = %d, y = %d\n", x, y);
    return 0;
}
$ gcc -o swap swap.c
$ ./swap
before: x = 1, y = 2
after:  x = 2, y = 1

Trace the door. swap(&x, &y) passes two pointer values — the addresses of x and y. Inside swap, a and b are copies of those addresses (call-by-value, still true, forever) — but copies of an address arrive pointing at the original objects. *a is x; *b is y; and the three-line dance through t exchanges their contents. The caller's variables changed, and that is not a violation of call-by-value: it is call-by-value, of an address.

This is the output parameter pattern, and it is everywhere once you know to see it. It is why scanf("%d", &n) is spelled with an & (the input function writes through the pointer you hand it — scanf fills your variables; we meet it properly in Chapter 7), why Chapter 2's mode_t masks could be passed for decoding, and why every C API that must hand back more than one value does it through pointer parameters. A function that wants to give you something it computed returns it; a function that wants to fill something of yours takes its address.

4.3 Arrays

An array is a fixed number of objects of one type, laid out contiguously — shoulder to shoulder in memory, no gaps:

int nums[4] = { 10, 20, 30, 40 };

declares four adjacent ints and initializes them; int more[4]; declares four uninitialized ones (garbage until written, §2.1); int zeros[4] = {0}; zeroes all four. Each element is reached by indexing: nums[0] is 10, nums[3] is 40. Indexing counts from zero for the same reason a ruler does: the index is the distance from the start — nums[0] is the element at the beginning, zero ints away; §4.4 turns that sentence into an identity.

Two facts about arrays and sizeof, and then the fact that changes everything:

The fact that changes everything is decay: in almost every expression, an array becomes a pointer to its first element. nums in an expression is &nums[0]. The exceptions are few and famous — sizeof nums (the whole array, as above) and &nums (the address of the array as one object) — and decay is why arrays and pointers feel like the same thing in C. They are not: the array is the storage; the pointer is a note about where it starts. A program to watch the seam:

/* arr.c -- arrays, sizes, and the pointer they become */
#include <stdio.h>

static void show(int *a, int len)
{
    printf("inside show: sizeof a = %zu (a pointer now, not the array)\n", sizeof a);
    for (int i = 0; i < len; i++)
        printf("  a[%d] = %d\n", i, a[i]);
}

int main(void)
{
    int nums[4] = { 10, 20, 30, 40 };
    int count = sizeof nums / sizeof nums[0];

    printf("in main: sizeof nums = %zu (the whole array)\n", sizeof nums);
    printf("in main: the element count is %d\n", count);
    show(nums, count);
    return 0;
}
$ gcc -o arr arr.c
$ ./arr
in main: sizeof nums = 16 (the whole array)
in main: the element count is 4
inside show: sizeof a = 8 (a pointer now, not the array)
  a[0] = 10
  a[1] = 20
  a[2] = 30
  a[3] = 40

The same object, two answers. In main, nums is the array — 16 bytes, and the count divides out. The moment nums is passed to a function — the moment it appears in an expression — it decays: show's a is a pointer, 8 bytes, and sizeof a tells you about pointers, not about your data. From which follows the rule that shapes the entire standard library: a function receiving an array knows nothing about its length — because it does not receive the array. It receives the address of the first element, and the length travels separately, as a second parameter: show(nums, count).

You have been reading this contract since Chapter 1: fgets(line, sizeof line, stdin) is (where to start, how much room, which stream) — pointer, size, context. Now you know why the size is a parameter and why sizeof appears at the call site (where the array is still an array) rather than inside some function (where it would be too late).

4.4 Pointer arithmetic

Because a pointer carries its element type, arithmetic on it moves in elements. The complete rule set is short:

A program that walks an array with a pointer instead of an index — and demonstrates the identity on its last line:

/* walk.c -- walking an array with a pointer */
#include <stdio.h>

int main(void)
{
    int nums[5] = { 2, 4, 6, 8, 10 };
    int *end = nums + 5;           /* one past the last element */

    for (int *p = nums; p < end; p++)
        printf("%s%d", p == nums ? "" : " ", *p);
    printf("\n");
    printf("nums[2] is %d and *(nums + 2) is %d\n", nums[2], *(nums + 2));
    return 0;
}
$ gcc -o walk walk.c
$ ./walk
2 4 6 8 10
nums[2] is 6 and *(nums + 2) is 6

The loop reads as a sentence: "for each p, starting at nums, while there is room, step forward" — and *p is "the element we've reached". The separator trick (p == nums ? "" : " ") prints no space before the first element, spaces between. And the last line prints the same element twice, once as nums[2] and once as *(nums + 2): one element, two spellings, by definition.

Now the boundary rules, which are the difference between correct code and the most famous crash class in the language. Pointer arithmetic is defined only for positions from the array's first element through exactly one past its last:

Position Compute? Compare? Dereference?
any element yes yes yes
one past the last yes yes never
before the first never — —

The one-past-the-end pointer is a deliberate gift: a stopping place — end in walk.c is nums + 5, which points at no element that exists, and that is fine, because nothing ever dereferences it; p < end only compares. Every bounded loop in this book is built on that corner. What is not fine is end's neighbors: dereferencing one-past-the-end, or computing any pointer before the start or beyond one-past-the-end, is undefined behavior — and "just one element over", the off-by-one error, is the single most common way C programs die. The arithmetic is cheap to write and unforgiving to get wrong; treat the bounds as part of the loop's sentence, and let ASan (§4.7) police the fences while you learn.

4.5 Passing arrays to functions

Decay has one more consequence worth its own section, because it is the shape of every interface you will ever write against: since a called function receives a pointer, arrays travel as (pointer, length) pairs. The library's contracts — fgets(buffer, size, stream), show(a, len) — are the pattern; write your own the same way:

static int sum(const int *a, int len)      /* read-only walk: §4.6 */
{
    int total = 0;
    for (int i = 0; i < len; i++)
        total += a[i];
    return total;
}

Three rules complete the picture. Read or write, the pointer decides: sum only reads, but a function with an int *a (no const) may write through it — a[i] = v; stores into the caller's array, exactly as swap stored into the caller's x. Never sizeof-hunt for the length inside a function: arr.c proved it reports the pointer's size; the length is a parameter or it is nowhere. And never return a pointer to a local — the reverse direction of the same wall. swap reached in because the caller's object outlives the call; a local dies when its function returns, and a pointer to it points at a ghost:

/* dangle.c -- wrong on purpose: returning a local's address */
#include <stdio.h>

int *where(void)
{
    int x = 42;
    return &x;                     /* x dies at the closing brace */
}

int main(void)
{
    int *p = where();

    printf("*p is %d\n", *p);      /* undefined behavior */
    return 0;
}

Compiled without -Werror (so the warning survives), GCC catches it statically — -Wreturn-local-addr rides with -Wall:

$ gcc -std=c17 -Wall -Wextra -g -o dangle dangle.c
dangle.c: In function 'where':
dangle.c:7:12: warning: function returns address of local variable [-Wreturn-local-addr]
    7 |     return &x;                     /* x dies at the closing brace */
      |            ^~

Run it anyway, and the undefined behavior gets its turn:

$ ./dangle
Segmentation fault
$ echo $?
139

On this machine, the leftover stack frame was already gone, and the very first read through p crashed. On another machine, the frame's remains may still hold 42 — and the program prints a plausible, wrong answer — or §2.1's garbage (32759). All of those outcomes are equally "correct", because undefined behavior promises nothing; what is stable is the cause, and the warning names it before anything runs. Returning a local's address is not a rule you memorize; it is a wall you can now see — automatic storage dies with its function (§2.1), and no pointer survives what it points at. (The chapter's one legitimate way out — asking for memory that outlives the call — is malloc, Chapter 5's subject.)

4.6 const and pointers

Chapter 3 handed a string to a function as const char *what and deferred the type. No more deferrals — const with pointers is a two-by-two grid, and the way to read any of it is right to left, from the name outward:

Declaration Reads as What is protected
const int *p p is a pointer to a const int the pointee: *p = 5; is rejected; p = other; is fine
int *const p p is a const pointer to int the pointer: p = other; is rejected; *p = 5; is fine
const int *const p p is a const pointer to a const int both

const is not a promise about storage — the object may be writable through some other pointer; and it is not security — a cast can shed it. It is a contract, read by the compiler: "no write through this name". Which makes it documentation that cannot rot, and it earns its keep in parameters. sum(const int *a, int len) announces in the prototype that it reads your array and will not write it; the caller of sum needs to read nothing to know it. The idiom extends naturally: const char *what is "a string I will only read"; a function that intends to modify your string takes char * and says so. A library that marks its read-only parameters const is telling you, in its interface, exactly what every call will touch.

The reading discipline, once more, because it generalizes past const: start at the name, read rightward for arrays/functions, leftward for pointers and qualifiers. char *argv[] — right to left — is "an array ([]) of pointers (*) to char", i.e. one pointer per argument, each pointing at a string's first character. That is the machinery of the chapter's last subject.

4.7 NULL, and the crash every C programmer knows

NULL is a pointer value defined to point at no object — a distinguished "nowhere", comparable to anything and dereferenceable by nothing. Since pointers are "zero is false" values (§3.1's promise), the idiom for "does this pointer point somewhere?" is direct:

if (p != NULL)      /* written by everyone as: */
if (p) {            /* ...use *p */ }

Both spellings mean the same thing; the bare form is the idiom and this book uses it from here on. The discipline that goes with it: a pointer that can be NULL is checked before it is dereferenced — fgets's return (§1.9) and getenv's (§1.15) were both "check first, then use", and now you know what the check was for.

And when the discipline slips — when NULL or a stale address is dereferenced — the kernel takes over, and it is worth meeting the crash as a designed mechanism, not folklore:

/* buggy.c -- wrong on purpose: dereferencing NULL */
#include <stdio.h>

int main(void)
{
    int *p = NULL;

    printf("about to dereference %p\n", (void *) p);
    *p = 42;                       /* the line the kernel objects to */
    printf("never gets here\n");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -g -o buggy buggy.c
$ ./buggy
about to dereference 0000000000000000
Segmentation fault
$ echo $?
139

(One honesty note about that transcript: in an interactive terminal the program's own line — "about to dereference" — appears first, because stdout is line-buffered there. In our scripted capture, with stdout redirected to a file, stdout was block-buffered (§1.10), the buffered line was still in the buffer when the kernel killed the process, and the log recorded only the shell's report. Same crash, same 139; buffering decided what reached the log. It is §1.10's quiet lesson, arriving with a crash attached — and it is the reason diagnostic output belongs on stderr, which is unbuffered.)

Walk the mechanism, because every part of it is something this book has already built. The address 0 is not mapped into your process (§4.1: your variables live at randomized, kernel-chosen addresses — page zero is deliberately left empty). The write *p = 42 therefore touches memory that is not yours, and the kernel — which manages each process's virtual address space and enforces its boundaries — delivers SIGSEGV, "segmentation violation", the signal from §1.7's division-by-zero family (signals arrive properly in Chapter 8). The default disposition of SIGSEGV is: kill the process. Bash, the parent, notices a child that died of signal 11 and reports it — the Segmentation fault line is bash's message, not the program's (run buggy under a script and the message changes shape; the exit status never does) — and $? records 139 = 128 + 11: death by signal n is reported as 128 + n, a formula worth keeping (§1.10 kept only the low 8 bits; the signal encoding lives in the high bit).

Two practical corollaries. First, "Segmentation fault (core dumped)" on Linux means the system wrote a core file — a post-mortem image of the process at the moment of death — which a debugger can open to find the exact faulting line (whether cores are written is a policy setting: ulimit -c, and on systemd systems, coredumpctl; Chapter 9 does the forensics). Second, the sanitizers catch this class before the kernel does: recompile with -fsanitize=address (§1.11's line) and the same run stops at the faulting access with a report naming the file, the line, and the access — on a full Linux toolchain; we note honestly that our capture machine's GCC install lacks the ASan runtime (ld: cannot find -lasan), so no report is shown here, and Chapter 9 returns to reading them properly.

One closing reframe. To a newcomer a segfault feels like failure; to a working programmer it is the system working — the kernel catching, in hardware, a process that broke its own boundaries, before it could touch anyone else's memory. The crash you should actually fear is dangle.c's other face: the pointer that happens to land on mapped memory and prints a plausible wrong number forever. Loud and dead beats silent and wrong; it is a theme the rest of this book keeps returning to.

4.8 The typeless pointer: void *

void * is a pointer with no element type attached: "an address, period". It cannot be dereferenced (* would not know what to produce) and does no arithmetic (nothing to scale by) — but it converts, without a cast, to and from any object pointer type. That makes it the standard currency for "a pointer to something whose type is not my business": malloc's return ("some memory — you know what you put in it", Chapter 5), qsort's comparator arguments (Chapter 5), printf's %p (why this book always casts to (void *)), and function parameters like memcpy's "copy from here to there". Its integer counterpart from §2.2, uintptr_t, is the round-trip type — (uintptr_t) p yields an integer wide enough to hold the address, for the rare APIs that traffic in addresses-as-numbers. void * trades type for generality and expects the type back at the edges; use it where interfaces demand it, not as a way to stop deciding.

4.9 argv, explained at last

Chapter 3 used main(int argc, char *argv[]); §4.6 read the declaration; here is the whole machine, on the table:

/* args.c -- argv, finally explained */
#include <stdio.h>

int main(int argc, char *argv[])
{
    printf("argc is %d\n", argc);
    for (int i = 0; i < argc; i++)
        printf("argv[%d] is \"%s\" at %p\n", i, argv[i], (void *) argv[i]);
    printf("argv[%d] is %p (the NULL sentinel)\n", argc, (void *) argv[argc]);
    return 0;
}
$ gcc -o args args.c
$ ./args one two 3
argc is 4
argv[0] is "E:\...\ch04\args.exe" at 000001701F3F7E70
argv[1] is "one" at 000001701F3F8150
argv[2] is "two" at 000001701F3F5750
argv[3] is "3" at 000001701F3F5770
argv[4] is 0000000000000000 (the NULL sentinel)

(Captured on a Windows shell, which passes the program's full path as argv[0] — abbreviated here to fit the page; a Linux shell passes the string you typed, ./args. The addresses are real and differ every run; argv[4]'s all-zero rendering is this platform's NULL — glibc on Linux prints (nil).)

Now the anatomy. argv is an array of pointers to char — one pointer per command-line word, each pointing at that word's first character (a string, Chapter 6's business: char sequence, '\0'-terminated). argv[0] is the program's own name, as the launcher spelled it. argv[1] through argv[argc - 1] are the arguments in order. And argv[argc] is NULL — a sentinel, so that a walker can find the end the same way walk.c's loop found its end, by meeting it:

for (char **p = argv; *p != NULL; p++)     /* p is a pointer to (pointer to char) */
    printf("argument: %s\n", *p);

Read that type right to left: p is a pointer to (char *) — a pointer to one of the array's elements, each of which is a pointer to a string. *p is the element: one char *, an argument. *p != NULL is the fence; p++ steps to the next word. No argc needed: the sentinel is the end marker, and this "terminated by a NULL pointer" convention is the same one the environment block uses (environ, Chapter 8's territory) — NULL-termination is a UNIX habit that runs deeper than strings.

One idiom to close on, because it is the last unpaid pointer debt. §2.5 promised that assignment-as-expression has a use worth the confusion; here it is, the two-idioms-in-one read loop of half the software ever written in C:

int c;
while ((c = getchar()) != EOF)     /* fetch one character, THEN test it */
    ...

The parentheses are mandatory (!= binds tighter than =, §2.5's table); the shape is "assign, then compare the result" — one line that fetches, stores, and decides, using a pointer-shaped fact (a condition is any scalar) and a constant (EOF, "end of file", a distinguished non-character value from stdio) that Chapter 7 explains. Chapter 6 uses exactly this loop, unremarked, until it is second nature.

4.10 Summary

A pointer is an address with a type attached — &x takes the address, *p reaches the object (read and write), and declarations read from the name outward. Addresses are values: printable (%p), copyable, different every run (ASLR), and arithmetic on them counts in elements — p + n steps n elements, p - q measures ptrdiff_t distance, a[i] is *(a + i), and all of it is defined only from the first element through one-past-the-end (computable, comparable, never dereferenceable). Call-by-value still holds everywhere — but a copied address still points at the original, which is swap, scanf's &, and the entire output-parameter pattern. Arrays are contiguous typed storage that decay into pointers to their first element in almost every expression — so sizeof a whole array works only where the array still is one, and functions receive (pointer, length) pairs, never the length by stealth; the library's every interface (fgets, show, sum) is that pair. const makes the read/write contract explicit in the prototype — pointee (const int *), pointer (int *const), or both — read right to left. NULL is "nowhere": check with if (p) before dereferencing; when the check is skipped the kernel answers with SIGSEGV, exit status 128 + 11 = 139, bash's "Segmentation fault", maybe a core dump — and ASan, where available, names the line before the kernel ever sees it. void * is the typeless address that interfaces trade in; argv is an array of pointers to strings whose argv[argc] is a NULL sentinel that makes even the end of the command line walkable. And returning a local's address is the wall in the other direction: automatic storage dies with its function, -Wreturn-local-addr names the sin, and the ghost it leaves behind is the quiet version of the same crash.

Next: Chapter 5 — structures and memory management: records that give data a shape (struct, members, nesting, arrays of structures), the machine's arithmetic underneath them (alignment, padding, offsetof), the heap that sizes data at run time (malloc, free, and the ownership discipline), qsort's generic machinery through void *, and a real growable array built as a three-file module.

4.11 Exercises

Recommended flags on; predict before running; remember that addresses differ every run, and read every diagnostic to the end.

  1. rotate3. Chapter 3's try_to_zero failed; swap succeeded. Write void rotate3(int *a, int *b, int *c) — a's value to b, b's to c, c's to a — and a main that proves it with before/after prints. Predict the output before running.

  2. Scale-check. Extend addr.c with double d and double *dp = &d, and char c and char *cp = &c. Before running, write down what (dp + 1) - dp and (cp + 1) - cp will print, and what the %p gap between dp and dp + 1 will be in bytes. Then run and check both predictions against the output.

  3. sum and fill. Write int sum(const int *a, int len) (§4.5's body, made real) and void fill(int *a, int len, int step) that stores step * i into a[i]. Demonstrate both from main on a five-element array — fill, then print, then sum — and note which parameter is const and why that matches what each does.

  4. Walk backwards. Change walk.c to print the array in reverse using for (int *p = end - 1; p >= nums; p--). Explain, in one sentence each: why end - 1 is the correct starting pointer; why p >= nums is the correct fence; and what would be wrong with continuing to p == nums - 1.

  5. Two loops over argv. Print all of argv twice from args.c's main: once by index (for i < argc), once by sentinel (for (char **p = argv; *p != NULL; p++)). Confirm identical output, then explain in two sentences why the second loop needs no argc at all.

  6. Crash forensics. Run buggy.c; verify 139 = 128 + 11 with echo $?. Then, on a Linux system: check ulimit -c (is a core configured?); compile with -fsanitize=address and read the report's first lines — what file, line, and access does it name? Finally, look at your own process's memory map with cat /proc/self/maps and find the [stack] line: that is where §4.1's randomized addresses live.

  7. The const grid, enforced. Declare const int *p;, int *const q = &n;, and const int *const r = &n; (with a real int n), then attempt each of *p = 1;, q = other;, *r = 1; as separate compiles with -Werror. Read each error message verbatim, and match it to a row of §4.6's table: the compiler is enforcing exactly the contract each declaration made.


Chapter 5 · Structures and Memory Management

So far every variable has held one thing: one number, one character, one pointer. But real data does not come in single values — it comes in records. A file is a name, a size, an owner, and nine permission bits (§2.5). A process is a pid, a state, an owner, an address space. The kernel's own task_struct — the record at the heart of every running process — has hundreds of members. C's word for a record is struct, and this chapter is about building them, passing them, laying them out byte by byte, and sorting them.

The second half of the chapter is the other thing real data does: it arrives in quantities you do not know when you write the program. A program that reads ten thousand log lines cannot declare char line[10000][256] and hope. The tool for "as much memory as the situation turns out to need" is malloc, and the discipline that keeps it honest — ownership, checking, freeing — is the difference between programs that run for years and programs that quietly eat a server.

This chapter also pays a remarkable pile of debts: the . and -> rows of §2.5's precedence table, Chapter 3's promise that a structure would come along whose shape earns recursion, Chapter 3's deferred extern, Chapter 4's "the one legitimate way out — asking for memory that outlives the call", and §4.8's qsort comparator. Same conventions as ever: programs in examples/ch05/, compiled under the recommended flags (one trap program, useafter.c, compiles with exactly its expected warning), every output block captured from a real run.

5.1 Structures: data with a shape

A structure packages values of different types under one name, as one object:

struct point {
    int x;
    int y;
};

This declares a type, not a variable: from now on struct point is a type name, as usable as int. The parts inside are members. A variable of the type is declared like any other, initialized with a brace list in member order, and its members are reached by name with the dot operator:

struct point a = { 1, 2 };
a.x = 5;                      /* the x member of a */

Structures nest — a member can be a structure — and whole structures assign and copy by value, member by member, which the first program demonstrates end to end:

/* point.c -- structures: data with a shape */
#include <stdio.h>

struct point {
    int x;
    int y;
};

static struct point make_point(int x, int y)
{
    struct point p = { x, y };
    return p;                    /* returned by value: a copy */
}

static void move_point(struct point *p, int dx, int dy)
{
    p->x += dx;                  /* p->x is (*p).x */
    p->y += dy;
}

int main(void)
{
    struct point a = make_point(1, 2);
    struct point b = a;          /* copies every member */

    move_point(&b, 10, 10);
    printf("a is (%d, %d)\n", a.x, a.y);
    printf("b is (%d, %d)\n", b.x, b.y);
    printf("sizeof(struct point) is %zu bytes\n", sizeof(struct point));
    return 0;
}
$ gcc -o point point.c
$ ./point
a is (1, 2)
b is (11, 12)
sizeof(struct point) is 8 bytes

Three observations, one per lesson. b = a copied all the members at once — structure assignment is a deep-by-value copy of the object (its bytes), so b moved and a stayed put: Chapter 3's call-by-value, now for whole records. move_point took a pointer — because passing the address of a 8-byte record is cheap, and because it reaches back, exactly as swap did in §4.2. And that p->x is new syntax doing Chapter 2's deferred work: the arrow operator p->x is exactly (*p).x — "the x of the thing p points at". Dot for values, arrow for pointers, same precedence row, always written without spaces (p->x, never p -> x).

One more piece of vocabulary: typedef gives a type a second name — typedef struct { int x; int y; } point; would let you write point a;. This book's rule, mirroring much of the Linux world: define structures with a tag (struct point), and reach for typedef when a module exports an opaque-ish handle type — a decision §5.7 will make for you.

5.2 Structures and functions

The rules of §3.6 apply to structures unchanged, with one engineering judgment added: size. A structure argument is copied, whole, on every call — right for small records, wasteful for large ones. The standing shapes:

And records rarely travel alone. An array of structures is C's oldest and most durable data design — the in-memory table — and it inherits every array fact from Chapter 4: contiguous, indexed, decaying to a pointer when passed, its length traveling beside it:

/* crew.c -- an array of structures, like a little database */
#include <stdio.h>

struct date {
    int day;
    int month;
    int year;
};

struct crew {
    const char *name;            /* points at a string literal (Chapter 6 owns strings) */
    struct date born;            /* a structure inside a structure */
    int uid;
};

static void print_crew(const struct crew *c)
{
    printf("%-8s uid %4d  born %d-%02d-%02d\n",
           c->name, c->uid, c->born.year, c->born.month, c->born.day);
}

int main(void)
{
    struct crew roster[] = {
        { "root",   {  1,  1, 1970 },     0 },
        { "daemon", {  7,  6, 1992 },     1 },
        { "linus",  { 28, 12, 1969 },  1000 },
    };
    int count = sizeof roster / sizeof roster[0];

    printf("the roster has %d members\n", count);
    for (int i = 0; i < count; i++)
        print_crew(&roster[i]);
    return 0;
}
$ gcc -o crew crew.c
$ ./crew
the roster has 3 members
root     uid    0  born 1970-01-01
daemon   uid    1  born 1992-06-07
linus    uid 1000  born 1969-12-28

Read the initializer as a little database: an array of records, each a brace list — name, nested date, uid — the count idiom (sizeof roster / sizeof roster[0]) from §4.3, the (pointer, length) loop from §4.4's world, and the %-8s/%4d/%02d field widths from §1.3 lining the table up. (The names point at string literals — Chapter 6 takes strings apart; today a literal in braces is simply a value.) This shape — array of structures, a count, functions over const pointers — is the design of half the software you use, from /proc parsers to configuration tables, and the rest of the chapter builds on it three times: sorting it (§5.6), laying it out (§5.4), and making it grow (§5.7).

5.3 A structure that points at itself

A structure member may be a pointer to the same structure type — the sentence that unlocks every linked data structure in C:

struct node {
    int value;
    struct node *next;         /* the link: NULL marks the end */
};

The self-pointing member is legal because the compiler needs only the declaration, not the size, to know "pointer to struct node" (§1.5: 8 bytes, whatever it points at). Chains of such nodes are linked lists: each node points at the next; the last points at NULL. And the walk over one is Chapter 4's sentinel loop, verbatim:

for (struct node *p = head; p != NULL; p = p->next)
    printf("%d ", p->value);

p is a pointer to the current node; p->next is the pointer to the following one; NULL is the fence — the same convention as argv's sentinel (§4.9), which is no coincidence: a char * array ending in NULL and a list of nodes ending in NULL are the same idea in different clothes. This is the structure whose shape earns recursion, as Chapter 3 promised: a list is either empty or a node followed by a list, and functions over it are written exactly that way — iteratively here, recursively in Exercise 4, where you will also meet the ownership question (§5.5) that lists raise more sharply than anything else in the language.

5.4 Layout: alignment and padding

A structure is not just its members — it is its members plus what the machine insists on adding between and after them. Every type has an alignment: the addresses it is fast (or, on some machines, possible) for the CPU to load it from — an int from a multiple of 4, a double from a multiple of 8, a char from anywhere. The compiler places each member at the next offset that satisfies its alignment, inserting padding bytes where needed, and rounds the whole structure's size up to a multiple of its largest member's alignment — so that in an array of structures (§5.2), every element's members stay aligned. A program that shows the machine's arithmetic, using offsetof — from <stddef.h>, it reports a member's byte offset within a structure:

/* pad.c -- alignment and padding: what the machine adds */
#include <stdio.h>
#include <stddef.h>              /* offsetof */

struct mixed {                  /* members in a scattered order */
    char  tag;                  /* 1 byte  */
    int   count;                /* 4 bytes */
    char  flag;                 /* 1 byte  */
};

struct sorted {                 /* same members, large to small */
    int   count;
    char  tag;
    char  flag;
};

int main(void)
{
    printf("sizeof(struct mixed)  = %zu\n", sizeof(struct mixed));
    printf("sizeof(struct sorted) = %zu\n", sizeof(struct sorted));
    printf("offset of tag  in mixed:  %zu\n", offsetof(struct mixed, tag));
    printf("offset of count in mixed:  %zu\n", offsetof(struct mixed, count));
    printf("offset of flag  in mixed:  %zu\n", offsetof(struct mixed, flag));
    printf("offset of tag  in sorted:  %zu\n", offsetof(struct sorted, tag));
    return 0;
}
$ gcc -o pad pad.c
$ ./pad
sizeof(struct mixed)  = 12
sizeof(struct sorted) = 8
offset of tag  in mixed:  0
offset of count in mixed:  4
offset of flag  in mixed:  8
offset of tag  in sorted:  4

Draw struct mixed as the machine lays it out, byte by byte:

offset:  0     1  2  3    4  5  6  7    8     9 10 11
member: tag | pad...... |   count    |  flag | pad....

tag takes offset 0; count needs a multiple of 4, so bytes 1–3 are padding; flag lands at 8; and the structure's size rounds 9 up to a multiple of 4 — three more pad bytes at the tail, present so that mixed[1].count in an array is aligned too. struct sorted — the same members, ordered largest first — needs one byte of internal padding (the tail round-off) instead of six, and is 8 bytes, not 12: a third of the structure was pure member-ordering waste. That is the practical rule, and the Linux kernel states it outright in its style guide: order members large to small (pointers and 8-byte values, then 4-byte, then chars) when size matters, and let -Wpadded (opt-in, like -Wshadow) report every pad byte your orderings create.

Two closing notes on layout. Never compare structures with == — it is a constraint violation; structures are compared member by member (a function like by_height, §5.6, is where such comparisons live). And there is a cousin of padding for bits: bit-fields — unsigned int busy : 1; declares a member exactly one bit wide, several of which pack into one word. They are the traditional spelling of flag registers and inode-style metadata, and they carry two cautions worth stating once: their order within a word is implementation-dependent (a wire-format contract wants §2.5's explicit masks instead), and you cannot take a bit-field's address (no & — there is no byte for it to name). Use them for flags in records; use masks when the layout is the data.

5.5 The heap: malloc, free, and ownership

Chapter 2 met storage durations; now the third one arrives, and with it the reason it exists. Recall the map so far: automatic storage (block-scoped variables) lives on the stack and dies with its block — the wall dangle.c ran into in §4.5; static storage (file scope) lives in the data segment for the whole program but its size is fixed at compile time. What is missing is memory that is sized at run time and outlives the function that asked for it. That is the heap, and its door is malloc:

void *malloc(size_t size);          /* "give me size bytes" */
void  free(void *ptr);              /* "I am done with this" */
void *calloc(size_t n, size_t size); /* "n elements, zeroed" */
void *realloc(void *ptr, size_t size); /* "resize this" -- §5.7 */

The contract, read carefully, has four clauses. It returns void * — §4.8's typeless address, converting implicitly to whatever you assign it to. It returns NULL if it cannot — on Linux, when the kernel refuses to grow your address space — and checking is not optional: a program that dereferences the unchecked result has written buggy.c from §4.7 with extra steps. The memory is uninitialized — garbage, §2.1's rule — unless you use calloc, which zeroes. And the memory is yours until you free it — no other event releases it: not the function returning, not the variable going out of scope, nothing. The first program, in full:

/* heap.c -- memory that lives until you say otherwise */
#include <stdio.h>
#include <stdlib.h>             /* malloc, free */

int main(void)
{
    int n = 5;
    int *scores;

    scores = malloc(n * sizeof *scores);
    if (scores == NULL) {
        fprintf(stderr, "heap: out of memory\n");
        return 1;
    }
    for (int i = 0; i < n; i++)
        scores[i] = (i + 1) * (i + 1);
    for (int i = 0; i < n; i++)
        printf("scores[%d] = %d\n", i, scores[i]);
    free(scores);
    return 0;
}
$ gcc -o heap heap.c
$ ./heap
scores[0] = 1
scores[1] = 4
scores[2] = 9
scores[3] = 16
scores[4] = 25

Every idiom in it is Chapter 4 come home to roost: malloc(n * sizeof *scores) — the count times the element size, with sizeof *scores so the expression survives a type change; scores == NULL checked before the first use; scores[i] — the returned pointer indexes like any array, because it points at contiguous memory (this is what "the heap" means: malloc's region, unrelated to the data structure of the same name); and free(scores) — exactly one, at the end.

"Exactly one" is the heart of heap discipline, and it is called ownership: every allocation has exactly one owner, responsible for freeing it exactly once. The rules that follow from it, all learned the hard way by everyone:

Now the two classic sins, each with its own texture of wrong. First, the leak — memory never freed:

/* leak.c -- wrong on purpose: memory that is never freed */
#include <stdio.h>
#include <stdlib.h>

int main(void)
{
    for (int round = 0; round < 3; round++) {
        int *p = malloc(100 * sizeof *p);
        if (p == NULL)
            return 1;
        p[0] = round;
        printf("round %d: allocated 100 ints, forgot to free\n", round);
    }
    return 0;                    /* 300 ints leak, silently */
}
$ gcc -o leak leak.c
$ ./leak
round 0: allocated 100 ints, forgot to free
round 1: allocated 100 ints, forgot to free
round 2: allocated 100 ints, forgot to free

Compiles clean under -Werror. Runs perfectly. Exits 0. That is the problem: the leak is invisible in every output the program can produce — three iterations leak 300 ints, and nothing complains. (A footnote of honesty: at exit the kernel reclaims the whole address space anyway, so a program that leaks and then exits looks fine — the leak that matters is in a program that keeps running: a server that leaks per request is dead in a week and looks healthy every step of the way.) Detection is a tool's job, not a printf's: on Linux, ASan includes LeakSanitizer, which reports the exact allocation stack of everything unreclaimed at exit — the same honesty clause as §4.7 applies here: our capture toolchain lacks the ASan runtime, so we state the behavior rather than show it, and Chapter 9 works the reports properly.

Second sin: the use-after-free — and here the compiler gets to show off, because modern GCC catches this one statically:

/* useafter.c -- wrong on purpose: using freed memory */
#include <stdio.h>
#include <stdlib.h>

int main(void)
{
    int *p = malloc(sizeof *p);

    if (p == NULL)
        return 1;
    *p = 42;
    printf("before free: *p is %d\n", *p);
    free(p);
    printf("after free:  *p is %d\n", *p);   /* undefined behavior */
    return 0;
}

Compiled without -Werror so the warning survives — and note it is a plain -Wall warning these days; the tools improve underneath you:

$ gcc -std=c17 -Wall -Wextra -g -o useafter useafter.c
useafter.c: In function 'main':
useafter.c:14:5: warning: pointer 'p' used after 'free' [-Wuse-after-free]
   14 |     printf("after free:  *p is %d\n", *p);   /* undefined behavior */
      |     ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
useafter.c:13:5: note: call to 'free' here
   13 |     free(p);
      |     ^~~~~~~

Run it anyway:

$ ./useafter
before free: *p is 42
after free:  *p is 1378178912
useafter exit status: 0

1378178912 — not 42, not garbage-of-the-week, but allocator bookkeeping: the freed block was already recycled into free-list metadata, and the read through the dead pointer surfaced the allocator's own notes. Exit status 0, "success" — §4.7's closing reframe, replayed on the heap: the use-after-free that prints a plausible number is worse than one that crashes, and the only reason this program confessed at all is that the heap reuses memory aggressively. The discipline is the ownership list above; the police are -Wuse-after-free (compile time) and ASan (run time, where available) — and the habit is treating a freed pointer exactly like a returned-local's address: dead.

5.6 qsort: generic code through void *

The standard library's sort — qsort, from <stdlib.h> — sorts any array, of any element type, which it can only do by combining every idea of the last two chapters: decay (it receives a void *, not an array), the (pointer, count, size) contract (count and element size travel as parameters, because from a void * it cannot know either), and §4.8's typeless pointer (your element type re-enters through your comparator). Its signature, read once, slowly:

void qsort(void *base, size_t nmemb, size_t size,
           int (*compar)(const void *, const void *));

base is the array to sort (in place — it returns nothing); nmemb is how many elements; size is each element's width; and the fourth parameter is a function pointer — "a pointer to a function that takes two const void * and returns int" — your code, handed over as data, called back for every comparison. (Function pointers get their full treatment with signal handlers, Chapter 8; today, "pass the function's name and it is called" is all the machinery needed.) The comparator's contract: return negative, zero, or positive as the first element sorts before, equal to, or after the second. Sorting an array of §5.2's structures by a member:

/* peaks.c -- qsort: sorting structures through void * */
#include <stdio.h>
#include <stdlib.h>             /* qsort */

struct peak {
    const char *name;
    unsigned int height;        /* meters */
};

static int by_height(const void *a, const void *b)
{
    const struct peak *pa = a;  /* void * converts to any object pointer */
    const struct peak *pb = b;

    if (pa->height < pb->height)
        return -1;
    if (pa->height > pb->height)
        return 1;
    return 0;
}

int main(void)
{
    struct peak list[] = {
        { "K2",            8611 },
        { "Everest",        8849 },
        { "Kangchenjunga", 8586 },
        { "Lhotse",        8516 },
    };
    int count = sizeof list / sizeof list[0];

    qsort(list, count, sizeof list[0], by_height);
    for (int i = 0; i < count; i++)
        printf("%-14s %u\n", list[i].name, list[i].height);
    return 0;
}
$ gcc -o peaks peaks.c
$ ./peaks
Lhotse         8516
Kangchenjunga  8586
K2             8611
Everest        8849

The comparator is the whole lesson. Its parameters arrive as const void * — the caller's element type erased for transit (§4.8) — and the first line of its body restores it: const struct peak *pa = a;, the implicit void * conversion from §4.8 doing exactly the job it exists for. It then answers the three-way question in the safest possible spelling — two comparisons and a constant, never arithmetic. And that spelling is a rule, because the tempting one-liner return pa->height - pb->height; is a trap: for unsigned members the subtraction wraps (§2.7), and for ints near the limits it is signed overflow — either way, a comparator that lies to qsort in rare corners produces a sort that is wrong in ways no test with small numbers will ever find. The two-if shape costs three lines and cannot lie; sort keys that are already guaranteed narrow (small enums) can afford subtraction, and nothing else can.

The output is sorted ascending, in place, and the list was never copied: qsort rearranged the caller's array. §4.8's promise — "qsort's comparator arguments (Chapter 5)" — is paid; what you own now is the standard pattern for any generic operation in C: erase the types at the boundary, carry (count, size) beside the pointer, and restore the type in a callback's first line.

5.7 A growable array, properly

Everything in the chapter now assembles into one design — the most-used dynamic data structure in working C: the growable array (a "vector"), built as a module in §3.8's three-file shape. The design insight is a structure (§5.1) that owns a heap block (§5.5) and carries its own bookkeeping: the data pointer, how many elements are in use, and how many fit before growing. As a header, in full — this is the interface, the ownership documents, and §3.8's declaration/definition discipline, all in nineteen lines:

/* intvec.h -- a growable array of ints */
#ifndef INTVEC_H
#define INTVEC_H

#include <stddef.h>             /* size_t */

typedef struct {
    int *data;                   /* the elements (NULL until first push) */
    size_t len;                 /* how many are in use */
    size_t cap;                 /* how many fit before growing */
} intvec;

/* Append value; returns 0, or -1 if memory ran out (v unchanged). */
int intvec_push(intvec *v, int value);

/* Free the elements and reset v. Safe on a zeroed intvec. */
void intvec_destroy(intvec *v);

#endif
/* intvec.c -- implementation of intvec */
#include <stdlib.h>             /* realloc, free */
#include "intvec.h"

int intvec_push(intvec *v, int value)
{
    if (v->len == v->cap) {                 /* out of room: grow */
        size_t ncap = v->cap == 0 ? 8 : v->cap * 2;
        int *ndata = realloc(v->data, ncap * sizeof *ndata);
        if (ndata == NULL)
            return -1;                      /* old data still valid */
        v->data = ndata;
        v->cap = ncap;
    }
    v->data[v->len] = value;
    v->len = v->len + 1;
    return 0;
}

void intvec_destroy(intvec *v)
{
    free(v->data);                          /* free(NULL) is a safe no-op */
    v->data = NULL;
    v->len = 0;
    v->cap = 0;
}

Walk intvec_push, line by line, because every line is a chapter's worth of decisions. When full, double — cap == 0 ? 8 : cap * 2: starting at eight and doubling makes a thousand appends cost a dozen reallocations total, amortized cheap, instead of one per push. realloc resizes in place when it can and moves when it must — it returns the (possibly new) address of a block holding the old contents, copied to the new location if it moved. Assign to a temporary first — the classic trap this code exists to demonstrate: v->data = realloc(v->data, ...) would overwrite the only pointer to the old block with NULL on failure, losing the data forever; the ndata temporary means failure leaves v completely intact. And note what failure does not do: it does not free, does not clear, does not half-apply — the old array stays valid, push returns -1, the caller decides. The element write is §5.5's heap arithmetic: v->data[v->len], index by the in-use count, then count it in. intvec_destroy is the ownership contract's other half — and its safety on a zeroed vector is earned, resting on free(NULL) being a no-op and on intvec v = { 0 }; (§2.1's zero-init) making an empty vector a legitimate object from its first breath.

The client — thin, per §3.6, and honest on failure per §3.9:

/* vec_main.c -- vec: fill a growable array and print it back */
#include <stdio.h>
#include <stdlib.h>             /* atoi */
#include "intvec.h"

int main(int argc, char *argv[])
{
    intvec v = { 0 };           /* zeroed: safe to destroy from moment one */
    int want = argc > 1 ? atoi(argv[1]) : 10;
    int rc = 0;

    for (int i = 0; i < want; i++) {
        if (intvec_push(&v, i * i) != 0) {
            fprintf(stderr, "vec: out of memory\n");
            rc = 1;
            break;
        }
    }
    printf("%zu squares:", v.len);
    for (size_t i = 0; i < v.len; i++)
        printf(" %d", v.data[i]);
    printf("\n");
    intvec_destroy(&v);
    return rc;
}
$ gcc -std=c17 -Wall -Wextra -Werror -g -o vec vec_main.c intvec.c
$ ./vec 8
8 squares: 0 1 4 9 16 25 36 49
$ ./vec 0
0 squares:
$ echo $?
0

Two mains of note. intvec v = { 0 }; — a structure initialized to all-zeroes — is the pattern that makes every dynamic structure safe to destroy (or push to, in designs where empty means zero) from the moment of its birth, and it is the correct C replacement for "constructor" thinking. And the failure path: push failing is not exceptional bookkeeping, it is a normal exit-status story — report on stderr, exit 1, destroy on the way out (note destroy runs on both paths, the success and the early exit; that is what ownership looks like in a main).

One more design rule, stated now because this is the chapter's last structure: the vector owns its buffer; nobody else holds pointers into it. When realloc moves the block, every pointer into the old block — &v.data[3], a saved element pointer, an iterator in another function — dies silently. A module that owns memory must either guarantee stability (moving is documented as invalidating) or hand out indices, which survive moves. intvec hands out indices (v.data[i]) and says so in its header; every long-lived C interface you will ever use makes this same choice, in writing, or regrets it.

5.8 extern: sharing data across files

Chapter 3's last deferral was this: functions are shared across files automatically, but variables are not — a definition in one .c is invisible in another unless declared there. The word is extern, and the pattern is exactly §3.8's, with data in place of functions: define once, in the implementation; declare everywhere, in the header. A module that owns a program-wide settings record:

/* opts.h -- the program's options, owned by opts.c */
#ifndef OPTS_H
#define OPTS_H

struct opts {
    int verbose;
    int max_lines;
};

extern struct opts opts;        /* defined in opts.c -- one definition */

void opts_defaults(void);
#endif
/* opts.c */
#include "opts.h"

struct opts opts;               /* THE definition: zero-initialized (§2.1) */

extern struct opts opts; announces "this name exists, with this type, defined somewhere the linker will find" — a declaration, not a definition; it creates no storage. The opts.c line without extern is the definition — there must be exactly one in the whole program (two, and the linker rejects the duplicate, exactly as with §3.8's double iseq). Every .c that includes opts.h reads and writes the same record. The standing caution travels with the power: global data is shared state, and shared state is where multi-module programs go wrong — the reason §5.7's intvec keeps its buffer private (static, §3.6, for helpers; plain module data only when the module is about the data) and hands out access through functions. Global variables are for program-wide facts — the name of the game, the parsed configuration — not for convenience plumbing between two functions that should be taking parameters.

5.9 Summary

Structures give data a shape: members under one name, nested, initialized with brace lists, copied whole by assignment and by-value passing — with pointers (->, (*p).x) as the efficient and mutable route, const struct * as the read-only contract, and arrays of structures plus a count as C's fundamental in-memory table. A structure member may point at its own type — the linked list, walked by Chapter 4's sentinel loop with NULL as the fence, recursion's natural habitat. Layout is not abstract: every type has an alignment, the compiler pads members to it and rounds the whole struct to the largest — offsetof shows the arithmetic, member order is a size decision (large-to-small, the kernel's rule), bit-fields pack flags but not wire formats, and structures are compared member by member, never with ==. The heap supplies what neither stack nor static storage can: memory sized at run time and owned across calls — malloc (check NULL, contents uninitialized), calloc (zeroed), realloc (resize, may move — assign to a temporary; failure must leave the old block alive), free (exactly once, of exactly what you were given, with free(NULL) a safe no-op) — all governed by ownership: one owner per allocation, documented in the interface, with leaks silent and use-after-free complicit (-Wuse-after-free and ASan as the police). qsort genericizes over all of it through void * and (count, size) parameters, calling back into a comparator that restores types in its first line and answers with two ifs and a constant — never subtraction. The growable array assembles the whole chapter into one module: a structure owning (data, len, cap), doubling capacity, realloc through a temporary, a zero-init that makes empty vectors legal objects, a destroy that frees and resets, and the stability rule — the module owns the buffer; hand out indices, not pointers. And extern completes the module system: one definition in the implementation, declarations in the header, global data reserved for global facts.

Next: Chapter 6 — strings: the '\0' convention underneath every quoted literal this book has printed, the <string.h> library (strlen, the copy family and why strcpy needs a safer successor, strcmp and its three-way truth), strtol paying Chapter 3's atoi debt, and the tools Chapters 1–5 kept promising.

5.10 Exercises

Recommended flags on; predict before running; the padding exercises reward paper first.

  1. Rectangles. Add struct rect { struct point ul, lr; }; (upper-left and lower-right corners) to point.c's world, plus int area(const struct rect *r). Initialize a rectangle, print its area, and — before running — predict sizeof(struct rect).

  2. Designated initializers. C99 lets initializers name their members: { .uid = 1000, .name = "linus", .born = { 28, 12, 1969 } }. Rewrite crew.c's roster with designated initializers in a different order than the declarations, confirm identical output, then delete .uid from one row and predict the printed value of that member before running (§2.1 has the answer).

  3. Predict the padding. On paper, lay out struct { char a; double b; char c; }; — every offset, every pad byte, the total size — then verify with offsetof and sizeof in a small program. Explain the tail padding to someone else; if you cannot, §5.4's array-alignment sentence is the missing sentence.

  4. A list, by hand, then honorably discharged. Using §5.3's struct node, build a three-node list (values 1, 2, 3) with malloc, walk-print it with the sentinel loop, then free it correctly — saving p->next before each free(p), and explaining in one sentence why freeing before saving next is a use-after-free (§5.5's second sin, in list clothing).

  5. Hunt the leak, on Linux. Run leak.c and confirm: no output betrays anything. Then recompile with -fsanitize=address (LeakSanitizer rides along) and read the report — it names the exact allocation line and its stack. Finally, fix the leak (free before the loop's next iteration) and confirm the report is silent. (Our capture machine lacked the ASan runtime; a Linux toolchain has it.)

  6. A bounds-checked getter. Add to intvec: int intvec_get(const intvec *v, size_t i, int *out) — 0 and *out set if i < v->len, −1 otherwise. It is §4.2's output-parameter pattern guarding §4.4's fence; demonstrate with an in-range and an out-of-range call.

  7. Two comparators. (a) Flip by_height to sort descending — predict the first line before running. (b) Compute, on paper, what pa->height - pb->height returns for the pair (1, 4294967295) as unsigned values, and explain in two sentences why that answer sorts those two peaks backwards. §5.6's two-if shape is now a rule you have seen break, not one you were told about.


Chapter 6 · Strings and Character Manipulation

Every printf("...") since Chapter 1, every argv word in Chapter 4, every line fgets ever handed this book — all of it was strings, borrowed on trust. The trust is repaid here. A string in C is one convention, learned in a minute and respected for a career: characters in memory, ended by a null character. On that convention stands the whole <string.h> library, the character classification of &lt;ctype.h>, and — because UNIX decided that its universal interchange format would be text — a healthy fraction of everything a Linux system does: configuration files, pipelines, /proc, the output of every tool you have ever piped into another.

The debts this chapter pays are by now a tall stack: the '\0' underneath every literal (§1.6); the repair of Chapter 1's nl.c, whose too-small buffer split long lines into numbered fragments (§1.9's "the honest repair waits for the string tools of Chapter 6"); atoi's honesty problem (§3.9); the getchar read loop "used unremarked" since §4.9; and qsort, which sorted ints and structs in Chapter 5 and now sorts words. Same conventions as ever: seven programs in examples/ch06/, six of them compiled clean under the recommended flags, one (copy.c) compiled with exactly its one expected warning — a warning worth a section of its own — and every output block below captured from a real run.

6.1 What a string is

A string is a sequence of characters in memory, terminated by the null character '\0' — byte value zero, the same one §1.9 promised would "return in Chapter 6". The terminator is part of the string's storage and no part of its length: "Linux" is 5 characters in 6 bytes, and every string function in the library — printf's %s, strlen, fgets — finds the end by looking for the zero. Nothing else marks it: no length field, no object header. The convention is the whole data structure.

index:  0   1   2   3   4   5    6    7
      +---+---+---+---+---+----+----+----+
word:  | L | i | n | u | x |\0 | \0 | \0 |   char word[8] = "Linux";
      +---+---+---+---+---+----+----+----+
                         ^    ^---------- sizeof word: the buffer (8)
      strlen stops here -|               (5: the characters, not the terminator)

The two lengths are the first thing to keep straight, and a program to hold them apart:

/* len.c -- a string, its terminator, and its two lengths */
#include <stdio.h>
#include <string.h>

int main(void)
{
    char word[8] = "Linux";       /* 6 bytes: L i n u x \0, plus 2 spare */
    char *lit = "literal";        /* a pointer to a string literal */

    printf("strlen(word) = %zu (characters, not counting '\\0')\n", strlen(word));
    printf("sizeof word  = %zu (the whole buffer)\n", sizeof word);
    printf("word[5]      = %d (the terminator, as an int)\n", word[5]);
    printf("strlen(lit)  = %zu\n", strlen(lit));
    printf("sizeof lit   = %zu (the pointer, Chapter 4)\n", sizeof lit);
    return 0;
}
$ gcc -o len len.c
$ ./len
strlen(word) = 5 (characters, not counting '\0')
sizeof word  = 8 (the whole buffer)
word[5]      = 0 (the terminator, as an int)
strlen(lit)  = 7
sizeof lit   = 8 (the pointer, Chapter 4)

Four facts, one per line of output. strlen counts characters, stopping at (not including) the terminator — the string's length. sizeof counts the whole object — the buffer, spare bytes and terminator included, and it works only where the array is still an array (§4.3's decay rule: sizeof lit measured the pointer, 8 bytes, because lit is one). word[5] is 0 — the terminator is an ordinary byte at an index you can print, inspect, and (as §6.7's repair will do) test. And there is no fifth fact: the array initialization char word[8] = "Linux" copied the literal's six bytes and zeroed the two spare ones (§2.1's array-initializer rule), which is why nothing past the terminator is garbage.

One property of the convention deserves its own warning, because it surprises everyone: strlen has to look. No length is stored anywhere, so strlen walks the string, byte by byte, until it finds the zero — a string of a million characters is a million-step strlen, and a string without a terminator is a strlen that walks off the end of the world (§4.4's fence, missing). The practical rules that follow: never call strlen in a loop condition when the length doesn't change (size_t n = strlen(s); once, then use n); and every place you write a string must guarantee room for the terminator — which is §6.4's whole subject.

6.2 Literals, arrays, and pointers to them

A string literal — the quoted form — is an anonymous, read-only array of characters in static storage: "Linux" is a char[6] somewhere in the program's data, '\0' included, placed there when the program was built. Two declarations that look alike use it in profoundly different ways:

char  copy[] = "Linux";   /* the literal's bytes COPIED into my array  */
const char *ref  = "Linux"; /* a pointer aimed AT the static original    */

The first declares an array — sized to fit the literal exactly (six bytes; Exercise 1 predicts sizeof copy) — whose contents are a modifiable copy. The second declares a pointer to the static original: cheap, shareable, and read-only in practice — the standard says writing through it is undefined behavior, and on modern systems the literal often lives in a read-only page, where the write is a §4.7 segfault wearing new clothes. Hence the spellings you have seen since Chapter 1: const char *name for "a string I will only read" (§4.6's contract — crew.c's names, cleanup.c's what), char buf[256] for "characters I intend to write". GCC's opt-in -Wwrite-strings flags the ambiguous middle — char *p = "literal";, writable-looking pointer, read-only reality — and modern code bases enable it; this book simply spells intent: const char * to read, an array to write.

6.3 The three-way truth: strcmp

Comparing strings is not comparing pointers. if (s == t) compares two addresses — true only if both names point at the same bytes, useless for "are these the same text?" — and comparing characters one at a time by hand is what the library is for. strcmp(s, t) walks both strings together and answers a three-way question, exactly the shape §5.6's comparator contract demanded:

A program that prints the truth in all three directions, then puts it to work sorting words with Chapter 5's qsort:

/* cmp.c -- strcmp's three-way truth, and qsort over an array of strings */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>

static int by_word(const void *a, const void *b)
{
    char *const *pa = a;          /* the element type is (char *) */
    char *const *pb = b;

    return strcmp(*pa, *pb);      /* negative / zero / positive */
}

int main(void)
{
    char *words[] = { "delta", "alpha", "bravo", "charlie" };
    int count = sizeof words / sizeof words[0];

    printf("strcmp(\"alpha\", \"delta\") is %d\n", strcmp("alpha", "delta"));
    printf("strcmp(\"delta\", \"alpha\") is %d\n", strcmp("delta", "alpha"));
    printf("strcmp(\"same\", \"same\") is %d\n", strcmp("same", "same"));

    qsort(words, count, sizeof words[0], by_word);
    for (int i = 0; i < count; i++)
        printf("%s\n", words[i]);
    return 0;
}
$ gcc -o cmp cmp.c
$ ./cmp
strcmp("alpha", "delta") is -1
strcmp("delta", "alpha") is 1
strcmp("same", "same") is 0
alpha
bravo
charlie
delta

Two notes on the output, both about promises. The numbers -1, 1, 0 are this library's answers; glibc typically returns the byte difference ('a' - 'd' is −3), and both are correct — only the sign is promised, so never test == 1 or == -1, only < 0, == 0, > 0. And the comparator is the chapter's centerpiece connection: by_word is §5.6's by_height with strings in place of unsigned heights — but look at what it does not contain. No two ifs, no subtraction, none of §5.6's arithmetic traps: strcmp already answers the three-way question in the only safe spelling, so the comparator is one line. (The char *const *pa = a conversion is §5.6's first line again, one pointer deeper: the array's elements are pointers, so the comparator's parameter points at a pointer.)

The searching cousins complete the family, both returning §4.2's most useful idiom — a pointer into your own string, or NULL: strchr(s, c) finds the first c in s; strstr(s, t) finds the first place t appears inside s. NULL means "not found", checked with if (p) (§4.7), and a non-NULL result is an index you can compute (p - s, §4.4's pointer subtraction) or a suffix you can print (%s from p — the convention means "the rest of it" is just "a string starting here").

6.4 Copies, and their bounds

Everything so far only read strings. Writing them is where C earned its reputation, and the whole subject is one sentence: a copy needs room, and only you know how much there is. The library's cast, worst to best:

strcpy(dst, src) copies all of src — terminator included — into dst, with no idea how big dst is, because it cannot: §4.3's rule says an array passed to a function is a pointer, and the size is not in the pointer. If src is longer than the destination's room, the copy runs past the end — a buffer overrun, §2.3's guard-refusing class, the original sin of C security. strcpy is correct and fast when the destination is provably big enough (a fixed-size literal into a large array); it is a loaded weapon anywhere else.

strncpy(dst, src, n) was meant to be the bounded fix, and its crack is famous: it copies at most n bytes, stopping early with the terminator if src fits — but if src is longer, it copies exactly n characters and no terminator at all. The destination then is not a string; strlen, %s, everything in §6.1 walks off the world. The correct two-line idiom patches it by hand, and every line is load-bearing:

strncpy(dst, src, sizeof dst - 1);   /* cut to fit... */
dst[sizeof dst - 1] = '\0';           /* ...and terminated BY HAND */

snprintf(dst, sizeof dst, "%s", src) is the modern discipline: bounded by the destination's real size and always terminated, because that is snprintf's contract — it writes at most size − 1 characters plus the terminator, truncating if it must. It is also the place where the compiler gets to show off, and this chapter's one expected warning is the evidence — captured, verbatim:

/* copy.c -- bounded copies: strncpy's crack, snprintf's discipline */
#include <stdio.h>
#include <string.h>

int main(void)
{
    char dst[8];
    char src[] = "a string much too long for dst";
    char a[16] = "hello";
    char b[16];

    strncpy(dst, src, sizeof dst - 1);   /* cut to fit... */
    dst[sizeof dst - 1] = '\0';           /* ...and terminated BY HAND */
    printf("strncpy + hand-termination: \"%s\"\n", dst);

    snprintf(dst, sizeof dst, "%s", src); /* bounded AND always terminated */
    printf("snprintf:                    \"%s\"\n", dst);

    memcpy(b, a, sizeof a);               /* fixed-size, byte-wise */
    printf("memcpy of the whole buffer:  \"%s\"\n", b);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -g -o copy copy.c
copy.c: In function 'main':
copy.c:16:32: warning: '%s' directive output may be truncated writing up to 30 bytes into a region of size 8 [-Wformat-truncation=]
   16 |     snprintf(dst, sizeof dst, "%s", src); /* bounded AND always terminated */
      |                                ^~   ~~~~
copy.c:16:5: note: 'snprintf' output between 1 and 31 bytes into a destination of size 8
   16 |     snprintf(dst, sizeof dst, "%s", src); /* bounded AND always terminated */
      |      ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

Read that warning with respect: the compiler traced the literal source length against the destination size, saw the truncation coming, and said so — before the program ever ran — naming both sizes. (Note also what it did not flag: the strncpy on the line above, whose identical truncation this build accepted in silence. The formatted path is instrumented; the raw one is on foot.) The program, warning acknowledged, is deliberately truncating — and truncation is sometimes the policy (a progress bar, a filename column) — so it runs and tells the truth about what it kept:

$ ./copy
strncpy + hand-termination: "a strin"
snprintf:                    "a strin"
memcpy of the whole buffer:  "hello"

Both bounded copies kept seven characters and a terminator — identical output, very different guarantees: snprintf's termination is always true; strncpy's is true because we wrote the second line. The third copy is a different tool entirely: memcpy(dst, src, n) copies exactly n bytes — §4.8's typeless "copy from here to there", no terminators, no characters, no strings: it is how structures, buffers, and fixed-size records move. Copying the whole 16-byte buffer moved "hello" and its terminator and its spare zeros, which is why the result prints. (For overlapping ranges — memcpy(&s[0], &s[2], n) — the answer is memmove, the same copy with license to overlap; memcpy with overlap is undefined.)

Three last pieces of the copy family. snprintf's return value is a gift: the number of characters the output would have needed — so if ((size_t) ret >= sizeof dst) detects truncation precisely, and ret is the length of the uncut version without ever building it. strdup(s) (POSIX, not ISO) is "copy this string and give me owned storage" — Chapter 5's malloc doing string duty: char *d = strdup(s); if (d == NULL) ... and, eventually, free(d) — the ownership rules of §5.5 unchanged. And the allocation answer to "how big is the buffer, then?" is getline ("man 3 getline"): it grows a malloc'd line to fit whatever arrives — the end of fixed buffers — and arrives with Chapter 7's stdio.

6.5 Characters: <ctype.h>

Between single characters and whole strings sits a small, immaculate library: &lt;ctype.h>, the classifiers and case-converters — isalpha, isdigit, isspace, ispunct, iscntrl, toupper, tolower, and friends ("man 3ctype" indexes them all). Its discipline is visible in every prototype: the functions take int, accept exactly the values unsigned char can hold or EOF, and answer int 0 or nonzero (a truth value in §3.1's sense — the comparison spellings if (isalpha(c)) and if (isalpha(c) != 0) are the same test). That "or EOF" clause is why the chapter's filter looks the way it does:

/* chars.c -- classify standard input, one kind at a time */
#include <stdio.h>
#include &lt;ctype.h>

int main(void)
{
    long counts[4] = { 0 };       /* letters, digits, space, other */
    int c;

    while ((c = getchar()) != EOF) {
        if (isalpha(c))
            counts[0]++;
        else if (isdigit(c))
            counts[1]++;
        else if (isspace(c))
            counts[2]++;
        else
            counts[3]++;
    }
    printf("%8ld letters\n", counts[0]);
    printf("%8ld digits\n", counts[1]);
    printf("%8ld space\n", counts[2]);
    printf("%8ld other\n", counts[3]);
    return 0;
}
$ gcc -o chars chars.c
$ printf 'Hello, world 42!\n' | ./chars
      10 letters
       2 digits
       3 space
       2 other

This is §4.9's promised loop — while ((c = getchar()) != EOF), the assign-then-test idiom — doing real work at last, and its shape is not style: getchar returns int precisely so it can return a character or EOF (−1, "end of file") without confusion, and the loop's != EOF is what guarantees isalpha a value it is licensed to take. The classification itself is an else-if cascade (§3.2) over long counters zero-initialized by §2.1's array rule, and the tally reads cleanly: ten letters, two digits, three whitespace (' ' twice, '\n' once — isspace counts them all), two of "other" (the comma and the exclamation point).

The rule that trips everyone: the same functions applied to bytes from an array need a cast. A char can be negative (§1.5's implementation-defined signedness); passed to isalpha as an int, a negative byte is out of the function's licensed domain — undefined behavior, and on real libraries it can index the tables off their ends. The idiom, always:

if (isdigit((unsigned char) line[i]))     /* cast INTO the licensed domain */
    ...

Case conversion is the same family with a result you keep: toupper(c) returns the upper-case equivalent (or c unchanged, if none), so the conversion idiom is store-back with the same cast — line[i] = toupper((unsigned char) line[i]);. Check the man page for the whole set; it is small, standard, and used daily.

6.6 Numbers from text

Chapter 3 needed command-line arguments as integers, took atoi on trust, and flagged the debt: garbage in, zero out, no complaint. The honest tool is strtol ("man 3 strtol", <stdlib.h>), and it is Chapter 4's whole design in one function:

long strtol(const char *s, char **endp, int base);

It converts the leading number in s — skipping white space, honoring a sign, reading in the given base (10 for plain text; 16 for hex; 0 for auto-detect, where a 0x prefix means hexadecimal and a leading 0 means octal) — and, through the second parameter, an output parameter (§4.2): a pointer to a char * that receives the address where parsing stopped. That end pointer is how strtol reports everything atoi swallowed. A program that shows both faces:

/* num.c -- converting text to numbers: atoi's debt, paid */
#include <stdio.h>
#include <stdlib.h>

int main(void)
{
    const char *good = "42 bottles";
    const char *bad = "x";
    char *end;
    long v;

    v = strtol(good, &end, 10);
    printf("strtol(\"%s\") = %ld, stopped at \"%s\"\n", good, v, end);

    v = strtol(bad, &end, 10);
    printf("strtol(\"%s\") = %ld, stopped at \"%s\"\n", bad, v, end);
    return 0;
}
$ gcc -o num num.c
$ ./num
strtol("42 bottles") = 42, stopped at " bottles"
strtol("x") = 0, stopped at "x"

Read the second line as the contract. "42 bottles" converted 42 and stopped where the digits ended — the end pointer is even useful prose ("what came after the number?"). "x" converted nothing: the result is 0 and the end pointer equals the start, end == bad, and that equality is the error signal — a clean, checkable "no digits found" that atoi never offered. The fully guarded parser (Exercise 6 builds it) is three checks: if (end == s) — no digits; errno = 0 before the call and if (errno == ERANGE) after — the number was too big and was clamped to LONG_MAX or LONG_MIN (<errno.h>; errno is Chapter 7's subject, met here as a checked global); and, if the argument must be entirely a number, if (*end != '\0') — trailing junk. For floating point, strtod is the same design; atoi, atof, and friends remain what they always were — convenient, silent, and never again used in this book's tools.

6.7 Chapter 1's debt, paid

The oldest promise in the book closes here. §1.9's nl.c numbered each buffer's worth of input, so a line longer than the buffer arrived in numbered fragments, and the honest repair "waits for the string tools of Chapter 6". The tools are on the table; here is the repair — nl2.c, with a deliberately tiny 16-byte buffer so the long-line case is exercised on every run:

/* nl2.c -- Chapter 1's nl, repaired: a long line gets one number */
#include <stdio.h>
#include <stdbool.h>
#include <string.h>

int main(void)
{
    char line[16];                 /* deliberately tiny, to exercise the fix */
    int n = 0;
    bool bol = true;               /* at the beginning of a line? */

    while (fgets(line, sizeof line, stdin) != NULL) {
        size_t len = strlen(line);

        if (bol) {
            n++;
            printf("%6d\t", n);
            bol = false;
        }
        printf("%s", line);
        if (len > 0 && line[len - 1] == '\n')
            bol = true;            /* the newline ended the line */
    }
    return 0;
}
$ gcc -o nl2 nl2.c
$ printf 'short one\nthis line is deliberately much longer than the tiny buffer\nand a last one\n' | ./nl2
     1  short one
     2  this line is deliberately much longer than the tiny buffer
     3  and a last one

The middle line is 57 characters — four fgets fills of the 15-characters-and-a-terminator buffer — and it received one number, exactly as cat -n would have given it. The machinery is the chapter in miniature. strlen(line) measures what fgets actually brought (§6.1, counting to the terminator fgets guaranteed). line[len - 1] == '\n' is the string tools' most-used idiom — the last character test — asking "did this chunk end the line?" (§2.4's floatcmp used the same shape on numbers; §6.8's exercises meet the twin idiom, stripping the newline). bool bol — Chapter 2's type, carrying the one bit of state the fix needs: am I at the beginning of a line? — set true by every chunk that ended in a newline, consumed by the first chunk of the next. And the flag is exactly §3.4's "escaping a loop's logic with a flag" pattern, moved into the data: bol says whether the input is between lines, the way end said whether a walk was between elements.

Compare honestly with Chapter 1: the old program was 8 lines and wrong about long lines; the repair is 25 and right. That ratio is not waste — it is the price of the guarantee, and the tools that paid it (strlen, the last-character test, a flag) are the same three that pay it in every line-oriented program you will ever write. The wc you pipe into, the log scanner you will maintain, the getline-based tools of Chapter 7: all of them are this loop, wearing different bodies.

6.8 Summary

A string is characters ended by '\0' — the convention, the storage, and the whole data structure; strlen counts to the terminator (and must walk to find it), sizeof measures the buffer, and every write must reserve room for the byte that ends it. Literals are static read-only arrays: char copy[] owns a modifiable duplicate, const char * borrows the original, and -Wwrite-strings polices the gap. Comparison is strcmp's three-way truth — sign only, never magnitude — which is why it slots into qsort's comparator as one line with all of §5.6's arithmetic traps designed out; strchr/strstr return a pointer into your string or NULL. Copies live or die by bounds: strcpy is provable-room-only, strncpy truncates without terminating (the hand-written second line is part of the idiom), snprintf is the bounded-and-always-terminated discipline — its return value measures the uncut length, -Wformat-truncation catches the collisions before running, memcpy/memmove move raw bytes (overlap needs memmove), and strdup buys owned storage. &lt;ctype.h> classifies and converts characters through int under a strict license — unsigned char values or EOF — which the getchar loop guarantees by construction and array indexing must guarantee with the cast. strtol converts text to numbers through an end-pointer output parameter, with end == start as the no-digits signal and errno's ERANGE as the overflow signal; atoi is retired. And the repair of Chapter 1's nl.c — strlen, the last-character test, a bool flag — is the loop shape under every line-oriented tool in the UNIX world.

Next: Chapter 7 — file I/O and system calls: the stdio machinery underneath everything this book has piped (FILE * streams, buffering made controllable, fread/fwrite, getline ending fixed buffers, fseek, and ferror paying §1.9's other deferred debt) — and, underneath that, the kernel's own interface: file descriptors, open, read, write, close, stat, and perror.

6.9 Exercises

Recommended flags on; predict before running; several of these want paper first.

  1. The exact-fit array. Add to len.c's world: char exact[] = "Linux"; — predict sizeof exact and strlen(exact) before running, then verify both plus exact[5]. Then print word[6] and word[7] and explain their values with §2.1's initializer rule.

  2. Palindromes. Using rev.c's two-pointer walk as the model, write bool is_palindrome(const char *s) — first character in step with the last, meeting in the middle. Test "level", "linux", and "", and predict the empty case before running (the p < q fence decides it).

  3. Two more comparators. (a) Flip cmp.c to sort descending — predict the first word before running. (b) Sort by length: a comparator using strlen(*pa) and strlen(*pb) in §5.6's two-if shape. Then explain, with §2.7's wrap arithmetic in hand, why return (int) (strlen(*pa) - strlen(*pb)); is a trap that a long-string test would eventually spring.

  4. The bounded copier, as a function. Write size_t copy_bounded(char *dst, size_t dsize, const char *src) — snprintf-based, always terminated — returning how many characters the source would have needed beyond what fit. Verify with dsize 1 (the destination becomes an empty string: the terminator still needs its byte), 8, and the source's exact length.

  5. Wider census. Extend chars.c with ispunct and iscntrl buckets, and a demo line containing a tab and a bell (\t, \a — say which bucket each lands in, and why \n chose isspace). Then change the source of characters from getchar to an fgets buffer classified by index — with the (unsigned char) cast — and explain in one sentence why the cast is needed there and not in chars.c.

  6. The guarded parser. Write bool parse_long(const char *s, long *out) — strtol with errno cleared first, then end == s rejected, ERANGE rejected, and trailing junk (*end != '\0') rejected; the result lands through the output parameter (§4.2). Predict, then verify: "42", "42 bottles" (rejected — trailing junk), "x", "9999999999999999999999" (rejected — ERANGE).

  7. The edge of the repair. Run printf 'no trailing newline' | ./nl2 and predict the number before you see it; explain the flag's state when the loop exits. Then add the fprintf(stderr, ...) total report from Chapter 1's Exercise 4 — count lines, not chunks — and finally predict which demo lines split with the buffer at 8, then verify.


Chapter 7 · File I/O and System Calls

Until now, this book's programs have lived on the three streams a process is born with: stdin, stdout, stderr, plus whatever the shell piped into them. This chapter opens named files — the durable kind, the kind with permissions and owners and sizes — and it opens them twice, because Linux gives you two doors into the same files, one above the other. The upper door is the standard library's stdio: FILE * streams, buffering, fread/fwrite — portable ISO C, the layer of every program so far. The lower door is the kernel's own interface, the system calls: open, read, write, close, stat — §1.12's section 2 man pages, arriving at last, with their file descriptors and their errno contract.

The Linux idea underneath both doors is worth naming, because it is one of the great designs in computing: everything is a file. A disk file is a file; so is a pipe, a terminal, a disk device, a network socket, and half of /proc — all opened, read, written, and closed through the same handful of calls, all reached by the same small integers. The tools this chapter builds work unchanged on all of them.

The debts paid here are the book's oldest: ferror, deferred since §1.9 ("the right check is ferror(stdin)"); the buffering mystery of §1.10, finally made controllable; EOF, promised "explained in Chapter 7" in §4.9; getline, deferred from Chapter 6; scanf's &, deferred from §4.2; and ssize_t, promised from §2.2. Same conventions: programs in examples/ch07/, compiled under the recommended flags, every output block captured from a real run — including, this chapter, two honest captures of this machine's platform quirks, which turn out to be lessons rather than embarrassments.

7.1 Two layers, one file

The layers, top to bottom:

 your program
      |
      |  FILE * streams  -- stdio: fopen, fread, fwrite, fclose, getline
      v
 [ a buffer ]          -- bytes gather here; one kernel call per buffer,
      |                   not one per byte (that is §1.10's mystery)
      |  file descriptors -- the kernel's door: open, read, write, close
      v
    the kernel --> the filesystem --> the disk
                  (and pipes, terminals, devices, /proc -- same calls)

A FILE * is an opaque handle: a structure the library owns (buffer, current position, error flags), whose internals you never touch. A file descriptor is much less — a small int, the index of an open file in your process's private table — and much more: the kernel's universal currency, the numbers 0, 1, 2 you met in §1.10 (standard input, output, error), and the type of every open file, socket, and pipe in the system. The choice of layer is engineering, not dogma: stdio when text, portability, and line-shaped work dominate; descriptors when you need what only the kernel offers — permissions on create, non-blocking reads, file metadata, sockets — or when the buffering would be in the way. Real programs mix them freely, and this chapter's centerpiece is one program built twice, once through each door, so the differences are seen rather than asserted.

7.2 Streams: fopen, fclose, and the mode string

fopen(path, mode) opens a stream on a named file, returning FILE * — or NULL, which is checked before anything else, as ever. The mode string is a small language:

Mode Meaning Notes
"r" read must exist; NULL if not
"w" write created or truncated to empty
"a" append created; writes go to the end, always
+ also the other direction "r+" reads and writes, no truncate
b binary no text translation — see below

Two facts about the mode string are worth a paragraph each. First, "b" exists for platforms that translate — where the "text" modes map the operating system's line convention onto C's \n. On Linux, nothing is ever translated, "b" is a quiet no-op, and text is binary — one of the system's cleanest decisions. (That is the Linux answer; §7.4's second capture shows a platform that does translate, at the descriptor layer, with the byte counts to prove it.) Second, fclose returns a value — 0, or EOF (the constant, −1) if anything went wrong — and checking it is not ceremony: the stream's buffer is flushed at close, so a full disk or a dropped network share surfaces here, after every fwrite already said "fine". The chapter's copy program checks every close, and the comment on the second one names why.

7.3 Blocks in, blocks out: fread/fwrite, ferror/feof

The workhorse pair moves blocks, not lines, and their contract reads like §4.5's array rule turned around: fread(ptr, size, nmemb, stream) reads up to nmemb items of size bytes each into ptr, returning how many complete items landed — short at end of file, zero at end or error, and the two "zero" cases are separated only by the flags the loop must not forget:

if (ferror(stream))   /* it stopped because something went WRONG   */
if (feof(stream))     /* it stopped because the file ended: normal */

§1.9's deferred debt, paid in full: after a read loop stops, ferror and feof say why — and only then; both are meaningless before a read has actually failed. Both flags live in the FILE object itself, which is why they are functions on the handle and not return values. Now the chapter's first copy tool, written through the stdio door:

/* fcopy.c -- copy a file, one buffer at a time (a tiny cp) */
#include <stdio.h>

int main(int argc, char *argv[])
{
    FILE *in, *out;
    char buf[4096];
    size_t n;

    if (argc != 3) {
        fprintf(stderr, "usage: fcopy src dst\n");
        return 2;
    }
    in = fopen(argv[1], "rb");
    if (in == NULL) {
        fprintf(stderr, "fcopy: cannot open %s\n", argv[1]);
        return 1;
    }
    out = fopen(argv[2], "wb");
    if (out == NULL) {
        fprintf(stderr, "fcopy: cannot open %s\n", argv[2]);
        fclose(in);
        return 1;
    }
    while ((n = fread(buf, 1, sizeof buf, in)) > 0) {
        if (fwrite(buf, 1, n, out) != n) {
            fprintf(stderr, "fcopy: write error\n");
            fclose(in);
            fclose(out);
            return 1;
        }
    }
    if (ferror(in)) {                  /* why did the read loop stop? */
        fprintf(stderr, "fcopy: read error\n");
        fclose(in);
        fclose(out);
        return 1;
    }
    fclose(in);
    if (fclose(out) != 0) {            /* buffered writes can fail HERE */
        fprintf(stderr, "fcopy: close error\n");
        return 1;
    }
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -Werror -g -o fcopy fcopy.c
$ ./fcopy fcopy.c fcopy.copy
$ cmp -s fcopy.c fcopy.copy && echo "fcopy: byte-identical"
fcopy: byte-identical

Read it as a checklist, because it is the shape of every correct file program you will ever write. Usage check first (§3.9). Every open is checked; failure cleans up what it opened — fclose(in) on the path where out failed is §3.5's cleanup discipline. The loop's unit is 1 byte × sizeof buf members — so n is bytes, and the contract "how many landed" is used as the count everywhere after. The write checks its short count: fwrite returns items written, and a partial write is an error here. When the loop stops, the flags are asked why — ferror first, because "error" and "end" need different answers, and §1.9's promise is finally kept on a real file. Both closes run; only the one whose flush could still fail is checked, with the comment explaining the asymmetry. And the transcript's cmp — byte-compare, silent on success — is the honest way to prove a copy tool: not "it printed no error", but "the bytes match".

7.4 The kernel's door: open, read, write, close

Below stdio sits the interface the kernel actually exports, and its vocabulary is Chapter 2's come home to roost. open(path, flags, mode) returns a file descriptor — a small non-negative int, or −1 on failure. The flags are bitmasks, OR-ed together exactly as §2.5 taught: O_RDONLY, O_WRONLY | O_CREAT | O_TRUNC, and on. The mode — consulted only when O_CREAT makes a new file — is a permission mask, spelled in octal, with the very bits §2.5's perm.c decoded: 0644 means rw-r--r-- (and the kernel's umask clears bits from it, which is Exercise 5's other half, back from Chapter 2).

read(fd, buf, count) and write(fd, buf, count) move count bytes through the descriptor and return ssize_t — §2.2's promised signed size — whose values are the verdicts this section exists to teach:

The first verdict is the one stdio was hiding: short reads and writes are normal — a terminal delivers per keypress, a pipe per writer's mood, a socket per packet — so a read loop treats its result as "what arrived", never as "what I asked for". (write can also report partial success; a production loop retries the remainder — Exercise 2 measures why buffering exists by defeating it.) And on failure, the kernel's error report is errno — a per-thread global from <errno.h> holding a small code — printed friendly by perror(prefix) ("prefix: message"), Chapter 6's checked global grown into its working clothes.

The same copy, through the kernel's door:

/* syscopy.c -- the same copy, straight through the kernel */
#include <stdio.h>
#include <fcntl.h>              /* open, O_RDONLY, O_CREAT, O_TRUNC */
#include <unistd.h>             /* read, write, close */

int main(int argc, char *argv[])
{
    int in, out;
    char buf[4096];
    ssize_t n;

    if (argc != 3) {
        fprintf(stderr, "usage: syscopy src dst\n");
        return 2;
    }
    in = open(argv[1], O_RDONLY);
    if (in < 0) {
        perror(argv[1]);
        return 1;
    }
    out = open(argv[2], O_WRONLY | O_CREAT | O_TRUNC, 0644);
    if (out < 0) {
        perror(argv[2]);
        close(in);
        return 1;
    }
    while ((n = read(in, buf, sizeof buf)) > 0) {
        if (write(out, buf, (size_t) n) != n) {
            perror("write");
            close(in);
            close(out);
            return 1;
        }
    }
    if (n < 0) {
        perror("read");
        close(in);
        close(out);
        return 1;
    }
    close(in);
    if (close(out) < 0) {
        perror("close");
        return 1;
    }
    return 0;
}

Structurally the twin of fcopy.c — same checklist, perror in place of hand-made messages, the verdict loop in place of the member-count loop — and it compiles clean and copies correctly. On this capture machine, though, the proof is where the second honest platform capture of the chapter lives:

$ gcc -std=c17 -Wall -Wextra -Werror -g -o syscopy syscopy.c
$ ./syscopy fcopy.c syscopy.copy
$ cmp fcopy.c syscopy.copy
fcopy.c syscopy.copy differ: byte 63, line 1
$ echo $?
1
$ wc -c fcopy.c syscopy.copy
1170 fcopy.c
1215 syscopy.copy
2385 total

The copy tool — the correct, checklist-perfect one — produced a file 45 bytes larger than its source, differing first at byte 63. The numbers convict the culprit: fcopy.c contains 45 newline characters, and 45 = 1215 − 1170, one extra byte per newline. This machine's descriptor layer, in its default text mode, translates every \n written into \r\n on disk; Linux never does — which is why §7.2's "b" is a no-op there and why the same program, unmodified, is byte-exact on Linux. And the frame completes with the companion capture:

$ head -c 5000 /dev/zero | tr '\0' x > nonewline.bin
$ ./syscopy nonewline.bin nonewline.copy
$ cmp -s nonewline.bin nonewline.copy && echo "syscopy on a newline-free file: byte-identical"
syscopy on a newline-free file: byte-identical

No newlines, no translation, no difference: the loop, the verdicts, and the cleanup were right all along — the platform's text translation was the whole story. Two lessons ride along for free: byte-exact tools must speak binary on platforms that translate ("b" above, O_BINARY at this layer, both absent on Linux because there is nothing to translate); and a byte-compare with cmp is the only proof of a copy tool that means anything. -Wall cannot catch a correct program translating honestly on the wrong platform; cmp can, and did.

7.5 stat: a file's facts

Every file carries its metadata — the kernel's own record of it: type, permissions, owner, size, timestamps. The call that fetches it is stat ("man 2 stat"), and it is Chapter 5 turned into a system call: a struct stat — the kernel's record type — filled through §4.2's output-parameter pattern, at syscall scale:

/* stat.c -- a file's facts, straight from the kernel */
#include <stdio.h>
#include <sys/stat.h>          /* stat, struct stat, S_ISDIR */

int main(int argc, char *argv[])
{
    struct stat st;
    const char *kind;

    if (argc != 2) {
        fprintf(stderr, "usage: stat file\n");
        return 2;
    }
    if (stat(argv[1], &st) != 0) {
        perror(argv[1]);
        return 1;
    }
    if (S_ISDIR(st.st_mode))
        kind = "directory";
    else if (S_ISREG(st.st_mode))
        kind = "regular file";
    else
        kind = "other";
    printf("file:        %s\n", argv[1]);
    printf("kind:        %s (mode %o)\n", kind, st.st_mode);
    printf("permissions: %04o\n", st.st_mode & 0777);
    printf("size:        %lld bytes\n", (long long) st.st_size);
    printf("owner uid:   %d\n", st.st_uid);
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -Werror -g -o stat stat.c
$ ./stat stat.c
file:        stat.c
kind:        regular file (mode 100666)
permissions: 0666
size:        832 bytes
owner uid:   0
$ ./stat .
file:        .
kind:        directory (mode 40777)
permissions: 0777
size:        4096 bytes
owner uid:   0

Read st.st_mode in octal, top to bottom, with Chapter 2 in hand: 100666 is a bit field, not one number — the high digits (10...) are the file type bits (S_IFREG, "regular file"), the low three digits (666) are the permission bits §2.5 decoded letter by letter. The directory answers 40777: type S_IFDIR in the high bits, 777 below. That is why the program prints the mode twice — whole (%o), then masked (st.st_mode & 0777, §2.5's bit-test arithmetic on the real thing): kind and permissions live in the same word, and the mask is how you take just yours. (The captured 666 and 777 are this machine's answers — files created here carry no execute bits and directories do not distinguish permissions; a Linux system answers 644 for a data file and 755 for a program, which is the shape the exercises explore.) st_size is printed through a (long long) cast because its type varies across systems — the §2.2 rule-of-thumb table's "avoid plain long" row, enforced by a printf. S_ISDIR and S_ISREG are the classifiers for the type bits (§6.5's isalpha, for file types); and lstat, the sibling call, reports the link itself where stat follows a symbolic link to its target — Exercise 3 makes the difference visible.

The stat call is also the design bridge to the rest of the system interface: kernel functions that fill your structures through output parameters — stat, later opendir/readdir, gettimeofday — are one repeated shape, and once you have read it here, you can call your way through half of section 2.

7.6 Buffering, controlled

§1.10 deferred the machinery of buffering to "a real tool in Chapter 7"; here is the control panel, and a program to make it visible:

/* bufdemo.c -- buffering made visible: stdout vs stderr, and fflush */
#include <stdio.h>

int main(void)
{
    printf("stdout line 1 (buffered)\n");
    printf("stdout line 2 (buffered)\n");
    fprintf(stderr, "stderr line 1 (unbuffered)\n");
    fflush(stdout);
    printf("stdout line 3, after fflush\n");
    return 0;
}

Captured with output redirected (so stdout is block-buffered — §1.10's rule):

$ ./bufdemo > bufdemo.out 2>&1
$ cat bufdemo.out
stderr line 1 (unbuffered)
stdout line 1 (buffered)
stdout line 2 (buffered)
stdout line 3, after fflush

The order in the file is not the order in the program. The stderr line — unbuffered by design — was written the instant it was printed. Lines 1 and 2 sat in the buffer until fflush(stdout) pushed them, after the stderr line had already gone. Line 3, printed after the flush, waited in the buffer until the program exited. Run interactively, the same program prints in program order — a terminal makes stdout line-buffered, so each \n flushes — and that difference is not a curiosity: it is why diagnostics belong on stderr (§1.10's rule, now demonstrated), why §4.7's crashing program lost its buffered last words, and why a program that exits abnormally — or dies by signal — should fflush(stdout) first if the unfinished product matters. The rest of the panel, one line each: setvbuf(stream, buf, mode, size) sets the policy before first use — _IONBF (none), _IOLBF (per line), _IOFBF (full); and the sizes above all are the stdio layer's — at the descriptor layer there is no buffering at all, which is §7.4's short reads and Exercise 2's slow loops.

7.7 Lines without limits: getline, earned

Chapter 6 ended with the fixed-buffer problem — fgets with buf[256] splits what does not fit — and deferred the allocation answer to getline. Honest chapter requires honest method: this capture machine's libc does not provide getline (the first attempt to use it was refused by the compiler, verbatim below), so the program below builds the mechanism by hand — and then getline arrives as exactly what it is: that mechanism, finished and debugged by a library.

First, the mechanism — a growable line reader, using Chapter 5's vector discipline (realloc through a temporary, doubling, ownership documented) with Chapter 6's string tools:

/* lines.c -- count lines, reading any length without a fixed buffer */
#include <stdio.h>
#include <stdlib.h>
#include <string.h>

/* Read one line of any length into *bufp (grown by realloc as needed).
   Returns 1 with a line, 0 at end of input, or -1 on error. */
static int read_line(FILE *f, char **bufp, size_t *capp)
{
    size_t len = 0;

    for (;;) {
        size_t room = *capp - len;

        if (room < 2) {                   /* keep room for a char + '\0' */
            size_t ncap = *capp ? *capp * 2 : 128;
            char *nb = realloc(*bufp, ncap);

            if (nb == NULL)
                return -1;               /* old buffer still valid (§5.7) */
            *bufp = nb;
            *capp = ncap;
            room = *capp - len;
        }
        if (fgets(*bufp + len, (int) room, f) == NULL) {
            if (len > 0)
                return 1;                /* final line, no newline */
            return (feof(f) && !ferror(f)) ? 0 : -1;
        }
        len += strlen(*bufp + len);
        if (len > 0 && (*bufp)[len - 1] == '\n')
            return 1;                     /* a complete line */
        if (feof(f))
            return 1;                     /* final line, no newline */
    }
}

int main(int argc, char *argv[])
{
    FILE *f = stdin;
    char *buf = NULL;
    size_t cap = 0;
    long count = 0;
    int rc;

    if (argc == 2) {
        f = fopen(argv[1], "r");
        if (f == NULL) {
            perror(argv[1]);
            return 1;
        }
    }
    while ((rc = read_line(f, &buf, &cap)) == 1)
        count++;
    free(buf);                            /* free(NULL) is safe (§5.5) */
    if (f != stdin)
        fclose(f);
    if (rc != 0) {
        fprintf(stderr, "lines: read error\n");
        return 1;
    }
    printf("%ld lines\n", count);
    return 0;
}

Walk read_line once — it is the densest function in the book, and every line is a chapter you have already read. The caller owns nothing at first: buf == NULL, cap == 0, so the first pass allocates (128) — realloc(NULL, n) is malloc's spelling (§5.7). room is always kept at 2 or more — space for a character and §6.1's terminator, because fgets needs to add one. The read aims at buf + len, the tail: fgets fills the unused end of the line, strlen of the tail adds only the new characters, and the loop keeps reading until the last character is '\n' — Chapter 6's last-character test, deciding "this line is finished". A line longer than the buffer is not an error but a doubling — §5.7's growth policy, through §5.7's temporary, with failure leaving the old buffer valid. The stop conditions return the trichotomy the header documents: 1 (a line, terminated or not), 0 (end of input — and only if feof and not ferror, §7.3's distinction), -1 (error or out of memory). main's half is ownership: the buffer is allocated by the callee, freed once by the caller — free(NULL) safe if no line was ever read (§5.5) — and the exit status tells the truth (§3.9).

Now the runs, including the two that exist to prove "any length":

$ gcc -std=c17 -Wall -Wextra -Werror -g -o lines lines.c
$ printf 'a\nb\n' | ./lines
2 lines
$ ./lines lines.c
64 lines
$ head -c 5000 /dev/zero | tr '\0' x > longline.txt; printf '\n' >> longline.txt
$ ./lines longline.txt
1 lines
$ printf 'one\ntwo\nthree' > nonl.txt
$ ./lines nonl.txt
3 lines

A 5000-character line — thirty-two times the old 256-byte fear — is one line, one count, one fgets after another until the newline arrives. And the file with no trailing newline counts its last line, exactly as Chapter 6's §6.7 promised a complete line tool should.

Now getline. The honest capture first — this is the compile that motivated the mechanism above, on this machine, verbatim:

$ gcc -std=c17 -Wall -Wextra -Werror -g -o lines lines.c
lines.c: In function 'main':
lines.c:19:12: error: implicit declaration of function 'getline' [-Wimplicit_function-declaration]
   19 |     while (getline(&line, &cap, f) != -1)
      |            ^~~~~~~

(The program whose main was then just that while loop; the machine's MinGW libc does not declare getline — POSIX 2008, present in glibc on every Linux distribution.) What the missing function is: the read_line you just walked, with fifteen fewer lines and one more guarantee. Its contract:

ssize_t getline(char **bufp, size_t *capp, FILE *f);

Bring buf == NULL, cap == 0 (or any buffer you own); it reads one line of any length — growing the buffer through realloc exactly as §5.7 taught — and returns the line's length (terminator included), or -1 at end of input or error. The buffer is yours: allocated by the callee, reused across calls, freed once at the end. On a Linux system, the whole mechanism above collapses:

while (getline(&buf, &cap, f) != -1)
    count++;
free(buf);

read_line's trichotomy becomes getline's −1; its growth policy is the library's; its '\n' test never needs writing, because getline keeps the newline and reports the length, letting you decide. The exercise set asks you to port lines.c both ways and keep the outputs byte-identical — which is the real lesson of the section: getline is not new machinery, it is your machinery, maintained by people who have seen more edge cases than any one book can print.

7.8 scanf, at last

§4.2 promised scanf a proper meeting. It is printf run backwards: a format string, conversions, and — where printf took values — pointers to places to put them, which is why every argument is an &-address or a buffer name. Its return value is the contract: the number of conversions that landed, 0 if the input did not match, EOF at end of input — checked, always, before using anything it filled:

int n;
if (scanf("%d", &n) != 1) {            /* count, not value (§3.9's rule) */
    fprintf(stderr, "usage: wanted a number\n");
    return 2;
}

The one standing rule if you use the %s conversion: give it a width — scanf("%9s", word) writes at most 9 characters plus the terminator, the bounded-copy discipline of §6.4 in a specifier; bare %s is unbounded strcpy with the keyboard attached. And the honest comparison, now that you can make it: scanf stops at the first mismatch, leaves the unread junk in the input, and reports that one conversion failed — while fgets + strtol (§6.6) reads a whole line, tells you exactly where parsing stopped, and leaves you in control. That is why the book's tools parse the second way, and why scanf earns its place in small structured input — and a section, not a career.

7.9 Moving around: fseek and friends

A stream over a regular file has a position — the byte index of the next read or write — and three calls move it: fseek(f, offset, whence) jumps (SEEK_SET from the start, SEEK_CUR from here, SEEK_END from the end — negative offsets legal, which is how "the last 100 bytes" is spelled); ftell(f) reports the current position; rewind(f) is fseek(f, 0, SEEK_SET) with the error flag quietly cleared — prefer the explicit fseek. Two cautions complete the picture: seeking is only meaningful on seekable files — on a pipe or a terminal, fseek fails (the kernel's ESPIPE, "illegal seek") — and mixing reads and writes on one stream requires an intervening fseek/fflush, a rule the standard states and the exercises let you collide with gently. At the descriptor layer the same service is lseek ("man 2 lseek"), with the same whence names and the same pipe allergy.

7.10 Summary

Files open through two doors: stdio streams — FILE * handles with a buffer and error flags, fopen's mode string (r/w/a, +, and b for platforms that translate; Linux never does), fclose's return value guarding the final flush — and the kernel's descriptors: open with OR-ed O_ flags and an octal mode (0644, masked by umask), small int handles where −1 is the only failure. Block I/O is fread/fwrite with a (buffer, item-size, count, stream) contract, short counts at end, and the loop-exit question answered by ferror/feof — §1.9's debt paid — with cmp as the only proof of a copy that means anything. The syscall layer's verdicts — ssize_t: bytes moved, 0 at end, −1 with errno — make short reads normal, and its error report is perror over errno, Chapter 6's checked global at work. stat fills the kernel's own record through the output-parameter pattern: st_mode is type bits over §2.5's permission bits (S_ISDIR, & 0777), st_size casts before printing, and lstat sees the link itself. Buffering is a policy: stderr unbuffered by design, stdout line-buffered on terminals and block-buffered elsewhere — the interleaving capture proved it — with fflush and setvbuf as the controls, and a flush before abnormal exits. The fixed-buffer era ended twice: once by mechanism — a hand-built growable line reader of §5.7's realloc discipline, doubling, tail-append fgets, and Chapter 6's last-character test — and once by library: getline, the same contract (−1 at end, buffer yours to free), absent on this capture machine (the compiler said so, in a captured error) and standard on Linux. scanf converts into pointers and returns the count that landed (%9s, never bare %s); fseek/ftell move through seekable streams only. And across everything: everything is a file — the same calls, the same descriptors, for disk, pipes, devices, and /proc.

Next: Chapter 8 — processes and signals: fork, exec, and wait (how a shell builds a pipeline, including the one this book has been drawing all along), the environment block and environ, exit's full story, and signals — SIGSEGV's family, kill, and handlers — where function pointers, promised since §5.6, finally report for duty.

7.11 Exercises

Recommended flags on; cmp remains the only proof of copies; several exercises want a real Linux system for their juiciest half.

  1. Progress and honesty. Extend fcopy.c to count bytes copied and report them to stderr at the end (fcopy: 1170 bytes); verify against wc -c. Then copy a directory as the source — on Linux, fopen may succeed where fread fails — and make the ferror path also call perror("read") (errno survives the flag check): what does the kernel say the problem was?

  2. Defeat the buffer, measure it. Make a 10 MB file (head -c 10000000 /dev/zero > big), and time syscopy with the buffer at 4096, then 1 (time ./syscopy big big2). Explain the difference in §7.6's terms; then, on Linux, count the syscalls with strace -c ./syscopy big big2 and check your explanation against the kernel's arithmetic.

  3. Symlinks, stat vs lstat. Add st_mtime printed with ctime(&st.st_mtime) (<time.h>) and a S_ISLNK branch; then ln -s fcopy.c link.c and predict the two outputs of ./stat link.c before running it once with stat, once with lstat. Verify the link's st_size is the length of the target's name, not the file.

  4. Both orders, then one. Run bufdemo interactively and redirected; write down both orders and explain each with §7.6's rules. Then put setvbuf(stdout, NULL, _IONBF, 0); first in main and predict — before running — that both contexts now print in program order.

  5. getline, ported. On a Linux system, replace read_line with getline (§7.7's snippet) and confirm all four captured counts byte-identical (2, 64, 1, 3). Then answer in one sentence each: why must the buffer be freed by this program, and why is cap passed by pointer?

  6. Atomic append. Write append.c: open with "a", write one line, close; verify with cat. Then write the descriptor edition with open(..., O_WRONLY | O_CREAT | O_APPEND) and read man 2 open's O_APPEND sentence about atomicity — explain, in two sentences, why two processes appending concurrently want that guarantee (and why "w" mode would be a disaster).

  7. A pocket hexdump. Write firstn file n — print the first n bytes, 16 per line, in x-pairs (%02x, §2.5's hex habit, field widths from §1.3), reading with fread. Verify against od -A d -t x1 file (first two lines match), and predict the output for n larger than the file before running.


Chapter 8 · Processes, Signals, and Inter-Process Communication

Chapter 7 crossed from the library to the kernel for files. This chapter crosses for the rest of the map: processes — programs in flight, each with its own memory, descriptors, environment, and identity — the signals that reach them asynchronously, from the kernel, from the terminal, and from each other — and the pipes that connect them. And it answers the question the book has been circling since §1.9 drew its first pipeline: when you type ls | sort, what actually happens? The answer is a five-step dance — pipe, fork, dup2, exec, wait — that this chapter performs by hand, so that the shell stops being magic and becomes something you could have written.

The debts paid here: SIGSEGV's family, named in §1.7 and dissected in §4.7, finally gets its whole taxonomy; the 128 + n exit formula meets its second and third appearances (a real SIGPIPE death, captured below); environ, deferred from §4.9, walks at last; the function pointers promised from §5.6 report for duty as signal handlers; and the pipeline — drawn in every chapter since the first — is assembled from its parts.

One scope note, stated up front, honestly, because this chapter is the system's deepest water: the process APIs are kernel features, and this book's capture machine — a Windows toolchain — provides neither the calls nor the headers. The evidence is a captured compile refusal, shown in §8.2. So this chapter's programs split in two: the library-level four (sighand, atexitdemo, abortdemo, envdemo) and one shell capture compile and run natively on this machine, and every one of their transcripts below is real; the kernel-level pair (forkwait, pipeline) is written against the man pages' contracts and marked predicted — with a genuine Linux kernel present on this machine (its WSL2, kernel 5.15, awaiting a compiler — Exercise 8 walks the install that turns both predictions into captures).

8.1 The process model

A process is a program in flight: the loaded code, its memory (§5.5's stack, heap, and data), its open file descriptors (§7.4 — inherited from its parent at birth), its environment (§4.9's promised environ), and its identity — a pid, the process ID, a positive integer unique at any moment. Two commands see this world:

$ ps -f
UID   PID  PPID  C STIME  TTY  ... CMD
dell 1234  1230  ...           -sh
dell 5678  1234  ...           gcc -std=c17 ...

PPID is the parent's pid, and it tells the whole story: every process is born from another, in a tree rooted at pid 1 — init or systemd — with your shell partway down and your commands its children. (And /proc, the everything-is-a-file filesystem of §7.1's promise: /proc/self is your process, /proc/5678 is that gcc, and /proc/5678/maps is the memory map Exercise §4.11 sent you to find.) In C, the identity is one call: getpid() from <unistd.h>, returning the caller's own pid — used for logging, for unique temp names, and for the parent–child choreography this chapter is about.

8.2 fork: one call, two returns

fork() ("man 2 fork") creates a new process by cloning the caller: same code, same memory (copy-on-write — the pages are shared until either side writes, which is why fork is cheap), same open descriptors, same environment. It is the strangest call in the book, and the whole strangeness is the return value: fork returns twice — once in each process.

The pattern that follows is carved in every UNIX program ever written — and one stdio gotcha rides inside it:

/* forkwait.c -- fork, wait: one call, two returns, one reaping */
#include <stdio.h>
#include <unistd.h>             /* fork, _exit, getpid */
#include <sys/types.h>
#include <sys/wait.h>          /* waitpid, WIFEXITED, WEXITSTATUS */

int main(void)
{
    pid_t pid;
    int status;

    printf("parent: pid %d, about to fork\n", (int) getpid());
    fflush(stdout);                        /* §7.6: or BOTH may print it */
    pid = fork();
    if (pid < 0) {
        perror("fork");
        return 1;
    }
    if (pid == 0) {                        /* the child's whole life */
        printf("child: pid %d, fork returned 0\n", (int) getpid());
        _exit(42);                         /* not exit(42) -- see §8.7 */
    }
    printf("parent: fork returned %d, the child's pid\n", (int) pid);
    if (waitpid(pid, &status, 0) < 0) {
        perror("waitpid");
        return 1;
    }
    if (WIFEXITED(status))
        printf("parent: child exited with %d\n", WEXITSTATUS(status));
    return 0;
}

Read the shape. Both processes continue from the call — the child is a copy of the parent at the moment of the fork, so it "returns" from fork too, holding 0, and runs the same subsequent code; the if (pid == 0) arm is the child's whole life, and _exit(42) ends it. The parent, holding the child's pid, falls through, then waits (§8.4). And the fflush before the fork is §7.6's lesson wearing process clothing: the child inherits the parent's stdio buffers, byte for byte — a buffered, unflushed line in either process becomes two lines on the terminal, one flushed by each process after the fork. Flush before forking anything that prints; Exercise 1 springs the gotcha on purpose.

Predicted output — the shape, with your pids (every run's differ, §4.1's ASLR cousin: the kernel hands out pids from an ever-changing allocator):

$ gcc -std=gnu17 -Wall -Wextra -Werror -g -o forkwait forkwait.c
$ ./forkwait
parent: pid 4242, about to fork
child: pid 4243, fork returned 0
parent: fork returned 4243, the child's pid
parent: child exited with 42

(The child line may land before or after the parent's second line — two processes, two timings, no promise of order; the last two lines are ordered, because the parent cannot print the exit status before it reaps it.) And the compile refusal that motivated the honesty note, captured on this machine verbatim:

$ gcc -std=gnu17 -Wall -Wextra -Werror -g -o forkwait forkwait.c
forkwait.c:5:10: fatal error: sys/wait.h: No such file or directory
    5 | #include <sys/wait.h>          /* waitpid, WIFEXITED, WEXITSTATUS */
      |          ^~~~~~~~~~~~
compilation terminated.

The header is the kernel's interface — <sys/wait.h> declares the system call wrappers, and where there is no such kernel, there is no such header. On a Linux system — including the WSL2 sitting on this very machine, once Exercise 8 installs its compiler — the program compiles clean and prints its four lines.

8.3 exec: becoming someone else

fork makes a copy of you; the exec family replaces you. execlp("sort", "sort", (char *) NULL) — or its siblings execv, execl, execvp, and friends ("man 3 exec", the family table is Exercise 7's drill) — loads a new program into the calling process: the old code, stack, and heap vanish, and the new program starts at its main. Three properties make the family work the way it does:

The most important sentence in the section: exec is how a shell runs any command. Your shell does not contain code for ls; it forks, execs, waits. Every command line you have ever typed is the next section's recipe.

8.4 wait: reaping, zombies, and the status word

A child that exits does not entirely disappear: the kernel keeps it — pid, exit status, and resource totals — until the parent asks, because somebody must deliver the child's exit status to whoever launched it. A dead-but-unreaped process is a zombie (ps shows its state as Z, <defunct>), and the asking is wait — in its precise form, waitpid(pid, &status, 0): wait for that child, deposit its status word into my int. The status word is a small bit-encoded record, and reading it is what the macros are for:

if (WIFEXITED(status))          /* the child called exit(n) — died normally */
    printf("exited with %d\n", WEXITSTATUS(status));   /* the low 8 bits (§1.10!) */
if (WIFSIGNALED(status))         /* the child died of a signal (§8.6) */
    printf("killed by signal %d\n", WTERMSIG(status)); /* the 128 + n source */

There is §4.7's formula, decoded at last: when a shell reports 139 for your segfaulted program, it is because the child died of WIFSIGNALED, WTERMSIG was 11, and the shell encodes signal-death as 128 + n for its own exit status — the convention is the shell's, built on the macros, which are built on the word, which is the kernel's record of how your program died. Zombies are what happens when nobody reads the record: a parent that never waits leaves dead children in the table until it exits — whereupon they are re-parented to pid 1, which reaps them; a long-lived parent that never waits can fill the kernel's process table with the unreaped (Exercise 2 manufactures a zombie on purpose and watches it in ps). The discipline, as with §5.5's heap: every child you fork, you wait — and waitpid, unlike bare wait, waits for the specific child, which is what a program with several children needs.

8.5 Pipes: the shell's |, by hand

A pipe is a one-way channel between two processes: a kernel buffer with a read end and a write end, created by pipe(pfd) — pfd[0] to read, pfd[1] to write, and the arithmetic is not a typo (§1.10's table again: 0 is the read-side convention, 1 the write-side). Two laws make pipes work. Reading an empty pipe blocks until data arrives — or until every write end is closed, which is EOF: the reader's read returns 0, the normal end-of-file of §7.4, which is why a writer that is done must close, and why orphans who keep a write end open forever keep readers waiting forever. Writing a pipe with no reader kills the writer: the kernel delivers SIGPIPE (signal 13), default disposition death — the writer-to-broken-pipe rule that makes head possible, and this chapter's second real capture:

$ bash -c 'yes | head -n 1 > /dev/null; echo "yes exit status: ${PIPESTATUS[0]}"'
yes exit status: 141

yes writes forever into the pipe; head reads one line and exits, closing its end; yes's next write finds no reader, the kernel sends SIGPIPE, and yes dies — status 141 = 128 + 13, §4.7's formula in the wild, captured on this machine natively. (PIPESTATUS is bash's array of each stage's exit status — $? alone reports only the last.) If the writer would rather know than die, it installs a SIGPIPE handler, or the pipe is opened with O_NONBLOCK — §8.6's machinery, Exercise 5's drill.

Now the recipe — the five steps a shell performs for every A | B, and the reason this chapter exists:

  1. pipe(pfd) — the channel exists before anyone is born;
  2. fork — the child inherits both ends;
  3. In the child: close the unneeded end, dup2(pfd[0], 0) — move the read end onto standard input (§7.4's dup2: make descriptor 0 be this pipe), then exec B, which keeps the descriptors;
  4. In the parent: close the unneeded end, write A's output into pfd[1], close it when done (the EOF law);
  5. waitpid for the child, and read its exit status — the last stage of a pipeline is what the shell reports.

The whole dance, as a program:

/* pipeline.c -- a shell pipeline by hand: pipe, fork, dup2, exec */
#include <stdio.h>
#include <unistd.h>             /* pipe, fork, dup2, close, _exit, execlp */
#include <sys/wait.h>

int main(void)
{
    const char *lines[] = { "pear", "apple", "mango", "banana", "cherry" };
    int pfd[2];
    pid_t pid;
    int status;

    if (pipe(pfd) != 0) {
        perror("pipe");
        return 1;
    }
    pid = fork();
    if (pid < 0) {
        perror("fork");
        return 1;
    }
    if (pid == 0) {                        /* the child becomes `sort` */
        close(pfd[1]);                     /* it will never write */
        dup2(pfd[0], 0);                   /* pipe's read end = stdin */
        execlp("sort", "sort", (char *) NULL);
        perror("execlp");                  /* reached only if exec failed */
        _exit(127);                        /* the shell's not-found code */
    }
    close(pfd[0]);                         /* the parent will never read */
    for (int i = 0; i < 5; i++)
        dprintf(pfd[1], "%s\n", lines[i]);
    close(pfd[1]);                         /* EOF: sort can finish */
    if (waitpid(pid, &status, 0) < 0) {
        perror("waitpid");
        return 1;
    }
    printf("parent: sort exited with %d\n",
           WIFEXITED(status) ? WEXITSTATUS(status) : -1);
    return 0;
}

Predicted output (the child's five sorted lines, then the parent's report — ordering fixed by the wait: sort writes everything before exiting, and the parent does not print until it reaps):

$ gcc -std=gnu17 -Wall -Wextra -Werror -g -o pipeline pipeline.c
$ ./pipeline
apple
banana
cherry
mango
pear
parent: sort exited with 0

Walk the closes, because every one is a law: the child closes pfd[1] before reading (if it held its own write end open, its stdin would never see EOF — the pipe would always have a writer: itself); the parent closes pfd[0] (§8.5's recipe step 4 — a stray read end in the parent is the classic dead-reader bug); the parent closes pfd[1] after the last line — that close is the EOF that lets sort finish. And dprintf — printf's descriptor-directed sibling, POSIX — is why this program compiles with -std=gnu17 and not the strict -std=c17 of the portable four: strict ISO hides the POSIX names ("man 3 feature_test_macros" is the whole subject); from the system-side chapters on, -std=gnu17 is the honest dialect. The same five-step recipe, applied recursively, is A | B | C — Exercise 3 builds the two-stage pipeline and byte-compares it against the shell's own.

8.6 Signals: small, asynchronous, strict

A signal is the kernel's way of saying stop and notice, out of band, at a moment you did not choose. What arrives is one small number — SIGINT is 2, SIGSEGV 11, SIGPIPE 13 — from one of a handful of sources: the terminal (ctrl-C sends SIGINT to the foreground process), the kernel itself (the hardware fault of §4.7's page-zero write is SIGSEGV; §1.7's divide-by-zero is SIGFPE; a child's exit quietly sends SIGCHLD to its parent), and other processes, through the kill(2) call — which is also a shell command, and how kill -TERM 4243 delivers signal 15 to that pid (nothing about kill implies death; it is a send, and kill -USR1 is often just a "please re-read your config file").

Each signal has a disposition — a rule for what happens — and there are exactly three rules, set per signal, per process:

Disposition Meaning
default the kernel's table: mostly terminate; some core dump (a §4.7 post-mortem file); a few ignore (SIGCHLD) or stop/continue (SIGSTOP/SIGCONT)
ignore nothing happens (SIG_IGN) — the parent of §8.7 uses it for SIGCHLD to opt out of zombies
handler your function runs, then execution resumes where it was interrupted

Two signals refuse all three rules — SIGKILL (9) and SIGSTOP (19) cannot be caught, blocked, or ignored — which is exactly why they exist: the operator's guaranteed off-switch, the one kill -9 that no buggy program can survive.

Registering a handler is where §5.6's promise lands: a signal handler is a function pointer, handed to the kernel. ISO C's spelling (which is why it compiles even on this capture machine) is signal(SIGINT, on_sigint); POSIX's better spelling is sigaction ("man 2 sigaction") — same idea, plus flags, the important one being SA_RESTART ("restart interrupted syscalls instead of failing them with EINTR" — a paragraph of its own that the man page tells better than this book would):

struct sigaction sa;

memset(&sa, 0, sizeof sa);
sa.sa_handler = on_sigint;
sigemptyset(&sa.sa_mask);
sa.sa_flags = SA_RESTART;
sigaction(SIGINT, &sa, NULL);

And now the discipline, which is the section's real subject — because a handler is not a normal function. It runs asynchronously: between any two machine instructions, possibly while the interrupted code was halfway through printf's own buffer manipulation. Hence the rules, learned industry-wide the hard way:

The chapter's handler program, self-testing — it raises the signal itself after three lines, so the demo is deterministic and no second terminal is needed:

/* sighand.c -- catching SIGINT: the flag pattern, done right */
#include <stdio.h>
#include <signal.h>

static volatile sig_atomic_t stop = 0;   /* the ONLY type a handler may touch */

static void on_sigint(int signo)
{
    (void) signo;                        /* the number is in the registration */
    stop = 1;                            /* set the flag; do nothing else */
}

int main(void)
{
    int i = 0;

    signal(SIGINT, on_sigint);
    printf("running; interrupt me (or wait for the self-test)\n");
    while (!stop) {
        printf("working %d\n", i);
        i++;
        if (i == 3)
            raise(SIGINT);               /* pretend the terminal sent ctrl-C */
    }
    printf("interrupt flag set: leaving cleanly\n");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -Werror -g -o sighand sighand.c
$ ./sighand
running; interrupt me (or wait for the self-test)
working 0
working 1
working 2
interrupt flag set: leaving cleanly

The capture is the whole pattern working: raise(SIGINT) delivers the signal — to this process, exactly as ctrl-C would — the handler runs between two loop iterations, sets stop, returns, and the loop — the program — notices and exits cleanly with its message, its buffers, and its exit status intact. (Run it interactively and press ctrl-C at any moment: same three lines of fate, at a time of your choosing. And then try kill -KILL on it, and meet the special pair in person — Exercise 4.) A program that wants to die gracefully on ctrl-C — an editor saving its file, a daemon closing its log — is exactly this program with the work changed.

8.7 exit's whole story: atexit, abort, and environ

§1.10 told exit's short story (the status, the low 8 bits). The long story has three parts. exit(n) is a process: run every atexit-registered handler — in LIFO order, stack-like — then flush and close every stream (§7.2's "fclose returns a value" promise, kept en masse), then terminate with status n. _exit(n) skips all of it: no handlers, no flush — terminate now, status n. The place for the brusque version is exactly one: inside a forked child that has just failed its exec — like §8.5's _exit(127) — because that child is a copy holding a copy of the parent's stdio buffers, and letting it exit would flush the parent's half-written buffer a second time (§8.2's gotcha, in its terminal form). Everywhere else, exit — or return from main, which is exit — is the polite spelling. The handlers, demonstrated:

/* atexitdemo.c -- exit's whole story: atexit handlers, in LIFO order */
#include <stdio.h>
#include <stdlib.h>

static void bye3(void)
{
    printf("third registered: runs FIRST (LIFO)\n");
}

static void bye2(void)
{
    printf("second registered\n");
}

static void bye1(void)
{
    printf("first registered: runs LAST\n");
}

int main(void)
{
    atexit(bye1);
    atexit(bye2);
    atexit(bye3);
    printf("main: returning 7\n");
    return 7;
}
$ gcc -std=c17 -Wall -Wextra -Werror -g -o atexitdemo atexitdemo.c
$ ./atexitdemo
main: returning 7
third registered: runs FIRST (LIFO)
second registered
first registered: runs LAST
$ echo $?
7

return from main is exit — handlers and all. abort() is the third exit: raise SIGABRT (signal 6), default disposition death with a core dump — the crash you call on yourself, for "the program's own state is so broken that continuing would be a lie" (the standard library itself calls it when it detects heap corruption — §5.5's double-free world). Its demo prints on stderr — §7.6's lesson, applied pre-mortem: the message survives:

/* abortdemo.c -- death by signal, on purpose: SIGABRT */
#include <stdio.h>
#include <stdlib.h>

int main(void)
{
    fprintf(stderr, "about to abort\n");   /* stderr: survives death (§7.6) */
    abort();                              /* raises SIGABRT; default: die */
    printf("never gets here\n");
    return 0;
}
$ gcc -std=c17 -Wall -Wextra -Werror -g -o abortdemo abortdemo.c
$ ./abortdemo
about to abort
$ echo $?
127

Captured on this machine, whose abort encoding is 127; on Linux the same program dies of signal 6 — bash prints Aborted, and $? records 134 = 128 + 6, §4.7's formula's third appearance in the chapter. (The counting, for the record: buggy.c's segfault was 139 = 128 + 11; §8.5's yes died 141 = 128 + 13; abortdemo is 134 = 128 + 6. One formula, three deaths, all genuine.)

And environ — §4.9's deferred promise, now walked: the environment is an array of strings, KEY=value one per entry, terminated by a NULL pointer — argv's twin convention, which is no accident, and the reason §4.9's sentinel loop walked it identically. The ISO face is getenv("USER") (a lookup); the full array is the global environ (<unistd.h> under POSIX, declared by this toolchain's <stdlib.h> — the comment in the program says so):

/* envdemo.c -- the environment block: argv's NULL-terminated twin */
#include <stdio.h>
#include <stdlib.h>             /* environ: declared here on this toolchain,
                                   and in <unistd.h> under POSIX */

int main(void)
{
    int n = 0;

    for (char **p = environ; *p != NULL; p++) {
        n++;
        if (n <= 3)
            printf("%s\n", *p);
    }
    printf("... %d environment variables\n", n);
    return 0;
}
$ gcc -std=gnu17 -Wall -Wextra -Werror -g -o envdemo envdemo.c
$ ./envdemo
ALLUSERSPROFILE=C:\ProgramData
APPDATA=C:\Users\DELL\AppData\Roaming
COMMONPROGRAMFILES=C:\Program Files\Common Files
... 93 environment variables

The capture is this machine's block (93 variables, Windows-named); a Linux login answers with USER=you, HOME=/home/you, SHELL=/bin/bash, PATH=..., TERM=... — different names, identical structure, and it is inherited: a child's environ is a copy of its parent's, which is how export works in your shell and how exec's e-suffix (§8.3) hands a child a different one. Modifying it is setenv/unsetenv (and the older, sharper putenv — read its man page's warning before touching it), changing your copy, which your future children will inherit. A -std=gnu17 program, note — environ is POSIX's name, and strict ISO hides it, exactly like §8.5's dprintf.

8.8 Summary

A process is a program in flight — pid, memory, descriptors, environment — born of a parent in a tree rooted at pid 1; fork clones it, returning 0 in the copy and the child's pid in the original, with copy-on-write making it cheap and the inherited stdio buffers making §7.6's pre-fork flush mandatory. exec replaces the running program — pid and descriptors survive, the caller survives only on failure (−1, and 127 is the shell's not-found convention); the letter suffixes are list/vector, PATH-search, fresh-environment. waitpid reaps: the kernel holds every dead child's status word until the parent asks — zombies are the unasked — and the word decodes with WIFEXITED/WEXITSTATUS and WIFSIGNALED/WTERMSIG, which is where the shell's 128 + n encoding comes from (139, 141, and 134, all captured or derived in this chapter). A pipe is a kernel buffer with two ends and two laws: a reader blocks until data or the last writer's close — the EOF discipline that the recipe's every close serves — and a writer into a readerless pipe dies of SIGPIPE (141, captured natively). The shell's A | B is five steps — pipe, fork, close/dup2/exec in the child, close/write/close in the parent, waitpid — performed by pipeline.c, with -std=gnu17 as the system-side dialect because strict ISO hides the POSIX names. Signals are asynchronous numbers from terminal, kernel, and other processes, with three possible dispositions — default, ignore, handler — except SIGKILL/SIGSTOP, which take no orders; handlers are function pointers registered through signal (ISO) or sigaction (POSIX, with SA_RESTART), and obey one discipline: volatile sig_atomic_t flags, nothing fancy, write not printf. exit runs atexit handlers LIFO then flushes; _exit does neither, and belongs in the failed-exec child; abort is self-inflicted SIGABRT (134 on Linux); and environ is the NULL-terminated KEY=value block — argv's twin, inherited per child, editable with setenv.

Next: the final chapter — system programming best practices, and the end of the book: the stack as the machine really lays it out (the frames of §3.7's recursion, finally drawn), the debugging method with gdb as a working skill (breakpoints on §8's child, a backtrace through a §4.7 core dump), the sanitizer and valgrind reports read line by line, make automating §1.4's pipeline, the preprocessor at full power — the book's practices collected into one checklist, and the view from the end of the road.

8.9 Exercises

Recommended flags on (-std=gnu17 for anything POSIX); the kernel-side exercises want a real Linux — this machine's WSL2 counts, once Exercise 8 installs its compiler.

  1. Two forkwait experiments. (a) Make the child die of a signal instead — replace _exit(42) with abort() — and give the parent the WIFSIGNALED/WTERMSIG branch; predict the parent's last line before running, then verify. (b) Delete the fflush, run with output redirected to a file, and predict — before looking — which line appears twice, and in which process each copy flushed.

  2. Zombie hour. Write zombify.c: fork a child that exits immediately, and have the parent sleep(10) before waitpid. Run it, and in another terminal, ps -l until you find the child in state Z (<defunct>). Explain what the kernel is holding and for whom; then add the waitpid and confirm the zombie never appears.

  3. Two-stage pipeline. Extend pipeline.c so the data runs sort | head -2: two pipes, three processes — fork an intermediate child that dup2s pipe one's read end onto its stdin and pipe two's write end onto its stdout, then execs head. Predict the exact output first, then verify byte-for-byte against printf 'pear\napple\nmango\nbanana\ncherry\n' | sort | head -2.

  4. The special pair, in person. Run ./sighand interactively and interrupt it with ctrl-C — the graceful exit, at your timing. Then kill -TERM it (no handler: default disposition, instant death), add a SIGTERM handler that also just sets the flag, and finally kill -KILL it mid-loop: the one signal that cannot be caught, and the reason it exists.

  5. SIGPIPE, precisely. (a) Reproduce §8.5's 141 capture yourself, and explain PIPESTATUS's necessity with $? alone. (b) Write wrpipe.c: fork; the child exits without reading; the parent loops writing single bytes and counting them, having installed a SIGPIPE handler that writes the count (a sig_atomic_t) to stderr and _exits — the flag pattern and the safe-write rule, both in one program.

  6. envgrep. Write envgrep prefix — print only the environment variables starting with prefix, using strncmp (§6.3's cousin; "man 3 strncmp" for its three-way return). Verify against env | grep '^PREFIX', and note the one thing strncmp can do that your loop must copy: a bounded comparison that cannot walk off a short string.

  7. The exec family, drilled. Rewrite §8.5's child using execv with the full path to sort (find it: command -v sort). Then build the family table — execl, execlp, execv, execvp, execle, execvpe — and write a one-sentence rule for choosing among the letters; verify each of your six sentences against "man 3 exec"'s own description.

  8. Make the predictions real. This machine's WSL2 runs a genuine Linux kernel as root, with no compiler yet: wsl -e sh -c 'apt-get install -y gcc libc6-dev' (a network install — your call), then wsl -e sh -c 'cd "/mnt/e/MfR/ITL-Books/The C Programming Language, 2nd Ed.PlentyofeBooks.net/examples/ch08" && gcc -std=gnu17 -Wall -Wextra -Werror -g -o forkwait forkwait.c && ./forkwait && gcc -std=gnu17 -Wall -Wextra -Werror -g -o pipeline pipeline.c && ./pipeline' — and compare, byte for byte, against §8.2's and §8.5's predicted transcripts: the shapes should match exactly; the pids will not, and §8.2 said why.


Chapter 9 · System Programming Best Practices, and the End of the Book

This is the last chapter, and it has three jobs. First, to pay the book's final deferred debts: the stack of §3.7's recursion, promised "explained properly in Chapter 9"; the debugger of §1.11, promised "a working skill"; the sanitizer reports of §4.7 and §5.5, promised "read line by line"; make, promised since §1.4 as the automation of the compile–link pipeline; and #define's "power", deferred twice with the words "until you know enough to miss it". Second, to collect the practices — the whole book's disciplines, gathered into one checklist you can hold against any program, including your own. Third, to say goodbye properly: what you now know, what this book honestly did not teach you, and where to go next.

A scope note in the book's established style, stated once for the chapter: this capture machine's toolchain ships no debugger, no make, and no sanitizer runtimes — all three facts probed, all three refusals already in evidence (-lasan in §4.7, -lubsan in this chapter's §9.3). So this chapter's transcripts split by kind: the programs — the macro demo, the undefined-behavior demo, the crash — compiled clean and ran here, and their outputs below are real captures; the gdb session and make session are built line-by-line from the tools' documented behavior and the Makefile's own recipes (a deterministic derivation: both print exactly what their inputs say), and each is labeled accordingly, with exercises that run them for real on any Linux system.

9.1 The stack, at last

§3.7 drew recursion as "a million-deep recursion is a crash" and deferred the reason. Here is the machine's answer. Automatic storage (§2.1 — block-scoped variables, function parameters, return addresses) lives in a region called the stack, and it is exactly what its name says: frames pushed one on top of another as functions call functions, popped as they return. Each frame holds one call's world — its parameters, its locals, and the return address that says where to resume — and main's frame is at the bottom:

   higher addresses
  +------------------------+
  | main's frame           |   total, p, argc, argv
  +------------------------+
  | count_file's frame     |   f, c, buf[4096], in_word      <-- pushed
  +------------------------+
  | (recursion: frame n)   |   v, len
  +------------------------+                                    <-- grows
  | (recursion: frame n+1) |
  +------------------------+
        ...downward...          the stack grows this way
  +------------------------+
  | the guard page         |   <-- unwritable: the fence
  +------------------------+
   lower addresses

The size of this region is finite and set by policy: the shell's ulimit -s reports it in kilobytes. Captured on this machine:

$ ulimit -s
2032

(two megabytes and a little; Linux distributions commonly answer 8192.) The guard page at the bottom fence is deliberately unwritable, so a stack that runs out does not silently trample other memory: the next frame lands on it, the write faults, and the kernel delivers §4.7's old friend SIGSEGV — a stack overflow, exit 139, exactly like a wild pointer but with an honest cause. §3.7's million-deep recursion dies here, on the fence, loudly — the good crash, not the silent one. The contrast completes the memory map of §5.5: the stack is limited, automatic, and fast (push a frame, pop a frame, no bookkeeping); the heap is enormous, manual, and owned (malloc, free, the discipline of §5.5); static storage (§2.1) sits in a fixed region for the whole run. Choosing storage is now a three-way engineering decision you can actually draw.

9.2 The debugging method

The tools matter less than the loop they serve, so the loop first — five steps, every time:

  1. Reproduce. A bug you cannot trigger on demand you cannot fix; find the smallest input that does it (§3.3's scripted sessions — printf '...' | ./prog — exist for exactly this).
  2. Read the diagnostic before touching anything. The compiler's warnings (§1.11), the shell's exit status (§1.10), the kernel's signal number (§8.4): the system has already told you a great deal.
  3. Locate, with a debugger. Break near the misbehavior, or run to the crash and walk the backtrace — the stack of §9.1, printed frame by frame, innermost (#0) first: how did execution get here?
  4. Form one hypothesis, test it, fix one thing. Not five things — one; then re-run the reproduction. A fix without a re-verified reproduction is a guess.
  5. Keep the reproduction. The input that found the bug becomes a test that keeps it fixed.

The debugger's working vocabulary is small — break main (stop on entry), run args (go), next/step (over / into), print x (inspect), bt (backtrace), quit — and §1.11's minimal session, grown up, debugs this chapter's own crash:

/* crash.c -- wrong on purpose: a bug worth a debugger */
#include <stdio.h>
#include <stdlib.h>

int main(int argc, char *argv[])
{
    long total = 0;
    int *p = NULL;

    if (argc > 1)
        total = strtol(argv[1], NULL, 10);   /* NULL end: no check (§6.6) */
    printf("total is %ld\n", total);
    *p = 42;                                 /* THE bug: page zero (§4.7) */
    return 0;
}

The crash itself, captured on this machine, all too familiar by now:

$ gcc -std=c17 -Wall -Wextra -Werror -g -o crash crash.c
$ ./crash 7
total is 7
Segmentation fault
$ echo $?
139

And the session — labeled as promised: written from gdb's documented behavior (this toolchain ships no debugger; Exercise 1 runs it for real), with the -g build of §1.11 attached:

$ gdb -q ./crash
(gdb) run 7
Starting program: ./crash 7
total is 7

Program received signal SIGSEGV, Segmentation fault.
main () at crash.c:13
13      *p = 42;
(gdb) bt
#0  main () at crash.c:13
(ggdb) print p
$1 = (int *) 0x0
(gdb) quit

Read what happened, because it is the whole method in five lines. The program ran under the debugger; the kernel's SIGSEGV stopped it at the faulting line — crash.c:13, the *p = 42, no guessing which of fifteen dereferences died; the backtrace showed the road there (#0 main — one frame, no mystery caller); and print p answered 0x0 — the hypothesis in one keystroke: the pointer is NULL, §4.7's page-zero write wearing a new coat. Fix, rebuild, re-run: step 4's loop. (Two extensions worth knowing: gdb ./prog core opens a §4.7 core dump — a saved image of the crashed process, the backtrace without re-running; and with §8's children, set follow-fork-mode child follows the fork instead of the parent — breakpoint in the child, exactly as promised in the last chapter.) And the honest word about printf debugging: it is not shameful, it is slower — the debugger already prints every variable without a rebuild; when you do add prints, they go to stderr (§7.6: they survive death and bypass the buffer), and they leave as asserts or debug-build lines (§9.6), not as cargo.

9.3 The sanitizers, read line by line

§1.11 introduced them; §4.7 and §5.5 promised their reports "read properly". The motivating capture first — undefined behavior, running silently, on this machine:

$ gcc -std=c17 -Wall -Wextra -Werror -g -o ub ub.c
$ ./ub
INT_MAX + 1 is -2147483648

That is §1.7's promised lie: INT_MAX + 1 is undefined, the compiler compiled it clean under -Werror (it is not required to diagnose it), and this machine's answer was the wrap to -2147483648. Another machine, another day, another number — or a deleted line: the optimizer, licensed to assume signed overflow cannot happen, may reason the loop for (int i = 0; i <= INT_MAX; i++) never ends and remove the check. Silent and wrong, §4.7's fear, on screen. The tool that turns this into a loud, located, one-line confession is UBSan — and the third honest refusal of the book, captured:

$ gcc -std=c17 -Wall -Wextra -g -fsanitize=undefined -o ubsan ub.c
ld.exe: cannot find -lubsan: No such file or directory
collect2.exe: error: ld returned 1 exit status

(-lasan in §4.7, -lubsan here: this toolchain ships no sanitizer runtimes; every Linux distribution's GCC does.) What the same run prints there — the report's anatomy, from the tools' documentation, so you can read one on sight:

valgrind (§1.11's third musketeer) does the same class of work without recompiling — valgrind ./prog — slower, but the binary you already shipped. The practice that makes all three multiply rather than add: debug builds under sanitizers, release builds under -O2, and the reproduction of §9.2's step 1 run under both — the sanitizer names the line; the optimizer shows what the license to assume "cannot happen" really costs.

9.4 Safe memory practice

§2.3 promised that "memory-safe code checks sizes with care bordering on paranoia"; here is the whole discipline, each rule with its chapter:

Paranoia, yes — but each rule is one line of code, and §9.3's tools police all of them automatically. The craft is writing the line anyway.

9.5 Building: make

§1.4 promised the compile–link pipeline would "matter later, when programs grow to many files"; §3.8 built it by hand (gcc -c each file, then link); the debt was always "what make automates". Here is the final chapter's capstone — a three-file tool in §3.8's exact shape, a pocket wc, with the book's whole arc in its three files: a structure record (§5.2), a header with a guarded prototype (§3.8), a client thin enough to read in one glance (§3.6):

/* count.h -- interface of count: lines, words, bytes */
#ifndef COUNT_H
#define COUNT_H

#include <stdio.h>
#include <stddef.h>

struct counts {                  /* a record, §5.2 */
    size_t lines;                /* newline characters seen */
    size_t words;                /* runs of letters/digits */
    size_t bytes;                /* raw bytes read */
};

/* Fill *c with the tallies of everything readable from f.
   Returns 0, or 1 if a read error intervened. */
int count_file(FILE *f, struct counts *c);

#endif
/* count.c -- implementation of count */
#include <stdio.h>
#include &lt;ctype.h>
#include <stdbool.h>
#include "count.h"

int count_file(FILE *f, struct counts *c)
{
    unsigned char buf[4096];
    bool in_word = false;

    c->lines = c->words = c->bytes = 0;
    for (;;) {
        size_t n = fread(buf, 1, sizeof buf, f);

        if (n == 0)
            return ferror(f) ? 1 : 0;         /* §7.3's distinction */
        c->bytes += n;
        for (size_t i = 0; i < n; i++) {
            unsigned char ch = buf[i];        /* §6.5's license */

            if (isalnum(ch)) {
                if (!in_word) {
                    c->words++;
                    in_word = true;          /* a word begins */
                }
            } else {
                in_word = false;              /* a word ends */
                if (ch == '\n')
                    c->lines++;
            }
        }
    }
}
/* count_main.c -- count: a pocket wc */
#include <stdio.h>
#include "count.h"

int main(int argc, char *argv[])
{
    struct counts c;
    FILE *f;

    if (argc != 2) {
        fprintf(stderr, "usage: count file\n");
        return 2;
    }
    f = fopen(argv[1], "rb");        /* "rb": bytes are bytes (§7.2) */
    if (f == NULL) {
        perror(argv[1]);
        return 1;
    }
    if (count_file(f, &c) != 0) {
        fprintf(stderr, "count: read error\n");
        fclose(f);
        return 1;
    }
    fclose(f);
    printf("%7zu %7zu %7zu %s\n", c.lines, c.words, c.bytes, argv[1]);
    return 0;
}

Every line is a chapter you have read: the usage check and exit 2 (§3.9), the checked open and close (§7.3), ferror at the loop's end (§7.3's debt, paid again, in a library), unsigned char through ctype's license (§6.5), a bool state machine for word runs (§3.4's flag pattern, at byte granularity), block reads with short counts normal (§7.4), and the (record, stream) contract of §5.2 feeding a pointer from main (§4.2). The build is a Makefile — make's script, in its own tiny language of targets, dependencies, and recipes:

# Makefile -- build `count` (§1.4's pipeline, automated)

CC      = gcc
CFLAGS ?= -std=gnu17 -Wall -Wextra -Werror -g
OBJECTS = count_main.o count.o

count: $(OBJECTS)
    $(CC) $(CFLAGS) -o count $(OBJECTS)

count_main.o: count_main.c count.h
    $(CC) $(CFLAGS) -c count_main.c

count.o: count.c count.h
    $(CC) $(CFLAGS) -c count.c

clean:
    rm -f count $(OBJECTS)

.PHONY: clean

Read it as three ideas. A rule is target: dependencies followed by a tab-indented recipe — the tab is not style, it is the grammar. make count finds count's rule, and builds only what is out of date: for each prerequisite, recursively the same question — and §1.4's pipeline runs itself, one gcc -c per changed .c, one link, skipping everything already newer than its sources. The dependency lines say the truth about the code — count_main.o depends on the header too (§3.8's "including your own interface"), which is why touch count.h rebuilds both objects and nothing else. ?= (§8.5's dialect note's cousin) assigns only if nothing already did — the command line and environment can override CFLAGS, which is how a project carries one flags policy while a debug build adds -fsanitize=address. And .PHONY: clean marks clean as a command, not a file — make clean always runs it, even if someone left a file named clean.

The session — labeled: this toolchain ships no make (probed twice, both subsystems), so the transcript is what make itself documents and what the recipes above literally say; Exercise 5 runs it for real:

$ make
gcc -std=gnu17 -Wall -Wextra -Werror -g -c count_main.c
gcc -std=gnu17 -Wall -Wextra -Werror -g -c count.c
gcc -std=gnu17 -Wall -Wextra -Werror -g -o count count_main.o count.o
$ ./count count.c
      33      58     688 count.c
$ make
make: 'count' is up to date.
$ touch count.h
$ make
gcc -std=gnu17 -Wall -Wextra -Werror -g -c count_main.c
gcc -std=gnu17 -Wall -Wextra -Werror -g -c count.c
gcc -std=gnu17 -Wall -Wextra -Werror -g -o count count_main.o count.o
$ make clean
rm -f count count_main.o count.o

(The three recipe lines are the Makefile's own, substituted; make's up to date. line is its documented message; and the count.c tallies are the one thing on this page not knowable without a run — Exercise 5 runs it and checks every column against wc, which is the honest test of the tool.) From here, make scales exactly as far as C does: variables, pattern rules (%.o: %.c), include of shared fragments — a day's reading in the GNU make manual, all of it the same three ideas.

9.6 The preprocessor, at full power

Two chapters deferred #define's "power" until you would miss it; here is what it is for — and the captured evidence of what it costs when spelled casually. A function-like macro takes parameters, textually substituted — the power (code generated per use, types erased) and the trap (code generated per use, types erased) in one construct:

/* dbg.c -- macros with parameters, and debug builds */
#include <stdio.h>

#define SQUARE_BAD(x)  x * x            /* no parentheses: the classic bug */
#define SQUARE(x)      ((x) * (x))      /* parens everywhere: the rule */

#ifdef DEBUG
#define DPRINTF(...)   printf(__VA_ARGS__)
#else
#define DPRINTF(...)   ((void) 0)       /* vanishes in the release build */
#endif

int main(void)
{
    printf("SQUARE_BAD(2 + 3) is %d\n", SQUARE_BAD(2 + 3));
    printf("SQUARE(2 + 3) is %d\n", SQUARE(2 + 3));
    DPRINTF("this line only exists in a -DDEBUG build\n");
    return 0;
}

Captured, both builds:

$ gcc -std=c17 -Wall -Wextra -Werror -g -o dbg dbg.c
$ ./dbg
SQUARE_BAD(2 + 3) is 11
SQUARE(2 + 3) is 25

$ gcc -std=c17 -Wall -Wextra -Werror -g -DDEBUG -o dbg2 dbg.c
$ ./dbg2
SQUARE_BAD(2 + 3) is 11
SQUARE(2 + 3) is 25
this line only exists in a -DDEBUG build

The first capture is the trap, sprung: SQUARE_BAD(2 + 3) is textually 2 + 3 * 2 + 3 — eleven, not twenty-five, because the macro's author forgot that the caller's expression is spliced into a larger one. Hence the discipline, total and non-negotiable: parenthesize the whole macro body, and every parameter, every time — ((x) * (x)); a macro body with ; wears do { ... } while (0); and never let a parameter be evaluated twice (SQUARE(i++) increments twice — a function call, the honest alternative, evaluates once). The rule of §1.6 and §2.6, kept all book long: macros for what functions cannot do, functions for everything else.

The second capture is the other half of the power — and a genuine best practice: the debug build. #ifdef DEBUG + DPRINTF(...) is this book's practice in two lines: with -DDEBUG, the diagnostic line exists and prints; without it, the preprocessor deletes the line before the compiler ever sees it — no cost, no leftover code, no risk of a debug print in production. __VA_ARGS__ (C99) is how a macro swallows printf's whole argument list; the ((void) 0) arm is the legally-discarded value. The same lever runs assert (<assert.h>): your invariants, checked on every line in debug builds, compiled away by -DNDEBUG in release — assert the invariant ("this pointer is not NULL, this index is in range"), never the recoverable error (a failed assertion is abort(), §8.7, death by design: a bug announcement, not error handling). And the third preprocessor practice, used since §3.8 and now seen whole: the include guard — #ifndef/#define/#endif — plus -I dir for your own include directories, and conditional compilation for portability, which this book exercised honestly from Chapter 1's LLP64 notes onward.

9.7 The practices, collected

The book's disciplines, one list, each with its home — hold this against any program, yours included:

  1. Compile with -std=c17 (or -std=gnu17 on the system side), -Wall -Wextra -Werror -g, always (§1.11, §8.5) — and read every diagnostic to the end.
  2. Check every return: NULL from allocators and openers (§5.5, §7.2), ssize_t verdicts (§7.4), fclose (§7.2), exec (§8.3), every one.
  3. Bounds are part of the loop's sentence (§4.4) — first element through one-past-the-end, size_t against size_t (§2.3).
  4. Diagnostics on stderr; the product on stdout (§1.10) — and flush before abnormal exits and forks (§7.6, §8.2).
  5. Exit statuses tell the truth: 0 success, 1 failure, 2 usage (§1.10, §3.9), 127 not-found (§8.3).
  6. Ownership is documented in the interface (§5.5, §5.7): one allocation, one owner, one free — and freed pointers are dead.
  7. const states the read/write contract (§4.6); structures are ordered large-to-small (§5.4) and compared member-by-member.
  8. Terminators are budgeted (§6.1, §6.4): every write reserves room for '\0'; snprintf over strcpy, always bounded, always terminated.
  9. Check for input where none was given: argc first (§3.9), if (p) before *p (§4.7), end == start for numbers (§6.6), ferror after every read loop stops (§7.3).
  10. Signal handlers touch nothing but volatile sig_atomic_t flags (§8.6), and write with write, not printf.
  11. Every child is waited for (§8.4), and every failure path cleans up what it opened — the goto ladder when the unwinding is real (§3.5).
  12. Debug builds carry sanitizers and asserts; release builds carry -O2 (§9.2, §9.3, §9.6) — and the bug's reproduction becomes its regression test.
  13. Honest limits beat silent lies — the running theme: the leak that crashes is kinder than the leak that doesn't (§5.5), the loud segfault than the plausible wrong number (§4.7), the refused compile than the phantom function (§7.7, §8.2).
  14. The build is code too: guarded headers, thin mains, one truth per fact (§3.8), and make to say it once (§9.5).
  15. Write for the reader — braces that show ownership (§3.1), loops that read as sentences (§3.4), names that say what things mean, and comments that explain why, never what.

9.8 What this book did not teach you

Honesty at the exit, as at the entrance. This book walked one road: standard C, then the UNIX system's files, processes, and signals. Three provinces share the border, each a next book in itself:

9.9 Conclusion: the view from here

Nine chapters ago, the book's first program was puts("Hello, Linux!") — one line, compiled with two keystrokes of flags, run with ./. The last program built by these pages was a three-file tool, compiled by a Makefile, structured as interface, implementation, and client, counting words out of a file through the kernel's own block-read interface. Between those two programs, you crossed the entire map the preface promised: the language first — types and operators (2), control flow and structure (3), pointers and arrays (4), structures and memory management (5), strings (6) — and then the system — files and the kernel's calls (7), processes and signals (8), and here, the practices that keep both halves honest (9).

More was built along the way than programs. Every chapter paid the previous one's debts in public: ferror promised in Chapter 1 and delivered in 7; argv promised as a spell and explained as machinery in 4; the line-numbering tool that split long lines in Chapter 1 and was repaired, with the full string toolkit, in 6; the pipeline drawn in every chapter since the first and finally assembled from pipe, fork, dup2, exec, and wait in 8. That structure was the book's second curriculum: C programs on Linux are not fragments — they are one long construction, each layer load-bearing, every abstraction a promise about memory that some later chapter keeps. You now know what the promises are, and what keeps them.

And the book kept one promise above all, deliberately, in a way no other treatment of C that we know of has attempted: it never showed you an output it did not capture, a claim it did not check, or a limit it did not name. Some of those names were this machine's — LLP64 against LP64, a libc without getline, a toolchain without sanitizers, an abort that exits 127, a descriptor layer that quietly doubled the newlines — and each became a lesson instead of an embarrassment, because that is what platform differences are: the machine's answers to questions the standard deliberately leaves open. Reading a real program's weirdnesses with that instinct — what question is this the answer to? — is the habit this book most wanted to leave you with. The classic 1978 text by Kernighan and Ritchie, whose lineage this whole project honors, taught the language with unmatched economy in 270 pages; this book took the longer road through the system underneath, because that is where a Linux C programmer actually lives.

So: write the tool. Spec it — usage line, exit codes, what it reads, what it prints — then build it thin-main'd and guarded-header'd, compile it under all the flags, run it under the sanitizers, debug it by the five steps, and hold it against §9.7's list. That loop is the whole book, one more time, and the last exercise is exactly it. The system is open — everything is a file, everything is documented in man, and everything you have just learned to read is running, right now, one ps away.

Go write something worth debugging.

9.10 Exercises

The last set. Recommended flags — all of them — and a real Linux system for the probed-out halves; this machine's WSL2 with gcc installed (§8.9's Exercise 8) covers every one.

  1. The session, for real. Run §9.2's crash under gdb yourself: break main, run 7, next to the faulting line, print p, bt — and compare your transcript against the chapter's, line by line. Then set a breakpoint on crash.c:13 before running, and finally fix the bug and re-verify. Bonus: set follow-fork-mode child, and breakpoint the child in Chapter 8's forkwait.

  2. Overflow the stack, on purpose. Write deep.c — a function that recurses with no base case, passing its depth along; run it; note the exit status and §8.4's decode. Then run it under gdb and read the backtrace: a wall of identical frames, §3.7's promise made visible. Vary ulimit -s (in a subshell, so your shell survives) and watch the crash depth change.

  3. The reports, one per chapter. On a Linux toolchain: ub.c under UBSan; Chapter 4's buggy.c and Chapter 5's useafter.c and leak.c under ASan — read each report's parts aloud (error line, access, allocation/free stacks, summary), then fix each and confirm silence. That is §9.3's anatomy, four times.

  4. The debug build, applied. Add DPRINTF-style lines and one assert to count.c (assert the interface's contract: f != NULL, c != NULL); verify the release build carries none of it, the -DDEBUG build prints it, and -DNDEBUG silences the asserts. Then add DPRINTF("block of %zu bytes\n", n) in the read loop and count the blocks on a 10 MB file.

  5. Make it real, then honest. Run §9.5's sequence on a real make: from clean, up-to-date, touch count.h, clean — and reconcile every line with the chapter's transcript. Then run ./count count.c and check all three columns against wc count.c (wc's word rule differs slightly from ours — where, and does it matter here?); finally, add a test target that runs that check and fails the build if it doesn't match.

  6. The audit. Take Chapter 6's nl2.c and audit it against §9.7's list, in writing: every rule, pass or fail (§1.9 flagged one debt it never paid — the ferror check after the loop — confirm the audit finds it). Fix every failure; re-verify with the long-line demo, byte-identical.

  7. The last exercise: your tool. Spec it (a usage line, a man-page-style SYNOPSIS comment in the header), build it (three files minimum, guarded header, a Makefile with clean), verify it (the recommended flags, a sanitizer pass, the five debugging steps on whatever you find), and audit it against §9.7 — in writing, all fifteen. When it passes: it is finished, and so is this book.