Build Tooling & Debugging Literacy

Build Systems & an Intro to CMake

Why this matters

Every exercise so far has been one file, compiled with one command. Real projects are dozens to thousands of files, built by people other than you, on machines other than yours. Two problems fall out of that immediately: how do separate files call each other’s functions, and how do you describe “compile all of this, in this order, with these settings” in a way that doesn’t depend on one person’s terminal history. The first is a language question (declarations vs. definitions – you already have the tools). The second is what build systems, and CMake specifically, are for.

Splitting one file into many is a declaration problem

The worked example marks where a real split would happen: the mean function is genuinely independent of main. In a real project, mean and any other statistics functions would live in stats.c, with their signatures listed in a header, stats.h:

// stats.h
double mean(const double *values, int count);

main.c would then #include "stats.h" and call mean without ever seeing its implementation. This is exactly the declaration/definition split from Module 3’s first lesson, just spread across files instead of positioned above/below each other in one file: the header supplies declarations so every file that includes it can compile calls to mean, while exactly one .c file supplies the definition. This is also why headers almost never contain function bodies for ordinary functions – if two .c files both included a header containing a full definition, you’d get the “redefinition” error from that same lesson, just triggered by an #include instead of a copy-pasted duplicate. A real stats.h would also be wrapped in the header guard from the last lesson (#ifndef STATS_H / #define STATS_H / … / #endif) – this is precisely the scenario those three lines protect against: stats.h ending up #included more than once while building the same file.

gcc main.c stats.c -o app runs the same four stages as before, just with two source files: main.c and stats.c are each preprocessed, compiled, and assembled into their own object files (main.o, stats.o) independently – neither file needs to see the other’s source, only declarations, to compile. Only at the link stage do they need to actually meet: the linker matches main.o’s reference to mean against stats.o’s definition of it, exactly like the undefined-reference lesson, except now “no definition found” would mean you forgot to pass stats.c to the compiler at all, not that you forgot to write the function.

Why not just type that gcc command every time?

For two files, typing the command by hand is fine. For a real project it isn’t: you’d need to know every source file, in what order, with what flags, for every platform, every time anything changes – and re-type all of it correctly. A build system is a tool that reads a description of your project once and figures out the right compiler invocations for you, including which files actually need recompiling after a change.

CMake is the most common one for C and C++. Critically, CMake itself doesn’t compile anything – it’s a generator. You describe your project in a CMakeLists.txt file, and CMake reads it and produces build files for whatever’s actually available on the machine (Makefiles on Linux, an Xcode project on macOS, Visual Studio files on Windows), which some other tool then executes to actually invoke the compiler.

cmake_minimum_required(VERSION 3.20)
project(stats_demo LANGUAGES C)

add_executable(stats_demo
    src/main.c
    src/stats.c
)

target_include_directories(stats_demo PRIVATE include)

Reading it top to bottom: cmake_minimum_required states the oldest CMake version this file assumes, so a too-old CMake fails loudly instead of misbehaving. project(...) names the project and its language. add_executable declares a build target – the name of a program to produce, and the list of source files that go into it (each still compiled separately, then linked together, exactly as above). target_include_directories tells the compiler where to look for headers #included with quotes, like "stats.h".

The normal workflow is two commands, run from the project’s root:

cmake -S . -B build
cmake --build build

The first configures: it reads CMakeLists.txt and generates build files into a build/ directory (kept separate from your source on purpose – an “out-of-source build” – so generated files never get mixed up with files you wrote, and you can delete build/ entirely to start clean). The second actually builds, invoking whatever the first command generated. You only need to re-run the first command when CMakeLists.txt itself changes; after that, cmake --build build alone re-compiles just the files that changed.

static at file scope: hiding a name from every other file

You’ve already seen static change a local variable’s lifetime (Module 3’s “stale state” exercise – making it persist across calls instead of being recreated each time). At file scope – applied to a function or a global variable, written outside any function – static means something related but distinct: it makes that name invisible to every other translation unit, even though the file it’s defined in can still use it freely.

// stats.c
static double running_total = 0.0;   // only visible inside stats.c

double mean(const double *values, int count) { ... }   // visible to other files (no static)

Without static, every function and global variable you define has external linkage by default – visible and linkable from any other .c file in the project, whether or not that’s what you intended. A helper function that’s only ever meant to be called from within stats.c itself, never from main.c, should be marked static: it documents “this is private to this file” and, as a real benefit, means another file can freely define its own same-named function without the linker ever seeing a conflict. This is the closest C gets to the encapsulation you may have seen in other languages as private – scoped to a whole file, since C has no smaller unit than that to scope it to.

Compiler flags worth knowing exist

Every gcc invocation in this course has run with default settings. Two flags are worth recognizing when you see them in a real project’s build command: -Wall -Wextra turns on a much broader set of compiler warnings than the default (many real bugs – an unused variable, a suspicious implicit conversion – show up only as warnings, not errors, unless you ask for them); -Werror upgrades every warning to a hard error, refusing to produce a program until they’re all fixed, which is exactly how the exercises in this course’s very first lesson caught mistakes that would otherwise have just been warnings. And below CMake, on Linux and macOS, the tool actually doing the building is very often Make, reading a Makefile – CMake’s generated build files, when you ask it for them, usually are a Makefile. You won’t write one by hand in this course, but the name is worth recognizing: wherever you see make typed in a terminal, a Makefile – generated by CMake or written directly – is what it’s reading.

There’s no runnable exercise for CMake itself here – it manages files on disk, which doesn’t fit a single-file browser exercise. The exercises below stay in one file, but focus entirely on the declaration-before-use and one-definition-rule habits that make splitting a real project across files work correctly once you do it.

Try it
Output will appear here.

Exercises