Why this matters
Every exercise so far has been one file, compiled with one command. Real projects are dozens to thousands of files, built by people other than you, on machines other than yours. Two problems fall out of that immediately: how do separate files call each other’s functions, and how do you describe “compile all of this, in this order, with these settings” in a way that doesn’t depend on one person’s terminal history. The first is a language question (declarations vs. definitions – you already have the tools). The second is what build systems, and CMake specifically, are for.
Splitting one file into many is a declaration problem
The worked example marks where a real split would happen: the mean function is genuinely independent of main.
In a real project, mean and any other statistics functions would live in stats.c, with their signatures listed
in a header, stats.h:
// stats.h
double mean(const double *values, int count);
main.c would then #include "stats.h" and call mean without ever seeing its implementation. This is exactly
the declaration/definition split from Module 3’s first lesson, just spread across files instead of positioned
above/below each other in one file: the header supplies declarations so every file that includes it can compile
calls to mean, while exactly one .c file supplies the definition. This is also why headers almost never
contain function bodies for ordinary functions – if two .c files both included a header containing a full
definition, you’d get the “redefinition” error from that same lesson, just triggered by an #include instead of a
copy-pasted duplicate. A real stats.h would also be wrapped in the header guard from the last lesson
(#ifndef STATS_H / #define STATS_H / … / #endif) – this is precisely the scenario those three lines
protect against: stats.h ending up #included more than once while building the same file.
Compiling multiple files: still preprocess → compile → assemble → link
gcc main.c stats.c -o app runs the same four stages as before, just with two source files: main.c and
stats.c are each preprocessed, compiled, and assembled into their own object files (main.o, stats.o)
independently – neither file needs to see the other’s source, only declarations, to compile. Only at the link
stage do they need to actually meet: the linker matches main.o’s reference to mean against stats.o’s
definition of it, exactly like the undefined-reference lesson, except now “no definition found” would mean you
forgot to pass stats.c to the compiler at all, not that you forgot to write the function.
Why not just type that gcc command every time?
For two files, typing the command by hand is fine. For a real project it isn’t: you’d need to know every source file, in what order, with what flags, for every platform, every time anything changes – and re-type all of it correctly. A build system is a tool that reads a description of your project once and figures out the right compiler invocations for you, including which files actually need recompiling after a change.
CMake is the most common one for C and C++. Critically, CMake itself doesn’t compile anything – it’s a
generator. You describe your project in a CMakeLists.txt file, and CMake reads it and produces build files for
whatever’s actually available on the machine (Makefiles on Linux, an Xcode project on macOS, Visual Studio files on
Windows), which some other tool then executes to actually invoke the compiler.
cmake_minimum_required(VERSION 3.20)
project(stats_demo LANGUAGES C)
add_executable(stats_demo
src/main.c
src/stats.c
)
target_include_directories(stats_demo PRIVATE include)
Reading it top to bottom: cmake_minimum_required states the oldest CMake version this file assumes, so a too-old
CMake fails loudly instead of misbehaving. project(...) names the project and its language. add_executable
declares a build target – the name of a program to produce, and the list of source files that go into it
(each still compiled separately, then linked together, exactly as above). target_include_directories tells the
compiler where to look for headers #included with quotes, like "stats.h".
The normal workflow is two commands, run from the project’s root:
cmake -S . -B build
cmake --build build
The first configures: it reads CMakeLists.txt and generates build files into a build/ directory (kept
separate from your source on purpose – an “out-of-source build” – so generated files never get mixed up with
files you wrote, and you can delete build/ entirely to start clean). The second actually builds, invoking
whatever the first command generated. You only need to re-run the first command when CMakeLists.txt itself
changes; after that, cmake --build build alone re-compiles just the files that changed.
static at file scope: hiding a name from every other file
You’ve already seen static change a local variable’s lifetime (Module 3’s “stale state” exercise – making it
persist across calls instead of being recreated each time). At file scope – applied to a function or a global
variable, written outside any function – static means something related but distinct: it makes that name
invisible to every other translation unit, even though the file it’s defined in can still use it freely.
// stats.c
static double running_total = 0.0; // only visible inside stats.c
double mean(const double *values, int count) { ... } // visible to other files (no static)
Without static, every function and global variable you define has external linkage by default – visible and
linkable from any other .c file in the project, whether or not that’s what you intended. A helper function
that’s only ever meant to be called from within stats.c itself, never from main.c, should be marked static:
it documents “this is private to this file” and, as a real benefit, means another file can freely define its own
same-named function without the linker ever seeing a conflict. This is the closest C gets to the encapsulation you
may have seen in other languages as private – scoped to a whole file, since C has no smaller unit than that to
scope it to.
Compiler flags worth knowing exist
Every gcc invocation in this course has run with default settings. Two flags are worth recognizing when you see
them in a real project’s build command: -Wall -Wextra turns on a much broader set of compiler warnings than the
default (many real bugs – an unused variable, a suspicious implicit conversion – show up only as warnings, not
errors, unless you ask for them); -Werror upgrades every warning to a hard error, refusing to produce a program
until they’re all fixed, which is exactly how the exercises in this course’s very first lesson caught mistakes
that would otherwise have just been warnings. And below CMake, on Linux and macOS, the tool actually doing the
building is very often Make, reading a Makefile – CMake’s generated build files, when you ask it for them,
usually are a Makefile. You won’t write one by hand in this course, but the name is worth recognizing: wherever
you see make typed in a terminal, a Makefile – generated by CMake or written directly – is what it’s reading.
There’s no runnable exercise for CMake itself here – it manages files on disk, which doesn’t fit a single-file browser exercise. The exercises below stay in one file, but focus entirely on the declaration-before-use and one-definition-rule habits that make splitting a real project across files work correctly once you do it.