The build¶
How libav* becomes a .wasm. The pipeline is four small scripts under build/:
toolchain.sh (the shared cross env), deps.sh (the external codec libraries), libav.sh
(FFmpeg's libav*), and driver.sh (the final link), orchestrated by a Dockerfile; this page
explains the parts that aren't obvious.
The toolchain (build/toolchain.sh)¶
Everything is built with the wasi-sdk, an
LLVM/clang toolchain targeting wasm32-wasip1 with a wasi-libc sysroot. The notable flags:
--target=wasm32-wasip1 --sysroot=…: cross-compile to WASI.- The WebAssembly feature set:
-mtail-call -mbulk-memory -msimd128 -mextended-const -mnontrapping-fptoint -msign-ext -mmutable-globals -mreference-types. These are kept in lock-step with what the afmpeg runtime enables; a module built with a feature the runtime doesn't allow won't instantiate. -Oz: optimise for size (the module is shipped over the wire).
Single-threaded by configuration¶
The reason this works at all (see Why libav-direct): the libraries
are configured --disable-pthreads --disable-w32threads --disable-os2threads. The threading
that blocks the FFmpeg CLI lives in fftools, which we don't build (--disable-programs).
We also --disable-asm (no wasm assembly path) and --disable-network.
setjmp / longjmp¶
FFmpeg's C uses setjmp/longjmp (notably in some codecs' error handling, and in libvpx's
encoder). clang lowers these for wasm with -mllvm=-wasm-enable-sjlj
-mllvm=-wasm-use-legacy-eh=false, which emits two host imports — env.__wasm_setjmp and
env.__wasm_longjmp, resolved at link time against the sysroot's libsetjmp.a. The runtime
provides the imports; afmpeg implements them with wazero's snapshotter, and the bundled tools/run
harness does the same. This is why a stock WASI runtime can't load the module but our setup can.
Why the module asks for an 8 MB stack¶
build/driver.sh links the wasm module with -Wl,-z,stack-size=8388608, lifting the data stack
from wasi-sdk's 64 KB default to 8 MB. Two things need it: the engine's op_process holds a
large fixed-size context on the stack (it is sized for the
job caps), and
FFmpeg's native encoders recurse deeply, the mpegvideo/mjpeg path most of all. At 64 KB the stack
overflows into a wasm trap rather than a clean error, which is a miserable failure to diagnose. The
native driver needs no equivalent: it uses the host's stack.
The wasi compatibility shims¶
WASI is a smaller world than POSIX, and FFmpeg assumes POSIX. wasi-libc deliberately omits a few functions ("WASI has no …"). We bridge the gap minimally:
config.hfixups:HAVE_SYSCTL,HAVE_MKSTEMP,HAVE_GETHRTIME,HAVE_SETRLIMITare forced off so FFmpeg takes its portable fallbacks.build/wasi-compat.h: force-included during the libav* build (via--extra-cflags, so it's baked intoconfig.mak) to declare functions wasi-libc's headers gate out (e.g.dup,tempnam), keeping the strict-C99 clang from erroring.build/wasi-compat.c: implements the symbols the link actually needs. For example WASI has nodup(2)and no way to allocate a fresh lowest-available fd, so it is a stub that fails withENOSYS, enough to satisfy the link (libavformat'sfile.oreferencesdup) while the code paths that would call it are never exercised, because the real file I/O goes over the mounted filesystem.
This shim layer is small and explicit, and it's exactly the kind of porting work that "owning a current FFmpeg build" means. It grows as new codecs/protocols pull in new corners of POSIX.
Not every POSIX gap is a link-time one. Some are runtime devices. libavutil's
av_get_random_seed(), for instance, reads /dev/urandom (its only compiled entropy source
here) and otherwise falls back to a clock()-jitter loop that never terminates under WASI,
so any format needing a random id (the Matroska muxer seeds track UIDs this way) would hang.
That gap is filled at the runtime layer, not here: the afmpeg host serves /dev/urandom
from its vfs bridge. See afmpeg's
vfs-bridge explanation.
The openh264 dependency¶
Both variants encode H.264 via openh264 (the GPL variant
additionally offers libx264). openh264 is C++ with a GNU-make build that doesn't know about wasm,
so build/deps.sh cross-compiles it with a few deliberate overrides: OS=linux (steers the
Makefile only; the C preprocessor never sees __linux__ for wasm), ARCH=generic USE_ASM=No
(portable C path), USE_STACK_PROTECTOR=No, and -fno-exceptions -fno-rtti. Three small wasm
adaptations make it build and run, all in build/openh264-wasi.patch:
- No
<sys/sysctl.h>/SCHED_FIFO: wasip1 lacks both; the patch teachesWelsThreadLibthe__wasi__case (CPU count is simply 1). - Single-threaded pthread/sem shim (
build/openh264-threads.c): wasip1 has no thread spawning. The encoder runs single-threaded (libav* is--disable-pthreads, so FFmpeg requests one thread), so the mutex/sem operations are no-op successes andpthread_createis never reached. The shim is archived intolibopenh264.aso it satisfies both FFmpeg's configure probe and the final link. - A 2-argument
ForceIntraFrame: openh264's C vtable (which FFmpeg calls through) declaresForceIntraFrame(self, bool), but its C++ method is(bool, int iLayerId = -1). On native ABIs the arity slip is harmless; wasm's strict indirect-call typing traps on it, so the patch drops the parameter (hardcoding the upstream-1default).
openh264's C++ runtime is pulled in at the engine link with -lc++ -lc++abi. The codec's
patent posture (self-compiled → outside Cisco's grant) is a licensing
matter, not a build one.
The artifact¶
build/driver.sh links src/driver.c + the compat shims against the libav* archives into
a single WASI command module. clang adds the _start/crt1 entry automatically; no
wasm-ld-only flags are needed. The result is dist/ffmpeg-wasi-<variant>.wasm.
The native driver (spec 0028)¶
The same engine also compiles to a native ELF: one build system, two targets, selected
by the TARGET env var. The wasm path above is TARGET=wasm (the default, byte-identical to
before the split). TARGET=native (build/Dockerfile.native) drops the wasi machinery (no
wasi-sdk target/sysroot, no SjLj lowering, no emulated libs, no single-thread shims) and instead
uses the host clang/gcc with real threads + SIMD and asm enabled. build/toolchain.sh,
build/deps.sh, build/libav.sh, and build/driver.sh each branch on TARGET; the deps are
built native (openh264/x264 with real asm, and per profile the software batch, plus x265/SVT-AV1
for full), and driver.sh links a native ELF (-DAFMPEG_NATIVE, --start-group, no wasm shims).
The native driver serves its filesystem I/O over an IPC bridge instead of WASI syscalls:
src/nativeio.c installs a custom seekable AVIOContext whose read/write/seek callbacks speak a
framed protocol over a Unix socket, so the driver's I/O crosses the caller's afero.Fs (in
afmpeg's pkg/afmpeg/native) exactly as the wasm build's WASI calls do, with no host disk. Even the
concat demuxer's per-segment opens route over the bridge: build/ffmpeg-concat-ioopen.patch (a
two-line FFmpeg patch applied in libav.sh) forwards the demuxer's io_open into its sub-contexts,
so a concat join of afero-only segments never touches host disk either. This is
"Backend B": threads + SIMD give
~50× (openh264) to ~170× (libx264) faster software encode, and the full profile adds the
HEVC/AV1 encoders that are
impractical in wasm. The native artifact is dist/driver, published as
ffmpeg-wasi-driver-linux-amd64-[<profile>-]<variant> and signed alongside the wasm modules.
# The native full/gpl driver (HEVC via x265 + AV1 via SVT-AV1):
docker build -f build/Dockerfile.native --build-arg VARIANT=gpl --build-arg PROFILE=full \
--target artifact -o dist-native .
Reproducibility¶
Every input is pinned: the wasi-sdk image in each Dockerfile, the upstream FFmpeg tag in
build/ffmpeg-version.txt, and each codec library by tag, commit or SHA-256 digest in
build/deps.sh. The tarball-fetched libraries are
verified against a hard-coded digest and every mirror that disagrees is rejected, so an altered
mirror cannot slip modified source into an artifact. build/versions.lock restates several of those
pins in one place as a record. No script reads it, so it is a summary to keep in step rather
than the mechanism. The build options reference lists every pin and
where its authoritative value lives.
A release tag is <FFMPEG_VERSION>-<build-rev> (e.g. n9.0.1-1); the build revision bumps when
the toolchain or config changes for the same upstream FFmpeg. The tag's version half must match
build/ffmpeg-version.txt, which is checked before any build job starts, so a release cannot name an
FFmpeg version that no merge request ever built (spec 0035).
What CI proves, and when¶
A merge request that touches build/** or src/** builds the lean pair for each target — four
artifacts, about 14 minutes — which is what proves the engine still compiles and links. The
intermediate and full profiles, 93% of the build time, build on a release tag, when a merge request
changes something that can alter what a non-lean profile contains (the dependency builder, the
component allowlist, the configure step, the pinned versions, the FFmpeg version, its patches, the
toolchain, or the engine pipeline file), or on demand with FULL_MATRIX=1. Everything
else (a docs edit, a dependency bump, a change to the release or signing jobs) runs only the fast
checks.
The build matrix lives in build/engine.gitlab-ci.yml rather than in .gitlab-ci.yml, and that is
what makes the sentence above true. changes: works on whole files, so while the ten jobs sat in
the root pipeline file, every edit to it rebuilt everything: over the 21 days to 2026-09-03, 54% of
this project's merge-request minutes went to merge requests that touched no engine file at all. A
merge request that rewires the build still rebuilds, because the file it must edit is under
build/.
Those builds share one project-scoped resource_group, so at most one engine compile runs at a
time. A full matrix therefore takes roughly 95 minutes of wall clock rather than the ~17 it would
take running wide. That is deliberate (spec 0035 D7): compiling FFmpeg ten times in parallel fills
every slot on the shared runners and starves every other project in the fleet, and wall clock on an
unattended pipeline is the cheaper thing to spend.
Signing and release stay tag-only. A merge request can build and measure; it can never sign or publish.
A size-budget gate (spec 0022) guards against accidental bloat: build/size-budget.txt sets a
per-artifact byte ceiling, and the size-budget CI job (build/check-size-budget.sh) prints each
artifact's size vs its budget, on merge requests as well as tags, since a size regression is worth
catching in review, and flags an overage. It is advisory (allow_failure) until the ceilings are
calibrated from real builds.