Skip to content

The build

How libav* becomes a .wasm. The pipeline is four small scripts under build/: toolchain.sh (the shared cross env), deps.sh (the external codec libraries), libav.sh (FFmpeg's libav*), and driver.sh (the final link), orchestrated by a Dockerfile; this page explains the parts that aren't obvious.

The toolchain (build/toolchain.sh)

Everything is built with the wasi-sdk, an LLVM/clang toolchain targeting wasm32-wasip1 with a wasi-libc sysroot. The notable flags:

  • --target=wasm32-wasip1 --sysroot=…: cross-compile to WASI.
  • The WebAssembly feature set: -mtail-call -mbulk-memory -msimd128 -mextended-const -mnontrapping-fptoint -msign-ext -mmutable-globals -mreference-types. These are kept in lock-step with what the afmpeg runtime enables; a module built with a feature the runtime doesn't allow won't instantiate.
  • -Oz: optimise for size (the module is shipped over the wire).

Single-threaded by configuration

The reason this works at all (see Why libav-direct): the libraries are configured --disable-pthreads --disable-w32threads --disable-os2threads. The threading that blocks the FFmpeg CLI lives in fftools, which we don't build (--disable-programs). We also --disable-asm (no wasm assembly path) and --disable-network.

setjmp / longjmp

FFmpeg's C uses setjmp/longjmp (notably in some codecs' error handling, and in libvpx's encoder). clang lowers these for wasm with -mllvm=-wasm-enable-sjlj -mllvm=-wasm-use-legacy-eh=false, which emits two host imports — env.__wasm_setjmp and env.__wasm_longjmp, resolved at link time against the sysroot's libsetjmp.a. The runtime provides the imports; afmpeg implements them with wazero's snapshotter, and the bundled tools/run harness does the same. This is why a stock WASI runtime can't load the module but our setup can.

Why the module asks for an 8 MB stack

build/driver.sh links the wasm module with -Wl,-z,stack-size=8388608, lifting the data stack from wasi-sdk's 64 KB default to 8 MB. Two things need it: the engine's op_process holds a large fixed-size context on the stack (it is sized for the job caps), and FFmpeg's native encoders recurse deeply, the mpegvideo/mjpeg path most of all. At 64 KB the stack overflows into a wasm trap rather than a clean error, which is a miserable failure to diagnose. The native driver needs no equivalent: it uses the host's stack.

The wasi compatibility shims

WASI is a smaller world than POSIX, and FFmpeg assumes POSIX. wasi-libc deliberately omits a few functions ("WASI has no …"). We bridge the gap minimally:

  • config.h fixups: HAVE_SYSCTL, HAVE_MKSTEMP, HAVE_GETHRTIME, HAVE_SETRLIMIT are forced off so FFmpeg takes its portable fallbacks.
  • build/wasi-compat.h: force-included during the libav* build (via --extra-cflags, so it's baked into config.mak) to declare functions wasi-libc's headers gate out (e.g. dup, tempnam), keeping the strict-C99 clang from erroring.
  • build/wasi-compat.c: implements the symbols the link actually needs. For example WASI has no dup(2) and no way to allocate a fresh lowest-available fd, so it is a stub that fails with ENOSYS, enough to satisfy the link (libavformat's file.o references dup) while the code paths that would call it are never exercised, because the real file I/O goes over the mounted filesystem.

This shim layer is small and explicit, and it's exactly the kind of porting work that "owning a current FFmpeg build" means. It grows as new codecs/protocols pull in new corners of POSIX.

Not every POSIX gap is a link-time one. Some are runtime devices. libavutil's av_get_random_seed(), for instance, reads /dev/urandom (its only compiled entropy source here) and otherwise falls back to a clock()-jitter loop that never terminates under WASI, so any format needing a random id (the Matroska muxer seeds track UIDs this way) would hang. That gap is filled at the runtime layer, not here: the afmpeg host serves /dev/urandom from its vfs bridge. See afmpeg's vfs-bridge explanation.

The openh264 dependency

Both variants encode H.264 via openh264 (the GPL variant additionally offers libx264). openh264 is C++ with a GNU-make build that doesn't know about wasm, so build/deps.sh cross-compiles it with a few deliberate overrides: OS=linux (steers the Makefile only; the C preprocessor never sees __linux__ for wasm), ARCH=generic USE_ASM=No (portable C path), USE_STACK_PROTECTOR=No, and -fno-exceptions -fno-rtti. Three small wasm adaptations make it build and run, all in build/openh264-wasi.patch:

  • No <sys/sysctl.h> / SCHED_FIFO: wasip1 lacks both; the patch teaches WelsThreadLib the __wasi__ case (CPU count is simply 1).
  • Single-threaded pthread/sem shim (build/openh264-threads.c): wasip1 has no thread spawning. The encoder runs single-threaded (libav* is --disable-pthreads, so FFmpeg requests one thread), so the mutex/sem operations are no-op successes and pthread_create is never reached. The shim is archived into libopenh264.a so it satisfies both FFmpeg's configure probe and the final link.
  • A 2-argument ForceIntraFrame: openh264's C vtable (which FFmpeg calls through) declares ForceIntraFrame(self, bool), but its C++ method is (bool, int iLayerId = -1). On native ABIs the arity slip is harmless; wasm's strict indirect-call typing traps on it, so the patch drops the parameter (hardcoding the upstream -1 default).

openh264's C++ runtime is pulled in at the engine link with -lc++ -lc++abi. The codec's patent posture (self-compiled → outside Cisco's grant) is a licensing matter, not a build one.

The artifact

build/driver.sh links src/driver.c + the compat shims against the libav* archives into a single WASI command module. clang adds the _start/crt1 entry automatically; no wasm-ld-only flags are needed. The result is dist/ffmpeg-wasi-<variant>.wasm.

The native driver (spec 0028)

The same engine also compiles to a native ELF: one build system, two targets, selected by the TARGET env var. The wasm path above is TARGET=wasm (the default, byte-identical to before the split). TARGET=native (build/Dockerfile.native) drops the wasi machinery (no wasi-sdk target/sysroot, no SjLj lowering, no emulated libs, no single-thread shims) and instead uses the host clang/gcc with real threads + SIMD and asm enabled. build/toolchain.sh, build/deps.sh, build/libav.sh, and build/driver.sh each branch on TARGET; the deps are built native (openh264/x264 with real asm, and per profile the software batch, plus x265/SVT-AV1 for full), and driver.sh links a native ELF (-DAFMPEG_NATIVE, --start-group, no wasm shims).

The native driver serves its filesystem I/O over an IPC bridge instead of WASI syscalls: src/nativeio.c installs a custom seekable AVIOContext whose read/write/seek callbacks speak a framed protocol over a Unix socket, so the driver's I/O crosses the caller's afero.Fs (in afmpeg's pkg/afmpeg/native) exactly as the wasm build's WASI calls do, with no host disk. Even the concat demuxer's per-segment opens route over the bridge: build/ffmpeg-concat-ioopen.patch (a two-line FFmpeg patch applied in libav.sh) forwards the demuxer's io_open into its sub-contexts, so a concat join of afero-only segments never touches host disk either. This is "Backend B": threads + SIMD give ~50× (openh264) to ~170× (libx264) faster software encode, and the full profile adds the HEVC/AV1 encoders that are impractical in wasm. The native artifact is dist/driver, published as ffmpeg-wasi-driver-linux-amd64-[<profile>-]<variant> and signed alongside the wasm modules.

# The native full/gpl driver (HEVC via x265 + AV1 via SVT-AV1):
docker build -f build/Dockerfile.native --build-arg VARIANT=gpl --build-arg PROFILE=full \
  --target artifact -o dist-native .

Reproducibility

Every input is pinned: the wasi-sdk image in each Dockerfile, the upstream FFmpeg tag in build/ffmpeg-version.txt, and each codec library by tag, commit or SHA-256 digest in build/deps.sh. The tarball-fetched libraries are verified against a hard-coded digest and every mirror that disagrees is rejected, so an altered mirror cannot slip modified source into an artifact. build/versions.lock restates several of those pins in one place as a record. No script reads it, so it is a summary to keep in step rather than the mechanism. The build options reference lists every pin and where its authoritative value lives.

A release tag is <FFMPEG_VERSION>-<build-rev> (e.g. n9.0.1-1); the build revision bumps when the toolchain or config changes for the same upstream FFmpeg. The tag's version half must match build/ffmpeg-version.txt, which is checked before any build job starts, so a release cannot name an FFmpeg version that no merge request ever built (spec 0035).

What CI proves, and when

A merge request that touches build/** or src/** builds the lean pair for each target — four artifacts, about 14 minutes — which is what proves the engine still compiles and links. The intermediate and full profiles, 93% of the build time, build on a release tag, when a merge request changes something that can alter what a non-lean profile contains (the dependency builder, the component allowlist, the configure step, the pinned versions, the FFmpeg version, its patches, the toolchain, or the engine pipeline file), or on demand with FULL_MATRIX=1. Everything else (a docs edit, a dependency bump, a change to the release or signing jobs) runs only the fast checks.

The build matrix lives in build/engine.gitlab-ci.yml rather than in .gitlab-ci.yml, and that is what makes the sentence above true. changes: works on whole files, so while the ten jobs sat in the root pipeline file, every edit to it rebuilt everything: over the 21 days to 2026-09-03, 54% of this project's merge-request minutes went to merge requests that touched no engine file at all. A merge request that rewires the build still rebuilds, because the file it must edit is under build/.

Those builds share one project-scoped resource_group, so at most one engine compile runs at a time. A full matrix therefore takes roughly 95 minutes of wall clock rather than the ~17 it would take running wide. That is deliberate (spec 0035 D7): compiling FFmpeg ten times in parallel fills every slot on the shared runners and starves every other project in the fleet, and wall clock on an unattended pipeline is the cheaper thing to spend.

Signing and release stay tag-only. A merge request can build and measure; it can never sign or publish.

A size-budget gate (spec 0022) guards against accidental bloat: build/size-budget.txt sets a per-artifact byte ceiling, and the size-budget CI job (build/check-size-budget.sh) prints each artifact's size vs its budget, on merge requests as well as tags, since a size regression is worth catching in review, and flags an overage. It is advisory (allow_failure) until the ceilings are calibrated from real builds.