Skip to content

checkpoint: embed the restore stub instead of exec'ing its build path - #270

Merged
congwang-mk merged 4 commits into
mainfrom
embed-restore-stub
Oct 2, 2026
Merged

congwang-mk merged 4 commits into
mainfrom
embed-restore-stub

Conversation

@congwang-mk

@congwang-mk congwang-mk commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Checkpoint restore only worked on the machine that built sandlock. build.rs compiled the restore stub into cargo's OUT_DIR and baked that absolute path into the crate (env!("RESTORE_STUB_PATH")), so release tarballs, PyPI wheels, and any distro package exec'd a path that does not exist on the user's host and failed with "restore-stub was not built". This blocks shipping sandlock as a distro package.

Fix

  • The stub is embedded with include_bytes! and handed to the child as a sealed memfd at fd 6 (STUB_FD, next to the existing CTRL/READY/GO fds).
  • The child runs it with execveat(fd, "", AT_EMPTY_PATH) through a new ChildEntry::ExecFd. No path is resolved, so the policy no longer gets a read grant on the stub, and restore does not depend on /proc inside the sandbox.
  • The memfd is created with MFD_EXEC (see below), so a vm.memfd_noexec=1 host does not seal it non-executable.
  • The stub closes fd 6 with the other control fds before reopening the checkpoint's fd table.
  • On arches without a restore engine, build.rs writes an empty placeholder so include_bytes! compiles; an empty stub reports restore as unavailable.

Chroot: exec by fd

The stub now reaches the child as a memfd, which has no path inside a chroot, and the chroot exec handler confined every exec to a path in the root. So execveat(fd, "", AT_EMPTY_PATH) failed with EACCES, both for restore and for any chroot workload calling fexecve on a memfd. Restore under chroot had only worked when the build directory happened to be visible inside the root.

The second commit makes the handler take the image from the guest's own fd:

  • An fd to a file inside the root goes through the path policy, as a path-based exec does.
  • Any other fd (a memfd, or a file inherited from the host) must be held readable. The guest could then copy the bytes into its own memfd anyway, so exec grants nothing new, and Landlock still gates EXECUTE on any real file.
  • The image is read through a fresh open file description, because the guest's shares its offset and the PT_INTERP scan seeks.

vm.memfd_noexec=1

The third commit adds memfd_create_exec, which asks for MFD_EXEC and retries with plain flags on kernels before 6.3 that reject it with EINVAL (their memfds are always executable). Both exec'd memfds use it: the restore stub, and the chroot exec handler's PT_INTERP-patched copy of dynamic binaries. The latter was created with flags 0, so on a vm.memfd_noexec=1 host every dynamic binary in a chroot failed to exec, on main too. At vm.memfd_noexec=2 memfd exec is forbidden outright, and restore and chroot dynamic exec still fail there. This could not be tested locally (changing the sysctl needs root).

Testing

  • test_restore_resumes_inside_a_chroot checkpoints and restores a program in a real rootfs; test_chroot_fexecve_runs_an_image_held_by_fd covers fexecve from an in-root file fd and from a memfd (new fexecve subcommand in rootfs-helper). Without the second commit, chroot restore and the memfd case fail with EACCES.
  • New unit test checks that the stub memfd is write-sealed and holds the embedded bytes; the synthetic-image stub tests now exec it through execveat like restore does.
  • Rust --lib (888), integration (519 + 13), FFI restore (3), and Python (430, excluding the guard tests) all pass locally.
  • With the build-tree out/restore-stub moved away, the Python restore tests still pass, which simulates an installed binary.

🤖 Generated with Claude Code

build.rs compiled the stub into cargo's OUT_DIR and baked that absolute
path into the crate, so restore only worked on the machine that built
sandlock. Release tarballs, PyPI wheels and any distro package exec'd a
path that does not exist on the user's host and failed with
"restore-stub was not built".

The stub is now embedded with include_bytes!, handed to the child as a
sealed memfd at fd 6, and exec'd with execveat(AT_EMPTY_PATH). No path
is resolved, so the policy no longer needs a read grant on the stub and
restore does not depend on /proc inside the sandbox. The memfd is
created with MFD_EXEC so a vm.memfd_noexec=1 host does not seal it
non-executable, and the stub closes fd 6 with the other control fds
before reopening the checkpoint's fd table.

Signed-off-by: Cong Wang <cwang@multikernel.io>
The chroot exec handler confined every exec to a path inside the root,
so execveat(fd, "", AT_EMPTY_PATH) of an fd with no such path failed
with EACCES. That broke fexecve of a memfd for any chroot workload, and
checkpoint restore under chroot, whose stub now arrives as a memfd.

An exec by fd now takes the image from the guest's own fd. When the fd
is a file inside the root, the path policy decides as before. Otherwise
the guest must hold it readable: it could then copy the bytes into a
memfd of its own anyway, so exec grants nothing new, and Landlock still
gates EXECUTE on any real file. The handler reads the image through a
fresh open file description, since the guest's shares its offset and
the PT_INTERP scan seeks.

Signed-off-by: Cong Wang <cwang@multikernel.io>
To make a dynamic binary load the image's own ld.so, the chroot exec
handler copies it into a memfd with PT_INTERP patched and execs that.
The memfd was created without MFD_EXEC, so on a host with
vm.memfd_noexec=1 the kernel sealed it non-executable and every
dynamic binary in a chroot failed to exec.

Both exec'd memfds, this one and the restore stub, now come from
memfd_create_exec, which asks for MFD_EXEC and falls back to plain
flags on kernels before 6.3 that reject it.

Signed-off-by: Cong Wang <cwang@multikernel.io>
The kernel opens a script's #! interpreter itself, against the host
root, so under chroot a script ran the host's interpreter or, with
Landlock granting only the image, failed to exec at all. Rewriting the
#! line to an injected fd would fix the lookup but hand the
interpreter /proc/self/fd/N as argv[0], which breaks busybox's sh and
Python's venv detection.

Exec a script through a small embedded trampoline instead. The kernel
runs it from a one-line launcher whose body names the interpreter, its
argument and the script; the trampoline then execs the interpreter by
path with the argv the kernel would have built, and that exec resolves
inside the image like any other. The handler parses the #! line as
fs/binfmt_script.c does and walks the interpreter chain first, so a
missing interpreter fails with ENOENT and a loop with ELOOP, as before.

Signed-off-by: Cong Wang <cwang@multikernel.io>
@congwang-mk
congwang-mk merged commit f6b84c8 into main Oct 2, 2026
17 checks passed
@congwang-mk
congwang-mk deleted the embed-restore-stub branch October 2, 2026 23:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant