[4/4] xtensa/esp32s3: Run user processes in a world they cannot escape - #19798
Conversation
c2ff339 to
5add8b3
Compare
|
@tmedicci @igrr @xiaoxiang781216 what do you think? |
🔗 Cross-repo PR dependenciesThe read-only Build run reported the following dependent PR(s) and fetched head SHA(s): CI run: https://github.com/apache/nuttx/actions/runs/31531219719 |
Why not let reviewer continue on #19795, #19796 and #19797? it's better to split your huge change into the small and self-contained pr, otherwise any concern will block your whole change to merge. |
I agree to split, however, i'd like to keep a PR with all the changes so they can test the final intended result. I updated the PR body to not mention that this pr replaces the others |
Yes, no problem. you can change this pr to draft for your testing and restore other to ready for review, so the maintainer could know which pr to review. |
❌ Cross-repo dependency could not be appliedThe Build report says the declared dependency PR(s) could not be applied, so CI did not run against the combined code: Reason: cherry-pick failed (if your PR has merge commits, rebase instead) CI run: https://github.com/apache/nuttx/actions/runs/31835808454 |
5add8b3 to
80665b5
Compare
❌ Cross-repo dependency could not be appliedThe Build report says the declared dependency PR(s) could not be applied, so CI did not run against the combined code:
Reason: cherry-pick failed (if your PR has merge commits, rebase instead) CI run: https://github.com/apache/nuttx/actions/runs/33196805203 |
80665b5 to
36a12ac
Compare
❌ Cross-repo dependency could not be appliedThe Build report says the declared dependency PR(s) could not be applied, so CI did not run against the combined code:
Reason: cherry-pick failed (if your PR has merge commits, rebase instead) CI run: https://github.com/apache/nuttx/actions/runs/36277511578 |
36a12ac to
8ebafac
Compare
🔗 Cross-repo PR dependenciesThe read-only Build run reported the following dependent PR(s) and fetched head SHA(s): CI run: https://github.com/apache/nuttx/actions/runs/36392888270 |
Separate the world split from the protected user image, give WORLD1 its own vector table and its own PMS permissions -- including the PSRAM -- clean up the user cache-MMU windows, and stop keeping the page pool mapped. Folds in: xtensa/esp32s3: separate the world split from the protected user image xtensa/esp32s3: give the unprivileged world its own vector table xtensa/esp32s3: give the unprivileged world its permissions xtensa/esp32s3: clean up the user cache-MMU windows xtensa/esp32s3: stop keeping the page pool mapped xtensa/esp32s3: give the PSRAM its own PMS permissions Assisted-by: Claude Code:claude-opus-5-5 Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
8ebafac to
d07feb3
Compare
🔗 Cross-repo PR dependenciesThe read-only Build run reported the following dependent PR(s) and fetched head SHA(s): CI run: https://github.com/apache/nuttx/actions/runs/36436011900 |
When an unprivileged task takes a fault the system cannot recover from, it now gets a fatal SIGSEGV and only that task ends. A fault in privileged code still panics. What decides it is the interrupted context, not the cause: the saved PS says whether the fault was taken in User Mode. A list of causes would leave every cause off the list as a way for a user task to stop the machine, and there are many -- a divide by zero, a privileged instruction, a load/store error, and an illegal instruction, which is how a refused fetch from kernel text arrives on this chip (TRM v1.8 p.699: a denied external-memory access is answered with 0xdeadbeaf instead of trapping). PS.UM is clear in a kernel thread, in a system call made on the user's behalf and in an interrupt handler, so those still panic. If the recoverable-fault dispatcher is enabled it still gets first refusal on causes 28, 29 and 20, the only ones re-executing can help. esp32s3_userfault_abort() records the exception frame as the task's context, dispatches SIGSEGV, and returns the redirected frame, so the vector's RFE resumes the task in the signal trampoline, whose default action exits it. CONFIG_ESP32S3_USERFAULT_ABORT enables it, default y wherever there is an unprivileged world, and selects SIG_DEFAULT and SIG_SIGKILL_ACTION. Verified on an ESP32-S3 DevKitC with a WROOM-2 module, esp32s3-devkit:kernel_oct: a user task that writes through NULL, reads a wild address, divides by zero, calls into a buffer of garbage or branches into kernel text is terminated on its own, while an unrelated task keeps running. Stack overflow is not contained. On the windowed ABI it faults inside the window overflow handler and arrives as a double exception with PS.UM already clear; guard pages are the answer, and separate work. Assisted-by: Claude Code:claude-opus-5-5 Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Reporting a fault can itself fault. syslog reaches memory the fault being reported may have made unreachable, so esp32s3_pagefault_dispatch() is re-entered from inside its own _alert() and never returns, and the console fills with the same half-printed line forever. Found under Espressif's QEMU, where PSRAM never initialises and the kernel build needs it; the board's PSRAM works, so hardware does not take this path. A fault repeating at the same address and PC is not helped by reporting it again, so the dispatcher tries three times and then halts with interrupts off. esp32s3_userfault_abort() clears the count through esp32s3_pagefault_clear_repeat(): reaching it means the fault was contained, so only unbroken recursion stops the machine, and three probes at one address do not halt a healthy system. Verified under QEMU: 12,958,521 bytes of output in 60 s before, four reports and a halt after. On an ESP32-S3 DevKitC, esp32s3-devkit:kernel_oct, three identical sandbox probes in one boot are all contained. Assisted-by: Claude Code:claude-opus-5-5 Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
The PMS grants and refuses physical addresses, so it never sees an access that no MMU entry translates. The cache answered such an access with zeros and raised nothing, and the task carried on with a value it never should have had. Enable EXTMEM_MMU_ENTRY_FAULT and route the Cache Invalid Access interrupt to the handler that already serves the PMS monitors. An unprivileged task that makes the access is terminated with SIGSEGV; a privileged one still panics. The latch is level triggered, so it is cleared with the others. Read the cause before the clear, so the log tells the two apart: a PMS violation is a refused translation, an MMU entry fault is an access that was never translated. Give the kernel_oct configuration the addresses that examples/sandbox needs to name its targets. Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
Two files each defined the same 16 byte constant, KSTACK_ALIGNMENT and SIGTRAMP_STACK_ALIGN. Use STACKFRAME_ALIGN, which arch/xtensa/include/irq.h already gives as 16, with the STACKFRAME_ALIGN_DOWN() of nuttx/irq.h. STACK_ALIGNMENT is not the name to use here. It is TLS_STACK_ALIGN when CONFIG_TLS_ALIGNED is set, which is the alignment of a thread stack and not of a frame. Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Marco Casaroli <marco.casaroli@gmail.com>
d07feb3 to
a78c5d8
Compare
🔗 Cross-repo PR dependenciesThe read-only Build run reported the following dependent PR(s) and fetched head SHA(s): CI run: https://github.com/apache/nuttx/actions/runs/36475333009 |
depends-on: [apache/nuttx-apps/pull/3721]
Summary
Part 4 of 4. This is the part that actually contains a user process.
#19795, #19796 and #19797 are merged: the ESP32-S3 has a kernel build, per-process address environments and
fork(). They do not give protection. PMS was programmed only fromesp32s3_userspace.c, built underCONFIG_BUILD_PROTECTEDalone, so a kernel build never split the worlds, never entered World 1 and never installed the monitor interrupt. A user process could read and write kernel memory and reach the registers that control its own mapping.This moves the world split into
esp32s3_isolation.c, built for any build that is not flat, andesp32s3_start()programs the worlds and the permissions beforenx_start().It also reports an access that no MMU entry translates. PMS grants and refuses physical addresses, so it never saw one: the cache answered with zeros and the task carried on with a value it should never have had.
EXTMEM_MMU_ENTRY_FAULTis enabled and the Cache Invalid Access interrupt goes to the same handler, which reads the cause before clearing the latch.An unprivileged task that makes either kind of access is killed with SIGSEGV. A privileged one still panics.
The branch is rebased on master, so it carries only its own five commits.
Impact
A flat build does not change. A protected build keeps the same permissions; only the file that sets them moved.
A kernel build on the ESP32-S3 becomes a security boundary, which it was not before.
Testing
ESP32-S3-DevKitC, WROOM-2 N32R8V,
esp32s3-devkit:kernel_oct.ostestruns to the end.examples/sandboxcarries the expected outcome per target, so it fails a build that refuses everything as well as one that permits everything.selfis the control:0x600c5000isDR_REG_MMU_TABLE, the registers that hold the mapping.A killed process gives back what it held. It allocates 64 KiB and opens files before faulting; memory goes 1441792 -> 2162688 -> 1441792 and no descriptor is left open. The test fails if the number never rises, because one that does not move proves nothing.
Three runs, twelve process deaths, same result each time.
tools/checkpatch.sh -c -u -m -gpasses.