feat(fspy): route preload allocations through a lock-free global allocator - #599
feat(fspy): route preload allocations through a lock-free global allocator#599wan9chi wants to merge 1 commit into
Conversation
…cator Alternative to #596 for the same problem: the preload library runs inside libc calls that programs may make from a signal handler or from the child of fork() in a multithreaded process, where libc malloc's lock may be held by a thread that is paused or gone. Where #596 hands each intercepted call its own bump arena, and converts call sites one at a time, this installs one lock-free allocator as the preload cdylib's #[global_allocator]. Every Rust allocation in the library is covered at once, with no call-site changes: power-of-two size classes carve blocks out of 1 MiB mmap'd slabs, freed blocks recycle through per-class Treiber free lists made ABA-resistant by a 40-bit generation tag, and larger or over-aligned requests map directly. Includes the same access-relative benchmark suite as #596 so the two approaches can be compared on identical workloads. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
fspy benchmarklinuxmacoswindows |
Comparison with #596Both PRs now have benchmark runs against the same base (
Reading:
RecommendationLand #596, close this one. It is cheaper on every row that moved, one quarter of the code (~350 vs ~1400 lines), and its lifetimes are compiler-enforced rather than unrestricted. The one thing this PR does better — covering allocations nobody has converted yet — is exactly what the remaining migration work in #596 addresses incrementally, and the numbers say that blanket coverage is not worth what it costs. Worth keeping from this branch: it is the only measurement showing what full allocator replacement costs, which is useful if the remaining conversions ever stall. 🤖 Generated with Claude Code |
Part of #605 — this lands the malloc-class fix (the largest of the hazards there); the lazy-dlsym, hot-path panic, TLS reentrancy, and posix_spawn-thread items remain follow-ups. The benchmark suite that prices this change merged in #602. ## Motivation The preload library runs inside libc calls such as `open`, `stat`, and `execve`. Programs are allowed to make these calls from a signal handler, or in the child of `fork()` in a program with many threads. In both situations, using libc's `malloc` can hang the program forever: the lock inside `malloc` may be held by a thread that is paused or no longer exists. The preload library still allocates through `malloc` today, so a traced program can hang in exactly these situations. ## What this does Adds a new crate, `sigsafe`: Unix syscall wrappers that are safe to call where libc is not — in signal handlers, in fork children, before libc has finished initializing. Its [README](https://github.com/voidzero-dev/vite-task/blob/claude/fspy-libc-async-signal-safe-129134/crates/sigsafe/README.md) states the three rules everything in it follows: syscalls only (never through libc on Linux), no locks and no hidden state, and no global allocation. **The no-libc rule is enforced at compile time.** rustix can be built with a libc backend, and anything in the dependency graph — including crates outside this repository — can select it; no build script can detect the feature-unification case. So `sigsafe`'s `lib.rs` references `rustix::runtime`, a module that exists only in rustix's raw-syscall build: selecting the libc backend makes the crate fail to compile instead of silently losing the guarantee. On top of the first wrappers (`mm::mmap_anonymous`, `mm::munmap`, `param::page_size`) sits `sigsafe::alloc`, allocation that never touches malloc, in three layers with only the top exposed: - `MmapAllocator` — every allocation asks the kernel for fresh memory pages through `sigsafe::mm`. It keeps no state of its own, so there is nothing a signal or a `fork()` can catch locked or half-written. - `ChunkPool` — keeps up to 64 freed 64 KiB chunks in a fixed array of atomic pointers, so the next call can reuse memory without asking the kernel again. Taking or returning a chunk is one atomic swap per slot, never a lock, and a thread that disappears mid-operation can strand at most the one chunk it held. - `alloc::arena()` — the only public function. It hands one intercepted call its own bump arena (a `bump_scope::Bump`) that draws chunks from the pool and returns them when the call ends. Values allocated in the arena cannot outlive the call; the borrow checker enforces it. Uses the arena in one place to start: `RawExec::to_c_str_array`, which builds the NULL-terminated argv/envp pointer arrays that an intercepted exec hands to the real call, then drops them when it returns. That temporary's lifetime is already exactly a bump arena's, so the change is nine lines and adds no `unsafe`. It has to come off malloc because exec runs in the child of `fork()` in multithreaded programs — `posix_spawn` forks then execs — where malloc's lock may be held by a thread that no longer exists. Both platforms take this path on every intercepted exec. The strings the array points at are still owned by `Exec` and still come from malloc, as does the rest of the preload; converting them is follow-up. This change establishes the crate, the layers, and the lifetime discipline in the smallest place all three apply. ## Benchmark The `access-relative` suite (#602) was added while this PR still used the arena for the fd-relative join, and it earned its keep twice: an early run showed **+29%** on Linux, which turned out to be the preload building without optimizations (fixed by #597), and the corrected runs showed the arena join costing ~+3% over `PathBuf::push` — which is why the join reverted and the arena moved to `execveat`. The investigation is written up in [this comment](#596 (comment) thread. With the join reverted, both suites should sit at baseline. ## Commits The first commit is an earlier version of the allocator — one lock-free size-class allocator installed as the preload's `#[global_allocator]` — kept so the two designs can be compared; #599 measured that design end to end and lost. The second commit replaces it with the arena design above. The third moves the allocator into the new `sigsafe` crate as `sigsafe::alloc`, adds `mm`/`param` and the compile-time backend enforcement, and adds the README. The fourth moves the arena use from the join to `execveat` and fixes the dangling pointer there. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Alternative to #596. Both PRs solve the same problem; this one exists so CI can price the two approaches against each other on identical workloads. One of them should be closed once the numbers are in.
Motivation
The preload library runs inside libc calls such as
open,stat, andexecve. Programs are allowed to make these calls from a signal handler, or in the child offork()in a program with many threads. In both situations, using libc'smalloccan hang the program forever: the lock insidemallocmay be held by a thread that is paused or no longer exists.The two approaches
This PR installs one
#[global_allocator]in the preload cdylib: power-of-two size classes carve blocks out of 1 MiB mmap'd slabs, freed blocks recycle through per-class Treiber free lists made ABA-resistant by a 40-bit generation tag, and requests larger than the biggest class (or over-aligned) map and unmap directly. All memory comes from anonymous mappings; libc malloc is never called, no locks are taken, and no thread-local state is used, so a thread that vanishes atfork()or is suspended by a signal cannot strand another.The trade-off in one line: the global allocator converts the whole library immediately but pays its cost on every allocation, while the arena is cheaper per allocation but only covers code that has been rewritten to use it.
Comparing
This PR carries the same
access-relativebenchmark suite as #596, so both report the same rows against the same base.accessprices the absolute-path lane (a borrowed pointer, no directory resolution);access-relativeprices the lane that resolves the working directory and joins a path — where #596's arena is actually used.Read the two benchmark comments together: #596 changes one hot call site, this one changes the allocator underneath all of them.
🤖 Generated with Claude Code