diff --git a/crates/rusty_alloc/CHANGELOG.md b/crates/rusty_alloc/CHANGELOG.md index d05ae3d..268250d 100644 --- a/crates/rusty_alloc/CHANGELOG.md +++ b/crates/rusty_alloc/CHANGELOG.md @@ -7,6 +7,32 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +### Fixed + +- **A failed re-commit no longer hands out unbacked memory (Windows, purging + on).** With `purge_delay >= 0` and `purge_decommits` on, a freed span is + decommitted, and reuse re-commits it — but the result of that commit was + discarded. Windows has no overcommit, so `MEM_COMMIT` fails once the + system's commit is exhausted, and the span was carved anyway: the first + store into it was an access violation in `page_extend`. One was seen in a + shipping consumer (`secure` on, `purge_delay = 0`, the machine at 98.5 % of + its commit limit). A failed re-commit is now a failed span: it stays on the + free list still marked purged, the allocator moves on, and true exhaustion + ends as an ordinary allocation failure. The default configuration + (`purge_delay = -1`) never purges and was not exposed. Details: + `docs/plans/recommit-failure-ignored.md`. +- **The same for whole segments.** `segment_free` and `huge_free` restore a + purged or guarded segment's commit and access before recycling it, and + dropped both results. A segment the OS will not restore is no longer + recycled: it is released to the OS, or, inside an arena, retired in place. + +### Added + +- `stats::commit_failures()`: a process-wide count of commits refused on + those paths, also printed by `stats::print_process` as `failed commits N`. + Non-zero after a fault in the page layer says the machine ran out of + commit. + ## [2.2.2](https://github.com/Remade-With-Rust/rusty_alloc/compare/rusty_alloc-v2.2.1...rusty_alloc-v2.2.2) - 2026-10-01 ### Fixed diff --git a/crates/rusty_alloc/UNSAFE.md b/crates/rusty_alloc/UNSAFE.md index 81350a7..a99fbad 100644 --- a/crates/rusty_alloc/UNSAFE.md +++ b/crates/rusty_alloc/UNSAFE.md @@ -23,7 +23,7 @@ unsafe in DEPENDENCIES, and ours are `libc` plus bindings-only `windows-sys`. | `alloc.rs` | 94 | The public entry points: raw-pointer reads on the malloc fast path (must never form `&mut` on the shared empty-heap sentinel), pointer-derived metadata on the free path (`segment_of`/`page_of`), block-content copies in the realloc family. **+15 on 2026-08-22**: the free path's two `asm!` sites — the memory-destination `used--` whose flags drive the retire branch (its `label` block is a separate item and carries its own `unsafe`), and the fused `cmp {tid}, fs:0` in both `free_inline` and `free_general` — plus `malloc_or`/`malloc_or_slow`, which give `operator new` a fast path whose miss is a tail call. Each asm reads or writes exactly one field it already had a valid pointer to; none widens what the surrounding code could already touch. **+2 on 2026-09-24 (instruction-count campaign, LEDGER):** `realloc` split into an inline null test and `unsafe fn realloc_live(p: NonNull, ..)` — one signature, and the `NonNull::new_unchecked` sits inside the block that already called it, on a pointer the line above proved non-null. The body is the old one unchanged; nothing new is dereferenced. **+10 on 2026-09-24 (round two, LEDGER):** (a) `realloc_move`, the moving arm shared by the growth and the shrink — one `unsafe fn`, its body block, and the two call-site blocks: the old inline move code, now reached with the copy length each arm already holds; (b) `zalloc` in `malloc`'s raw-read shape — its fast-path block reads through `hb`, which may be the sentinel, exactly as `malloc` does (no `&mut`, no write unless a block was popped, which proves the heap is real), and the cold `unsafe fn zalloc_slow` that performs the sentinel test the fast path dropped (signature + block); (c) `free_local_owned` / `free_local_no_xheap`, `owner_heap` split so its heap-creating fallback is a tail call — two signatures and their two blocks, the same loads and the same `free_local_at` as before; (d) `malloc_aligned_pow2`, an `unsafe fn` only because its caller must have proven the alignment a power of two (a violation is not memory-unsafe — the bound test still refuses a mask at or above half a segment). Nothing new is dereferenced anywhere; each carries its SAFETY line | 2026-09-24 | | `heap.rs` | 77 | Owner-thread page/queue manipulation under raw pointers (no two `&mut Page` may coexist), the aligned-allocation peek, span carving. **+3 on 2026-08-19**: `try_unlink_huge_segment` split out of `remove_huge_segment`. **+15 on 2026-08-22**: `malloc_generic` split into a small entry plus `malloc_generic_walk`, with `grow_front`, `try_guarded` and `drain_delayed` as cold arms — each split adds an `unsafe fn` signature and its block while dereferencing nothing the single function did not — and the immortal `EMPTY_DELAYED` sentinel that lets the heartbeat read its list without a null test. **+2 on 2026-09-07 (P4e, §2.15 of `docs/plans/small-metal.md`)**: the reclamation fixes — one `unsafe` for the periodic `generic_collect` sweep and one for the reclaim-and-retry that runs before `malloc_generic` returns null. Both call `collect_inner`, which allocates nothing and touches only this heap’s own pages on the owner thread; each carries its SAFETY line. `malloc_generic` itself became a safe wrapper over the renamed `malloc_generic_once`, so the split added no signature. **+2 on 2026-09-08:** the medium collect-and-retry ahead of the heartbeat -- one `unsafe` around `page_collect` + `page_pop` on the bin queue front, one reading `free_is_zero` off the page just popped. Both are the SAME operations `malloc_generic_walk` performs on the same page a few lines later, on the owner thread under the heap lock: the block moves earlier, nothing new is dereferenced. Each carries its SAFETY line. **+2 on 2026-09-16 (Openheimer):** the two refuse arms in `malloc_aligned_at_slow` when `bins::aligned_at_from` reports that `block + offset + align` would wrap for a caller-chosen offset — each is one `free_local` of the block just allocated on this thread, which has not escaped. **+1 on 2026-09-24 (instruction-count campaign, LEDGER):** the medium collect-and-retry's `page_collect` + `page_pop` block became two blocks so the `saw_remote_free` latch is stored between them — same two calls on the same page, owner thread, in the same order; nothing new is dereferenced. **+1 on 2026-09-24 (round two):** a medium miss grows the front page directly — one block around the `capacity < reserved` test and the `grow_front` call, on the same page the block above just collected, which is `grow_front`'s whole contract **+2 on 2026-09-25 (CURIOSITY ROUND FOUR):** the medium arm pops the front page BEFORE collecting it — one `page_pop` block and the `free_is_zero` read beside it, the same two operations the collect-then-pop below already performs on the same owner-thread page | 2026-09-25 | | `init.rs` | 42 | Thread/heap lifecycle: the initial-exec TLS slot (`global_asm!` + fs-relative asm reads), thread-pointer register reads (`fs:0`/`gs:0x30`/`tpidrro_el0`), heap-box creation/teardown, the abandonment path run inside platform TLS destructors. **+1 on 2026-09-09 (`firmware-what-is-left.md` §7.3): an ATTRIBUTE, not an operation** — `unsafe(link_section = ".rodata.…")` on `EMPTY_HEAP_BOX`, bare metal only, so the never-written sentinel lives in flash instead of costing 1,752 bytes of RAM. The token is counted because the census is a text search; the contract it rests on (raw reads only, no page ever stores its address) is the one the `Sync` impl below already carries, and a write would now fault against the flash cache instead of silently landing. `create_heap`'s template copy from it is a `ptr::read` inside the block that already existed. **+5 on 2026-09-16 (Openheimer, OH-rusty_alloc-13):** thread exit now abandons EVERY heap the dying thread still owns, not only the one in `done_slot` — a first-class heap that was never `heap_delete`d, or one installed with `set_default_heap`, used to stay DELAYED under a dead owner forever. `take_heap_owned_by` unlinks one such box under `HEAPS_LOCK` (one block: the registry walk); `thread_done` reads `owner_tid` off its own box and calls the split-out `unsafe fn thread_done_one` for the primary box and for each extra one (three sites plus the signature). Nothing new is dereferenced that the single function did not; each carries its SAFETY line | 2026-09-16 | -| `segment.rs` | 35 | Segment/page metadata addressing: the mask trick (`segment_of`), `page_of`'s contract-based indexing — **the bound is now PROVED for every in-segment offset by `proofs.rs` (Kani), not merely asserted** — span tiling, purge/recommit | 2026-08-19 | +| `segment.rs` | 43 | Segment/page metadata addressing: the mask trick (`segment_of`), `page_of`'s contract-based indexing — **the bound is now PROVED for every in-segment offset by `proofs.rs` (Kani), not merely asserted** — span tiling, purge/recommit. The baseline had drifted to 36 without a row edit; recorded now. **+2 on 2026-10-04 (`docs/plans/recommit-failure-ignored.md`):** `restore_for_reuse` — an `unsafe fn` and its block — is the guard-lift and re-commit `segment_free` and `huge_free` each did inline, moved into one function so both can act on its answer: a segment whose memory the OS will not make accessible and committed again is released whole or retired in place, never recycled. It touches the same range the inline code did; `segment_free`'s own block became the expression that calls it (no net change). `span_recommit` now returns whether the span is backed (same block, same call). **+5 the same day, all `#[cfg(test)]`:** `a_failed_recommit_is_a_failed_span_not_an_unbacked_one` writes and frees the blocks it allocated (four blocks) and frees them at the end (one); each carries its SAFETY line | 2026-10-04 | | `prim/windows.rs` | 34 | OS FFI: VirtualAlloc family, FLS destructors, QPC, BCryptGenRandom. **+3 on 2026-09-16 (Openheimer):** `range_is_reserved` (`zeroed` `MEMORY_BASIC_INFORMATION` + `VirtualQuery`, the same OS answer `unix.rs` gets from `mincore`), and `alloc_aligned` now refuses a garbage alignment and a `size + align` that wraps BEFORE it reserves anything, and releases a re-reservation that the OS placed anywhere but the aligned address it asked for instead of returning it misaligned (`VirtualFree`, one block). **+1 on 2026-09-24 (`docs/plans/youslowbro.md` §4):** `getenv` — one `GetEnvironmentVariableA` into a caller buffer of exactly the length passed, the allocation-free environment read that replaced `std::env::var`'s owned strings in the options pass (251 allocations through the global allocator on every process's first allocation, now none) | 2026-09-24 | | `page.rs` | 38 | Free-list links written into dead blocks, the lock-free `xthread_free` four-state protocol (loom-modeled in `tests/loom_xthread.rs`), the immortal `EMPTY_PAGE` sentinel — **+1 on 2026-09-09: the same `unsafe(link_section)` attribute as `init.rs`, placing it in flash on bare metal; its free list is permanently null, so nothing writes it**. **+8 on 2026-08-22**: `page_link_local` split out of `page_push_local` so the caller can decrement `used` in one memory-destination RMW, `page_collect_impl` const-generic over whether it also writes the protocol flag, and `USED_OFFSET` — whose value is asserted against `offset_of!(Page, used)` by a unit test, because an asm operand is not type-checked and a field reordering would silently decrement the wrong bytes | 2026-08-22 (free campaign) | | `prim/unix.rs` | 33 | OS FFI: mmap family, madvise/decommit, pthread keys, /dev/urandom. **+1 on 2026-09-16 (Openheimer, OH-201):** `range_is_reserved` asks `mincore` whether a caller-supplied `manage_os_memory` range is mapped at all before it becomes an arena — one FFI call on a page-aligned probe. **+1 on 2026-09-24 (`docs/plans/youslowbro.md` §4):** `getenv` — `libc::getenv` and a byte-by-byte read of the returned C string up to its NUL or the caller buffer's end, nothing written through it; the same call upstream's prim makes, with the same standing caveat about a concurrent `setenv` **+2 on 2026-09-25:** `env_for_each` — an `unsafe extern "C"` declaration of `environ` and one block that walks it to its NULL terminator, handing each NUL-terminated entry on as a pointer; read only, with the same concurrent-`setenv` caveat as `getenv`. **+2 on 2026-10-01 (macOS):** the Apple arm of `range_is_reserved` — an `unsafe extern "C"` declaration of `mach_vm_region` and the `mach_task_self_` global, and one call per region: a read-only query of our own task into a `packed(4)` `vm_region_basic_info_64` whose 36-byte size is asserted at compile time and is exactly the word count passed, because XNU's `mincore` succeeds on unmapped ranges | 2026-10-01 | @@ -31,7 +31,7 @@ unsafe in DEPENDENCIES, and ours are `libc` plus bindings-only `windows-sys`. | `rusty_alloc_api/src/lib.rs` | 16 | The `GlobalAlloc`/`Allocator` impls forwarding layouts to the core; `unsafe` is inherent to those traits' contracts . **+1 on 2026-09-24:** `GlobalAlloc::alloc`'s over-aligned arm calls `alloc::malloc_aligned_pow2`, an `unsafe fn` whose one added precondition — a power-of-two alignment — `Layout` guarantees **+1 on 2026-09-25:** the over-aligned `realloc` arm reads `usable_size(ptr)` to keep a block in place — `ptr` is a live non-null block of ours by the `GlobalAlloc` contract | 2026-09-25 | | `os.rs` | 12 | The prim-layer wrapper: commit/decommit/protect plumbing | 2026-08-08 | | `prim/mock.rs` | 8 | Miri-only mock OS backend (never shipped; `cfg(miri)`) | 2026-08-06 | -| `arena.rs` | 15 | Lock-free chunk bitmap claim/verify, recycled-chunk scrubbing (the 0.1.0-alpha.2 UAF fix lives here: `wait_no_remote_in_flight` on every recycle path). The baseline had drifted to 13 without a row edit; recorded now. **+2 on 2026-09-16 (Openheimer, `docs/plans/openheimer-run.md`):** `arena_register` claims its slot with a CAS instead of a load-then-store, and the two new blocks are the REFUSE arm — `os::free` of the descriptor and of the OS mapping when the table is full or the slot was lost — so a refused arena leaks nothing. Both free memory this function allocated a few lines earlier and nothing else has seen; each carries its SAFETY line | 2026-09-16 | +| `arena.rs` | 16 | Lock-free chunk bitmap claim/verify, recycled-chunk scrubbing (the 0.1.0-alpha.2 UAF fix lives here: `wait_no_remote_in_flight` on every recycle path). The baseline had drifted to 13 without a row edit; recorded now. **+2 on 2026-09-16 (Openheimer, `docs/plans/openheimer-run.md`):** `arena_register` claims its slot with a CAS instead of a load-then-store, and the two new blocks are the REFUSE arm — `os::free` of the descriptor and of the OS mapping when the table is full or the slot was lost — so a refused arena leaks nothing. Both free memory this function allocated a few lines earlier and nothing else has seen; each carries its SAFETY line. **+1 on 2026-10-04 (`docs/plans/recommit-failure-ignored.md`):** `owns` — the range scan `chunk_free` already performs, asked without clearing the used bit, so the release path can tell an arena chunk (retire in place) from a reservation of its own (release to the OS). It reads each live descriptor's `base` and `chunks_live` and nothing else | 2026-10-04 | | `prim/fixed.rs` | 37 | **New 2026-09-07 (P1 of `docs/plans/small-metal.md`).** The fixed-region backend for a target with no OS: memory is a `&'static mut [u8]` handed over once. **6 of the 18 are `unsafe fn` signatures the prim seam requires** (`alloc`/`free`/`commit`/`decommit`/`reset`/`protect`) whose *bodies contain no unsafe operation at all* — the free list is `AtomicUsize` arrays under a spin lock and the pointers are built with the safe `with_exposed_provenance_mut`, so this backend adds **zero** unsafe dereferences to the shipped crate. The other 12 are in `#[cfg(test)]`: two `&raw mut` static-region handoffs and ten calls through the `unsafe fn` seam, each with its SAFETY line. **+1 on 2026-09-07 (P2):** the region test became two-sided — where the shipped geometry refuses a segment-sized request from a 512 KiB region, the small profile SERVES one, so the test now frees it too. Audited at the site; the module is `allow(dead_code)` and unreachable on every platform that has an arm above it. **+7 on 2026-09-07 (P4b, §2.9/§2.10 of the same plan):** the two-ended `place` rule added ZERO unsafe to shipped code — `place` is a pure arithmetic `fn` and the scan around it is unchanged — and all seven are in `#[cfg(test)]`: one `slice::from_raw_parts_mut` carving the `REGION_ALIGN`-aligned window out of `BACKING` (replacing a `&mut *ptr` that a `repr(align(65536))` static would have needed, which rustc 1.97.1 on MSVC cannot compile), one `ptr::add` to reach that window, and five calls through the `unsafe fn` seam in the placement assertions and in `greedy_segments`, which allocates segments until refusal and frees every one before returning. Each carries its SAFETY line **+1 on 2026-09-09 (morning, `#21`):** the two-sided alignment test frees the SEGMENT_SIZE-aligned page it is served when the region straddles a boundary — `#[cfg(test)]`, banked without a row here at the time; recorded now. **+5 on 2026-09-09 (`docs/plans/finished/region-alignment-bug.md` §7): the module ships its first unsafe OPERATIONS.** (1) `unsafe impl Sync for Region` and (2) the `&mut *self.bytes.get()` in `Region::give` — the once-only handoff the Janus seam has carried since 2.0.0, moved here so the alignment travels with it; a module-wide `REGION_GIVEN` swap precedes the `&mut`, so a second `give` on any instance is refused before it could alias. (3) `unsafe impl Sync for FirstHeapBox`, the bare-metal static holding the first heap's descriptor (handed out once by `take_first_heap_box`, on the one thread such a build has). The other two are `#[cfg(test)]`: a `from_raw_parts_mut` carving the report's misaligned base out of a static for the refusal probe, and a `Box::new_zeroed().assume_init()` allocating a `Region` in place, because materialising a 64 KiB-aligned value on the stack first faults on Windows. Each carries its SAFETY line. **+2 on 2026-09-09 (`docs/plans/finished/region-alignment-dissolve.md`): zero new unsafe in shipped code** — segments now stride from the region's base, and every piece of that is safe: `stride_base` is a relaxed load, `place` takes an `origin` and does arithmetic, `install_region` aligns the base up. The two are `#[cfg(test)]`, the `--cfg ra_aligned_region` arm of the two-sided alignment test (an `alloc` and a `free` through the `unsafe fn` seam: the 2.0.4 straddle case, kept under the knob that restores the mask), and the `Box::new_zeroed().assume_init()` moved under that same cfg — the default arm is a plain `static Region` now, which a 16-byte-aligned type can be on every host toolchain. Each carries its SAFETY line. **+4 on 2026-09-09 (`docs/plans/finished/esp32-large-alloc-ceiling.md`): all four are `#[cfg(test)]`, and shipped code gained none.** The large-allocation ceiling reported from `rusty_zstd` is reproduced against the real extent allocator by `greedy_dedicated`, which serves `huge_alloc`-shaped reservations until the region refuses and frees every one before returning (two calls through the `unsafe fn` seam), plus one `alloc`/`free` pair modelling the whole segment a first small allocation claims. The fix itself -- the `ra_segment_size` geometry knob, `dedicated_segments`, `region_for_allocs` and `region_capacity` -- is arithmetic over the existing atomics and adds no unsafe operation at all. Each carries its SAFETY line | 2026-09-09 | | `prim/wasm.rs` | 6 | `memory.grow` linear-memory backend | 2026-08-06 | | `options.rs` | 17 | Env parsing at init, registered-hook invocation. **+1 on 2026-09-24: `#[cfg(test)]` only** — `std::env::set_var` (an `unsafe fn` in edition 2024) seeding two probe variables for the in-place lookup test. Shipped code gained none: the pass itself went from 251 allocations to zero by building its keys in a stack buffer and reading through `prim::getenv` **+10 on 2026-09-25:** the one-walk environment pass (Linux) — three `unsafe fn` helpers (`strip`, `match_name`, `copy_value`) that read a NUL-terminated entry byte by byte and stop at the first mismatch or its NUL, and the blocks that call them; nothing is written through an entry | 2026-09-25 | diff --git a/crates/rusty_alloc/src/arena.rs b/crates/rusty_alloc/src/arena.rs index d140764..3300cd7 100644 --- a/crates/rusty_alloc/src/arena.rs +++ b/crates/rusty_alloc/src/arena.rs @@ -539,6 +539,43 @@ pub fn chunk_free_n(p: *mut u8, n: usize) -> bool { false } +/// Whether `p` lies inside any arena's live chunks — the question +/// [`chunk_free`] answers as a side effect, asked without freeing anything. +/// +/// A segment whose memory cannot be made usable again must not be returned +/// to its arena (arena memory is committed once, at reservation, and handed +/// out as-is), and a chunk inside an arena's reservation cannot be released +/// to the OS on its own either. Such a segment is retired in place, and this +/// is how the release path tells the two cases apart. +#[allow(clippy::needless_range_loop)] // indexed scan over a fixed atomic table +pub fn owns(p: *const u8) -> bool { + if crate::FIXED_REGION { + return false; + } + let addr = p.addr(); + let n = ARENA_COUNT.load(Ordering::Acquire).min(MAX_ARENAS); + for id in 0..n { + let a = ARENAS[id].load(Ordering::Acquire); + if a.is_null() { + continue; + } + // SAFETY: live descriptor; only its base and live-chunk count are read. + unsafe { + let chunks = (*a).chunks_live.load(Ordering::Acquire); + let Some(span) = chunks.checked_mul(SEGMENT_SIZE) else { + continue; + }; + let Some(end_addr) = (*a).base.addr().checked_add(span) else { + continue; + }; + if addr >= (*a).base.addr() && addr < end_addr { + return true; + } + } + } + false +} + /// Return a chunk to its arena. True when the address belonged to one. #[allow(clippy::needless_range_loop)] // indexed scan over a fixed atomic table pub fn chunk_free(p: *mut u8) -> bool { diff --git a/crates/rusty_alloc/src/os.rs b/crates/rusty_alloc/src/os.rs index ab683fa..9b6b430 100644 --- a/crates/rusty_alloc/src/os.rs +++ b/crates/rusty_alloc/src/os.rs @@ -197,10 +197,44 @@ pub unsafe fn free(block: OsBlock) -> Result<(), PrimError> { /// # Safety /// Range must lie within a live block from [`alloc_aligned`], page-aligned. pub unsafe fn commit(ptr: *mut u8, size: usize) -> Result { + #[cfg(all(test, feature = "std"))] + if test_hooks::take_commit_failure() { + return Err(test_hooks::INJECTED); + } // SAFETY: forwarded contract. unsafe { prim::commit(ptr, page_align_up(size)) } } +/// Test-only fault injection: make this thread's next `n` commits fail, the +/// way `VirtualAlloc(MEM_COMMIT)` fails on Windows once the system's commit is +/// exhausted. Thread-local, so a parallel test's commits are never taken. +#[cfg(all(test, feature = "std"))] +pub(crate) mod test_hooks { + use core::cell::Cell; + + /// The synthetic error an injected failure returns. + pub const INJECTED: super::PrimError = 0xFA11; + + std::thread_local! { + static FAIL_COMMITS: Cell = const { Cell::new(0) }; + } + + /// Fail this thread's next `n` commits. + pub fn fail_next_commits(n: usize) { + FAIL_COMMITS.with(|c| c.set(n)); + } + + pub(super) fn take_commit_failure() -> bool { + FAIL_COMMITS.with(|c| { + let n = c.get(); + if n > 0 { + c.set(n - 1); + } + n > 0 + }) + } +} + /// Decommit a page-aligned sub-range. Returns whether recommit is required. /// /// # Safety diff --git a/crates/rusty_alloc/src/segment.rs b/crates/rusty_alloc/src/segment.rs index 997f0b1..5607c85 100644 --- a/crates/rusty_alloc/src/segment.rs +++ b/crates/rusty_alloc/src/segment.rs @@ -405,6 +405,34 @@ pub fn segment_alloc(arena_id: i32) -> Result<*mut Segment, PrimError> { Ok(seg) } +/// Make a dead segment's memory fit for a next tenant: lift guard-page +/// protection and re-commit purged spans (see `purged_any` / `guarded`). +/// Recycled memory is handed out as-is, so skipping this re-tenants an +/// inaccessible page (the M8 P0). Returns whether the memory is usable; when +/// the OS refuses either step (Windows refuses `MEM_COMMIT` once the system's +/// commit is exhausted), the caller must not recycle the segment, and the +/// failure is counted in [`crate::stats::commit_failures`]. +/// +/// # Safety +/// `seg` must be a live segment with no live blocks or references into it. +unsafe fn restore_for_reuse(seg: *mut Segment) -> bool { + // SAFETY: per contract; the range lies inside the segment's reservation. + unsafe { + if !((*seg).purged_any || (*seg).guarded) { + return true; + } + let base = seg.cast::().add(HEADER_SLICES * SEGMENT_SLICE_SIZE); + let bytes = (*seg).total_size - HEADER_SLICES * SEGMENT_SLICE_SIZE; + if os::protect(base, bytes, false).is_err() || os::commit(base, bytes).is_err() { + crate::stats::commit_failed(); + return false; + } + (*seg).purged_any = false; + (*seg).guarded = false; + true + } +} + /// Release an empty Normal segment (caller has unlinked it from the heap): /// back to its arena when it came from one, else to the OS. /// @@ -416,19 +444,16 @@ pub unsafe fn segment_free(seg: *mut Segment) -> Result<(), PrimError> { unsafe { wait_no_remote_in_flight(seg) }; segment_map::unregister(seg); // SAFETY: seg is live and empty per the contract. - unsafe { - if (*seg).purged_any || (*seg).guarded { - // Restore full commitment AND accessibility before the memory can - // be re-tenanted from an arena (see purged_any / guarded). - let base = seg.cast::().add(HEADER_SLICES * SEGMENT_SLICE_SIZE); - let bytes = (*seg).total_size - HEADER_SLICES * SEGMENT_SLICE_SIZE; - let _ = os::protect(base, bytes, false); - let _ = os::commit(base, bytes); - (*seg).purged_any = false; - (*seg).guarded = false; + let usable = unsafe { restore_for_reuse(seg) }; + if !usable { + // The memory could not be made accessible and committed again, so it + // must not be re-tenanted. Outside an arena, releasing it to the OS + // (below) needs neither; inside one, it cannot be released alone, so + // it stays marked used for the life of the process. + if crate::arena::owns(seg.cast()) { + return Ok(()); } - } - if crate::arena::chunk_free(seg.cast()) { + } else if crate::arena::chunk_free(seg.cast()) { return Ok(()); } // wasm: the slice pool is the free list (F2) — `expose_provenance` so a @@ -436,7 +461,12 @@ pub unsafe fn segment_free(seg: *mut Segment) -> Result<(), PrimError> { #[cfg(all(target_arch = "wasm32", not(miri)))] // SAFETY: seg is live per the contract; reading total_size. unsafe { - if crate::slice_pool::free_range(seg.cast::().expose_provenance(), (*seg).total_size) { + if usable + && crate::slice_pool::free_range( + seg.cast::().expose_provenance(), + (*seg).total_size, + ) + { return Ok(()); } } @@ -656,11 +686,23 @@ pub unsafe fn span_alloc(seg: *mut Segment, slices: usize) -> (*mut Page, bool) while !s.is_null() { let len = (*s).slice_count as usize; if len >= slices { - span_list_remove(seg, s); let idx = page_index(seg, s); // Re-commit BEFORE splitting so both halves are backed; the // remainder inherits the (now cleared) purged state. - span_recommit(seg, idx, len); + // + // A failed re-commit is a failed allocation, never a span + // handed out unbacked: Windows has no overcommit, so + // `MEM_COMMIT` fails when the system's commit is exhausted, + // and the page layer's first store would then fault on + // reserved memory (an access violation in `page_extend`, + // `docs/plans/recommit-failure-ignored.md`). The span stays + // on the free list, still marked purged; null sends the + // caller to another segment, and a fresh one fails its own + // commit cleanly. + if !span_recommit(seg, idx, len) { + return (ptr::null_mut(), false); + } + span_list_remove(seg, s); if len > slices { // The remainder starts at a new slice: re-mark all of it. span_mark_free(seg, idx + slices, len - slices, 1); @@ -828,19 +870,25 @@ pub unsafe fn purge_free_spans(seg: *mut Segment) { } } -/// Re-commit a span that was purged while free (no-op otherwise). +/// Re-commit a span that was purged while free (no-op otherwise). Returns +/// whether the span is backed. On failure the span stays marked purged, so +/// nothing forgets that its memory is not there. /// /// # Safety -/// `[idx, idx+len)` is a span of `seg` being handed to a caller. -unsafe fn span_recommit(seg: *mut Segment, idx: usize, len: usize) { +/// `[idx, idx+len)` is a span of `seg` about to be handed to a caller. +unsafe fn span_recommit(seg: *mut Segment, idx: usize, len: usize) -> bool { // SAFETY: caller contract; range lies inside the segment reservation. unsafe { if !(*seg).pages[idx].purged { - return; + return true; } - (*seg).pages[idx].purged = false; let area = page_area(seg, idx); - let _ = os::commit(area, len * SEGMENT_SLICE_SIZE); + if os::commit(area, len * SEGMENT_SLICE_SIZE).is_err() { + crate::stats::commit_failed(); + return false; + } + (*seg).pages[idx].purged = false; + true } } @@ -1035,16 +1083,20 @@ pub unsafe fn huge_free(seg: *mut Segment) -> Result<(), PrimError> { // memory must be handed back in a USABLE state: lift any guard-page // protection and restore commitment first. Skipping this recycles an // inaccessible page into the next tenant (the M8 P0). - if (*seg).total_size.is_multiple_of(SEGMENT_SIZE) { - if (*seg).guarded || (*seg).purged_any { - let base = seg.cast::().add(HEADER_SLICES * SEGMENT_SLICE_SIZE); - let bytes = (*seg).total_size - HEADER_SLICES * SEGMENT_SLICE_SIZE; - let _ = os::protect(base, bytes, false); - let _ = os::commit(base, bytes); - (*seg).guarded = false; - (*seg).purged_any = false; - } - if crate::arena::chunk_free_n(seg.cast(), (*seg).total_size / SEGMENT_SIZE) { + // + // When that restore fails, the memory must not be recycled at all: + // released to the OS if it is a reservation of its own, retired in + // place (left marked used) if it is chunks of an arena. + // Only chunk-multiple segments can reach an arena and need the + // restore; a ragged one is released whole (or pooled, on wasm). + let chunked = (*seg).total_size.is_multiple_of(SEGMENT_SIZE); + let usable = !chunked || restore_for_reuse(seg); + if chunked { + if !usable { + if crate::arena::owns(seg.cast()) { + return Ok(()); + } + } else if crate::arena::chunk_free_n(seg.cast(), (*seg).total_size / SEGMENT_SIZE) { return Ok(()); } } @@ -1052,7 +1104,12 @@ pub unsafe fn huge_free(seg: *mut Segment) -> Result<(), PrimError> { // the slice pool — the F2 fix itself. Chunk-multiple ones only // reach here when no arena owns them, and the pool takes those too. #[cfg(all(target_arch = "wasm32", not(miri)))] - if crate::slice_pool::free_range(seg.cast::().expose_provenance(), (*seg).total_size) { + if usable + && crate::slice_pool::free_range( + seg.cast::().expose_provenance(), + (*seg).total_size, + ) + { return Ok(()); } let block = os::OsBlock { @@ -1064,3 +1121,84 @@ pub unsafe fn huge_free(seg: *mut Segment) -> Result<(), PrimError> { os::free(block) } } + +#[cfg(all(test, feature = "std"))] +mod recommit_tests { + use super::*; + use crate::alloc::{collect, free, malloc, stats}; + use crate::os::test_hooks::fail_next_commits; + use crate::types::MEDIUM_PAGE_SLICES; + + /// A purged span whose re-commit FAILS must not be handed out. + /// + /// The MATA desktop app died with an access violation writing a heap + /// address in `page_extend` (2026-10-04, `secure` on, `purge_delay = 0`, + /// the machine at 98.5 % of its commit limit). `span_recommit` dropped + /// the result of `os::commit`, and on Windows that call fails when the + /// system's commit is exhausted, so a span whose re-commit failed was + /// carved anyway and the page layer's first store hit reserved memory + /// (`docs/plans/recommit-failure-ignored.md`, mechanism A). + /// + /// Without the fix the second allocation below lands back in the purged + /// span: on Windows, where the purge is a real `MEM_DECOMMIT`, writing it + /// is that access violation; everywhere, the "not in the purged span" + /// assertion fails. The last step checks the span still knows it is + /// purged, so a later re-commit that succeeds makes it usable again. + #[test] + fn a_failed_recommit_is_a_failed_span_not_an_unbacked_one() { + // A span `span_free` will purge: at least MEDIUM_PAGE_SLICES slices. + let span = SEGMENT_SLICE_SIZE * MEDIUM_PAGE_SLICES + SEGMENT_SLICE_SIZE; + crate::options::set(15, 0); // purge_delay: purge at once + let pin = malloc(span); + let p1 = malloc(span); + assert!(!pin.is_null() && !p1.is_null()); + assert_eq!( + segment_of(pin), + segment_of(p1), + "setup: the pin must keep p1's segment alive" + ); + // SAFETY: live blocks of `span` bytes. + unsafe { + core::ptr::write_bytes(pin, 1, span); + core::ptr::write_bytes(p1, 1, span); + } + let purges = stats().purges; + // SAFETY: p1 is live and freed once. + unsafe { free(p1) }; + collect(true); + assert!(stats().purges > purges, "setup: p1's span was not purged"); + + let failures = crate::stats::commit_failures(); + fail_next_commits(1); + let p2 = malloc(span); + fail_next_commits(0); + assert!( + !p2.is_null(), + "a failed re-commit must move on, not fail outright here" + ); + assert_eq!( + crate::stats::commit_failures(), + failures + 1, + "the re-commit was attempted, failed, and was counted" + ); + let in_p1 = (p1.addr()..p1.addr() + span).contains(&p2.addr()); + assert!(!in_p1, "a span whose re-commit failed was handed out"); + // SAFETY: p2 is a live block of `span` bytes; the write is the check. + unsafe { core::ptr::write_bytes(p2, 2, span) }; + + // The span kept its `purged` mark, so with commit available again + // whichever allocation reaches it re-commits it before use. + let p3 = malloc(span); + assert!(!p3.is_null()); + // SAFETY: live block of `span` bytes. + unsafe { core::ptr::write_bytes(p3, 3, span) }; + + // SAFETY: all live, each freed once. + unsafe { + free(p3); + free(p2); + free(pin); + } + crate::options::set(15, -1); // restore the shipped default + } +} diff --git a/crates/rusty_alloc/src/stats.rs b/crates/rusty_alloc/src/stats.rs index 6ddb755..94614f6 100644 --- a/crates/rusty_alloc/src/stats.rs +++ b/crates/rusty_alloc/src/stats.rs @@ -6,6 +6,29 @@ use crate::heap::Stats; #[cfg(feature = "std")] use crate::options::out_fmt; +use core::sync::atomic::{AtomicUsize, Ordering}; + +/// Process-wide count of OS commits (or protection restores) that failed on a +/// path that would otherwise have handed out, or recycled, memory that is not +/// backed or not accessible. Unlike +/// the per-heap counters it is global: a failed commit is a property of the +/// machine (Windows has no overcommit and refuses `MEM_COMMIT` when the +/// system's commit is exhausted), and the number a crash report needs is the +/// process's. Non-zero after an access violation in the page layer says the +/// machine ran out of commit, not that a span skipped its re-commit +/// (`docs/plans/recommit-failure-ignored.md`). +static COMMIT_FAILURES: AtomicUsize = AtomicUsize::new(0); + +/// Record one failed commit (cold: only reached when the OS said no). +#[cold] +pub(crate) fn commit_failed() { + COMMIT_FAILURES.fetch_add(1, Ordering::Relaxed); +} + +/// How many OS commits have failed in this process (see `COMMIT_FAILURES`). +pub fn commit_failures() -> usize { + COMMIT_FAILURES.load(Ordering::Relaxed) +} /// Sum the counters of every registered heap (`mi_stats_merge` semantics — /// our counters are always per-heap, so the merged view is computed on read). @@ -63,11 +86,12 @@ pub fn print_process() { let (elapsed, user, sys, rss, peak_rss, commit, peak_commit, faults) = process_info(); out_fmt(&std::format!( "process: elapsed {elapsed} ms, user {user} ms, sys {sys} ms, rss {} KiB (peak {}), \ - commit {} KiB (peak {}), faults {faults}\n", + commit {} KiB (peak {}), faults {faults}, failed commits {}\n", rss / 1024, peak_rss / 1024, commit / 1024, peak_commit / 1024, + commit_failures(), )); } diff --git a/docs/plans/recommit-failure-ignored.md b/docs/plans/recommit-failure-ignored.md new file mode 100644 index 0000000..15e83e5 --- /dev/null +++ b/docs/plans/recommit-failure-ignored.md @@ -0,0 +1,224 @@ +# An unbacked page reached `page_extend` on Windows, and a re-commit whose failure is ignored + +**Status:** one crash in a shipping consumer, **not reproduced**; mechanism A +**fixed in the working tree 2026-10-04, uncommitted** (section 8); which of the +two mechanisms caused the crash is still **open** · +**Date:** 2026-10-04 · **Seen on:** 1.1.6 (`secure` on, `purge_delay = 0`) · +**Code checked against:** HEAD `86980ec` (2.2.2) · **Platform:** Windows 11, +x86-64 · **From:** the MATA desktop session (`mata-master`, branch +`hc-0-connectivity`) + +--- + +## 0. The answer in one paragraph + +The MATA desktop app died with an access violation **writing** a heap address +inside `page_extend`: the allocator handed out a page whose memory was not +backed. The app runs with immediate purging and decommit, and the machine was +1.9 GB from its commit limit. There are two candidate mechanisms and the +evidence does not choose between them. **(A)** `span_recommit` discards the +result of `os::commit`, and on Windows a commit can fail, so a span whose +re-commit failed is used anyway. That is a defect in the source whether or not +it caused this crash. **(B)** A purged span reached a reuse path that does not +re-commit it at all, which is the family `docs/LEDGER.md` records under M8. +Section 5 says what to change for (A) and section 6 what would tell the two +apart. + +## 1. What was seen + +| | | +|---|---| +| Process | `desktop.exe` (MATA desktop, debug build), global allocator `rusty_alloc` 1.1.6 through `mata-alloc` | +| Configuration | `secure` feature on; `mata_alloc::configure(Profile::LongLived)` sets `purge_delay = 0`; `purge_decommits` left at its default, 1 | +| When | 2026-10-04 14:16:21 local, about seven minutes after launch | +| Exception | `0xC0000005`, access violation **writing** `0x0000017ab8310d80` | +| Faulting code | `desktop.exe+0x412f3` = `rusty_alloc::page::block_set_next` (`page.rs:309`) | +| System memory when the dump was written | commit charge 125.8 GB, limit 127.7 GB, peak 128.4 GB; 1.7 GB of 31.7 GB physical memory available | +| Process memory when the dump was written | 2.61 GB committed (pagefile usage), 0.13 GB working set | + +The call path, innermost first. It is an ordinary small allocation, a `String` +being built by `serde_json`: + +``` +rusty_alloc::page::block_set_next page.rs:309 +rusty_alloc::page::page_extend page.rs:1113 +rusty_alloc::heap::Heap::malloc_generic_walk +rusty_alloc::heap::Heap::malloc_generic heap.rs:457 +rusty_alloc::alloc::malloc_slow alloc.rs:322 +rusty_alloc::alloc::malloc alloc.rs:222 +rusty_alloc_api::...::alloc rusty_alloc-api lib.rs:146 +alloc::str::to_owned <- serde_json parse_str <- SyncEntry::deserialize +``` + +The fault address is a normal heap address, not null and not a wild value: +`rax = rdx = 0x17ab8310d80`, the block being linked, and `rcx = 0`. + +**Method.** The Windows Application event log (events 1000 and 1001) gave the +exception code and offset. `llvm-symbolizer --relative-address` against the +binary and its PDB resolved the offset. The local crash dump +(`%LOCALAPPDATA%\CrashDumps\desktop.exe.12424.dmp` on the machine that ran it) +gave the faulting address, the registers, the stack and the memory figures: +exception stream 6, thread list stream 3, system memory stream 21, process +counters stream 22. The stack listing is a scan of the faulting thread's stack +for return addresses inside the image, so it can include a stale frame; the +frames above are the ones that form a consistent chain. The source lines are +1.1.6's, read from the cargo registry copy. + +**The previous build.** The build before this one, with the allocator and its +configuration unchanged, ran for 31 minutes on the same machine an hour earlier +and was closed normally. So this is not a crash on every run. + +## 2. Mechanism A: the re-commit result is dropped + +Line numbers are 1.1.6 / HEAD. + +1. **Purging is on and it decommits.** `purge_delay = 0` makes a freed span + purge at once. With `purge_decommits = 1`, `os::purge` calls `decommit`, + which on Windows is `VirtualFree(MEM_DECOMMIT)` (`prim/windows.rs:163` / + `:173`). The span's first page is marked `purged`. +2. **Reuse re-commits, and ignores the answer.** `span_recommit` + (`segment.rs:768` / `:843`; called from the first-fit loop at `:663` in + HEAD): + + ```rust + (*seg).pages[idx].purged = false; + let area = page_area(seg, idx); + let _ = os::commit(area, len * SEGMENT_SLICE_SIZE); + ``` + + `os::commit` returns `Result`. On Windows it is + `VirtualAlloc(ptr, size, MEM_COMMIT, PAGE_READWRITE)` (`prim/windows.rs:152` + / `:153`). Windows has no overcommit, so that call fails when the system has + no commit left. The `purged` flag has already been cleared, so nothing + remembers that the span is not backed. +3. **The page layer writes to it.** `page_extend` (`page.rs:1113`) links the + page's fresh blocks with `block_set_next`. The first store lands on reserved, + uncommitted memory. + +Two more call sites drop the same result, on the path that hands a whole +segment back for re-tenanting: `segment.rs:394` and `:905` in 1.1.6, `:426` and +`:1043` in HEAD. Both also drop the result of `os::protect(base, bytes, false)` +on the line before. + +## 3. Mechanism B: a reuse path that does not re-commit + +`docs/LEDGER.md` (around lines 2910 and 3100 at HEAD) describes this family: +the M8 P0 was guard pages recycled while still no-access, and a later forced +purge "reaches spans whose reuse path does not re-commit them — the M8 defect +exactly (Windows `MEM_DECOMMIT` faults on touch)". The fault here has that +shape too. `span_recommit` reads the `purged` flag of the span's **first** +page only, so any way for a decommitted range to sit behind a first page whose +flag is clear would skip the commit. That was not traced for this note: no such +path has been shown, and none has been ruled out. With `secure` on, a guard +page left no-access is a third way to the same fault. + +## 4. What is established and what is not + +**Established:** +- The exception, its address and its kind (a write), the function it happened + in, and the call path. +- The consumer runs with immediate purging and decommit. +- The three `let _ = os::commit(...)` sites exist in 1.1.6 and in HEAD. +- The system's commit charge was 98.5% of its limit when the dump was written, + and had exceeded the current limit earlier (peak above limit). + +**Not established:** +- Which mechanism it was. The dump has no memory-region list, so the state of + the faulting page is not recorded, and nothing logs the `VirtualAlloc` result. +- For A: 1.9 GB of commit was still free when the dump was written, and a span + re-commit is far smaller than that. A needs the headroom to have been gone at + the instant of the call. The machine was being pushed to its limit by other + processes throughout, so that is plausible and unproven. +- For B: nothing beyond the shape of the fault and the family's history. + +## 5. Suggested change for A + +Treat a failed re-commit as a failed allocation. + +- `span_recommit` returns whether the span is backed. On failure it leaves + `purged` set and the caller does not use the span: put it back on the free + list and either try the next span or fail the allocation. The public entry + then returns null, and Rust's `handle_alloc_error` ends the process with + "memory allocation of N bytes failed". +- The two re-tenanting sites do the same: a segment that cannot be made + accessible and committed again is released, not handed on. +- The `os::protect(.., false)` results on those paths get the same treatment, + because a span left no-access fails in the same way. + +The process still ends when the machine is truly out of memory. What changes: +the report names the real cause, callers that allocate fallibly +(`try_reserve`) get the null they asked for, and the allocator no longer hands +out memory it does not have. + +## 6. Tests, and telling A from B + +- **Mock prim.** `prim/mock.rs` makes commit a bookkeeping no-op that always + succeeds. Add a switch that makes the next N commits fail, and have the mock + track which ranges are decommitted so a write to one is detectable. Then: + allocate, free with `purge_delay = 0`, fail the next commit, allocate again. + Today that returns a pointer into a decommitted range; it should return null. + The same range tracking catches B without any commit failing: every pointer + handed out must lie in a committed range. +- **Windows, real pages.** A Job Object with `JOB_OBJECT_LIMIT_PROCESS_MEMORY` + set just above the test's current commit makes `MEM_COMMIT` fail on demand + without exhausting the machine. Same sequence; expect null and no access + violation. +- **In the field.** A counter of failed commits, readable through the stats + interface, would settle it on the next crash: non-zero means A. The MATA app + can also be made to write dumps with the memory-region list + (`MiniDumpWithFullMemoryInfo`), which records whether the faulting page was + reserved or no-access. + +## 7. Who is exposed + +- A: any consumer on Windows with purging on (`purge_delay >= 0`) and + `purge_decommits = 1`, when the machine runs out of commit. The process does + not have to be large: this one held 2.6 GB. +- B, if it exists: the same configuration, at any time. +- The default configuration (`purge_delay = -1`) never sets `purged`, so + `span_recommit` returns early. `mata-alloc` turns purging on in + `configure(Profile::LongLived)`, which the MATA desktop app calls. +- Unix was not examined for this note. + +## 8. What was changed for A (2026-10-04, uncommitted) + +- **`span_recommit` returns whether the span is backed** (`segment.rs`). On a + failed `os::commit` it leaves `purged` set, counts the failure, and + `span_alloc` returns null with the span still on the free list. Both callers + in `heap.rs` already treat null as "this segment has no room" and move on; a + fresh segment then fails its own commit as an ordinary allocation failure. +- **`restore_for_reuse`** (`segment.rs`) is the guard-lift + re-commit that + `segment_free` and `huge_free` each did inline with both results dropped. + When either step fails, the segment is not recycled: released to the OS if + it is a reservation of its own (release needs neither step), retired in + place, its used bit left set, if it is chunks of an arena (arena memory is + committed once at reservation and handed out as-is, and a chunk cannot be + released alone). `arena::owns` tells the two apart. `huge_free` restores only + chunk-multiple segments, as before. +- **`stats::commit_failures()`**, a process-wide counter, also printed by + `print_process` as `failed commits N`. Non-zero after a fault in the page + layer means the machine ran out of commit (A); zero points at B. +- **Test:** `segment::recommit_tests::a_failed_recommit_is_a_failed_span_not_an_unbacked_one`, + through a `cfg(test)`-only, thread-local "fail the next N commits" switch in + `os::commit` (the miri-only mock in `prim/mock.rs` would not reach native + runs, and on Windows the purge is a real `MEM_DECOMMIT`, which is what makes + the test reproduce the fault). **Poisoned:** with the check in `span_alloc` + disabled, the test fails at "a span whose re-commit failed was handed out". +- Verified on Windows: the lib and 10 integration suites pass (88 tests), + `secure` on as well (71), clippy `-D warnings` with `secure`, `cargo fmt`, + `cargo check --target wasm32-unknown-unknown`, and the unsafe census (965 -> + 973, recorded in `UNSAFE.md`: +3 shipped, +5 test-only). + +**Not covered by a test:** the two release sites. Exercising them needs a +whole segment freed while purged, with the failure injected on that path; the +logic is small and reviewed, but it has not been reproduced. + +**Still open:** B. The counter and a `MiniDumpWithFullMemoryInfo` dump from +the MATA app are what will tell A from B on the next crash. + +**Context on the crash itself:** at 14:16 the same machine was also running +the `bench/alloc-eval` replays of a recorded Endless Sky battle in WSL (about +2 GB per run), and the WSL VM restarted twice around then from host memory +exhaustion. That load very likely contributed to the 98.5 % commit charge the +dump recorded, which makes A's precondition more plausible than section 4 +assumed. It does not rule out B. diff --git a/tools/unsafe-baseline.txt b/tools/unsafe-baseline.txt index d6f4dc5..f8c4f0d 100644 --- a/tools/unsafe-baseline.txt +++ b/tools/unsafe-baseline.txt @@ -1,5 +1,5 @@ 106 crates/rusty_alloc/src/alloc.rs - 15 crates/rusty_alloc/src/arena.rs + 16 crates/rusty_alloc/src/arena.rs 77 crates/rusty_alloc/src/heap.rs 42 crates/rusty_alloc/src/init.rs 2 crates/rusty_alloc/src/lib.rs @@ -13,7 +13,7 @@ 6 crates/rusty_alloc/src/prim/wasm.rs 34 crates/rusty_alloc/src/prim/windows.rs 2 crates/rusty_alloc/src/random.rs - 36 crates/rusty_alloc/src/segment.rs + 43 crates/rusty_alloc/src/segment.rs 3 crates/rusty_alloc/src/stats.rs 16 crates/rusty_alloc_api/src/lib.rs 15 crates/rusty_alloc_bench/src/kernels.rs