diff --git a/Cargo.toml b/Cargo.toml index 8e6f5869..9b72af53 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -84,6 +84,10 @@ required-features = ["vulkan", "ash/loaded"] name = "d3d12-buffer-winrs" required-features = ["d3d12"] +[[example]] +name = "d3d12-stomp" +required-features = ["d3d12"] + [[example]] name = "metal-buffer" required-features = ["metal"] diff --git a/README.md b/README.md index 0eaff06c..04c65a4a 100644 --- a/README.md +++ b/README.md @@ -74,9 +74,15 @@ let mut allocator = Allocator::new(&AllocatorCreateDesc { device: ID3D12DeviceVersion::Device(device), debug_settings: Default::default(), allocation_sizes: Default::default(), + stomp: None, }); ``` +For development builds, `stomp: Some(StompSettings::default())` turns on stomp detection +for resources created with `Allocator::create_resource()`: each resource is packed against +a never-resident guard page so that a GPU write past its end page-faults and DRED names the +resource. See `d3d12::StompSettings` for modes and limitations. + ## Simple d3d12 allocation example ```rust diff --git a/docs/d3d12-stomp-allocator.md b/docs/d3d12-stomp-allocator.md new file mode 100644 index 00000000..15eef344 --- /dev/null +++ b/docs/d3d12-stomp-allocator.md @@ -0,0 +1,325 @@ +# D3D12 stomp allocator: implementation plan + +Debug-only mode for the D3D12 backend that places every guarded buffer so that a GPU +write past its end (or, optionally, before its start) hits a page that has never been +resident. The GPU page-faults and the device is removed with `DXGI_ERROR_DEVICE_HUNG`. +Freed resources can be quarantined so that use-after-free faults too. + +This document is written so that it can be executed top to bottom without further +context. Sections 1, 3, 4 and 6 describe the design as implemented in +`src/d3d12/stomp.rs` after the spike; section 8 records the design that was tried +first and why it was rejected. + +## 1. Mechanism and why it is shaped this way + +Facts that constrain the design. Do not re-derive them. + +1. Residency in D3D12 is per heap. A heap created with + `D3D12_HEAP_FLAG_CREATE_NOT_RESIDENT` and never passed to `MakeResident` has no + physical backing. Any GPU access to it page-faults. `Evict` is only a hint and must + not be used for this purpose (Jesse Natalie, Microsoft). +2. Unmapped (NULL) tiles in reserved resources do not fault on tiled tier 2 hardware: + reads return zero, writes are dropped. They are useless as guards. +3. D3D12 gives no control over where a heap lands in GPU virtual address (VA) space. + Only reserved resources give VA control. Section 8 has the measurements that show + heap adjacency cannot be relied on. +4. Textures have no GPU VA and all view accesses are hardware bounds-checked. The + unbounded write paths in D3D12 are root descriptors (raw VA), acceleration structure + builds and scratch, DXR shader tables and ExecuteIndirect arguments. Copy commands + are rejected by the runtime at `Close()` when they overrun, so they never reach a + guard. Guarding buffers is what catches real stomps. +5. `Allocator::allocate()` returns a heap and an offset for the caller's own + `CreatePlacedResource`. That path cannot be guarded and stays untouched. + Only `Allocator::create_resource()` is affected. +6. `UpdateTileMappings` is a command queue operation. Stomp mode therefore needs a + queue and a fence; the allocator CPU-waits after each mapping update so the + resource is usable on any queue when `create_resource` returns. + +Resulting layout for one stomp buffer, all inside one reserved resource: + +``` +tile 0 : head guard (optional; mapped to the never-resident heap) +tiles h .. h+n : payload (mapped to a 64KB-aligned range of a normal allocator block) +tile h+n : tail guard (default on; mapped to the never-resident heap) +``` + +`h` is 1 with a head guard, else 0; `n = ceil(size / 64KB)`. The never-resident heap is +one 64KB heap per memory type, created lazily, mapped with +`D3D12_TILE_RANGE_FLAG_REUSE_SINGLE_TILE` so every guard tile of every resource shares +it. Adjacency is by construction, so this holds under churn and across threads. + +A reserved resource always starts on a tile boundary, so the tail guard is 64KB precise +by default. With `StompSettings::payload_alignment` set, the payload is packed against +the tail guard at `Resource::offset()` (rounded down to that alignment) and the caller +adds the offset to every VA and view. Ignoring the offset only widens the slack. + +Textures become reserved resources with `D3D12_TEXTURE_LAYOUT_64KB_UNDEFINED_SWIZZLE` +and exact dimensions, all tiles mapped to allocator memory, no guard tiles. They gain +use-after-free detection only. + +Quarantine: on `free_resource`, all tiles of a stomp resource are remapped to the +never-resident heap and the `ID3D12Resource` is kept alive until the allocator drops. +A stale descriptor or stale VA then faults. VA is retained, memory is returned to the +sub-allocator immediately. + +## 2. Spike (done, `examples/d3d12-stomp.rs`) + +Measured on an NVIDIA RTX 5070 Ti, resource heap tier 2, tiled resources tier 3: + +- Access to a `CREATE_NOT_RESIDENT` heap removes the device with + `DXGI_ERROR_DEVICE_HUNG`. Root UAV write, root SRV read and texture guard access all + fault. In-bounds and slack accesses complete. +- DRED does not attribute the fault on this driver: `PageFaultVA 0`, no allocation + nodes, even for a bogus VA. Breadcrumbs work. Treat naming as informational. +- `CopyBufferRegion` cannot overrun: the runtime rejects it at `Close()` with + `E_INVALIDARG`. +- After the second device removal in one process, `D3D12CreateDevice` keeps failing + with `DEVICE_HUNG`. Fault tests run each body in a child process. +- Heap VA adjacency (the first design, section 8) fails under churn. + +## 3. Public API (`src/d3d12/stomp.rs`, re-exported from `gpu_allocator::d3d12`) + +```rust +#[derive(Clone, Debug)] +pub struct StompSettings { + /// Queue that executes the tile mapping updates. Use a dedicated queue (COPY is fine). + pub queue: ID3D12CommandQueue, + pub mode: StompMode, + /// Seed for `StompMode::Random`. Same seed, same sequence of decisions. + pub seed: u64, + /// Guard tile after the payload. Default true: most stomps run forward. + pub tail_guard: bool, + /// Guard tile before the payload. Default false. + pub head_guard: bool, + /// Pack the payload against the tail guard, rounded down to this power of two (<= 65536). + /// None keeps offset() == 0. + pub payload_alignment: Option, + /// Keep freed resources alive with all tiles mapped to the never-resident heap. + pub quarantine: bool, +} +impl StompSettings { pub fn new(queue: ID3D12CommandQueue) -> Self } // All, tail on, head off, None, quarantine on + +#[derive(Clone, Copy, Debug)] +pub enum StompMode { All, Random { probability: f32 }, OptIn, Filter(fn(&ResourceCreateDesc<'_>) -> bool) } + +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct StompLayout { + pub resource_va: u64, // GetGPUVirtualAddress() of the reserved resource, 0 for textures + pub offset: u64, // payload byte offset inside the resource + pub payload_tiles: u32, + pub head_guard: bool, + pub tail_guard: bool, +} +impl StompLayout { + pub fn head_guard_va(&self) -> Option; // resource_va + pub fn tail_guard_va(&self) -> Option; // resource_va + (head + payload_tiles) * 64KB +} + +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +pub struct StompStatistics { pub guarded: usize, pub unguarded_fallbacks: usize, pub quarantined: usize } +``` + +- `AllocatorCreateDesc::stomp: Option`; `None` keeps today's behaviour. +- `ResourceCreateDesc::stomp: Option`; `Some` overrides the mode. +- `Resource::offset()`, `Resource::is_stomp_guarded()` (buffers with a guard), + `Resource::stomp_layout()`. +- `Allocator::stomp_statistics()`, `Allocator::committed_statistics()`. +- `AllocationError::InvalidStompSettings(String)` from `Allocator::new` for + `probability` outside `0.0..=1.0`, both guards off, `payload_alignment` not a power of + two or above 65536, and tiled resources tier 0. +- `guarded` counts every resource that took the stomp path, textures included; + `unguarded_fallbacks` counts refusals (multisampled, 3D below tiled tier 3, and + `CpuToGpu` / `GpuToCpu` because reserved resources cannot be `Map()`ed, verified: + `E_INVALIDARG`), which go to the normal path with one `warn!`. + +## 4. Internals + +`src/d3d12/mod.rs` keeps only the hooks: the `stomp` fields on the descs, `Resource` +(`stomp: Option`, `offset`), `Allocator` (`stomp: Option`) and +`MemoryType` (`guard_heap: Option`); `StompState::new` in `Allocator::new`; +`try_create_resource_stomp` at the top of `create_resource`; `release_resource` in +`free_resource`; and the drop order in `Drop for Allocator` (quarantine, then memory +blocks, then guard heaps). Shared helpers stay in `mod.rs`: `create_heap`, `heap_flags`, +`find_memory_type_index`, `create_placed`. + +`src/d3d12/stomp.rs` holds everything else: + +- Pure functions, unit-tested without a device: `xorshift64`, `stomp_decide`, + `stomp_geometry(size, tail_guard, head_guard, payload_alignment) -> (payload_tiles, + offset)`, `validate_stomp_settings`. +- `StompState { settings, rng, tiled_tier, fence, fence_value, guarded, + unguarded_fallbacks, quarantine: Vec }`. +- `MemoryType::guard_heap()`: lazy 64KB heap with the category flags plus + `CREATE_NOT_RESIDENT`, never made resident. +- `Allocator::create_resource_stomp`: + 1. Buffers: `Width = (h + n + t) * 64KB`, `Alignment = 0`. Textures: `Layout = + 64KB_UNDEFINED_SWIZZLE`, `Alignment = 0`, no guards. + 2. `CreateReservedResource` (or `CreateReservedResource2` on Device10/Device12 for + barrier layouts and castable formats), same error mapping as `create_placed`. + 3. Textures: `GetResourceTiling` for `NumTilesForEntireResource`. + 4. Payload memory through the existing `allocate()` with `size = payload_tiles * 64KB`, + `alignment = 64KB`, so leak reports, visualizer and `free` are unchanged. + 5. `UpdateTileMappings` on `settings.queue`: payload region to the allocation's heap at + `offset / 64KB`; head and tail tiles to the guard heap with `REUSE_SINGLE_TILE`. + 6. `queue.Signal(fence)` then `fence.SetEventOnCompletion(value, HANDLE::default())`, + which blocks on a null handle. On error, free the allocation and return the error. + 7. `SetName(desc.name)`, return `Resource { allocation: Some, stomp: Some(layout), offset }`. +- `Allocator::release_resource` (free side): with quarantine on, remap all + `h + n + t` tiles to the guard heap, wait, push the resource into `quarantine`; + otherwise drop it. Then the normal `free(allocation)`. + +## 5. Example: `examples/d3d12-stomp.rs` + +DRED on, allocator with `StompSettings::new(copy_queue)`, 4000-byte buffer named +`"stomp me"`, root UAV compute write at `tail_guard_va()`, execute, fence wait, expect +device removal, print the reason and whatever DRED attributes. Exit 0 when the guard +access removed the device. Falls back to WARP with a note that WARP does not fault. + +## 6. Tests + +Three layers. Layers 1 and 2 run on CI (`windows-latest` has WARP). Layer 3 needs real +hardware because WARP does not page-fault. + +### 6.1 Layer 1: unit tests, no device (`src/d3d12/stomp.rs`, `mod tests`) + +Geometry table, `(size, tail, head, payload_alignment) -> (payload_tiles, offset)`: + +| size | tail | head | alignment | tiles | offset | +|-------|-------|-------|-----------|-------|---------------| +| 4000 | true | false | Some(8) | 1 | 61536 | +| 4000 | true | false | None | 1 | 0 | +| 4000 | true | true | Some(256) | 1 | 65536 + 61440 | +| 4000 | false | true | Some(8) | 1 | 65536 | +| 65536 | true | false | Some(8) | 1 | 0 | +| 65537 | true | false | Some(8) | 2 | 65528 | +| 70000 | true | false | Some(256) | 2 | 60928 | +| 1 | true | false | Some(65536) | 1 | 0 | +| 1 | true | false | Some(8) | 1 | 65528 | + +Property check over 10k random sizes and alignments, with and without head guard: +`offset % alignment == 0`, `offset >= head * 64KB`, `offset + size <= (head + tiles) * +64KB`, slack under `alignment` when `Some`, and `offset == head * 64KB` when `None`. + +Random mode bounds and determinism, decision matrix for all modes and overrides, +settings validation (probability, both guards off, alignment 0 / 3 / 131072), +`heap_flags` mapping. + +### 6.2 Layer 2: device-backed integration tests (`tests/d3d12_stomp.rs`) + +Gated `#![cfg(all(windows, feature = "d3d12"))]`. Helpers: `make_device()` with the +debug layer, hardware else WARP, panic if neither; `assert_no_d3d12_errors()` sweeping +`ID3D12InfoQueue` for `ERROR`/`CORRUPTION` at the end of every test; a copy queue for +`StompSettings::new`. Each test creates its own device and allocator. + +1. `plain_allocator_unchanged`: no stomp settings; `stomp_statistics()` is `None`, + `stomp_layout()` is `None`, `offset() == 0`. +2. `all_mode_stomps_buffer`: 4000-byte `GpuOnly` buffer; `allocation.is_some()`, + `is_stomp_guarded()`, layout has `tail_guard`, no `head_guard`, `payload_tiles == 1`, + `offset == 0`; `guarded == 1`, `unguarded_fallbacks == 0`. +3. `width_covers_tiles`: `GetDesc().Width == (payload_tiles + 1) * 65536`. +4. `tail_guard_va_follows_payload`: `tail_guard_va() == resource_va + 65536`; + `GetGPUVirtualAddress() == resource_va`. +5. `payload_alignment_packs_against_tail`: `payload_alignment: Some(256)`; `offset == + 61440`, `offset + 4000 <= 65536`, slack under 256. +6. `head_guard_layout`: `head_guard: true, tail_guard: true`; `head_guard_va() == + resource_va`, payload starts at tile 1 (`offset == 65536` with `None` alignment), + `tail_guard_va() == resource_va + 2 * 65536`. +7. `head_only`: `head_guard: true, tail_guard: false`; `offset == 65536`, + `tail_guard_va()` is `None`. +8. `random_zero_guards_none` / `random_one_guards_all`: 50 buffers each. +9. `random_is_deterministic`: two allocators, same seed, identical decision vectors. +10. `opt_in_mode` and `Some(false)` under `All`. +11. `filter_mode`: buffers stomped, textures on the normal path. +12. `texture_takes_stomp_path`: 256x256 texture; `stomp_layout().is_some()`, + `is_stomp_guarded() == false`, `resource_va == 0`, `payload_tiles == + GetResourceTiling` count, `GetDesc()` dimensions unchanged, `Layout == + 64KB_UNDEFINED_SWIZZLE`. +13. `rt_texture_takes_stomp_path`: same with `ALLOW_RENDER_TARGET` and a clear value. +14. `msaa_falls_back`: 4x MSAA render target on the normal path, `unguarded_fallbacks == 1`. +15. `cpu_heaps_fall_back_and_stay_mappable`: `CpuToGpu` and `GpuToCpu` buffers under + `All` take the normal path (`unguarded_fallbacks` incremented) and `Map()` works. + Reserved resources reject `Map()` with `E_INVALIDARG`, measured. +16. `committed_request_takes_stomp_path`. +17. `free_quarantines`: create 20, free all; `quarantined == 20`, `generate_report()` + shows zero live allocations, no D3D12 errors. +18. `free_without_quarantine`: `quarantine: false`; `quarantined == 0` after frees. +19. `churn_no_errors`: 500 random-size buffers (1 byte to 4MB), ring of 32; `guarded == + 500`, `unguarded_fallbacks == 0`, every resource `is_stomp_guarded()`, no D3D12 + errors. This is the invariant plan B could not meet. +20. `drop_order_is_safe`: drop the allocator with quarantined resources, with a live + stomp resource (expect the "not freed" warning), and after freeing everything. +21. `rename_and_report_unaffected`. +22. `tiled_tier_required`: skip unless the adapter reports tier 0 (WARP is tier 3). +23. `threads_do_not_break_guards`: 4 threads each with its own allocator on the same + device, 100 buffers each; all guarded. + +### 6.3 Layer 3: fault tests, hardware only (`tests/d3d12_stomp_fault.rs`) + +All `#[ignore]`, early return unless `GPU_ALLOCATOR_STOMP_FAULT_TESTS=1`, each body in a +child process (device creation fails after two removals in one process), run with +`--test-threads=1`. Root-descriptor compute shaders under `tests/shaders/` with the dxc +command in each `.hlsl` header. + +1. `tail_overrun_faults`: root UAV write at `tail_guard_va()`; expect removal. +2. `in_bounds_write_does_not_fault`: write at `resource_va + offset + size - 4`. +3. `slack_write_does_not_fault`: with `payload_alignment: Some(8)`, write at + `tail_guard_va() - 4` (in the slack or the payload), no fault; write at + `tail_guard_va()`, fault. +4. `head_underrun_faults`: `head_guard: true`; root SRV read at `head_guard_va()`. +5. `use_after_free_faults`: quarantine on; keep a root UAV to `resource_va` of a buffer, + free it, write; expect removal. With `quarantine: false` the same write must not + fault (memory may be reused, but no page is missing). +6. `texture_use_after_free_faults`: keep a descriptor to a stomp texture, free it, + sample it in a compute shader; expect removal. +7. `dred_attribution`: informational; print `PageFaultVA` and allocation node names. + +### 6.4 CI wiring + +- `tests/d3d12_stomp.rs` runs on the existing `windows-latest` job via + `cargo test --workspace --all-targets --features d3d12,...` on WARP. +- Layer 3: + `GPU_ALLOCATOR_STOMP_FAULT_TESTS=1 cargo test --features d3d12 --test d3d12_stomp_fault -- --ignored --test-threads=1`. + +## 7. Documentation + +The `StompSettings` doc comment covers what is caught and what is not, the offset +contract, `GetDesc().Width` and `CopyResource`, quarantine, the queue requirement, DRED +enablement and its limits on NVIDIA, and WARP. `src/lib.rs` points at `StompSettings` +under the D3D12 setup section; regenerate `README.md` with `cargo readme > README.md`. + +## 8. Rejected first design: adjacent heaps with tight placed resource alignment + +The first implementation gave every stomp resource its own heap, placed the resource at +the end of that heap (byte-exact with `D3D12_RESOURCE_FLAG_USE_TIGHT_ALIGNMENT`, Agility +SDK 1.618.1 and newer), created a 64KB `CREATE_NOT_RESIDENT` heap right after it, and +verified adjacency by reading the GPU VA of throwaway placed buffers, retrying on a miss. +No offset contract, no queue, exact `Width`. Measurements on NVIDIA: + +- Raw heap pairs (payload then guard, dropped each iteration): 0.0 percent adjacent in + steady state. Freed VA is reused by size class; a 64KB guard lands in any 64KB hole. +- Guard heap sized equal to the payload: 96.9 percent for 64KB payloads, 7.8 percent + for random 64KB to 16MB payloads. +- Through the allocator, 500 buffers of 1 byte to 4MB with a ring of 32 live: 55 of 500 + guarded with 0 retries, 72 with 4, 105 with 16. Payloads that are 2MB multiples never + got an adjacent guard even on fresh VA. +- The driver's VA allocator is process-wide across devices and threads, so heaps + created on another thread land between payload and guard. + +A guard that holds for the first allocation and then silently degrades is worse than +none. Tight alignment is unusable with reserved resources (spec: `E_INVALIDARG`), so it +plays no part in the current design. + +## 9. Verification checklist + +``` +cargo fmt --all -- --check +cargo clippy --workspace --all-targets --no-default-features --features d3d12,std -- -D warnings +cargo test --workspace --all-targets --no-default-features --features d3d12,std +cargo doc --no-deps --workspace --all-features --document-private-items +cargo readme > README.md && git diff --quiet README.md +cargo run --example d3d12-stomp --features d3d12 # on a real GPU, expect exit code 0 +GPU_ALLOCATOR_STOMP_FAULT_TESTS=1 cargo test --features d3d12 --test d3d12_stomp_fault -- --ignored --test-threads=1 # real GPU only +``` + +Also run the existing `d3d12-buffer-winrs` example to confirm the non-stomp path is +untouched. diff --git a/examples/d3d12-buffer-winrs.rs b/examples/d3d12-buffer-winrs.rs index b5f5354c..62e1d421 100644 --- a/examples/d3d12-buffer-winrs.rs +++ b/examples/d3d12-buffer-winrs.rs @@ -93,6 +93,7 @@ fn main() -> Result<()> { device: ID3D12DeviceVersion::Device(device.clone()), debug_settings: Default::default(), allocation_sizes: Default::default(), + stomp: None, }) .unwrap(); diff --git a/examples/d3d12-stomp.rs b/examples/d3d12-stomp.rs new file mode 100644 index 00000000..15bd2cba --- /dev/null +++ b/examples/d3d12-stomp.rs @@ -0,0 +1,228 @@ +//! Stomp mode demo. See `docs/d3d12-stomp-allocator.md` and [`StompSettings`]. +//! +//! Creates a stomp-guarded buffer named "stomp me" and writes to its tail guard page through +//! a root UAV. Root descriptors are not bounds checked, which is exactly the kind of access +//! stomp mode exists to catch. The GPU page-faults and the device is removed. Exit code 0 +//! when that happened, 1 otherwise. +//! +//! DRED page fault attribution is printed when the driver provides it; NVIDIA currently +//! reports the removal but no fault VA. WARP does not page-fault, so on WARP the demo only +//! prints the placement. +use gpu_allocator::{ + d3d12::{ + Allocator, AllocatorCreateDesc, ID3D12DeviceVersion, ResourceCategory, ResourceCreateDesc, + ResourceStateOrBarrierLayout, ResourceType, StompSettings, + }, + MemoryLocation, +}; +use windows::{ + core::{Interface, Result}, + Win32::{ + Foundation::HANDLE, + Graphics::{ + Direct3D::D3D_FEATURE_LEVEL_11_0, + Direct3D12::*, + Dxgi::{ + Common::{DXGI_FORMAT_UNKNOWN, DXGI_SAMPLE_DESC}, + CreateDXGIFactory2, IDXGIAdapter1, IDXGIFactory6, DXGI_ADAPTER_FLAG_SOFTWARE, + DXGI_ERROR_NOT_FOUND, + }, + }, + }, +}; + +/// Compute shader that stores 4 bytes through a root UAV bound at an arbitrary GPU VA. +/// Source and compile command in `tests/shaders/`. +const ROOT_UAV_WRITE_CS: &[u8] = include_bytes!("../tests/shaders/root_uav_write.dxil"); + +fn enable_dred() { + let mut settings: Option = None; + match unsafe { D3D12GetDebugInterface(&mut settings) } { + Ok(()) => unsafe { + let settings = settings.unwrap(); + settings.SetPageFaultEnablement(D3D12_DRED_ENABLEMENT_FORCED_ON); + settings.SetAutoBreadcrumbsEnablement(D3D12_DRED_ENABLEMENT_FORCED_ON); + }, + Err(e) => println!("DRED unavailable ({e}), the fault will not be attributed"), + } +} + +/// Hardware adapter first, WARP as fallback. Returns `(device, adapter name, is_warp)`. +fn create_device() -> Result<(ID3D12Device, String, bool)> { + let factory: IDXGIFactory6 = unsafe { CreateDXGIFactory2(Default::default()) }?; + for idx in 0.. { + let adapter: IDXGIAdapter1 = match unsafe { factory.EnumAdapters1(idx) } { + Ok(a) => a, + Err(e) if e.code() == DXGI_ERROR_NOT_FOUND => break, + Err(e) => return Err(e), + }; + let desc = unsafe { adapter.GetDesc1() }?; + if desc.Flags & DXGI_ADAPTER_FLAG_SOFTWARE.0 as u32 != 0 { + continue; + } + let mut device: Option = None; + if unsafe { D3D12CreateDevice(&adapter, D3D_FEATURE_LEVEL_11_0, &mut device) }.is_ok() { + let name = String::from_utf16_lossy(&desc.Description) + .trim_end_matches('\0') + .to_string(); + return Ok((device.unwrap(), name, false)); + } + } + let warp: IDXGIAdapter1 = unsafe { factory.EnumWarpAdapter() }?; + let mut device: Option = None; + unsafe { D3D12CreateDevice(&warp, D3D_FEATURE_LEVEL_11_0, &mut device) }?; + Ok((device.unwrap(), "WARP".into(), true)) +} + +fn create_queue(device: &ID3D12Device, ty: D3D12_COMMAND_LIST_TYPE) -> Result { + unsafe { + device.CreateCommandQueue(&D3D12_COMMAND_QUEUE_DESC { + Type: ty, + ..Default::default() + }) + } +} + +/// Runs one dispatch of the root-UAV-write kernel at `va` on a direct queue and waits. +fn write_at(device: &ID3D12Device, va: u64) -> Result<()> { + unsafe { + let root_signature: ID3D12RootSignature = + device.CreateRootSignature(0, ROOT_UAV_WRITE_CS)?; + let pso: ID3D12PipelineState = + device.CreateComputePipelineState(&D3D12_COMPUTE_PIPELINE_STATE_DESC { + pRootSignature: std::mem::ManuallyDrop::new(Some(root_signature.clone())), + CS: D3D12_SHADER_BYTECODE { + pShaderBytecode: ROOT_UAV_WRITE_CS.as_ptr().cast(), + BytecodeLength: ROOT_UAV_WRITE_CS.len(), + }, + ..Default::default() + })?; + let queue = create_queue(device, D3D12_COMMAND_LIST_TYPE_DIRECT)?; + let allocator: ID3D12CommandAllocator = + device.CreateCommandAllocator(D3D12_COMMAND_LIST_TYPE_DIRECT)?; + let list: ID3D12GraphicsCommandList = + device.CreateCommandList(0, D3D12_COMMAND_LIST_TYPE_DIRECT, &allocator, None)?; + list.SetComputeRootSignature(&root_signature); + list.SetPipelineState(&pso); + list.SetComputeRootUnorderedAccessView(0, va); + list.Dispatch(1, 1, 1); + list.Close()?; + queue.ExecuteCommandLists(&[Some(list.cast()?)]); + let fence: ID3D12Fence = device.CreateFence(0, D3D12_FENCE_FLAG_NONE)?; + queue.Signal(&fence, 1)?; + // A null event blocks the calling thread until the fence reaches the value. + fence.SetEventOnCompletion(1, HANDLE::default()) + } +} + +/// Names DRED associates with the faulting VA, when the driver attributes page faults. +fn dred_page_fault(device: &ID3D12Device) -> Option<(u64, Vec)> { + let dred: ID3D12DeviceRemovedExtendedData = device.cast().ok()?; + let out = unsafe { dred.GetPageFaultAllocationOutput() }.ok()?; + let mut names = Vec::new(); + for head in [ + out.pHeadExistingAllocationNode, + out.pHeadRecentFreedAllocationNode, + ] { + let mut node = head; + while !node.is_null() { + let n = unsafe { &*node }; + if !n.ObjectNameW.is_null() { + names.push(unsafe { n.ObjectNameW.to_string() }.unwrap_or_default()); + } + node = n.pNext; + } + } + Some((out.PageFaultVA, names)) +} + +fn main() -> Result<()> { + env_logger::Builder::from_env(env_logger::Env::default().default_filter_or("info")).init(); + enable_dred(); + let (device, name, is_warp) = create_device()?; + println!("adapter: {name}"); + + // Tile mappings are enqueued on this queue and CPU-waited, so keep it off the main queue. + let mapping_queue = create_queue(&device, D3D12_COMMAND_LIST_TYPE_COPY)?; + let mut allocator = Allocator::new(&AllocatorCreateDesc { + device: ID3D12DeviceVersion::Device(device.clone()), + debug_settings: Default::default(), + allocation_sizes: Default::default(), + stomp: Some(StompSettings::new(mapping_queue)), + }) + .expect("Allocator::new"); + + let desc = D3D12_RESOURCE_DESC { + Dimension: D3D12_RESOURCE_DIMENSION_BUFFER, + Alignment: 0, + Width: 4000, + Height: 1, + DepthOrArraySize: 1, + MipLevels: 1, + Format: DXGI_FORMAT_UNKNOWN, + SampleDesc: DXGI_SAMPLE_DESC { + Count: 1, + Quality: 0, + }, + Layout: D3D12_TEXTURE_LAYOUT_ROW_MAJOR, + Flags: D3D12_RESOURCE_FLAG_NONE, + }; + let resource = allocator + .create_resource(&ResourceCreateDesc { + name: "stomp me", + memory_location: MemoryLocation::GpuOnly, + resource_category: ResourceCategory::Buffer, + resource_desc: &desc, + castable_formats: &[], + clear_value: None, + initial_state_or_layout: ResourceStateOrBarrierLayout::ResourceState( + D3D12_RESOURCE_STATE_COMMON, + ), + resource_type: &ResourceType::Placed, + stomp: None, + }) + .expect("create_resource"); + + let layout = resource.stomp_layout().expect("stomp path"); + println!("is_stomp_guarded: {}", resource.is_stomp_guarded()); + println!("stomp_layout: {layout:#x?}"); + println!("stomp_statistics: {:?}", allocator.stomp_statistics()); + + if is_warp { + println!("WARP does not page-fault on never-resident memory, stopping here"); + allocator.free_resource(resource).unwrap(); + std::process::exit(1); + } + + let guard_va = layout.tail_guard_va().expect("tail guard"); + println!("writing 4 bytes at the tail guard {guard_va:#x}"); + if let Err(e) = write_at(&device, guard_va) { + println!("submission error: {e}"); + } + + let removed = match unsafe { device.GetDeviceRemovedReason() } { + Ok(()) => { + println!("device NOT removed: the guard access went unnoticed"); + false + } + Err(reason) => { + println!("device removed: {reason}"); + // Fault data is gathered asynchronously by the runtime; give it a moment. + std::thread::sleep(std::time::Duration::from_millis(500)); + match dred_page_fault(&device) { + Some((va, names)) if names.iter().any(|n| n == "stomp me") => { + println!("DRED page fault VA {va:#x} named \"stomp me\"") + } + Some((va, names)) => println!( + "DRED page fault VA {va:#x}, allocations {names:?}: this driver does not \ + attribute the fault" + ), + None => println!("DRED page fault output unavailable"), + } + true + } + }; + // Freeing after device removal is still well-formed on the CPU side. + allocator.free_resource(resource).unwrap(); + std::process::exit(if removed { 0 } else { 1 }); +} diff --git a/src/d3d12/mod.rs b/src/d3d12/mod.rs index 6ead72c5..41f144cc 100644 --- a/src/d3d12/mod.rs +++ b/src/d3d12/mod.rs @@ -18,6 +18,10 @@ use windows::Win32::{ }, }; +mod stomp; +use stomp::StompState; +pub use stomp::{StompLayout, StompMode, StompSettings, StompStatistics}; + #[cfg(feature = "visualizer")] mod visualizer; #[cfg(feature = "visualizer")] @@ -56,6 +60,10 @@ pub struct ResourceCreateDesc<'a> { pub clear_value: Option<&'a D3D12_CLEAR_VALUE>, pub initial_state_or_layout: ResourceStateOrBarrierLayout, pub resource_type: &'a ResourceType<'a>, + /// Per-resource stomp override. `Some(true)` forces a guard, `Some(false)` never guards, + /// `None` defers to [`StompSettings::mode`]. Ignored when the allocator was created without + /// [`AllocatorCreateDesc::stomp`]. + pub stomp: Option, } #[derive(Clone, Copy, Debug, PartialEq, Eq)] @@ -164,6 +172,8 @@ pub struct AllocatorCreateDesc { pub device: ID3D12DeviceVersion, pub debug_settings: AllocatorDebugSettings, pub allocation_sizes: AllocationSizes, + /// Debug-only stomp detection, see [`StompSettings`]. `None` disables it at zero cost. + pub stomp: Option, } pub enum ResourceType<'a> { @@ -188,12 +198,31 @@ pub struct Resource { pub memory_location: MemoryLocation, memory_type_index: Option, pub size: u64, + stomp: Option, + offset: u64, } impl Resource { pub fn resource(&self) -> &ID3D12Resource { self.resource.as_ref().expect("Resource was already freed.") } + + /// Byte offset of the payload inside [`Self::resource()`]. Always `0` except for stomp + /// buffers with [`StompSettings::payload_alignment`] set, see the offset contract there. + pub fn offset(&self) -> u64 { + self.offset + } + + /// `true` when this resource was created by stomp mode with at least one guard tile. + /// `false` for normal resources and for stomp textures, which only get quarantine. + pub fn is_stomp_guarded(&self) -> bool { + self.stomp.is_some_and(|s| s.head_guard || s.tail_guard) + } + + /// Tile layout when this resource was created by stomp mode. + pub fn stomp_layout(&self) -> Option { + self.stomp + } } impl Drop for Resource { @@ -258,6 +287,41 @@ impl Allocation { } } +fn heap_flags(heap_category: HeapCategory) -> D3D12_HEAP_FLAGS { + match heap_category { + HeapCategory::All => D3D12_HEAP_FLAG_NONE, + HeapCategory::Buffer => D3D12_HEAP_FLAG_ALLOW_ONLY_BUFFERS, + HeapCategory::RtvDsvTexture => D3D12_HEAP_FLAG_ALLOW_ONLY_RT_DS_TEXTURES, + HeapCategory::OtherTexture => D3D12_HEAP_FLAG_ALLOW_ONLY_NON_RT_DS_TEXTURES, + } +} + +fn create_heap( + device: &ID3D12Device, + size: u64, + heap_properties: &D3D12_HEAP_PROPERTIES, + flags: D3D12_HEAP_FLAGS, + alignment: u64, +) -> Result { + let desc = D3D12_HEAP_DESC { + SizeInBytes: size, + Properties: *heap_properties, + Alignment: alignment, + Flags: flags, + }; + + let mut heap = None; + match unsafe { device.CreateHeap(&desc, &mut heap) } { + Err(e) if e.code() == E_OUTOFMEMORY => Err(AllocationError::OutOfMemory), + Err(e) => Err(AllocationError::Internal(format!( + "ID3D12Device::CreateHeap failed: {e}" + ))), + Ok(()) => heap.ok_or_else(|| { + AllocationError::Internal("ID3D12Heap pointer is null, but should not be.".into()) + }), + } +} + #[derive(Debug)] struct MemoryBlock { heap: ID3D12Heap, @@ -272,34 +336,13 @@ impl MemoryBlock { heap_category: HeapCategory, dedicated: bool, ) -> Result { - let heap = { - let mut desc = D3D12_HEAP_DESC { - SizeInBytes: size, - Properties: *heap_properties, - Alignment: D3D12_DEFAULT_MSAA_RESOURCE_PLACEMENT_ALIGNMENT as u64, - ..Default::default() - }; - desc.Flags = match heap_category { - HeapCategory::All => D3D12_HEAP_FLAG_NONE, - HeapCategory::Buffer => D3D12_HEAP_FLAG_ALLOW_ONLY_BUFFERS, - HeapCategory::RtvDsvTexture => D3D12_HEAP_FLAG_ALLOW_ONLY_RT_DS_TEXTURES, - HeapCategory::OtherTexture => D3D12_HEAP_FLAG_ALLOW_ONLY_NON_RT_DS_TEXTURES, - }; - - let mut heap = None; - let hr = unsafe { device.CreateHeap(&desc, &mut heap) }; - match hr { - Err(e) if e.code() == E_OUTOFMEMORY => Err(AllocationError::OutOfMemory), - Err(e) => Err(AllocationError::Internal(format!( - "ID3D12Device::CreateHeap failed: {e}" - ))), - Ok(()) => heap.ok_or_else(|| { - AllocationError::Internal( - "ID3D12Heap pointer is null, but should not be.".into(), - ) - }), - }? - }; + let heap = create_heap( + device, + size, + heap_properties, + heap_flags(heap_category), + D3D12_DEFAULT_MSAA_RESOURCE_PLACEMENT_ALIGNMENT as u64, + )?; let sub_allocator: Box = if dedicated { Box::new(DedicatedBlockAllocator::new(size)) @@ -324,6 +367,8 @@ struct MemoryType { heap_properties: D3D12_HEAP_PROPERTIES, memory_type_index: usize, active_general_blocks: usize, + /// Stomp guard: one tile, never resident. Created on first use. + guard_heap: Option, } impl MemoryType { @@ -507,6 +552,7 @@ pub struct Allocator { debug_settings: AllocatorDebugSettings, memory_types: Vec, allocation_sizes: AllocationSizes, + stomp: Option, } impl Allocator { @@ -514,6 +560,11 @@ impl Allocator { &self.device } + /// Committed resource statistics, one per memory type. + pub fn committed_statistics(&self) -> impl Iterator { + self.memory_types.iter().map(|m| &m.committed_allocations) + } + pub fn new(desc: &AllocatorCreateDesc) -> Result { // Perform AddRef on the device let device = desc.device.clone(); @@ -533,6 +584,15 @@ impl Allocator { let is_heap_tier1 = options.ResourceHeapTier == D3D12_RESOURCE_HEAP_TIER_1; + let stomp = match &desc.stomp { + None => None, + Some(settings) => Some(StompState::new( + settings, + options.TiledResourcesTier, + &device, + )?), + }; + let heap_types = [ ( MemoryLocation::GpuOnly, @@ -605,6 +665,7 @@ impl Allocator { num_allocations: 0, total_size: 0, }, + guard_heap: None, }, ) .collect::>(); @@ -614,9 +675,29 @@ impl Allocator { device, debug_settings: desc.debug_settings, allocation_sizes: desc.allocation_sizes, + stomp, }) } + fn find_memory_type_index( + &self, + location: MemoryLocation, + resource_category: ResourceCategory, + ) -> Result { + self.memory_types + .iter() + .position(|memory_type| { + let is_location_compatible = + location == MemoryLocation::Unknown || location == memory_type.memory_location; + + let is_category_compatible = memory_type.heap_category == HeapCategory::All + || memory_type.heap_category == resource_category.into(); + + is_location_compatible && is_category_compatible + }) + .ok_or(AllocationError::NoCompatibleMemoryTypeFound) + } + pub fn allocate(&mut self, desc: &AllocationCreateDesc<'_>) -> Result { let size = desc.size; let alignment = desc.alignment; @@ -644,20 +725,9 @@ impl Allocator { return Err(AllocationError::InvalidAllocationCreateDesc); } - // Find memory type - let memory_type = self - .memory_types - .iter_mut() - .find(|memory_type| { - let is_location_compatible = desc.location == MemoryLocation::Unknown - || desc.location == memory_type.memory_location; - - let is_category_compatible = memory_type.heap_category == HeapCategory::All - || memory_type.heap_category == desc.resource_category.into(); - - is_location_compatible && is_category_compatible - }) - .ok_or(AllocationError::NoCompatibleMemoryTypeFound)?; + let memory_type_index = + self.find_memory_type_index(desc.location, desc.resource_category)?; + let memory_type = &mut self.memory_types[memory_type_index]; memory_type.allocate( &self.device, @@ -779,9 +849,91 @@ impl Allocator { } } + fn create_placed( + &self, + heap: &ID3D12Heap, + offset: u64, + resource_desc: &D3D12_RESOURCE_DESC, + desc: &ResourceCreateDesc<'_>, + ) -> Result { + let mut result: Option = None; + if let Err(e) = unsafe { + match (&self.device, desc.initial_state_or_layout) { + (_, ResourceStateOrBarrierLayout::ResourceState(_)) + if !desc.castable_formats.is_empty() => + { + return Err(AllocationError::CastableFormatsRequiresEnhancedBarriers) + } + ( + ID3D12DeviceVersion::Device12(device), + ResourceStateOrBarrierLayout::BarrierLayout(initial_layout), + ) => { + let resource_desc1 = Self::d3d12_resource_desc_1(resource_desc); + device.CreatePlacedResource2( + heap, + offset, + &resource_desc1, + initial_layout, + None, + Some(desc.castable_formats), + &mut result, + ) + } + (_, ResourceStateOrBarrierLayout::BarrierLayout(_)) + if !desc.castable_formats.is_empty() => + { + return Err(AllocationError::CastableFormatsRequiresAtLeastDevice12) + } + ( + ID3D12DeviceVersion::Device10(device), + ResourceStateOrBarrierLayout::BarrierLayout(initial_layout), + ) => { + let resource_desc1 = Self::d3d12_resource_desc_1(resource_desc); + device.CreatePlacedResource2( + heap, + offset, + &resource_desc1, + initial_layout, + None, + None, + &mut result, + ) + } + (_, ResourceStateOrBarrierLayout::BarrierLayout(_)) => { + return Err(AllocationError::BarrierLayoutNeedsDevice10) + } + (device, ResourceStateOrBarrierLayout::ResourceState(initial_state)) => device + .CreatePlacedResource( + heap, + offset, + resource_desc, + initial_state, + None, + &mut result, + ), + } + } { + if e.code() == DXGI_ERROR_DEVICE_REMOVED { + return Err(AllocationError::Internal(format!( + "ID3D12Device::CreatePlacedResource DEVICE_REMOVED: {:?}", + unsafe { self.device.GetDeviceRemovedReason() } + ))); + } + return Err(AllocationError::Internal(format!( + "ID3D12Device::CreatePlacedResource failed: {e}" + ))); + } + + Ok(result.expect("Allocation succeeded but no resource was returned?")) + } + /// Create a resource according to the provided parameters. /// Created resources should be freed at the end of their lifetime by calling [`Self::free_resource()`]. pub fn create_resource(&mut self, desc: &ResourceCreateDesc<'_>) -> Result { + if let Some(result) = self.try_create_resource_stomp(desc) { + return result; + } + match desc.resource_type { ResourceType::Committed { heap_properties, @@ -867,20 +1019,9 @@ impl Allocator { let allocation_info = Self::resource_allocation_info(&self.device, desc); - let memory_type = self - .memory_types - .iter_mut() - .find(|memory_type| { - let is_location_compatible = desc.memory_location - == MemoryLocation::Unknown - || desc.memory_location == memory_type.memory_location; - - let is_category_compatible = memory_type.heap_category == HeapCategory::All - || memory_type.heap_category == desc.resource_category.into(); - - is_location_compatible && is_category_compatible - }) - .ok_or(AllocationError::NoCompatibleMemoryTypeFound)?; + let memory_type_index = + self.find_memory_type_index(desc.memory_location, desc.resource_category)?; + let memory_type = &mut self.memory_types[memory_type_index]; memory_type.committed_allocations.num_allocations += 1; memory_type.committed_allocations.total_size += allocation_info.SizeInBytes; @@ -891,7 +1032,9 @@ impl Allocator { resource: Some(resource), size: allocation_info.SizeInBytes, memory_location: desc.memory_location, - memory_type_index: Some(memory_type.memory_type_index), + memory_type_index: Some(memory_type_index), + stomp: None, + offset: 0, }) } ResourceType::Placed => { @@ -909,76 +1052,12 @@ impl Allocator { let allocation = self.allocate(&allocation_desc)?; - let mut result: Option = None; - if let Err(e) = unsafe { - match (&self.device, desc.initial_state_or_layout) { - (_, ResourceStateOrBarrierLayout::ResourceState(_)) - if !desc.castable_formats.is_empty() => - { - return Err(AllocationError::CastableFormatsRequiresEnhancedBarriers) - } - ( - ID3D12DeviceVersion::Device12(device), - ResourceStateOrBarrierLayout::BarrierLayout(initial_layout), - ) => { - let resource_desc1 = Self::d3d12_resource_desc_1(desc.resource_desc); - device.CreatePlacedResource2( - allocation.heap(), - allocation.offset(), - &resource_desc1, - initial_layout, - None, - Some(desc.castable_formats), - &mut result, - ) - } - (_, ResourceStateOrBarrierLayout::BarrierLayout(_)) - if !desc.castable_formats.is_empty() => - { - return Err(AllocationError::CastableFormatsRequiresAtLeastDevice12) - } - ( - ID3D12DeviceVersion::Device10(device), - ResourceStateOrBarrierLayout::BarrierLayout(initial_layout), - ) => { - let resource_desc1 = Self::d3d12_resource_desc_1(desc.resource_desc); - device.CreatePlacedResource2( - allocation.heap(), - allocation.offset(), - &resource_desc1, - initial_layout, - None, - None, - &mut result, - ) - } - (_, ResourceStateOrBarrierLayout::BarrierLayout(_)) => { - return Err(AllocationError::BarrierLayoutNeedsDevice10) - } - (device, ResourceStateOrBarrierLayout::ResourceState(initial_state)) => { - device.CreatePlacedResource( - allocation.heap(), - allocation.offset(), - desc.resource_desc, - initial_state, - None, - &mut result, - ) - } - } - } { - if e.code() == DXGI_ERROR_DEVICE_REMOVED { - return Err(AllocationError::Internal(format!( - "ID3D12Device::CreatePlacedResource DEVICE_REMOVED: {:?}", - unsafe { self.device.GetDeviceRemovedReason() } - ))); - } - return Err(AllocationError::Internal(format!( - "ID3D12Device::CreatePlacedResource failed: {e}" - ))); - } - - let resource = result.expect("Allocation succeeded but no resource was returned?"); + let resource = self.create_placed( + unsafe { allocation.heap() }, + allocation.offset(), + desc.resource_desc, + desc, + )?; let size = allocation.size(); Ok(Resource { name: desc.name.into(), @@ -987,6 +1066,8 @@ impl Allocator { size, memory_location: desc.memory_location, memory_type_index: None, + stomp: None, + offset: 0, }) } } @@ -996,11 +1077,13 @@ impl Allocator { pub fn free_resource(&mut self, mut resource: Resource) -> Result<()> { // Explicitly drop the resource (which is backed by a refcounted COM object) // before freeing allocated memory. Windows-rs performs a Release() on drop(). - let _ = resource + let d3d12_resource = resource .resource .take() .expect("Resource was already freed."); + self.release_resource(&resource, d3d12_resource); + if let Some(allocation) = resource.allocation.take() { self.free(allocation) } else { @@ -1071,9 +1154,14 @@ impl Drop for Allocator { // Because Rust drop rules drop members in source-code order (that would be the // ID3D12Device before the ID3D12Heaps nested in these memory blocks), free - // all remaining memory blocks manually first by dropping. + // all remaining memory blocks manually first by dropping. Quarantined reserved + // resources still map into the guard heaps, so they go first. + if let Some(stomp) = &mut self.stomp { + stomp.quarantine.clear(); + } for mem_type in self.memory_types.iter_mut() { mem_type.memory_blocks.clear(); + mem_type.guard_heap = None; } } } diff --git a/src/d3d12/stomp.rs b/src/d3d12/stomp.rs new file mode 100644 index 00000000..d88b750d --- /dev/null +++ b/src/d3d12/stomp.rs @@ -0,0 +1,877 @@ +//! Debug-only stomp detection for the D3D12 backend, see [`StompSettings`]. +//! +//! Everything stomp-specific lives here. `mod.rs` only carries the hooks: the `stomp` fields +//! on the descs, `Resource`, `Allocator` and `MemoryType`, and the calls into +//! [`Allocator::try_create_resource_stomp()`] and [`Allocator::release_resource()`]. + +use alloc::vec::Vec; + +use log::warn; +use windows::{ + core::PCWSTR, + Win32::{ + Foundation::HANDLE, + Graphics::{Direct3D12::*, Dxgi::DXGI_ERROR_DEVICE_REMOVED}, + }, +}; + +use super::{ + create_heap, heap_flags, AllocationCreateDesc, Allocator, ID3D12DeviceVersion, MemoryType, + Resource, ResourceCreateDesc, ResourceStateOrBarrierLayout, +}; +use crate::{AllocationError, MemoryLocation, Result}; + +/// Tile size of reserved resources, also the heap granularity used by stomp mode. +const TILE: u64 = D3D12_TILED_RESOURCE_TILE_SIZE_IN_BYTES as u64; + +/// Debug-only stomp detection for resources created with [`Allocator::create_resource()`]. +/// +/// A stomp resource is a reserved (tiled) resource. Its payload tiles are mapped to memory +/// from the normal sub-allocator; the guard tiles before and/or after the payload are mapped +/// to a 64KB heap that was created with `D3D12_HEAP_FLAG_CREATE_NOT_RESIDENT` and is never +/// made resident. A GPU access to a guard tile page-faults and removes the device with +/// `DXGI_ERROR_DEVICE_HUNG`. Adjacency of payload and guard is guaranteed by construction, so +/// this works under allocation churn and from multiple threads. +/// +/// # What is caught +/// +/// Accesses that the hardware does not bounds-check: root descriptor SRV/UAV/CBV reads and +/// writes, acceleration structure builds and their scratch buffers, DXR shader tables and +/// `ExecuteIndirect` arguments. Descriptor-table views are clamped by the hardware and never +/// reach the guard. Copy commands are rejected by the runtime when they overrun, they are not +/// caught here either. +/// +/// Buffers get a tail guard (default) and optionally a head guard. Textures get no guard: +/// D3D12 exposes no unbounded write path into textures. Both get use-after-free detection +/// through `quarantine`. +/// +/// Only [`MemoryLocation::GpuOnly`] resources are stomped. Reserved resources cannot be +/// `Map()`ed, so `CpuToGpu` and `GpuToCpu` resources take the normal path (counted in +/// [`StompStatistics::unguarded_fallbacks`]), as do multisampled resources and 3D textures +/// below tiled resources tier 3. +/// +/// [`Allocator::allocate()`] is not affected: the caller creates the placed resource there. +/// +/// # Offset contract +/// +/// A reserved resource always starts on a 64KB tile, so a tail guard is only 64KB precise by +/// default: a buffer of 4000 bytes has 61536 bytes of slack before the guard. Set +/// `payload_alignment` to get byte precision: the payload is then packed against the tail +/// guard and starts at [`Resource::offset()`], rounded down to that alignment. The caller +/// must add that offset to `GetGPUVirtualAddress()`, `BufferLocation`, `FirstElement` and +/// copy offsets. Ignoring the offset only widens the slack, nothing breaks. +/// +/// The reserved buffer is `Width = tiles * 64KB`, so `GetDesc().Width` is larger than +/// requested and `CopyResource` between a stomp buffer and a normal one fails validation. +/// +/// # Quarantine +/// +/// With `quarantine` on, [`Allocator::free_resource()`] remaps every tile of the resource to +/// the never-resident heap and keeps the `ID3D12Resource` alive until the allocator is dropped. +/// A stale descriptor or stale GPU VA then faults instead of reading whatever reused the +/// memory. This keeps GPU virtual address space, not memory. +/// +/// # Queue +/// +/// Tile mappings are enqueued on `queue` and CPU-waited before the resource is returned. Pass +/// a dedicated queue (a `D3D12_COMMAND_LIST_TYPE_COPY` queue is fine) so the wait does not +/// stall your main queue. +/// +/// # Diagnosing a fault +/// +/// Enable DRED before creating the device: +/// +/// ```ignore +/// let dred: ID3D12DeviceRemovedExtendedDataSettings = D3D12GetDebugInterface()?; +/// dred.SetPageFaultEnablement(D3D12_DRED_ENABLEMENT_FORCED_ON); +/// dred.SetAutoBreadcrumbsEnablement(D3D12_DRED_ENABLEMENT_FORCED_ON); +/// ``` +/// +/// Not every driver attributes the fault: NVIDIA returns `PageFaultVA 0` and no allocation +/// nodes, breadcrumbs still show the offending command. Resources are named after +/// [`ResourceCreateDesc::name`] for the drivers that do. WARP does not fault at all. +/// +/// Development builds only: every stomp resource costs a reserved resource, tile mapping +/// calls and a fence wait. +#[derive(Clone, Debug)] +pub struct StompSettings { + /// Queue that executes the tile mapping updates. Use a dedicated queue. + pub queue: ID3D12CommandQueue, + pub mode: StompMode, + /// Seed for [`StompMode::Random`]. Same seed, same sequence of decisions. + pub seed: u64, + /// Guard tile after the payload. Default `true`: most stomps run forward. + pub tail_guard: bool, + /// Guard tile before the payload. Default `false`. + pub head_guard: bool, + /// Pack the payload against the tail guard, rounded down to this power of two (at most + /// 65536). `None` keeps the payload at the start of its tiles with `offset() == 0`. + pub payload_alignment: Option, + /// Keep freed resources alive with all tiles mapped to the never-resident heap. + pub quarantine: bool, +} + +impl StompSettings { + /// Mode `All`, tail guard only, no payload alignment, quarantine on. + pub fn new(queue: ID3D12CommandQueue) -> Self { + Self { + queue, + mode: StompMode::All, + seed: 0x9E37_79B9_7F4A_7C15, + tail_guard: true, + head_guard: false, + payload_alignment: None, + quarantine: true, + } + } +} + +/// Which resources get guarded. A per-resource [`ResourceCreateDesc::stomp`] of `Some` +/// overrides the mode. +#[derive(Clone, Copy, Debug)] +pub enum StompMode { + /// Every resource. + All, + /// Each resource independently with this probability in `0.0..=1.0`. + Random { probability: f32 }, + /// Only resources with [`ResourceCreateDesc::stomp`] set to `Some(true)`. + OptIn, + /// Caller decides per resource. A closure that captures nothing coerces to this, + /// e.g. `StompMode::Filter(|d| d.resource_category == ResourceCategory::Buffer)`. + Filter(fn(&ResourceCreateDesc<'_>) -> bool), +} + +/// Tile layout of a stomp resource. Tiles are 64KB. With a head guard, tile 0 is the guard +/// and the payload starts at tile 1; the tail guard is the tile after the payload. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct StompLayout { + /// `GetGPUVirtualAddress()` of the reserved resource, `0` for textures. + pub resource_va: u64, + /// Byte offset of the payload inside the resource, see [`Resource::offset()`]. + pub offset: u64, + /// Tiles backed by allocator memory. + pub payload_tiles: u32, + pub head_guard: bool, + pub tail_guard: bool, +} + +impl StompLayout { + fn head_tiles(&self) -> u32 { + u32::from(self.head_guard) + } + + fn total_tiles(&self) -> u32 { + self.head_tiles() + self.payload_tiles + u32::from(self.tail_guard) + } + + /// GPU VA of the tail guard tile, `None` without one or for textures. + pub fn tail_guard_va(&self) -> Option { + (self.tail_guard && self.resource_va != 0) + .then(|| self.resource_va + u64::from(self.head_tiles() + self.payload_tiles) * TILE) + } + + /// GPU VA of the head guard tile, `None` without one or for textures. + pub fn head_guard_va(&self) -> Option { + (self.head_guard && self.resource_va != 0).then_some(self.resource_va) + } +} + +/// Counts since the allocator was created. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq)] +pub struct StompStatistics { + /// Resources that took the stomp path (guarded buffers and quarantine-only textures). + pub guarded: usize, + /// Resources the stomp path refused (multisampled, or 3D below tiled tier 3). + pub unguarded_fallbacks: usize, + /// Freed resources kept alive with all tiles mapped to the never-resident heap. + pub quarantined: usize, +} + +pub(super) struct StompState { + settings: StompSettings, + rng: u64, + tiled_tier: D3D12_TILED_RESOURCES_TIER, + fence: ID3D12Fence, + fence_value: u64, + guarded: usize, + unguarded_fallbacks: usize, + pub(super) quarantine: Vec, +} + +fn xorshift64(state: &mut u64) -> u64 { + let mut x = *state; + x ^= x << 13; + x ^= x >> 7; + x ^= x << 17; + *state = x; + x +} + +fn stomp_decide( + mode: StompMode, + per_resource: Option, + rng: &mut u64, + desc: &ResourceCreateDesc<'_>, +) -> bool { + if let Some(forced) = per_resource { + return forced; + } + match mode { + StompMode::All => true, + StompMode::OptIn => false, + StompMode::Filter(f) => f(desc), + StompMode::Random { probability } => { + // 24 random bits, uniform in `0.0..1.0`, so `0.0` never and `1.0` always selects. + let unit = (xorshift64(rng) >> 40) as f32 / (1u64 << 24) as f32; + unit < probability + } + } +} + +/// Payload tile count and payload byte offset for a stomp buffer of `size` bytes. With +/// `tail_guard` and a `payload_alignment` the payload is packed against the tail guard, +/// rounded down to that alignment; otherwise it starts on its first tile. +fn stomp_geometry( + size: u64, + tail_guard: bool, + head_guard: bool, + payload_alignment: Option, +) -> (u32, u64) { + let payload_tiles = ((size + TILE - 1) / TILE) as u32; + let slack = match payload_alignment { + Some(alignment) if tail_guard => { + (u64::from(payload_tiles) * TILE - size) & !(alignment - 1) + } + _ => 0, + }; + (payload_tiles, u64::from(head_guard) * TILE + slack) +} + +fn validate_stomp_settings( + mode: StompMode, + tail_guard: bool, + head_guard: bool, + payload_alignment: Option, +) -> Result<()> { + if let StompMode::Random { probability } = mode { + if !(0.0..=1.0).contains(&probability) { + return Err(AllocationError::InvalidStompSettings(format!( + "probability {probability} is not in 0.0..=1.0" + ))); + } + } + if !tail_guard && !head_guard { + return Err(AllocationError::InvalidStompSettings( + "at least one of tail_guard or head_guard must be enabled".into(), + )); + } + if let Some(alignment) = payload_alignment { + if !alignment.is_power_of_two() || alignment > TILE { + return Err(AllocationError::InvalidStompSettings(format!( + "payload_alignment {alignment} must be a power of two <= {TILE}" + ))); + } + } + Ok(()) +} + +impl StompState { + pub(super) fn new( + settings: &StompSettings, + tiled_tier: D3D12_TILED_RESOURCES_TIER, + device: &ID3D12Device, + ) -> Result { + validate_stomp_settings( + settings.mode, + settings.tail_guard, + settings.head_guard, + settings.payload_alignment, + )?; + if tiled_tier.0 < D3D12_TILED_RESOURCES_TIER_1.0 { + return Err(AllocationError::InvalidStompSettings( + "tiled resources unsupported".into(), + )); + } + let fence = unsafe { device.CreateFence(0, D3D12_FENCE_FLAG_NONE) }.map_err(|e| { + AllocationError::Internal(format!("ID3D12Device::CreateFence failed: {e}")) + })?; + Ok(Self { + settings: settings.clone(), + // xorshift never leaves state 0. + rng: settings.seed | 1, + tiled_tier, + fence, + fence_value: 0, + guarded: 0, + unguarded_fallbacks: 0, + quarantine: Vec::new(), + }) + } +} + +impl MemoryType { + fn guard_heap(&mut self, device: &ID3D12Device) -> Result<&ID3D12Heap> { + if self.guard_heap.is_none() { + self.guard_heap = Some(create_heap( + device, + TILE, + &self.heap_properties, + heap_flags(self.heap_category) | D3D12_HEAP_FLAG_CREATE_NOT_RESIDENT, + TILE, + )?); + } + Ok(self + .guard_heap + .as_ref() + .expect("guard heap was just created")) + } +} + +impl Allocator { + /// Stomp counts since creation, `None` without stomp mode. + pub fn stomp_statistics(&self) -> Option { + self.stomp.as_ref().map(|s| StompStatistics { + guarded: s.guarded, + unguarded_fallbacks: s.unguarded_fallbacks, + quarantined: s.quarantine.len(), + }) + } + + fn should_stomp(&mut self, desc: &ResourceCreateDesc<'_>) -> bool { + match &mut self.stomp { + Some(state) => stomp_decide(state.settings.mode, desc.stomp, &mut state.rng, desc), + None => false, + } + } + + /// Hook for [`Allocator::create_resource()`]: `Some` when the resource was created (or + /// failed) on the stomp path, `None` when the normal path should handle it. + pub(super) fn try_create_resource_stomp( + &mut self, + desc: &ResourceCreateDesc<'_>, + ) -> Option> { + if !self.should_stomp(desc) { + return None; + } + let state = self.stomp.as_mut().expect("stomp state exists"); + let d = desc.resource_desc; + let tier3 = state.tiled_tier.0 >= D3D12_TILED_RESOURCES_TIER_3.0; + let refused = if d.SampleDesc.Count > 1 { + Some("multisampled") + } else if d.Dimension == D3D12_RESOURCE_DIMENSION_TEXTURE3D && !tier3 { + Some("3D below tiled resources tier 3") + } else if matches!( + desc.memory_location, + MemoryLocation::CpuToGpu | MemoryLocation::GpuToCpu + ) { + // Reserved resources cannot be `Map()`ed (E_INVALIDARG), so a CPU-visible + // stomp resource would be useless to the caller. + Some("CPU-visible") + } else { + None + }; + match refused { + Some(why) => { + warn!("Stomp: `{}` is {why}, not guarded", desc.name); + state.unguarded_fallbacks += 1; + None + } + None => Some(self.create_resource_stomp(desc)), + } + } + + /// Hook for [`Allocator::free_resource()`]: quarantines a stomp resource when enabled, + /// otherwise releases it. + pub(super) fn release_resource(&mut self, resource: &Resource, d3d12_resource: ID3D12Resource) { + let quarantine = self.stomp.as_ref().is_some_and(|s| s.settings.quarantine); + match (resource.stomp, &resource.allocation) { + (Some(layout), Some(allocation)) if quarantine => { + let memory_type_index = allocation.memory_type_index; + if let Err(e) = self.quarantine(d3d12_resource, layout, memory_type_index) { + warn!( + "Stomp: could not quarantine `{}`, releasing instead: {e}", + resource.name + ); + } + } + _ => drop(d3d12_resource), + } + } + + fn wait_for_mappings(&mut self) -> Result<()> { + let state = self + .stomp + .as_mut() + .expect("stomp state exists when stomping"); + state.fence_value += 1; + unsafe { + state + .settings + .queue + .Signal(&state.fence, state.fence_value) + .map_err(|e| { + AllocationError::Internal(format!("ID3D12CommandQueue::Signal failed: {e}")) + })?; + // A null event blocks the calling thread until the fence reaches the value. + state + .fence + .SetEventOnCompletion(state.fence_value, HANDLE::default()) + .map_err(|e| { + AllocationError::Internal(format!( + "ID3D12Fence::SetEventOnCompletion failed: {e}" + )) + }) + } + } + + /// Map `tiles` tiles starting at tile `first` of `resource` to `heap` at `heap_tile`. + /// With `reuse_single_tile` every tile maps to the same heap tile (the guard). + fn map_tiles( + &self, + resource: &ID3D12Resource, + first: u32, + tiles: u32, + heap: &ID3D12Heap, + heap_tile: u32, + reuse_single_tile: bool, + ) { + let queue = &self + .stomp + .as_ref() + .expect("stomp state exists") + .settings + .queue; + let coordinate = D3D12_TILED_RESOURCE_COORDINATE { + X: first, + Y: 0, + Z: 0, + Subresource: 0, + }; + let region = D3D12_TILE_REGION_SIZE { + NumTiles: tiles, + ..Default::default() + }; + let range_flags = if reuse_single_tile { + D3D12_TILE_RANGE_FLAG_REUSE_SINGLE_TILE + } else { + D3D12_TILE_RANGE_FLAG_NONE + }; + unsafe { + queue.UpdateTileMappings( + resource, + 1, + Some(&coordinate), + Some(®ion), + heap, + 1, + Some(&range_flags), + Some(&heap_tile), + Some(&tiles), + D3D12_TILE_MAPPING_FLAG_NONE, + ) + }; + } + + fn create_reserved( + &self, + resource_desc: &D3D12_RESOURCE_DESC, + desc: &ResourceCreateDesc<'_>, + ) -> Result { + let clear_value: Option<*const D3D12_CLEAR_VALUE> = + desc.clear_value.map(|v| -> *const _ { v }); + let mut result: Option = None; + if let Err(e) = unsafe { + match (&self.device, desc.initial_state_or_layout) { + (_, ResourceStateOrBarrierLayout::ResourceState(_)) + if !desc.castable_formats.is_empty() => + { + return Err(AllocationError::CastableFormatsRequiresEnhancedBarriers) + } + ( + ID3D12DeviceVersion::Device12(device), + ResourceStateOrBarrierLayout::BarrierLayout(initial_layout), + ) => device.CreateReservedResource2( + resource_desc, + initial_layout, + clear_value, + None, + Some(desc.castable_formats), + &mut result, + ), + (_, ResourceStateOrBarrierLayout::BarrierLayout(_)) + if !desc.castable_formats.is_empty() => + { + return Err(AllocationError::CastableFormatsRequiresAtLeastDevice12) + } + ( + ID3D12DeviceVersion::Device10(device), + ResourceStateOrBarrierLayout::BarrierLayout(initial_layout), + ) => device.CreateReservedResource2( + resource_desc, + initial_layout, + clear_value, + None, + None, + &mut result, + ), + (_, ResourceStateOrBarrierLayout::BarrierLayout(_)) => { + return Err(AllocationError::BarrierLayoutNeedsDevice10) + } + (device, ResourceStateOrBarrierLayout::ResourceState(initial_state)) => device + .CreateReservedResource(resource_desc, initial_state, clear_value, &mut result), + } + } { + if e.code() == DXGI_ERROR_DEVICE_REMOVED { + return Err(AllocationError::Internal(format!( + "ID3D12Device::CreateReservedResource DEVICE_REMOVED: {:?}", + unsafe { self.device.GetDeviceRemovedReason() } + ))); + } + return Err(AllocationError::Internal(format!( + "ID3D12Device::CreateReservedResource failed: {e}" + ))); + } + + Ok(result.expect("Allocation succeeded but no resource was returned?")) + } + + /// Stomp path of [`Self::create_resource()`]: reserved resource, payload tiles mapped to + /// sub-allocator memory, guard tiles mapped to the never-resident heap. + fn create_resource_stomp(&mut self, desc: &ResourceCreateDesc<'_>) -> Result { + let settings = &self.stomp.as_ref().expect("stomp state exists").settings; + let is_buffer = desc.resource_desc.Dimension == D3D12_RESOURCE_DIMENSION_BUFFER; + // Textures get no guard tiles: no unbounded write path exists for them. + let (head_guard, tail_guard) = if is_buffer { + (settings.head_guard, settings.tail_guard) + } else { + (false, false) + }; + + let mut resource_desc = *desc.resource_desc; + resource_desc.Alignment = 0; + let (mut payload_tiles, offset) = if is_buffer { + let (tiles, offset) = stomp_geometry( + resource_desc.Width, + tail_guard, + head_guard, + settings.payload_alignment, + ); + resource_desc.Width = + u64::from(u32::from(head_guard) + tiles + u32::from(tail_guard)) * TILE; + (tiles, offset) + } else { + resource_desc.Layout = D3D12_TEXTURE_LAYOUT_64KB_UNDEFINED_SWIZZLE; + (0, 0) + }; + + let resource = self.create_reserved(&resource_desc, desc)?; + if !is_buffer { + let mut num_subresource_tilings = 0; + let mut subresource_tiling = D3D12_SUBRESOURCE_TILING::default(); + unsafe { + self.device.GetResourceTiling( + &resource, + Some(&mut payload_tiles), + None, + None, + Some(&mut num_subresource_tilings), + 0, + &mut subresource_tiling, + ) + }; + } + if payload_tiles == 0 { + return Err(AllocationError::InvalidAllocationCreateDesc); + } + let layout = StompLayout { + // Textures have no VA; asking would only raise a debug layer warning. + resource_va: if is_buffer { + unsafe { resource.GetGPUVirtualAddress() } + } else { + 0 + }, + offset, + payload_tiles, + head_guard, + tail_guard, + }; + + let allocation = self.allocate(&AllocationCreateDesc { + name: desc.name, + location: desc.memory_location, + size: u64::from(payload_tiles) * TILE, + alignment: TILE, + resource_category: desc.resource_category, + })?; + + let mapped = (|| { + self.map_tiles( + &resource, + layout.head_tiles(), + payload_tiles, + &allocation.heap, + (allocation.offset / TILE) as u32, + false, + ); + if head_guard || tail_guard { + let guard_heap = self.memory_types[allocation.memory_type_index] + .guard_heap(&self.device)? + .clone(); + if head_guard { + self.map_tiles(&resource, 0, 1, &guard_heap, 0, true); + } + if tail_guard { + self.map_tiles( + &resource, + layout.head_tiles() + payload_tiles, + 1, + &guard_heap, + 0, + true, + ); + } + } + self.wait_for_mappings() + })(); + if let Err(e) = mapped { + drop(resource); + self.free(allocation)?; + return Err(e); + } + + // Best effort: this is the name DRED reports on a page fault. + let wide_name: Vec = desc.name.encode_utf16().chain(Some(0)).collect(); + let _ = unsafe { resource.SetName(PCWSTR::from_raw(wide_name.as_ptr())) }; + + self.stomp.as_mut().expect("stomp state exists").guarded += 1; + Ok(Resource { + name: desc.name.into(), + allocation: Some(allocation), + resource: Some(resource), + size: u64::from(payload_tiles) * TILE, + memory_location: desc.memory_location, + memory_type_index: None, + stomp: Some(layout), + offset, + }) + } + + /// Remap every tile of a freed stomp resource to the never-resident heap and keep the + /// resource alive, so stale descriptors and stale GPU VAs fault instead of reading reused + /// memory. + fn quarantine( + &mut self, + resource: ID3D12Resource, + layout: StompLayout, + memory_type_index: usize, + ) -> Result<()> { + let guard_heap = self.memory_types[memory_type_index] + .guard_heap(&self.device)? + .clone(); + self.map_tiles(&resource, 0, layout.total_tiles(), &guard_heap, 0, true); + self.wait_for_mappings()?; + self.stomp + .as_mut() + .expect("stomp state exists") + .quarantine + .push(resource); + Ok(()) + } +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::{ + d3d12::{HeapCategory, ResourceCategory, ResourceType}, + MemoryLocation, + }; + use windows::Win32::Graphics::Dxgi::Common::{DXGI_FORMAT_UNKNOWN, DXGI_SAMPLE_DESC}; + + const BUFFER_DESC: D3D12_RESOURCE_DESC = D3D12_RESOURCE_DESC { + Dimension: D3D12_RESOURCE_DIMENSION_BUFFER, + Alignment: 0, + Width: 4000, + Height: 1, + DepthOrArraySize: 1, + MipLevels: 1, + Format: DXGI_FORMAT_UNKNOWN, + SampleDesc: DXGI_SAMPLE_DESC { + Count: 1, + Quality: 0, + }, + Layout: D3D12_TEXTURE_LAYOUT_ROW_MAJOR, + Flags: D3D12_RESOURCE_FLAG_NONE, + }; + + fn desc(category: ResourceCategory) -> ResourceCreateDesc<'static> { + ResourceCreateDesc { + name: "test", + memory_location: MemoryLocation::GpuOnly, + resource_category: category, + resource_desc: &BUFFER_DESC, + castable_formats: &[], + clear_value: None, + initial_state_or_layout: ResourceStateOrBarrierLayout::ResourceState( + D3D12_RESOURCE_STATE_COMMON, + ), + resource_type: &ResourceType::Placed, + stomp: None, + } + } + + #[test] + fn geometry_table() { + // (size, tail, head, payload_alignment) -> (payload_tiles, offset) + let cases = [ + ((4000, true, false, Some(8)), (1, 61536)), + ((4000, true, false, None), (1, 0)), + ((4000, true, true, Some(256)), (1, 65536 + 61440)), + ((4000, false, true, Some(8)), (1, 65536)), + ((65536, true, false, Some(8)), (1, 0)), + ((65537, true, false, Some(8)), (2, 65528)), + ((70000, true, false, Some(256)), (2, 60928)), + ((1, true, false, Some(65536)), (1, 0)), + ((1, true, false, Some(8)), (1, 65528)), + ]; + for ((size, tail, head, alignment), expected) in cases { + assert_eq!( + stomp_geometry(size, tail, head, alignment), + expected, + "size {size} tail {tail} head {head} alignment {alignment:?}" + ); + } + } + + #[test] + fn geometry_properties() { + let mut rng = 1; + for _ in 0..10_000 { + let size = xorshift64(&mut rng) % (8 << 20) + 1; + let alignment = [8, 16, 256, 4096, 65536][(xorshift64(&mut rng) % 5) as usize]; + let head = xorshift64(&mut rng) % 2 == 0; + let (tiles, offset) = stomp_geometry(size, true, head, Some(alignment)); + let payload_end = u64::from(u32::from(head) + tiles) * TILE; + assert_eq!(offset % alignment, 0); + assert!(offset >= u64::from(head) * TILE); + assert!(offset + size <= payload_end); + assert!(payload_end - (offset + size) < alignment); + + let (tiles, offset) = stomp_geometry(size, true, head, None); + assert_eq!(offset, u64::from(head) * TILE); + assert!(offset + size <= u64::from(u32::from(head) + tiles) * TILE); + } + } + + fn draws(probability: f32, seed: u64, n: usize) -> Vec { + let mut rng = seed | 1; + let d = desc(ResourceCategory::Buffer); + (0..n) + .map(|_| stomp_decide(StompMode::Random { probability }, None, &mut rng, &d)) + .collect() + } + + #[test] + fn random_bounds() { + assert!(draws(0.0, 7, 10_000).iter().all(|&b| !b)); + assert!(draws(1.0, 7, 10_000).iter().all(|&b| b)); + let hits = draws(0.5, 7, 100_000).iter().filter(|&&b| b).count(); + assert!((45_000..=55_000).contains(&hits), "{hits}"); + } + + #[test] + fn random_is_deterministic() { + assert_eq!(draws(0.5, 42, 1000), draws(0.5, 42, 1000)); + } + + #[test] + fn xorshift_never_sticks_at_zero() { + // Seed 0 is what `Allocator::new` remaps with `| 1`. + let seed: u64 = 0; + let mut state = seed | 1; + for _ in 0..1000 { + xorshift64(&mut state); + assert_ne!(state, 0); + } + } + + #[test] + fn decision_matrix() { + let mut rng = 1; + let buffer = desc(ResourceCategory::Buffer); + let texture = desc(ResourceCategory::OtherTexture); + let filter = StompMode::Filter(|d| d.resource_category == ResourceCategory::Buffer); + + assert!(stomp_decide(StompMode::All, None, &mut rng, &buffer)); + assert!(!stomp_decide( + StompMode::All, + Some(false), + &mut rng, + &buffer + )); + assert!(!stomp_decide(StompMode::OptIn, None, &mut rng, &buffer)); + assert!(stomp_decide( + StompMode::OptIn, + Some(true), + &mut rng, + &buffer + )); + assert!(stomp_decide(filter, None, &mut rng, &buffer)); + assert!(!stomp_decide(filter, None, &mut rng, &texture)); + assert!(stomp_decide(filter, Some(true), &mut rng, &texture)); + assert!(stomp_decide( + StompMode::Random { probability: 0.0 }, + Some(true), + &mut rng, + &buffer + )); + } + + #[test] + fn settings_validation() { + let invalid = |r: Result<()>| matches!(r, Err(AllocationError::InvalidStompSettings(_))); + // `StompSettings::new` defaults. + assert!(validate_stomp_settings(StompMode::All, true, false, None).is_ok()); + assert!(validate_stomp_settings(StompMode::All, true, false, Some(256)).is_ok()); + assert!(validate_stomp_settings(StompMode::All, true, false, Some(65536)).is_ok()); + for probability in [1.5, -0.1, f32::NAN] { + assert!(invalid(validate_stomp_settings( + StompMode::Random { probability }, + true, + false, + None + ))); + } + assert!(invalid(validate_stomp_settings( + StompMode::All, + false, + false, + None + ))); + for alignment in [0, 3, 131072] { + assert!(invalid(validate_stomp_settings( + StompMode::All, + true, + false, + Some(alignment) + ))); + } + } + + #[test] + fn heap_flags_mapping() { + assert_eq!(heap_flags(HeapCategory::All), D3D12_HEAP_FLAG_NONE); + assert_eq!( + heap_flags(HeapCategory::Buffer), + D3D12_HEAP_FLAG_ALLOW_ONLY_BUFFERS + ); + assert_eq!( + heap_flags(HeapCategory::RtvDsvTexture), + D3D12_HEAP_FLAG_ALLOW_ONLY_RT_DS_TEXTURES + ); + assert_eq!( + heap_flags(HeapCategory::OtherTexture), + D3D12_HEAP_FLAG_ALLOW_ONLY_NON_RT_DS_TEXTURES + ); + } +} diff --git a/src/lib.rs b/src/lib.rs index 2eef751b..301be077 100644 --- a/src/lib.rs +++ b/src/lib.rs @@ -84,12 +84,18 @@ //! device: ID3D12DeviceVersion::Device(device), //! debug_settings: Default::default(), //! allocation_sizes: Default::default(), +//! stomp: None, //! }); //! # } //! # #[cfg(not(feature = "d3d12"))] //! # fn main() {} //! ``` //! +//! For development builds, `stomp: Some(StompSettings::default())` turns on stomp detection +//! for resources created with `Allocator::create_resource()`: each resource is packed against +//! a never-resident guard page so that a GPU write past its end page-faults and DRED names the +//! resource. See `d3d12::StompSettings` for modes and limitations. +//! //! # Simple d3d12 allocation example //! //! ```no_run @@ -104,6 +110,7 @@ //! # device: ID3D12DeviceVersion::Device(device), //! # debug_settings: Default::default(), //! # allocation_sizes: Default::default(), +//! # stomp: None, //! # }).unwrap(); //! //! let buffer_desc = Direct3D12::D3D12_RESOURCE_DESC { diff --git a/src/result.rs b/src/result.rs index 50bfff33..aae5234a 100644 --- a/src/result.rs +++ b/src/result.rs @@ -22,6 +22,8 @@ pub enum AllocationError { CastableFormatsRequiresEnhancedBarriers, #[error("Castable formats require at least `Device12`")] CastableFormatsRequiresAtLeastDevice12, + #[error("Invalid StompSettings: {0}")] + InvalidStompSettings(String), } pub type Result = ::core::result::Result; diff --git a/tests/d3d12_stomp.rs b/tests/d3d12_stomp.rs new file mode 100644 index 00000000..705ed13e --- /dev/null +++ b/tests/d3d12_stomp.rs @@ -0,0 +1,898 @@ +//! Device-backed tests for D3D12 stomp mode. Runs on real hardware when present, else on +//! WARP, so it also runs on CI. WARP does not page-fault; the fault tests live in +//! `d3d12_stomp_fault.rs`. +#![cfg(all(windows, feature = "d3d12"))] + +use std::mem::size_of; + +use gpu_allocator::{ + d3d12::{ + AllocationCreateDesc, Allocator, AllocatorCreateDesc, ID3D12DeviceVersion, Resource, + ResourceCategory, ResourceCreateDesc, ResourceStateOrBarrierLayout, ResourceType, + StompMode, StompSettings, + }, + AllocationError, MemoryLocation, +}; +use windows::{ + core::Interface, + Win32::Graphics::{ + Direct3D::D3D_FEATURE_LEVEL_11_0, + Direct3D12::*, + Dxgi::{ + Common::{ + DXGI_FORMAT, DXGI_FORMAT_R8G8B8A8_UNORM, DXGI_FORMAT_UNKNOWN, DXGI_SAMPLE_DESC, + }, + CreateDXGIFactory2, IDXGIAdapter, IDXGIAdapter1, IDXGIFactory6, + DXGI_ADAPTER_FLAG_SOFTWARE, DXGI_CREATE_FACTORY_FLAGS, DXGI_ERROR_NOT_FOUND, + }, + }, +}; + +const TILE: u64 = D3D12_TILED_RESOURCE_TILE_SIZE_IN_BYTES as u64; + +struct TestDevice { + device: ID3D12Device, +} + +/// Debug layer on, first hardware adapter, else WARP. Panics without any device: a silent +/// skip would hide breakage. +fn make_device() -> TestDevice { + let mut debug: Option = None; + if unsafe { D3D12GetDebugInterface(&mut debug) }.is_ok() { + unsafe { debug.unwrap().EnableDebugLayer() }; + } + + let factory: IDXGIFactory6 = + unsafe { CreateDXGIFactory2(DXGI_CREATE_FACTORY_FLAGS(0)) }.expect("DXGI factory"); + + for idx in 0.. { + let adapter: IDXGIAdapter1 = match unsafe { factory.EnumAdapters1(idx) } { + Ok(a) => a, + Err(e) if e.code() == DXGI_ERROR_NOT_FOUND => break, + Err(e) => panic!("EnumAdapters1: {e}"), + }; + let desc = unsafe { adapter.GetDesc1() }.unwrap(); + if desc.Flags & DXGI_ADAPTER_FLAG_SOFTWARE.0 as u32 != 0 { + continue; + } + let mut device: Option = None; + if unsafe { D3D12CreateDevice(&adapter, D3D_FEATURE_LEVEL_11_0, &mut device) }.is_ok() { + println!("adapter: {}", wide(&desc.Description)); + return TestDevice { + device: device.unwrap(), + }; + } + } + + let warp: IDXGIAdapter = unsafe { factory.EnumWarpAdapter() }.expect("WARP adapter"); + let mut device: Option = None; + unsafe { D3D12CreateDevice(&warp, D3D_FEATURE_LEVEL_11_0, &mut device) } + .expect("no hardware adapter and WARP device creation failed"); + println!("adapter: WARP"); + TestDevice { + device: device.unwrap(), + } +} + +fn wide(s: &[u16]) -> String { + let len = s.iter().position(|&c| c == 0).unwrap_or(s.len()); + String::from_utf16_lossy(&s[..len]) +} + +/// Fail on any ERROR or CORRUPTION message the debug layer stored for this device. +/// Warnings are printed so mapping onto the never-resident heap can be inspected. +fn assert_no_d3d12_errors(device: &ID3D12Device) { + let Ok(queue) = device.cast::() else { + println!("no ID3D12InfoQueue, debug layer unavailable"); + return; + }; + let mut errors = Vec::new(); + for i in 0..unsafe { queue.GetNumStoredMessages() } { + let mut len = 0usize; + let _ = unsafe { queue.GetMessage(i, None, &mut len) }; + // u64 storage keeps the D3D12_MESSAGE header aligned. + let mut buf = vec![0u64; (len.max(size_of::()) + 7) / 8]; + let msg = buf.as_mut_ptr().cast::(); + if unsafe { queue.GetMessage(i, Some(msg), &mut len) }.is_err() { + continue; + } + let msg = unsafe { &*msg }; + let text = unsafe { std::ffi::CStr::from_ptr(msg.pDescription.cast()) } + .to_string_lossy() + .into_owned(); + if msg.Severity == D3D12_MESSAGE_SEVERITY_ERROR + || msg.Severity == D3D12_MESSAGE_SEVERITY_CORRUPTION + { + errors.push(text); + } else if msg.Severity == D3D12_MESSAGE_SEVERITY_WARNING { + println!("d3d12 warning: {text}"); + } + } + assert!( + errors.is_empty(), + "D3D12 debug layer errors:\n{}", + errors.join("\n") + ); +} + +fn tiled_tier(device: &ID3D12Device) -> D3D12_TILED_RESOURCES_TIER { + let mut options = D3D12_FEATURE_DATA_D3D12_OPTIONS::default(); + unsafe { + device.CheckFeatureSupport( + D3D12_FEATURE_D3D12_OPTIONS, + <*mut D3D12_FEATURE_DATA_D3D12_OPTIONS>::cast(&mut options), + size_of::() as u32, + ) + } + .unwrap(); + options.TiledResourcesTier +} + +fn buffer_desc(width: u64) -> D3D12_RESOURCE_DESC { + D3D12_RESOURCE_DESC { + Dimension: D3D12_RESOURCE_DIMENSION_BUFFER, + Alignment: 0, + Width: width, + Height: 1, + DepthOrArraySize: 1, + MipLevels: 1, + Format: DXGI_FORMAT_UNKNOWN, + SampleDesc: DXGI_SAMPLE_DESC { + Count: 1, + Quality: 0, + }, + Layout: D3D12_TEXTURE_LAYOUT_ROW_MAJOR, + Flags: D3D12_RESOURCE_FLAG_NONE, + } +} + +fn tex_desc( + dimension: D3D12_RESOURCE_DIMENSION, + width: u64, + height: u32, + depth: u16, + format: DXGI_FORMAT, + flags: D3D12_RESOURCE_FLAGS, + samples: u32, +) -> D3D12_RESOURCE_DESC { + D3D12_RESOURCE_DESC { + Dimension: dimension, + Alignment: 0, + Width: width, + Height: height, + DepthOrArraySize: depth, + MipLevels: 1, + Format: format, + SampleDesc: DXGI_SAMPLE_DESC { + Count: samples, + Quality: 0, + }, + Layout: D3D12_TEXTURE_LAYOUT_UNKNOWN, + Flags: flags, + } +} + +fn tex2d_desc(flags: D3D12_RESOURCE_FLAGS, samples: u32) -> D3D12_RESOURCE_DESC { + tex_desc( + D3D12_RESOURCE_DIMENSION_TEXTURE2D, + 256, + 256, + 1, + DXGI_FORMAT_R8G8B8A8_UNORM, + flags, + samples, + ) +} + +/// A dedicated copy queue for the tile mapping updates, as the docs recommend. +fn copy_queue(device: &ID3D12Device) -> ID3D12CommandQueue { + unsafe { + device.CreateCommandQueue(&D3D12_COMMAND_QUEUE_DESC { + Type: D3D12_COMMAND_LIST_TYPE_COPY, + ..Default::default() + }) + } + .unwrap() +} + +fn settings(td: &TestDevice) -> StompSettings { + StompSettings::new(copy_queue(&td.device)) +} + +fn try_allocator( + td: &TestDevice, + stomp: Option, +) -> Result { + Allocator::new(&AllocatorCreateDesc { + device: ID3D12DeviceVersion::Device(td.device.clone()), + debug_settings: Default::default(), + allocation_sizes: Default::default(), + stomp, + }) +} + +fn stomp_allocator(td: &TestDevice, settings: StompSettings) -> Allocator { + try_allocator(td, Some(settings)).unwrap() +} + +fn plain_allocator(td: &TestDevice) -> Allocator { + try_allocator(td, None).unwrap() +} + +fn create<'a>( + alloc: &mut Allocator, + name: &str, + desc: &D3D12_RESOURCE_DESC, + location: MemoryLocation, + stomp: Option, + resource_type: &ResourceType<'a>, + clear_value: Option<&D3D12_CLEAR_VALUE>, +) -> Resource { + alloc + .create_resource(&ResourceCreateDesc { + name, + memory_location: location, + resource_category: desc.into(), + resource_desc: desc, + castable_formats: &[], + clear_value, + initial_state_or_layout: ResourceStateOrBarrierLayout::ResourceState( + D3D12_RESOURCE_STATE_COMMON, + ), + resource_type, + stomp, + }) + .unwrap() +} + +fn create_tex(alloc: &mut Allocator, name: &str, desc: &D3D12_RESOURCE_DESC) -> Resource { + create( + alloc, + name, + desc, + MemoryLocation::GpuOnly, + None, + &ResourceType::Placed, + None, + ) +} + +fn create_buffer( + alloc: &mut Allocator, + name: &str, + width: u64, + location: MemoryLocation, + stomp: Option, +) -> Resource { + create( + alloc, + name, + &buffer_desc(width), + location, + stomp, + &ResourceType::Placed, + None, + ) +} + +fn va(r: &Resource) -> u64 { + unsafe { r.resource().GetGPUVirtualAddress() } +} + +// 1 +#[test] +fn plain_allocator_unchanged() { + let td = make_device(); + let mut alloc = plain_allocator(&td); + assert!(alloc.stomp_statistics().is_none()); + + let r = create_buffer(&mut alloc, "plain", 4000, MemoryLocation::GpuOnly, None); + assert!(r.allocation.is_some()); + assert!(r.stomp_layout().is_none()); + assert!(!r.is_stomp_guarded()); + assert_eq!(r.offset(), 0); + assert_eq!(unsafe { r.resource().GetDesc() }.Width, 4000); + assert!(alloc.capacity() > 0); + assert_eq!(alloc.generate_report().allocations.len(), 1); + alloc.free_resource(r).unwrap(); + assert_eq!(alloc.generate_report().allocations.len(), 0); + assert_no_d3d12_errors(&td.device); +} + +// 2 +#[test] +fn all_mode_guards_buffer() { + let td = make_device(); + let mut alloc = stomp_allocator(&td, settings(&td)); + let r = create_buffer(&mut alloc, "stomp", 4000, MemoryLocation::GpuOnly, None); + // Payload lives in a normal sub-allocator block. + assert!(r.allocation.is_some()); + assert!(r.stomp_layout().is_some()); + assert!(r.is_stomp_guarded()); + assert_eq!(alloc.generate_report().allocations.len(), 1); + let stats = alloc.stomp_statistics().unwrap(); + assert_eq!((stats.guarded, stats.unguarded_fallbacks), (1, 0)); + alloc.free_resource(r).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +// 3 +#[test] +fn width_and_size() { + let td = make_device(); + let mut alloc = stomp_allocator(&td, settings(&td)); + let r = create_buffer(&mut alloc, "tail", 4000, MemoryLocation::GpuOnly, None); + assert_eq!(unsafe { r.resource().GetDesc() }.Width, 2 * TILE); + assert_eq!(r.size, TILE); + assert_eq!(r.stomp_layout().unwrap().payload_tiles, 1); + alloc.free_resource(r).unwrap(); + + let mut both = stomp_allocator( + &td, + StompSettings { + head_guard: true, + ..settings(&td) + }, + ); + let r = create_buffer(&mut both, "head+tail", 65537, MemoryLocation::GpuOnly, None); + assert_eq!(unsafe { r.resource().GetDesc() }.Width, 4 * TILE); + assert_eq!(r.size, 2 * TILE); + assert_eq!(r.stomp_layout().unwrap().payload_tiles, 2); + both.free_resource(r).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +// 4 +#[test] +fn tail_guard_va() { + let td = make_device(); + let mut alloc = stomp_allocator(&td, settings(&td)); + let r = create_buffer(&mut alloc, "tail", 4000, MemoryLocation::GpuOnly, None); + let layout = r.stomp_layout().unwrap(); + assert_eq!(layout.resource_va, va(&r)); + assert_ne!(layout.resource_va, 0); + assert_eq!(layout.offset, 0); + assert_eq!(r.offset(), 0); + assert_eq!(layout.tail_guard_va(), Some(va(&r) + TILE)); + assert_eq!(layout.head_guard_va(), None); + alloc.free_resource(r).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +// 5 +#[test] +fn payload_alignment_packs_against_tail() { + let td = make_device(); + for (alignment, expected_offset) in [(None, 0), (Some(8), 61536), (Some(256), 61440)] { + let mut alloc = stomp_allocator( + &td, + StompSettings { + payload_alignment: alignment, + ..settings(&td) + }, + ); + let r = create_buffer(&mut alloc, "packed", 4000, MemoryLocation::GpuOnly, None); + let layout = r.stomp_layout().unwrap(); + assert_eq!(r.offset(), expected_offset, "{alignment:?}"); + assert_eq!(layout.offset, expected_offset); + let payload_end = va(&r) + r.offset() + 4000; + let tail = layout.tail_guard_va().unwrap(); + assert!(payload_end <= tail); + let slack = tail - payload_end; + match alignment { + Some(a) => assert!(slack < a, "slack {slack} >= {a}"), + None => assert_eq!(slack, TILE - 4000), + } + alloc.free_resource(r).unwrap(); + } + assert_no_d3d12_errors(&td.device); +} + +// 6 +#[test] +fn head_guard_va() { + let td = make_device(); + let mut alloc = stomp_allocator( + &td, + StompSettings { + head_guard: true, + ..settings(&td) + }, + ); + let r = create_buffer(&mut alloc, "head+tail", 4000, MemoryLocation::GpuOnly, None); + let layout = r.stomp_layout().unwrap(); + assert_eq!(r.offset(), TILE); + assert_eq!(layout.head_guard_va(), Some(va(&r))); + assert_eq!(layout.tail_guard_va(), Some(va(&r) + 2 * TILE)); + assert!(r.is_stomp_guarded()); + alloc.free_resource(r).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +// 7 +#[test] +fn head_only() { + let td = make_device(); + let mut alloc = stomp_allocator( + &td, + StompSettings { + head_guard: true, + tail_guard: false, + payload_alignment: Some(8), + ..settings(&td) + }, + ); + let r = create_buffer(&mut alloc, "head", 4000, MemoryLocation::GpuOnly, None); + let layout = r.stomp_layout().unwrap(); + // No tail guard, so payload_alignment has nothing to pack against. + assert_eq!(r.offset(), TILE); + assert_eq!(layout.head_guard_va(), Some(va(&r))); + assert_eq!(layout.tail_guard_va(), None); + assert_eq!(unsafe { r.resource().GetDesc() }.Width, 2 * TILE); + alloc.free_resource(r).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +fn count_stomped(td: &TestDevice, probability: f32, seed: u64, n: usize) -> Vec { + let mut alloc = stomp_allocator( + td, + StompSettings { + mode: StompMode::Random { probability }, + seed, + ..settings(td) + }, + ); + let resources: Vec<_> = (0..n) + .map(|i| { + create_buffer( + &mut alloc, + &format!("r{i}"), + 1024, + MemoryLocation::GpuOnly, + None, + ) + }) + .collect(); + let stomped = resources + .iter() + .map(|r| r.stomp_layout().is_some()) + .collect(); + for r in resources { + alloc.free_resource(r).unwrap(); + } + assert_no_d3d12_errors(&td.device); + stomped +} + +// 8 +#[test] +fn random_zero_guards_none() { + let td = make_device(); + assert!(count_stomped(&td, 0.0, 1, 50).iter().all(|&s| !s)); +} + +// 8 +#[test] +fn random_one_guards_all() { + let td = make_device(); + assert!(count_stomped(&td, 1.0, 1, 50).iter().all(|&s| s)); +} + +// 9 +#[test] +fn random_is_deterministic() { + let td = make_device(); + let a = count_stomped(&td, 0.5, 42, 100); + let b = count_stomped(&td, 0.5, 42, 100); + assert_eq!(a, b); + assert!( + a.iter().any(|&s| s) && a.iter().any(|&s| !s), + "0.5 should mix" + ); +} + +// 10 +#[test] +fn opt_in_mode() { + let td = make_device(); + let mut alloc = stomp_allocator( + &td, + StompSettings { + mode: StompMode::OptIn, + ..settings(&td) + }, + ); + let normal = create_buffer(&mut alloc, "none", 1024, MemoryLocation::GpuOnly, None); + let forced = create_buffer( + &mut alloc, + "some", + 1024, + MemoryLocation::GpuOnly, + Some(true), + ); + assert!(normal.stomp_layout().is_none()); + assert!(forced.stomp_layout().is_some()); + alloc.free_resource(normal).unwrap(); + alloc.free_resource(forced).unwrap(); + + let mut all = stomp_allocator(&td, settings(&td)); + let opted_out = create_buffer(&mut all, "out", 1024, MemoryLocation::GpuOnly, Some(false)); + assert!(opted_out.stomp_layout().is_none()); + all.free_resource(opted_out).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +// 11 +#[test] +fn filter_mode() { + let td = make_device(); + let mut alloc = stomp_allocator( + &td, + StompSettings { + mode: StompMode::Filter(|d| d.resource_category == ResourceCategory::Buffer), + ..settings(&td) + }, + ); + let buffer = create_buffer(&mut alloc, "buf", 1024, MemoryLocation::GpuOnly, None); + let tex = create_tex(&mut alloc, "tex", &tex2d_desc(D3D12_RESOURCE_FLAG_NONE, 1)); + assert!(buffer.stomp_layout().is_some()); + assert!(tex.stomp_layout().is_none()); + alloc.free_resource(buffer).unwrap(); + alloc.free_resource(tex).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +fn texture_stomped( + td: &TestDevice, + flags: D3D12_RESOURCE_FLAGS, + clear: Option<&D3D12_CLEAR_VALUE>, +) { + let mut alloc = stomp_allocator(td, settings(td)); + let desc = tex2d_desc(flags, 1); + let r = create( + &mut alloc, + "tex", + &desc, + MemoryLocation::GpuOnly, + None, + &ResourceType::Placed, + clear, + ); + let got = unsafe { r.resource().GetDesc() }; + assert_eq!((got.Width, got.Height, got.Flags), (256, 256, flags)); + assert_eq!(got.Layout, D3D12_TEXTURE_LAYOUT_64KB_UNDEFINED_SWIZZLE); + let layout = r.stomp_layout().unwrap(); + // Textures get no guard tiles, only quarantine. + assert!(!r.is_stomp_guarded()); + assert!(!layout.head_guard && !layout.tail_guard); + assert_eq!(layout.resource_va, 0); + assert_eq!(layout.offset, 0); + // 256x256 RGBA8 is exactly 4 tiles. + assert_eq!(layout.payload_tiles, 4); + assert_eq!(r.size, 4 * TILE); + assert!(r.allocation.is_some()); + assert_eq!(alloc.stomp_statistics().unwrap().guarded, 1); + alloc.free_resource(r).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +// 12 +#[test] +fn texture_2d_stomped() { + let td = make_device(); + texture_stomped(&td, D3D12_RESOURCE_FLAG_NONE, None); +} + +// 13 +#[test] +fn rt_texture_with_clear_value_stomped() { + let td = make_device(); + let clear = D3D12_CLEAR_VALUE { + Format: DXGI_FORMAT_R8G8B8A8_UNORM, + Anonymous: D3D12_CLEAR_VALUE_0 { + Color: [0.0, 0.5, 1.0, 1.0], + }, + }; + texture_stomped(&td, D3D12_RESOURCE_FLAG_ALLOW_RENDER_TARGET, Some(&clear)); +} + +// 14 +#[test] +fn msaa_falls_back() { + let td = make_device(); + let mut alloc = stomp_allocator(&td, settings(&td)); + let desc = tex2d_desc(D3D12_RESOURCE_FLAG_ALLOW_RENDER_TARGET, 4); + let r = create_tex(&mut alloc, "msaa", &desc); + assert!(r.allocation.is_some() && r.stomp_layout().is_none()); + let stats = alloc.stomp_statistics().unwrap(); + assert_eq!((stats.guarded, stats.unguarded_fallbacks), (0, 1)); + alloc.free_resource(r).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +// 15 +#[test] +fn texture_3d_tier_dependent() { + let td = make_device(); + let tier = tiled_tier(&td.device); + let mut alloc = stomp_allocator(&td, settings(&td)); + let desc = tex_desc( + D3D12_RESOURCE_DIMENSION_TEXTURE3D, + 64, + 64, + 64, + DXGI_FORMAT_R8G8B8A8_UNORM, + D3D12_RESOURCE_FLAG_NONE, + 1, + ); + let r = create_tex(&mut alloc, "tex3d", &desc); + let stats = alloc.stomp_statistics().unwrap(); + if tier.0 >= D3D12_TILED_RESOURCES_TIER_3.0 { + println!("tiled tier {}: 3D stomped", tier.0); + assert!(r.stomp_layout().is_some()); + // 64^3 RGBA8 = 1MB = 16 tiles. + assert_eq!(r.stomp_layout().unwrap().payload_tiles, 16); + assert_eq!((stats.guarded, stats.unguarded_fallbacks), (1, 0)); + } else { + println!("tiled tier {}: 3D falls back", tier.0); + assert!(r.stomp_layout().is_none() && r.allocation.is_some()); + assert_eq!((stats.guarded, stats.unguarded_fallbacks), (0, 1)); + } + alloc.free_resource(r).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +// 16 +#[test] +fn cpu_heaps_fall_back_and_stay_mappable() { + // Reserved resources cannot be `Map()`ed (E_INVALIDARG), so CPU-visible resources are + // refused by the stomp path and must still map through the normal path. + let td = make_device(); + let mut alloc = stomp_allocator(&td, settings(&td)); + for (i, (name, location)) in [ + ("upload", MemoryLocation::CpuToGpu), + ("readback", MemoryLocation::GpuToCpu), + ] + .into_iter() + .enumerate() + { + let r = create_buffer(&mut alloc, name, 4000, location, None); + assert!(r.stomp_layout().is_none() && r.allocation.is_some()); + assert_eq!(alloc.stomp_statistics().unwrap().unguarded_fallbacks, i + 1); + let mut ptr = std::ptr::null_mut(); + unsafe { r.resource().Map(0, None, Some(&mut ptr)) }.unwrap(); + assert!(!ptr.is_null()); + let bytes = unsafe { std::slice::from_raw_parts_mut(ptr.cast::(), 4000) }; + bytes.fill(0xAB); + assert_eq!(bytes[3999], 0xAB); + unsafe { r.resource().Unmap(0, None) }; + alloc.free_resource(r).unwrap(); + } + assert_no_d3d12_errors(&td.device); +} + +// 17 +#[test] +fn committed_request_takes_stomp_path() { + let td = make_device(); + let mut alloc = stomp_allocator(&td, settings(&td)); + let heap_properties = D3D12_HEAP_PROPERTIES { + Type: D3D12_HEAP_TYPE_DEFAULT, + ..Default::default() + }; + let r = create( + &mut alloc, + "committed", + &buffer_desc(4000), + MemoryLocation::GpuOnly, + None, + &ResourceType::Committed { + heap_properties: &heap_properties, + heap_flags: D3D12_HEAP_FLAG_NONE, + }, + None, + ); + assert!(r.is_stomp_guarded()); + assert!(r.allocation.is_some()); + for s in alloc.committed_statistics() { + assert_eq!(s.num_allocations, 0); + } + alloc.free_resource(r).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +fn free_many(td: &TestDevice, quarantine: bool) { + let mut alloc = stomp_allocator( + td, + StompSettings { + quarantine, + ..settings(td) + }, + ); + let resources: Vec<_> = (0..20) + .map(|i| { + create_buffer( + &mut alloc, + &format!("b{i}"), + 4000, + MemoryLocation::GpuOnly, + None, + ) + }) + .collect(); + let tex = create_tex(&mut alloc, "tex", &tex2d_desc(D3D12_RESOURCE_FLAG_NONE, 1)); + assert_eq!(alloc.generate_report().allocations.len(), 21); + assert_eq!(alloc.stomp_statistics().unwrap().quarantined, 0); + for r in resources { + alloc.free_resource(r).unwrap(); + } + alloc.free_resource(tex).unwrap(); + assert_eq!(alloc.generate_report().allocations.len(), 0); + let stats = alloc.stomp_statistics().unwrap(); + assert_eq!(stats.guarded, 21); + assert_eq!(stats.quarantined, if quarantine { 21 } else { 0 }); + for s in alloc.committed_statistics() { + assert_eq!((s.num_allocations, s.total_size), (0, 0)); + } + drop(alloc); + assert_no_d3d12_errors(&td.device); +} + +// 18 +#[test] +fn free_with_quarantine() { + let td = make_device(); + free_many(&td, true); +} + +// 19 +#[test] +fn free_without_quarantine() { + let td = make_device(); + free_many(&td, false); +} + +fn churn(td: &TestDevice, iterations: usize) { + let mut alloc = stomp_allocator(td, settings(td)); + let mut ring = std::collections::VecDeque::new(); + let mut seed = 0x1234_5678_u64; + for i in 0..iterations { + seed ^= seed << 13; + seed ^= seed >> 7; + seed ^= seed << 17; + let width = 1 + seed % (4 * 1024 * 1024); + let r = create_buffer( + &mut alloc, + &format!("churn{i}"), + width, + MemoryLocation::GpuOnly, + None, + ); + assert!(r.is_stomp_guarded(), "iteration {i}"); + let layout = r.stomp_layout().unwrap(); + assert_eq!( + layout.tail_guard_va(), + Some(va(&r) + u64::from(layout.payload_tiles) * TILE) + ); + ring.push_back(r); + if ring.len() > 32 { + alloc.free_resource(ring.pop_front().unwrap()).unwrap(); + } + } + for r in ring { + alloc.free_resource(r).unwrap(); + } + let stats = alloc.stomp_statistics().unwrap(); + println!("churn: {stats:?}"); + assert_eq!( + (stats.guarded, stats.unguarded_fallbacks, stats.quarantined), + (iterations, 0, iterations) + ); + assert_eq!(alloc.generate_report().allocations.len(), 0); +} + +// 20 +#[test] +fn churn_all_guarded() { + let td = make_device(); + churn(&td, 500); + assert_no_d3d12_errors(&td.device); +} + +// 21 +#[test] +fn churn_from_threads() { + let handles: Vec<_> = (0..4) + .map(|_| { + std::thread::spawn(|| { + let td = make_device(); + churn(&td, 100); + assert_no_d3d12_errors(&td.device); + }) + }) + .collect(); + for h in handles { + h.join().unwrap(); + } +} + +// 22 +#[test] +fn drop_order_is_safe() { + let td = make_device(); + { + let mut alloc = stomp_allocator(&td, settings(&td)); + let r = create_buffer(&mut alloc, "freed", 4000, MemoryLocation::GpuOnly, None); + alloc.free_resource(r).unwrap(); + } + { + let mut alloc = stomp_allocator(&td, settings(&td)); + let live = create_buffer(&mut alloc, "live", 4000, MemoryLocation::GpuOnly, None); + drop(alloc); + drop(live); // warns "not freed", must not crash + } + assert_no_d3d12_errors(&td.device); +} + +// 23 +#[test] +fn rename_and_report_unaffected() { + let td = make_device(); + let mut alloc = stomp_allocator(&td, settings(&td)); + let desc = buffer_desc(1024); + let mut allocation = alloc + .allocate(&AllocationCreateDesc::from_d3d12_resource_desc( + &td.device, + &desc, + "before", + MemoryLocation::GpuOnly, + )) + .unwrap(); + alloc.rename_allocation(&mut allocation, "after").unwrap(); + assert_eq!(alloc.generate_report().allocations[0].name, "after"); + alloc.report_memory_leaks(log::Level::Debug); + alloc.free(allocation).unwrap(); + assert_no_d3d12_errors(&td.device); +} + +// 24 +#[test] +fn invalid_settings_rejected() { + let td = make_device(); + let invalid = |s: StompSettings| { + matches!( + try_allocator(&td, Some(s)), + Err(AllocationError::InvalidStompSettings(_)) + ) + }; + assert!(invalid(StompSettings { + mode: StompMode::Random { probability: 1.5 }, + ..settings(&td) + })); + assert!(invalid(StompSettings { + tail_guard: false, + head_guard: false, + ..settings(&td) + })); + assert!(invalid(StompSettings { + payload_alignment: Some(3), + ..settings(&td) + })); + assert!(try_allocator(&td, Some(settings(&td))).is_ok()); + assert_no_d3d12_errors(&td.device); +} + +// 25 +#[test] +fn tiled_tier_unsupported_rejected() { + let td = make_device(); + let tier = tiled_tier(&td.device); + if tier.0 >= D3D12_TILED_RESOURCES_TIER_1.0 { + println!("tiled tier {}, nothing to test", tier.0); + return; + } + assert!(matches!( + try_allocator(&td, Some(settings(&td))), + Err(AllocationError::InvalidStompSettings(_)) + )); +} diff --git a/tests/d3d12_stomp_fault.rs b/tests/d3d12_stomp_fault.rs new file mode 100644 index 00000000..4d95e6f2 --- /dev/null +++ b/tests/d3d12_stomp_fault.rs @@ -0,0 +1,693 @@ +//! Hardware-only fault tests for D3D12 stomp mode. See `docs/d3d12-stomp-allocator.md` +//! section 6.3. +//! +//! Every test removes the device, so each creates its own device and the file must run +//! single-threaded. All tests are `#[ignore]` and additionally return early unless +//! `GPU_ALLOCATOR_STOMP_FAULT_TESTS=1` is set. WARP does not page-fault, so the tests also +//! return early when no hardware adapter is present. +//! +//! ```text +//! GPU_ALLOCATOR_STOMP_FAULT_TESTS=1 cargo test --features d3d12 --test d3d12_stomp_fault -- --ignored --test-threads=1 --nocapture +//! ``` +//! +//! Findings on NVIDIA (RTX 5070 Ti, September 2026): a root descriptor access to a guard +//! tile removes the device with `DXGI_ERROR_DEVICE_HUNG`, but DRED reports page fault VA 0 +//! and no allocation names, so attribution is informational only. Oversized +//! `CopyBufferRegion` is rejected by the runtime at `Close()` with `E_INVALIDARG`, so copies +//! cannot be used to overrun; the tests touch guard tiles through root UAV/SRV descriptors, +//! which the hardware does not bounds check. +//! +//! The allocator's tile mappings go through a COPY queue while every dispatch here runs on a +//! DIRECT queue, so each test also proves the CPU wait after mapping makes the resource +//! usable on another queue. +#![cfg(all(windows, feature = "d3d12"))] + +use gpu_allocator::{ + d3d12::{ + Allocator, AllocatorCreateDesc, ID3D12DeviceVersion, Resource, ResourceCategory, + ResourceCreateDesc, ResourceStateOrBarrierLayout, ResourceType, StompSettings, + }, + MemoryLocation, +}; +use windows::{ + core::{Interface, Result}, + Win32::{ + Foundation::HANDLE, + Graphics::{ + Direct3D::D3D_FEATURE_LEVEL_11_0, + Direct3D12::*, + Dxgi::{ + Common::{DXGI_FORMAT_R8G8B8A8_UNORM, DXGI_FORMAT_UNKNOWN, DXGI_SAMPLE_DESC}, + CreateDXGIFactory2, IDXGIAdapter1, IDXGIFactory6, DXGI_ADAPTER_FLAG_SOFTWARE, + DXGI_ERROR_NOT_FOUND, + }, + }, + }, +}; + +const TILE: u64 = D3D12_TILED_RESOURCE_TILE_SIZE_IN_BYTES as u64; + +/// Source and compile commands in `tests/shaders/`. Root signatures are embedded. +const ROOT_SRV_READ_CS: &[u8] = include_bytes!("shaders/root_srv_read.dxil"); +const ROOT_UAV_WRITE_CS: &[u8] = include_bytes!("shaders/root_uav_write.dxil"); +const TABLE_SRV_TEXTURE_LOAD_CS: &[u8] = include_bytes!("shaders/table_srv_texture_load.dxil"); + +fn enabled() -> bool { + if std::env::var("GPU_ALLOCATOR_STOMP_FAULT_TESTS").as_deref() != Ok("1") { + eprintln!("skipped: set GPU_ALLOCATOR_STOMP_FAULT_TESTS=1 to run fault tests"); + return false; + } + true +} + +fn enable_dred() { + let mut settings: Option = None; + if unsafe { D3D12GetDebugInterface(&mut settings) }.is_ok() { + let settings = settings.unwrap(); + unsafe { + settings.SetPageFaultEnablement(D3D12_DRED_ENABLEMENT_FORCED_ON); + settings.SetAutoBreadcrumbsEnablement(D3D12_DRED_ENABLEMENT_FORCED_ON); + } + } +} + +/// Hardware device with DRED enabled, `None` when only WARP is available. Retries for a few +/// seconds: right after a device removal the adapter can refuse new devices while it resets. +fn hardware_device() -> Option { + enable_dred(); + let mut last_err = None; + for attempt in 0..20 { + if attempt > 0 { + std::thread::sleep(std::time::Duration::from_millis(500)); + } + let factory: IDXGIFactory6 = unsafe { CreateDXGIFactory2(Default::default()) }.ok()?; + for idx in 0.. { + let adapter: IDXGIAdapter1 = match unsafe { factory.EnumAdapters1(idx) } { + Ok(a) => a, + Err(e) if e.code() == DXGI_ERROR_NOT_FOUND => break, + Err(e) => panic!("EnumAdapters1: {e}"), + }; + let desc = unsafe { adapter.GetDesc1() }.ok()?; + if desc.Flags & DXGI_ADAPTER_FLAG_SOFTWARE.0 as u32 != 0 { + continue; + } + let mut device: Option = None; + match unsafe { D3D12CreateDevice(&adapter, D3D_FEATURE_LEVEL_11_0, &mut device) } { + Ok(()) => { + eprintln!( + "adapter: {}", + String::from_utf16_lossy(&desc.Description).trim_end_matches('\0') + ); + return device; + } + Err(e) => last_err = Some(e), + } + } + if last_err.is_none() { + break; // No hardware adapter at all, no point retrying. + } + } + match last_err { + Some(e) => eprintln!("skipped: hardware adapter refused device creation: {e}"), + None => eprintln!("skipped: no hardware adapter, WARP does not page-fault"), + } + None +} + +const CHILD_ENV: &str = "GPU_ALLOCATOR_STOMP_FAULT_CHILD"; + +/// Runs `body` in a child process. After the second device removal in one process the +/// adapter keeps refusing new devices with `DXGI_ERROR_DEVICE_HUNG`, so every test gets a +/// process of its own. In the child, `setup()` returns the device. +fn isolated(test_name: &str, body: fn()) { + if std::env::var_os(CHILD_ENV).is_some() { + body(); + return; + } + if !enabled() { + return; + } + let status = std::process::Command::new(std::env::current_exe().unwrap()) + .args([ + "--exact", + test_name, + "--ignored", + "--nocapture", + "--test-threads=1", + ]) + .env(CHILD_ENV, "1") + .status() + .expect("spawn child test process"); + assert!( + status.success(), + "child process for `{test_name}` failed: {status}" + ); +} + +/// `Some(device)` only inside an `isolated` child with hardware present. +fn setup() -> Option { + hardware_device() +} + +/// An `#[ignore]`d test whose body runs in its own process, see [`isolated`]. +macro_rules! fault_test { + ($name:ident, $body:block) => { + #[test] + #[ignore] + fn $name() { + isolated(stringify!($name), || $body); + } + }; +} + +fn create_queue(device: &ID3D12Device, ty: D3D12_COMMAND_LIST_TYPE) -> ID3D12CommandQueue { + unsafe { + device.CreateCommandQueue(&D3D12_COMMAND_QUEUE_DESC { + Type: ty, + ..Default::default() + }) + } + .expect("CreateCommandQueue") +} + +/// Default stomp settings on a dedicated COPY queue; adjust fields on the result. +fn settings(device: &ID3D12Device) -> StompSettings { + StompSettings::new(create_queue(device, D3D12_COMMAND_LIST_TYPE_COPY)) +} + +fn stomp_allocator(device: &ID3D12Device, settings: StompSettings) -> Allocator { + Allocator::new(&AllocatorCreateDesc { + device: ID3D12DeviceVersion::Device(device.clone()), + debug_settings: Default::default(), + allocation_sizes: Default::default(), + stomp: Some(settings), + }) + .expect("Allocator::new") +} + +fn buffer_desc(width: u64, flags: D3D12_RESOURCE_FLAGS) -> D3D12_RESOURCE_DESC { + D3D12_RESOURCE_DESC { + Dimension: D3D12_RESOURCE_DIMENSION_BUFFER, + Alignment: 0, + Width: width, + Height: 1, + DepthOrArraySize: 1, + MipLevels: 1, + Format: DXGI_FORMAT_UNKNOWN, + SampleDesc: DXGI_SAMPLE_DESC { + Count: 1, + Quality: 0, + }, + Layout: D3D12_TEXTURE_LAYOUT_ROW_MAJOR, + Flags: flags, + } +} + +fn tex2d_desc(width: u64, height: u32) -> D3D12_RESOURCE_DESC { + D3D12_RESOURCE_DESC { + Dimension: D3D12_RESOURCE_DIMENSION_TEXTURE2D, + Alignment: 0, + Width: width, + Height: height, + DepthOrArraySize: 1, + MipLevels: 1, + Format: DXGI_FORMAT_R8G8B8A8_UNORM, + SampleDesc: DXGI_SAMPLE_DESC { + Count: 1, + Quality: 0, + }, + Layout: D3D12_TEXTURE_LAYOUT_UNKNOWN, + Flags: D3D12_RESOURCE_FLAG_NONE, + } +} + +fn create( + allocator: &mut Allocator, + name: &str, + desc: &D3D12_RESOURCE_DESC, + category: ResourceCategory, +) -> Resource { + allocator + .create_resource(&ResourceCreateDesc { + name, + memory_location: MemoryLocation::GpuOnly, + resource_category: category, + resource_desc: desc, + castable_formats: &[], + clear_value: None, + initial_state_or_layout: ResourceStateOrBarrierLayout::ResourceState( + D3D12_RESOURCE_STATE_COMMON, + ), + resource_type: &ResourceType::Placed, + stomp: None, + }) + .expect("create_resource") +} + +fn compute_pso( + device: &ID3D12Device, + blob: &[u8], +) -> Result<(ID3D12RootSignature, ID3D12PipelineState)> { + let root_signature: ID3D12RootSignature = unsafe { device.CreateRootSignature(0, blob) }?; + let pso = unsafe { + device.CreateComputePipelineState(&D3D12_COMPUTE_PIPELINE_STATE_DESC { + pRootSignature: std::mem::ManuallyDrop::new(Some(root_signature.clone())), + CS: D3D12_SHADER_BYTECODE { + pShaderBytecode: blob.as_ptr().cast(), + BytecodeLength: blob.len(), + }, + ..Default::default() + }) + }?; + Ok((root_signature, pso)) +} + +/// Direct queue plus fence, with root-descriptor read and write kernels. +struct Gpu { + device: ID3D12Device, + queue: ID3D12CommandQueue, + fence: ID3D12Fence, + value: u64, + read: (ID3D12RootSignature, ID3D12PipelineState), + write: (ID3D12RootSignature, ID3D12PipelineState), + texture_load: (ID3D12RootSignature, ID3D12PipelineState), + srv_heap: ID3D12DescriptorHeap, + scratch: ID3D12Resource, +} + +/// Outcome of one GPU submission. +#[derive(Debug)] +enum Outcome { + Completed, + /// `Close()` or `ExecuteCommandLists` refused the work; the device is still alive. + Rejected(windows::core::Error), + /// `GetDeviceRemovedReason` failed after the submission. + DeviceRemoved { + reason: windows::core::Error, + page_fault_va: u64, + dred_names: Vec, + }, +} + +impl Gpu { + fn new(device: &ID3D12Device) -> Self { + let queue = create_queue(device, D3D12_COMMAND_LIST_TYPE_DIRECT); + let fence = unsafe { device.CreateFence(0, D3D12_FENCE_FLAG_NONE) }.expect("CreateFence"); + let read = compute_pso(device, ROOT_SRV_READ_CS).expect("read pso"); + let write = compute_pso(device, ROOT_UAV_WRITE_CS).expect("write pso"); + let texture_load = + compute_pso(device, TABLE_SRV_TEXTURE_LOAD_CS).expect("texture load pso"); + let srv_heap = unsafe { + device.CreateDescriptorHeap(&D3D12_DESCRIPTOR_HEAP_DESC { + Type: D3D12_DESCRIPTOR_HEAP_TYPE_CBV_SRV_UAV, + NumDescriptors: 1, + Flags: D3D12_DESCRIPTOR_HEAP_FLAG_SHADER_VISIBLE, + NodeMask: 0, + }) + } + .expect("CreateDescriptorHeap"); + let mut scratch: Option = None; + unsafe { + device.CreateCommittedResource( + &D3D12_HEAP_PROPERTIES { + Type: D3D12_HEAP_TYPE_DEFAULT, + ..Default::default() + }, + D3D12_HEAP_FLAG_NONE, + &buffer_desc(256, D3D12_RESOURCE_FLAG_ALLOW_UNORDERED_ACCESS), + D3D12_RESOURCE_STATE_COMMON, + None, + &mut scratch, + ) + } + .expect("scratch"); + Self { + device: device.clone(), + queue, + fence, + value: 0, + read, + write, + texture_load, + srv_heap, + scratch: scratch.unwrap(), + } + } + + /// Records into a fresh list, executes, waits, then classifies what happened. + fn run(&mut self, record: impl FnOnce(&ID3D12GraphicsCommandList)) -> Outcome { + let submitted: Result<()> = (|| unsafe { + let allocator: ID3D12CommandAllocator = self + .device + .CreateCommandAllocator(D3D12_COMMAND_LIST_TYPE_DIRECT)?; + let list: ID3D12GraphicsCommandList = self.device.CreateCommandList( + 0, + D3D12_COMMAND_LIST_TYPE_DIRECT, + &allocator, + None, + )?; + record(&list); + list.Close()?; + self.queue.ExecuteCommandLists(&[Some(list.cast()?)]); + self.value += 1; + self.queue.Signal(&self.fence, self.value)?; + // A null event blocks until the fence reaches the value (or the device dies). + self.fence + .SetEventOnCompletion(self.value, HANDLE::default())?; + Ok(()) + })(); + match unsafe { self.device.GetDeviceRemovedReason() } { + Ok(()) => match submitted { + Ok(()) => Outcome::Completed, + Err(e) => Outcome::Rejected(e), + }, + Err(reason) => { + std::thread::sleep(std::time::Duration::from_millis(500)); + let (page_fault_va, dred_names) = dred_page_fault(&self.device); + Outcome::DeviceRemoved { + reason, + page_fault_va, + dred_names, + } + } + } + } + + fn write_at(&mut self, va: u64) -> Outcome { + let (rs, pso) = self.write.clone(); + self.run(|list| unsafe { + list.SetComputeRootSignature(&rs); + list.SetPipelineState(&pso); + list.SetComputeRootUnorderedAccessView(0, va); + list.Dispatch(1, 1, 1); + }) + } + + fn read_at(&mut self, va: u64) -> Outcome { + let (rs, pso) = self.read.clone(); + let scratch_va = unsafe { self.scratch.GetGPUVirtualAddress() }; + self.run(|list| unsafe { + list.SetComputeRootSignature(&rs); + list.SetPipelineState(&pso); + list.SetComputeRootShaderResourceView(0, va); + list.SetComputeRootUnorderedAccessView(1, scratch_va); + list.Dispatch(1, 1, 1); + }) + } + + /// Writes an SRV for `texture` into the descriptor heap. Done before the texture is freed + /// so the descriptor goes stale, exactly like a real use-after-free. + fn make_texture_srv(&self, texture: &ID3D12Resource) { + unsafe { + self.device.CreateShaderResourceView( + texture, + None, + self.srv_heap.GetCPUDescriptorHandleForHeapStart(), + ) + } + } + + /// Loads texel (0,0) through whatever SRV is in the descriptor heap. + fn load_texture(&mut self) -> Outcome { + let (rs, pso) = self.texture_load.clone(); + let heap = self.srv_heap.clone(); + let table = unsafe { heap.GetGPUDescriptorHandleForHeapStart() }; + let scratch_va = unsafe { self.scratch.GetGPUVirtualAddress() }; + self.run(|list| unsafe { + list.SetDescriptorHeaps(&[Some(heap.clone())]); + list.SetComputeRootSignature(&rs); + list.SetPipelineState(&pso); + list.SetComputeRootDescriptorTable(0, table); + list.SetComputeRootUnorderedAccessView(1, scratch_va); + list.Dispatch(1, 1, 1); + }) + } +} + +fn dred_page_fault(device: &ID3D12Device) -> (u64, Vec) { + let Ok(dred) = device.cast::() else { + return (0, vec![]); + }; + let Ok(out) = (unsafe { dred.GetPageFaultAllocationOutput() }) else { + return (0, vec![]); + }; + let mut names = Vec::new(); + for head in [ + out.pHeadExistingAllocationNode, + out.pHeadRecentFreedAllocationNode, + ] { + let mut node = head; + while !node.is_null() { + let n = unsafe { &*node }; + if !n.ObjectNameW.is_null() { + names.push(unsafe { n.ObjectNameW.to_string() }.unwrap_or_default()); + } + node = n.pNext; + } + } + (out.PageFaultVA, names) +} + +fn assert_removed(outcome: Outcome, expected_name: &str) { + match outcome { + Outcome::DeviceRemoved { + reason, + page_fault_va, + dred_names, + } => { + eprintln!("device removed: {reason}"); + if dred_names.iter().any(|n| n == expected_name) { + eprintln!("DRED named `{expected_name}` at VA {page_fault_va:#x}"); + } else { + // Informational: NVIDIA reports VA 0 and no names, see file header. + eprintln!( + "DRED did not attribute the fault (VA {page_fault_va:#x}, names {dred_names:?})" + ); + } + } + other => panic!("expected device removal, got {other:?}"), + } +} + +fn assert_completed(outcome: Outcome) { + match outcome { + Outcome::Completed => {} + other => panic!("expected completion, got {other:?}"), + } +} + +fault_test!(tail_overrun_faults, { + let Some(device) = setup() else { return }; + let mut allocator = stomp_allocator(&device, settings(&device)); + let desc = buffer_desc(4000, D3D12_RESOURCE_FLAG_NONE); + let resource = create( + &mut allocator, + "stomp-tail", + &desc, + ResourceCategory::Buffer, + ); + assert!(resource.is_stomp_guarded()); + let layout = resource.stomp_layout().unwrap(); + let guard_va = layout.tail_guard_va().unwrap(); + assert_eq!( + guard_va, + layout.resource_va + TILE, + "one payload tile, then the guard" + ); + + let mut gpu = Gpu::new(&device); + assert_removed(gpu.write_at(guard_va), "stomp-tail"); + allocator.free_resource(resource).unwrap(); +}); + +fault_test!(in_bounds_write_does_not_fault, { + let Some(device) = setup() else { return }; + let mut allocator = stomp_allocator(&device, settings(&device)); + let desc = buffer_desc(4000, D3D12_RESOURCE_FLAG_ALLOW_UNORDERED_ACCESS); + let resource = create(&mut allocator, "in-bounds", &desc, ResourceCategory::Buffer); + assert!(resource.is_stomp_guarded()); + let va = unsafe { resource.resource().GetGPUVirtualAddress() } + resource.offset(); + assert_eq!( + resource.offset(), + 0, + "default settings keep the payload at offset 0" + ); + + let mut gpu = Gpu::new(&device); + assert_completed(gpu.write_at(va)); + // Last 4 bytes of the payload. + assert_completed(gpu.write_at(va + 4000 - 4)); + assert_completed(gpu.read_at(va)); + allocator.free_resource(resource).unwrap(); +}); + +fault_test!(head_underrun_faults, { + let Some(device) = setup() else { return }; + let mut settings = settings(&device); + settings.head_guard = true; + settings.tail_guard = false; + let mut allocator = stomp_allocator(&device, settings); + let desc = buffer_desc(4000, D3D12_RESOURCE_FLAG_NONE); + let resource = create( + &mut allocator, + "stomp-head", + &desc, + ResourceCategory::Buffer, + ); + assert!(resource.is_stomp_guarded()); + let layout = resource.stomp_layout().unwrap(); + assert_eq!( + layout.offset, TILE, + "payload starts after the head guard tile" + ); + assert_eq!(resource.offset(), TILE); + assert_eq!(layout.head_guard_va(), Some(layout.resource_va)); + assert_eq!(layout.tail_guard_va(), None); + let payload_va = layout.resource_va + layout.offset; + + let mut gpu = Gpu::new(&device); + // First bytes of the payload are fine. + assert_completed(gpu.read_at(payload_va)); + // 64 bytes before the payload, inside the head guard. + assert_removed(gpu.read_at(payload_va - 64), "stomp-head"); + allocator.free_resource(resource).unwrap(); +}); + +fault_test!(slack_boundary_default_is_tile_precise, { + let Some(device) = setup() else { return }; + let mut allocator = stomp_allocator(&device, settings(&device)); + let desc = buffer_desc(4000, D3D12_RESOURCE_FLAG_NONE); + let resource = create(&mut allocator, "slack", &desc, ResourceCategory::Buffer); + let layout = resource.stomp_layout().unwrap(); + assert_eq!(layout.offset, 0); + let guard_va = layout.tail_guard_va().unwrap(); + + let mut gpu = Gpu::new(&device); + // Past the requested 4000 bytes but still inside the payload tile: slack, no fault. + assert_completed(gpu.write_at(layout.resource_va + 4000)); + assert_completed(gpu.write_at(guard_va - 4)); + assert_removed(gpu.write_at(guard_va), "slack"); + allocator.free_resource(resource).unwrap(); +}); + +fault_test!(slack_boundary_with_payload_alignment_is_byte_precise, { + let Some(device) = setup() else { return }; + let mut settings = settings(&device); + settings.payload_alignment = Some(8); + let mut allocator = stomp_allocator(&device, settings); + let desc = buffer_desc(4000, D3D12_RESOURCE_FLAG_NONE); + let resource = create(&mut allocator, "packed", &desc, ResourceCategory::Buffer); + let layout = resource.stomp_layout().unwrap(); + assert_eq!(layout.offset, TILE - 4000); + assert_eq!(resource.offset(), TILE - 4000); + let payload_va = layout.resource_va + layout.offset; + let guard_va = layout.tail_guard_va().unwrap(); + assert_eq!( + payload_va + 4000, + guard_va, + "payload ends exactly at the guard" + ); + + let mut gpu = Gpu::new(&device); + assert_completed(gpu.write_at(payload_va)); + // Last 4 bytes of the payload. + assert_completed(gpu.write_at(payload_va + 4000 - 4)); + // First byte past the payload is the guard. + assert_removed(gpu.write_at(guard_va), "packed"); + allocator.free_resource(resource).unwrap(); +}); + +fault_test!(buffer_use_after_free_faults, { + let Some(device) = setup() else { return }; + let mut allocator = stomp_allocator(&device, settings(&device)); + let desc = buffer_desc(4000, D3D12_RESOURCE_FLAG_ALLOW_UNORDERED_ACCESS); + let resource = create(&mut allocator, "stale", &desc, ResourceCategory::Buffer); + let va = unsafe { resource.resource().GetGPUVirtualAddress() } + resource.offset(); + + let mut gpu = Gpu::new(&device); + assert_completed(gpu.write_at(va)); + + allocator.free_resource(resource).unwrap(); + let stats = allocator.stomp_statistics().unwrap(); + assert_eq!(stats.quarantined, 1); + // Same VA, resource kept alive by the quarantine, all tiles now never-resident. + assert_removed(gpu.write_at(va), "stale"); +}); + +fault_test!(texture_use_after_free_faults, { + let Some(device) = setup() else { return }; + let mut allocator = stomp_allocator(&device, settings(&device)); + let desc = tex2d_desc(64, 64); + let resource = create( + &mut allocator, + "stale-tex", + &desc, + ResourceCategory::OtherTexture, + ); + assert!( + resource.stomp_layout().is_some(), + "texture took the stomp path" + ); + assert!(!resource.is_stomp_guarded(), "textures have no guard tile"); + + let mut gpu = Gpu::new(&device); + gpu.make_texture_srv(resource.resource()); + assert_completed(gpu.load_texture()); + + allocator.free_resource(resource).unwrap(); + assert_eq!(allocator.stomp_statistics().unwrap().quarantined, 1); + // Stale descriptor: the resource is alive but every tile maps to the never-resident heap. + assert_removed(gpu.load_texture(), "stale-tex"); +}); + +fault_test!(freed_without_quarantine_is_released, { + let Some(device) = setup() else { return }; + let mut settings = settings(&device); + settings.quarantine = false; + let mut allocator = stomp_allocator(&device, settings); + let desc = buffer_desc(4000, D3D12_RESOURCE_FLAG_ALLOW_UNORDERED_ACCESS); + let resource = create(&mut allocator, "released", &desc, ResourceCategory::Buffer); + let va = unsafe { resource.resource().GetGPUVirtualAddress() }; + + let mut gpu = Gpu::new(&device); + assert_completed(gpu.write_at(va)); + allocator.free_resource(resource).unwrap(); + assert_eq!(allocator.stomp_statistics().unwrap().quarantined, 0); + // Nothing to assert about faults: the VA is unmapped or reused, driver dependent. + eprintln!("stats after release: {:?}", allocator.stomp_statistics()); +}); + +fault_test!(oversized_copy_is_rejected_at_close, { + // Documents why the other tests use root descriptors: the runtime bounds-checks copies. + let Some(device) = setup() else { return }; + let mut allocator = stomp_allocator(&device, settings(&device)); + let dst = create( + &mut allocator, + "copy-dst", + &buffer_desc(4000, D3D12_RESOURCE_FLAG_NONE), + ResourceCategory::Buffer, + ); + let src = create( + &mut allocator, + "copy-src", + &buffer_desc(TILE * 3, D3D12_RESOURCE_FLAG_NONE), + ResourceCategory::Buffer, + ); + // The reserved buffer is 2 tiles wide (payload plus tail guard); copy 3 tiles into it. + let dst_width = unsafe { dst.resource().GetDesc() }.Width; + assert_eq!(dst_width, 2 * TILE); + + let mut gpu = Gpu::new(&device); + let outcome = gpu.run(|list| unsafe { + list.CopyBufferRegion(dst.resource(), 0, src.resource(), 0, dst_width + TILE) + }); + match outcome { + Outcome::Rejected(e) => eprintln!("runtime rejected the oversized copy: {e}"), + Outcome::DeviceRemoved { .. } => { + eprintln!("driver executed the oversized copy and faulted") + } + Outcome::Completed => panic!("oversized copy neither rejected nor faulted"), + } + allocator.free_resource(src).unwrap(); + allocator.free_resource(dst).unwrap(); +}); diff --git a/tests/shaders/root_srv_read.dxil b/tests/shaders/root_srv_read.dxil new file mode 100644 index 00000000..8e11f497 Binary files /dev/null and b/tests/shaders/root_srv_read.dxil differ diff --git a/tests/shaders/root_srv_read.hlsl b/tests/shaders/root_srv_read.hlsl new file mode 100644 index 00000000..4b1319bf --- /dev/null +++ b/tests/shaders/root_srv_read.hlsl @@ -0,0 +1,13 @@ +// Reads 4 bytes through a root SRV at whatever GPU VA the caller binds (no bounds check), +// then writes the value through a root UAV so the read cannot be optimised away. +// +// Compiled with: +// "C:/Program Files (x86)/Windows Kits/10/bin/10.0.26100.0/x64/dxc.exe" -T cs_6_0 -E main -Fo root_srv_read.dxil root_srv_read.hlsl +ByteAddressBuffer src : register(t0); +RWByteAddressBuffer dst : register(u0); + +[RootSignature("SRV(t0), UAV(u0)")] +[numthreads(1, 1, 1)] +void main() { + dst.Store(0, src.Load(0)); +} diff --git a/tests/shaders/root_uav_write.dxil b/tests/shaders/root_uav_write.dxil new file mode 100644 index 00000000..17157170 Binary files /dev/null and b/tests/shaders/root_uav_write.dxil differ diff --git a/tests/shaders/root_uav_write.hlsl b/tests/shaders/root_uav_write.hlsl new file mode 100644 index 00000000..c833c96a --- /dev/null +++ b/tests/shaders/root_uav_write.hlsl @@ -0,0 +1,11 @@ +// Stores 4 bytes through a root UAV at whatever GPU VA the caller binds (no bounds check). +// +// Compiled with: +// "C:/Program Files (x86)/Windows Kits/10/bin/10.0.26100.0/x64/dxc.exe" -T cs_6_0 -E main -Fo root_uav_write.dxil root_uav_write.hlsl +RWByteAddressBuffer dst : register(u0); + +[RootSignature("UAV(u0)")] +[numthreads(1, 1, 1)] +void main() { + dst.Store(0, 0xDEADBEEF); +} diff --git a/tests/shaders/table_srv_texture_load.dxil b/tests/shaders/table_srv_texture_load.dxil new file mode 100644 index 00000000..4c32712e Binary files /dev/null and b/tests/shaders/table_srv_texture_load.dxil differ diff --git a/tests/shaders/table_srv_texture_load.hlsl b/tests/shaders/table_srv_texture_load.hlsl new file mode 100644 index 00000000..f324e84e --- /dev/null +++ b/tests/shaders/table_srv_texture_load.hlsl @@ -0,0 +1,13 @@ +// Loads texel (0,0) of a Texture2D bound through a descriptor table and stores it through a +// root UAV, so a stale SRV to a quarantined (never-resident) texture is actually dereferenced. +// +// Compiled with: +// "C:/Program Files (x86)/Windows Kits/10/bin/10.0.26100.0/x64/dxc.exe" -T cs_6_0 -E main -Fo table_srv_texture_load.dxil table_srv_texture_load.hlsl +Texture2D tex : register(t0); +RWByteAddressBuffer dst : register(u0); + +[RootSignature("DescriptorTable(SRV(t0)), UAV(u0)")] +[numthreads(1, 1, 1)] +void main() { + dst.Store(0, asuint(tex.Load(int3(0, 0, 0)).x)); +}