What is the problem this feature will solve?
TL;DR: Companion to a libuv/libuv proposal for UV_THREADPOOL_AFFINITY (libuv/libuv#5306). Same synthetic benchmark, isolating what CPU-affinity placement buys on a Ryzen 9 5900X across 4 independent runs. Naive whole-process pinning -- the only option available today -- roughly triples p99 event-loop delay, reproduced every run. Selectively pinning just Node's own threads helps too, but modestly (4-8%). Full results and caveats near the bottom.
Node processes that also run CPU-heavy native addons (ML inference, DSP, video encoding, BLAS/oneDNN, etc.) can see their event loop starved under contention with native addon-managed thread pools scheduled across the same CPUs, even when aggregate CPU utilization doesn't look saturated.
I ran into this building a native Node.js ML inference runtime using N-API and TensorFlow's C API, where contention between an addon-managed thread pool and Node's own threads was measurable at higher concurrency.
Tools like taskset apply to the whole process, including threads a native addon creates. Pinning the process to protect the event loop also pins the addon's compute threads to the same CPUs, which is the contention this is meant to avoid. The same applies to cgroup cpuset isolation: cgroups draw a boundary at the process/container level, and a native addon runs inside the same process as Node's main thread, so there's no process-boundary tool that can separate the two. The useful abstraction is affinity for Node's own main thread, independent of addon-created threads.
What is the feature you are proposing to solve the problem?
An opt-in environment variable for the main thread specifically:
NODE_MAIN_THREAD_AFFINITY=0,1
Named with the NODE_ prefix rather than UV_, since this would be set by Node's own bootstrap code rather than read by libuv itself. UV_-prefixed variables in this ecosystem are libuv's to own (UV_THREADPOOL_SIZE, UV_USE_IO_URING); this is closer in spirit to NODE_OPTIONS, NODE_DEBUG, and similar Node-owned configuration.
A candidate point in the bootstrap sequence would be early in InitializeOncePerProcess() or in NodeMainInstance::Run() (src/node_main_instance.cc), before addon loading begins. I'd leave the exact placement to maintainer judgment rather than prescribe it here.
This is scoped strictly to the initial runtime thread running the root environment's event loop. A Worker has a programmatic construction point (new Worker(filename, { ... })) where options like this could reasonably be added later, but the main thread has no equivalent JS-level API to configure before it starts running, which is why an env var is the only lever available here.
I've filed a companion proposal at libuv/libuv for UV_THREADPOOL_AFFINITY, covering the libuv threadpool workers (fs, dns.lookup, some crypto, zlib). That piece belongs in libuv since libuv owns thread creation for the pool; the main thread doesn't, since libuv never creates it, Node just calls uv_run() from whatever thread it happens to be on. This issue is scoped to that main-thread piece specifically. Link: libuv/libuv#5306
Thread affinity inheritance (worker_threads and native addons)
Worth being explicit about a real OS-level side effect rather than assuming it away. On Linux, a thread created via pthread_create() inherits the CPU affinity mask of the thread that created it by default, unless the creator explicitly overrides it (this mirrors fork()'s documented behavior of inheriting the parent's CPU affinity mask, and has caused real bugs in other language runtimes that assumed otherwise). This has two concrete implications here:
worker_threads: a new Worker(...) call on the main thread spawns its underlying OS thread from a call chain that originates on the main thread, so by default it would inherit the pinned mask. Proposed default: reset a new Worker's affinity back to the process's original (pre-pin) mask at creation time, rather than let it inherit silently. The reasoning is about avoiding surprise, not about what Workers are conceptually for: worker_threads are commonly used for genuinely CPU-bound work, so silently constraining every new Worker to whatever narrow set the main thread happens to be pinned to would be an unexpected and easy-to-miss performance regression for code that never opted into this feature. A user who does want a Worker pinned can still do so explicitly from within the worker's own startup code.
- Native addons loaded via N-API: if an addon spins up its own thread pool during
Init() (which runs on the main thread during require()) without setting its own explicit affinity, those threads would likewise inherit the pinned main-thread mask rather than being free to use the rest of the machine, which partially undercuts the isolation this proposal is meant to provide. Libraries that set their own explicit affinity (oneDNN, MKL, via their own env vars) aren't affected in practice, but this isn't guaranteed for every addon. One option worth discussing: bracket native addon loading by resetting to the original/unpinned mask immediately before loading an addon, then re-applying the main-thread pin afterward, so addon-spawned threads inherit the wide mask rather than the narrow one. This wouldn't cover threads an addon creates lazily after load time, so it's a partial mitigation at best.
Emscripten/WASM modules aren't a separate case: WASM executes on whichever JS thread runs it, and Emscripten's pthreads support implements each WASM-side thread as a worker_threads Worker under the hood, so the reset-by-default behavior above would apply there too.
Composition with addon-level affinity
Affinity is per-thread and the last writer wins, so this proposal composes with native libraries that manage their own placement (OpenMP via OMP_PLACES/OMP_PROC_BIND, KMP_AFFINITY, GOMP_CPU_AFFINITY, library-specific affinity options) only if the two CPU sets are disjoint. If a library pins its threads to CPUs that overlap the reserved set, the two contend for the same cores and neither can migrate away, which is worse than no pinning. Nothing here coordinates the two sides; the deployer has to configure both. That is a limitation rather than a bug: this is a building block for reserving CPUs for Node's own threads, and full isolation from an addon also depends on that addon's placement being configured.
A related path: threads an addon creates lazily from a libuv pool worker (for example, code run via napi_async_work) inherit the pool's mask rather than the main thread's. That case is covered in the companion libuv issue (libuv/libuv#5306).
On cross-platform consistency
Opt-in and platform-limited, similar in shape to libuv's own UV_USE_IO_URING: gated behind an environment variable, existing behavior preserved wherever it doesn't apply. uv_thread_setaffinity() is documented as unsupported on macOS and non-atomic on Windows, so I'd propose this be a documented no-op (or explicit startup warning) on platforms where the underlying call isn't reliable.
Insufficient CPU capacity (e.g. small cloud VMs)
Worth calling out separately from platform support: on environments with very few total vCPUs, which is common on budget cloud and serverless tiers, a requested affinity set can be technically satisfiable while providing no real isolation benefit, since Node's main thread, the libuv threadpool, and any addon-managed threads all end up sharing the same one or two cores regardless. In that case, applying the pin is strictly worse than doing nothing, since it removes the OS scheduler's ability to time-slice freely for no isolation gain. Proposed handling: compare the resolved CPU set against total available capacity (the same cgroup/container-aware signal os.availableParallelism() already uses); if the pinned set would cover essentially all available capacity, warn that the environment doesn't have enough CPU headroom for meaningful isolation and skip applying the restriction.
Security scope
Worth addressing directly rather than waiting for it to come up: CPU/cache placement can be a genuine attack enabler. Cache-timing side channels (Flush+Reload, Prime+Probe) generally require co-residency in the same last-level cache, and on multi-CCD/multi-cluster hardware that means a specific cache domain, not just the same host. Code that can choose its own placement has a lever for engineering that co-residency deliberately.
This proposal doesn't meaningfully change that threat model:
- It's read once at process startup from an environment variable set by whoever launches the process, the same trust boundary
NODE_OPTIONS and UV_THREADPOOL_SIZE already occupy, not something untrusted code running inside the process can influence.
- The topology this would act on (core counts, cache-sharing groups) is already readable by any unprivileged process on the host via
/proc/cpuinfo and sysfs, independent of whether Node exposes affinity control. This doesn't disclose new information, it lets a deployer act on information that's already available.
- A native addon can already call
pthread_setaffinity_np directly today with no dependency on this proposal, so the underlying placement capability isn't new.
I'd propose this stay startup-only and env-var-driven permanently, rather than growing into a runtime JS API (e.g. something callable as process.setAffinity()). A dynamic API would be a materially different and worse threat model, since it would let arbitrary JS running inside a multi-tenant process (a plugin system, a sandboxing platform) actively position itself against other code sharing that process for exactly this kind of attack. Worth treating that boundary as a deliberate design constraint here, not an oversight to be closed later.
Scope / open questions
Strictly opt-in, not a default. Given the bootstrap-ordering and platform questions above, I'd suggest this land initially without a permanent compatibility guarantee (documented as subject to change) rather than settling the final name/format upfront.
Also worth flagging as a known limitation: this covers Node's main thread and, via the companion libuv issue, the libuv threadpool, but not V8's own background threads (GC and background compilation). Those are sized via the existing --v8-pool-size flag, but not placed, V8's background thread pool has no affinity control today, independent of both this proposal and the libuv one. Full isolation from those would likely require excluding the pinned cores at the cgroup/cpuset level outside Node entirely, which is worth surfacing here rather than implying this proposal achieves complete isolation on its own.
Also open:
- What should happen if a requested CPU is unavailable?
- No-op or warning on unsupported platforms?
- Should this be Linux-first given libuv's platform differences here?
I'd plan to include CI coverage as part of any implementation PR: a test asserting graceful fallback when an invalid/unavailable CPU is requested, and a test on a constrained (e.g. 1-2 vCPU) runner verifying the insufficient-capacity warning fires rather than a hard failure.
Benchmark results
I built a synthetic benchmark isolating the effect of affinity specifically, separate from the event-loop-blocking fix in the ONNX Runtime PR mentioned below: a CPU-heavy native addon (N pthreads doing sustained floating-point work, no sleeps) contending with Node's main thread and libuv threadpool (8 concurrent crypto.pbkdf2 calls), with taskset -p approximating this proposal and the companion libuv one from outside the process. Full methodology and code: affinity-bench.
Machine: Ryzen 9 5900X, 12C/24T, two 6-core CCDs each with its own 32MB L3, reported by Linux as a single NUMA node. The two CPU sets below are L3 cache domains discovered at runtime, not a NUMA split. 45s per condition, run four times end to end to find out whether single-run numbers were representative or noise.
Delay added above the 10ms monitorEventLoopDelay sampling floor, p99, across all 4 runs (mean ± population stdev):
| Condition |
N |
Mean |
Stdev |
Range |
vs baseline |
| 1. Baseline, no taskset |
24 |
0.97ms |
±0.02ms |
0.94-0.99ms |
— |
| 2. Whole process pinned to the Node-set mask (naive) |
24 |
2.96ms |
±0.03ms |
2.94-3.00ms |
+204% |
| 3. Selective: Node's threads / addon's threads on separate sets |
24 |
0.90ms |
±0.02ms |
0.88-0.92ms |
-8% |
| 4. Addon capped to the addon-set's core count, no pinning |
12 |
0.93ms |
±0.02ms |
0.90-0.96ms |
-4% |
| 5. Same as 4, plus selective pinning |
12 |
0.90ms |
±0.02ms |
0.88-0.92ms |
-7% (-3% vs cond. 4) |
At n=4, p99 turned out tight: every condition's stdev is about 0.02-0.03ms against means of roughly 1-3ms, a 2-3% coefficient of variation. Condition 2 vs 1 is the clearest signal and it held up under repetition: pinning the whole process to the Node-set mask while the addon still spawns one thread per logical CPU roughly triples p99 added delay in all four independent runs, with almost no spread (2.94, 2.95, 2.94, 3.00ms), since it crams the addon onto the same CPUs as Node's main thread and threadpool instead of separating anything. That's direct support for the "why process-level affinity isn't sufficient" argument above.
Selective pinning alone (3 vs 1) and capping the addon's own thread count with no pinning at all (4 vs 1) land in the same modest 4-8% range, consistently. Worth stating plainly rather than glossing over: on this workload and machine, a meaningful share of the benefit comes from giving the main thread and threadpool room to run at all, not from placement specifically. Holding thread count fixed, adding selective pinning on top of sizing (5 vs 4) moves p99 by only about 3%.
max is where repeating the run mattered: coefficients of variation of 33-80% on max across the same four runs, versus 2-3% on p99, with condition 4 alone ranging from 3.08ms to 18.85ms added delay across identical runs. I don't think a single-run max supports a tail-latency claim, and four runs is what made that visible rather than assumed. p95/p99 are what I'd trust from this benchmark; max would need considerably more repeated trials, reported as a distribution, before I'd lean on it.
Net: pinning helps, reproducibly, at p99, but modestly, 4-8%, and a comparable share of that comes from sizing the addon's pool down rather than from placement itself. The clean, dramatic-looking part of this data is process-wide pinning being bad, not selective pinning being a big win, and I'd rather present it that way than oversell the median-case numbers.
Context
This came out of work on Isidorus, a native Node.js ML inference runtime, where CPU-heavy native execution made thread placement relevant to event-loop latency.
I also encountered a related event-loop starvation problem while authoring ONNX Runtime Node.js PR #28646, which moves napi_async_work execution off the JavaScript/event-loop path. In the benchmarks for that work, maximum event-loop stall at concurrency 8 went from 129 ms to <2 ms after the change. To be clear about what that number does and doesn't show: it demonstrates the value of moving blocking work off the event loop, not thread affinity specifically. The affinity benchmark above is a separate measurement, isolating placement from that scheduling fix, and shows a real but considerably smaller effect, which is consistent with the two being different mechanisms rather than the same one under a different name.
There's also precedent for the general pattern in production: Alibaba's Noslate/Anode project, a Node.js distribution for serverless, has a merged patch enabling CPU affinity for the Node main thread specifically, for the same context-switch-avoidance reasoning.
I'm holding off on a PR here until there's some signal on the libuv side (libuv/libuv#5306), since the isolation this is meant to provide only fully holds if both land -- pinning just the main thread without the threadpool doesn't match what the benchmark above actually shows. Happy to prototype the Linux implementation once that's clearer.
What alternatives have you considered?
It's already possible today to pin the whole process from a small --require'd native addon calling pthread_setaffinity_np (there's precedent for this in small existing npm packages that wrap process-level affinity). Two reasons this doesn't fully substitute for a core-level option:
- Timing. A
--require'd addon runs after V8 initialization, which spins up its own background GC and compilation threads, and potentially after the libuv default loop already exists. Setting affinity inside InitializeOncePerProcess() (or equivalent) happens earlier in the bootstrap sequence than any userland hook can reach, which matters if the goal is isolating the earliest threads Node creates.
- Granularity. A userland addon pinning
pthread_self() at that point still only achieves process-wide affinity for whatever's been created so far. It doesn't provide a way to keep Node's threads on one CPU set and a native addon's own thread pool on a different set within the same process. That's the same limitation as taskset, just wrapped in a native module.
What is the problem this feature will solve?
TL;DR: Companion to a libuv/libuv proposal for
UV_THREADPOOL_AFFINITY(libuv/libuv#5306). Same synthetic benchmark, isolating what CPU-affinity placement buys on a Ryzen 9 5900X across 4 independent runs. Naive whole-process pinning -- the only option available today -- roughly triples p99 event-loop delay, reproduced every run. Selectively pinning just Node's own threads helps too, but modestly (4-8%). Full results and caveats near the bottom.Node processes that also run CPU-heavy native addons (ML inference, DSP, video encoding, BLAS/oneDNN, etc.) can see their event loop starved under contention with native addon-managed thread pools scheduled across the same CPUs, even when aggregate CPU utilization doesn't look saturated.
I ran into this building a native Node.js ML inference runtime using N-API and TensorFlow's C API, where contention between an addon-managed thread pool and Node's own threads was measurable at higher concurrency.
Tools like
tasksetapply to the whole process, including threads a native addon creates. Pinning the process to protect the event loop also pins the addon's compute threads to the same CPUs, which is the contention this is meant to avoid. The same applies to cgroupcpusetisolation: cgroups draw a boundary at the process/container level, and a native addon runs inside the same process as Node's main thread, so there's no process-boundary tool that can separate the two. The useful abstraction is affinity for Node's own main thread, independent of addon-created threads.What is the feature you are proposing to solve the problem?
An opt-in environment variable for the main thread specifically:
Named with the
NODE_prefix rather thanUV_, since this would be set by Node's own bootstrap code rather than read by libuv itself.UV_-prefixed variables in this ecosystem are libuv's to own (UV_THREADPOOL_SIZE,UV_USE_IO_URING); this is closer in spirit toNODE_OPTIONS,NODE_DEBUG, and similar Node-owned configuration.A candidate point in the bootstrap sequence would be early in
InitializeOncePerProcess()or inNodeMainInstance::Run()(src/node_main_instance.cc), before addon loading begins. I'd leave the exact placement to maintainer judgment rather than prescribe it here.This is scoped strictly to the initial runtime thread running the root environment's event loop. A
Workerhas a programmatic construction point (new Worker(filename, { ... })) where options like this could reasonably be added later, but the main thread has no equivalent JS-level API to configure before it starts running, which is why an env var is the only lever available here.I've filed a companion proposal at libuv/libuv for
UV_THREADPOOL_AFFINITY, covering the libuv threadpool workers (fs, dns.lookup, some crypto, zlib). That piece belongs in libuv since libuv owns thread creation for the pool; the main thread doesn't, since libuv never creates it, Node just callsuv_run()from whatever thread it happens to be on. This issue is scoped to that main-thread piece specifically. Link: libuv/libuv#5306Thread affinity inheritance (worker_threads and native addons)
Worth being explicit about a real OS-level side effect rather than assuming it away. On Linux, a thread created via
pthread_create()inherits the CPU affinity mask of the thread that created it by default, unless the creator explicitly overrides it (this mirrorsfork()'s documented behavior of inheriting the parent's CPU affinity mask, and has caused real bugs in other language runtimes that assumed otherwise). This has two concrete implications here:worker_threads: anew Worker(...)call on the main thread spawns its underlying OS thread from a call chain that originates on the main thread, so by default it would inherit the pinned mask. Proposed default: reset a new Worker's affinity back to the process's original (pre-pin) mask at creation time, rather than let it inherit silently. The reasoning is about avoiding surprise, not about what Workers are conceptually for:worker_threadsare commonly used for genuinely CPU-bound work, so silently constraining every new Worker to whatever narrow set the main thread happens to be pinned to would be an unexpected and easy-to-miss performance regression for code that never opted into this feature. A user who does want a Worker pinned can still do so explicitly from within the worker's own startup code.Init()(which runs on the main thread duringrequire()) without setting its own explicit affinity, those threads would likewise inherit the pinned main-thread mask rather than being free to use the rest of the machine, which partially undercuts the isolation this proposal is meant to provide. Libraries that set their own explicit affinity (oneDNN, MKL, via their own env vars) aren't affected in practice, but this isn't guaranteed for every addon. One option worth discussing: bracket native addon loading by resetting to the original/unpinned mask immediately before loading an addon, then re-applying the main-thread pin afterward, so addon-spawned threads inherit the wide mask rather than the narrow one. This wouldn't cover threads an addon creates lazily after load time, so it's a partial mitigation at best.Emscripten/WASM modules aren't a separate case: WASM executes on whichever JS thread runs it, and Emscripten's pthreads support implements each WASM-side thread as a
worker_threadsWorker under the hood, so the reset-by-default behavior above would apply there too.Composition with addon-level affinity
Affinity is per-thread and the last writer wins, so this proposal composes with native libraries that manage their own placement (OpenMP via
OMP_PLACES/OMP_PROC_BIND,KMP_AFFINITY,GOMP_CPU_AFFINITY, library-specific affinity options) only if the two CPU sets are disjoint. If a library pins its threads to CPUs that overlap the reserved set, the two contend for the same cores and neither can migrate away, which is worse than no pinning. Nothing here coordinates the two sides; the deployer has to configure both. That is a limitation rather than a bug: this is a building block for reserving CPUs for Node's own threads, and full isolation from an addon also depends on that addon's placement being configured.A related path: threads an addon creates lazily from a libuv pool worker (for example, code run via
napi_async_work) inherit the pool's mask rather than the main thread's. That case is covered in the companion libuv issue (libuv/libuv#5306).On cross-platform consistency
Opt-in and platform-limited, similar in shape to libuv's own
UV_USE_IO_URING: gated behind an environment variable, existing behavior preserved wherever it doesn't apply.uv_thread_setaffinity()is documented as unsupported on macOS and non-atomic on Windows, so I'd propose this be a documented no-op (or explicit startup warning) on platforms where the underlying call isn't reliable.Insufficient CPU capacity (e.g. small cloud VMs)
Worth calling out separately from platform support: on environments with very few total vCPUs, which is common on budget cloud and serverless tiers, a requested affinity set can be technically satisfiable while providing no real isolation benefit, since Node's main thread, the libuv threadpool, and any addon-managed threads all end up sharing the same one or two cores regardless. In that case, applying the pin is strictly worse than doing nothing, since it removes the OS scheduler's ability to time-slice freely for no isolation gain. Proposed handling: compare the resolved CPU set against total available capacity (the same cgroup/container-aware signal
os.availableParallelism()already uses); if the pinned set would cover essentially all available capacity, warn that the environment doesn't have enough CPU headroom for meaningful isolation and skip applying the restriction.Security scope
Worth addressing directly rather than waiting for it to come up: CPU/cache placement can be a genuine attack enabler. Cache-timing side channels (Flush+Reload, Prime+Probe) generally require co-residency in the same last-level cache, and on multi-CCD/multi-cluster hardware that means a specific cache domain, not just the same host. Code that can choose its own placement has a lever for engineering that co-residency deliberately.
This proposal doesn't meaningfully change that threat model:
NODE_OPTIONSandUV_THREADPOOL_SIZEalready occupy, not something untrusted code running inside the process can influence./proc/cpuinfoand sysfs, independent of whether Node exposes affinity control. This doesn't disclose new information, it lets a deployer act on information that's already available.pthread_setaffinity_npdirectly today with no dependency on this proposal, so the underlying placement capability isn't new.I'd propose this stay startup-only and env-var-driven permanently, rather than growing into a runtime JS API (e.g. something callable as
process.setAffinity()). A dynamic API would be a materially different and worse threat model, since it would let arbitrary JS running inside a multi-tenant process (a plugin system, a sandboxing platform) actively position itself against other code sharing that process for exactly this kind of attack. Worth treating that boundary as a deliberate design constraint here, not an oversight to be closed later.Scope / open questions
Strictly opt-in, not a default. Given the bootstrap-ordering and platform questions above, I'd suggest this land initially without a permanent compatibility guarantee (documented as subject to change) rather than settling the final name/format upfront.
Also worth flagging as a known limitation: this covers Node's main thread and, via the companion libuv issue, the libuv threadpool, but not V8's own background threads (GC and background compilation). Those are sized via the existing
--v8-pool-sizeflag, but not placed, V8's background thread pool has no affinity control today, independent of both this proposal and the libuv one. Full isolation from those would likely require excluding the pinned cores at the cgroup/cpuset level outside Node entirely, which is worth surfacing here rather than implying this proposal achieves complete isolation on its own.Also open:
I'd plan to include CI coverage as part of any implementation PR: a test asserting graceful fallback when an invalid/unavailable CPU is requested, and a test on a constrained (e.g. 1-2 vCPU) runner verifying the insufficient-capacity warning fires rather than a hard failure.
Benchmark results
I built a synthetic benchmark isolating the effect of affinity specifically, separate from the event-loop-blocking fix in the ONNX Runtime PR mentioned below: a CPU-heavy native addon (N pthreads doing sustained floating-point work, no sleeps) contending with Node's main thread and libuv threadpool (8 concurrent
crypto.pbkdf2calls), withtaskset -papproximating this proposal and the companion libuv one from outside the process. Full methodology and code: affinity-bench.Machine: Ryzen 9 5900X, 12C/24T, two 6-core CCDs each with its own 32MB L3, reported by Linux as a single NUMA node. The two CPU sets below are L3 cache domains discovered at runtime, not a NUMA split. 45s per condition, run four times end to end to find out whether single-run numbers were representative or noise.
Delay added above the 10ms
monitorEventLoopDelaysampling floor, p99, across all 4 runs (mean ± population stdev):At n=4, p99 turned out tight: every condition's stdev is about 0.02-0.03ms against means of roughly 1-3ms, a 2-3% coefficient of variation. Condition 2 vs 1 is the clearest signal and it held up under repetition: pinning the whole process to the Node-set mask while the addon still spawns one thread per logical CPU roughly triples p99 added delay in all four independent runs, with almost no spread (2.94, 2.95, 2.94, 3.00ms), since it crams the addon onto the same CPUs as Node's main thread and threadpool instead of separating anything. That's direct support for the "why process-level affinity isn't sufficient" argument above.
Selective pinning alone (3 vs 1) and capping the addon's own thread count with no pinning at all (4 vs 1) land in the same modest 4-8% range, consistently. Worth stating plainly rather than glossing over: on this workload and machine, a meaningful share of the benefit comes from giving the main thread and threadpool room to run at all, not from placement specifically. Holding thread count fixed, adding selective pinning on top of sizing (5 vs 4) moves p99 by only about 3%.
maxis where repeating the run mattered: coefficients of variation of 33-80% on max across the same four runs, versus 2-3% on p99, with condition 4 alone ranging from 3.08ms to 18.85ms added delay across identical runs. I don't think a single-run max supports a tail-latency claim, and four runs is what made that visible rather than assumed. p95/p99 are what I'd trust from this benchmark;maxwould need considerably more repeated trials, reported as a distribution, before I'd lean on it.Net: pinning helps, reproducibly, at p99, but modestly, 4-8%, and a comparable share of that comes from sizing the addon's pool down rather than from placement itself. The clean, dramatic-looking part of this data is process-wide pinning being bad, not selective pinning being a big win, and I'd rather present it that way than oversell the median-case numbers.
Context
This came out of work on Isidorus, a native Node.js ML inference runtime, where CPU-heavy native execution made thread placement relevant to event-loop latency.
I also encountered a related event-loop starvation problem while authoring ONNX Runtime Node.js PR #28646, which moves
napi_async_workexecution off the JavaScript/event-loop path. In the benchmarks for that work, maximum event-loop stall at concurrency 8 went from 129 ms to <2 ms after the change. To be clear about what that number does and doesn't show: it demonstrates the value of moving blocking work off the event loop, not thread affinity specifically. The affinity benchmark above is a separate measurement, isolating placement from that scheduling fix, and shows a real but considerably smaller effect, which is consistent with the two being different mechanisms rather than the same one under a different name.There's also precedent for the general pattern in production: Alibaba's Noslate/Anode project, a Node.js distribution for serverless, has a merged patch enabling CPU affinity for the Node main thread specifically, for the same context-switch-avoidance reasoning.
I'm holding off on a PR here until there's some signal on the libuv side (libuv/libuv#5306), since the isolation this is meant to provide only fully holds if both land -- pinning just the main thread without the threadpool doesn't match what the benchmark above actually shows. Happy to prototype the Linux implementation once that's clearer.
What alternatives have you considered?
It's already possible today to pin the whole process from a small
--require'd native addon callingpthread_setaffinity_np(there's precedent for this in small existing npm packages that wrap process-level affinity). Two reasons this doesn't fully substitute for a core-level option:--require'd addon runs after V8 initialization, which spins up its own background GC and compilation threads, and potentially after the libuv default loop already exists. Setting affinity insideInitializeOncePerProcess()(or equivalent) happens earlier in the bootstrap sequence than any userland hook can reach, which matters if the goal is isolating the earliest threads Node creates.pthread_self()at that point still only achieves process-wide affinity for whatever's been created so far. It doesn't provide a way to keep Node's threads on one CPU set and a native addon's own thread pool on a different set within the same process. That's the same limitation astaskset, just wrapped in a native module.