diff --git a/changelog.d/11321-fast-emit-per-function-containment.md b/changelog.d/11321-fast-emit-per-function-containment.md
new file mode 100644
index 0000000000..6c8dd78ccc
--- /dev/null
+++ b/changelog.d/11321-fast-emit-per-function-containment.md
@@ -0,0 +1,8 @@
+- **codegen: a function over the optimized machine-pipeline budget no longer demotes its whole codegen unit.** `PERRY_LL_FAST_EMIT_MAX_INSTRS` (600k on x86-64, 100k elsewhere) used to switch the *unit's* `TargetMachine` to LLVM's O0 machine pipeline, because the optimization level is a per-module property. On the Claude Code bundle with #11179, one ~976k-instruction factory closure (89 % `gc.relocate`) took ~950 ordinary sibling functions to O0 with it; with this change it is re-lowered to 48,418 instructions and nothing in the bundle leaves the optimized machine pipeline. A function over the budget now takes, in order:
+ 1. **re-lowered**: a statepoint function is sent back to codegen through the existing typed RS4GC retry (`Rs4gcBudgetCause::MachineBudget`), lowered with its GC roots in a shadow frame, and its unit is compiled again at the same level through the *optimized* machine pipeline. Most of such a function is `gc.relocate` fan-out, which the shadow frame does not have.
+ 2. **contained**: a function still over the budget is moved into a module of its own after the IR pipeline (`inprocess/fast_emit_split.rs`) and partially linked back with `ld -r` the way codegen units already are. It is emitted by the optimized machine pipeline with **FastISel** instruction selection (O0 only past four times the budget): on the cc closure that is +2.6 % code on x86-64 (+1.3 % arm64) where O0 was ×23 (×49), at 55 s instead of 94 s of `llc` CPU. Every other function in the unit keeps the optimized machine pipeline. Locals the cut severs (internal functions, private constants, module globals) are promoted to external **hidden** symbols under a name made unique by a hash of the unit's function set. Each half carries its own stack map.
+ 3. **whole unit**: only where the unit cannot be split (COFF targets, or a host that cannot partially link the target's objects, e.g. a cross-architecture ELF target) does the whole unit take the bounded machine (same tier rule), which is the old behaviour. An alias/ifunc in the unit or a moved function in a comdat also declines the split, with a log line.
+- The compile prints one `perry: machine code: …` summary line when anything left the optimized tier (`perry_codegen::machine_tiers`), so a regression is visible without per-unit logs.
+- `PERRY_LL_FAST_EMIT_DUMP=
` (diagnostic, not a cache input) keeps the bitcode of each contained module, and of a unit before a machine-budget re-lowering, for `llc` study.
+- `optimize_and_emit_module*` now return one emission per part; `linker::finish_native_emission` finishes and joins them.
+- Validation: `cargo test --release -p perry-codegen` (lib + all `tests/` suites) green on Linux x86-64; every `test_gap_*.ts` (1010 files) compiles to byte-identical objects on main and on this branch, so no function in the gap corpus reaches the budget. Sabotage: disabling the split, dropping the promotion, leaving the moved body on the sibling side, and disabling the re-lowering each turn their tests red.
diff --git a/crates/perry-codegen/src/codegen/spec_preserve_none_tests.rs b/crates/perry-codegen/src/codegen/spec_preserve_none_tests.rs
index b576baee6f..473e56e679 100644
--- a/crates/perry-codegen/src/codegen/spec_preserve_none_tests.rs
+++ b/crates/perry-codegen/src/codegen/spec_preserve_none_tests.rs
@@ -435,7 +435,7 @@ fn the_clone_entry_is_shrink_wrapped_frameless() {
crate::codegen::helpers::native_stack_roots_enabled(),
)
.expect("in-process -O3 -S pipeline");
- let asm = String::from_utf8(asm_bytes).expect("assembly is UTF-8");
+ let asm = String::from_utf8(asm_bytes.concat()).expect("assembly is UTF-8");
// The label line for the clone (Mach-O prefixes `_`; ELF does not).
let mut lines = asm.lines();
diff --git a/crates/perry-codegen/src/inprocess.rs b/crates/perry-codegen/src/inprocess.rs
index f8fe02b633..f8165f5c7c 100644
--- a/crates/perry-codegen/src/inprocess.rs
+++ b/crates/perry-codegen/src/inprocess.rs
@@ -17,6 +17,7 @@
//! IR and flags this pipeline produces objects byte-identical to Homebrew
//! clang 22's `clang -c`.
+mod fast_emit_split;
mod optimize_emit;
use optimize_emit::optimize_and_emit;
@@ -168,7 +169,7 @@ pub fn compile_ll_to_object_inprocess(
clang_style_args: &[String],
module_name: &str,
native_roots: bool,
-) -> Result> {
+) -> Result>> {
let (opt, mcpu_native, explicit_cpu, mllvm, emit_asm) = interpret_plan_args(clang_style_args)?;
// Same guard as the external `opt` path (`linker::rs4gc_funclet_refusal`):
// rewrite-statepoints-for-gc crashes on WinEH funclet pads, and here the
@@ -295,12 +296,17 @@ pub(crate) fn parse_ir_text<'ctx>(
/// Interpret plan argv (same grammar as `compile_ll_to_object_inprocess`) and
/// run verify -> pass pipeline -> object emission on an already-built module.
/// The native construction path calls this directly.
+///
+/// Returns one emission per part: normally one, two when fast-emit
+/// containment moved over-budget functions into a module of their own (see
+/// `fast_emit_split`). `linker::finish_native_emission` turns the parts into
+/// one object.
pub(crate) fn optimize_and_emit_module(
module: &inkwell::module::Module<'_>,
effective_target: &str,
clang_style_args: &[String],
native_roots: bool,
-) -> Result> {
+) -> Result>> {
optimize_and_emit_module_with_stats(
module,
effective_target,
@@ -319,7 +325,7 @@ pub(crate) fn optimize_and_emit_module_with_stats(
clang_style_args: &[String],
native_roots: bool,
stats: Option<&mut UnitCodegenStats>,
-) -> Result> {
+) -> Result>> {
let (opt, mcpu_native, explicit_cpu, mllvm, emit_asm) = interpret_plan_args(clang_style_args)?;
optimize_and_emit(
module,
@@ -398,18 +404,32 @@ fn module_instruction_census(
/// Per-function instruction ceiling for LLVM's optimized machine pipeline.
///
/// This budget is checked *after* the requested `default` IR pipeline has
-/// completed. It changes neither JS lowering nor middle-end optimization; it
-/// only asks the target machine to use its O0 instruction-selection,
-/// live-interval and register-allocation pipeline for a unit containing an
-/// extreme generated function.
+/// completed. It changes neither JS lowering nor middle-end optimization. A
+/// function over it takes, in order (see `crate::machine_tiers`):
+///
+/// 1. **Re-lowered**: a statepoint function goes back to codegen, is lowered
+/// with its GC roots in a shadow frame, and its unit is compiled again at
+/// the same level through the *optimized* machine pipeline. Such a
+/// function's size is mostly RS4GC relocation fan-out (163,100 of the
+/// 227,108 instructions in the `@babel/parser` closure below are
+/// `gc.relocate`), and that fan-out is what the machine pipeline is
+/// super-linear in.
+/// 2. **Contained**: a function still over the budget (or one that never had
+/// statepoints) is moved into a module of its own (`fast_emit_split`) and
+/// only it is emitted through LLVM's O0 machine pipeline. Every other
+/// function of its unit keeps the optimized one.
+/// 3. **Whole unit**: only where the unit cannot be split — COFF, or a host
+/// that cannot partially link the target's objects — the whole unit takes
+/// the O0 machine pipeline, which is the behaviour before containment.
///
-/// **The demotion is a whole-unit act, so the budget must not be set where
-/// ordinary functions pay for it.** A `TargetMachine`'s optimization level is
-/// a per-module property: LLVM has no per-function escape from the optimized
-/// machine pipeline (`optnone` reaches instruction selection and the optional
-/// machine passes, but *not* LiveIntervals or the greedy register allocator —
-/// measured below), so every ordinary function sharing the unit with one
-/// extreme function is emitted through the O0 machine pipeline too.
+/// Why the budget exists at all: a `TargetMachine`'s optimization level is a
+/// per-module property, and LLVM has no per-function escape from the
+/// optimized machine pipeline (`optnone` reaches instruction selection and
+/// the optional machine passes, but *not* LiveIntervals or the greedy
+/// register allocator — measured below). Before containment (tier 2) every
+/// ordinary function sharing the unit with one extreme function was emitted
+/// through the O0 machine pipeline too, which is why the numbers below were
+/// taken per unit.
///
/// Measured on `@babel/parser`'s unit 0, LLVM 22 / x86-64 / `-Os` IR pipeline:
/// one 227,108-instruction closure (163,100 of those are `gc.relocate`) and
@@ -422,13 +442,15 @@ fn module_instruction_census(
/// | `optnone` on the closure only | 2,070,326 B | 621,693 B | 1.382 MiB | 9.5 s | 518 MiB |
/// | the same unit *without* the closure | 1,448,633 B | — | 1.382 MiB | 6.4 s | 208 MiB |
///
-/// So the siblings are pure loss: the fallback costs them 2.06 MiB of machine
-/// code (168 of 282 functions change) to save ~6 s, and their emitted code is
-/// byte-for-byte what a unit without the extreme function produces as soon as
-/// the unit keeps the optimized pipeline. The `optnone` row is why this is a
-/// budget and not a per-function demotion: it frees the siblings but bounds
-/// neither time (9.5 s of 10.0 s) nor memory (518 MiB — *above* the -O2 arm),
-/// because the greedy allocator still runs on the demoted function.
+/// So the siblings were pure loss: the whole-unit fallback cost them 2.06 MiB
+/// of machine code (168 of 282 functions change) to save ~6 s, and their
+/// emitted code is byte-for-byte what a unit without the extreme function
+/// produces as soon as the unit keeps the optimized pipeline — which is what
+/// containment now gives them. The `optnone` row is why containment moves
+/// the function into its own module instead of stamping it: `optnone` frees
+/// the siblings but bounds neither time (9.5 s of 10.0 s) nor memory
+/// (518 MiB — *above* the -O2 arm), because the greedy allocator still runs
+/// on the demoted function.
///
/// On x86-64 the ceiling is therefore set above the whole measured
/// population of extreme generated functions rather than immediately below
@@ -525,7 +547,7 @@ thread_local! {
/// Thread-local budget seam; mutating the process environment would race the
/// other LLVM tests in this binary.
#[cfg(test)]
-fn with_test_fast_emit_budget(cap: usize, run: impl FnOnce() -> T) -> T {
+pub(crate) fn with_test_fast_emit_budget(cap: usize, run: impl FnOnce() -> T) -> T {
with_test_fast_emit_budget_value(FastEmitBudget::Cap(cap), run)
}
@@ -544,6 +566,61 @@ fn with_test_fast_emit_budget_value(budget: FastEmitBudget, run: impl FnOnce(
run()
}
+/// How far past the budget a function may be and still keep the optimized
+/// machine pipeline with FastISel instruction selection ([`MachineTier`]).
+const FAST_ISEL_TIER_FACTOR: usize = 4;
+
+/// The bounded machine configuration an over-budget function is emitted with.
+///
+/// Measured on the Claude Code bundle's 975,886-instruction factory closure
+/// (#11179; 870,626 of those instructions are `gc.relocate`), `llc` 22 on the
+/// post-`default` IR, one sample each, `.text` of the function alone:
+///
+/// | machine configuration | x86-64 `.text` | x86-64 CPU | x86-64 RSS | arm64 `.text` | arm64 CPU | arm64 RSS |
+/// |---|---|---|---|---|---|---|
+/// | optimized (`-O2`) | 280,366 B | 93.8 s | 1.22 GB | 223,568 B | 59.8 s | 1.49 GB |
+/// | optimized + FastISel (this tier) | 287,654 B | 55.1 s | 1.21 GB | 226,512 B | 56.3 s | 1.49 GB |
+/// | O0 (the old fallback) | 6,584,974 B | 20.1 s | 1.28 GB | 10,952,192 B | 12.6 s | 1.65 GB |
+///
+/// On x86-64 95 % of the optimized pipeline's time on that function is
+/// SelectionDAG instruction selection (84.5 of 89.4 s under `-time-passes`;
+/// the greedy allocator took 0.4 s), and FastISel is exactly the part of the
+/// O0 pipeline that removes it. So this tier costs +2.6 % code where O0 costs
+/// ×23 (×48 on arm64), and it is a per-`TargetMachine` switch, so it needs
+/// no process-global `cl::opt`.
+///
+/// O0 stays as the backstop only for a function more than
+/// [`FAST_ISEL_TIER_FACTOR`] times over the budget, which is where the
+/// measurements stop: the degradation grows with the excess instead of
+/// stepping to ×23 at the budget.
+#[derive(Debug, Clone, Copy, PartialEq, Eq)]
+pub enum MachineTier {
+ /// The requested optimization level with FastISel instruction selection:
+ /// the optimized register allocator and machine passes are kept.
+ FastIsel,
+ /// LLVM's O0 machine pipeline (FastISel + the fast register allocator).
+ O0,
+}
+
+impl MachineTier {
+ pub(crate) fn for_function(instructions: usize, cap: usize) -> Self {
+ if instructions <= cap.saturating_mul(FAST_ISEL_TIER_FACTOR) {
+ MachineTier::FastIsel
+ } else {
+ MachineTier::O0
+ }
+ }
+
+ fn describe(self) -> &'static str {
+ match self {
+ MachineTier::FastIsel => {
+ "the optimized machine pipeline with FastISel instruction selection"
+ }
+ MachineTier::O0 => "LLVM's O0 machine pipeline",
+ }
+ }
+}
+
/// One extreme function which selected bounded machine-code emission, and how
/// many defined functions in its unit are demoted along with it.
#[derive(Debug, Clone, PartialEq, Eq)]
@@ -551,20 +628,39 @@ pub struct FastEmitFallback {
pub name: String,
pub instructions: usize,
pub cap: usize,
- /// Defined functions in the unit — the size of the collateral, since the
- /// machine pipeline is selected per module and not per function.
+ /// Defined functions in the unit — the size of the collateral when the
+ /// unit could not be split, since the machine pipeline is selected per
+ /// module and not per function.
pub unit_functions: usize,
+ /// Whether the over-budget functions were moved into a module of their
+ /// own (`fast_emit_split`), so only they took the bounded pipeline and
+ /// every other function in the unit kept the optimized one.
+ pub contained: bool,
+ /// The bounded machine configuration it was emitted with.
+ pub tier: MachineTier,
}
impl std::fmt::Display for FastEmitFallback {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
+ let how = self.tier.describe();
+ if self.contained {
+ return write!(
+ f,
+ "`{}` has {} instructions after IR optimization, above the optimized \
+ machine-pipeline budget {}; keeping the requested IR optimization, then \
+ emitting this function alone through {how} to bound instruction selection. \
+ Every function of its {}-function unit that is under the budget keeps the \
+ optimized machine pipeline. Override with PERRY_LL_FAST_EMIT_MAX_INSTRS= \
+ (raise) or =0 (disable).",
+ self.name, self.instructions, self.cap, self.unit_functions
+ );
+ }
write!(
f,
"`{}` has {} instructions after IR optimization, above the optimized machine-pipeline \
budget {}; keeping the requested IR optimization, then emitting this unit — all {} \
- of its defined functions, not only this one — through LLVM's O0 machine pipeline to \
- bound instruction selection, live intervals and register allocation. LLVM selects \
- that pipeline per module, so the siblings are demoted too and grow: shrinking this \
+ of its defined functions, not only this one — through {how}. The unit could not \
+ be split for this target, so the siblings take that pipeline too: shrinking this \
function is what lifts the whole unit back. Override with \
PERRY_LL_FAST_EMIT_MAX_INSTRS= (raise) or =0 (disable).",
self.name, self.instructions, self.cap, self.unit_functions
@@ -610,6 +706,8 @@ fn fast_emit_fallbacks(
instructions,
cap,
unit_functions: defined,
+ contained: false,
+ tier: MachineTier::for_function(instructions, cap),
})
.collect()
}
@@ -652,6 +750,12 @@ pub(crate) enum Rs4gcBudgetCause {
/// RS4GC finished, but its relocation fan-out made the rewritten body too
/// large for the normal optimization pipeline.
PostRewrite { post_instructions: usize },
+ /// The IR pipeline finished, but the optimized function is over the
+ /// machine-pipeline budget ([`DEFAULT_FAST_EMIT_MAX_INSTRS_X86_64`]).
+ /// Re-lowering it onto a shadow frame removes the statepoint relocations
+ /// that make up most of such a function, so it can keep the optimized
+ /// machine pipeline instead of the bounded one.
+ MachineBudget { instructions: usize },
}
#[derive(Debug, Clone, PartialEq, Eq)]
@@ -961,6 +1065,13 @@ fn rewrite_budget_message(violation: &Rs4gcBudgetViolation, retry: bool) -> Stri
violation.name, violation.cap
)
}
+ Rs4gcBudgetCause::MachineBudget { instructions } => format!(
+ "`{}` has {instructions} instructions after IR optimization, above the optimized \
+ machine-pipeline budget {}; most of a statepoint function this size is relocation \
+ fan-out, so {outcome}. Override with PERRY_LL_FAST_EMIT_MAX_INSTRS= (raise) or \
+ =0 (disable).",
+ violation.name, violation.cap
+ ),
}
}
diff --git a/crates/perry-codegen/src/inprocess/fast_emit_split.rs b/crates/perry-codegen/src/inprocess/fast_emit_split.rs
new file mode 100644
index 0000000000..7f6187818e
--- /dev/null
+++ b/crates/perry-codegen/src/inprocess/fast_emit_split.rs
@@ -0,0 +1,464 @@
+//! Per-function fast-emit containment.
+//!
+//! A `TargetMachine` selects its machine pipeline for a whole module, so the
+//! bounded machine pipeline chosen for one extreme function (see
+//! [`super::DEFAULT_FAST_EMIT_MAX_INSTRS_X86_64`]) used to reach every
+//! ordinary function in its codegen unit too. On the Claude Code bundle one
+//! 988k-instruction factory closure demoted 950 siblings.
+//!
+//! This module moves the over-budget functions into a module of their own
+//! after the IR pipeline has run, so each half is emitted by its own target
+//! machine: the siblings by the unit's optimized one, the extreme functions
+//! by the bounded one. The two emissions become two objects that the caller
+//! partially links (`linker::merge_unit_objects`, the step that already joins
+//! codegen units), so the rest of the backend never sees two objects.
+//!
+//! What has to hold across the cut:
+//!
+//! * Every internal symbol that the moved functions reference (functions,
+//! string constants, module globals) is defined on the sibling side and
+//! declared on the contained side. It is promoted to external linkage with
+//! hidden visibility under a name made unique by a hash of the unit's
+//! function set, so no two units — and no two Perry modules — can define
+//! the same promoted name. Hidden keeps it out of the final image's
+//! dynamic symbol table and keeps its references PC-relative.
+//! * A moved function that was internal is promoted the same way, because its
+//! callers stay on the sibling side.
+//! * Each object carries its own statepoint stack map for exactly the
+//! functions it defines, and the partial link concatenates the compact
+//! `.perry_gcmap` sections the way it already does for codegen units.
+//! * Module-level `asm` (the Mach-O `.no_dead_strip` for the stack map) is
+//! kept on both sides; `llvm.used`-style appending arrays and every global
+//! initializer stay on the sibling side only.
+//!
+//! Shapes that this cut cannot express safely — an alias or ifunc anywhere in
+//! the unit, a moved function in a comdat — decline containment, and the unit
+//! falls back to whole-unit bounded emission exactly as before.
+
+use std::collections::HashSet;
+
+use inkwell::module::Module;
+use llvm_sys::comdat::{LLVMGetComdat, LLVMSetComdat};
+use llvm_sys::core::*;
+use llvm_sys::prelude::*;
+use llvm_sys::{LLVMLinkage, LLVMTypeKind, LLVMVisibility};
+
+/// Suffix infix of every promoted name, so a symbol table shows where it came
+/// from.
+pub(super) const PROMOTED_INFIX: &str = ".perry.fe.";
+
+/// Why a unit could not be split. The caller logs it and keeps the old
+/// whole-unit behaviour.
+#[derive(Debug, Clone, PartialEq, Eq)]
+pub(super) struct SplitDeclined(pub String);
+
+impl std::fmt::Display for SplitDeclined {
+ fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
+ f.write_str(&self.0)
+ }
+}
+
+/// Whether this host can partially link objects for `effective_target`.
+///
+/// The two halves are joined by `ld -r` (`linker::merge_unit_objects`). COFF
+/// has no relocatable link (codegen units are archived there instead, and an
+/// archive cannot nest inside another unit's archive), and a host linker
+/// only reads its own object format — and, for GNU ld, its own architecture.
+/// Everywhere else the unit keeps whole-unit bounded emission.
+pub(super) fn split_emission_supported(effective_target: &str) -> bool {
+ let arch = effective_target
+ .split('-')
+ .next()
+ .unwrap_or(effective_target);
+ let apple = effective_target.contains("apple");
+ if effective_target.contains("windows") {
+ return false;
+ }
+ if cfg!(target_os = "macos") {
+ return apple;
+ }
+ if cfg!(target_os = "linux") {
+ let host = std::env::consts::ARCH;
+ let same_arch = match arch {
+ "x86_64" | "x86_64h" | "amd64" => host == "x86_64",
+ "aarch64" | "arm64" => host == "aarch64",
+ _ => false,
+ };
+ return !apple && same_arch;
+ }
+ false
+}
+
+fn value_name(value: LLVMValueRef) -> String {
+ let mut len = 0usize;
+ let ptr = unsafe { LLVMGetValueName2(value, &mut len) };
+ if ptr.is_null() || len == 0 {
+ return String::new();
+ }
+ let bytes = unsafe { std::slice::from_raw_parts(ptr as *const u8, len) };
+ String::from_utf8_lossy(bytes).into_owned()
+}
+
+fn set_value_name(value: LLVMValueRef, name: &str) {
+ unsafe { LLVMSetValueName2(value, name.as_ptr() as *const _, name.len()) };
+}
+
+fn is_local(linkage: LLVMLinkage) -> bool {
+ matches!(
+ linkage,
+ LLVMLinkage::LLVMInternalLinkage | LLVMLinkage::LLVMPrivateLinkage
+ )
+}
+
+fn is_definition(global: LLVMValueRef) -> bool {
+ unsafe { LLVMIsDeclaration(global) == 0 }
+}
+
+fn functions(module: LLVMModuleRef) -> Vec {
+ let mut out = Vec::new();
+ let mut f = unsafe { LLVMGetFirstFunction(module) };
+ while !f.is_null() {
+ out.push(f);
+ f = unsafe { LLVMGetNextFunction(f) };
+ }
+ out
+}
+
+fn global_variables(module: LLVMModuleRef) -> Vec {
+ let mut out = Vec::new();
+ let mut g = unsafe { LLVMGetFirstGlobal(module) };
+ while !g.is_null() {
+ out.push(g);
+ g = unsafe { LLVMGetNextGlobal(g) };
+ }
+ out
+}
+
+/// FNV-1a over every defined function name (module order) and the moved
+/// set. Deterministic, so the object cache and reproducible builds see the
+/// same promoted names every time; distinct per unit, because no two units
+/// define the same set of functions.
+fn unit_token(module: LLVMModuleRef, moved: &[LLVMValueRef]) -> u64 {
+ let mut hash: u64 = 0xcbf2_9ce4_8422_2325;
+ let mut eat = |bytes: &[u8]| {
+ for b in bytes.iter().chain(std::iter::once(&0u8)) {
+ hash ^= u64::from(*b);
+ hash = hash.wrapping_mul(0x0000_0100_0000_01b3);
+ }
+ };
+ for f in functions(module) {
+ if is_definition(f) {
+ eat(value_name(f).as_bytes());
+ }
+ }
+ eat(b"--moved--");
+ for f in moved {
+ eat(value_name(*f).as_bytes());
+ }
+ hash
+}
+
+/// Every internal/private global value a moved function body references,
+/// through instruction operands and nested constant expressions (not through
+/// other globals' initializers, which stay on the sibling side).
+fn locals_referenced_by(function: LLVMValueRef) -> Vec {
+ let mut seen_constants: HashSet = HashSet::new();
+ let mut found: Vec = Vec::new();
+ let mut found_set: HashSet = HashSet::new();
+ let mut stack: Vec = Vec::new();
+
+ let mut visit_operand = |value: LLVMValueRef, stack: &mut Vec| {
+ if value.is_null() {
+ return;
+ }
+ unsafe {
+ if !LLVMIsAGlobalValue(value).is_null() {
+ if is_local(LLVMGetLinkage(value)) && found_set.insert(value) {
+ found.push(value);
+ }
+ } else if !LLVMIsAConstant(value).is_null() && seen_constants.insert(value) {
+ stack.push(value);
+ }
+ }
+ };
+
+ unsafe {
+ if LLVMHasPersonalityFn(function) != 0 {
+ visit_operand(LLVMGetPersonalityFn(function), &mut stack);
+ }
+ let mut bb = LLVMGetFirstBasicBlock(function);
+ while !bb.is_null() {
+ let mut inst = LLVMGetFirstInstruction(bb);
+ while !inst.is_null() {
+ let n = LLVMGetNumOperands(inst);
+ for i in 0..n.max(0) as u32 {
+ visit_operand(LLVMGetOperand(inst, i), &mut stack);
+ }
+ inst = LLVMGetNextInstruction(inst);
+ }
+ bb = LLVMGetNextBasicBlock(bb);
+ }
+ while let Some(constant) = stack.pop() {
+ let n = LLVMGetNumOperands(constant);
+ for i in 0..n.max(0) as u32 {
+ visit_operand(LLVMGetOperand(constant, i), &mut stack);
+ }
+ }
+ }
+ found
+}
+
+/// Give a local global value external linkage, hidden visibility and a
+/// unit-unique name.
+fn promote(global: LLVMValueRef, token: u64, anon: &mut usize) {
+ let base = value_name(global);
+ let base = if base.is_empty() {
+ *anon += 1;
+ format!("anon{}", *anon)
+ } else {
+ base
+ };
+ set_value_name(global, &format!("{base}{PROMOTED_INFIX}{token:016x}"));
+ unsafe {
+ LLVMSetLinkage(global, LLVMLinkage::LLVMExternalLinkage);
+ LLVMSetVisibility(global, LLVMVisibility::LLVMHiddenVisibility);
+ }
+}
+
+/// Remove a function's body in place, keeping its type,
+/// attributes, calling convention, GC strategy and visibility — the things a
+/// caller's code generation reads from the callee.
+///
+/// The C API has no `deleteBody`, so this is what `DeleteDeadBlocks` does:
+/// replace every instruction's uses with poison, erase the instructions
+/// (dropping their references to blocks and values), then delete the empty
+/// blocks. The caller settles the linkage: a declaration may not be local.
+pub(super) fn strip_body(function: LLVMValueRef) {
+ unsafe {
+ let mut bb = LLVMGetFirstBasicBlock(function);
+ while !bb.is_null() {
+ let mut inst = LLVMGetFirstInstruction(bb);
+ while !inst.is_null() {
+ let ty = LLVMTypeOf(inst);
+ if LLVMGetTypeKind(ty) != LLVMTypeKind::LLVMVoidTypeKind
+ && !LLVMGetFirstUse(inst).is_null()
+ {
+ LLVMReplaceAllUsesWith(inst, LLVMGetPoison(ty));
+ }
+ inst = LLVMGetNextInstruction(inst);
+ }
+ bb = LLVMGetNextBasicBlock(bb);
+ }
+ let mut bb = LLVMGetFirstBasicBlock(function);
+ while !bb.is_null() {
+ let mut inst = LLVMGetLastInstruction(bb);
+ while !inst.is_null() {
+ let prev = LLVMGetPreviousInstruction(inst);
+ LLVMInstructionEraseFromParent(inst);
+ inst = prev;
+ }
+ bb = LLVMGetNextBasicBlock(bb);
+ }
+ loop {
+ let bb = LLVMGetFirstBasicBlock(function);
+ if bb.is_null() {
+ break;
+ }
+ LLVMDeleteBasicBlock(bb);
+ }
+ if LLVMHasPersonalityFn(function) != 0 {
+ LLVMSetPersonalityFn(function, std::ptr::null_mut());
+ }
+ LLVMSetComdat(function, std::ptr::null_mut());
+ }
+}
+
+/// A declaration may only be external (or extern_weak). Visibility is kept:
+/// a hidden declaration is what lets the backend address its definition in
+/// the other half PC-relatively.
+fn declaration_linkage(global: LLVMValueRef) {
+ unsafe {
+ if LLVMGetLinkage(global) != LLVMLinkage::LLVMExternalWeakLinkage {
+ LLVMSetLinkage(global, LLVMLinkage::LLVMExternalLinkage);
+ }
+ }
+}
+
+/// Split `module` in place: after this returns `Ok(Some(contained))`,
+/// `module` defines every function except `moved`, and `contained` defines
+/// exactly `moved` and declares everything they reference.
+///
+/// `Ok(None)`: every defined function in the unit is in `moved`, so there is
+/// nothing to protect and the unit is emitted whole by the bounded machine.
+///
+/// `Err` means nothing was cut and `module` is still one emittable unit (the
+/// only mutation that may already have happened is the promotion of locals,
+/// which is correct for a single object too).
+pub(super) fn split_moved_functions<'ctx>(
+ module: &Module<'ctx>,
+ moved_names: &[String],
+) -> Result