Skip to content

Commit 92b4fcd

Browse files
committed
G1 falsifier ANSWERED: the widening WAS the cost — 6.75x on the re-chain
The plan pre-registered this and left it open through both N2 and N3: "build `gt_u8_to_mask`, re-run both probes; if neither the re-chain nor the #308 crossover moves, the widening was not the cost and G1 drops in priority." The re-chain moved. `hex_tenant_mq_probe` carried its own admission at the widening site -- "widening costs 4x the bandwidth, so `n_gen` below is an UPPER bound on the reveal term" -- and the T1 addition that would replace the bound with a measurement has now landed. This adds the NATIVE u8 arm beside the widened i32 one rather than replacing it, so both are timed in ONE process on ONE dataset. | tier | M1b: 6 masks | | | coal: one re-chain | | | |---|---|---|---|---|---|---| | | widened i32 | native u8 | ratio | widened i32 | native u8 | ratio | | v4 / AVX-512 | 40810 ns | 5064 ns | **8.06x** | 6789 ns | 1006 ns | **6.75x** | | v3 / AVX2 | 43798 ns | 5638 ns | 7.77x | 7189 ns | 1101 ns | 6.53x | In the probe's own cost-model units (v4, x=4) a maneuver goes from **1.02 maintained steps to 0.15** -- a re-chain used to cost a whole maintained step and now costs about a seventh of one. Three things about the method, because the number is only as good as they are: * **Bit-identity is asserted BEFORE either arm is timed.** A timing comparison between two operations that do not produce the same answer measures nothing; both `assert_eq!`s must pass or the probe aborts. * **Same process, same data, same run.** Neither tier reproduces the 8.9 us the plan quotes for this re-chain (v4 6789 ns, v3 7189 ns), so that historical absolute came from a build this one does not reproduce. The RATIO is unaffected by that drift precisely because both arms are measured side by side rather than across runs -- which is why the native arm was ADDED rather than swapped in. * **The widened column's own materialization is OUTSIDE the timed region** (it is built once, up front). So 6.75x is the steady-state sweep cost only, and the conservative reading -- it favours the widened arm. **The ratio is nearly tier-independent (6.75x vs 6.53x), which says the win is a WIDTH effect, not an ISA effect.** Arithmetic that bounds it: for N elements the i32 path issues N/16 compares reading 4N bytes, the u8 path N/64 compares reading N bytes -- 4x fewer instructions AND 4x less memory. Observed 6.5-8x exceeds either alone. The plausible remainder is the packing: the i32 path must shift-and-OR each 16-bit group into its word, while at u8 one chunk IS one whole word and the packing disappears entirely. That attribution is a CONJECTURE consistent with the numbers, not a separate measurement. **Half the falsifier remains unrun, and it is blocked, not skipped.** `r2il_column_scan_probe` is the other pre-registered half (its own header names both gaps: no u8 comparator, no u64 range comparator -- N2 and N3 respectively). It needs a column dump from `r2sleigh-lift`'s `win32_census`, which needs a Win32 PE binary; none exists in this container. Substituting a synthetic dump would produce a number shaped like the falsifier's answer without being it, and that probe's own docs insist on "a real lift rather than a synthetic stream". So G2's measurement stays open. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X6y3drwKSE2zSgoexheLFX
1 parent 1c4e45f commit 92b4fcd

1 file changed

Lines changed: 50 additions & 5 deletions

File tree

‎examples/hex_tenant_mq_probe.rs‎

Lines changed: 50 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -53,8 +53,8 @@ use std::time::Instant;
5353

5454
use ndarray::simd::ternlog::{AND3, OR2_AND};
5555
use ndarray::simd::{
56-
gt_i32_to_mask, mask_and, mask_set_range, mask_shift_morton, mask_ternlog_assign, popcount_batch_u64,
57-
ternary_match_u32_to_mask, MortonDir,
56+
gt_i32_to_mask, gt_u8_to_mask, mask_and, mask_set_range, mask_shift_morton, mask_ternlog_assign,
57+
popcount_batch_u64, ternary_match_u32_to_mask, MortonDir,
5858
};
5959

6060
// ── counting allocator: the 0k instrument ────────────────────────────────────
@@ -472,9 +472,23 @@ fn main() {
472472
})
473473
.collect();
474474
let thr: u8 = 96; // ~62% of rails permeable
475-
// Column views for the SIMD compare. The T1 compare is i32-wide today; a
476-
// u8 column compare is a T1 addition (stated, not hidden). Widening costs
477-
// 4× the bandwidth, so `n_gen` below is an UPPER bound on the reveal term.
475+
476+
// Column views for the SIMD compare, BOTH widths, so the widening cost is
477+
// measured in one process instead of argued across two runs.
478+
//
479+
// This probe's original note read: *"The T1 compare is i32-wide today; a
480+
// u8 column compare is a T1 addition (stated, not hidden). Widening costs
481+
// 4× the bandwidth, so `n_gen` below is an UPPER bound on the reveal
482+
// term."* That T1 addition landed (N2/G1, `gt_u8_to_mask`), so the upper
483+
// bound can be replaced by a measurement — which is the falsifier the plan
484+
// pre-registered for G1: build the u8 comparator, re-run this probe, and
485+
// **if the re-chain does not move, the widening was never the cost.**
486+
//
487+
// The rail byte is the NATIVE column: `r[2 * d]` is already `u8`, and the
488+
// `as i32` below is the whole widening.
489+
let perm_cols_u8: Vec<Vec<u8>> = (0..DIRS)
490+
.map(|d| rails.iter().map(|r| r[2 * d]).collect())
491+
.collect();
478492
let perm_cols: Vec<Vec<i32>> = (0..DIRS)
479493
.map(|d| rails.iter().map(|r| r[2 * d] as i32).collect())
480494
.collect();
@@ -487,6 +501,27 @@ fn main() {
487501
gt_i32_to_mask(&perm_cols[d], thr as i32, &mut elig[d]);
488502
}
489503
});
504+
505+
// The same six masks off the NATIVE u8 columns. Gate first, time second:
506+
// a timing comparison between two operations that do not produce the same
507+
// answer measures nothing, so bit-identity is asserted before any number
508+
// is printed.
509+
let mut elig_u8: [Vec<u64>; DIRS] = std::array::from_fn(|_| vec![0u64; WORDS]);
510+
for d in 0..DIRS {
511+
gt_u8_to_mask(&perm_cols_u8[d], thr, &mut elig_u8[d]);
512+
assert_eq!(elig_u8[d], elig[d], "u8 and widened-i32 eligibility masks differ on rail {d}");
513+
}
514+
let t_gen_u8 = timed(|| {
515+
for d in 0..DIRS {
516+
gt_u8_to_mask(&perm_cols_u8[d], thr, &mut elig_u8[d]);
517+
}
518+
});
519+
println!(
520+
"M1b generation NATIVE u8: 6 masks = {t_gen_u8:.0} ns ({:.0} ns/mask, {:.2} ns/row) — {:.2}× the widened i32 arm",
521+
t_gen_u8 / DIRS as f64,
522+
t_gen_u8 / (DIRS * N) as f64,
523+
t_gen / t_gen_u8
524+
);
490525
println!(
491526
"M1b generation: 6 eligibility masks from 6 columns = {:.0} ns ({:.0} ns/mask, {:.2} ns/row)",
492527
t_gen,
@@ -724,6 +759,16 @@ fn main() {
724759
c / t_tern,
725760
c / (4.0 * t_tern + n_term)
726761
);
762+
let mut m_new_u8 = vec![0u64; WORDS];
763+
gt_u8_to_mask(&perm_cols_u8[0], thr, &mut m_new_u8);
764+
assert_eq!(m_new_u8, m_new, "u8 and widened-i32 re-chain masks differ");
765+
let c_u8 = timed(|| gt_u8_to_mask(&perm_cols_u8[0], thr, &mut m_new_u8));
766+
println!(
767+
"[coal] one re-chain NATIVE u8 (gt_u8 sweep) = {c_u8:.0} ns = {:.1} ternlogq passes = {:.2} maintained steps at x=4 → {:.2}× vs widened",
768+
c_u8 / t_tern,
769+
c_u8 / (4.0 * t_tern + n_term),
770+
c / c_u8
771+
);
727772
println!(
728773
"[M2] speed change x→x±1 costs one ternlogq pass ({t_tern:.0} ns); x→x±k costs k passes — linear, no cliff"
729774
);

0 commit comments

Comments
 (0)