Skip to content

[WIP 0.71.0] numpy.polynomial, np.random parity, numpy.ma, np.emath, 48 new np.* functions, System.Numerics.Tensors interop, NumPy-parity fixes - #631

Draft
Nucs wants to merge 300 commits into
masterfrom
journey4
Draft

Nucs wants to merge 300 commits into
masterfrom
journey4

Conversation

@Nucs

@Nucs Nucs commented Sep 23, 2026 •

Copy link
Copy Markdown
Member

NumSharp 0.71.0

🧭 Overview

  • NumPy 2.x API coverage 87.1% -> 95.5% (488 -> 535 of the 560 headline APIs), and 494 of the 560 are oracle-verified against committed NumPy 2.4.2 contracts - every figure from the Supported Features Dashboard.
    • np.polynomial.* - the numpy.polynomial package, 172 of its 193 functions for all six bases, bit-exact.
    • np.* - +48 functions.
    • np.random.* - every NumPy bit generator, the whole Generator distribution surface, and a legacy RandomState that is NumPy's own sampler code line for line, all byte-identical streams.
    • np.ma.* - the masked-array module; np.emath.* - the automatic-domain math module.
  • DType is NumPy's two-level dtype model - NDArray.dtype returns a numpy.dtype-shaped descriptor, with np.dtypes, byte order and datetime64 / timedelta64 descriptors.
  • Reductions on NumPy's own schedules - max, min, ptp, nanmax, nanmin, argmax and the fused np.evaluate reductions match NumPy bit for bit, down to the sign of a ±0 tie and a NaN's payload, and most of them run faster.
  • ~60 NumPy-parity and memory-safety fixes, 14 measured speedups and 14 breaking changes - every line traced to one of 138 cited commits out of the release's 348.
  • 1 new NuGet package - NumSharp.Interop.System.Numerics.Tensors, a zero-copy bridge to .NET's Tensor<T> / TensorSpan<T>.
  • 407,913 NumPy 2.4.2 oracle test cases (0.70.0 shipped with 116K+) and 16,818 test methods.

Detailed Breakdown

Show All

📦 New NuGet Packages

One optional companion package ships for the first time, co-versioned with NumSharp 0.71.0 (it depends on NumSharp).

  • NumSharp.Interop.System.Numerics.Tensors - zero-copy interop between NDArray and System.Numerics.Tensors (Tensor<T>, TensorSpan<T>, ReadOnlyTensorSpan<T>); all 15 NumSharp dtypes cross as their own CLR type, Half, Complex, decimal and char included.
    • Export - AsTensorSpan<T>() views an NDArray of any non-negative-stride layout (contiguous, sliced, transposed, strided, broadcast) as a TensorSpan<T> over the same unmanaged buffer, through a disposable handle that keeps the buffer pinned; ToTensor<T>() copies into a dense Tensor<T>.
    • Import - AsNDArray() views a Tensor<T> as an NDArray without copying (writes go through to the tensor); ToNDArray() copies a Tensor<T>, TensorSpan<T> or ReadOnlyTensorSpan<T>.
    • The exported spans feed the BCL's SIMD TensorPrimitives directly, with no fill loop.
    • Limits - a negative-stride view is refused, because the BCL forbids negative strides (copy it first); the copies stop at the managed-array size limit; live export and import counters help leak checks.
    • Depends - System.Numerics.Tensors 10.x only, a managed BCL package. Its tensor types are experimental, so the bridge marks its own API experimental too and you opt in once. No engine backend and no native asset: Core stays 100% managed.

📊 Dashboards & Docs

  • Supported Features Dashboard - now reports, for every API, whether its behavior is oracle-verified (committed NumPy 2.4.2 contracts resolve to it, directly or through a reviewed alias) next to whether it is available. Headline 95.5% available (535/560), 88.2% oracle-verified (494/560).
    • An oracle filter in the API explorer and an evidence pane per API: contract and error-contract counts, dtypes, layouts, corpus files and host pins.
    • The catalog also covers what lies beyond the headline - NumPy's submodules (numpy.ma, numpy.emath, numpy.polynomial, numpy.char, numpy.strings, ...) and the members of its objects (nditer, dtype, finfo, the random classes): 1,137 of 2,519 APIs available in that expanded scope, the dashboard's default view.
  • Tests & Oracle Dashboard - 407,913 committed NumPy 2.4.2 oracle test cases over 793 operations (0.70.0: 116K+), and 16,818 test declarations.
  • Benchmark Dashboard - benchmark evidence for 452 of 518 benchmarkable APIs (the denominator grew with this release's new APIs).
  • Documentation site - new guides, whose code samples run as tests:

✨ New APIs & Modules

  • np.polynomial - the numpy.polynomial package for all six bases (polynomial, chebyshev, legendre, laguerre, hermite, hermite_e) plus polyutils: 172 of its 193 functions, bit-exact with NumPy 2.4.2 against generated oracles (139,882 cases, 0 excused); {p} below is poly / cheb / leg / lag / herm / herme - 57545200, 38ed8da7, ac8608d9, 38526155, e2ff9477, 0c02740c (+ 0896cbf8, d494e83d).
    • {p}val, {p}val2d, {p}val3d, {p}grid2d, {p}grid3d, {p}valnd - evaluation.
    • {p}add, {p}sub, {p}mulx, {p}mul, {p}div, {p}pow, {p}fromroots - series arithmetic.
    • {p}der, {p}int - calculus, with every m / k / lbnd / scl / axis form.
    • {p}vander, {p}vander2d, {p}vander3d - Vandermonde matrices.
    • {p}companion, {p}roots - companion matrices and roots (the roots compute through the OpenBLAS backend's eigvals).
    • {p}trim, {p}line, {p}domain, {p}zero, {p}one, {p}x, the five X2poly / poly2X pairs - helpers, module constants and basis conversions.
    • as_series, trimseq, trimcoef, getdomain, mapparms, mapdomain - polyutils.
    • Speed - x4.65-x37.8 NumPy's speed by benchmark geomean per family (evaluation x7.5, additive x8.5, calculus x7.6, series algebra x37.8, Vandermonde x10.2, roots x4.65).
    • Still to come: fitting, Gauss quadrature, chebpts / chebinterpolate, polyvalfromroots, format_float and the six series classes.
  • np.random bit-generator family - NumPy's engines, seeding and stream splitting, byte-identical to NumPy 2.4.2 and replayed overload by overload (493 members, 5 engines x 10 seeds) - c1e0feb5, 65bdea7b (+ 0e39714b, eebe85a6).
    • PCG64DXSM, Philox, SFC64 - new bit generators; PCG64.advance, jumped, random_raw, state - engine controls; MT19937 is now a NumPy BitGenerator too.
    • SeedSequence.spawn, spawn_key, pool_size, state, BitGenerator.spawn, Generator.spawn, ISeedSequence - seeding and independent streams.
    • default_rng - every NumPy seed form (integers, BigInteger, arrays, SeedSequence, a bit generator, a Generator, a RandomState, null for fresh entropy).
    • np.random.PCG64(42), np.random.Generator(...), np.random.SeedSequence(...) and the other classes - NumPy's module spellings; get_bit_generator, set_bit_generator.
  • Generator distributions - NumPy's whole numpy.random.Generator surface on NumPy's modern algorithms, 29 samplers byte-identical to default_rng - c0613b2f.
    • beta, chisquare, f, noncentral_chisquare, noncentral_f, standard_t, standard_cauchy, vonmises - the gamma family and friends.
    • pareto, weibull, power, laplace, gumbel, logistic, lognormal, rayleigh, wald, triangular - one transform per draw.
    • binomial, negative_binomial, poisson, zipf, geometric, hypergeometric, logseries - discrete.
    • multinomial, dirichlet, multivariate_normal (svd / eigh / cholesky), multivariate_hypergeometric (marginals / count) - multivariate.
    • pareto / power stay within 2 ULP / ceil(2/a)+3 ULP of NumPy, whose C runtime expm1 has no portable twin; multivariate_normal is byte-identical with the OpenBLAS backend and may pick the other eigenvector sign without it.
  • Array-valued distribution parameters - the 29 parameterized samplers of both RandomState and Generator take NDArray parameters, broadcast against size as NumPy does, byte-identical streams - 39d446e0 (+ 5f204884).
    • integers, randint - array bounds and BigInteger bounds, so g.integers(0, BigInteger.Pow(2, 64), dtype: np.uint64) ports verbatim - 0e39714b.
  • Legacy RandomState gains NumPy's form over any bit generator (np.random.RandomState(new PCG64(42))), the dict get_state(legacy: false) / set_state, tomaxint, ranf and sample - f87b4d04.
  • np.ma - NumPy's masked-array module on a new NDMaskedArray type, bit-exact vs NumPy 2.4.2 and gated by a 68,860-case differential corpus - 5f1476f5 (+ ba15bad4, bf66bc0f, b33b6515, 3fb1d897, e0d6a81c, 67f8f96e, 7a262afc).
    • The ufunc wrappers, masked reductions, operators, creation, extrema and manipulation - the core.
    • average, median, dot, the set operations, sort, unique, cov, corrcoef, convolve, correlate, polyfit, apply_along_axis, apply_over_axes - the extras.
    • The masked indexer, hard masks and byte-exact masked str / repr; only the structured/record-dtype functions are missing.
  • np.emath - automatic-domain sqrt, log, log2, log10, logn, power, arccos, arcsin, arctanh (an out-of-domain input gives the complex result instead of NaN) - a3666a5f.
  • 48 new np.* functions, NumPy 2.4.2 parity, each gated by the differential-fuzz oracle -
    • hypot (correctly rounded: float32/float16 bit-exact, float64 bit-exact on ~91% of inputs and 1 ULP more accurate on the rest), divmod, fmod, remainder, float_power, heaviside, spacing, frexp, ldexp, gcd, lcm, fabs, signbit, fix - elementwise math ufuncs - 5c667b09, df232348, 44636747, 6e0e5ea2, ade8d6b5, 83e0e43e, a15f0ee4, 68a84d3c, 46f5f651.
    • sinc, i0, real_if_close, unwrap - special functions and signal helpers - c3be0b7f, a0c6e2fe, 877a36b2, b1937c51.
    • bartlett, blackman, hamming, hanning, kaiser - window functions, bit-identical - 0e9dbc42.
    • trapezoid, gradient - numerical integration and differentiation, bit-exact - 22d5f4d3 (+ f6037bc8, 4898b40a).
    • histogram, histogram_bin_edges, histogramdd, histogram2d - histograms - 2693d83e (+ 0e80bc89).
    • packbits, unpackbits, bitwise_count - bit packing and population count - 3936fd0d, bab51edb.
    • piecewise, putmask, put_along_axis - indexed and conditional assignment - 76b2ab9d, 86cc28b8, f417827e.
    • vectorize, frompyfunc, apply_along_axis, apply_over_axes, broadcast_shapes - functional and shape helpers - b1841830, 53b7a149.
    • array_equiv, may_share_memory, shares_memory (with TooHardError) - array predicates - 38f9a14e, 4cd3b72f.
    • logspace, geomspace - log-scale ranges, bit-exact - d48737ec (+ 3a367520).
    • binary_repr, base_repr, typename - string helpers - c2f129b5, 7cd38a97.
  • Existing APIs gain NumPy's parameters -
    • maximum, minimum, fmax, fmin - out= / where= - e5afe84c.
    • left_shift, right_shift, logical_and, logical_or, logical_xor, logical_not, modf - out= / where= / dtype= - 3384a0b5.
    • array_equal - equal_nan= - 38f9a14e.
    • indices, indices_sparse - long[] / Shape dimensions (np.indices(arr.shape)), and mgrid past 2^31 elements - 07c7b94c.
  • np.evaluate / NDExpr - the fused-expression DSL covers most of the ufunc surface and every reduction, NumPy-typed and bit-exact (29,563-case oracle tier) - 6175a017, bd0091fb, e1400974, 2f02d65d, eb1921e1, fbef1c4e, c5422cfc, dcbe7079.
    • Maximum, Minimum, FMax, FMin, Fmod, CopySign, NextAfter, LogAddExp, LogAddExp2, Hypot, Heaviside, Gcd, Lcm, LeftShift, RightShift - binary nodes.
    • Positive, Conjugate, Fabs, Spacing, SignBit, IsPosInf, IsNegInf, BitwiseCount, Rint, Real, Imag, Angle, Cast, Round(decimals) - unary nodes.
    • LogicalAnd, LogicalOr, LogicalXor, Select, Clip, and the < / > / <= / >= operators on NDExpr - logic and selection.
    • Any, All, CountNonzero, NanSum, NanProd, Ptp, NanMin, NanMax, ArgMax, ArgMin, NanMean, Var, Std (ddof), the weighted Average, axis forms and keepdims - reductions.
    • dtype=, casting=, order=, where= - keywords; expr.Compile() returns a reusable CompiledExpression.
    • If, Switch, Mux, Saturate, NanTo, Coalesce, Relu, LeakyRelu, Elu, Sigmoid, Swish, Softplus, Gelu, Step, Nand, Implies, IsClose, Between, Bucketize, Lerp and 18 more - 38 decision and activation combinators that fuse like any node - d3057b39 (+ 7aba9b4a).
  • DType system - NumPy's NEP 40-50 two-level dtype model (dtype classes and dtype instances), the descriptor every dtype parameter and NDArray.dtype now use - 896f02b2, c916c579, 72e4d29c.
    • np.dtypes (Float64DType and the other dtype classes), np.datetime_data - the class module.
    • num, kind, char, itemsize, str, descr, byteorder, newbyteorder, isbuiltin, isnative, flags, type_num (NPY_TYPES) - the full numpy.dtype surface.
    • Non-native byte orders (np.dtype(">i4")) and datetime64 / timedelta64 descriptors that parse, print, promote and cast-check (no datetime storage yet).
    • promote_types, result_type, can_cast, issubdtype, isdtype - DType overloads with NEP 50 weak scalars.
  • Typed iteration and .NET interop -
    • np.flat<T>(a) / a.flatiter.AsTyped<T>() - typed by-ref C-order flat iteration - fa2e8521.
    • np.ndenumerate<T>(a).AsRef(), np.ndindex(...).AsSpans() - allocation-free enumeration - afb0d5ce.
    • nd.Unsafe.Span<T>(), Memory<T>(), Bytes(), Pointer<T>(), TryGetSpan<T>(), TryGetMemory<T>() - zero-copy BCL views of the buffer (they do not keep it alive: hold the NDArray); nd.AsEnumerable<T>() and the typed iterators' ToArray() / CopyTo(Span<T>) - safe copies - 37a112be.
  • Py.RegisterNumSharpCodec() (NumSharp.Interop.pythonnet) - the one-call codec registration, spelled on pythonnet's own Py (a C# 14 extension member) - f18aaa23 (+ b648ee53).
  • NumSharp.Collections - ConcurrentOrderedDictionary<TKey,TValue>, ConcurrentOrderedCompactDictionary<TKey,TValue> (about a third of the memory) and the lock-free OrderedDictionary<TKey,TValue>: insertion-ordered, index-addressable dictionaries, plus a vendored NumSharp.Collections.Concurrent.ConcurrentDictionary<TKey,TValue> with by-ref access - 39e5ba17, 47c12604, ba5c347d (+ 04677203, 873900b1, 478b5310).

🧩 ndarray surface

  • ToArray<T>() - a T other than the dtype converts like astype (it threw ArrayTypeMismatchException), and every layout copies in one pass (~5.8x faster than before on the measured set) - 73514657.
  • ToMuliDimArray<T>() / (Array)nd - one pass over any layout; float16, complex128, a T other than the dtype (converted like astype), 0-d and empty decimal arrays now work (they threw) - c243201a.
  • ToJaggedArray<T>() - any rank (rank 6 came back wrong, deeper ranks threw), any layout (a transposed view gave wrong values, a column slice read out of bounds) and a T other than the dtype; 10.7x faster with 7.7x less memory - 4231d2d0 (+ 6f2ec473, ac89847a).

⚡ Performance

Ratios are NumPy ÷ NumSharp - higher is better (x2 = twice NumPy's speed); xLOW->xHIGH spans the worst→best measured cell across sizes and dtypes.

  • x0.70->x3.1 - np.max / np.min / np.ptp on NumPy's exact reduction schedules, slab reductions folding eight reduced indices per pass (axis cells on non-C layouts were x0.28-x1.0; L3-bound 4M cells of 1- and 4-byte lanes still trail) - 1d79ce61, f8f91b3d.
  • x0.41->x4.7 - np.nanmax / np.nanmin (float32/float64 axis cells were x0.15-x0.50, flat C now x2.89 / float32 x3.33; float16/complex128 non-C axes run NumPy's scalar loop at x0.41-x0.69) - de6333dd, 268d91fd.
  • x0.77->x106 - np.argmax / np.argmin along an axis walk the strides instead of a per-output div/mod walk (axis 0: float64 x4.4-x8.8, integers x10-x16, bool x95-x106; memory-bound contiguous rows at 4M x0.77-x1.04), and the 64-bit flat argmax runs NumPy's single-pass tournament - ceac36f2.
  • x0.89->x2.65 - np.mean / np.sum along an axis with a large output: the identity seed and the mean's division run as typed block passes (np.mean(a, axis=0) of 3x1M float32 was x0.29; mean(axis=1) over 3-wide rows stays below NumPy) - 3f22d30a.
  • x1.43->x1.83 - np.convolve long products without the OpenBLAS backend (1000x1000 float64 was ~x0.9, complex128 x0.25) - 38526155.
  • x1.15->x5.9 - complex128 divide, multiply and integer-exponent power, now on NumPy's own formulas - dc9caff5, 338d260a, 3e7d6b33.
  • x1.26->x4.0 - np.modf plain, out= and promoted calls (out= at 10M was x0.40) - 3384a0b5.
  • x1.0->x16 - np.array_equal / np.array_equiv fused early-exit compare (x15-x16 at N=1000, parity at 100K, instant on an early mismatch) - 44604193.
  • x1.05->x6.4 - every np.random bit generator fills through a bulk layer: all 105 engine x op x size cells beat NumPy - aa9905b7.
  • x1.08->x5.58 - legacy RandomState samplers (hypergeometric was x0.96, zipf x1.12, multinomial x0.97; the samplers bound by the C runtime's pow / log / exp stay at x1.17-x1.49) - 910a6480.
  • x0.65->x17.6 - np.evaluate comparison, logical, where and bool trees vectorize instead of running the whole tree scalar (a > 0.5 was x0.16) - 939d0636.
  • x1.0->x22.2 - np.evaluate mixed-dtype trees vectorize through lane groups (i4*2+f8 x14.2 and f4*f8+f4 x22.2 at 100K; 1K cells sit at the per-call floor) - 3631fcc7.
  • x0.86->x53 - np.evaluate reductions stream block by block instead of materializing their child, and Any / All stop at the deciding block (any(a > b, axis=0) was x0.37; count_nonzero(a > b) at 100K sits at the read-bandwidth ceiling) - ee639862, 0f8819b7, 129fbaba, 78692d91, ceac36f2, 0f78f6e2.
  • x2.06 - a small np.evaluate call skips the iterator like NumPy's trivial loop: a*b+c at n=8 runs in 255 ns vs NumPy's 526 (out=: 75 ns) - 30dbfe0d (+ df799381).

🎯 Parity & Fixes

  • np.max / np.min / np.ptp, flat and along an axis, run NumPy's own reduction schedules - the sign of a ±0 tie and a canonical-vs-payload NaN now match on every non-broadcast layout - 1d79ce61 (+ 129fbaba, f8f91b3d).
  • np.nanmax / np.nanmin are NumPy's fmax.reduce / fmin.reduce bit for bit: complex128 skips NaN (it propagated it), float16 on non-C layouts read the wrong elements, and an all-NaN float16 slice returns its first NaN - de6333dd (+ 268d91fd).
  • complex128 arithmetic bit-exact with NumPy - divide (1 ULP off on ~33% of operands), multiply (NumPy's fused simd_cmul), power with an integer exponent (up to ~1e14 ULP off) and abs (NumPy's fused simd_cabsolute; 35.5% of values were off) - dc9caff5, 338d260a, 3e7d6b33, d1b89c11.
  • percentile / quantile / median bit-exact - NumPy's two-branch _lerp in each dtype's arithmetic, median as the mean of the middle pair, and a view at a non-zero offset no longer reads the wrong data - d2c8d26e.
  • np.round / np.around with decimals != 0 port PyArray_Round - negative decimals threw, complex values stayed unrounded, float32 rounded in double, and the result dtypes follow NumPy - 6db8238a.
  • sum / prod of a size-0 or size-1 array widen the accumulator by NEP 50 like any other size - 23ea2ad9.
  • np.array_equal(a, a) with a NaN in a returned True; NumPy returns False unless equal_nan - 38f9a14e.
  • int64-vs-uint64 comparisons are exact on every path (they went through float64) - 53602da9.
  • float16 mod / floor_divide / fmod by zero give ±inf like NumPy (they gave NaN), and float16 add / subtract / multiply / divide on the out= / where= / strided paths compute in float32 like NumPy's HALF loops (they double-rounded through double) - df232348, b53a0cc1.
  • float16 signalling NaNs stay signalling across every float16 cast route, uint64 at or above 2^63 converts to float64/float32 with one rounding on .NET 8 (it rounded twice, inside mixed uint64/float arithmetic and comparisons too), and maximum / minimum of float16 with float32/float64 keep the NaN's bits - ac89847a.
  • np.linalg.eig / eigvals of a float32 matrix with complex eigenvalues return NumPy's complex64 values (they kept double precision) - 0c02740c.
  • Use-after-free fixed: np.real / np.imag / .real / .imag, view(dtype) and getfield did not keep their owner's buffer alive, so after the owner was disposed (as inside eigvals) the next allocation overwrote them; resize(refcheck: true) now sees such views - 59890501.
  • np.convolve / np.correlate without the backend are byte-exact on float32 positions under 32 terms and complex128 positions under 8 (NumPy's scalar sdot / zdotu regimes), and complex products with inf/NaN follow NumPy (convolve([1+0j], [inf+0j]) is nan+nanj) - 38526155, 058f4f27.
  • np.arctan2(complex) corrupted the heap and crashed the process; it raises NumPy's TypeError - dc2fc075.
  • Complex input to cbrt, floor, ceil, trunc, deg2rad, rad2deg, floor_divide and mod raises NumPy's TypeError (it raised NotSupportedException, and a zero-size complex input returned an empty array); IncorrectTypeException now derives from TypeError - 9a1eeb7e, 2e205a1f.
  • np.modf accepts int, float16 and bool input like NumPy (it rejected them), float16 byte-exact - 3384a0b5.
  • np.trace of uint8/uint16/char accumulates in uint64 like NumPy; ndarray.flat of a 0-d array has shape (1,); np.take, ndarray.item() and len() of a 0-d array raise NumPy's verbatim texts - e84fd2f1.
  • np.argsort of an empty array with sort axis >= 2 overran the heap - 74656f13.
  • A boolean-mask selection owns its result (flags.owndata), like NumPy - 32162b64.
  • ! (logical not) on a non-contiguous array returned wrong values - 76b2ab9d.
  • Memory leaks fixed in np.choose, tofile, tobytes, ogrid, poly1d, polyfit, multivariate_normal, apply_along_axis / apply_over_axes, flatiter slice/fancy access, cross-dtype SetData and the float16 sum / mean temporaries; NDArray.ReplaceData could free another array's buffer - 1dd23bbe, 6f2ec473, 7e60327d, e84fd2f1.
  • Generator parity (it shipped in 0.70.0) - 1855ea9e.
    • integers with a high past the dtype returned constant data (integers(0, 257, dtype: np.uint8) was all zeros); it raises NumPy's error.
    • Concurrent draws on one bit generator tore its state (72-99% of 4-thread values belonged to no seed's stream); draws hold the generator's lock, as in NumPy.
    • choice without a size returns one element (it returned shape (1,)); permuted follows memory order and casts into out= safely; a -0.0 scale, shape or range is rejected; a () size keeps float32.
  • Legacy RandomState parity - c1e0feb5, f87b4d04.
    • An unseeded RandomState() took its seed from Environment.TickCount, so instances made in the same millisecond were identical; it uses OS entropy.
    • choice ignored replace=False and searched the probability CDF from the wrong side; randint(dtype: bool) crashed.
    • Every sampler is NumPy's legacy-distributions.c line for line: gamma(shape<1), binomial, negative_binomial, poisson, f, pareto, power, standard_cauchy, multinomial, hypergeometric, zipf, geometric and vonmises used other algorithms.
  • SeedSequence - a typed null seed array seeded the empty sequence (the same stream on every run; now fresh entropy, NumPy's None), and SeedSequence(new[] { -1 }) silently seeded 0xFFFFFFFF instead of raising - 0e39714b, 65bdea7b.
  • result_type of four or more operands dropped every other pair (result_type(int8, int8, float64, int8) was int8; it is float64) - 896f02b2.
  • np.evaluate - flat and axis Sum / Mean / Prod / Min / Max follow NumPy's own schedules bit for bit (they matched the value only), a flat Mean of an integer child is NumPy's buffered float64 sum, NEP 50 typing gaps are closed, and two closures over one lambda no longer share a kernel - 53602da9, e1400974, 129fbaba, df799381.
  • .NET 8: string slicing (nd["1:3"]) re-parsed its pattern on every call because the static Regex cache thrashed - 9-18 µs -> 0.6-2.8 µs per parse - 28fa9d79.

🧰 Testing & Tooling

  • Differential-fuzz oracle 116K+ -> 406,172 corpus cases: new tiers for every numpy.polynomial unit (139,882 cases), np.ma (68,860), np.evaluate (29,563), np.emath, ndarray instance members, and complex128 / Decimal coverage, with a dtype-spread gate (at least 4 dtypes per op) and an applicability gate - 57545200, 67f8f96e, 7a262afc, e84fd2f1, 55a2a6a7, 192f456c, 1dd23bbe.
  • The random-API oracle replays every public np.random overload against NumPy 2.4.2 by its exact C# signature (493 members, 5 engines x 10 seeds, the full post-call state), with coverage gates and a nightly soak under 10 fresh seeds - 20a5b60d, 1c7ac693, 8f70af7c, ecdab2fb.
  • The coverage join credits each API from the oracle contracts that name it, strictly: a new op key or corpus file that resolves to nothing fails the docs build - fd058f03.

💥 Breaking Changes

  • DType replaces System.Type as the dtype currency - c916c579, 68a02258.
    • NDArray.dtype returns DType (was Type), and np.float64, np.int32 and the rest are DType descriptors (were Type); DType -> Type is explicit (a.dtype.type), and a.dtype.Name is a.dtype.name.
    • Each dtype parameter takes one DType: the Type / NPTypeCode twins are gone (a Type, NPTypeCode or dtype string still converts implicitly), and linspace / nancumsum's typeCode: is dtype:.
    • The promotion family (result_type, promote_types, find_common_type, common_type, ...) returns DType, and isdtype rejects issubdtype's vocabulary ("floating"), as NumPy does.
  • finfo.dtype / iinfo.dtype return DType (were NPTypeCode), and the pythonnet dtype maps take DType - ea080d8e.
  • Casting-rule strings are case-sensitive like NumPy ("Safe" raises ValueError) - 896f02b2.
  • On Windows, np.int32 / np.int64 report the local NumPy's LLP64 type numbers and chars (7 / 'l' and 9 / 'q'; they were hard-coded to Linux's 5 / 'i' and 7 / 'l') - 72e4d29c.
  • A write to a read-only array raises NumPy's ValueError (was NumSharpException); it derives from ArgumentException, so catch (NumSharpException) no longer catches it - f319b657.
  • np.maximum / np.minimum / np.fmax / np.fmin take out as the third positional argument (NumPy's order), not dtype; pass dtype: by name - e5afe84c.
  • np.logical_and / logical_or / logical_xor / logical_not return NDArray (was NDArray<bool>), and np.modf's tuple field Intergral is Integral - 3384a0b5.
  • ndarray.shape = ... raises NumPy's AttributeError when the layout cannot be reshaped in place (it silently copied); use reshape - 6f2ec473.
  • np.nanmax / np.nanmin of an empty reduction raise NumPy's "zero-size array to reduction operation fmax which has no identity" (they returned a 0-d NaN) - de6333dd.
  • new MT19937(42) seeds like NumPy's np.random.MT19937(42) (through SeedSequence), not like RandomState(42); mt._legacy_seeding(42) is the old stream. MT19937's non-NumPy members (NextUInt32, NextDouble, NextInt, NextLong, Next, NextBytes, Seed, SeedByArray, SetState, Key, Pos, Clone) are gone, and Pcg64StateData is PCG64.State - c1e0feb5.
  • Seeded legacy RandomState streams of gamma(shape<1), binomial, negative_binomial, poisson, f, pareto, power, standard_cauchy, multinomial, hypergeometric, zipf, geometric, vonmises (and gumbel / logistic / rayleigh at scale 0) change to NumPy's values; binomial(int n) is binomial(long n); inputs NumPy accepts (dirichlet([]), a NaN alpha, negative_binomial(p: 0)) no longer throw - f87b4d04.
  • Legacy integers are NumPy's LP64 C long: randint defaults to int64 (was int32), random_integers / permutation / choice / multinomial return int64, NumPyRandom.Seed is a uint (was int), and SeedSequence.generate_state returns an NDArray (was Array) - eebe85a6.
  • NumPy's parameter names: Generator(bit_generator) (was bitGenerator), PCG64(seed) and the other bit generators (was seedSeq, now any ISeedSequence), and default_rng(seed) (was bitGenerator / generator) - named-argument calls must be renamed - eebe85a6, 1c7ac693.
  • The non-NumPy NumPyRandom.uniform(NDArray low, NDArray high, DType dtype) is removed; two-array calls bind NumPy's broadcasting overload (float64; .astype() the result for another dtype) - 39d446e0.

Nucs added 30 commits September 13, 2026 15:43
…ate-cap scenarios

Converts the differential-validation sweep (2873 recipe cases + 1408 flag-state
cases, all bit-exact vs NumPy 2.4.2) into permanent regression coverage for the
scenarios the original 14 tests did not pin:

- CandidateCap_ThresholdMatchesNumPy — NumPy's own test case
  (test_mem_overlap.py): x[:, ::2, ::3] vs x[:, ::3, ::2] genuinely overlap but
  the exact Diophantine DFS needs >2 candidates, so shares_memory raises
  TooHardError at max_work 0..2 and decides True at >=3 (byte-identical threshold
  to NumPy); may_share_memory folds every undecided budget into True.
- BroadcastViews_ShareByCollapsedExtent — a broadcast view's byte extent is only
  the source row it repeats (stride-0 axes add nothing), so two broadcasts of the
  same row share while broadcasts of disjoint rows fail even the bounds check;
  also pins that broadcast views are read-only (WRITEABLE=false).
- ReadOnlyView_SharesLikeWriteable — the WRITEABLE flag is irrelevant to sharing
  (the functions only read strides).
- FContiguousLayouts_ShareOnlyWhenAView — a transpose is an F-contiguous VIEW
  (OWNDATA=false, shares) while asfortranarray of a C-contiguous 2-D array is a
  fresh F-contiguous COPY (OWNDATA=true, does not share).
- MaxWork_LongMaxValue_IsValidHugeBudget — NumPy's max_work=10**100 OverflowError
  is a Py_ssize_t/bignum artifact unrepresentable in C# long; long.MaxValue is the
  analogous extreme and is a valid effectively-unlimited budget.
- IntIndexZeroD_SharesInNumSharp_ButNumPyReturnsScalarCopy ([Misaligned]) —
  documents that a single-int index is a 0-d VIEW in NumSharp (shares) but a
  scalar COPY in NumPy (shares nothing); an indexing-semantics divergence, NOT a
  shares_memory difference. The reshape(()) 0-d view shares in both.

Validation summary (all vs NumPy 2.4.2, no code change needed): flags
(C/F/O/W/num) match across 26 layout builders x 4 dtypes; every shares_memory /
may_share_memory result matches across ~4300 recipe/flag/pairwise/max_work
assertions; the only divergences are the two documented above (int-index 0-d
construction, 10**100 budget), neither a function-behavior difference.
…nt (validation-found parity fix)

Adversarial validation of the logspace/geomspace family (bit-constructed special-value differential vs NumPy
2.4.2: 310 logspace cases + 234 geomspace-real cases + 14 complex + error taxonomy) surfaced ONE real
divergence: geomspace(NaN, 1.0, 4) gave [nan, nan, nan, 1.0] where NumPy gives [nan, nan, nan, nan].

Root cause: the real core computed out_sign as `start < 0 ? -1 : 1`, which yields 1.0 for a NaN start
(np.sign(NaN) is NaN, not 1), and it wrote the ORIGINAL start/stop at the endpoints as a shortcut for
`x_r * out_sign`. That shortcut is only valid when out_sign is ±1 — for a NaN start NumPy's trailing
`result *= out_sign` turns EVERY element (the finite stop endpoint included) into NaN.

Fix (both changes keep the ±1 paths BIT-IDENTICAL — double negation is exact):
  - out_sign = np.sign(start): `double.IsNaN(start) ? start : start < 0 ? -1 : 1` (start!=0 guaranteed,
    -0.0 already rejected), so a NaN start yields a NaN out_sign.
  - the loop now computes the rotated sequence (startR at 0, stopR at the endpoint, 10**exp in between) and
    multiplies EVERY element by out_sign uniformly, exactly as NumPy does — startR*(+1)==start and
    startR*(-1)==start bit-for-bit for finite/inf inputs, and NaN everywhere for a NaN start.

A NaN STOP with a finite start was already correct (the stop endpoint IS NaN). +inf/-inf starts stay correct
(sign(±inf)=±1). The complex core already multiplied endpoints by out_sign, so it needed no change.

Validation result after the fix: logspace 310/310 + geomspace-real 234/234 BIT-EXACT, complex 14/14 allclose
(worst-rel 1.0e-15), errors match (ValueError verbatim; AxisError type+text + documented .NET paramName
suffix). Pinned by 3 new unit tests (Geomspace_NaNStart_IsAllNaN, Geomspace_Infinities,
Logspace_SpecialValues); the class is 23/23 green net8+net10, creation oracle corpus unchanged (its geomspace
cases use positive values, unaffected).
…2.4.2

Implements np.fabs, the FLOAT-ONLY element-wise absolute value (NumPy's
`fabs` ufunc, `TD(flts, f='fabs', astype={'e':'f'})` — float loops `efdg`
only). It is the float sibling of the existing abs/absolute and completes
the absolute-value family (copysign/hypot/signbit/nextafter/spacing already
shipped).

Design — fabs IS abs on the float loops (both clear the IEEE sign bit), so
it REUSES the fully-SIMD `UnaryOp.Abs` kernel rather than adding a new op:
`Default.Fabs` -> `ExecuteUnaryOp(nd, UnaryOp.Abs, floatType, out, where,
ufuncName:"fabs")`. It differs from abs in exactly two dtype-level ways:

  * PROMOTES bool/int to float (NEP50 tier via the shared
    ResolveUnaryFloatReturnType: bool/i8/u8->f16, i16/u16/char->f32,
    i32/u32/i64/u64->f64, f16/f32/f64/decimal preserved) — where abs
    preserves int. The promoting path casts int->float BEFORE the abs, so
    fabs(int.MinValue) is the exact float magnitude, never a wrapped
    integer abs.
  * has NO complex loop (abs maps complex->magnitude; fabs REJECTS complex),
    with NumPy's exact three error paths (all probed 2.4.2):
      - no dtype=       -> TypeError "ufunc 'fabs' not supported for the
                           input types, ... casting rule ''safe''"
      - dtype=<float>   -> "Cannot cast ufunc 'fabs' input from complex128
                           to <float> with casting rule 'same_kind'"
      - dtype=complex/int -> "No loop matching the specified signature and
                           casting was found for ufunc fabs"
    where=-bad wins over the complex-input error (ValidateWhereMask first).

BIT-EXACT with NumPy 2.4.2 across all 12 NumPy-representable non-complex
dtypes x 26 layouts, verified over curated edges + 4096-element random
sweeps: -0.0->+0.0, +/-inf->+inf, and a negative/payload-bearing NaN whose
sign bit is cleared while the payload is preserved
(0xfff8...abcdef -> 0x7ff8...abcdef).

New seam: `ExecuteUnaryOp`/`ExecuteUnaryUfuncInto` gained an optional
`ufuncName` override (defaulted, non-breaking) so a kernel reused under
another ufunc name reports the right name in the out= cast error —
UfuncName(UnaryOp.Abs) is "absolute", so without this fabs's out= error
would leak "absolute".

Perf (NPY/NS, Release, best-of-15): faster than NumPy on every measured
cell — NumPy's fabs is a SCALAR CRT loop (no SIMD dispatch, unlike
absolute), while NumSharp rides Vector.Abs. float16 4.5-24.8x, float64
1.75-13.9x, int32 1.6-10.8x, out= paths 2.5-14x. The sole sub-1.5x cell
(fresh-alloc float32@100K, 1.19x) is the shared allocation floor — its
out= variant is 12.2x, confirming the kernel is optimal.

Oracle: fabs added to UNARY_OPS (gen_oracle.py) + OpRegistry, cascading
into the unary tier (338 value cases, all -> f16/f32/f64), Char coverage
(uint16 proxy), and errors_full (26 complex-input TypeError cases). The
surface gate auto-covers it (direct corpus op key). errors_full was NOT
regenerated wholesale (that surfaces PRE-EXISTING spacing/fmod complex
error-message divergences from other ops, whose committed corpus is stale)
— only the 26 fabs cases were appended onto HEAD's content, keeping the
change fabs-isolated.

Files: Math/np.fabs.cs, Backends/Default/Math/Default.Fabs.cs,
Backends/TensorEngine.cs (abstract Fabs), the two ExecuteUnary* seam edits;
test/NumSharp.Tests/Math/np.fabs.Test.cs (21 tests). All fabs gates green:
unary tier bit-exact, errors_full green, 21 unit tests + 176 sibling
abs/fix/signbit/ufunc-overload tests pass.
… oracle, typed vector emission, per-root compiled program)

Brings the four commits of `exprs` (docs 5e65ed9, P0 53602da, P1 939d063, P3 dcbe707; plan
docs/plans/ndexpr-evaluate.md) onto master:

  * P0 — correctness + the differential gate: int8 zero-push (Where/LogicalNot over int8 threw),
    Abs(complex) → float64, exponent-array sign check, complex Min/Max reductions, typed literals
    (bool/uint/ulong/Half/decimal/Complex/char spellings with NEP50 weak/strong kinds), exact
    int64/uint64 comparisons, engine dispatch on the first array's engine, Call dtype map to all 15;
    the evaluate.jsonl oracle tier (14,702 cases — the unfused NumPy chain evaluated node by node is
    the oracle for a fused tree) with MisalignedRegistry E1–E5.
  * P1 — "vector v2": one lane dtype W per kernel, Boolean-typed nodes as lane masks, the fused
    inner-loop shell (DirectILKernelGenerator.InnerLoop.Fused.cs: bool operand expansion, packed bool
    output, hoisted stride-0 operands, AVX2 gather), so comparisons / where / min-max / logical /
    predicate trees vectorize (a>0.5 0.16× → 1.3–1.45× NumPy; leaky relu 17.6× at 100K).
  * P3 — the per-root compiled program (NDExprProgram), CompiledExpression handle
    (NDExpr.Compile()/Compile(types)), N-ary single-allocation input broadcast with NumPy's
    every-operand error text; a hoisted tree 1006 → ~410 ns / 584 B at n = 8.

Conflicts (both additive, same anchors as the branch point): OracleSurfaceCoverageTests.cs — the
sibling-owned list keeps master's new entries and drops "evaluate" (the evaluate.jsonl tier now
covers it); gen_oracle.py — both appended generators kept (gen_windows gets its own `return cases`,
the shared one closes gen_evaluate) and both dispatch arms with the union of the mode list.

The follow-up commit on top closes the two residuals the branch left (the rebuilt-tree spelling and
the 0-d operand cost) with a structural program cache and hoisted parameters.
…l parameters — the rebuilt-tree spelling costs what a hoisted tree costs

Why np.evaluate impacted performance (measured, n = 8, ns / managed B per call, pinned P-core,
DOTNET_TC_CallCountingDelayMs=0; probe benchmark/fusion/probes/fc_probe.cs):

  * the natural spelling `np.evaluate((NDExpr)a * b + c)` builds a NEW root every call. On the old
    NDExpr that cost 995 / 3056 (bind clone + Dictionary typing pass + StringBuilder kernel key +
    broadcast + iterator) against 349 / 944 for the unfused `a*b+c` chain, so fused lost to the chain
    it replaces on 10 of 15 probe rows at 1K. The Phase-3 per-root program cache (branch `exprs`)
    served a HOISTED tree at 405 / 592 but made the rebuilt spelling WORSE — 1580 / 3136 — because a
    fresh root misses the per-root slot by construction and `NDExprProgram.Build` now also ran the
    vector plan on every miss. Every in-repo consumer (np.sinc, the windows, trapezoid, gradient)
    spells its tree inline.
  * a 0-d input (`a*b+k`, the plan's own "parameter form" for a runtime scalar) rode the iterator
    as a stride-0 operand: 681 / 1376 prebuilt (1255 / 3752 old NDExpr) — ~250 ns of NDIter stride-0
    construction per call — AND 8–10 % per element on BOTH kernel paths (the fused shell's per-load
    broadcast branch on the SIMD path; the strided fallback's per-element stride multiply on the
    scalar path: the hanning tree over a 0-d M-1 ran 408 µs vs 372 µs literal at 100K). The
    alternative, a LITERAL, bakes the value into the IL and the kernel key: np.hanning(M) compiled
    one kernel per distinct M (20 lengths → 20 kernels, 3.6 ms), each a permanent kernel-cache entry.
  * a latent correctness bug in the same key: a Call node's signature omitted the delegate SLOT for
    captured delegates, so two closures over one lambda body shared a kernel and the second silently
    ran with the first's captured state.

What landed:

  1. Global structural program cache (NDExpr.Structure.cs, NDExprProgramCache in NDExpr.Program.cs).
     Every node folds its identity into a 64-bit order-sensitive hash (HashStructure) while
     collecting the distinct array leaves in BindArrays' first-visit order, so an ArrayNode hashes
     as the InputNode index it binds to; the hash + dtype signature + 0-d mask + ForceScalar picks a
     candidate in a process-wide ConcurrentDictionary, and the hit is VERIFIED node by node against
     the candidate's bound tree (StructureEquals: kinds, ops, literal bits, reduction axis/keepdims,
     Call kind/method/slot, operand dedup) — a hash-only match would run the wrong kernel. Two
     allocation-free walks (a [ThreadStatic] operand scratch, reset before return) replace bind +
     typing + string key. One entry per hash (a genuine collision evicts, never mis-pairs); 4096-entry
     cap with a wholesale clear (the compiled kernels stay in the kernel cache). The per-root slot
     stays as the zero-walk path and now publishes (program, inputs) as one immutable object.
     NDExprProgram no longer holds NDArrays — it is shared across every root of one structure.
  2. 0-d inputs as kernel PARAMETERS (NDExpr.Params.cs; plan item 3.3 done properly). np.evaluate
     copies each 0-d input's element into the kernel's aux block (16-byte slots; after the
     accumulator slot for reductions) and the kernel prologue — a new optional `prologue` action on
     CompileFusedInnerLoop, and inline in the two reduce kernels — loads it ONCE into a local: the
     scalar for the scalar body, its broadcast Vector<W> (a lane mask for a bool in W-mode) for the
     vector body. InputNode reads it through NDExprCompileContext.LocalOfInput, so it costs exactly
     what a literal costs per element, the iterator never sees it, and one kernel serves every value.
     Typing is untouched (a 0-d array is the strong scalar it always was). The mask is part of the
     program identity (hash, verification, kernel cache key suffix `|p0110`), re-validated on the
     per-root fast path (ndarray.resize of a 0-d array re-resolves; a Compile()d handle refuses with
     a clear message), and empty when EVERY input is 0-d (0-d result; iterator path unchanged). A
     parameter is read BEFORE the pass, so one aliasing `out` keeps its original value (NumPy's
     COPY_IF_OVERLAP semantics). A dtype-pinned positional Compile(types) handle knows no shapes and
     streams everything (a 0-d operand still works there, as a stride-0 stream).
  3. Call slot identity: the slot joins the kernel signature for BOTH slot-backed kinds, and
     DelegateSlots dedups registration by identity (target reference + method; targets by reference)
     so a field-held delegate rebuilt into a fresh tree per call keeps one slot and one kernel — a
     closure ALLOCATED per call is a new identity (documented: hold delegates in fields).
  4. np.windows: every M-dependent value (M-1, alpha, beta, i0(beta)) rides as a 0-d operand — one
     kernel per window kind for every M, bit-identical (the windows.jsonl tier and the 22 unit tests
     are unchanged and green).

Measured after (same probe):
  * rebuilt `a*b+c` 464–508 / 864 (old 995 / 3056; P3-only 1580 / 3136); prebuilt 397–414 / 592;
    `out=` rebuilt 353–363 / 456; `Sum(a*b)` rebuilt 368–388 / 704 (was 715); `Where(a>b,a,b)`
    rebuilt 448–456 / 880 (was 865); `a*b+k` 0-d prebuilt 388 / 624 (was 681 / 1376), rebuilt
    481 / 896 (was 1304 / 3904).
  * fused ≥ unfused on 13 of 15 rows at 1K with the rebuilt spelling (was 5/15; P3-only 6/15). The two
    losses, maximum(a,b) 0.59× and abs(a) 0.63×, are single-op trees against the direct kernel's lower
    fixed cost (plan P6.3).
  * consumers: sinc tree at 1K 4926 → 3385 ns (unfused chain 5601); hanning tree 5165 → 4281;
    np.bartlett(1000) 2591 → 1533; np.hanning(1000) 5681 → 5033; the hanning tree over a 0-d M-1
    now runs at literal speed per element (371.7 vs 372.1 µs at 100K, 3.72 vs 3.72 ms at 1M); 20
    distinct hanning lengths compile 0 kernels (0.46 ms) where the literal compiled 20 (3.64 ms).

Gates: NDEvaluateProgramTests (21 — per-root + structural cache identity: child order, op, literal
value/kind/sign bits, dedup pattern, dtype signature, reduce kind/axis/keepdims, embedded vs
positional sharing, capacity clear, Parallel.For thread-safety, Call closure/instance slots),
NDEvaluateParamTests (12 — every one of the 15 dtypes on both kernel paths, bool lane masks in W-
and byte-mode, flat + axis reductions, mixed-dtype promotion, a parameter aliasing out, the all-0-d
fallback, the resize guard, positional forms, complex/decimal 16-byte slots, two parameters), the
Oracle FuzzMatrix evaluate.jsonl tier (14,702 cases; its pp_scalar_* / scalar_0d layouts drive the
parameter path bit-exact vs NumPy 2.4.2) + windows.jsonl. Full suites: NumSharp.Tests 15,246 passed
and Oracle 153 passed with the identical 25 + 2 pre-existing failures the clean master base shows
(demo/port tests + the deferred oracle-surface classification), verified by TRX name diff.

Docs: .claude/CLAUDE.md "Fused Expressions" (the cache, the parameter rule "prefer a 0-d array over a
literal for a runtime scalar", the Call slot rule), docs/plans/ndexpr-evaluate.md §0.1 (6.1 + 3.3
landed; both residuals closed with numbers). benchmark/fusion/probes/fc_probe.cs is the probe that
produced every number above (sections D fixed cost / G consumer trees / C 1K rows).
…_equiv (EqualityScan)

The array-equality family computed all-equal as np.all(a == b): the comparison
kernel writes a full boolean temp (size x 1 byte) that the reduction then re-reads,
and — like NumPy — NEITHER side early-exits on the first mismatch. This adds a fused
single-pass fast path for the common case (same shape, same dtype, dense-contiguous,
equal_nan=false) that replaces both passes with ONE early-exiting SIMD scan and no
intermediate allocation.

EqualityScan (Backends/Kernels/EqualityScan.cs) — sibling of FiniteScan, same design
(generic Vector<T> body, JIT bakes the V128/V256/V512 width per instantiation; the
Address + offset*itemsize base rule for contiguous slices; 4x-unrolled with a
one-vector remainder + scalar tail):
  - Value vs byte comparison is load-bearing. Integer family + bool + char: value ==
    byte equality, so the whole buffer is scanned as raw bytes through ONE Vector<byte>
    path. Float/double/complex CANNOT byte-compare — two NaNs share a bit pattern yet
    must compare UNEQUAL (equal_nan=false: NaN != NaN) and -0.0 vs +0.0 differ in bits
    yet must compare EQUAL — so those use IEEE value comparison via Vector.EqualsAll
    (false for a NaN lane, true for signed zero), complex128 by reinterpreting each
    element as two interleaved doubles (NaN component -> unequal, +/-0 components equal).
  - Per-group early-exit: a mismatch near the front returns almost instantly — a win
    NEITHER the old composition NOR NumPy has.
  - TryAllEqual(a, b, out result) gates eligibility (same shape, same dtype, both C- or
    both F-contiguous, not Half/Decimal) and returns false for everything else, so
    broadcast / mixed-dtype / strided / Half / Decimal / equal_nan=true all fall back to
    the proven np.all(a == b) composition unchanged.

Wired at the STATIC np.* layer (Logic/np.array_equal.cs, Logic/np.array_equiv.cs) — the
public NumPy-API entry points — rather than the instance NDArray.array_equal, keeping the
change entirely in these two files + the new kernel (the instance method's composition is
unchanged and still correct for its rarer direct callers).

Correctness (bit-identical to the composition and to NumPy 2.4.2): the 800-case in-process
differential and the 160-case dtype x layout matrix from the family's landing both stay
green with the fast path engaged; a focused adversarial probe (NaN / +/-0 / inf / complex-
NaN / complex-+/-0 / both-F-contiguous / offset slice, all 13 fast-path dtypes) matches the
composition on every case; Half and Decimal correctly fall back. FuzzMatrix Logic tier
(the committed array_equal oracle cases) stays bit-exact through the fused path.

Perf (in-process A/B, fused vs the old composition, best-of-51, Release):
  N=1000    int32/f64  9.0x / 10.0x   (small arrays — the most common array_equal use;
                                        avoids the whole bool-temp + reduction setup)
  N=100000  int32/f64  1.27x / 1.12x
  N=10000000 int32/f64 1.19x / 1.03x  (memory-bandwidth bound: reading the two operands
                                        dominates, the bool round-trip is ~11% — same
                                        ceiling NumPy hits)
  mismatch@0 (any size)  effectively instant (early-exit) vs a full scan
Versus NumPy 2.4.2 (best-of-51): ~15-16x at N=1000 (NumSharp's low fixed overhead vs
Python dispatch), ~1.3-2.0x at 10M int32/f64, ~parity at 100K (cross-process regime
noise; the in-process A/B above is the trustworthy signal), and instant vs NumPy's full
scan on an early mismatch. The genuine-broadcast array_equiv cell is unchanged (it falls
back to the composition, whose broadcast-Compare-kernel ceiling is a shared-infra matter
tracked separately).

Gate: test/NumSharp.Tests/Logic/np.array_equal.FastPath.Test.cs (11 — NaN/signed-zero
byte-traps, complex, infinity, early-exit at many positions, both-F-contiguous, all
fast-path dtypes, offset slice, Half/Decimal fallback, array_equiv equal-shape vs
broadcast) + the existing family unit tests + the FuzzMatrix Logic tier, net8.0/net10.0.
…ufuncName param

Follow-up to 5f020a0. The initial fabs reused UnaryOp.Abs and threaded an
optional `ufuncName` string through ExecuteUnaryOp/ExecuteUnaryUfuncInto so
the out= cast error would report "fabs" instead of UfuncName(UnaryOp.Abs)
== "absolute". That parameter had exactly ONE caller (Default.Fabs) passing
a constant — a single-use knob added to the shared unary executor that every
unary ufunc flows through.

Replaced with a dedicated `UnaryOp.Fabs`, matching the codebase convention
that every operation is its own UnaryOp value with its own UfuncName entry
(there is no other name-override anywhere). The op drives BOTH kernel
selection and the error ufunc name, so a distinct enum is the natural
decoupling: UfuncName(UnaryOp.Fabs) => "fabs", and the KERNEL is Abs's —
Fabs aliases Abs at the 6 non-complex emit sites (scalar EmitAbsCall, the
Vector.Abs name-map, the SIMD gate, the f16-contiguous selector, decimal
scalar, f16 scalar). The 3 `Abs && InputType == Complex` sites are NOT
aliased — fabs rejects complex before dispatch, so it never reaches them.

This keeps fabs's specialness in the per-op dispatch layer (where op-specific
behavior already lives) and restores ExecuteUnaryOp/ExecuteUnaryUfuncInto to
their pristine signatures. Kernel cache: Fabs compiles its own byte-identical
kernels — a negligible one-time JIT cost, no correctness impact.

Behavior is unchanged: bit-exact sweep (17 dtype/edge cases incl. NaN
sign/payload) + behavioral suite (30 cases: layouts, dtype=, out=, where=,
the 3 complex error paths) both reproduce identically; 21 fabs unit tests +
61 abs/ufunc-overload sibling tests green (Abs unaffected by the shared-site
aliases); oracle Unary + ErrorsFull tiers green. Corpus and OpRegistry are
untouched (fabs routes through UnaryOp.Fabs -> same output bytes).
…ibling-owned)

np.array_equiv (added in 68081a7) is a new public np.* surface that
OracleSurfaceCoverageTests flagged as "unclassified public surface", redding the
FuzzMatrix gate. Classify it as SiblingOwned: its VALUE path is identical to
array_equal's — both reduce all(a == b) — and array_equal is already a direct corpus
op (logic.jsonl via ALLCLOSE_OPS), so the elementwise comparison + all-reduction is
fuzzed across every dtype/layout. array_equiv's only distinguishing behavior is the
broadcast-shape gate (shape-consistent vs exact), which the single-operand-per-slot
differential corpus cannot express (it reconstructs each operand independently and
cannot stage a genuine (M,)-vs-(N,M) broadcast pair) — the same structural reason the
broadcast predicates and set routines are sibling-owned. Gated instead by the dedicated
np.array_equiv.Test.cs suite, verified against NumPy 2.4.2.

Surface gate now green (net10.0).
…d — multi-output, in-tree reductions, scans, gather, vectorizable Call, fixed-cost dispatch, cache management, AOT, tree API

docs/plans/ndexpr-capabilities.md is the second plan for the fused-expression engine: everything
the current plan (P2/P4/P5/P6, unchanged) does not cover, written right after the structural
program cache + hoisted parameters landed (45c6840, 670afdd) and grounded in that substrate
(the one-output shell contract, the 16-byte aux block, NDExprProgram / NDExprProgramCache, the
phase-free host). Fourteen items, each with motivation (measured where it exists), API + examples,
the NumPy-chain or scalar-loop contract, the design in terms of the existing files, edge cases,
gates and acceptance numbers, effort and dependencies; an ordering table (C9 → C10 → C1 → C2 → C3
→ C4 → C5 → C8 → C6/C7 → C11 → C13 → C12, C14 folded into each):

  C1  multi-output evaluation (K roots, one pass, reference-CSE; reduce forests) — AdamW 3 passes → 1
  C2  reductions inside trees as phases (flat result = hoisted parameter, axis result = operand) — softmax/layernorm/standardize in one call
  C3  scan / recurrence (`Scan(body, init)` + `Prev`) / `Shift` + `Diff` as a bind-time view rewrite
  C4  gather `Take(table, idx, mode)` — table operand in aux, AVX2 gather lanes, NumPy take modes
  C5  vectorizable `Call` twins + partial vectorization with an honest cost model (kaiser gains little; erf/mod trees 1.5–2×)
  C6  macro nodes (sigmoid/softplus/relu/gelu/…, parity by composition) + hooks for the Cephes ufuncs of docs/plans/scipy.md
  C7  `Complex(re, im)` (documented inf/NaN difference from `re + 1j*im`)
  C8  np.* overloads over NDExpr, `Eval()`; and WHY there must be no implicit NDExpr→NDArray (`expr == null` would evaluate)
  C9  direct kernel dispatch for identical contiguous operands (skip NDIter, ~110 ns of the 410) + all-0-d trees
  C10 bounded kernel cache + bit-exact literal keys (NaN payloads currently collapse under "R") + analyzer NDW019 (literal from a variable → NDArray.Scalar)
  C11 per-broadcast-pattern SIMD loop copies + running pointers in the strided scalar fallback (the measured 8–10 % branch/imul tax)
  C12 AOT: unfused-chain fallback (bit-identical by definition, also a no-Python differential gate) now; source-generated kernels optional
  C13 public tree API (ToString / visitor / structural equality) + ONNX export through NumSharp.Interop.OnnxRuntime
  C14 the gates: oracle grammar tokens per node, fusion benchmark rows for fixed cost / parameter form / 1K, docs ledger

Non-goals recorded (RNG in kernels, threading beyond P6.4, a GPU engine, full numexpr syntax) and the
three risks (semantic drift without a stated chain, kernel-count growth, the shell's one-output
assumption). docs/plans/ndexpr-evaluate.md gets a pointer to the new plan above its status table.
Implements NumPy 2.4.2's np.put_along_axis(arr, indices, values, axis) as the
exact scatter mirror of take_along_axis: for every position in the (broadcast)
iteration space it reads one index and writes one value into `arr` along `axis`.
In-place (mutates arr, returns void); `axis` is required (no default), matching
NumPy's signature.

Two IL kernels (Backends/Kernels/Direct/DirectILKernelGenerator.PutAlongAxis.cs):
  * PutAlongAxisValidate — a dtype-agnostic pass over the index odometer that
    resolves the advanced-index negative wrap and bounds-checks EVERY index
    BEFORE any store, reproducing NumPy's PyArray_MapIterCheckIndices atomicity:
    an out-of-bounds index leaves arr completely untouched (probed 2.4.2 — a valid
    write preceding the bad index in C-order does NOT land).
  * PutAlongAxisScatter(elemBytes) — the bounds-check-free scatter (indices now
    known in range), a whole-array strided odometer reusing take's exact
    iteration-shape / arrStrides / non-axis-broadcast machinery, so J != M,
    negative wrap, index broadcast and the identical fancy-index IndexError all
    fall out. Byte-width-keyed element copy covers all 15 dtypes. Splitting
    validation into its own pass both guarantees the atomicity contract and keeps
    the scatter hot loop branch-free on the bound.

Setter-only behaviours, all probed against NumPy 2.4.2:
  * values is BROADCAST — not cycled — to the indexing result shape (right-aligned,
    extra leading size-1 dims stripped; the assignment broadcast is more lenient
    than broadcast_to), cast to arr's dtype (assignment cast: floats truncate
    toward zero, over/underflow wraps — the put/place/putmask convention). Where
    several positions collapse onto one element (a size-1 arr dim, or duplicate
    indices), the LAST write in C-order wins. A mismatch raises NumPy's verbatim
    "shape mismatch: value array of shape ... could not be broadcast to indexing
    result of shape ...".
  * axis=None reproduces np.array(arr.flat)'s view-vs-copy split, which is
    load-bearing: a C-contiguous arr is written back through the flat view (aliased
    storage), while a non-contiguous arr raises read-only (NumPy's flat copy is
    read-only there — the dtype check still fires first).
  * COPY_IF_OVERLAP snapshots values when it may alias arr — put_along_axis(a,
    reversing_idx, a) reverses a slice via a copy, not through overwritten reads.
  * Validation ORDER mirrors NumPy exactly: axis -> dtype -> ndim -> writeable ->
    non-axis broadcast -> value broadcast -> per-index bounds.

Perf (NPY/NS, Release, best-of-15): the argsort/argmax-along-axis reconstruction
(put_along_axis's purpose) is 1.68-2.18x — NumPy builds _make_along_axis_idx's
arange grids + a MapIter, NumSharp is a direct odometer; the degenerate axis=None
flat 1-D case is ~1.03-1.11x, the memory-bandwidth ceiling the whole scatter
family (put/place/putmask) hits.

Gates:
  * Indexing/PutAlongAxisTests.cs (41 unit tests, from probed NumPy 2.4.2 output):
    core semantics, the docstring example, axis=None contig-writeback /
    non-contig read-only / 0-d, value broadcast (full/col/row/scalar/leading-1-
    strip/conflict), J!=M, negative wrap, duplicate/broadcast last-wins, all 15
    dtypes, dtype cast (trunc/NaN/inf/overflow), COPY_IF_OVERLAP, OOB atomicity,
    F/transposed/neg-stride/sliced write-through, and full error parity.
  * 48 put_along_axis cases in the groupa differential-fuzz tier — bit-exact vs
    NumPy 2.4.2 across int32/float64/uint8/complex128 (OpRegistry + gen_oracle
    wired; the oracle-surface coverage gate auto-classifies put_along_axis via the
    corpus).

See Indexing/np.put_along_axis.cs.
…/_type_check_impl.py)

Ports NumPy 2.4.2's `np.typename(char)`, which is literally `_namefromtype[char]`:
a pure array-protocol type-CODE → human-description dict lookup ('i' → "integer",
'D' → "complex double precision", 'S1' → "character"). There is NO NDArray operand,
no dtype loop and no kernel — it is the type-check sibling of the scalar→string
formatters binary_repr/base_repr, so it lives in its own file and is classified
sibling-owned in the oracle surface guard, gated by a dedicated unit suite.

Implementation (src/NumSharp.Core/Creation/np.typename.cs):
- `public static string typename(string @char)` — `@char` mirrors NumPy's `char`
  parameter name (positional call `np.typename("i")` ports verbatim).
- Backed by a private `Dictionary<string,string> _namefromtype` reproducing NumPy's
  22 entries VERBATIM, with StringComparer.Ordinal so the lookup is CASE-SENSITIVE
  ('s' misses, 'S' hits) — a load-bearing parity property.
- A miss (unknown/empty/multi-char/differently-cased code, or a null argument ≙
  Python None) raises the existing `KeyError` whose message is the Python repr of the
  argument ('x', '', None), reproduced via PyLiteral.Repr — the same NumPy-parity repr
  helper used elsewhere. null is guarded EXPLICITLY: Dictionary.TryGetValue throws on a
  null key rather than reporting a miss, and np.typename(None) must surface as
  KeyError("None"), not an NRE/ArgumentNullException.

Fidelity notes (every point probed against NumPy 2.4.2):
- 'S1' → "character" and a bare 'S' → "string" are DISTINCT entries (footgun preserved).
- 'g'/'G' read "long precision" / "complex long double precision" — NumPy 2.4.2's
  ACTUAL wording, NOT the older "long double precision" the stale docstring still prints.
- There is deliberately NO 'e' (half) key — NumPy omits it, so typename("e") raises
  exactly as NumSharp does. Parity, not a gap.

Verification:
- Differential probe: NumSharp bit-identical to NumPy 2.4.2 across all 22 codes + 12
  error/edge cases (unknown code, empty, multi-char, dtype names, case, null) INCLUDING
  every KeyError message.
- Perf: 2.9 ns/call vs NumPy 76.0 ns/call → 26× (NPY/NS), an O(1) C# dict lookup vs a
  Python function call; no array/dtype/size axis to sweep.
- Gates: np.typename.Test.cs (8 tests) green on net8.0 + net10.0; OracleSurfaceCoverage
  gate green (typename classified sibling-owned, self-retire check passes) on both TFMs.

Scope: isolated to a new impl file, a new test file, and one SiblingOwned entry — the
unrelated in-progress DType work already dirty in the tree is left untouched.
…(non-contiguous multi-OOB indices)

Validation of the take_along_axis <-> put_along_axis family (10,200 randomized
differential cases vs NumPy 2.4.2 + 142 metamorphic layout invariants, 0
unexplained) confirmed put_along_axis shares take_along_axis's documented benign
[Misaligned] class: for a NON-contiguous indices array with MULTIPLE out-of-bounds
values, the reported offending index VALUE follows the validation odometer's
C-order while NumPy's follows its MapIter order. Error type/axis/size always match;
argsort/argmax output (contiguous) is exact, so it never arises in practice.
Documents the edge to match take_along_axis's coverage.
…urally blind to

The API-coverage generator compared only FIVE NumPy surfaces — the top-level
`numpy` (`np.__all__`), `numpy.ndarray` (`dir()` minus `_`-prefixed), and the
`__all__` of `numpy.random`/`numpy.linalg`/`numpy.fft`. Everything else NumPy
ships publicly was invisible to the artifact: it could not appear even as a gap.
That is exactly how whole families stayed off the radar rather than reading as
"missing" — repeated ad-hoc probes kept rediscovering ~570 public callables the
tool never considered.

Fix (Python side only; the C# reflector is unchanged):

- New out-of-headline "extended surfaces" pass (`extended_surface_rows`) that
  sweeps the public submodules the headline scope excludes and emits them as
  `in_default_scope=false` rows, so a scan SEES the family without moving the
  headline percentage (which would be dishonest for subsystems NumSharp
  intentionally lacks). Covered: `numpy.emath` (= lib.scimath),
  `numpy.polynomial.{polynomial,chebyshev,legendre,hermite,hermite_e,laguerre}`,
  `numpy.lib.{stride_tricks,array_utils,recfunctions,format}`, `numpy.ma`,
  `numpy.char`, `numpy.strings`, `numpy.rec`, `numpy.testing`, `numpy.ctypeslib`,
  plus the ndarray interop dunders (`__array__`/`__array_interface__`/`__dlpack__`
  /…) that the `_`-prefix filter otherwise hides.

- Each extended row carries a DISPOSITION so the summary ranks real opportunities
  above non-goals: `candidate` (implementable, unimplemented — emath,
  polynomial.*, stride_tricks, array_utils), `subsystem` (needs a NumSharp
  subsystem that does not exist — ma, char/strings, rec, recfunctions), `tooling`
  (Python-runtime tooling with no analog — testing, ctypeslib, lib.format),
  `interop` (ndarray array-protocol hooks).

- Honest crediting: an extended member is "available" ONLY when NumSharp exposes
  a same-name member on a [ModuleName] facade for that EXACT submodule (e.g. an
  eventual np.emath); it auto-credits the day such a facade lands. A top-level np
  namesake with different semantics (`np.sqrt` vs `emath.sqrt(-1)==1j`) is noted
  in prose, never counted — closing the emath/char false-positive class.

- New `summary.extendedSurfaces` block + a "## Extended NumPy submodules" table in
  summary.md (ranked candidate->subsystem->tooling->interop, with notable-missing
  samples and a note on the absent Polynomial/Chebyshev/…/MaskedArray/chararray/
  recarray classes). Headline scope, `by_surface`/`by_category`, and every
  existing row remain byte-identical — extended rows only append.

- manifest gains `extended_surfaces`; GENERATOR_VERSION 1.6.0 -> 1.7.0; README
  documents the extended scope and the remaining object-method boundary.

Also removed the now-stale `numpy.remainder` alias from overrides.json (np.remainder
matches directly since the divmod/fmod/remainder work; the generator was emitting a
"delete this alias" warning every run).

Verified: generator runs clean; denominator intact (total=560, by_surface sums to
560); 0 extended rows leak into default scope; 570 extended rows catalogued across
18 submodule surfaces. Headline moved 530->533 only from concurrent tree work
(i0/real_if_close/typename landed), not from this change.

Note: object-method surfaces remain out of scope even after this fix (they are
methods on objects, not module-level names) — a separate deeper probe found the
ufunc-method protocol (`np.add.reduce`/`.reduceat`/`.at`/`.outer`/`.accumulate`
over 106 ufuncs; NumSharp has no ufunc object model), modern RNG class methods
(Generator.spawn/multivariate_hypergeometric, SeedSequence.spawn, BitGenerator/
PCG64/MT19937 random_raw/jumped/spawn, RandomState.tomaxint), and unmodelled
exception/warning types (ComplexWarning, RankWarning, ModuleDeprecationWarning,
VisibleDeprecationWarning). These are documented as the enumeration boundary.
…its real lane

Port of NumPy 2.4.2 numpy/lib/_type_check_impl.py::real_if_close, the last missing
member of the real/imag complex-inspection family (real, imag, angle, conjugate/conj,
iscomplex, isreal, iscomplexobj, isrealobj already shipped; its See-Also — real, imag,
angle — is complete).

Behaviour (all probed against NumPy 2.4.2):
- non-complex input -> returned UNCHANGED (the SAME instance), as NumPy returns `a`.
- complex input: if EVERY imaginary part is strictly within `tol` of zero, return the
  float64 real lane (a write-through VIEW == NumPy's a.real); otherwise return the
  complex array unchanged (also the same instance).
- tol > 1 -> interpreted in machine-epsilon multiples (float64 eps 2.220446049250313e-16,
  the sole complex dtype's eps, so the default tol=100 resolves to ~2.22e-14); tol <= 1 ->
  absolute tolerance; tol <= 0 -> NOTHING collapses (|imag| >= 0 is never strictly < 0),
  even for an exactly-zero imaginary part.
- STRICT `<`: an imaginary part exactly == tol does not collapse.
- NaN / +/-inf imaginary parts prevent the collapse (their comparison with tol is false),
  matching np.all over absolute(a.imag) < tol.
- an EMPTY complex array COLLAPSES (np.all([]) is vacuously true) -> empty float64.
- the band test is on the imaginary part ONLY: a NaN REAL part collapses fine and the
  result carries that NaN in the real lane.

Implementation — ImagCloseScan (Backends/Kernels/ImagCloseScan.cs), modelled on FiniteScan.
NumPy computes np.all(np.absolute(a.imag) < tol) as a strided read + an `absolute` temp +
a boolean-`<` temp + a reduce; we FUSE all of it into ONE early-exit streaming pass over
the imaginary lane with no intermediate allocation. The imaginary lane is ALWAYS float64
(NumSharp has one complex dtype), so the kernel is float64-only — no per-dtype switch. A
System.Numerics.Complex is two contiguous doubles, so a C-/F-contiguous complex array
deinterleaves the imaginary lanes with one vshufpd (control 0b1111 surfaces all four
imaginary parts per shuffle; order is irrelevant for a band reduction) then Vector256.Abs +
Vector256.LessThanAll (a lane that is NaN, +/-inf or >= tol exits immediately). Strided /
transposed / negative-stride / broadcast views take an incremental-offset odometer whose
innermost run reuses the dense scan (unit / reversed complex stride) or AVX2-gathers the
strided imaginary doubles, falling to a scalar walk for pathological strides that overflow
the int32 gather index. On collapse, np.real(a) yields the write-through float64 real-lane
view (matching NumPy's a.real, including read-only-ness for a broadcast source).

Perf (NPY/NS, best-of-25, Release, warm) — FASTER than NumPy on every case x size:
- collapse (full scan)        1K 6.2x   100K 14.9x  10M 4.7x
- no-collapse, big imag @EnD   1K 30x    100K 15.5x  10M 4.7x
- no-collapse, big imag @start effectively free (NumSharp early-exits; NumPy's 3-pass
  composition has no early exit)
- strided ::2 collapse         1K 6.4x   100K 4.6x   10M 2.9x
Min 2.9x, comfortably above the 1.5x bar. The win is structural: one fused early-exit pass
vs NumPy's absolute + `<` + all three passes.

Gates:
- test/NumSharp.Tests/Math/np.real_if_close.Test.cs (20 tests): collapse/no-collapse,
  all tol modes, NaN/inf imag, strict-`<` boundary, -0.0 imag, empty vacuous-all, 0-d,
  2-D, transposed / negative-stride / strided layouts, write-through view semantics, and
  the >= 8-element SIMD deinterleave body.
- Differential-fuzz tier real_if_close.jsonl (360 cases, RunCorpus — host-INDEPENDENT and
  byte-exact, since the result is pure copies of stored bits): both outcomes across 8 tol
  values x 6 imaginary patterns (tiny / one-big / nan / inf / boundary / nan-real) x
  dense / F-contiguous / negative-stride / strided-inner-gather / broadcast / 0-d / empty
  layouts, plus non-complex passthrough. 360/360 bit-exact vs NumPy 2.4.2 on net8.0/net10.0.
  Wired through gen_oracle.py (gen_real_if_close + a `real_if_close` mode), OpRegistry.cs and
  FuzzCorpusTests.cs; as a DIRECT corpus op it satisfies the surface-coverage gate.
…s (np.typename KeyError parity)

Validation of np.typename (bdd4df0) via a 221-input, encoding-safe differential vs NumPy 2.4.2
surfaced exactly ONE divergence class — 11 cases, all the KeyError MESSAGE for NON-PRINTABLE
characters. NumPy's KeyError text is repr(char), and CPython's str repr escapes control chars and
non-printable format/separator code points as \xNN / \uNNNN / \UNNNNNNNN; PyLiteral.Repr emitted
them raw. Every other case (65 valid values + 145 KeyError messages incl. quote selection
"it's"->double-quoted, the \n \t \r \\ short escapes, and printable unicode like u-umlaut/euro/CJK/
emoji rendered verbatim) was already bit-identical.

Root fix in the SHARED helper PyLiteral.ReprString, whose own doc claims it renders "the way
Python's repr() would" and was simply incomplete:
- Iterate CODE POINTS (Rune) instead of UTF-16 units, so a non-printable supplementary character
  escapes as ONE \UNNNNNNNN (its scalar value), not two lone-surrogate \uNNNN.
- New IsPythonPrintable(codePoint) = CPython's str.isprintable: non-printable iff the Unicode
  category is Other (Cc/Cf/Cs/Co/Cn) or Separator (Zs/Zl/Zp), with U+0020 SPACE the sole
  printable exception.
- Non-printables emit \xNN (<0x100) / \uNNNN (<0x10000) / \UNNNNNNNN (>=0x10000), lowercase hex,
  matching CPython. Printable characters (incl. non-ASCII) and the existing \\ \n \r \t + quote
  escaping are unchanged.

Blast radius: PyLiteral.Repr is used ONLY in error messages (np.typename's KeyError and NpyFormat's
FormatException texts), never in a valid .npy header write path, and its output is byte-identical
for every printable input (dtype descriptors, header dicts, shape lists) — so this is strictly more
faithful with zero behavior change on realistic input. np.typename itself is UNCHANGED; it already
routed its KeyError message through PyLiteral.Repr.

Verification:
- The full 221-input differential now reports 0 mismatches through the REAL np.typename (via the
  real PyLiteral), including every control char, NBSP (U+00A0), soft hyphen (U+00AD), zero-width
  space (U+200B) and a non-printable supplementary code point (U+E0001) — .NET's UnicodeCategory
  agrees with CPython on every tested cell.
- np.typename.Test.cs gains 2 methods (QuotesAndShortEscapes, EscapesNonPrintables_LikePythonRepr),
  10 tests total, green on net8.0 + net10.0. Inputs use \u/\U escapes (never raw control bytes) so
  the source stays legible and unambiguous.
- NpyOracle (the other PyLiteral.Repr consumer) is 13/13 green on both TFMs — no error-message
  regression.

Scope: only PyLiteral.cs + np.typename.Test.cs, committed with --only so the shared working tree's
concurrent DType and np.i0/np.real_if_close work is left untouched.
…y float precision

Public wrapper for the modified Bessel function of the first kind order 0. The
scalar float64 helper (BesselI0 / _i0A / _i0B / Chbevl) already existed in
np.windows.cs as np.kaiser's private taper; this exposes NumPy's np.i0 over it
and adds the float32 / float16 / decimal precision variants required for parity.

Port of NumPy 2.4.2 numpy/lib/_function_base_impl.py::i0:
    x = asanyarray(x)
    if x.dtype.kind == 'c': raise TypeError("i0 not supported for complex values")
    if x.dtype.kind != 'f': x = x.astype(float)      # bool/int/char → float64
    x = abs(x); return piecewise(x, [x <= 8.0], [_i0_1, _i0_2])

TWO load-bearing behaviours, both probed and easy to get wrong:

 (1) DTYPE FOLLOWS THE INPUT FLOAT PRECISION, not a blanket float64. NumPy's
     piecewise/_chbevl chain runs in the input dtype under NEP50 weak-scalar
     promotion (the float64 coefficients adopt the array's dtype), so i0(float32)
     is computed IN float32 — i0(0f)=0.99999994, NOT 1.0 — and i0(float16) IN
     float16. A float64-computed-then-cast result diverges from NumPy by up to
     3–5 ULP on ~50% of float16/float32 inputs, so each float precision needs its
     OWN native path. bool/int/Char → float64; Decimal preserved via the double
     bridge (NumSharp extension); Complex refused with NumPy's exact TypeError.

 (2) The float32 exp must be NumPy's OWN kernel, not MathF.Exp. NumPy's float32
     i0 calls the float32 exp loop (simd_exp_FLOAT), which differs from the
     ~correctly-rounded MathF.Exp on ~39% of inputs; the float32 path uses
     NDFloatMath.Exp (the ported kernel). float16 exp/sqrt are BCL Half.Exp/
     Half.Sqrt (byte-identical to NumPy's half loop); float64 uses Math.Exp/Sqrt
     (== win-amd64 ucrtbase, already bit-exact at f8). Every chbevl step is pure
     add/sub/mul, IEEE-exact per op at each width.

VERIFIED 0 bit-diffs vs NumPy 2.4.2 across 10,518 adversarial float64 AND float32
inputs (specials / ±inf / NaN / subnormals / the x≤8 branch split / overflow to
+inf at |x|≳710) and ALL 65,536 float16 bit patterns.

IMPLEMENTATION — one fused np.evaluate pass per precision (the np.kaiser pattern):
the cephes routine is a data-dependent Clenshaw recurrence over 30/25 coefficients
with a per-element |x|≤8 branch. It CANNOT be an unrolled NDExpr tree — Clenshaw
reuses b0 as both b1 and b2, so a no-CSE tree explodes Fibonacci-like and
stack-overflows the emitter. Instead the whole scalar routine rides ONE
NDExpr.Call node, so np.evaluate drives it through NDIter: one pass, every memory
layout (C/F/strided/reversed/broadcast/sliced), no intermediate arrays (NumPy
materializes ~30). The per-precision delegate is a static readonly field so the
compiled kernel is cached once (a method group would JIT a fresh kernel per call).
A 4-arm precision-domain switch selects the bit-exact native helper — mirroring
the kernel's own EmitUnaryHalfOperation/decimal/complex split, not a mechanical
per-NPTypeCode copy; bool/int/Char fall to the double arm where NDExpr.Call
converts the operand to double at the call edge (NumPy's astype(float)).

PERF (NPY/NS, Release, best-of-30 warm): float64 — the dtype real Bessel work
uses — WINS at every size (3.8× at 1K, 1.9× at 100K, 9.4× at 10M) because the
single fused compute-bound pass beats NumPy's ~35 memory-bound array passes.
float32 wins at 1K (~1.0×) and 10M (4.4×) but is ~0.84× at 100K, and float16 is
~0.4–0.5×, because at cache-resident sizes NumPy's SIMD passes stay in L2/L3 and
beat a per-element scalar Clenshaw — and float16 additionally has NO BCL vector
arithmetic (the ceiling Half hits library-wide). Beating those two cells would
need i0 as a first-class SIMD ILKernelGenerator op (a 55-coefficient vector
Clenshaw + width-specific NumPy exp), disproportionate for a function NumPy itself
documents as "not a proper ufunc — use scipy". Correctness is the gate and is met.

Gates:
- Math/np.i0.Test.cs (15) — docstring values, the x≤8 branch split + overflow,
  even-symmetry, ±inf/NaN→NaN, native-precision f32 (bit pin 0x3F7FFFFF for
  i0(0f)) / f16, int/bool→f64, decimal preserve + overflow-throws, complex→
  IncorrectTypeException (verbatim message), null, 0-d/empty/2-D shape, and
  non-contiguous layout self-consistency.
- Oracle "i0" tier (338 cases) — gen_oracle.py i0 mode + OpRegistry case +
  FuzzCorpusTests.I0(); HOST-PINNED (RunHostLibmCorpus, like Unary/Sinc — f64/f16
  ride host libm), Inconclusive off-Windows. 338/338 bit-exact on win-amd64.
  OracleSurfaceCoverageTests now classifies np.i0; zero-leak gate confirms the
  fused pass strands no pooled intermediate.
A full adversarial differential of np.real_if_close against live NumPy 2.4.2
(74 crafted cases: tol=NaN/+-inf/subnormal/huge, special reals/imags, ranks
3D/4D, genuine-F/transpose/negstride/strided/offset/broadcast/newaxis layouts,
all 12 non-complex dtypes, every empty shape, 0-d) came back 74/74 BYTE-IDENTICAL,
and the view-semantics probe confirmed writeable write-through on a normal
collapse, a read-only view on a broadcast collapse, no input mutation, and
0-d-from-indexing handling — all matching NumPy. No implementation bug found.

This encodes the edges the 360-case oracle tier does not reach as permanent
regression guards (20 -> 28 tests):
- TolPositiveInfinity_CollapsesEveryFiniteImag: tol=+inf => eps*inf=inf collapses
  every FINITE imag, but an infinite imag is still not < inf (no collapse).
- TolNaN_And_TolSubnormal_DoNotCollapse: NaN>1 is false (stays NaN, nothing < NaN);
  subnormal tol is absolute and nothing is within it.
- ImagBoundary_JustUnderCollapses_JustOverDoesNot: the two doubles adjacent to
  eps*100 split on the strict `<`.
- SpecialRealParts_SurviveCollapse_BitExact: NaN / +-inf / -0.0 (sign bit) reals
  survive into the float64 lane (the band test is imag-only).
- Rank3_And_Rank4_CollapseAndNoCollapse: the strided odometer beyond 2-D.
- FContiguous_And_OffsetSlice_Collapse: genuine asfortranarray dense path + a
  non-zero Shape.offset slice.
- BroadcastReadOnly_Collapse_ReturnsReadOnlyView: matches NumPy's read-only a.real.
- DoesNotMutateInput: real_if_close never writes to the array.

All 28 pass on net8.0 and net10.0; the 360-case FuzzMatrix tier stays green.
…lved C-integer numbers

Stage A follow-up (docs/plans/dtype-system.md §2.2). Adds NumPy's C-API
type-number identity ALONGSIDE NumSharp's storage/kernel discriminator
NPTypeCode — the exact split NumPy draws between its C `type_num` and the
descriptor it keys — and fixes the fixed-width integer numbers to be
byte-identical to the LOCAL NumPy.

NPY_TYPES enum (Creation/np.dtype.cs)
  A byte-identical mirror of NumPy's C `NPY_TYPES` (numpy/_core/include/
  numpy/ndarraytypes.h), sibling to the existing NPY_TYPECHAR: every member
  and value is exactly NumPy's, so (int)NPY_CDOUBLE is 15 on every platform.
  Carries the members NumSharp has no storage for (NPY_CFLOAT/OBJECT/VOID/
  LONGDOUBLE/CLONGDOUBLE, the sentinels) so type_num is a whole-enum mirror
  and From() can reject them precisely, plus the user-range extras
  NUMSHARP_DECIMAL (256 = NPY_USERDEF) and NUMSHARP_CHAR (257) — exactly as
  NumPy numbers a user dtype.

DType.type_num + From(NPY_TYPES) + implicit cast (DTypes/DType.cs)
  - type_num => (NPY_TYPES)Meta.TypeNum: the typed view riding alongside
    typecode (NPTypeCode). DType.num stays `int` (NumPy's .num is an int).
  - From(NPY_TYPES) is the PyArray_DescrFromType reverse, deliberately
    many-to-one on the C integers: NumSharp has ONE type per width where
    NumPy keeps intc/long/longlong, so every C-integer number maps to the
    NumSharp type of that width+signedness — NPY_INT and NPY_LONGLONG resolve
    on ALL platforms, NPY_LONG/ULONG follow C `long` via CLongIs32Bit.
    Numbers with no NumSharp class (complex64, object, void, extended
    precision, sentinels) throw NotSupportedException, matching how the
    string parser refuses a dtype NumSharp cannot represent.
  - implicit operator DType(NPY_TYPES): the fifth dtype spelling, so
    np.zeros(3, NPY_TYPES.NPY_DOUBLE) binds the one DType overload like
    Type/NPTypeCode/string. A value type, so no null-literal binding hazard
    (unlike the barred ==(DType,string)); the reverse is the type_num
    property, not an operator, to keep DType's back-conversions unambiguous.

Platform-resolved 32/64-bit integer (num, char) (DTypes/DTypeRegistry.cs)
  np.dtype('int32')/('int64') resolve to whichever C integer is that width
  (verified against numpy 2.4.2): LP64 (Linux/macOS, C long = 64-bit)
  int32 == NPY_INT(5,'i'), int64 == NPY_LONG(7,'l'); LLP64 (Windows / any
  32-bit process, C long = 32-bit) int32 == NPY_LONG(7,'l'), int64 ==
  NPY_LONGLONG(9,'q'). NumSharp's single Int32/UInt32/Int64/UInt64 are now
  registered with the LOCAL platform's (num, char) via the existing
  CLongIs32Bit switch — the same switch the LongDType/LongLongDType name
  aliases already used, so the nums are now CONSISTENT with those aliases
  (they were hard-coded LP64 before, inconsistent on Windows). Everything
  else stays platform-independent.

NPTypeCode is UNTOUCHED — it remains the storage/kernel discriminator every
backend switch and IL generator dispatches on; NPY_TYPES is the NumPy-facing
identity served alongside it into the kernels (Stage C).

Tests (platform-aware)
  - DTypeDescriptorTests.Builtin_Surface remaps the integer DataRows to the
    local platform and now asserts type_num == (NPY_TYPES)num.
  - New NpyTypes_ImplicitCast_And_From_WidthCanonicalized: the implicit cast,
    dtype:/np.zeros binding, From() C-integer width-folding on any platform,
    NPY_HALF/CDOUBLE, type_num round-trip, the user-range extras, and
    rejection of complex64/object/void.
  - NpDtypesModuleTests remaps the integer classes' expected nums and
    retargets Registry_Register_RejectsDuplicates onto NPY_SHORT(3), which is
    registered on every platform (NPY_INT=5 is unused on LLP64/Windows).

Gate: full DType/promotion/parity suite green, dtype_text fuzz tier +
OracleSurfaceCoverage green.
ORT interop: non-tensor OrtValue readers (sequence/map/string) + 3 fixtures, coverage 128->148. Conflicts: took branch versions of the 4 ORT-specific docs; in CLAUDE.md kept master's MLNet row and took the branch's updated OnnxRuntime row.
…d 2-8x faster

Port of numpy/lib/_function_base_impl.py::unwrap (NumPy 2.4.2). Unwraps a
signal by changing deltas larger than max(discont, period/2) to their
period-complement so adjacent differences never exceed period/2 (the default
period 2*pi / discont pi unwraps a radian phase by adding 2*k*pi).

NumPy is a pure composition; NumSharp mirrors it operation-for-operation so
NEP50 promotion, broadcasting and every edge case fall out of the existing
parity-tested building blocks (diff / mod / cumsum / where / abs):

    dd    = diff(p, axis)
    ddmod = mod(dd - interval_low, period) + interval_low
    if boundary_ambiguous: ddmod[(ddmod==interval_low)&(dd>0)] = interval_high
    ph    = ddmod - dd ; ph[|dd| < discont] = 0
    up    = p.astype(result_type(dd, period)) ; up[1:] = p[1:] + cumsum(ph, axis)

API — two overloads mirror NumPy's KEYWORD-ONLY `period` (a Python int and a
Python float are distinct at the call site, and the C# port needs the same
distinction because it decides the output dtype):
  * unwrap(NDArray p, double? discont=null, int axis=-1, double period=2*pi)
      the default / float path (radian phase). Every float use.
  * unwrap(NDArray p, long period, double? discont=null, int axis=-1)
      the integer-preserving path (NumPy's period=<int>): np.unwrap(a, period: 4).
  Resolution: `np.unwrap(a, period: 4)` binds the long overload (int->long beats
  int->double), `period: 4.0` the double one. DOCUMENTED wrinkle: a bare
  positional int second arg `np.unwrap(a, 4)` binds period=4, NOT discont=4
  (period is keyword-only in NumPy, so ported code writes `period:`); set discont
  positionally with a double or by name.

Integer vs float path. NumPy branches on issubdtype(result_type(dd, period),
integer): a FLOAT result (any float input, OR an integer input with a float
period) uses interval_high=period/2 with an always-ambiguous boundary; an
INTEGER result (integer/bool input AND integer period) uses period//2 (floor)
with the boundary ambiguous only for an even period. For an INTEGER input the
two paths compute the SAME values (an integer difference never lands in the
(period//2, period/2] gap) and differ only in OUTPUT DTYPE — integer period
keeps the input's integer dtype (bool -> int64, promoted because a weak int is
higher-kind than bool), float period yields float64.

Dtype/edge parity (all probed against NumPy 2.4.2):
  * float16/float32/float64 preserve width (weak scalars never widen the float),
    Decimal preserved (NumSharp extension), integer/bool + float period -> float64.
  * Unsigned integer (and Char, the uint16 proxy) with an INTEGER period, or a
    signed integer too narrow for the derived interval, reproduce NumPy's
    OverflowError ("Python integer N out of bounds for <dtype>"): interval_low is
    negative and cannot be cast to the dtype. Checked up front via
    NDExprTypeRules.CheckIntLiteralFits (interval_low before period, the order
    NumPy hits them) so it fires for EMPTY inputs too, exactly as NumPy's does.
  * Complex -> TypeError ("ufunc 'remainder' not supported ...") — mod has no
    complex loop, matching NumPy verbatim.
  * 0-D -> ArgumentException ("diff requires input that is at least one
    dimensional"); out-of-range axis -> AxisError reporting the original axis;
    single-element / empty along the axis returns the dtype-cast copy.
  * Every layout (C/F/strided/reversed/transposed/broadcast) read through strides
    by np.diff; result is a fresh, writeable, C-contiguous copy.

Implementation. The whole elementwise chain dd -> ph_correct (subtract, mod,
add, optional boundary Where, subtract, discont Where) is built as ONE NDExpr
and produced in a single fused np.evaluate pass where NumPy allocates ~6
intermediate arrays; only three array passes remain (diff, fused ph_correct,
cumsum). Scalars ride as NEP50-weak NDExpr constants: on the float path a weak
double never widens the operand's float dtype, on the integer path a weak int
keeps the integer dtype and (pre-checked above) triggers the OverflowError.
There is no hand-written per-element loop — every loop lives inside np.diff /
np.evaluate / np.cumsum. [NDScoped] + explicit view-before-owner disposal keeps
it leak-clean (0 escapes in the corpus leak gate).

Perf (NPY/NS, Release, best-of-7, warm; higher = NumSharp faster):
  float64  1K 2.65x  100K 7.86x  1M 6.76x  10M 2.68x
  float32  1K 2.62x  100K 8.38x  1M 7.79x  10M 4.73x
  int64/p4 1K 1.98x  100K 3.10x  1M 2.99x  10M 2.09x
Every measured cell >= 1.5x — the fusion halves NumPy's memory traffic.

Gates:
  * differential-fuzz "unwrap" tier (test/oracle/gen_oracle.py gen_unwrap +
    OpRegistry + FuzzCorpusTests.Unwrap): 702 cases over SCAN layouts x every
    dtype x {default/float/discont/even+odd integer period} x axes, PORTABLE
    (pure arithmetic, no libm), bit-exact vs NumPy 2.4.2 (NaN tokenized). Complex
    (TypeError) and unsigned integer-period (OverflowError) raise in-generator
    and are skipped there, gated by the unit tests instead. Char woven via the
    uint16 char_tier (float-period only). OracleSurfaceCoverageTests auto-classifies
    unwrap as a corpus op (no explicit registration needed).
  * test/NumSharp.Tests/Math/np.unwrap.Test.cs (20): the NumPy doc examples, the
    float/integer path distinction, discont, axis, dtype coverage, edges and the
    full error taxonomy.
…gn/power/arccos/arcsin/arctanh

Port of NumPy 2.4.2 numpy.emath (numpy.lib.scimath) — the branch-cut-aware siblings of the
ordinary ufuncs, reachable as np.emath.* exactly like Python (np.emath.sqrt(-1) == 1j). New
facade EmathModule ([ModuleName("np.emath")]) exposed via the lowercase np.emath property, the
np.fft/np.random house shape.

WHAT IT IS
Where np.sqrt(-1) is nan, np.emath.sqrt(-1) is the complex value; likewise log/log2/log10 for
negative reals and arccos/arcsin/arctanh for |x|>1. The promotion is a WHOLE-ARRAY decision: if
ANY element is out of the real-valued domain the ENTIRE array is promoted to complex128, then the
standard ufunc runs — matching NumPy's _fix_real_lt_zero / _fix_real_abs_gt_1 / _fix_int_lt_zero.

IMPLEMENTATION (pure composition + one fused scan)
Each of the 9 functions is a line-for-line port of numpy/lib/_scimath_impl.py: a domain scan, an
optional cast to complex128 (or, for power's exponent, a weak *1.0 float promotion), then the
existing np.* ufunc. Because the numerics ARE the already-fuzzed np.sqrt/log/arccos/power kernels,
results are bit-identical to NumPy on every real path and every bit-exact complex-unary path.
 - logn(n,x) = log(x)/log(n) after both operands are domain-fixed (a negative base OR value → complex).
 - power(x,p): negative base → complex base; negative exponent → float exponent (so a negative
   integer exponent does not hit np.power's integer-domain error). Weak *1.0 preserves float32 width.
 - A COMPLEX input skips the scan entirely: NumPy's isreal(x)&(x<0) is false for a complex element,
   and _tocomplex on an already-complex128 array is dtype-idempotent, so the complex ufunc on the
   untouched input is the identical (cheaper) result.

FUSED DOMAIN SCAN (specialized fast path, EmathDomainScan.cs)
NumPy computes the trigger as any(x<0) / any(abs(x)>1) — a full bool (and, for abs, a full abs)
temp plus a reduce. This fuses the comparison+reduction into ONE early-exit streaming pass with no
intermediate allocation for the hot CONTIGUOUS float32/float64 case (4x-unrolled, byte-mask OR so
the block test is immune to NaN/±0 float-compare hazards); every other dtype/layout falls back to
the correct-by-construction composition (np.any(x<0) / np.any(np.abs(x)>1)), which reuses the
already-fuzzed abs/comparison kernels and so inherits e.g. NumPy's abs(int.MinValue) overflow.
Measured: the fused scan is ~4 ms for a non-triggering 80 MB float64 pass vs the composition's
~6 ms (x<0) / ~24 ms (abs>1 — an 80 MB abs temp), i.e. up to 6x cheaper.

PARITY (150-case differential vs NumPy 2.4.2, canonicalized to isolate dtype-width)
138/150 bit-exact (values AND dtype). The 12 remaining are pre-documented, unavoidable divergences,
NOT bugs:
 - F1 (10 cases): NumSharp has no complex64, so a triggering float32/float16/small-int input yields
   complex128 (double precision) where NumPy yields complex64 (single precision). The VALUE is
   NumSharp's correct double-precision answer (== NumPy computing the same input in the float64
   domain), just more precise — the library-wide "no complex64" divergence (issue #569).
 - F5 (2 cases): complex power with a NON-integer exponent takes Complex.Pow, allclose to NumPy's
   npy_cpow but not bit-exact (the documented complex-power ULP divergence). Integer exponents are
   bit-exact.
Every real path and every complex128-domain function (sqrt/log/log10/log2/arccos/arcsin/arctanh)
is bit-exact, including edges: -0.0 does NOT trigger (-0.0<0 false); a NaN never triggers; +inf
triggers abs>1; the |x|==1 boundary stays real (arctanh(1)=inf); arccos(1+0j) reproduces the -0.0
imaginary sign bit-for-bit; empty/scalar(0-d)/2D/reversed/transposed/strided all correct.

PERF (NPY/NS, Release, best-of-21 warm)
>=1.5x across sizes/dtypes: 1K 1.5-6.2x, 100K 1.9-4.1x, 1M 2.6-7.8x. At 10M the transcendental
cells sit ~1.4-2.3x (memory/compute-bound on the ufunc itself; the fused scan adds only a ~4 ms
tax and NumSharp's bit-exact bare ufuncs already beat NumPy) — the same physical ceiling documented
for the arcsinh/gcd/float_power family.

GATING
Dedicated suite test/NumSharp.Tests/Math/np.emath.Test.cs (36 tests, all probed against NumPy
2.4.2) — the SiblingOwned pattern (bmat/histogram/array_equiv), since emath is pure composition of
ops already under the FuzzMatrix corpus gate. np.emath is a new module facade the surface guard's
four enumerated types (np/linalg/fft/random) do not cover, and the emath property getter is a
filtered special-name, so OracleSurfaceCoverageTests stays green (verified). A dedicated emath
corpus tier is a natural follow-up once the shared oracle files settle.

Files: src/NumSharp.Core/Math/np.emath.cs, src/NumSharp.Core/Backends/Kernels/EmathDomainScan.cs,
test/NumSharp.Tests/Math/np.emath.Test.cs.
…validation

Validation (897-case layout-aware differential vs NumPy 2.4.2: all 13 NumPy dtypes × the full
layout matrix {c/f/T/reversed/strided/offset/newaxis} × edge pools {NaN/±inf/-0.0/subnormal/
boundary} × logn/power two-operand combos, plus char/decimal self-consistency) confirmed emath
introduces ZERO divergence of its own — 775 bit-exact, and every one of the 122 remaining cases
traces to a documented/pre-existing property of the underlying ufunc:
 - 109 F1 (no complex64 → complex128 double precision; NumSharp's value is the correct
   float64-domain answer),
 - 2 F5 (complex power non-integer exponent, allclose),
 - 4 complex arccos ≤1-ULP — PRE-EXISTING in plain np.arccos (np.arccos(-0.5+0j) real part is
   already 1 ULP from NumPy WITHOUT emath), inherited by emath.arccos via delegation,
 - 7 np.power dtype-width (np.power(int32, int64-scalar)→int32) — the pre-existing C#-int/NEP50
   house convention, present in plain np.power; VALUES correct.

The committed class-doc overclaimed "every complex path is bit-exact except power". Refined to
state precisely: complex sqrt/log/log10/log2 are bit-exact; arccos/arcsin/arctanh inherit the
complex-unary ≤~1-ULP envelope (a property of np.arccos itself, not emath); power non-integer
exponent is allclose (F5). No behavior change — documentation accuracy only.
…pper family

Port of numpy.ma's ufunc layer into NumSharp as np.ma (a lowercase facade
property, the np.fft/np.random house shape), all in ONE file
src/NumSharp.Core/Ma/MaskedArray.cs. The session's explicit trio abs/absolute/add
are the ufunc-wrapper mechanism, so "closely related" = the whole ufunc family +
the MaskedArray substrate they require.

WHAT
- MaskedArray type: an NDArray of DATA paired with an optional boolean NDArray
  MASK (null == NumPy's nomask fast path, keeping an all-valid masked array as
  cheap as the NDArray). data/mask/shape/dtype/ndim/size/filled/ToString; implicit
  NDArray->MaskedArray; MaskedConstant `masked` singleton.
- np.ma substrate: nomask, masked, MaskType, getdata, getmask, getmaskarray,
  is_mask, isMaskedArray, filled, array/masked_array, default_fill_value.
- The three ufunc wrappers, ports of NumPy's _MaskedUnaryOperation /
  _MaskedBinaryOperation / _DomainedBinaryOperation (numpy/ma/core.py):
  * Unary domain-free (mask passes THROUGH): abs, absolute, negative, fabs,
    conjugate, angle, around, floor, ceil, exp, sin, cos, sinh, cosh, tanh,
    arctan, arcsinh, logical_not.
  * Unary domained (mask ORs ~isfinite + invalid-input domain): sqrt(>=0),
    log/log2/log10(>0), tan(poles), arcsin/arccos([-1,1]), arccosh(>=1),
    arctanh((-1,1)).
  * Binary (mask = OR of the two operand masks): add, subtract, multiply,
    arctan2, hypot, equal, not_equal, less, less_equal, greater, greater_equal,
    logical_and/or/xor, bitwise_and/or/xor.
  * Domained binary (mask ORs ~isfinite + safe-divide domain |a|*tiny >= |b|):
    divide, true_divide, floor_divide, remainder, mod, fmod.

HOW — no new kernels
numpy.ma adds no arithmetic: each wrapper calls the existing np.* op on the raw
DATA and only does mask bookkeeping (logical-OR, ~isfinite, domain compares,
copyto fill-back), all IL/NDIter-backed ops already in Core. Every
layout/dtype/promotion/SIMD path of the base op is inherited; no struct kernel is
introduced. Masked positions get the input data restored into .data (NumPy's
copyto(result, d/da, where=m)) — unary fill-back uses same_kind casting (caught on
failure), binary uses unsafe, domained zeroes then re-adds the numerator where it
casts safely — matching numpy/ma/core.py exactly.

TRAPS encoded (all cost a compile error or silent wrong answer)
- NDArray == null / != null do ELEMENTWISE comparison (return NDArray), never a
  reference check — every mask-null test uses `is null` / `is not null`.
- np.abs/np.add/... are overloaded, so a method group won't convert to Func<> —
  every op is wrapped in a lambda (this also resolves the equal/less
  (NDArray,object) vs (object,NDArray) overload ambiguity when both args are
  arrays).
- np.less has a 0-D-scalar-RHS broadcast quirk; the domain predicates are built
  from greater/greater_equal/less_equal/logical_not instead. NaN differences
  between less(x,v) and !(x>=v) are absorbed by the ~isfinite term every domained
  op already ORs in, so the final mask stays bit-identical.
- NDArray.shape is long[] (not int[]).
- No np.any() gate on the array fill-back: copyto with an all-false mask is a
  no-op, and the scan would cost a full extra O(n) pass; NumPy's unary path
  likewise fills unconditionally.

PARITY — probed vs NumPy 2.4.2
21 cases / 83 numeric values BIT-IDENTICAL + all masks identical, across
unary(domain-free+domained), binary, domained-binary, comparisons, scalar/plain
operands, filled, and masked-position fill-back.

PERF (NPY/NS, best-of-21, warm, ~1/7 masked; >1 = NumSharp faster)
add 100K/1M/10M = 3.8/2.5/1.2x, abs = 1.1/2.4/1.0x, divide = 1.65/1.67/1.35x,
sqrt = 1.5/1.4/1.1x. Beats NumPy at 100K/1M everywhere; 10M is DRAM-bandwidth
bound — the mask fill-back pass is inherent to masked semantics and NumPy pays it
too (same documented ceiling as putmask/place/signbit f64). Future lever: fuse
base-op + fill-back into one np.evaluate where(mask,d,op(d)) pass to cut ~40%
traffic and clear 1.5x at 10M.

SCOPE
The ufunc-family slice only. numpy.ma's reductions (_frommethod: sum/mean/...),
creation (_convert2ma: zeros/arange/...) and extras.py are separate mechanisms,
deliberately out of scope for this session.

Gate: test/NumSharp.Tests/Ma/MaskedArrayTests.cs (10 tests, probed vs NumPy
2.4.2). OracleSurfaceCoverageTests green (np.ma is a facade property like
np.fft/np.emath, not enumerated). Not oracle-corpus-testable — MaskedArray is a
new type absent from gen_oracle.
…ma, manip

Extends the np.ma module (same file) well beyond the ufunc family toward the full
numpy.ma surface. All values/masks probed bit-exact vs NumPy 2.4.2 (reduction
differential 37 cases, wave-2b 18, wave-2c 20).

REDUCTIONS (filled(identity).op(axis) + result-cell masked iff its whole slice was
masked, per NumPy _check_mask_axis=mask.all(axis); scans keep the per-position mask):
sum, prod/product, mean, min, max, ptp, var, std, count, cumsum, cumprod, argmin,
argmax, all, any, anom/anomalies + minimum_fill_value/maximum_fill_value. Fully-masked
scalar reductions return the  singleton. mean/var restore 0 (not 0/0=NaN) at
fully-masked slices, matching NumPy's masked-arithmetic .data.

OPERATORS on MaskedArray: + - * / unary-, and ordering < > <= >= (compose the ufuncs;
mask=OR). ==/!= deliberately NOT overloaded (would shadow reference equality and the
 elementwise trap) — use np.ma.equal/not_equal. Plus instance reduction methods
(a.sum()/a.mean()/…).

CREATION (fresh, unmasked): zeros, ones, empty, zeros_like, ones_like, empty_like,
arange, identity, indices.

masked_* CONSTRUCTORS: masked_where (+ the family masked_equal/not_equal/greater/
greater_equal/less/less_equal/inside/outside/invalid/values/object). masked_values fills
with value FIRST then isclose-masks (NumPy semantics). masked_less routes a<v as v>a to
dodge the np.less 0-D-scalar-RHS bug (NaN-exact).

EXTREMA/POWER/WHERE/ROUND: maximum/minimum = ma.where(compare(a,b),a,b) (masked where +
masked compare — masked slot yields the other operand, NaN falls out); power masks
non-finite results + sets their data to fill_value; where = filled(cond,False) then
np.where on data and masks; round/around preserve the mask.

SHAPE/MANIP (same transform on data AND mask): ravel, flatten, reshape, transpose,
swapaxes, moveaxis, squeeze, expand_dims, repeat, take, diag, diagflat, atleast_1d/2d/3d,
concatenate, stack, hstack, vstack, dstack, column_stack, compressed.

Gate: MaskedArrayTests 18 (10 ufunc + 8 new).
…statistics

Extends np.ma (same file) with the extras.py surface and sorting, all probed
bit-exact vs NumPy 2.4.2 (wave-2d 24 cases, wave-2e sort/argsort/unique).

STATISTICS / PRODUCTS: count_masked, masked_all, masked_all_like, average
(uniform + weighted; sum(a*w)/sum(w) over unmasked), median (flat + unmasked-axis),
ediff1d, allequal, allclose, dot (masked->0, result masked iff no valid pair),
inner, outer, vander (masked rows zeroed, returns plain ndarray).

SET / MEMBERSHIP: isin, in1d (masked element -> not-a-member, unmasked result).

CLUMPS / EDGES: clump_masked, clump_unmasked (contiguous mask-run slices),
flatnotmasked_edges (first/last unmasked index), flatnotmasked_contiguous.

2-D ROW/COL: compress_rows, compress_cols (drop rows/cols with any masked -> plain
array), mask_rows, mask_cols (mask entire rows/cols with any masked).

SORT: sort (reorders BOTH data and mask by argsort(filled-with-largest) so masked
entries land at the END keeping their original data, exactly as NumPy — not the
fill value a plain sort-of-filled would leave), argsort (masked treated as largest),
unique (sorted unique unmasked values + one trailing masked entry). Mask hardness
harden_mask/soften_mask/shrink_mask accepted for parity (NumSharp masks have no
hard/soft state).

TRAPS: np.isclose returns NDArray<bool> — reassignment needs an NDArray-typed local;
maximum/minimum = ma.where(compare(a,b),a,b) so a masked slot yields the other operand
and NaN falls out; masked_values fills with value FIRST then isclose-masks; ma.vander
returns a PLAIN array with zeroed masked rows.

DEFERRED (documented; need machinery NumSharp lacks or record dtypes): median with an
explicit axis on a masked array (masked sort core), the set operations
intersect1d/union1d/setxor1d/setdiff1d and full unique(return_index/inverse), cov/
corrcoef (pairwise-complete masked covariance), polyfit (needs LAPACK lstsq —
backend-only), apply_along_axis/apply_over_axes, notmasked_edges/notmasked_contiguous
(nested per-axis), correlate/convolve (propagate_mask), mvoid/MaskedIterator/mr_/
mrecords (record dtypes / iterator objects). The trailing masked entry's DATA in
ma.unique is an arbitrary hidden value and may differ from NumPy (position is masked).

Gate: MaskedArrayTests 22 (probed vs NumPy 2.4.2).
…float16, complex preserved)

Found by a comprehensive dtype x layout differential vs NumPy 2.4.2 (253 cases across
all 13 NumPy-representable dtypes x 6 layouts x edge cases): np.ma.mean force-cast every
result to float64, but NumPy's masked mean is `dsum * 1. / cnt`, whose `* 1.` (a Python
float64) promotes float32->float64 while float16 is computed in float32 then cast BACK to
float16, and complex128 is preserved. So mean of a complex masked array must stay
complex128 (was silently float64) and mean of float16 must return float16.

Fix: mean resolves an explicit compute dtype (complex128 for complex, float32 for
float16, else float64) and an output dtype (float16 for float16 input, complex128 for
complex, else float64) instead of a blanket .astype(float64).

VALIDATION RESULT (this fix in): all 253 result dtype-signatures, masks, and
ERR/masked-scalar error-parity match NumPy; 252/253 values bit-exact. The one value
difference is std@float16 (NumSharp 1.247219 vs NumPy 1.247393) where NumSharp is MORE
accurate — NumPy centers with an f16-rounded mean, NumSharp centers in f8 — an excused
"prefer-precise" divergence per the codebase precision policy, not a bug.

Gate: MaskedArrayTests 23 (added Mean_ResultDtype_MatchesNumPy). No regression:
Statistics 307/307, surface guard green, both TFMs green.
…g-usage parity)

Validation of "is MaskedArray casting/usage the same as NumPy" surfaced a real bug and
two API gaps; all fixed and probed against NumPy 2.4.2.

BUG: `ma + nd` and `nd + ma` (and the other arithmetic/ordering operators mixing a
MaskedArray with a plain NDArray) were CS0034-AMBIGUOUS compile errors — the implicit
NDArray->MaskedArray conversion made both the (…, MaskedArray) and (…, object) operator
overloads applicable with no better candidate. In NumPy both work and yield a mask-aware
MaskedArray. FIX: explicit (MaskedArray, NDArray) and (NDArray, MaskedArray) overloads for
+ - * / and < > <= >= ; an exact NDArray parameter beats the conversions, so they resolve
cleanly and return a mask-aware MaskedArray. Verified values+mask+dtype match NumPy for
ma+nd / nd+ma / ma-nd / ma*nd / ma/nd / ma<nd / ma+scalar.

GAPS CLOSED:
- MaskedArray.astype(dtype, copy=true): casts the DATA, PRESERVES the mask (the mask is
  boolean/dtype-independent) — NumPy's MaskedArray.astype. Was missing.
- np.ma.asarray / np.ma.asanyarray (object, dtype=null): the explicit ndarray->MaskedArray
  entry points — keep an incoming mask, wrap a plain array as unmasked, optional dtype cast.

CASTING MODEL (documented divergence, architectural): NumPy's MaskedArray IS-A ndarray
subclass, so regular np.* accept it and stay mask-aware via __array_wrap__ dispatch.
NumSharp's MaskedArray is a SEPARATE class (no subclass, no dispatch hook), so the
mask-aware ops live on np.ma.* (mirroring numpy.ma.*) and regular np.* do NOT accept a
MaskedArray. The implicit conversion is NDArray->MaskedArray (unmasked wrap == np.ma.asarray),
the REVERSE of NumPy's subclass upcast. Mask-drop to a plain array is explicit and matches
NumPy: ma.data == np.asarray(ma), np.ma.filled(ma)/ma.filled(), np.ma.getdata(ma).

KNOWN benign nuance (not chased): comparison OPERATORS' masked-position .data. NumSharp's
`ma > nd` delegates to np.ma.greater (fills da at masked -> hidden .data True), matching
NumPy's MODULE np.ma.greater; NumPy's `>` OPERATOR leaves the raw comparison (False) there.
NumPy is INTERNALLY INCONSISTENT on this (operator != module), and both .filled() and the
repr are identical — the differing value is masked/unobservable. NumSharp is self-consistent
(operator == module).

Gate: MaskedArrayTests 25 (added Operators_MixedWithNDArray_AreMaskAware,
Astype_And_Asarray_MatchNumPy), green net8.0 + net10.0.
…ticle per NumPy fundamentals article

Mirrors NumPy's "NumPy fundamentals" user-guide set (doc/source/user/basics.*.rst)
as a new DocFX section under docs/website-src/docs/, converting each NumPy article
into a NumSharp one (C# API, NumSharp semantics, honest about divergences). The
section is wired into docs/website-src/docs/toc.yml as an expanded node placed after
"NDArray".

New landing + 7 new articles (docs/website-src/docs/fundamentals/):
- index.md                — "Fundamentals and usage" landing; maps each article to its
                            NumPy source, states the conversion philosophy (1:1 parity,
                            C# code, documented divergences), and points to the deeper
                            existing guides.
- array-creation.md       — the six creation mechanisms (sequences, intrinsic funcs,
                            join/replicate, disk, raw bytes, random). Pins the
                            int32-vs-int64 gotcha (np.array(int[])=int32 vs
                            np.arange=int64) and the weak-scalar-vs-strong-array
                            downcast difference (C# int[] wraps; a Python list raises).
- indexing.md             — basic (views) / advanced (copies) / boolean / flat / assign;
                            leads with the string-slice syntax (a["1:5:2"]) and the
                            raw-int[]-sole-index-is-fancy gotcha (coordinate access via
                            nd.GetData(coords)).
- io.md                   — the byte-exact .npy/.npz stack (save/load/savez/load_npy/
                            load_npz/mmap_mode + dtype map), text I/O (loadtxt/savetxt/
                            fromstring), raw binary (fromfile/tofile/frombuffer); states
                            genfromtxt/fromregex are not implemented and why.
- copies-and-views.md     — view vs copy, basic->view / advanced->copy, reshape/ravel/
                            flatten, broadcast read-only, detection via arr.@base /
                            arr.Storage.IsView / np.shares_memory, and the C# "+= is not
                            in-place" rule.
- strings-and-bytes.md    — the Char dtype (2-byte UTF-16, not NumPy's 1-byte S1) and
                            why str_/bytes_/void/StringDType are not ported; use .NET
                            string/string[]/byte[] + np.frombuffer.
- structured-arrays.md    — why record/structured dtypes are not implemented and the
                            .NET-idiomatic replacements (parallel NDArrays, struct/record
                            arrays, np.frombuffer for packed binary records).
- ufuncs.md               — elementwise model, out=/where=/dtype=, broadcasting, NEP 50
                            casting, reductions + the reduce-upcast rule (sum(int32)->
                            int64); notes NumSharp exposes direct functions instead of
                            ufunc-object methods (np.add.reduce -> np.sum, ...) and the
                            np.evaluate fused-expression extension.

Reused under the section (no duplication — they are already the NumSharp conversions):
- Data types  -> existing docs/website-src/docs/dtypes.md
- Broadcasting -> existing docs/website-src/docs/broadcasting.md

Design notes:
- Existing top-level "Dtypes"/"Broadcasting" TOC entries are folded into the new
  section (files unmoved, so inbound links are unaffected); the section otherwise adds
  only new pages.
- Every article cross-links to the deeper how-to guides (NDArray, Getting & Setting
  Values, Iterating & Enumerating, Compliance) rather than duplicating them, and to the
  upstream NumPy article it converts.
- Facts verified against src (np.arange integer default = int64; np.vander/np.indices/
  np.frombuffer/np.copy exist; np.load -> object with typed load_npy/load_npz; coordinate
  access via nd.GetData; view detection via @base/Storage.IsView; np.shares_memory).

Validation: toc.yml parses as YAML; 82/82 relative file links resolve; 16/16 cross/self
anchors resolve; code fences balanced.
Adds the four masked-array set operations from numpy.ma.extras' arraysetops
(`__all__`), the last user-facing group of that family still missing. All four
are in NumPy's public surface and were previously DEFERRED. Result is ALWAYS a
MaskedArray, matching NumPy.

Semantics (numpy.ma: masked values are equal only to one another and to no real
value, so every masked element in the inputs collapses to at most ONE trailing
masked entry that sorts last):
  - intersect1d: sorted values UNMASKED in BOTH inputs; +masked iff BOTH inputs
    carry a masked element (masked == masked).
  - union1d:     sorted union of both inputs' unmasked values; +masked iff EITHER
    input carries a masked element.
  - setxor1d:    symmetric difference of the unmasked values; +masked iff exactly
    one input carries a masked element (XOR of masked-presence).
  - setdiff1d:   ar1's unmasked values absent from ar2; +masked iff ar1 has a
    masked element and ar2 does NOT (masked in ar2 removes it). Directional.

Implementation — a REFORMULATION, not a transcription of NumPy's
unique/concatenate/sort/masked-index composition (NumSharp's MaskedArray has no
indexing/slicing, and no MaskedArray->NDArray subclass dispatch): run the plain,
already-optimized np.<setop> over the UNMASKED values (via the existing
`compressed`), then a boolean masked-presence rule
(`HasMasked = mask != null && np.any(mask)`) decides the single trailing masked
slot. No new kernel — the shared substrate helpers `HasMasked`,
`WithTrailingMasked` (mirrors `unique`'s append), and `SetOpDtype` sit beside the
existing set/unique code. Verified BIT-EXACT (filled values + mask + dtype) vs
NumPy 2.4.2 across 24 differential cases: masked/unmasked/all-masked/empty/
disjoint/dtype-promotion(int+float->float64)/nan/duplicates/assume_unique, both
operand orders.

Three traps encoded (each a silent-wrong-answer otherwise):
  - Values are cast to `np.result_type(dataA, dataB)` explicitly. Without it,
    np.union1d(empty_i64, empty_i64) collapses to float64 (from the empty
    concatenate) and the all-masked / empty-result cases would lose the int64
    dtype NumPy preserves. astype is a no-op copy on the non-empty path (the set
    op already promoted), so it only bites the empty case — no truncation risk.
  - The trailing masked datum is a zero of the result dtype: it is hidden by the
    mask, and NumPy's own raw datum there is likewise arbitrary; `.filled()` (the
    observable value) then matches NumPy's default fill.
  - Every mask-null test uses `is null`/`is not null` — `== null` on an NDArray
    runs the ELEMENTWISE comparison (the file-wide ma trap).

Perf (NPY/NS, best-of, warm; higher = NumSharp faster):
  - 1M: union 3.8x, intersect 2.5x, setdiff 3.2x, setxor 2.9x — NumPy's ma
    subclass overhead (Python-level unique/sort/concatenate over the MaskedArray
    subclass + mask bookkeeping) dominates at scale.
  - 100K: ~1.1-1.65x — the masked wrapper overhead is thin (compressed 0.125ms;
    ~13% over the raw np.<setop>); the case is bound by the shared scalar-radix
    sort core (the documented vqsort gap), not by anything in these wrappers.
    Closing 100K uniformly is a shared-sort-core change, not a property of these
    four functions.

Tests: MaskedArrayTests +4 methods (Intersect1d/Union1d/Setxor1d/Setdiff1d
_MatchesNumPy), all probed against NumPy 2.4.2. Full class 29/29 green on both
net8.0 and net10.0.
…e NumSharp article per NumPy advanced article

Mirrors NumPy's "Advanced usage and interoperability" user-guide section
(user/index.rst: c-info, f2py, dev/underthehood, basics.interoperability) as a new
DocFX section under docs/website-src/docs/, converting each NumPy article into a
NumSharp one. Where NumPy's advanced story is about dropping into C/Fortran, NumSharp's
is the opposite — Core is 100% managed with no P/Invoke — so each article reframes the
NumPy topic honestly for a managed runtime.

New landing + 3 new articles (docs/website-src/docs/advanced/):
- index.md                 — "Advanced usage and interoperability" landing; maps each
                             article to its NumPy source and states the "no C-API" thesis
                             (managed seams + runtime IL/SIMD instead of C).
- extending-numsharp.md    — C-API analog. The managed extension seams: typed/unboxed
                             iteration (np.nditer<T>/nditer_chunks<T>/nd.Unsafe), fused
                             custom ops (np.evaluate/NDExpr incl. Call), the general NDIter,
                             runtime IL kernels (ILKernelGenerator/DirectILKernelGenerator),
                             the TensorEngine.Blas backend seam, NumSharp.Build weaver, and
                             np.multithreading. Notes frompyfunc/vectorize are absent.
- native-backends.md       — F2PY analog. Pure-managed default (no native dep/P/Invoke);
                             the one seam IBlasBackend on TensorEngine.Blas; the flagship
                             NumSharp.Interop.OpenBLAS (byte-identical products +
                             LAPACK factorisations, bundled binary, Enable/Disable/Info,
                             thread/coretype parity levers); how to write your own backend
                             + the 3 review-pinned rules (read Blas into a local, failed
                             Enable is a no-op, element-strides not bytes); the getenv trap;
                             the when-native-vs-managed table.
- under-the-hood.md        — internals analog. The buffer+metadata split; UnmanagedStorage
                             + ARC (Release/Abandon); Shape readonly struct + ArrayFlags
                             (O(1) flags); the strides-in-elements(internal)-vs-bytes(public
                             nd.strides) gotcha; views as metadata-only reinterpretation;
                             broadcast stride-0 read-only; C-order indexing + the matrix/image
                             convention discussion; how NDIter/IL-gen traverse; inspection.

Reused under the section (nested, not duplicated — it is already the NumSharp conversion):
- Interoperability -> existing docs/website-src/docs/interop/index.md (the one-buffer
  contract + every bridge). The existing top-level "Interoperability" TOC node (index + 9
  package pages) is moved to nest under the new section, mirroring NumPy's combined caption;
  files unmoved, so inbound links are unaffected.

toc.yml: the standalone "Interoperability" node is replaced by "Advanced usage and
interoperability" containing the 3 new articles + the nested Interoperability subsection.

Facts verified against src: nd.strides is in BYTES (NumPy parity) while Shape.strides is
in elements; np.frompyfunc/np.vectorize do not exist; nd.Unsafe.{Span,Memory,Bytes,Pointer}
+ np.multithreading exist; IBlasBackend/ISlidingDotBackend/OpenBlasEngine surface per
CLAUDE.md; ArrayFlags values from Shape.cs.

Validation: toc.yml parses as YAML; 54/54 relative file links resolve; 2/2 anchors resolve;
code fences balanced. (The em-dash-in-heading anchor into ufuncs.md was dropped in favor of a
plain page link to stay slugger-independent.)
Nucs added 23 commits September 27, 2026 09:54
…4, SeedSequence/Generator/BitGenerator members carry NumPy's types, NumPy's seed-string rule (no octal), get/set_bit_generator and the module seed of a swapped engine

An audit of the whole random suite against NumPy 2.4.2: every public member's name, parameters, defaults,
return and error types, and every integer the suite accepts, stores or returns, checked against the type NumPy
uses (the C `long`, `npy_intp`, `Py_ssize_t`, `uint32_t`, or a Python int).

Method
- Surface: NumPy's signatures (inspect + the .pyx sources) against a C# reflection dump of Generator,
  NumPyRandom (RandomState), BitGenerator/PCG64/PCG64DXSM/Philox/SFC64/MT19937, SeedSequence, ISeedSequence,
  ISpawnableSeedSequence, SeedlessSeedSequence and the state classes, plus a per-parameter integer-type table.
- Behaviour: two NumPy 2.4.2 probes, each run on Windows (LLP64: C long = int32) and on Linux x86-64 in WSL
  (LP64: int64). A C# replay of the same ~250 cells through NumSharp gave 226 identical to Linux NumPy; the 10
  that differ are the deliberate divergences listed below. Two dozen extra NumSharp-only cells were then checked
  against NumPy and all matched.
- Windows vs Linux NumPy differ ONLY in the legacy dtype (int32 vs int64), in the bounds past 2**31 that the
  Windows build rejects, and in three C-long-dependent message texts. The VALUES are identical wherever both
  builds accept the input.

1. The legacy C long is ONE width: int64 (LP64), NumPyRandom.LegacyLong
   NumSharp already modelled the LP64 long for the discrete samplers (int64_to_long). randint, random_integers,
   permutation, choice and multinomial were the odd ones out: they returned the win-amd64 int32 and rejected
   bounds past 2**31, a mix no NumPy build produces. Now:
   - randint (both overloads) defaults to int64 — mtrand maps dtype=int to np.dtype("long"). The draws are
     unchanged: a range that fits 32 bits takes the same buffered 32-bit masked sampler at either width.
     randint(0, 2**40) and randint(0UL, 2**63) now draw Linux NumPy's values; 2**63 + 1 is "high is out of
     bounds for int64".
   - random_integers = randint(low, high + 1, dtype='l') in int64; random_integers(0, long.MaxValue) reaches the
     int64 maximum (high + 1 = 2**63 is exact).
   - permutation(int/long) returns int64 and goes through LegacyPermutation(BigInteger) + ArangeLength, a port
     of PyArray_ArangeObj's _calc_length / _arange_safe_ceil_to_intp. The length is ceil(double(n)) checked
     against (double)NPY_MAX_INTP == 2**63. A stop whose double rounds to exactly 2**63 (the band
     [2**63 - 512, 2**63 + 1024]) casts to INT64_MIN on x86, an EMPTY range, on Windows and Linux alike (probed).
     Larger stops are "Maximum allowed size exceeded"; 2**62 is "array is too big".
   - choice: the indices are int64. The weighted path is astype(long). A uint64 population past int64 with
     replacement is randint(0, pop) in int64: 2**63 draws, 2**63 + 5 and 2**64 - 1 are NumPy's "high is out
     of bounds for int64". Without replacement it is permutation(pop)[:size]: the empty band leaves too few
     indices, which is NumPy's "cannot reshape array of size 0 into shape (3,)". The weighted
     no-replacement `found` buffer is long[].
   - multinomial: n is a long (NumPy's `long n`) and the counts are int64. multinomial(2**40, [.5, .5]) draws,
     where Windows NumPy raises "Python int too large to convert to C long".
   - The overflow texts that depend on the C long's width follow LP64 NumPy (probed on Linux): Philox's
     list-key overflow is "Python int too large to convert to C long" (Windows: "int too big to convert").
     SeedSequence.spawn's uint32 overflow is "value too large to convert to uint32_t" (Windows: "Python int too
     large to convert to C unsigned long").

2. Sizes and seeds
   - dirichlet/hypergeometric/logseries/pareto/power/multivariate_normal: the single-integer `size` overload
     takes a long (one npy_intp dimension; it was int).
   - random_sample/random/ranf/sample(Shape): NumPy takes ONE size argument. default is size=None (a scalar),
     Shape.Scalar is size=() (0-d), anything else the shape. These sit beside the params long[] spellings.
   - NumPyRandom.Seed is a uint. It was an int, so a valid legacy seed in [2**31, 2**32 - 1] read back
     negative. Every setter stores the seed itself.

3. get_bit_generator / set_bit_generator and the module seed of a swapped engine
   - get_bit_generator() returns the engine; set_bit_generator(bitgen) swaps it under the new engine's lock and
     discards the cached Gaussian (NumPy's _initialize_bit_generator). null is NumPy's AttributeError "'NoneType'
     object has no attribute 'capsule'", and the engine is left as it was ([Misaligned]: NumPy assigns first
     and leaves its singleton holding None).
   - NumPy's MODULE seed differs from the RandomState METHOD. After a swap to a non-MT19937 engine,
     np.random.seed(x) re-seeds the engine in place as `engine.state = type(engine)(x).state`, with the
     SeedSequence rules: seed(2**40) is legal, [] is legal, a negative seed is "expected non-negative
     integer". The cached Gaussian is KEPT; probed: the next standard_normal returns the pair's cached half.
     TrySeedSwappedSingleton implements exactly that, for np.random itself only (ReferenceEquals). Any other
     RandomState over such an engine still raises "can only re-seed a MT19937 BitGenerator", NumPy's method
     rule. The fresh engine comes from the new internal BitGenerator.NewSeeded (the existing type(self)(seed)
     factory behind spawn). Seed is not touched on this path.

4. Generator and the bit generators
   - Generator(bit_generator): NumPy's parameter name (was bitGenerator). null is the AttributeError NumPy's
     constructor raises reading bit_generator.capsule (was ArgumentNullException). _bit_generator and
     _poisson_lam_max are public, as in NumPy.
   - choice(long a, size, replace, p, int axis, bool shuffle): the axis slot was missing, so a ported
     positional call bound its fifth argument (NumPy's axis) to shuffle. permutation(long x, int axis = 0)
     accepts and ignores the axis for an integer, as NumPy does.
   - PCG64/PCG64DXSM/Philox/SFC64/MT19937(ISeedSequence seed): NumPy's parameter name (was seedSeq). null is
     NumPy's seed=None, a fresh OS-entropy SeedSequence (was ArgumentNullException).

5. SeedSequence: NumPy's member types and coercion
   - spawn_key is BigInteger[] (a tuple of Python ints; a key element of any size is legal). pool_size is a
     long (Py_ssize_t), n_children_spawned a uint (uint32_t), pool a uint32 NDArray. The full constructor is
     (object entropy, object spawn_key = null, long pool_size = 4, uint n_children_spawned = 0); a pool_size
     past Array.MaxLength is OutOfMemoryException (NumPy's MemoryError from np.zeros).
   - generate_state(long n_words, DType dtype) returns an NDArray (NumPy's ndarray, in unmanaged storage, so the
     count is not capped at a managed array's length); the engines' internal fast paths share one FillState. A
     uint64 request is DOUBLED before NumPy's np.zeros, so a doubled count past npy_intp (n >= 2**62 or
     n < -2**62) is "Maximum allowed dimension exceeded" where a single-width 2**62 is "array is too big".
   - spawn_key coercion = NumPy's tuple(spawn_key) + _coerce_to_uint32_array, stored as the flat Python ints the
     coercion reads. Any iterable is accepted, and a string key is a sequence of one-character strings
     (tuple("12") is ('1', '2')). Nested sequences and integer arrays flatten; the pool is identical because
     ((1, 2),) and (1, 2) coerce to the same words. Bools count as ints. NumPy's refusals are verbatim: a 0-d
     array key is "iteration over a 0-d array", a bool array "object of type 'numpy.bool' has no len()", a
     float "seed must be integer", None "object of type 'NoneType' has no len()", an int "'int' object is not
     iterable".
   - ParseSeedString = NumPy's string rule, for entropy AND spawn keys. "0x…" is int(x, 16); a string that
     starts with a decimal digit (DECIMAL_RE.match) is int(x) in base 10; anything else is "unrecognized seed
     string". This fixes a real stream divergence: the entropy coercion read a leading-0 string as OCTAL, so
     SeedSequence(["012"]) seeded from 10 where NumPy seeds from 12, and every stream built on it differed.
     SeedSequence(["08"]) raised where NumPy reads 8, and "abc" reported "invalid literal" where NumPy says
     "unrecognized seed string". The existing SeedSequence_EntropyCoercion_Errors pinned the octal reading; it
     is corrected from NumPy's own output.
   - ISeedSequence.generate_state(long, DType) returns an NDArray (NumPy's protocol). SeedWords32/64 read the
     custom sequence's array and refuse a wrong dtype or a short array; NumPy reads PyArray_DATA unchecked.

Oracle
- gen_oracle.py: _RND_INT64_CAST += randint/permutation/choice/multinomial, and generator_parity's
  random_integers is recorded widened to int64 (the Windows authoring host's long is int32). Comment blocks
  updated.
- Regenerated random_parity(.jsonl, _host) and generator_parity(.jsonl, _host) with Windows NumPy 2.4.2. Exactly
  30 cases changed, all int32 -> int64 with identical values: random_parity choice 6 / permutation 4 /
  randint 12, random_parity_host multinomial 3, generator_parity random_integers 5. No other byte changed, and
  no legacy int32 result is left. Re-generated on the rebased base: byte-identical.
- OracleSurfaceCoverageTests classifies get_bit_generator/set_bit_generator as non-stream API; LeakCatalogue
  measures both on the fixture's own RandomState (a self-swap: np.random is never touched).
  OpRegistry.Generator passes choice's shuffle by name (the new axis slot); the random-isolation test reads
  Seed as uint.

Tests
- New RandomSampling/RandomTypeParity.Test.cs: 21 tests, every value and message from Linux NumPy 2.4.2.
  - RandomTypeParityTests: every legacy integer is int64; randint LP64 bounds; choice populations past int32;
    permutation's arange edges; multinomial counts past int32; Seed as uint; random_sample(Shape);
    long size overloads.
  - SeedSequence tests: member types; spawn-key coercion; entropy strings; generate_state's long counts and
    errors; a custom ISeedSequence returning NDArray words (and refused bad arrays).
  - Generator tests: constructor/attributes; the axis slot; the bit generators' `seed` parameter and null.
  - RandomSingletonBitGeneratorTests ([DoNotParallelize], restores the singleton's engine, state and Seed):
    get/set_bit_generator; set(null); the module seed for all four non-MT engines; the kept Gaussian; only the
    singleton re-seeds, and MT19937 keeps the legacy validation.
- Updated the pins of the old int32 legacy dtype, and reads of int64 data through GetInt32 / ArraySlice<int> /
  Data<int>, to int64: OpenBugs.Random, LegacyRandomState, RandomState.AuditFixes, multinomial, randint,
  random_integers, AuditV2. RandomIntegers_ByteExact_AndBounds now draws the LP64 bounds instead of
  expecting the int32 rejection.
- WhereSimdTests.Where_Simd_Int32_Correctness and the benchmark int32 operand factories (BenchmarkBase,
  SimdVsScalarBenchmarks) relied on randint's old int32 default; they pass int32 explicitly so they keep
  exercising int32.

Deliberate divergences (documented, pinned where observable)
- SeedSequence.spawn past 2**32 children: NumPy's uint32 loop counter wraps and never ends (MemoryError,
  probed); NumSharp raises the OverflowError NumPy's += would ([Misaligned], BitGeneratorFamily).
- set_bit_generator(null) leaves the engine; get_state() on a non-MT19937 engine refuses where NumPy warns and
  returns the dict (a C# method cannot change its return type; get_state(legacy: false) is the dict).
- spawn_key reprs show the flattened ints ((12,) for ('12',), (1, 2) for ((1, 2),)); the pools are identical.
  A 0-d or multi-dimensional uint32 array nested INSIDE a spawn key is flattened where NumPy's concatenate
  refuses it.
- capsule/cffi/ctypes (CPython FFI) have no .NET meaning; NumPy's module-level functions (get_bit_generator,
  default_rng, ranf, …) live on the np.random singleton; spawn(int n_children) stays int (the result is an
  array).

Verification
- Unit suite: 17,477 passed (net10.0) / 17,476 (net8.0), 0 failed.
- Oracle suite: 195/195 on both frameworks (FuzzMatrix replay of the regenerated corpora, the surface guard,
  the leak audit).
- The whole solution builds with 0 errors (benchmarks, examples, interop and analyzer projects included).
- Mutation check, each re-introduced alone and restored, 7/7 killed:
  - LegacyLong = Int32: 5 tests.
  - The octal seed-string rule: 3.
  - No module-seed branch: 2.
  - No uint64 dimension check: 1.
  - A string spawn key not iterable: 1.
  - A truncated Seed: 1.
  - A Gaussian reset on the engine re-seed: 1.
- Rebased onto da84683f (it touched .claude/CLAUDE.md and gen_oracle.py in other regions; no conflict) and
  re-verified there.

Docs: compliance.md (the uniform LP64 long, get/set_bit_generator + the module seed, SeedSequence's types and
string rule), CLAUDE.md's Random section.

Note for the parallel random-oracle work in the main tree (docs/plans/random-oracle-coverage.md, uncommitted
when this landed): its reflected inventory (random_surface.json) predates these signature changes, and a
Windows-authored random_api corpus must record the legacy ints widened to int64, as gen_oracle.py does.
…verload replayed against NumPy 2.4.2 by exact C# signature, 5 engines x 10 seeds, state after every call; 8 NumSharp divergences fixed

The random world's oracle so far checked one C# overload per NumPy call pattern, on PCG64 (Generator) and the
legacy MT19937 only, at two seeds. This commit starts the plan in docs/plans/random-oracle-coverage.md: every public
member of the random suite, invoked BY ITS EXACT C# SIGNATURE, compared with NumPy 2.4.2 on every engine and ten fixed
seeds, with the receiver's full state recorded after each call. P0-P3 are done here (NumPyRandom, NumPyRandom.State,
NativeRandomState, Generator); P4 (bit generators, SeedSequence, state classes, default_rng) and the gates/soak/docs
(P5-P7) follow.

How a case works
- test/oracle/random_surface.json is the committed inventory (439 members over 20 types, one canonical signature
  each: Type.name(p1,p2), params long[], .ctor, .get/.set, .field, op_Implicit). The G1 test
  (Fuzz/RandomApi/RandomApiSurfaceTests) reflects the surface and fails on any drift; NUMSHARP_WRITE_RANDOM_SURFACE=1
  rewrites the file. It caught the parallel 9aac6ee9 (random type parity) mid-work: 8 new members, 13 changed
  signatures (pareto(double,int) -> pareto(double,long), ...).
- test/oracle/gen_random_oracle.py runs NumPy and writes the corpus. Every inventory member must be claimed by a
  family or an exemption (--partial relaxes that during the build-out). A case names its overload (params.sig), its
  NumPy-level member (params.member, the coverage join's parameter), the receiver (RandomState legacy-seeded or over
  ENGINE, Generator(ENGINE), a bare engine, a state tuple/dict of another receiver, or none), a priming (u32 buffered
  half, gauss cache, raw3), a repeat count, and every parameter by name - an omitted optional is the C# default path
  and NumPy is called without it, so a default that differs shows up.
- Arguments: doubles by IEEE bits, integers as decimal text, Shape as dims/none, DType by name, NDArray operands
  through layout_catalog (views derived from a 1-D C-contiguous base), arrays, objects (bit generators, bit
  generator states, legacy tuples and dicts of other receivers or spelled out, Python int/str) built identically on
  both sides and fresh per case, aliasing (permuted(x, out=x) passes one operand to both), typed generic arguments
  (randn<T> is NumPy's cast of the float64 draw to T). A value no C# parameter of that type can carry (2**64 for a
  ulong) is not a case.
- Observations: arrays (dtype/shape/bytes), Python scalars, text, objects (bit generator: type + state text;
  RandomState/Generator: str + state; tuple/dict/bit-generator states as canonical text), sequences (spawn), and the
  receiver's canonical state after the call (MT19937 key by SHA-256; PCG64/DXSM/Philox/SFC64 in full, buffered half
  included; RandomState adds has_gauss and the gauss bits). OS-entropy words are masked identically on both sides
  (seed(), RandomState(), null seed arrays). Operands a member mutates (shuffle, out=) are observed after the call.
- The C# harness (Fuzz/RandomApi/RandomApiHarness.cs) resolves the signature by reflection and invokes exactly that
  member; RandomApiObservation renders NumSharp's result the same way; RandomApiDivergences holds the intended,
  bounded divergences (keyed on member AND a bound or condition, printed by every replay).

Tiers (test/NumSharp.Tests.Oracle/Fuzz/corpus)
- random_api.jsonl        6,511  libm-free members; hard on every host (also replayed green under Linux .NET/glibc)
- random_api_host.jsonl  15,733  samplers whose transform consumes libm; win-amd64 authored, Inconclusive elsewhere
- random_api_mvn.jsonl      705  multivariate_normal on both APIs through NumPy's own scipy-openblas at one thread
                                 (the linalg_parity pin; byte-exact), win-amd64
- random_api_lp64.jsonl     206  answers only a 64-bit C long gives (randint(0, 2**40), choice(2**40), tomaxint,
                                 geometric(5e-324), zipf past LONG_MAX, ...): the generator re-runs the legacy C-long
                                 families under Linux NumPy in WSL and a case moves here when Linux's answer leaves
                                 int32 or the platforms disagree on raising (ids are content-stable; the merge refuses
                                 an id naming different cases on the two platforms). Hard on x64 Windows and Linux.
Coverage: 29 samplers on both APIs (every overload, size forms, omitted optionals, nulls, scalar variants including
NumPy's rejections, priming, repeated calls, array layouts and dtypes); rand/randn (loose dims and Shape), randn<T>
over 13 dtypes, random_sample/random/ranf/sample, standard_cauchy/exponential/normal, tomaxint, bernoulli (NumSharp's
own, as NumPy's random_sample(size) < p), bytes (negative lengths included), choice, permutation, shuffle,
multinomial, dirichlet, randint (every dtype at its edges), random_integers, get/set_bit_generator,
_bit_generator, _poisson_lam_max, str, get_state/set_state (tuple, dict, bare bit-generator state, rejections),
seed (all 8), the RandomState factories, and the state objects' fields/properties/constructors. Generator: random,
standard_normal/exponential/gamma/cauchy with dtype/method/out= in every out layout NumPy distinguishes, integers
(endpoint, every dtype), choice (N-D along an axis, shuffle=False), permutation, shuffle, permuted (out=x), bytes,
dirichlet (both algorithms), multinomial (array n, 2-D pvals), multivariate_hypergeometric (marginals/count),
multivariate_normal (svd/eigh/cholesky), spawn, attributes, the constructor.

NumSharp divergences found and fixed (each pinned in test/NumSharp.Tests/RandomSampling/RandomOracleFindings.Test.cs)
1. rand(default(Shape)) threw ArgumentNullException; it is NumPy's rand(): one draw, 0-d.
2. None arguments threw NullReferenceException in legacy choice/permutation/shuffle and Generator
   choice/permutation/permuted/shuffle. np.asarray(None) is a 0-d object array, so each now raises what a 0-d argument
   gets, verbatim (permuted including its out checks and the dtype('O') cast refusal).
3. A null seed array raised "Seed must be non-empty" (seed(int[]/long[]/uint[]), RandomState(...),
   MT19937._legacy_seeding(...)); NumPy's seed(None) takes OS entropy. An EMPTY array still raises.
4. has_gauss was a bool, so set_state(..., has_gauss=2, ...) read back 1; NumPy keeps the C int until the cached
   Gaussian is consumed. _hasGauss is now that int.
5. Generator.permuted's safe-cast error said "array data" for a 0-d source; NumPy's copyto says "scalar".
6. Legacy multinomial leaked its output on the n < 0 error path (NumPy allocates np.zeros before checking n; NumSharp
   mirrored the order but never released the array). Caught by the leak sweep replaying the new tier.
7. vonmises(mu=NaN) returned -NaN on net8.0 only: .NET 8's double % returns the DEFAULT NaN (0xfff8..., sign set) for
   a NaN dividend where C's fmod (NumPy) and .NET 10 propagate the input NaN. Both vonmises wraps pass a NaN through.
   The same % is the float np.fmod/np.mod kernel on net8.0; the ordinary oracle tokenizes NaN bits, so that is
   recorded in the plan (finding 9) for a separate look.
8. (harness) An omitted `Shape size = default` was passed as Activator.CreateInstance(typeof(Shape)), which runs
   Shape's parameterless constructor (Shape.Scalar, size=()) where the compiler passes default(Shape) (size=None);
   DefaultOf now uses an uninitialized struct.

Intended divergences (documented in the plan, excused with a key and a condition)
- Generator.pareto/power: Sun's s_expm1 vs the win-amd64 CRT, <= 2 / ceil(2/a)+3 ULP, stream positions identical.
- NumPyRandom.get_state() on a non-MT19937 engine: NumPy warns and returns the dict; the tuple-typed overload raises
  a ValueError naming get_state(legacy: false).
- set_state / RandomState(NativeRandomState) with pos outside [0, 624]: NumPy stores it and its next draw reads past
  the key (undefined behavior); NumSharp refuses.
Not in the corpus by design: randint(dtype='l') (np.dtype('l') is the host's C long in both libraries), and
multivariate_hypergeometric(method='count') near sum(colors)=10**9 (an 8 GB index array in both).

Wiring
- FuzzCorpusTests.RandomApi.cs: RandomApi / RandomApiHostLibm / RandomApiMultivariateNormal / RandomApiLp64; every
  failure is listed (the full list to the test output), the slowest cases are reported.
- FuzzCorpus.Expected gains Result, State and After.
- UndisposedIntermediateTests routes random_api cases through the harness (they were counted as "threw" before) with
  a random_api:<member> coverage key, and LeakSurfaceCoverageTests credits np.random / Generator members through it.
- coverage/oracle_map.json: the random_api key (param member, template numpy.random.{value}, 61 reviewed overrides:
  RandomState members also credit the np.random module function, Generator members do not, constructors map to the
  class rows, bernoulli to numsharp.random.bernoulli, attributes without a catalog row to nothing), host pins for the
  three pinned tiers, and the four aliases the new direct contracts made stale (random/ranf/sample onto
  random_sample) deleted as the join requires. coverage/test_oracle_evidence.py's normal pin names the new key.
  Oracle-verified headline APIs: 486 -> 493 of 560.

Verification (in a detached worktree at 9aac6ee9 plus exactly these changes; the shared tree carries a parallel
session's uncommitted polynomial work)
- NumSharp.Tests: 17,483 / 17,482 passed on net10.0 / net8.0 (0 failed).
- NumSharp.Tests.Oracle: 220 passed on net10.0 and net8.0 with NumPy's own OpenBLAS bound (5 bundled-binary skips).
- WSL Linux .NET: random_api and random_api_lp64 green (host and mvn Inconclusive by design).
- python coverage/generate_coverage.py (strict join) and pytest coverage/test_oracle_evidence.py: green.
…ses, SeedSequence, SeedlessSeedSequence and default_rng replayed against NumPy 2.4.2 by exact C# signature (+1,888 cases, whole inventory claimed); 3 NumSharp divergences fixed

P4 of docs/plans/random-oracle-coverage.md. After P2/P3 (879791da) covered NumPyRandom, NumPyRandom.State,
NativeRandomState and Generator, the remaining 162 inventory members are generated here, so every one of the
439 public members of the random world is now either replayed against NumPy 2.4.2 by its exact C# overload or
exempt with a reason (3: NativeRandomState(byte[]), NumPyRandom.Seed, BitGenerator.lock).

What the new families generate (test/oracle/gen_random_oracle.py)
------------------------------------------------------------------
- Engine constructors, all five engines, every overload: long (the 10 fixed seeds and the edges 0, 2**32,
  2**63-1, -1), ulong (2**63, 2**64-1), BigInteger (2**64, 2**100, 2**128+5, -1), int[] / long[] / uint[]
  (empty, one word, negative; null = NumPy's None), ISeedSequence (SeedSequence with every keyword -
  spawn_key, pool_size, n_children_spawned - a list entropy, a 2**100 entropy, and the seedless one, which
  NumPy refuses with "seedless SeedSequences cannot generate state"; null = OS entropy), and the
  parameterless one (OS entropy). OS-entropy cases mask the drawn words (MT19937 key, PCG64/DXSM state+inc,
  Philox key, SFC64 state) and the seed sequence's pool/repr on both sides.
- Philox(seed, counter, key) in every combination NumPy distinguishes: counter as int / list / uint64 array /
  2**200 / negative, key as int / list / array / 2**128-1 / 2**128 (refused) / negative, key with counter,
  seed and key together (refused), a SeedSequence seed, string/float/negative seeds, all None.
- advance (0, 1, 7, 2**64, 2**128-1, 2**128, 2**200, -1: NumPy wraps the delta), jumped (the new engine is
  built as NumPy builds it - a fresh OS-entropy seed sequence, then the copied and jumped state: its seed
  sequence is masked, the state exact; negative and 0 counts; BigInteger counts on PCG64/DXSM/Philox; MT19937
  kept to small counts because NumPy loops one polynomial jump per count).
- state getters and setters, typed (ENGINE.state) and base (BitGenerator.state): a state from another
  seed/priming, a fresh one, null (NumPy's "state must be a dict"), another engine's state (base setter only -
  the typed setter cannot even be handed one in C#), and per engine: MT19937 pos 0/624/625/-1 and 10/624/625-word
  or unset keys; PCG64 explicit 128-bit words, an even increment, a buffered half, has_uint32=2; Philox explicit
  words, buffer_pos 4/5/-1, 3-word counter, 3-word key, unset counter/key/buffer, 3-word buffer; SFC64 explicit,
  3 words, 1 word (NumPy broadcasts it into all four), none, unset.
- MT19937._legacy_seeding (the int path with its 2**32-1 bound, init_by_array for every array type, an empty
  array's "Seed must be non-empty", None as OS entropy).
- BitGenerator's random_raw (size None / () / (2,3) / (0,) / omitted, output=False including NumPy's quirk of
  drawing sum(size) words - (2,3) draws 5 - and a 1000-word draw), seed_seq and spawn (0, 1, 3, -2) on every
  engine and on a legacy-seeded MT19937 (no seed sequence: "The underlying SeedSequence does not implement
  spawning.").
- The five State classes: every getter against NumPy's state dict, every setter (valid, edge and unset values -
  the state object holds whatever it is given; the engine validates), every constructor (full, defaults
  omitted, an unset array, parameterless), BitGeneratorState.bit_generator, and MT19937.State's conversion from
  the legacy tuple (NumPy's MT19937.state setter translating a tuple; another algorithm's tuple refused with that
  setter's text).
- SeedSequence: the 8 constructors (int / list / nested list / uint32 array / int64 ndarray / 0-d ndarray /
  str in NumPy's decimal and 0x forms and refused forms / float / negative / bool / None; spawn_key as list, big
  int, str, nested list, int (refused), ndarray, 0-d ndarray (refused), negative; pool_size 8 and 3 (refused);
  n_children_spawned up to 2**32-1), entropy / spawn_key / pool_size / n_children_spawned / pool / state / repr
  on four receivers, generate_state (0, 1, 4, 1000 words; uint32, uint64, null = the default, omitted, int32 and
  float64 refused; -1), spawn (0, 1, 3, -1: NumPy's OverflowError after an empty loop; repeated calls continue
  the numbering).
- SeedlessSeedSequence (constructor, generate_state refused, spawn returning itself) and the ISeedSequence /
  ISpawnableSeedSequence interface members dispatched to both implementations.
- default_rng, all 13 overloads: integer / array seeds as above, each engine wrapped as is (fresh and primed),
  each Generator passed through, RandomState's engine shared (legacy-seeded, primed, and over every engine),
  seed sequences, int64/uint32/empty/float/negative/strided/2-D ndarrays, the object overload with every kind
  of value, and null (OS entropy) on every reference-typed overload.

Observations and harness
------------------------
- Engines and Generators are observed with their seed sequence too (seedseq text: children counter, the mixed
  pool's SHA-256, NumSharp's repr - which reproduces NumPy's - as the last, maskable field); a seed sequence,
  the spawn_key tuple, the state dict and the entropy object are compared through NumPy's repr; an engine's
  typed State and random_raw(output=False)'s None have their own observation kinds.
- The C# side decodes the new receivers (a typed State, a RandomState's engine, a seed sequence built from the
  same spec as NumPy's, the seedless one) and object arguments (explicit engine states, seed-sequence specs,
  plain Python values for `object` parameters, NDArray operands for them).
- Mapping decisions, recorded in plan §7: a uint[] SeedSequence/engine/default_rng seed stands for NumPy's
  uint32 ndarray (NumSharp prints it as array([...], dtype=uint32)); every C# array in the LEGACY seeding
  (seed, RandomState(...), MT19937._legacy_seeding) is a Python list - a list takes init_by_array at any
  length, where a one-element ndarray would be squeezed to the scalar seed; a null reference is None.

NumSharp divergences fixed (each pinned in RandomOracleFindings.Test.cs)
------------------------------------------------------------------------
1. Engine state setters answered an unset word array with NumSharp's own text ("state['state']['key'] must be
   a sequence of 624 integers"). NumPy's setters subscript the array, so CPython's text is the contract:
   "'NoneType' object is not subscriptable" (MT19937, Philox) and, for SFC64's broadcast into its words,
   "int() argument must be a string, a bytes-like object or a real number, not 'NoneType'". Philox now checks
   in NumPy's read order (counter[i] and key[i] interleaved for i < 4, then the buffer), so an unset key is
   reported before a short counter's missing fourth word.
2. default_rng with a null typed argument threw (BitGenerator: AttributeError "'NoneType' object has no
   attribute 'capsule'"; NDArray, NumPyRandom: ArgumentNullException) or returned null (Generator). NumPy's
   default_rng(None) is a fresh OS-entropy PCG64 Generator, which the object overload already returned.
3. default_rng's parameters were named bitGenerator / generator / randomState on three overloads; NumPy's only
   parameter is `seed`, so a ported default_rng(seed=...) did not bind. Renamed to `seed` (no in-repo caller
   used the old names; the inventory was regenerated by the G1 gate).

Intended divergences added to RandomApiDivergences (plan §7), each keyed on a condition and printed per replay
------------------------------------------------------------------------------------------------------------------
- Philox.state with a negative buffer_pos: NumPy stores it and its next draw reads the word BEFORE the buffer
  (probed: random_raw() then returns 4294967295, memory outside the array); NumSharp refuses with a ValueError.
  Same class as the MT19937 position guard already documented. Keyed on the refusal text and a negative value.
- SeedSequence built with a str / nested-sequence / ndarray spawn_key: NumPy keeps tuple(spawn_key) as given
  and prints it (('1', '2'), ([1, 2],), (np.int64(3), np.int64(4))); NumSharp stores the flattened Python ints
  its coercion reads (the typed BigInteger[] spawn_key, a deliberate decision documented on ToSpawnKey) and
  prints (1, 2). Excused only when the observation matches everywhere except the repr's spawn_key= line - the
  children counter and the mixed pool's SHA-256 must be identical, so every stream built on the sequence is too.

Coverage join (coverage/oracle_map.json, random_api key only)
-------------------------------------------------------------
Constructor members are now the class name alone (Generator, RandomState, MT19937, SeedSequence, default_rng -
the catalog's rows, so the template maps them without overrides); BitGenerator's members invoked on an engine
also credit numpy.random.BitGenerator.<member>; get/set_bit_generator and RandomState's str() credit their
numsharp.random.* extension rows (NumPy has them outside numpy.random.__all__). 154 members, 84 overrides,
9 mapping to nothing (private/NumSharp-only: _legacy_seeding, _poisson_lam_max, SeedlessSeedSequence.*, ...).
Oracle-verified headline 494 of 560, expanded 829 of 2,519; every random catalog row with a NumSharp member is
now oracle-verified except the six .lock rows and numsharp.random.Seed (exempt), the numpy.random module row
itself and the abstract numpy.random.BitGenerator class row.

Corpus
------
random_api.jsonl 6,511 -> 8,399; random_api_host.jsonl 15,733, random_api_mvn.jsonl 705 and
random_api_lp64.jsonl 206 regenerate byte-identical. test/oracle/random_surface.json regenerated by the G1 gate
for the renamed default_rng parameters.

Verification (detached worktree at 879791da + these changes, because a parallel session's uncommitted polynomial
work keeps the shared tree's build red)
-----------------------------------------------------------------------------------------------------------------
- Full Oracle suite: net10.0 220 passed / 5 inconclusive by design, net8.0 the same (the pinned tiers with
  NUMSHARP_OPENBLAS_LIBRARY bound to NumPy's own scipy-openblas).
- Full unit suite (TestCategory!=OpenBugs&TestCategory!=HighMemory): net10.0 17,487 passed, net8.0 17,486.
- Linux .NET 10 under WSL (glibc): RandomApi and RandomApiLp64 green, host and mvn tiers Inconclusive by design.
- coverage/generate_coverage.py and the coverage unit tests (32) green.
- Every P4 replay is printed with its documented excuses: 8x MT19937 position guard, 4x Philox buffer_pos guard,
  3x flattened spawn-key repr.

Remaining (plan §8): P5 gates G2-G6 (overload, parameter, engine, seed and state-presence coverage computed from
the corpus against the inventory), P6 the nightly soak with 10 fresh seeds, P7 floors, docs and final
verification.
… the 24 module constants, bit-exact with NumPy 2.4.2 (polyseries 15,950 cases, 0 excused); CPython machine-number lane, fused scalarmath mapparms, CPython NaN operand priority

Plan unit U1 of docs/plans/numpy-polynomial.md (54 names), for all six bases at once:
  {p}add, {p}sub, {p}trim, {p}line                      (polynomial, chebyshev, legendre, laguerre, hermite, hermite_e)
  {p}domain, {p}zero, {p}one, {p}x                      (the 24 module constants)
  np.polynomial.polyutils: as_series, trimseq, trimcoef, getdomain, mapparms, mapdomain
(format_float belongs to the printing unit U11.)

ENGINE
- Polynomial/Package/NDPolyNumber.cs - PolyNumber: one operand of numpy.polynomial's Python-level scalar code.
  It is a Python scalar, a NumPy scalar or an ndarray, and Binary/Negate/NotZero/LessZero dispatch exactly as
  Python's operator protocol does between THREE arithmetics:
    Python o Python  -> CPython (exact ints, long_true_divide, 3.12 _Py_c_quot, ZeroDivisionError texts)
    NumPy scalars    -> scalarmath (NEP 50, the NAIVE complex product, inf/nan instead of raising)
    any ndarray      -> a ufunc (the fused complex product; a 0-d result comes back as a scalar)
  pycomplex op np.float64 stays CPython (complex's methods accept a float subclass).
  C# boundary = the house NEP 50 map:
    bool/integers/float/double/Complex/BigInteger -> Python scalars
    Half/char/decimal                             -> NumPy scalars
    NDArray (0-d too) / typed C# arrays           -> ndarrays
    object[]/IList -> Python lists;  ValueTuple -> Python tuples
  FromValue<T> classifies a statically typed value without boxing. MakeArray is np.array discovery (polyline(1, 2)
  is int64; Python ints past uint64 are NumPy's object array -> NotSupportedException).
- Polynomial/Package/NDPolySeries.cs - the functions, NumPy's line order and error order throughout:
  - _add/_sub update the LONGER operand in place (c2 on a tie); _sub with len(c1) <= len(c2) negates c2, then adds c1.
  - trimseq returns the input itself or the VIEW seq[:k].
  - trimcoef checks tol < 0 first; c[:1]*0 when nothing survives.
  - getdomain reduces a fresh copy (the +-0/NaN answers are schedule-dependent).
  - mapparms indexes its domains and never converts them.
  - mapdomain converts x only when it is not int/float/complex/np.generic (a bool x becomes a 0-d array), then runs
    ONE fused np.evaluate pass with NEP 50-typed 0-d parameters.
- Backends/Kernels/Direct/DirectILKernelGenerator.PolySeries.cs - every element loop is an IL kernel:
    trim scan (+n = the input itself, -k = the view seq[:k]), NumPy's in-place combine, the tolerance scan
    (|c| > tol in NumPy's comparison dtype; complex via ComplexAbs), cast, one-element scalarmath ops
    (add/sub/mul/div x naive/simd/loop_scalar complex products, negate, != 0 / < 0, box),
    and the fused mapparms kernel (below).
  Scalar kernels sit behind flat per-dtype slot arrays (the ConcurrentDictionary lookup cost ~25 ns of a ~25 ns op).
- Constants: one shared, WRITEABLE instance each (NumPy's module attributes - the same object on every access, so a
  write persists), detached from every NDScope, holding one extra ARC reference so a caller's Dispose() cannot free
  them. Plan decision D6 is updated accordingly; [Misaligned] M6 "read-only constants" is retired.

FAST PATHS - each proven against the exact general lane, never different
- Generic tuple overloads mapparms<T0..T3>((T0,T1), (T2,T3)) and mapdomain<T0..T3>(double | Complex | NDArray x, ...):
  a C# tuple literal binds them (identity beats boxing). One per x kind is REQUIRED: with only the NDArray one,
  mapdomain(complexX, (-1, 1), (0, 2)) is CS0121-ambiguous against mapdomain(Complex, object, object).
- CPython machine-number lane (NDPolySeries.PyNum). Operands are Python ints within long, floats and complexes, from
  tuples AND lists. It runs CPython's arithmetic on machine numbers, bails to the exact BigInteger lane when an int
  intermediate leaves long (overflow-checked subtract/add, Math.BigMul), and uses long_true_divide's < 2^53 fast
  path. The object facades try it first, before opening a scope.
- Fused scalarmath kernel for two 1-D ndarray domains of one non-bool dtype (the ABCPolyBase.domain/window case):
  the six operations run in one call, intermediates wrap/round in the dtype through locals of its CLR type (stloc
  to an int8 local truncates exactly like the per-op narrowing store), and the true division of integer dtypes is
  float64 via EmitConvertTo.
- mapdomain's two 0-d parameters are reused per thread (a hoisted 0-d input is read once, before the pass).
- {p}line opens its NDScope only when an operand is, or converts to, an ndarray.

CPYTHON NaN OPERAND PRIORITY (found by the lane-vs-reference property test)
When both operands are NaN, x86 returns the first source operand's NaN, and RyuJIT swaps commutative a + b / a * b
for register allocation - so two call sites of one C# expression disagreed. The found case:
  pu.mapdomain(nan, (0.5, inf), (-4.98e17, 0.22)):  off is -inf/inf's default (negative) NaN, scl*x the positive x
  NaN; CPython's float_add returns the RIGHT one (0x7ff8...), and the general lane returned 0xfff8...
Probed on the corpus's CPython 3.12 (MSVC), per operator and operand order:
  float_add                         -> RIGHT operand's NaN (generic and specialized BINARY_OP_ADD_FLOAT alike)
  float_sub, float_div              -> LEFT
  float_mul                         -> LEFT once the call site is specialized (BINARY_OP_MULTIPLY_FLOAT); RIGHT on a
                                       site's first, generic execution - interpreter state, so the steady state is modelled
  _Py_c_sum/_diff/_prod/_quot       -> LEFT, every operation
  _Py_c_quot's NaN-divisor branch   -> Py_NAN = the POSITIVE 0x7ff8... (was .NET's double.NaN 0xfff8... - a latent bug)
All Python arithmetic now goes through PyScalar.FloatAdd/FloatSub/FloatMul/FloatDiv/ComplexSum/ComplexDiff/
ComplexProd/ComplexQuotient with explicit NaN tests. The polyeval tier (U3 folds Python-scalar subtrees through the
same PyScalar.Apply) stays 16,606/16,606.

ORACLE + GATES
- gen_oracle.py polyseries -> polyseries.jsonl, 15,950 cases, 0 excused (133 complex64/object cells skipped, #569).
  Covered:
  - add/sub: the full dtype-pair matrix on the power basis (a subset on the others) x trim patterns x layouts,
    0-d and broadcast operands, Python lists;
  - trimcoef tolerances;
  - mapparms over Python tuples/lists: ints past 2^53/2^63/2^64, long overflow inside a difference or product,
    0 / -5 == -0.0, both ZeroDivisionError texts;
  - mapparms over same-dtype 1-D ndarray domains of EVERY dtype: wrapping ints, float specials, strided/reversed views;
  - mapdomain for every x dtype x domain form, including the lane's hand-over forms;
  - {p}line over Python and NumPy scalars;
  - four FACETS per constant (value, identity, writeable, owndata) - a constant has no argument to vary.
- OpRegistry.PolySeries.cs replays every mapparms/mapdomain case through THREE routes, which must agree to the byte
  (or raise the same type and text) before NumPy is compared: the object overload, the generic overload (tuples
  rebuilt element-typed by reflection) and NDPolySeries.MapParmsGeneral (the exact reference). Planted-bug checks:
    - dropping the lane's product-overflow bail -> 27 red
    - corrupting the list lane's long read -> 205 red
    - feeding the fused kernel the wrong numerator -> 129 red
- test/NumSharp.Tests/Polynomial/PolynomialSeriesTests.cs (83): NumPy-probed literals, plus
  MachineNumberLane_AgreesWithTheExactGeneralLane_BitForBit - 400 seeded draws x 7 type mixes x 8 routes against
  the general lane. It is the test that found the NaN-priority bug and a PyNum.ToObject bug: an uncast switch over
  long/double/Complex has the natural type Complex, so every real result was boxed as a Complex.
- Gates brought along:
  - UndisposedIntermediateTests.Properties reads the six submodule singletons (the constants' leak audit);
  - OracleCoverageStrengthTests carries one reviewed FixedDtypeOps reason for the 24 constants (its static ctor);
  - OracleSurfaceCoverageTests / LeakSurfaceCoverageTests include polyutils and properties;
  - coverage/oracle_map.json gains the polyutils prefix rule;
  - coverage/test_oracle_evidence.py synthesizes the constants' NumSharp-only property rows, as the real generator
    emits them (generate_coverage.py runs clean: 298,814 contracts).

VERIFIED
  - Oracle suite: 226/226 on net10.0 and net8.0.
  - Unit suite: net10.0 17,758 and net8.0 17,757 passed, 0 failed.
  - coverage/test_oracle_evidence.py: 15/15.

PERF (NPY/NS, higher = NumSharp faster; pinned to CPUs 4-7, Release, best-of-7)
  polyadd/polysub/chebadd 10-pt    8.7-10.6x      polyadd 100K       13.7x
  polytrim                         12.1-12.6x     trimseq            4.1x
  as_series                        3.1-3.5x       getdomain          2.7-3.9x
  mapparms tuples                  4.3x           mapparms arrays    8.3x
  mapparms lists                   2.4x           mapparms complex   2.7x
  mapdomain scalar                 4.2-5.6x       mapdomain 1K       2.8-3.6x
  mapdomain 100K                   4.5-18x
  {p}line                          0.87-0.89x   <- the NDArray allocation floor: ~230 ns to create and dispose a
                                                   2-element array, while NumPy's whole np.array([off, scl]) is
                                                   ~200 ns. Not the algorithm (measured; the allocation paths
                                                   differ by <40 ns, within noise).

Docs: .claude/CLAUDE.md (U1 section + source-file row), docs/plans/numpy-polynomial.md (U1 as-built, D6, M6 retired,
section 0.1 status), test/NumSharp.Tests.Oracle/Fuzz/README.md (polyseries tier).
…meter-name gate against NumPy's signatures; the gaps they found closed (+132 cases); 2 NumSharp divergences fixed

P5 of docs/plans/random-oracle-coverage.md. The corpus built in P1-P4 (879791da, 844580a4) claims every overload of the
random world; this commit makes CI PROVE the claim and keep proving it: new FuzzMatrix gates measure the committed
corpus against the reflected inventory (RandomApiSurface, itself pinned by G1), so an overload, a parameter, an
engine or a seed that no case exercises now turns CI red instead of going unnoticed.

The gates (test/NumSharp.Tests.Oracle/Fuzz/RandomApi/RandomApiCoverageTests.cs)
-------------------------------------------------------------------------------
They read the corpus only - each line of the four random_api* tiers parsed once and reduced to signature, receiver,
outcome, state presence and per-argument value keys (an operand by the SHA-256 of its serialized form); the replay
tests stay the ones that compare values.
- G2 overloads: every inventory member has a case by its exact signature, or an exemption with a reason
  (NativeRandomState(byte[]), NumPyRandom.Seed get/set, BitGenerator.lock); an exempt member has NO case; every
  case and every exemption names an inventory member.
- G3 parameters, per overload: every optional parameter omitted AND passed a value other than its declared default;
  every required parameter at least two distinct values; every nullable parameter null (explicit, or omitted where
  the default is null) and non-null; every params array lengths 0 and 2+.
- G3 enumerations, per overload: every value NumPy accepts appears in a case NumPy answers, and a value outside the
  set in a case NumPy rejects - Generator.integers / legacy randint dtypes (9 each), the float fillers' and
  standard_gamma's float64/float32, generate_state's uint32/uint64, multivariate_normal's method and check_valid,
  multivariate_hypergeometric's method. standard_exponential(method=) is listed with "no rejection": NumPy takes
  every method other than 'zig' as the inverse transform.
- G3 parameter NAMES: every method/constructor parameter of the NumPy-mirroring types (NumPyRandom, Generator,
  BitGenerator, the five engines, SeedSequence, SeedlessSeedSequence, the two interfaces) carries one of NumPy's
  names for that member, in NumPy's order, so a ported keyword call binds. NumPy's side is a new committed table,
  test/oracle/random_numpy_signatures.json (123 members, written by gen_random_oracle.py from inspect.signature or
  the Cython docstring's first line); NumPy *args members (rand, randn, ranf, sample) accept any names;
  NumSharp-only members (bernoulli, the three ToString) are allow-listed; RandomState(NativeRandomState), NumSharp's
  restoring constructor, is exempt by signature.
- G4 engines: Generator members on all five engines, legacy members on the legacy MT19937 and each engine, the
  BitGenerator / BitGeneratorState / NumPyRandom.State members over every engine's receiver, and every member taking
  an engine-bearing argument (a bit generator, a Generator, a RandomState) takes one of every engine.
- G5 seeds: per receiver kind and engine with answered stream cases, all ten fixed seeds answered.
- G6 state: every answered case on a stateful receiver records the state after the call.
Every table carries its reasons and is self-retiring: an exempt member that gains a case, an exempt rule the corpus
meets, an allow-listed member NumPy turns out to have, or a stale signature-table entry fails the gate. The np.random
factories (RandomState(...), default_rng(...)) are receiver-independent: G4/G5 do not sweep their receiver, G6 checks
it is left untouched, and G4 demands every engine among their engine-bearing arguments.

What the first run found, and the generator changes that close it (test/oracle/gen_random_oracle.py)
---------------------------------------------------------------------------------------------------
- Reference parameters never passed null: the string enumerations (multivariate_normal method/check_valid,
  multivariate_hypergeometric method, standard_exponential method), multivariate_normal's managed double[] /
  double[,] mean and cov, the legacy samplers' typed int[]/long[] size arrays, the params dimension arrays of
  rand/randn/random/random_sample/ranf/sample, and Philox.State's key and buffer. Each now has a null case (Python's
  None on NumPy's side; a null params array is no dimensions at all). The encoder learned null for T[,].
- State constructors whose required parameters took one value: a second full value set per engine, and null for every
  array parameter one at a time.
- geometric / logseries array overloads answered NO engine x seed base case: the 1-D base [p, p+1, p+2] left the
  probability domain, so every engine x seed base case (50 per Generator overload, 60 per legacy one) was NumPy's
  domain error and the streams were never recorded. The samplers now carry a step (0.2 / 0.1) for their array forms.
- set_state(dict) was answered on engine receivers under one seed only: a dict of the receiver's own engine is now
  restored on every receiver x seed (the MT19937 base, and the other engines' rejections, stay).
- The factories recorded no receiver state: RandomState(...) and default_rng(...) now record it (proving a factory
  draws nothing from the RandomState it is called on).
- default_rng(object) passed one engine per kind: now a bit generator, a Generator and a RandomState of every engine
  (and the legacy RandomState), plus the seedless sequence.
- Explicit enumerated values the gate demanded: multivariate_normal method='svd', standard_gamma dtype float64
  (and float16, rejected).

NumSharp divergences fixed (pinned in RandomOracleFindings.Test.cs)
-------------------------------------------------------------------
1. A null params dimension array threw NullReferenceException in rand, randn, random_sample, random, ranf and sample
   (np.random.rand.cs / np.random.randn.cs). A null array only arises from an explicit null - the port of None - so it
   is no dimensions: NumPy's rand() / randn() / random_sample(size=None), one draw, 0-d, same stream position
   (RandomState(42).rand(null) = 0.3745401188473625, randn(null) = 0.4967141530112327).
2. RandomState(BitGenerator) named its parameter bit_generator; NumPy's RandomState(seed) takes the engine as `seed`
   (np.random.cs). Renamed, like default_rng in 844580a4. Every other method/constructor of the NumPy-mirroring types
   already carried NumPy's names in NumPy's order.

Corpus and inventory
--------------------
random_api.jsonl 8,399 -> 8,485, random_api_host.jsonl 15,733 -> 15,779, random_api_mvn.jsonl 705 -> 730,
random_api_lp64.jsonl 206 (unchanged). test/oracle/random_surface.json regenerated by the G1 gate (the renamed
parameter). The coverage join is unchanged (the member set is the same).

Verification (detached worktree at c5ef9461 - the parallel session's polynomial U1 landed during P5 - plus these
changes)
------------------------------------------------------------------------------------------------------------------
- Full Oracle suite: 228 passed / 5 inconclusive by design on net10.0 and on net8.0 (pinned tiers with
  NUMSHARP_OPENBLAS_LIBRARY bound to NumPy's own scipy-openblas); the 7 new gates green.
- Full unit suite (TestCategory!=OpenBugs&TestCategory!=HighMemory): 17,519 passed on net10.0, 17,518 on net8.0.
- Linux .NET 10 under WSL: RandomApi, RandomApiLp64 and every gate green (host and mvn tiers Inconclusive there by
  design).
- coverage/generate_coverage.py and the 32 coverage tests green.
- A verification trap found on the way (scratch tooling, recorded in the plan's state log): the worktree sync copied
  files with `cp -p`, keeping an edited file's OLDER timestamp, so MSBuild's incremental check skipped the compile and
  a replay ran the previous code. The sync now copies changed files only, with a fresh timestamp; every result above
  was re-run after the fix.

Remaining (plan §8): P6 the nightly soak with 10 fresh seeds; P7 floors, docs and the final verification.
…(Linux LP64 job + windows-latest generation and replay on net10.0/net8.0), soak replay test, committed tier floors

P6 of docs/plans/random-oracle-coverage.md. The committed random_api* tiers replay every overload of the random world
under 10 FIXED seeds on every push; the nightly soak now regenerates the same families under 10 seeds drawn fresh
every night and replays them through the same harness, so a stream divergence only some seeds reach still surfaces.

.github/workflows/fuzz-soak.yml - two new jobs (the corpus needs two NumPys)
------------------------------------------------------------------------------
- random-api-lp64 (ubuntu-latest): draws 10 distinct seeds in [0, 2**32) - the legacy RandomState(int) domain every
  family runs - none of them one of the committed fixed seeds (those replay on every push already); a new
  workflow_dispatch input `random_api_seeds` pins them to replay a night. Runs gen_random_oracle.py --lp64-only: Linux
  NumPy's answers for the legacy integer families whose Windows answer depends on a 32-bit C long, uploaded as an
  artifact. The seeds are the job's output.
- random-api-soak (windows-latest, the host the libm tier is authored on; needs random-api-lp64): generates all four
  tiers under those seeds with --lp64-rows (the Linux rows merged exactly as the committed corpus merges them from WSL),
  builds the Oracle project, and replays the directory on net10.0 AND net8.0 (the committed corpus showed one
  runtime-only divergence, .NET 8's double % NaN, so fresh seeds could too) through FuzzCorpusTests.RandomApiSoak,
  binding NumPy's own scipy-openblas from numpy.libs for the multivariate-normal tier (content-pinned; that tier goes
  Inconclusive, not red, if it ever stops matching). Evidence (manifest.json + trx) is uploaded every night; the whole
  corpus on failure.
- The header comment names the new jobs.

test/oracle/gen_random_oracle.py - soak mode
--------------------------------------------
- --lp64-rows PATH: merge the rows of an --lp64-only run made elsewhere (same seeds; the ids are content-derived, and a
  row naming a different case than the Windows one still aborts the merge) instead of running the WSL sub-run.
- --seeds is validated: distinct integers in [0, 2**32) (a repeat would generate id-colliding duplicates).
- The output directory is created on demand (a CI job passes a fresh path).
- An --out directory gets manifest.json: seeds, NumPy and Python versions, platform, the LP64 source, per-tier counts.
- The module docstring describes the soak pipeline.

test/NumSharp.Tests.Oracle/Fuzz/FuzzCorpusTests.RandomApi.cs
------------------------------------------------------------
- RandomApiSoak (FuzzMatrix, DoNotParallelize): replays every tier of the directory in NUMSHARP_RANDOM_SOAK_DIR with
  the same harness and comparator as the committed tiers, each tier under its own host gating (the libm tier on
  Windows, the LP64 tier on x64 Windows/Linux, the mvn tier through the pinned OpenBLAS, disabled afterwards), all
  tiers' divergences collected before failing. Inconclusive when the variable is unset (every per-push run); a missing
  tier is a failure (the pipeline broke), not a skip. Prints the manifest.
- RunRandomApiCorpus takes an optional directory (the soak's) and now enforces committed floors of its own,
  RandomApiMinCases: random_api 8,300, host 15,500, mvn 700, LP64 120 (generated: 8,485 / 15,779 / 730 / 206) - a
  truncated or partial regeneration fails the replay instead of passing on fewer cases. The soak corpora are held to
  the same floors; the LP64 floor leaves more room because how many cases need Linux's 64-bit-long answer varies a
  little with the seeds. (Kept in this file rather than the shared MinCases table other sessions edit.)

Dry run (this machine, 10 fresh seeds 926935571, 1468300824, 1968923127, 2214447923, 2274677550, 3076328221,
3260894170, 3702504375, 3941943068, 4114009221; LP64 rows from WSL fed through --lp64-rows exactly as CI feeds the Linux
job's): 25,225 cases (8,510 / 15,778 / 730 / 207), RandomApiSoak green on net10.0 and net8.0. Full Oracle suite green on
both frameworks at d34ff6c0 + this change (228 passed; the soak test Inconclusive without its variable). The LP64
--lp64-only mode was also run from a clean output path under WSL (directory created on demand).

Correction to d34ff6c0's title: the P5 corpus delta was +157 cases (random_api +86, host +46, mvn +25), not +132.

Remaining (plan §8): P7 docs (Fuzz README, CLAUDE.md's differential-fuzz section, the compliance page) and the final
verification.
…iance page describe the oracle; stale random carve claims corrected; plan complete

P7 of docs/plans/random-oracle-coverage.md — the last phase. The oracle itself landed in 879791da (P0-P3), 844580a4
(P4), d34ff6c0 (P5 gates) and 780fe625 (P6 soak + floors); this commit documents it and closes the plan.

test/NumSharp.Tests.Oracle/Fuzz/README.md
-----------------------------------------
- New section "The random-API oracle (`random_api` tiers)": what it covers (439 members of the random world,
  G1-pinned inventory), signature-addressed cases (reflection-invoked, omitted optionals = declared defaults, NumPy
  called without them), receivers and primings, observations (with the post-call state on every stateful case and
  entropy masks), the four tiers and their gating (portable 8,485 / host 15,779 / mvn 730 / LP64 206, floors), the
  gates G2-G6 incl. the parameter-name gate, the nightly soak, the intended divergences, and how to regenerate.
- The generator joins the file-tree listing.
- The `gen_random_parity` carve-list row was stale: since 2026-09-25 only multivariate_normal stays carved (byte-
  identical only through a LAPACK backend's gesdd); the other seven samplers were uncarved and their pins are ordinary
  tests (OpenBugs.Random.cs says so). The row now says that.

.claude/CLAUDE.md (differential-fuzz section)
---------------------------------------------
- The "Products & random streams" bullet carried the same stale claim ("seven public samplers plus gamma(shape<1)
  remain carved and [OpenBugs]-pinned"); corrected.
- New "Random-API oracle" bullet: generator, plan, inventory, signature-addressed replay, tiers, gates, nightly soak
  jobs and the regeneration order after a surface change (NUMSHARP_WRITE_RANDOM_SURFACE=1 on G1 first).

docs/website-src/docs/compliance.md (Random Number Generation)
--------------------------------------------------------------
- New bullet: the whole random surface is oracle-checked overload by overload (every legacy sampler on the legacy
  MT19937 and RandomState(engine) per engine, every Generator member on all five engines, 10 fixed seeds, the full
  post-call state compared, gates proving the coverage, nightly 10 fresh seeds, ~25,000 cases, the intended
  differences named).
- Two corrections the oracle's P4 cases proved: "Any ISeedSequence can seed a bit generator (SeedlessSeedSequence
  included)" was wrong — NumPy and NumSharp both REFUSE a seedless sequence as a seed ("seedless SeedSequences cannot
  generate state"); and default_rng(RandomState) wraps whatever engine the RandomState has, not only MT19937. Also:
  every default_rng overload's parameter is `seed`, and a null argument of any type is default_rng(None) (both fixed
  in 844580a4).

docs/plans/random-oracle-coverage.md
------------------------------------
P7 checked; the final state-log entry checks the acceptance criteria off one by one (1 every overload — G2; 2 every
parameter — G3; 3 every generator type — G4; 4 ten fixed seeds in the corpus — G5, ten fresh seeds nightly — P6;
5 the post-call state — G6; 6 NumPy's errors verbatim — every error case; 7 every gate a FuzzMatrix test); the member
count in criterion 1 updated (439 at P7, 402 when the plan was written).

Final verification (detached worktree at 780fe625; this commit changes documentation only)
------------------------------------------------------------------------------------------
- Full Oracle suite: 228 passed / 6 inconclusive by design on net10.0 and net8.0 (the soak test without its variable,
  host-pinned tiers on this host all ran).
- Full unit suite (TestCategory!=OpenBugs&TestCategory!=HighMemory): 17,519 passed on net10.0, 17,518 on net8.0.
- Determinism: `python test/oracle/gen_random_oracle.py` regenerates all four committed tiers and both committed JSON
  side files byte-identical.
…random_api_oracle_map.py)

The `random_api` key of coverage/oracle_map.json was built during P1-P5 of docs/plans/random-oracle-coverage.md by a
scratch script that never reached the repository, so the next session to change the random member set would have had
to rediscover its rules while the strict coverage join (coverage/oracle_evidence.py) failed on the new member. It is
now test/oracle/random_api_oracle_map.py, fully documented:

- reads every params.member of the committed random_api* tiers and the coverage catalog
  (coverage/generated/coverage.json, written by coverage/generate_coverage.py);
- writes an override exactly where the template `numpy.random.{value}` is not the whole answer: a template id that is
  not a catalog row (mapped to its real row through a hand-reviewed SPECIAL table, else to nothing), RandomState
  members additionally crediting the np.random module function of the same name, BitGenerator members invoked on an
  engine additionally crediting numpy.random.BitGenerator.<member>;
- rewrites only the `random_api` line (inserted between `r_` and `ravel_f` when absent) and validates the JSON.

Re-run on the current corpus it reproduces the committed entry value for value (154 members, 84 overrides, 9 mapping
to nothing); the only change to coverage/oracle_map.json is the note naming the script. The Fuzz README's
regeneration paragraph and the plan's §6.8 point at it.

Verified: coverage/generate_coverage.py (2,851 catalog rows, oracle-verified headline 494/560) and the 32 coverage
unit tests green on b236c701 + this change.
…and bind; integers/randint array and BigInteger bounds; np.random.<Class> factories; typed null seeds are None; oracle over 493 members

Task (verbatim): "/np-function Another pass and validating and confirming random apis are whole and fully supported
including edge cases". Plan and state: docs/plans/random-oracle-coverage.md (new §8 phase "Wholeness pass", §9 top
entry, §10 items 15-20).

WHY A SECOND PASS WAS NEEDED
===========================
The random-API oracle (879791da..3d4392e6) replays every overload of the random world by EXACT C# signature, invoked
through reflection. That proves what each overload DOES; it cannot prove that NumPy's call, spelled verbatim in C#,
COMPILES and BINDS that overload. So a scratchpad harness (not committed: scratchpad/v2 harness.py, callforms.py,
edges.py, edges_int.py, lp64_compare.py) wrote NumPy spellings in C#, compiled each case in isolation (a compile error
is recorded per case, not per file), ran it, and compared the observation (kind, itemsize, shape, float64 hex bits)
with NumPy 2.4.2's:
  - 960 call forms: for every member, the required arguments only, each keyword alone, all keywords, positional
    prefixes, NumPy's defaults spelled out, and the None forms;
  - 579 edge values (NaN/inf/huge/zero/negative parameters, degenerate shapes, axis errors, big populations);
  - 254 array- and big-bound integers/randint cases.
1,793 spellings in all; the 721 legacy ones were also re-judged against Linux NumPy (WSL ~/np242/bin/python), because
NumSharp models the LP64 C long.

FINDINGS, ALL FIXED (plan §10 items 15-20)
==========================================
15. NumPy's size-only call did not compile for nine legacy samplers. normal, uniform, exponential, poisson, gumbel,
    laplace, logistic, lognormal and rayleigh default every parameter in NumPy (normal(loc=0.0, scale=1.0,
    size=None)), but NumSharp split each into a defaulted one-draw overload WITHOUT size and a size overload WITHOUT
    defaults, so `normal(size: 3)` and `uniform(high: 5.0, size: 2)` matched neither. The size overload now carries
    NumPy's defaults and the one-draw overload's parameters are REQUIRED (each file's remark says why), so the two never
    compete; normal() and normal(0.0, 1.0) still draw one 0-d value.
    Files: np.random.{exponential,gumbel,laplace,logistic,lognormal,poisson,rayleigh,uniform,randn}.cs.
16. numpy.random's classes were not reachable through the module. np.random.PCG64(42),
    np.random.Generator(np.random.PCG64(seed)) and np.random.SeedSequence(42) — NumPy's idiomatic spellings — did not
    compile: NumSharp's np.random is a NumPyRandom INSTANCE and only `new PCG64(42)` existed. New
    RandomSampling/np.random.classes.cs: 50 NumPyRandom factory methods mirroring EVERY public constructor overload of
    MT19937, PCG64, PCG64DXSM, Philox, SFC64 (each (), (long), (ulong), (BigInteger), (int[]), (long[]), (uint[]),
    (ISeedSequence); Philox also (object seed, object counter, object key)), SeedSequence ((), the integer/array forms,
    (object entropy, object spawn_key, long pool_size, uint n_children_spawned)) and Generator(BitGenerator). Same
    parameters, defaults and overload priorities as the constructors, so a call binds the constructor `new X(...)`
    would pick. No BitGenerator factory: the class is abstract and NumPy refuses too ("BitGenerator is a base class and
    cannot be instantized").
    TRAP: inside NumPyRandom these methods SHADOW the class names in expression position and in crefs. Static member
    access is now spelled global::NumSharp.Generator.Log1p / .ScalarPopulation / .ValidateChoiceProbabilities,
    global::NumSharp.MT19937.LegacySeeded / .ValidateLegacyArray, and crefs cref="NumSharp.X" (np.random.cs,
    np.random.choice.cs, np.random.default_rng.cs, np.random.rayleigh.cs, NumPyRandom.Broadcast.cs,
    NumPyRandom.LegacyDistributions.cs). Type positions (new PCG64(...), `is MT19937`, declarations) are unaffected.
17. A bare `null` seed — NumPy's None — was ambiguous (CS0121) between the int[]/long[]/uint[]/ISeedSequence
    overloads (and BitGenerator for RandomState) of every engine constructor, SeedSequence, default_rng,
    RandomState(...), seed(...) and MT19937._legacy_seeding(...): PCG64(None), default_rng(None), seed(None) had no C#
    spelling. [OverloadResolutionPriority(1)] now names the overload that takes it: the engines' (ISeedSequence seed)
    constructors, SeedSequence(uint[]), default_rng(ISeedSequence), RandomState(BitGenerator seed), seed(uint[]),
    MT19937._legacy_seeding(uint[]), and the matching factories. New Assembly/OverloadResolutionPriorityAttribute.cs
    polyfills the attribute for net8.0 (#if !NET9_0_OR_GREATER; net9.0+ uses the BCL type). The C# 13 compiler
    recognizes it by full name on the member's metadata, so it is honored ACROSS assemblies — verified by the net8.0
    test build compiling `new PCG64(null)` and friends.
18. A typed null seed ARRAY was the empty list, not None. new SeedSequence((int[])null) — and through it every engine's
    and default_rng's array overloads — seeded SeedSequence([]), the SAME stream on every run, where a null means
    Python's None (fresh OS entropy) everywhere else. The oracle's entropy cases mask the state words on BOTH sides, so
    a deterministic "entropy" passed them. SeedSequence(int[]/long[]/uint[]) now forward to the object constructor
    (: this((object)entropy)), which reads null as None. The pin asserts two constructions differ.
19. `size: default` bound the wrong overload. It is the C# spelling of NumPy's explicit size=None (`size: null` cannot
    convert to the Shape struct). The int[]/long[]/long size shims of dirichlet, hypergeometric, logseries,
    multinomial, multivariate_normal, pareto and power made it ambiguous, and for multivariate_normal it bound the
    `long` shim: size = 0, an EMPTY (0, 2) result where NumPy draws one vector — silently.
    [OverloadResolutionPriority(-1)] on every shim (each with a remark); callers that pass a shim still bind it.
20. integers/randint lacked NumPy's array bounds and Python-int bounds: g.integers(lows, highs),
    rs.randint(0, highs, dtype=np.uint8) and the full-range idiom g.integers(0, 2**64, dtype=np.uint64) had no overload.
    New:
      Generator.integers(NDArray low, NDArray high = null, Shape size = default, DType dtype = null, bool endpoint = false)
      Generator.integers(BigInteger low, BigInteger? high = null, Shape size = default, DType dtype = null, bool endpoint = false)
      NumPyRandom.randint(NDArray low, NDArray high = null, Shape size = default, DType dtype = null)
      NumPyRandom.randint(BigInteger low, BigInteger? high = null, Shape size = default, DType dtype = null)
    New RandomSampling/BoundedIntegers.Broadcast.cs (BoundedIntegers is now a partial class): a port of
    numpy/random/_bounded_integers.pyx.in's _rand_<dtype> for array bounds —
      - ResolveDtype first; a zero-size request returns before the bounds are read (NumPy's order);
      - the one-argument swap (high=None: [0, low)); two 0-d bounds take the SCALAR path through Python's int() (so a
        0-d None/complex/NaN/inf raises Python's TypeError/ValueError/OverflowError text);
      - the small dtypes (<= 32-bit, bool): low/high/ordering checks exactly where NumPy makes them (each skipped when
        np.can_cast already guarantees it), NumPy's messages ('low is out of bounds for uint8', 'high <= 0',
        'low >= high', and the '<' / '>' / '>=' not supported ... 'NoneType' TypeErrors), then FORCECAST to the up type;
      - the 64-bit dtypes: a safe cast, or NumPy's element-by-element Python-int() conversion with range checks — whose
        output is a layout-KEEPING empty_like, so 64-bit float bounds in a non-C layout come out SCRAMBLED exactly as in
        NumPy (KeepOrderMemoryIndex reproduces K order: C-contig -> none, F -> reversed, else stable |stride| sort);
      - the multi-iterator mismatch text with the output as arg 2; NumPy's quirk that a size SMALLER than the bounds'
        broadcast takes the first positions (no error);
      - the draw loop under the bit generator's lock, one bounded draw per position in C order, the 8/16/32-bit buffer
        (buf, bcnt) carried ACROSS positions as random_buffered_bounded_* do; the legacy randint on the masked sampler,
        the Generator on Lemire's; views and temporaries disposed in finally, the output on failure.
    The BigInteger overloads clamp to +/-2**100 (ClampToInt128) and take the scalar path, so NumPy's out-of-bounds texts
    appear beyond any dtype ('high is out of bounds for int64').

THE ORACLE GREW WITH THE SURFACE
================================
- test/oracle/random_surface.json (G1 inventory): 439 -> 493 members (NumPyRandom 188 -> 240: +50 factories, +2 randint;
  Generator 86 -> 88: +2 integers). test/oracle/random_numpy_signatures.json: 123 -> 130 (the seven module classes).
- test/oracle/gen_random_oracle.py:
    - emit_engine_ctor_overload(...) factored out of make_engine_ctor_family, plus fam_philox_ctor3,
      emit_seedseq_ctor_overload and emit_gen_ctor_overload, so the constructor families and the new
      make_module_class_family (every factory overload on the constructors' arguments, receiver RandomState, the
      receiver's state recorded to prove it is never drawn from) share one emitter; NP_SPECIAL entries for the factories;
    - integer_array_variants(gen): per-dtype bounds at, past and inside each dtype's range (an A() helper picks
      int64/uint64/float64 so each bound is representable; exact-float over/under constants 2**64+4096, -2**63-2048),
      negative, many/mixed/single/scalar-path forms, NaN/inf/None, endpoint forms (Generator), bool bounds (legacy),
      every layout through `derived`, the size-smaller-than-broadcast quirk; integer_big_extra() for BigInteger bounds;
      fam_gen_integers / fam_randint dispatch on the first parameter's type (NDArray -> the array sweep, BigInteger ->
      the scalar variants plus 2**63 / 2**64-1 and the big extras);
    - merge_lp64, two rules the array path needed: a PORTABLE-tier value difference between Windows and Linux NumPy now
      counts as LP64-dependent (no libm there, so the difference IS the C long: float bounds in a non-C layout take the
      64-bit element-wise conversion on Linux only — values in range, no error), and a Windows row flagged LP64 is
      DROPPED when Linux NumPy raised at an earlier receiver of the same (signature, variant) (validation precedes the
      draw and depends on neither engine nor seed; e.g. randint(low=[nan, 1.0], high=5): Windows' 32-bit path draws,
      Linux's 64-bit path raises int(nan)).
- Corpus: random_api.jsonl 8,485 -> 9,820 (+1,335), random_api_host.jsonl 15,779 -> 15,795, random_api_mvn.jsonl 730
  (byte-identical), random_api_lp64.jsonl 206 -> 234. 435 -> 489 distinct signatures (+ 4 reasoned exemptions = 493).
- FuzzCorpusTests.RandomApi.RandomApiMinCases: the portable floor 8,300 -> 9,600 (the old one would have let the whole
  ~1,300-case new family vanish); host/mvn/LP64 floors unchanged. A 10-fresh-seed soak generation (seeds 1057881256,
  4042940814, 2559081809, 1006913854, 2147695899, 4266922423, 3234014135, 1945494982, 4021612840, 3891615676) gave
  9,861 / 15,797 / 730 / 231 — every floor clears — and RandomApiSoak replays it green on net10.0 and net8.0.

GATES
=====
- RandomApiCoverageTests.ReceiverIndependent: the seven factories join RandomState and default_rng (a factory builds
  from its arguments, never from the receiver) for G4/G5.
- OracleSurfaceCoverageTests.EveryPublicNumpySurface_IsCoveredOrExplicitlyClassified: a new last route — a NumPyRandom
  method the random-API oracle replays by exact signature (RandomApiReplayedMethods, read off the shared CorpusSurvey's
  ParamSignature: the raw sig="NumPyRandom.<name>(" pair) is covered. The factories are the first members that need
  it (no stream draw exercises them); the older classifications keep precedence and meaning. Its summary line now
  prints random_api_methods (65).
- LeakSurfaceCoverageTests: np.random.<Class> factories are credited through random_api:<Class>, the key the tier
  records both the constructor cases and the factory cases under (both replayed by the leak sweep; G2 proves each
  factory overload has cases of its own). Leak surface 1,521 / 1,521 (7 via this route).

TESTS
=====
New test/NumSharp.Tests/RandomSampling/RandomApiWholeness.Test.cs (11 tests; every one a COMPILE-TIME proof of the
NumPy spelling first, then NumPy 2.4.2's values; legacy integers against the LP64 Linux build):
LegacyAllDefaultSamplers_TakeASizeOnly, LegacyAllDefaultSamplers_TakeAnyKeywordSubset,
ModuleClasses_AreFactoriesOfTheConstructors, NullLiteralSeeds_BindAndDrawFreshEntropy,
TypedNullSeedArrays_DrawFreshEntropy_NotTheEmptyList, SizeDefault_BindsTheNumPyShapedOverload,
Integers_ArrayBounds_MatchNumPy, Integers_ArrayBounds_ReproduceNumPysBroadcastQuirks,
Randint_ArrayBounds_MatchLinuxNumPy, Integers_ArrayBounds_RaiseNumPysErrors, Integers_BigIntegerBounds_MatchNumPy.

DOCS
====
docs/plans/random-oracle-coverage.md (§1 count, §2 post-P7 inventory note, §7 null-mapping row, §8 phase, §9 entry,
§10 items 15-20); test/NumSharp.Tests.Oracle/Fuzz/README.md ("Call shapes are not the replay's job", counts, the LP64
move/drop rules); .claude/CLAUDE.md (the random-API bullet's counts + the call-shape note; a Random-section paragraph
on the verbatim spellings, the shadowing trap, size: default, and the new bound forms); docs/website-src/docs/
compliance.md (np.random.<Class> factories, null literal/typed-null seeds, integers/randint array + BigInteger bounds,
the oracle bullet's counts and the call-shape pass).

VERIFICATION (isolated worktree on HEAD 3d4392e6 + this pass's files only; parallel sessions' WIP excluded)
===========================================================================================================
- Full unit suite (TestCategory!=OpenBugs&TestCategory!=HighMemory): net10.0 17,751 passed / net8.0 17,750, 0 failed.
- Full Oracle suite: 228 passed / 6 skipped (the unstaged bundled-OpenBLAS packaging tests + the soak) on BOTH
  frameworks, with NUMSHARP_OPENBLAS_LIBRARY = numpy.libs' scipy-openblas (the mvn tier byte-exact).
- Linux .NET (WSL, glibc, net10.0): the portable and LP64 tiers + G2-G6 + the name gate green (host/mvn/soak skipped by
  design there).
- Fresh-seed soak (above) replayed green on net10.0 and net8.0.
- Coverage: python coverage/generate_coverage.py passes its STRICT oracle join (300,350 contracts; headline
  oracle-verified 494/560). The seven numpy.random.<Class> rows moved from availability "type" (credited only because
  a same-named C# type exists) to "exact" (callable through the module), each oracle-verified "direct" through
  random_api:<Class>; BitGenerator stays "type". coverage tests 32/32, node dashboard tests 15/15,
  audit_documentation.py (2,640 links), test/inventory/generate_test_inventory.py.
- Documentation audit of every new/added C# declaration (scratchpad doccheck.py/doccheck2.py): /// present, every
  <param>, <returns> on every non-void member, <typeparam> on generics — 104 declarations inspected, 0 missing.

RESIDUAL HARNESS DIFFERENCES — ALL EXPLAINED, NONE A NUMSHARP DEFECT
====================================================================
vs Windows NumPy: callforms 113 = 79 `size: null` forms (CS0037: Shape is a struct; `size: default` compiles and
matches) + 22 multivariate_normal (harness ran without the OpenBLAS backend; the random_api_mvn tier covers it
byte-exactly with it) + 6 tomaxint (C long; match Linux NumPy) + 2 seed() entropy + 4 random_raw() (NumPy returns a
Python int, which the harness labels i8 below 2**63; value-identical 0-d uint64). vs Linux NumPy (legacy): callforms 54
= 36 of the same size: null forms + 12 mvn + 2 entropy + 4 vonmises (glibc libm, 1 ULP; NumSharp matches Windows'
ucrt). edges 18 vs Windows / 5 vs Linux: the harness spelled NumPy's np.nan as C#'s double.NaN (the NEGATIVE NaN —
NumPy fed -NaN returns 0xfff8... too, probed), vonmises(inf) under glibc, NumPy's documented zipf(a >= 1025) hang,
choice(2**40, replace=False) (NumPy hangs under Linux overcommit / MemoryError on Windows; NumSharp
OutOfMemoryException — the library-wide MemoryError text), the house AxisError suffix, dirichlet with a NESTED LIST
(NumPy's ndarray alpha gives "object too deep for desired array", NumSharp's text), and a harness spelling of
np.empty's order. edges_int 19 vs Windows (all legacy C long) / 0 vs Linux.

OUT OF SCOPE, REPORTED (not changed here)
=========================================
np.empty: NumPy's positional np.empty((3, 4), np.int64, 'F') has no C# spelling — NumSharp's overloads are
(Shape, DType, string device) and (Shape, char order, DType); the char 'F' does not compile in third position and the
string "F" binds `device` and throws "Device not understood". `order:` by name works.
…o mask smear for Lemire, uint8 Lemire threshold table (same stream); measured shortfall on narrow widths documented

Follow-up to 1a367afe (the wholeness pass that added integers/randint(NDArray/BigInteger ...)). The np-function bar is
>= 1.5x NumPy on every variation, and the new array path had not been benchmarked. Everything below was measured
P-core-pinned (affinity 0xFFFF on the i9-13900K: never an E-core, and more than one CPU so the background tier-up JIT
is not starved), every case warmed 40 calls past the 30-call tier-up, best-of-21 rounds, the NumPy 2.4.2 twin pinned
the same way through psutil, NumSharp BASE (a detached worktree at 1a367afe) and NEW alternated in fresh processes.
Scratch harness: scratchpad/v2 bench_pin.cs.tmpl / bench_pin.py / bench_ab.sh.

WHAT CHANGED (all stream-identical — verified by the oracle replay below)
- BoundedIntegers.Broadcast.cs DrawLoop: the NDIter chunk's pointers and BYTE strides are read once per chunk into
  locals. Read through the BroadcastWalk object inside the loop they were reloaded for every position, because the
  virtual NextUInt32/NextUInt64 calls may, as far as the JIT can tell, write the object's fields.
- DrawOne's 64-bit case passes the GenMask smear only to the masked (legacy) sampler; Lemire ignores the mask (NumPy's
  own Lemire callers pass 0), and GenMask is a call per position.
- BoundedIntegers.cs LemireUInt8: NumPy's `(UINT8_MAX - rng) % rng_excl` is read from LemireThreshold8, a 256-entry
  table built once from that same expression in the same byte arithmetic. Identical threshold for every reachable rng
  (entry 255 is never read: the full range returns the raw byte before Lemire), so identical accept/reject decisions and
  an identical stream. The "rare" branch holding the division runs for rng_excl/256 of the draws — 78% at rng 199.

MEASURED (NPY/NS, higher = NumSharp faster; base -> new)
  case                          1K            100K          1M
  gen int64 arr,arr             0.94 -> 1.02  1.81 -> 2.09  1.68 -> 1.83
  gen uint8 0,arr varying       1.30 -> 1.48  1.11 -> 1.30  1.11 -> 1.29
  gen uint8 0,arr const 200     1.36 -> 1.55  1.12 -> 1.35  1.11 -> 1.24
  gen int32 arr,arr             1.60 -> 1.71  1.31 -> 1.32  1.15 -> 1.25
  gen int32 0,arr (<= 1e5)      1.83 -> 1.90  1.44 -> 1.50  1.36 -> 1.50
  gen float64 bounds, 100       4.94 -> 5.38  5.57 -> 6.23  5.24 -> 5.61
  gen column broadcast          2.48 -> 2.72  2.41 -> 2.84  2.56 -> 2.90
  rs  int64 arr,arr             2.60 -> 2.41  1.45 -> 1.30  1.44 -> 1.42   (masked; within this host's spread)
  rs  uint8 0,arr varying       1.29 -> 1.39  1.02 -> 1.11  1.01 -> 1.10
  rs  uint8 0,arr const 200     1.30 -> 1.38  1.02 -> 1.09  0.99 -> 1.10
  rs  int32 0,arr (<= 1e5)      1.50 -> 1.60  1.18 -> 1.26  1.15 -> 1.28
  gen uint8 scalar 0,200 (PRE-EXISTING path, the table's side effect)
                                1.64 -> 1.91  1.00 -> 1.19  1.00 -> 1.21

WHAT DID NOT WORK (measured, reverted — recorded so nobody repeats it)
- AggressiveInlining on DrawOne alone: neutral within noise.
- The per-position loop extracted into its own method (outside DrawLoop's three EH regions: try/catch/finally, lock,
  using) + DrawOne inlined: uint8 up (gen const 1M 1.21 -> 1.46x) but int64 DOWN (1.92 -> 1.68x), float-bound and
  column-broadcast cells down too.
- AggressiveInlining on the by-ref sub-word helpers (BufferedUInt8/16, LemireUInt8/16) + GenMask: worse everywhere
  that mattered (int32 array 4.75 -> 5.53 ms, legacy uint8 6.33 -> 6.98 ms).
The Tier1 listing of the extracted loop (DOTNET_JitDisasm) shows why: PGO inlines EVERY width's path plus the
devirtualized PCG64 step (guarded devirtualization of the virtual NextUInt32) into one body — a 248-byte frame with
buf/bcnt and the engine state spilled to the stack. With one loop serving five widths, every inlining decision helps one
width's register allocation and hurts another's.

STILL BELOW THE 1.5x BAR (faster than NumPy on every cell; minimum 1.02x, int64 at 1K where fixed cost dominates)
uint8 (gen 1.24-1.55x, legacy 1.09-1.39x), int32 at 100K/1M (1.25-1.50x), legacy int64 at 100K/1M (~1.3-1.4x), and the
pre-existing scalar uint8 fills (1.02-1.21x at >= 100K). The known fix is per-width chunk loops — the scalar path's own
`Fill` -> `FillUInt8/16/32/64/Bool` shape and NumPy's per-dtype `_rand_<dtype>_broadcast` — and, for 32/64-bit, bulk word
draws per chunk as `Bounded32Chunk` does. Deferred: the np-function rule forbids per-dtype functions / per-dtype switch
dispatch, so it needs the user's call. Recorded in .claude/CLAUDE.md (Random, item 5).

VERIFICATION (isolated worktree = HEAD 1a367afe + these edits)
- Full Oracle suite: 228 passed / 6 skipped (unstaged bundled-OpenBLAS packaging tests + the soak) on net10.0 AND
  net8.0 — the random_api / random_parity / generator_parity tiers replay integers/randint at every dtype, engine and
  seed byte-exact, the uint8/int8 fills the threshold table touches included.
- RandomSampling unit namespace: 1,445 / 1,445 on net10.0 and net8.0.
…routes, mapdomain affine kernel, getdomain blocked min/max pass, AVX2 int->float casts (int64/uint64 exponent splice), SIMD strided word copy; polyseries oracle 15,950 -> 18,437 cases, 0 excused

Audit of the numpy.polynomial additive family + polyutils committed in c5ef9461 ({p}add/sub/trim/line for the
six bases, the 24 module constants, polyutils as_series/trimseq/trimcoef/getdomain/mapparms/mapdomain) for full
NumPy 2.4.2 parity and >= 1.5x NumPy speed on every variation.

WHY: the U1 corpus never went past a dozen coefficients, so every oracle cell ran the scalar IL kernels. A long
series (>= 64 elements) is where speed matters and where every vector route below lives, so parity of those routes
had no gate, and the scalar routes were slower than NumPy on long converted/strided series.

== Parity: corpus (test/oracle/gen_oracle.py polyseries -> polyseries.jsonl, 18,437 cases, 0 excused) ==

- (K) long series, 1,458 cases: add/sub/legsub across the 64-element house threshold (63/64/65, 70/64, 64/70,
  71/3, 257/100 ... 1,000/1,000, 1,003/5) in float64/float32/float16/complex128, twelve mixed dtype pairs,
  contiguous / stride-2 / reversed operand pairs, trailing zeros trimmed back below the threshold, a cancelling
  tail, specials (NaN, +-inf, +-0, subnormals, overflowing sums) in the updated prefix AND the untouched tail;
  as_series / trimcoef / getdomain / mapdomain at 100 and 1,000 points over every dtype and layout.
- (L) complex64 ARRAY loops of mapdomain, 552 cases: a complex64 off/scl (mapparms of a float16/float32 domain and
  a Python complex tuple) with every narrow x, a Python complex tuple with a float16/float32 x; lengths 1, 5, 37
  and strided/reversed. NumSharp has no complex64 dtype (#569): the loops run on float32 kernels with NumPy's
  product forms (simd_cmul at every length for a trivially iterable call - even ONE element, probed), the result
  is carried as complex128 holding the float32-exact values, and the replay VALUE-compares the up-cast pairs
  (Complex64ValuesMatch). 752 complex64-result cases in the tier; 109 cells still skipped (complex64 arrays
  outside mapdomain, object arrays).
- (M) block and staging boundaries, 45 cases (appended after L so every earlier case keeps its counter): the
  1,024-element blocks of the in-place combine (int32+float64 both ways, float64 negated, float32 stride-2, int16
  reversed, float16+float32 at 2,100/2,150), as_series' blocked strided cast, mapdomain's affine route over whole
  blocks + a partial last one (contiguous 8/16/64-bit ints at 2,100 into float64, int8/uint8 at 1,030 into float32,
  float16 at 600; int64 reversed, int16/uint8/int32/float32 stride-2 at 2,100), and the fused pass's staging of a
  transposed 2-D x (96x90 = 8,640 points, past the 8,192 threshold).
- (N) getdomain across the blocked pass's windows, 72 cases (appended after M): every dtype the pass serves crosses an
  8 KB window boundary (window + 37 elements) contiguous / stride-2 / reversed, plus stride 3 for int64 / float64 /
  uint8 (AVX2 gather, scalar sub-word copy); and the schedule facts the carried accumulators must keep, at float64 and
  float32 over two windows with no scalar tail: a -0.0 early in the HIGHEST lane against a +0.0 late in lane 0 (NumPy
  returns the EARLY zero - its horizontal cascade lets the higher lane win the tie, where a sequential fold returns
  the late one; mirrored for min), and one NaN with a payload that comes back canonical from a window boundary inside
  the vector section and with its payload from the scalar tail.
- Every previous line is preserved byte for byte (verified: all 18,365 earlier cases identical); the generator is
  deterministic (re-run = same sha).
- MisalignedRegistry: a zero-excuse guard for the polynomial ops, so a drift fails instead of being excused.
- FuzzCorpusTests floor for polyseries.jsonl 15,000 -> 18,000, so a regeneration that silently lost K-N fails.

== Routes (all byte-identical to the scalar kernels; NDPolySeries.cs) ==

- {p}add/{p}sub: CombineViaHouseKernels - NumPy's own steps (converted copy of the in-place target, c2 = -c2 when
  the subtrahend is at least as long, c[:nb] OP= other) run by house SIMD kernels, block by block (1,024 elements)
  through L1 scratch, when PreferHouseCombine/VectorReadable say every step is a vector loop (contiguous or
  sub-word/<=4-byte strided operands with a vectorized cast; float16 always - the scalar kernel emulates it).
- as_series' conversion copy: CopyInto - memcpy / house contiguous cast / SIMD strided copy (+ cast in L1 blocks).
- getdomain: ONE blocked pass for every layout of a float64 / float32 / integer x (TryDomainBlocked +
  MinMaxBlocked<T>). NumPy reduces a fresh C-contiguous copy with simd_reduce_c; the pass takes that copy 8 KB at a
  time - read in place from a unit-stride x of the reduced dtype, packed into an L1 scratch by CopyInto otherwise -
  and folds BOTH the min and the max while each window is in L1, with the engine's own exact schedule
  (NumPyMinMaxReduce.FoldGroups for full windows, Finish for the last). Bit-identical to one call over the whole copy
  by the argument DefaultEngine.StreamMinMax rests on: windows hold whole 8-vector groups counted after element 0, and
  per lane NumPy's group tree + single-vector loop is one ordered fold of an associative op, which the carried
  accumulator continues unchanged. An integer x is order-free (no +-0, no NaN), so a reversed one is handed over as
  the forward run it covers (new PolySeriesView.Reversed) and folded in place; a float x keeps its order. uint64
  packs CONVERTED to float64 (the house exponent-splice cast) and folds as float64 - what NumPy does, exact by
  monotonicity - because AVX2 has no unsigned 64-bit compare: the emulated uint64 folds cost 20.4 us per 100K on an
  L1 block where the cast + float64 folds cost ~14; int64 stays int64 (NumPy's tree fold, 13.9 us, beat a fused
  one-read fold 18.3 and a 3n/2 pair fold 16.0 in a probe). float16 / char decline (no exact schedule) to the
  copy-then-reduce route, which also serves when NDExpr.DisableExactMinMax is set. It replaced allocating the copy and
  running two engine reductions over it: 100K stride-2 float64 / int64 / uint64 ~37 / ~45 / ~50 -> ~23 / ~30 / ~31 us;
  16 points ~0.9 -> ~0.3 us (the engine reductions' fixed cost); 100K contiguous float64 / int32 11.2 / 6.2 -> 9.4 /
  4.8 us, same process (one pass over memory instead of two). Every one of 187 route comparisons (11 dtypes x 16 /
  1,000 / 2,100 / 100K / 1M x 4 layouts, same process, each route a whole warm batch) returned identical bytes; on
  the 170 served cells the blocked pass was never slower (1.00-4.9x; float16's 17 decline). Test hooks NDPolySeries.DisableBlockedDomain / BlockedDomainRuns.
- mapdomain: MapDomainAffine for a float64/float32/complex128 loop shared by both operations over 1-D (any stride)
  or C-contiguous points - converted 8 KB at a time into an L1 scratch (CopyInto), then PolyAffineDouble /
  PolyAffineSingle (multiply THEN add, two roundings - never an FMA; pinned by a test whose inputs separate the two)
  or PolyComplex128Affine. FusedPassIsFaster keeps the fused np.evaluate pass for a contiguous x on a vector widen
  edge or of bool, and for a strided x already of the loop dtype or on an edge from 8,192 points (where that pass
  packs it). Measured in ONE process, interleaved, via the new test hook NDPolySeries.DisableAffineMapDomain:
  fused/affine geomean 1.34 at 16 points, 1.53-1.62 at 1,000, 1.36-1.46 at 100K over 249 shapes, worst cell 0.94.
  complex64 loops: MapDomainComplex64 on the float32 kernels (PolyComplex64ScaleReal/AffineReal/AddScalarReal).
  The remaining fused shapes stage their points with ONE direct SIMD cast (not astype, whose setup was most of a
  1,000-point call) above measured break-evens (StagedContiguousMinPoints: 512 float16 / 1,024 float32 loop /
  4,096 float64 loop).
- [SkipLocalsInit] on CopyInto / CombineViaHouseKernels / MapDomainAffine: their stack scratch is always written
  before it is read, and zeroing 8 KB per BLOCK cost more than converting the block (mapdomain at 16 points
  1.03 -> 1.34x the fused pass once removed).

== New house kernels (DirectILKernelGenerator; library-wide through TryGetCastKernel, bit-exact) ==

- Cast.IntToFloat.cs: AVX2 {int8, uint8, int16, uint16, char, int32, uint32} -> {float32, float64} (pmovsx/pmovzx
  from memory + cvtdq2pd/ps; uint32 via the sign-bias trick; uint32 -> float32 through the exact float64, ONE
  rounding) and int64/uint64 -> float64 via the exponent-bias splice (high part in the mantissa of 2^84 / 3*2^67,
  low part in 2^52's, biases subtracted exactly, ONE final RNE add) - 2.2-3.3x the scalar loop; 134M random values
  (every magnitude, ties at 2^53..2^64) identical to cvtsi2sd / Converts.ToDouble(ulong); NumPy-probed pins.
  The generic emitter's int->float strategies had measured at SCALAR speed (0.19-0.37 ns/element on an L1 block).
- Cast.WordCopy.cs: 4/8-byte same-type (or same-size int reinterpret) strided copy - SIMD reverse (stride -1),
  deinterleave (stride 2), AVX2 gather otherwise, per-row memcpy for unit stride; bit copies only (a float<->int
  pair is a value cast and is refused).
- Cast.Half.cs: float16 -> float64 exactly as NumPy's ToDoubleBits (a signalling NaN stays signalling where
  cvtps2pd / (double)Half quiet it), contiguous (Giesen + cvtps2pd, inf/NaN groups rewritten) and strided.
- EmitConvertTo: float32 <-> float64 through packed cvtps2pd/cvtpd2ps - the scalar cvtss2sd/cvtsd2ss MERGE into
  their destination register, a false dependency that made a scalar loop latency-bound (128 vs 37 us).
- MixedType: IsVectorizedSameTypeBinary (which same-dtype binary kernels are vector loops, float16 included).
- PolySeries: long-series cast routing slots (GetPolyContiguousCast/GetPolyStridedCast), the complex64 kernels,
  PolyComplex128Affine, PolyAffineDouble/Single.

== Tests ==

- PolynomialSeriesTests (38): + GetDomain_BlockedPass_MatchesTheCopyRoute_BitForBit (10 served dtypes x 13 lengths
  across every width's window boundaries x contiguous / stride-2 / stride-3 / reversed / reversed-stride-2 x data built
  to expose a schedule slip - sign-mixed zeros, the lane-order zero pair, one NaN payload on the seed / a window
  boundary / inside / the tail, uint64 values that collide as float64; every call must take the pass; float16 must
  decline); GetDomain_BlockedPass_FoldsEveryElement (a UNIQUE min and max planted at every position around each window
  edge, the seed and the tail, every served dtype and layout - this is what caught a planted skip-8-per-window mutant
  that the route-agreement test's repeating extremes let through); GetDomain_BlockedPass_NumPyPins (NumPy-probed
  lane-order -0.0 / +0.0 and canonical-vs-payload NaN across windows, contiguous and stride-2);
  GetDomain_BlockedPass_DeclinesWhenExactSchedulesAreDisabled. Planted-bug checks: reversed floats treated as
  order-free -> red (float32 reversed -0.0 came back +0.0); 8,000-byte windows (a partial group dropped per window) ->
  red via FoldsEveryElement.
  + MapDomain_AffineRoute_MatchesTheFusedPass_BitForBit (13 dtypes x 5 lengths across
  the block boundaries x contiguous/stride-2/reversed/2-D/transposed x float64/float32/complex128 loops, 1,092
  calls, the run counter proves the route served >780 of them) and MapDomain_AffineRoute_RoundsMultiplyAndAdd-
  Separately (NumPy-probed 5.551115123125783e-17, plus 4,099 random points at both widths where FMA would differ);
  the complex64 scalar test.
- IntToFloatCastParityTests (11): every 8/16-bit value, NumPy-probed int32/uint32 -> float32 and int64/uint64 ->
  float64 pins (ties, 2^63+1025, 2^64-1, the 48-bit splice seam), sampled ranges, all 65,536 float16 patterns ->
  float64 in contiguous/reversed/stride-3 layouts (sNaN pin), resolver checks.
- WordCopyStridedParityTests (2): every inner stride (+-1, +-2, +-3, 7, broadcast 0) over 4/8-byte dtypes, raw
  bits (sNaN, negative qNaN payloads, -0.0, -inf); the resolver admits bit copies only.
- Randomized batches (scratchpad generator, NumPy 2.4.2): 380K cases replayed green (fz21-23, fz3, fz31-32,
  fz41-44). The oldest batch fz2 fails only on its 75 f_contiguous_2d as_series cases - a stale-generator artifact
  (that version serialized an F-ordered base in C order, fixed in the generator before every later batch, which
  carry ~150 such cases each and pass); fz2 without them passes (19,925 cases).

== Verification ==

- Unit: net10.0 17,790 passed / net8.0 17,789 passed, 0 failed (11 skipped each).
- Oracle: net10.0 233 / net8.0 233 passed (1 skipped each), 0 failed - re-run after the polyseries floor raise.
- Earlier in the audit: oracle tiers Poly / Polyeval / Polyseries / Astype_Full / Astype_Smoke / Conversion green on
  net10.0 and net8.0; the polyseries replay (18,437 cases) re-run green after section N.

== Perf (NPY/NS, P-cores pinned, best of 7, NumPy re-measured back to back, 867 cells) ==

dtypes x 16 / 1,000 / 100K (+1M) x contiguous / stride-2 / reversed, geomean 8.47x (the full run, plus a back-to-back
re-run of the 90 getdomain cells after the blocked pass):
- add/sub 264 cells 1.70-24.3x (geo 6.7); trim 174 cells 5.5-515x (geo 24.6); trimseq 84 cells 3.8-490x (geo 18.2)
- getdomain 90 cells 2.13-104x (geo 8.3) - was 1.26-75x (geo 4.7) before the blocked pass
- mapdomain 183 cells 1.96-33x (geo 4.9); as_series 42 cells 1.11-10.3x (geo 3.9); mapparms 9 cells 1.18-8.7x
- {p}line 21 cells 0.80-1.03x
26 cells stay under 1.5x, each a measured floor:
- {p}line (21): the NDArray allocation floor (~230 ns to create and dispose a 2-element array; NumPy's whole
  np.array([off, scl]) is 260-430 ns).
- as_series of two 100K arrays (4, 1.11-1.36x): memcpy - both sides copy the same 2 x 800 KB, and two plain
  Buffer.MemoryCopy calls into already-allocated buffers take 45.5 us of the call's 44.7 (float64).
- mapparms(float64 array, (0, 2)) (1, 1.18x, 0.35 vs 0.42 us): the exact general lane's dispatch - four element reads
  (~60 ns) and six PolyNumber.Binary operations (~25 ns each, a >100-byte struct through every step) for a mix of
  NumPy scalars and Python ints. The lever is a typed program per argument-kind signature, a refactor of the exact lane
  left for its own change.
Measurement notes: 100K cells move 1.5-2x between runs with the allocator's page state (NumPy's own getdomain went
53.8 -> 67.5 us between two runs), so route A/Bs ran in ONE process; and an A/B that alternates routes batch by batch
lets the copy route's allocations evict the other route's working set (the blocked getdomain read 39-43 us
interleaved, 30 us alone) - each route was timed as a whole warm batch and confirmed standalone.

== Docs ==

.claude/CLAUDE.md U1 section (long-series routes, new kernels, traps), docs/plans/numpy-polynomial.md (status line
+ U1 audit bullet), test/NumSharp.Tests.Oracle/Fuzz/README.md (polyseries sections K/L/M).
…e in-place IL kernel per basis; polycalc oracle 21,646 cases bit-exact, 606 benchmark cells min 2.02x geomean 8.71x NumPy

Plan unit U4 of docs/plans/numpy-polynomial.md: polyder/polyint, chebder/chebint, legder/legint, lagder/lagint,
hermder/hermint, hermeder/hermeint (12 names) under np.polynomial.{polynomial,chebyshev,legendre,laguerre,hermite,
hermite_e}, with every NumPy 2.4.2 parameter: m, k, lbnd, scl, axis. Bit-exact with NumPy 2.4.2 for every dtype
NumPy has (and decimal / char), every memory layout and every argument kind.

ENGINE
------
- src/NumSharp.Core/Polynomial/Package/NDPolyCalc.cs (new): the driver. It is NumPy's Python prologue in NumPy's
  statement order, which is also NumPy's error order: np.array(c, ndmin=1, copy=True), int/bool -> float64, k
  wrapping and padding, _as_int checks, "Too many integration constants", lbnd/scl must be scalars,
  normalize_axis_index, moveaxis, the per-order loop.
- src/NumSharp.Core/Backends/Kernels/Direct/DirectILKernelGenerator.PolyCalculus.cs (new): the kernel.
  * The six recurrences are DATA (PolyCalcRoutines: head steps, a j loop, tail steps; each statement is NumPy's
    source expression token for token, never re-associated).
  * ONE whole-array DirectILKernelGenerator kernel per PolyCalcKey(basis, der/int, dtype T, source dtype,
    ScalarMath (1-D), Scale).
  * Each stage is its own DynamicMethod (root, vector load, scalar load, in-place scale, recurrence), so the JIT's
    per-method inline budget covers one stage's lane helpers.
- One buffer, in place, for any order.
  * A derivative writes der[q] one row below c[q]; the result is rows [m, n).
  * An integral writes tmp[q] one row above c[q], with m spare rows above the loaded series.
  * Every value is still produced by NumPy's statement sequence.
- Columns (positions in c.shape[1:]) never meet, so the kernel takes them in ~512 KB blocks through EVERY order of
  a derivative.
  * The vector loop over columns uses the house lane kinds (U3's ILKernelGenerator.Polynomial.Lanes.cs), then a
    scalar tail.
  * The per-j weak Python ints are converted once per (block, j): exactly through double, then the house cast
    (float16: 2*(j+1) past 2048 rounds to even, past 65504 is inf, as in NumPy).
  * A complex divisor's Smith preparation is hoisted with them.
- The load stage converts int/bool series to float64 and applies c *= scl in the same pass. It reads the caller's
  layout in place when a row's columns sit at one stride (every 1-D and 2-D series, C-mergeable N-D); otherwise
  NDIter copies first. A scl the kernel cannot fold (an array scl on a derivative, or a strong 0-d scl that
  promotes the series) runs as NumPy's in-place ufunc per order (ScaleRoute.House).

NUMPY SEMANTICS REPRODUCED
--------------------------
- Three arithmetics, per statement.
  * c *= scl is always an array op: simd_cmul with operands (c, scl) — (inf+1j) *= 1 is inf+nanj.
  * A 1-D series' recurrence statements are NumPy SCALARMATH (the naive complex product): the ScalarMath key.
  * An N-D series' statements are ufuncs.
- The integral's tmp[0] += k[i] - {p}val(lbnd, tmp) runs through U3's NDPolyEval.Val and U1's PolyNumber.
  * For a 1-D series it is scalarmath, then setitem's cast: a complex value into a real series keeps its real
    part (NumPy's ComplexWarning).
  * For an N-D series it is an in-place ufunc on tmp[0]'s row: the same_kind cast check first, then NpyIter's
    broadcast rules.
- A 1-D float16/float32 series with a Python-complex lbnd makes NumPy evaluate {p}val in complex64 SCALARMATH.
  U3's evaluation kernel computes complex in complex128, so NDPolyCalc.PyVal1D interprets U3's weak-x step trees
  (PolySteps.RewriteForWeakX) with PolyNumber instead, which emulates complex64 scalars exactly.
- Result objects are NumPy's (the oracle records C_CONTIGUOUS / F_CONTIGUOUS / OWNDATA of each result):
  * every growing order returns an np.moveaxis VIEW of a fresh C-order buffer (OWNDATA false);
  * m == 0 returns the K-order copy (OWNDATA true; an F series stays F);
  * der with m >= n returns c[:1]*0, laid out by NpyIter's KEEPORDER vote (Shape.MultiSortedStridePerm) and moved
    back — except hermeder, which returns it UNMOVED (NumPy's quirk);
  * an integral whose every order took the n == 1 zero branch returns a view of the moved copy.
  NDPolyEval.LikeKeepOrder became internal for the K-order copy.
- Errors, verbatim, in NumPy's order.
  * A complex scl is caught at the first c *= scl, so it is reported before an empty series' IndexError.
  * An array constant for a 1-D series is setitem's ValueError ("setting an array element with a sequence.")
    into a REAL series, but complex()'s TypeError ("only 0-dimensional arrays can be converted to Python
    scalars") into a COMPLEX one — probed for every array shape and dtype.
  * For an N-D series, a complex correction is the in-place add's UFuncTypeError. It names complex64 for a
    float16/float32 series with a Python complex lbnd/k (InPlace's operandIsComplex64).

ORACLE
------
test/oracle/gen_oracle.py gen_polycalc -> test/NumSharp.Tests.Oracle/Fuzz/corpus/polycalc.jsonl: 21,646 cases,
floor 21,000, 0 excused, replayed by OpRegistry.PolySeries.cs.
- (A) every dtype x length 0-13 x order 0 .. n+1
- (B) weak and strong scl kinds, 1-D and N-D
- (C) array scl on derivatives
- (D) integration constants: scalars, lists, typed arrays, 0-d, N-D rows
- (E) lbnd kinds x 1-D / N-D x float dtypes
- (F) N-D series at every axis x C/F/transposed/strided/reversed layouts, values AND result flags; broadcast, 0-d
- (G) specials
- (H) full-mantissa complex, 1-D scalarmath vs N-D ufunc
- (I) long 1-D series, widths around every lane count, >4096-column blocks
- (J) argument errors in NumPy's order
- (K) Python-list series
- (L) the n == 1 zero branch, with flags
- float16 constants past 2048 and past 65504 (2,100 / 33,000 / 66,000 coefficients)
Char rides the uint16 proxy.

Planted-bug check: 10 mutants, 10 killed. The kernel mutants:
- M1: the 1-D scale as the naive product — 279 red cases
- M3: the derivative loop's last j — 2,234
- M4: the integral loop's first j — 7,981
- M6: the block width — 8, only the >4096-column cells see it
- M7: the strided load — 1,146
- M8: vector negation — 612
- M10: later-order scaling — 549
The driver mutants:
- M9: the correction's sign — 3,190
- M11: the complex64 error name — 60
- M12: the complex64 {p}val — 63
The coverage generator's strict join resolves all 12 new op keys as verified-direct from polycalc.jsonl.

TESTS
-----
test/NumSharp.Tests/Polynomial/PolynomialCalculusTests.cs (new, 20). It covers:
- NumPy's TestIntegral/TestDerivative ported, made BITWISE where NumPy's statement sequence makes them exact:
  multi-order integrals vs repeated single orders, per-lane axis results;
- NumPy byte dumps for float64/float32 (weak vs strong 0-d scl)/float16/complex128 (1-D vs N-D);
- int32/bool/char series (char = uint16 bytes);
- decimal in decimal arithmetic;
- every error text; the result layouts;
- vector lanes for every dtype (PolyCalcVectorFallbacks == 0);
- column blocks vs column slices; direct load vs house copy vs a contiguous copy;
- thread-safe first use.

Full suites: NumSharp.Tests 17,810 / 17,809 passed, 0 failed (net10.0 / net8.0). NumSharp.Tests.Oracle 234 passed,
1 skipped, 0 failed on both.

PERF (NPY/NS, higher is better)
-------------------------------
benchmark/polynomial/polycalc_{numpy.py,bench.cs,report.py} (new), committed summary polycalc_results.md. 606 cells,
every one SHA-256-checked bit-exact against NumPy before it was timed; both sides pinned to 0xF0; C# Release,
--no-cache, DOTNET_TC_CallCountingDelayMs=0, drained tier-1 warm-up.
  1-D series (n 4..1000, m 1/3)   min 3.69  geomean 25.30
  parameters (scl/k/lbnd kinds)   min 2.47  geomean  7.72
  N-D float64 (every axis, 2/3-D) min 2.02  geomean  4.76
  dtypes (N-D and 1-D)            min 2.14  geomean  7.97
  layouts                         min 2.91  geomean  7.26
  orders m = 1/2/5/10             min 3.10  geomean  6.61
  small N-D                       min 3.87  geomean 12.35
  ALL                             min 2.02  geomean  8.71

Open lever, measured, NOT done: float16 N-D runs ~30x slower than float32 (still 2.1-2.6x NumPy). U3's float16
lane kind rounds every op back onto the f16 grid (narrow + widen) and the store narrows again. A cheaper in-float
round buys only ~1.25x; a float32 working buffer narrowed once at the end is the real fix. Recorded in the plan.

HARNESS FIX
-----------
OpRegistry.PolySeries.cs CalcFacet: a "facet": "flags" case REPLACES the result with its [C, F, OWNDATA] flags,
so the registry must dispose the result it drops. Undisposed, the leak gate (UndisposedIntermediateTests) read one
escaped buffer per flags case: 96 families, all 57-case N-D flag cells.

TRAPS (also in CLAUDE.md "Polynomial package - calculus family (U4)")
--------------------------------------------------------------------
- describe() serializes an operand's BASE in C order. An F layout must be a C base of base.T plus a .T view; an
  F base broke 1,311 cases in the first oracle run.
- A quoted --filter passed from Git Bash through cmd //c reaches dotnet with LITERAL quotes. A TestCategory!=
  filter then matches everything and a ClassName~ filter nothing; read the filter from a file (set /p).
- complex setitem of an array raises TypeError, not ValueError.

DOCS
----
- .claude/CLAUDE.md: the U4 section and the Core Source Files row.
- docs/plans/numpy-polynomial.md: U4 as built; status 102 of 193 names.
- test/NumSharp.Tests.Oracle/Fuzz/README.md: the polycalc tier and the regenerate line.
- np.polynomial.cs: the PolynomialModule implemented-families list.
- The {p}der / {p}int facades: fully documented, per NumPy.
- benchmark/polynomial/.gitignore: the calc TSVs.
…ported (NDPolySequence), object-typed U3 arguments, trimseq/mapdomain list overloads, widened-scale calculus kernels; oracle polyeval 19,216 / polyseries 18,647 / polycalc 27,526 bit-exact

WHY
---
A second /np-function pass asked for every U1/U3/U4 API to be whole: every argument kind a NumPy caller passes,
every edge. Replaying the C# boundary kinds against NumPy 2.4.2 (three probes, ~330 calls) found the conversions
of Python SEQUENCES wrong or missing in several places, U3's facades refusing anything but NDArray, and the
calculus scale losing NumPy's wider multiply loop when a strong scalar promotes the series. This commit makes
every argument conversion of the package one port of np.array's coercion and closes each gap with oracle rows.

WHAT CHANGED
------------
1. np.array's coercion, ported (src/NumSharp.Core/Polynomial/Package/NDPolySequence.cs, new).
   Statement for statement array_coercion.c: PyArray_DiscoverDTypeAndShape_Recursive + update_shape + the fill.
   - The first leaf reached fixes the dims; a leaf at another depth or a sequence of another length is ragged:
     NumPy's verbatim ValueError "setting an array element with a sequence. The requested array has an
     inhomogeneous shape after N dimensions. The detected shape was (...) + inhomogeneous part.", N and the
     shape being what the walk still agreed on.
   - An EMPTY sequence ends the dims without voting on the dtype ([] is float64 (0,); [np.zeros(0, int8), []]
     is int8 (2, 0)); an array leaf that cannot be assigned raises the broadcast ValueError.
   - Leaf dtypes are each leaf's DEFAULT descriptor (Python int -> int64, uint64 up to 2^64-1; NumPy scalars
     and arrays keep theirs), promoted STRONGLY - array coercion is not NEP 50 ([np.float32(1), 1.0] is float64).
   - str / None / a Python int past uint64 / unknown objects are refused (NotSupportedException) AFTER the walk,
     so a ragged input still reports NumPy's ValueError.
   - MakeTuple builds a ValueTuple of any arity (8+ nest through TRest) for tuple-shaped results.
   - The C# map (header of NDPolyNumber.cs): ITuple is a Python tuple; object[] and every other C# array whose
     elements are not a NumSharp dtype (jagged double[][], NDArray[], BigInteger[]) and any other IList/
     IEnumerable are lists, an object[,] a list of rows; a typed dtype array and Memory<T> of a dtype are
     ndarrays; bool/integers/BigInteger/float/double/Complex Python scalars; Half/char/decimal NumPy scalars.
   Every array_like the package converts now goes through it: {p}der/{p}int's c, k, lbnd, scl; {p}val's x and
   c; mapdomain's x; as_series' items; {p}add/{p}sub's operands (PolyNumber.FromObject routes sequences to it -
   np.asanyarray converts only ONE level of an object[] and discovers a C# int as int32 where Python's is int64).

2. U3 evaluation facades take NumPy's argument kinds (np.polynomial.{polynomial,chebyshev,legendre,laguerre,
   hermite,hermite_e}.cs, NDPolyEval.cs).
   - {p}val(object x, object c, bool tensor = true), {p}val2d/val3d/grid2d/grid3d(object ...),
     {p}valnd(object[] pts, object c): lists, tuples, nested sequences, typed arrays and scalars for every
     ordinate and for c.
   - A Python-scalar ordinate becomes a 0-d array of its discovered dtype (NOT ndmin=1: NumPy's result shape
     is (), 120 polyeval cases caught the (1,) form); c converts with ndmin=1.
   - _valnd/_gridnd convert every ordinate first, then check shapes with NumPy's texts ("x, y, z are
     incompatible" / "x, y are incompatible" / "ordinates are incompatible").
   - A ragged c is reported before a ragged x (NumPy converts c first).
   - Conversion ownership: a converted argument is the call's intermediate. U3 disposes it by hand (Val,
     ValNd, GridNd incl. the error paths - a ragged second ordinate leaked the partial result before); U1/U4
     facades are [NDScoped].

3. polyutils list overloads (np.polynomial.polyutils.cs, NDPolySeries.cs).
   - trimseq(object[] seq) returns the list ITSELF or a new list of the same element type (NumPy's list
     slice); trimseq(object seq) takes an NDArray, a str (itself), a tuple (MakeTuple), an array-like
     (asanyarray, the wrapper disposed when the result is a view), any IEnumerable (object[]); None is
     "object of type 'NoneType' has no len()", a non-sequence "object of type '<name>' has no len()".
   - trimseq's zero test is Python's item != 0 over every item kind, NumPy's truth-value errors for array
     items of size 0 or > 1.
   - mapdomain(object[] x, ...) and mapdomain<T0..T3>(object[] x, (T0,T1), (T2,T3)).
   - OVERLOAD TRAP: C# binds an NDArray overload over an object one for ANY C# array argument (NDArray is the
     better conversion target through the implicit Array -> NDArray conversion), so a list needs its own
     object[] overload or it silently becomes an ndarray of the wrong dtype (or fails to convert).

4. Calculus: NumPy's widened scale loop (NDPolyCalc.cs, DirectILKernelGenerator.PolyCalculus.cs,
   ILKernelGenerator.Polynomial.Lanes.cs).
   - A strong scalar that PROMOTES the series (float64 against float32; float32/float64 against float16) makes
     NumPy's `c *= scl` run the WIDER multiply loop and cast back. PolyCalcKey carries ScaleLoop
     (WidenedScale / HasWidenedScale) and the load stage calls PolyLaneOps.F32ScaleF64 / HalfScaleF64 /
     HalfScaleF32 (vector) and their *Scalar twins.
   - float64 -> float16 must round ONCE (npy_double_to_half): the vector path rounds the float64 product to
     float32 with ROUND-TO-ODD (RoundToOddF32: truncate + sticky bit), then narrows to float16
     to-nearest-even, exact because 24 >= 11 + 2 bits; a plain f64 -> f32 -> f16 chain double-rounds at
     float16 ties. A NaN coefficient keeps its own payload, quieted.
   - An ARRAY scl and promotions without a widened kernel (decimal) run ScaleInPlace =
     np.multiply(c, scl, out: c, dtype: promote(series, scl)); the explicit dtype is load-bearing (the house
     ufunc computed a float32 array times a 0-d float64 with a float32 out in FLOAT32 - 44 corpus cells red).
   - k's items are converted only when their order uses them, AFTER that order's {p}val(lbnd, tmp) (Python's
     left-to-right k[i] - {p}val(...)); a str lbnd is refused only where NumPy evaluates at it; np.ndim runs on
     lbnd and scl before anything converts; a BigInteger past float64 raises CPython's "int too large to
     convert to float" (OverflowException), one past float32/float16 is inf.
   - Mutants W1-W4 on the widened scale (drop round-to-odd, narrow in the loop dtype, wrong loop, skip the
     widening) each turned 44-73 cases red.

5. Oracle (test/oracle/gen_oracle.py, OpRegistry.Polynomial.cs, OpRegistry.PolySeries.cs, FuzzCorpusTests.cs).
   - polyeval.jsonl 17,326 -> 19,216: section J (Python-sequence x - tuples, nested lists, empty, ragged,
     NumPy-scalar items - and Python ints past int64 as BigInteger) and K (every ordinate form x c in
     {float64 array, float32 array, list, ragged} for val2d/val3d/grid2d/grid3d/valnd).
   - polyseries.jsonl 18,437 -> 18,647: section O (tuples / nested sequences through as_series, add/sub,
     mapparms/mapdomain) and P (26 trimseq sequence cases: lists, tuples, nested, mixed item kinds, None,
     non-sequences, array items of size 0 and > 1).
   - polycalc.jsonl 21,646 -> 27,526: section M (argument kinds for c/k/lbnd/scl), N (zero-size, 5-D and
     extreme-int series), O (widened scale, incl. float16 double-rounding hazard scales built by
     _pc_f16_hazards: every finite float16 against scales like 1 + 2^-11 +- 2^-40 - random scales never land
     within ~2^-24 of a float16 tie).
   - 0 excused; no new op keys (the coverage oracle-evidence join is unchanged); floors polyeval 19100,
     polyseries 18600, polycalc 26400.
   - The registry passes c/x specs as objects (the NDArray casts are gone) and decodes tuples of any arity.

6. Unit tests: test/NumSharp.Tests/Polynomial/PolynomialArgumentKindsTests.cs (new, 21) - the C# boundary
   kinds with NumPy-probed bytes and texts: array-like c for {p}val (list, nested, Half), ragged c reported
   before ragged x, array-like ordinates (val2d float64 vs grid2d float32 results), the list overload of
   mapdomain, trimseq keeping the sequence kind, calculus argument kinds and the widened scale.

7. Benchmark: polycalc section A (114 argument-form cells: list / nested / row-list series, a float64 0-d scl
   widening float32/float16 series, array scl per column/row/coefficient). Full rerun, both sides pinned to
   0xF0, every cell SHA-256-checked against NumPy: 720 cells, min 1.71x, geomean 7.62x (section A min 1.71x
   = the float16 widened-scale N-D cells, the known open float16 lever; N-D float64 min 1.82x, host drift
   from 2.02x at delivery). polycalc_results.md regenerated.

8. Docs: CLAUDE.md (U3 argument kinds, U1 overload trap, U4 argument kinds / widened scale / conversion
   ownership / traps, perf line), docs/plans/numpy-polynomial.md (counts, wholeness bullet, perf table),
   Fuzz/README.md (new sections, counts, floors).

VERIFICATION
------------
- NumSharp.Tests (CI filter) net10.0 17,831 passed / net8.0 17,830 passed, 0 failed.
- NumSharp.Tests.Oracle (CI filter) 234/234 on net10.0 and net8.0 (FuzzMatrix incl. the polynomial tiers, the
  leak gate UndisposedIntermediateTests, the strength/surface/spread gates).
- Leak probe over every new overload: the only escapes it reported were arrays the probe itself built;
  implicit C#-side conversions at a call site (double[] -> NDArray) belong to the caller.

TRAPS (recorded in CLAUDE.md)
-----------------------------
- A lone object[] passed to a params object[] parameter IS the params array, not its one item.
- A ?: or switch expression whose arms are object[] and NDArray is typed NDArray through the implicit
  Array -> NDArray conversion: cast every arm to object.
- Python read_text/write_text on Windows turn this repo's LF files into CRLF working copies; edit through the
  Edit tool or in binary mode.
… 6 bases + X2poly/poly2X; NumPy's Python replayed over a pooled arena, np.convolve's dotfunc regimes modelled, blocked long products, fused chebmulx IL kernel; polyalgebra oracle 26,163 bit-exact + 186 host-pinned BLAS-bound, 404 bench cells min 1.45x geomean 34.9x NumPy

Plan unit U2 of docs/plans/numpy-polynomial.md: the 40 series-algebra names of numpy.polynomial.
- {p}mulx, {p}mul, {p}div (tuple quo, rem), {p}pow (int and Python-float powers, maxpower), {p}fromroots for polynomial,
  chebyshev, legendre, laguerre, hermite, hermite_e.
- X2poly / poly2X for the five non-power bases.
- Facades: np.polynomial.<module>.<name>, [NDScoped]. The package is now 142 of 193 names (U1 + U3 + U4 + U2).

HOW IT RUNS
NumPy runs these as Python loops over SHORT series. Every statement is one of: as_series, _add/_sub, a {p}mulx, an
ARRAY op with a Python int or NumPy scalar, np.convolve, trimseq, or scalarmath on one element. The engine
(Polynomial/Package/NDPolyAlgebra.cs + NDPolyAlgebra.Bases.cs) replays that Python statement for statement, in NumPy's
order and dtypes, over raw series (PolySer: pointer, length, dtype, and whether NumPy returns it as a view). Each
statement runs through the kernel NumPy's statement runs through in NumSharp:
- array ops: the house ufunc kernels behind flat slot fronts (GetPolyHouseBinaryKernel over GetMixedTypeKernel:
  simd_cmul, CDOUBLE_divide's Smith, the HALF loop). The slot fronts exist because the engine issues dozens of calls
  per Python-level step, and a ~25 ns dictionary lookup would dominate a dozen-element op. The U1 TrimLen / Combine /
  Cast kernel getters gained the same fronts (PolyCastSlot holds its exact pair next to the kernel: dtype codes fold
  modulo 32);
- as_series / _add / _sub: U1's substrate. NDPolySeries.CombineInto was split out of AddSub so arena series combine into
  caller-owned memory; PolySeriesView.Raw views arena memory;
- {p}mulx: integral-layout calculus kernels (PolyCalcKey.Mulx + PolyCalcRoutines.GetMulx: the six mulx routines as
  data, prd one row above c); chebmulx of >= 2 terms the fused kernel below;
- np.convolve: NDArray.SlidingCorrelateInto, the pointer-level core np.convolve itself runs, including the byte-parity
  OpenBLAS route;
- scalarmath: U1's one-element kernels (the naive complex product, NumPy's scalar division).

Dtypes are carried per series, because NumPy's change mid-algorithm:
- legmul/lagmul/hermmul/hermemul with a one-term shorter factor bind the Python int `c1 = 0`, which as_series makes
  float64 [0.], so a float16/float32 product turns float64;
- _div computes the remainder against the int unit series [0]*i + [1], so a float32 remainder is float64 while the
  quotient stays float32 (polydiv keeps float32 in both, its remainder a VIEW);
- poly2X starts from res = 0 (float64); fromroots' lines are np.array([-r, 1]) (strong array coercion);
- {p}pow(c, 0) is [1] in c's dtype.

THE ARENA (PolyArena)
A per-thread bump allocator. A series costs a pointer bump, not an NDArray (~200 ns each, the whole cost of a 1-to-1
NDArray port). Keep moves a function's result down to its entry mark, so memory is bounded by the live series. Blocks
come from SizeBucketedBufferPool, and growth blocks (powers of two) go back to it when the outermost call exits. Fresh
OS memory cost ~150 us of first-touch page faults per 800 KB per call on long products, more than the convolution.
Only the final result becomes an NDArray, with NumPy's ownership (a trimseq slice is a view, owndata false).
NDPolySequence gained a flat fast path (TryFlatScalars) for an object[] of boxed doubles or ints.

NP.CONVOLVE'S DOTFUNC, MODELLED PER DTYPE (Math/NDArray.SlidingDot.cs)
The power and Chebyshev products are np.convolve, whose per-position dotfunc is scipy-openblas' ?dot: a plain scalar
sum below its vector block, a reordered vector kernel from there on. Each scalar regime is now reproduced exactly:
- ddot: a sequential sum below 16 terms (DotSimd);
- sdot: float32 products summed in a DOUBLE below 32 terms (SdotManaged; probed 0 of 9,300 differ, a float32 sum
  differed on up to 72%);
- zdotu: four separate double sums below 8 terms (ZdotuManaged; 0 of 2,100 differ, the naive per-term product on up
  to 90%).
float16's HALF_dot is a sequential float32 sum at every length. So without the backend every position shorter than the
vector block is byte-exact, and only vector-regime positions are bounded-ULP; with NumSharp.Interop.OpenBLAS they run
the same ?dot NumPy calls, byte-exact.

LONG PRODUCTS WITHOUT THE BACKEND (Math/NDArray.SlidingDot.Long.cs, new)
Blocked kernels over exactly the vector-regime positions:
- B consecutive outputs share ONE loop over the kernel: a broadcast of k[s], one load per vector, FMAs into eight
  independent chains (two accumulator sets over alternating s);
- the core reads the data directly; the <= B-1 head/tail steps read tiny zero-padded edge windows;
- a padding lane adds 0*k[s], which is 0 only for a FINITE kernel, so a kernel holding inf/NaN declines to the
  per-position kernels;
- float16 gets an EXACT blocked kernel (sequential per lane, separate multiply and add, over house-widened float32
  copies): every float16 position is byte-exact.
Masked edge lanes measured 12% slower at 1000x1000; padded ramp copies cost O(n2) per call. np.convolve 1000x1000
(paired): float64 1.83x NumPy (was ~0.9x), complex128 1.43x (was 0.25x).

THE FUSED CHEBMULX KERNEL (Backends/Kernels/Direct/DirectILKernelGenerator.PolyAlgebra.cs, GetPolyChebMulxKernel)
chebmulx is NumPy's one mulx that is not a loop but three array statements:
    tmp = c[1:]/2; prd[2:] = tmp; prd[0:-2] += tmp
So NumPy streams it. Two earlier routes fell short at 10000 terms:
- the calculus kernel, which vectorizes over COLUMNS and divides once per row: 0.43x float16, 0.74x float32;
- the three statements through house kernels: 1.05x complex128, 1.48x float64.
The statements' combined effect is prd[j] = c[j-1]/2 + c[j+1]/2, computed in ONE IL pass:
- a root DynamicMethod (head, scalar remainder, tail: the literal scalar ops) calls a separate vector stage, so the
  root's scalar helpers cannot exhaust the stage's JIT inline budget;
- the vector stage uses U3's lane kinds and halves by a MULTIPLY by 0.5. x/2 and x*0.5 are the same exact half,
  rounded once, for every binary float, NaNs included; vdivpd would bound the pass at one division per element;
- float16 halves exactly in float32, then PolyLaneOps.HalfHalveOnGrid rounds the half to the f16 grid (magic number:
  0.75 + |y| has float32 ulp 2^-24, the f16 subnormal step; the sign is ORed back). The add is a PLAIN float32 add,
  rounded once by the store: NumPy's HALF_add;
- complex uses Smith's branch with the divisor 2+0j prepared once (CDivPrep/CDivBy: rat = +0, scl = 0.5);
- decimal is scalar-only;
- the kernel reads a contiguous same-dtype series in place (no as_series copy), or the converted copy in prd[1:]
  (every element is read before the pass writes its slot).
Result: 5.5-6.5x NumPy at 10000 terms. The first float16 version used the lane kind's generic Bin, which rounds every
op to the grid (narrow + widen + NaN blends): 3.7 ns/element, fully inlined (0 calls in JitDisasm), so the cost was
instruction count. HalfHalveOnGrid took it to 0.9 ns.

ORACLE
test/oracle/gen_oracle.py polyalgebra writes three files:
- polyalgebra.jsonl: 26,163 cases, sections A-L, 0 excused, portable;
- polyalgebra_parity.jsonl: 186 BLAS-bound products;
- polyalgebra_parity.host.jsonl: the pin.
The generator patches np.convolve while a case runs (_PAConvRecorder: lengths, dtype, finiteness) and routes a case
with a vector-regime product to the host-pinned file. That file is replayed with the OpenBLAS backend at threads=1 and
is Inconclusive off the pinned host (MatmulParityPin, like matmul_parity).
Sections:
- A dtype x length, and the mul/div dtype-pair matrix;
- B full-mantissa values across every scalar/vector boundary;
- C trim and special patterns;
- D layouts, 0-d, broadcast;
- E Python-typed arguments (scalars, lists, tuples, big ints, nested);
- F errors in NumPy's check order (the empty-message ZeroDivisionError);
- G pow / maxpower kinds;
- H long series across the managed/BLAS boundary;
- I result flags (trimseq views);
- J fromroots' root kinds (one zero per list);
- K underflow / overflow;
- L the fused chebmulx kernel's vector stage: specials at 40/101 terms, every float16 pattern below 2^-12 plus
  inf/NaN/extremes as c[j-1] and as c[j+1], float32/float64 subnormals.
Char rides the uint16 proxy.
OpRegistry.PolyAlgebra.cs replays both files. {p}pow binds NumPy's argument kinds: an in-range int to the int overload,
a float or huge int to the double one, maxpower only when present. BlasBackendDelta gained the six convolving ops and
the polyalgebra tier: 6,522 affected, 6,480 identical outcomes, 42 flips, all byte-checked against NumPy (14 are
NaN-payload-only).
Planted-bug checks on section L:
- dropping the float16 grid rounding: 2/26,163 red (only section L reaches it);
- a plain complex multiply instead of Smith: 5 red.

TESTS
Polynomial/PolynomialAlgebraTests.cs (20):
- NumPy's TestArithmetic (mulx/mul/div/pow) and TestMisc (fromroots/conversions) x 6 ported;
- dtype-quirk and trim/view byte dumps probed from NumPy 2.4.2;
- decimal and char; the C# spellings of {p}pow; the arena (growth blocks released, per-thread);
- a long product without the backend within the Dot2 error bound;
- the fused chebmulx kernel: vector stage vs scalar-only emission over every series dtype, lengths 2-70,
  specials/subnormals, both layouts, no fallback; and float16 over EVERY bit pattern vs a model of NumPy's HALF loop;
- two [Misaligned] divergences:
  - fromroots with mixed +0.0/-0.0 roots: NumPy's x86-simd-sort order of equal keys is CPU-dependent; NumSharp puts
    -0.0 first, which decides a zero coefficient's sign;
  - a non-finite complex product without the backend: zdotu's vector kernel mixes real and imaginary lanes.
Both new chebmulx tests go red on the grid-rounding mutant.

VALIDATION
- Oracle full suite 236/236 on net10.0 and net8.0 (RandomApiSoak skipped).
- Unit suite, CI filter: net10.0 17,851 passed / 11 skipped; net8.0 17,850 / 11.
- Interop live-NumPy convolve/correlate/signal parity: 34/34.
- coverage/test_oracle_evidence.py OK.
- A bit-identity probe of the fused chebmulx kernel vs the calculus mulx kernel: 35,896 checks, 0 differences
  (float16 exhaustive; 16 NaN-payload-only).

PERF (benchmark/polynomial/polyalg_{numpy.py,bench.cs,report.py}, committed summary polyalg_results.md)
404 cells: 371 bit-exact (SHA-256) and 33 BLAS-bound checked to a relative 1e-12. Min 1.45x, geomean 34.88x NumPy.
- mulx: min 2.41, geomean 23.6
- mul float64: min 1.98, geomean 29.0
- mul other dtypes: min 1.45, geomean 26.9
- div: min 7.99, geomean 61.2
- pow: min 5.26, geomean 26.3
- fromroots: min 6.19, geomean 31.9
- conversions: min 40.5, geomean 91.8
- Python-list arguments: min 5.04, geomean 36.0
Method: both sides pinned to 0xF0, never concurrent; DOTNET_TC_CallCountingDelayMs=0 and a drained tier-1 warm-up;
best-of-rounds; OPENBLAS_NUM_THREADS=1.

CEILING
complex128 1000x1000 polymul is 1.45x (1.49x paired). A complex product is four real multiply-adds, one 256-bit FMA
per complex MAC on both sides (zdotu's AVX2 microkernel does exactly that), and the blocked kernel runs at ~90% of FMA
peak, so only NumPy's per-position overhead is left to win. Rejected:
- Gauss's three-multiply product: it loses componentwise accuracy on real x complex imaginary parts;
- FFT convolution: its error is normwise, not per coefficient.

TRAPS (also in .claude/CLAUDE.md, the U2 section)
- A float16 lane kind's generic Bin rounds every op to the f16 grid. Round only where NumPy's float16 array holds the
  value, and let the store's narrow be the add's rounding.
- Emit a power-of-two halving as a multiply: vdivpd bounds a streaming pass.
- Random operands never reach a vector stage's rare branches: build the edge set (section L).
- Slow SIMD with no calls in JitDisasm is instruction count, not inlining.

Docs: .claude/CLAUDE.md (a new "series algebra (U2)" section; the correlate/convolve section gained the dotfunc models
and the blocked long products; a Core Source Files row), docs/plans/numpy-polynomial.md (U2 DELIVERED block, section
0.1 count), test/NumSharp.Tests.Oracle/Fuzz/README.md (the polyalgebra tiers).
…wer (CPython int() + NumPy's comparisons), str/object series refused in NumPy's error order, zdotu's C99 result + CDOUBLE_dot's plain loop for complex np.convolve/np.correlate; polyalgebra 28,144 + parity 139, polyseries 18,698, groupa 601 bit-exact, 452 bench cells min 1.57x geomean 37.8x NumPy

WHAT
----
A second pass over U2, the numpy.polynomial series algebra (8d71581e: {p}mulx / {p}mul / {p}div / {p}pow /
{p}fromroots for the six bases, X2poly / poly2X for the five non-power bases). The question was whether every
argument NumPy 2.4.2 accepts is accepted, answers byte-for-byte, and fails with NumPy's error in NumPy's order.

A probe generator crossed every C# argument kind of the house boundary map (NDPolyNumber's header) with every U2
function in all six modules. It added pow / maxpower kinds, error-order pairs (a bad series AND a bad power, a bad
divisor AND a str dividend, ...), non-finite complex products, and np.convolve / np.correlate with one-element
operands of every stride. That is 6,675 NumPy-generated cases, replayed through the public facades by reflection.
It found three gaps, all closed here. Two things remain, and both are inherent (see REMAINING).

GAP 1 - {p}pow took only int / double / int? for pow and maxpower
-----------------------------------------------------------------
NumPy reads both with three Python statements, after as_series (polyutils._pow and chebpow's copy of it):
    power = int(pow)
    if power != pow or power < 0: raise ValueError("Power must be a non-negative integer.")
    elif maxpower is not None and power > maxpower: raise ValueError("Power is too large")
So WHICH values pass, and which error a bad one raises, is CPython's int() and rich comparison, plus NumPy's
comparison ufunc when an argument is a NumPy scalar or an ndarray.

- New src/NumSharp.Core/Polynomial/Package/NDPolyPowerArgument.cs (PolyPowerArgument) ports the three statements
  for every C# value, all probed against 2.4.2.
  * Int(pow) is CPython's int():
    - a bool is 0/1; a double / float / Half (np.float16) / decimal truncates toward zero;
    - NaN raises ValueError "cannot convert float NaN to integer", +-inf OverflowError "cannot convert float
      infinity to integer";
    - a str parses with int()'s grammar ("3", " 3 ", "3_0"; "x" and "" raise "invalid literal for int() with
      base 10: 'x'");
    - a Complex / null / list / tuple raises TypeError "int() argument must be a string, a bytes-like object or a
      real number, not 'complex'" ('NoneType' / 'list' / 'tuple');
    - an ndarray of one or more dims (an NDArray, a typed C# array, a Memory<T>) raises TypeError "only
      0-dimensional arrays can be converted to Python scalars" whatever its size; a 0-d NDArray converts its
      element.
  * DiffersFromInt: an integer value always equals its int(); a float, Half or decimal differs exactly when it had a
    fractional part; a str always differs (an int never equals a str, so a parsed "3" still fails).
  * Exceeds(power, maxpower):
    - exact for bool / every integer type / BigInteger / char / decimal;
    - a Python float compares EXACTLY (CPython's int-vs-float): NaN and +inf are never exceeded, -inf always is;
    - np.float16 (Half) and float / complex ndarrays compare in THAT dtype (NEP 50). The weak int goes to a double
      first ("int too large to convert to float" at 2**1024), then to float16 / float32, so 2049 > np.float16(2048)
      is False;
    - complex compares by NumPy's CGT, (xr > yr && !isnan(xi) && !isnan(yi)) || (xr == yr && xi > yi);
    - a bool ndarray compares in int64 ("int too big to convert" past int64); an integer ndarray compares exactly;
    - the ndarray's comparison result is then tested for truth. It needs exactly one element, else NumPy's two
      ValueError texts ("The truth value of an empty array is ambiguous..." / "...more than one element is
      ambiguous..."). The conversion's OverflowError comes first, whatever the size;
    - a Complex / str / list / tuple / unknown object raises TypeError "'>' not supported between instances of
      'int' and '...'".
    None of this loops over data: an ndarray argument has one element (read directly) or fails by its size alone.
    HalfLess narrows through the house double->half cast kernel (npy_double_to_half, overflow to inf).
- Every facade (np.polynomial.{polynomial,chebyshev,legendre,laguerre,hermite,hermite_e}.cs) gained two [NDScoped]
  overloads:
  * {p}pow(object c, object pow): the default limit, i.e. NumPy's maxpower=16 (NDPolyAlgebra.DefaultMaxPower, boxed
    once), or None for polypow;
  * {p}pow(object c, object pow, object maxpower): null is None.
  C# binds int to the int overload and long / ulong / uint / double / float to the double overload, so only the
  kinds those cannot take reach the object overloads. Every NumPy spelling still compiles and binds.
- NDPolyAlgebra.Pow(basis, c, object pow, object maxpower) runs NumPy's order:
  as_series -> int(pow) -> `power != pow or power < 0` -> `power > maxpower` -> the object-series refusal (GAP 2) ->
  a power beyond int64 raises OverflowException (NumPy would multiply until memory runs out) -> PowProduct (the
  former body of PowChecked, split out so the two entry points share it).

GAP 2 - a None / str series was refused on conversion, before NumPy's own errors
--------------------------------------------------------------------------------
polyutils.as_series first converts every argument, then checks each one's size and dims, then trims, and only
THEN takes np.common_type. When that fails, its except path looks for ANY object array:
- with one, NumPy computes with Python objects (None, a non-numeric object, a Python int past uint64);
- without one it raises ValueError "Coefficient arrays have no common type". A str array (a list holding a str)
  lands here, and so does a bool series.
NumSharp refused a str or None on conversion, so an empty or 2-D sibling argument, div's zero divisor, or pow's power
checks never got to report their own error first.

- NDPolySequence.cs: ToArray = ToArrayOrNonNumeric(seq, out nonNumeric) ?? throw nonNumeric.Refusal.
  * ToArrayOrNonNumeric returns null with NumPy's array SHAPE for a walk that ended refused rather than ragged.
    The shape is a PolyNonNumericArray: dims, whether a leaf was an OBJECT (not a str), and the refusal. A ragged
    sequence still raises the inhomogeneous-shape ValueError.
  * Discovery gained ObjectLeaf. ScalarLeaf sets it for refused items that are not strings, and so does the
    huge-int catch.
- NDPolySeries.cs:
  * PolySeriesView gained IsObject and Refusal. The ctor takes them as optional parameters, and Slice / Reversed
    carry them over.
  * NonNumericArray(dims, isObject, refusal) is a view with NumPy's len / ndim / size, so Validate reports empty /
    not-1-d exactly as for a numeric array.
  * AsCoefficientArray: null becomes a one-element object array (NumPy's np.array(None)). A sequence goes through
    ToArrayOrNonNumeric. A scalar FromObject cannot convert becomes an object view.
  * TryCommonType: an object view means "no common type, compute with objects" (false). A str view raises the
    no-common-type ValueError, and so does a common_type TypeError.
  * CommonType = TryCommonType, else throw FirstObjectRefusal (whose doc says it throws InvalidOperationException
    when there is no object view).
- NDPolyAlgebra.cs:
  * Div converts both operands, validates both, then calls TryCommonType. On the object path NumPy's next
    statement `if c2[-1] == 0: raise ZeroDivisionError` still runs before any object arithmetic: a NUMERIC divisor
    whose trimmed last coefficient is zero (False included) raises the bare DivideByZeroException, and anything
    else throws the refusal. So polydiv([None], [False]) is ZeroDivisionError, as in NumPy.
  * Pow (int / double / object) raises the refusal only AFTER the power checks. NumPy computes with Python objects
    even for power 0 or 1, which return object arrays.
  * Exception docs are rewritten throughout ("an object series ..., refused where NumPy starts computing with it";
    a str series is the common-type ValueError).
- NDPolyEval.cs (ArrayOf) and NDPolyCalc.cs (Coefficients) throw `v.Refusal ?? <old message>`. U3 and U4 compute at
  once, so NumPy's error order there is unchanged: refuse on conversion.
- Facade docs fixed ("power seriess" -> "power series"; the NotSupportedException / ValueError texts per the rule
  above).

GAP 3 - np.convolve's complex dot, the two things NumPy's CDOUBLE_dot does that NumSharp did not model
----------------------------------------------------------------------------------------------------
(a) zdotu's result construction. scipy-openblas builds zdotu's result with C99 complex arithmetic,
    openblas_make_complex_double(re, im) = re + im*_Complex_I, so the real part becomes re + im*0: a NaN when the
    imaginary part is infinite or NaN. This was verified against the DLL NumPy loads (scipy_cblas_zdotu_sub64_):
    np.convolve([1+0j], [inf+0j]) is nan+nanj where the plain sums give inf+nanj.
    - NDArray.SlidingDot.cs: ZdotuResult(re, im) = { sum0 = 0; sum0 += re + im*0.0; sum1 = 0; sum1 += im }.
      ZdotuManaged returns it in both regimes: ZdotuResult(d0 - d1, d2 + d3) in the scalar regime (the four
      separate sums), ZdotuResult(sr, si) in the vector regime.
    - NDArray.SlidingDot.Long.cs: FmaBlockComplex stores go through a vector ZdotuResult(Vector256<double>). The
      imaginary lanes are permuted onto the real lanes and multiplied by zero, the product is added, and the
      imaginary lanes are BLENDED back from the input (Avx.Blend 0b1010). im + re*0 would poison an imaginary part
      with an infinite real part, which zdotu does not do.
(b) When NumPy calls cblas at all. CDOUBLE_dot uses cblas only when blas_stride holds for BOTH operands: a positive
    stride that is a multiple of the itemsize. PyArray_Correlate converts with NPY_ARRAY_DEFAULT, which copies only
    a non-contiguous operand or one of another dtype. A size-1 array is C-contiguous whatever its stride, so a
    ONE-element operand arrives with its own stride. np.convolve's kernel is v[::-1], so a fresh one-element
    kernel has stride -16: blas_stride refuses it, and CDOUBLE_dot runs its plain loop instead:
    sumr/sumi from 0.0, then `sumr += ar*br - ai*bi; sumi += ar*bi + ai*br`, with no result poisoning.
    - NDArray.SlidingDot.cs:
      * DotOperandBlasable(x, retType, reversed): false only for a one-element operand of the result dtype whose
        (possibly reversed) stride is not positive. A cast or multi-element operand is always a fresh positive
        copy.
      * SlidingCorrelate / SlidingCorrelateInto gained complexDotViaBlas (default true). False runs
        SlidingComplexPlain / CdoubleDotPlain and skips the backend, as NumPy skips cblas.
    - NDArray.Convolve.cs decides from the CALLER's arrays, before MaterializeForSliding loses the stride:
      DotOperandBlasable(a, fwd) && DotOperandBlasable(v, reversed).
    - NDArray.Correlate.cs: conj(v) is always a fresh array, so only `a` decides.
    - NDPolyAlgebra.cs Convolve (U2's arena) passes complexDotViaBlas: v.N != 1. Both factors are fresh arrays of
      one dtype there, so only np.convolve's reversed one-element kernel is non-blasable.
Effect: complex products holding inf / NaN are now managed-exact below the vector block.
- PolynomialAlgebraTests.ComplexNonFiniteProduct_WithoutBackend_MatchesZdotu is no longer [Misaligned]. It asserts
  NumPy's bytes: the polymul inf four-term, the one-term factor, and polypow([inf+1j], 3).
- _PAConvRecorder.blas_bound no longer counts non-finite values, and counts complex only when the dot is >= 8 terms
  long. 47 cases moved from the host-pinned tier to the portable one.
- The groupa tier's new complex convolve / correlate block gates np.convolve / np.correlate with specials and
  one-element operands (fresh, reversed, stepped, negative-stepped, stride-0).

ORACLE (NumPy 2.4.2, 0 excused)
-------------------------------
- polyalgebra.jsonl 26,163 -> 28,144 (+1,981). New section M of gen_polyalgebra:
  * argkind 198: every C# argument kind of every function;
  * maxkind 870 + maxedge 96: pow / maxpower kinds, the float16 / float32 rounding edges, the truth-value and
    conversion errors;
  * objorder 350: the deferred refusal against every error NumPy reports first;
  * cspecial / cspecial1 / cspecial11 / cspecial1r / cspecialr 420: non-finite complex products, one-element
    factors of both orientations;
  * plus the 47 moved cases (pat 41, rk 4, ext 2).
- polyalgebra_parity.jsonl 186 -> 139. The host pin (polyalgebra_parity.host.jsonl) is unchanged.
- polyseries.jsonl 18,647 -> 18,698 (+51, section Q "objorder": U1's add / sub deferred refusal).
- groupa.jsonl 364 -> 601 (+237, the complex convolve / correlate block).
- gen_oracle.py:
  * _ps_enc encodes None and np.float16;
  * blas_bound as above;
  * a maxpower is spec-encoded unless it is None or a plain number;
  * _pa_object_land keeps only outcomes NumPy reaches BEFORE any object arithmetic. It drops an object-array result
    and any TypeError other than int()'s, the limit comparison's and len()'s, because NumSharp has no object dtype
    to compute with;
  * gen_polyalgebra M1-M4, gen_polyseries Q, and the gen_groupa complex block.
- OpRegistry.PolySeries.cs: Decode handles "none" (null) and "npscalar" (float16 -> Half via
  BitConverter.UInt16BitsToHalf); Raw(name) exposes a parameter's JSON.
- OpRegistry.PolyAlgebra.cs PolyPow binds as C# source would. A plain limit (absent / null / number) with a power
  that is a long in int range goes to the int overload; a long / ulong / double power to the double overload.
  Everything else, and every spec-encoded limit, goes to the object overloads.
- FuzzCorpusTests MinCases floors: groupa 590, polyseries 18,690, polyalgebra 28,100, polyalgebra_parity 135.
- BlasBackendDelta: 7,064 affected ordinary cases, 7,001 identical outcomes, 63 flips byte-checked against NumPy
  (was 6,522 / 6,480 / 42).

TESTS
-----
- New test/NumSharp.Tests/Polynomial/PolynomialAlgebraArgumentKindsTests.cs (5 tests, NumPy-probed bytes and texts):
  * SeriesSpellings_AreNumPysArrayLikes;
  * Pow_AnyPowerArgument_IsPythonsInt;
  * Pow_AnyMaxpower_ComparesAsNumPy;
  * ObjectAndStrSeries_RefusedWhereNumPyComputes;
  * ComplexDot_IsZdotuOrThePlainLoop.
- PolynomialAlgebraTests: the non-finite complex product is exact (above); the class doc is updated.

PLANTED-BUG CHECK (each planted, run, reverted; the tree is the verified one)
------------------------------------------------------------------------------
- drop zdotu's result poisoning: GroupA 127 + Polyalgebra 25 red;
- the backend path taken before the complexDotViaBlas decision: BlasBackendDelta 72 failures;
- np.convolve's reversed flag wrong: GroupA 45 red;
- U2's complexDotViaBlas forced true: Polyalgebra 15 red;
- the float16 limit compared exactly (no rounding of the power): 12 red;
- div's object-path zero test removed + pow's refusal moved before the power checks: 54 red.

RESULTS (this tree)
-------------------
- NumSharp.Tests (CI filter): net10.0 17,856 passed, net8.0 17,855 passed, 0 failed.
- NumSharp.Tests.Oracle (CI filter): net10.0 and net8.0 236 passed each, 0 failed (RandomApiSoak skipped: it
  needs a nightly soak directory).
- NumSharp.Tests.Interop, live NumPy with the backend (FullyQualifiedName~Polynomial|SlidingDot|Convolve|Correlate):
  21/21.

PERF (benchmark/polynomial/polyalg_*, NPY/NS, both sides pinned to 0xF0, every cell checked against NumPy)
---------------------------------------------------------------------------------------------------------
- New section K, pow / maxpower argument kinds through the object-typed overloads (10 coefficients, power 3): a
  bool / np.float16 / 0-d array power; a float / np.float16 / 0-d / one-element array limit, and None. 48 cells,
  min 5.11x, geomean 38.9x. The int() / comparison port costs ~20 ns.
- Whole sheet: 452 cells, min 1.57x, geomean 37.80x (404 cells, min 1.45x, geomean 34.9x at delivery). 412 cells are
  bit-exact; 40 are BLAS-vector-regime products, which are byte-exact only with the backend.
- polyalg_bench.cs: a PowFn helper and `record struct PowArgs` (declared at the end of the file, after the top-level
  statements) decoding JSON null / bool / number and {"f16"} / {"nd0"} / {"nd1"}. polyalg_numpy.py: section K,
  cell(extra_py=...). polyalg_report.py: the K entry. polyalg_results.md regenerated.

REMAINING (inherent, documented)
--------------------------------
- NumSharp has no object dtype. Wherever NumPy computes with Python objects (it returns object arrays of Python ints
  / None, or raises its object arithmetic's TypeError), NumSharp raises NotSupportedException, at the point NumPy
  starts computing.
- A vector-regime product can overflow to inf in one summation order and not in the other: 3 of 1,510 probe cases.
  Byte parity for that regime needs NumSharp.Interop.OpenBLAS, as before.

DOCS
----
- .claude/CLAUDE.md: the correlate / convolve dotfunc paragraph (zdotu's C99 result, CDOUBLE_dot's plain loop,
  DotOperandBlasable); the argument-kinds note in U4; the U2 section (counts, the wholeness-pass paragraph, one
  documented divergence left, perf, two new traps); the source-files row (NDPolyPowerArgument.cs).
- docs/plans/numpy-polynomial.md: the header row and the U2 block.
- test/NumSharp.Tests.Oracle/Fuzz/README.md: the polyalgebra section, section M, polyseries Q, the FIXED list, the
  backend-delta counts.

TRAPS (for the next pass)
-------------------------
- A one-element array's stride is observable. NumPy's dot passes a size-1 operand through with whatever stride it
  has, and cblas refuses a non-positive one, so polymul(c, [x]) and polymul(c, [x, y]) reach different complex dot
  code. MaterializeForSliding loses that stride: decide DotOperandBlasable from the caller's array first.
- A corpus case that NumPy only reaches through object arithmetic cannot be gated. Filter the generator's candidates
  by NumPy's OUTCOME (_pa_object_land), not by hand; the filter keeps exactly the errors NumPy raises before it
  computes.
- Inside namespace NumSharp.Tests.*, `Math.Pow` binds to the namespace NumSharp.Tests.Math (CS0234). Spell
  System.Math.Pow.
- Bash heredocs lose one level of backslash escaping. A script holding `\\` or '\n' must be written with the Write
  tool, not a heredoc.
…ses; one IL kernel per basis on the calculus step language, fused outer product, streamed (non-temporal) products past 32 MiB; polyvander oracle 16,612 bit-exact, 480 bench cells min 1.55x geomean 10.43x NumPy

Plan unit U5 of docs/plans/numpy-polynomial.md: polyvander/polyvander2d/polyvander3d, chebvander*, legvander*,
lagvander*, hermvander*, hermevander* (18 names). The package now has 160 of its 193 names.

NUMPY'S STATEMENTS, IN NUMPY'S ORDER (which is also its error order)
--------------------------------------------------------------------
1-D (numpy/polynomial/{polynomial,chebyshev,legendre,laguerre,hermite,hermite_e}.py, NumPy 2.4.2):
  ideg = pu._as_int(deg, "deg")      operator.index; TypeError "deg must be an integer, received {format(deg,'')}"
  if ideg < 0: ValueError "deg must be non-negative"
  x = np.array(x, copy=None, ndmin=1) + 0.0     bool/int/char -> float64; -0.0 -> +0.0; sNaN quieted, payload kept
  v = np.empty((ideg + 1,) + x.shape)           "Maximum allowed dimension exceeded" at >= 2^63-1 rows,
                                                AllocationGuard's "array is too big; ..." below, MemoryError past it
  v[0] = x*0 + 1; v[1] = <x | 1 - x | x*2>; for i in range(2, ideg+1): v[i] = <recurrence>
  return np.moveaxis(v, 0, -1)                  a VIEW (OWNDATA false): F-contiguous for 1-D points, C and F for a
                                                scalar / degree 0 / empty points, neither for N-D points
2-D / 3-D (polyutils._vander_nd / _vander_nd_flat): len(deg) ("object of type 'int' has no len()", "len() of unsized
object") -> "Expected N dimensions of degrees, got K" -> np.asarray(tuple(points)) + 0.0 (strong promotion: float32 +
int32 is float64, float16 + int8 float16; a ragged stack is np.array's inhomogeneous ValueError) -> per dimension its
degree checks and its own matrix allocation, and from the second dimension on the product's allocation -> the empty
stack's reshape error ("cannot reshape array of size 0 into shape (0,newaxis)", Shape.ConvertShapeToString) -> the
outer product flattened, column (a*(dy+1)+b)[*(dz+1)+c] = V_x[a]*V_y[b](*V_z[c]).

format(deg, '') is NOT str(): a 0-d array formats through its Python scalar (np.float16(1000) -> "1000.0"), an ndim>=1
array through array_str, a list/tuple through its items' REPR ([np.float16(1e+03), array([1.5]), (2,)]).
operator.index takes bool, every integer width, BigInteger, char and a 0-d integer array; a 0-d bool is refused.

ENGINE
------
- Polynomial/Package/NDPolyVander.cs - the driver (the statements above, in order). A 2-D / 3-D stack of same-shape
  NDArrays is read IN PLACE per coordinate, each loaded from its own dtype (NumPy's two-step `stack + 0.0` rounds at
  most once, the same value), so no stacked copy is made; anything else goes through NDPolySequence's np.array
  coercion. Non-flat point layouts are copied by NDIter first. Scratch comes from SizeBucketedBufferPool; the result
  is the house ARC contract (a view over the fresh buffer, working buffers disposed however the call ends).
- Polynomial/Package/NDPolyIndexArgument.cs - PolyIndexArgument.AsInt / Format / Repr (operator.index + format()).
- Backends/Kernels/Direct/DirectILKernelGenerator.PolyVander.cs - PolyVanderRoutines: the six recurrences as DATA in
  U4's step language (PolyCalcRoutine), run forward, NumPy's source expressions token for token (legvander's
  `v[i-1]*x*(2i-1)` is ((v[i-1]*x)*(2i-1)); lagvander's `2*i - 1 - x` is ((2i-1) - x)); x and chebvander's x2 live in
  per-block scratch slots (hermvander's x2 IS v[1], read back from that row). One kernel per (basis, dtype, 1/2/3-D,
  source dtypes, streamed), four DynamicMethod stages so each stays inside the JIT's inline budget:
    LOAD      d[i] = T(s[i*sCol]) + T(0) - the conversion and NumPy's `+ 0.0` in one pass (house lane kinds for a
              contiguous source, the scalar path for any stride; a bool source stays scalar: NumPy reads ANY nonzero
              byte as True)
    RECUR     the routine's head (v[0]; v[1] when ideg > 0) and loop over a block of points
    PRODUCT   (2-D / 3-D) each output row straight into the final layout; 3-D through a scratch row holding
              V_x[a]*V_y[b], NumPy's rounded temporary
    ROOT      walks the point blocks (48 KB 1-D, 192 KB N-D budgets, <= 4096 points, multiples of 16)
  Every per-dimension matrix stays in L1/L2 while its rows stream out; NumPy materializes three to five temporaries
  per degree.
- Shared-code changes: PolyCalcLoad/PolyCalcStore gained a Scratch flag and PolyCalcRowPointer resolves a scratch slot
  (DirectILKernelGenerator.PolyCalculus.cs - the calculus kernels never bind scratch, and an unbound slot throws
  InvalidOperationException at emission); ArrayFormatter.PythonComplexRepr and NDPolyCalc.View became internal.

THE OUTER PRODUCT'S NaN PRIORITY
--------------------------------
NumPy's FLOAT/DOUBLE_multiply (MSVC, win-amd64, probed at every loop length) returns the SECOND operand's NaN when both
are NaN; x86 mulpd/mulps return the first's, and RyuJIT may swap a commutative multiply's operands. The outer product
is the one place two independently sourced NaNs meet, so PolyVanderOps.Mul{128,256,512} + scalar Mul blend
explicitly - but only for a vector whose product holds a NaN lane (a NaN-free product had no NaN operand, so the plain
product IS NumPy's: one compare on the common path). float16 imposes the priority in its lane kind; complex128 keeps
the house simd_cmul (the oracle tokenizes complex NaN payloads - value NaN either way).

STREAMED PRODUCTS FOR RESULTS >= 32 MiB
---------------------------------------
PolyVanderKey.NonTemporal (driver: NDPolyVander.NonTemporalMinBytes = 32 MiB, 2-D / 3-D only). Each product row runs
a scalar head up to the store's alignment (row starts are only element-aligned: the row stride is npts*itemsize),
then aligned non-temporal vector stores (PolyVanderOps.StoreNonTemporal256 = vmovntdq for float64 / float32 /
complex128; StoreNonTemporalHalf8 = 16-byte movntdq after the house RTNE narrow for float16; no store for decimal or a
non-256-bit kind -> normal stores), then the scalar tail; the stage ends with sfence. The bytes are identical. A result
past a last-level cache is evicted before its consumer reads it anyway, and a normal store reads every line first:
  (100000, 121) float64, fresh pages (above the 64 MiB pool cap):  ~12 ms streamed vs 15-18 ms normal
  (66000, 121) float64, reused pooled buffer:                        1.63 ms vs 3.68 ms
About 10.4 ms of the fresh case is demand-zero page faults (0.44 us per 4 KB page on the dev host), which NumPy pays as
well. The 1-D form never streams: its recurrence reads its output rows back. Test hooks NonTemporalOverride (force
on/off) and BlockOverride (tiny blocks), both thread-static.

ORACLE - polyvander.jsonl (gen_oracle.py polyvander, 16,612 cases, 0 excused)
------------------------------------------------------------------------------
8,224 vander / 4,986 vander2d / 3,402 vander3d; 846 errors (390 ValueError, 456 TypeError, NumPy's texts); 840 flags
facets ([C_CONTIGUOUS, F_CONTIGUOUS, OWNDATA]); 600 char (uint16 proxy). Sections:
  (A) every dtype x length 0-33 x degree 0-14          (B) full-mantissa values around every lane width
  (C) special values (_pv_special: sNaN / negative-payload NaN, +-inf, +-0, subnormals, max patterns)
  (D) layouts of x (strided, reversed, offset, F / transposed N-D, broadcast, 0-d) with values AND flags, zero sizes
  (E) float16 Python-int constants past 2048 (round to even) and past 65504 (inf) - degrees 2,100 / 33,000 / 65,600
  (F) 2-D / 3-D at every dtype x point shape x degree pair/triple, 14 mixed-dtype pairs, special + full-mantissa values
      in EVERY coordinate (the NaN-pair products), per-coordinate layouts, flags, Python-typed points
  (G) 54 degree kinds, 20 x kinds, 37 / 14 degree containers, points' errors in NumPy's order
  (I) inputs longer than one block (1,700-9,000 points, strided)    (H) char relabel
Wired through OpRegistry.PolySeries.cs (CalcFacet) + OpRegistry.PolyVander.cs (a C# long in int range binds the int
overload, every other degree kind the object one). Two planted kernel bugs turned 2,403 and 117 cases red.
The coverage join (coverage/generate_coverage.py) resolves every new key: the 18 rows are oracle-verified (direct).

UNIT TESTS - test/NumSharp.Tests/Polynomial/PolynomialVanderTests.cs (16)
--------------------------------------------------------------------------
NumPy's TestVander x 6 ported (column i is {p}val of the unit series; vander2d/3d @ c.flat is val2d/val3d; shapes;
negative degree); byte dumps per dtype family; NaN bits (sNaN quieting, second-operand product NaN); degree kinds and
texts; degree-before-points and the str / object refusals; nd degree containers; dtypes; decimal; view flags; the
int / object overloads and layouts; forced small blocks; no vector fallback; and the two streamed-product gates:
forced on vs off at every basis x dtype x point count 1..1001 x block (every row-head alignment; a planted
skipped-head bug turns it red at `poly float64 n=2`), and the automatic path on 38.7 MB / 33.6 MB results (SHA-256).

PERF - benchmark/polynomial/polyvander_{numpy.py,bench.cs,report.py}, summary polyvander_results.md
---------------------------------------------------------------------------------------------------
480 cells, each result SHA-256-checked against NumPy before timing; both sides pinned to 0xF0, never concurrent.
  D 1-D float64 1..1M points x deg 3/10/30   min 2.23  geo 16.26
  T point dtypes @1K/100K                     min 1.65  geo  7.84
  L layouts @100k                             min 6.95  geo  8.49
  N N-D points                                min 7.81  geo  9.33
  V 2-D products 1..100K x (1,1)..(10,10)     min 1.55  geo 11.73
  W 3-D products 1..100K x (1,1,1)..(5,5,5)   min 2.40  geo 12.80
  M mixed dtypes @10k                         min 2.33  geo  4.73
  S small calls                               min 5.10  geo 10.18
  A argument forms                            min 2.54  geo  7.15
  all                                         min 1.55  geo 10.43
The floor is the six (100000, 121) 2-D cells, 1.55-1.66x (NumPy 21.9-24.2 ms vs 14.1-14.9 ms; before streaming
1.20-1.37x): ~10.4 ms of each side is page faults. The float16 cells (1.65-2.7x) are U4's open float16 lever (the
lane kind rounds every op to the f16 grid). The biggest cells run best-of-9 on both sides: fault-bound timings swing
with the OS's zeroed-page list (the same kernel measured 11.8-16.2 ms across best-of-9 runs).

VERIFICATION
------------
NumSharp.Tests (CI filter) net10.0 17,872 passed / 11 skipped, net8.0 17,871 / 11; NumSharp.Tests.Oracle net10.0 and
net8.0 237 passed / 1 skipped each (FuzzMatrix incl. Polyvander, the leak and native-allocation gates);
coverage/test_oracle_evidence.py 15 OK; the coverage generator's strict join clean.

TRAPS (in .claude/CLAUDE.md "Polynomial package - Vandermonde family (U5)")
---------------------------------------------------------------------------
- A random corpus never pairs two NaNs in a product's operands: it took special values in BOTH coordinates, and that
  is what found the NaN priority.
- A valid huge degree over an object x: NumPy computes `x + 0.0` with Python objects BEFORE the dimension error (where
  NumSharp refuses the object stack), so only degree errors that come first are corpus cases.
- NumPy hangs on a huge degree over EMPTY points whose byte count does not overflow (its Python loop runs anyway);
  never put one in a generator. NumSharp returns at once.
- A page-fault-bound cell measures the host: A/B such a kernel change in ONE process, configurations interleaved.
- A file-based `dotnet run` reuses its cached NumSharp.Core build: --no-cache after every Core edit (a NaN fix
  "failed" its first re-check on the stale build).

DOCS: .claude/CLAUDE.md (new U5 section + Core Source Files row), docs/plans/numpy-polynomial.md (U5 DELIVERED block,
summary line), test/NumSharp.Tests.Oracle/Fuzz/README.md (polyvander tier + regenerate line),
benchmark/polynomial/.gitignore (numpy_vander_times*.tsv).
…puted (BigInteger points -> per-dimension numbers, mixed-dtype outer products), char degree text, typed-collection / nested flat-list coercion in one pass (List<double> 0.70x -> 13.6-18.3x NumPy); polyvander oracle 31,180 bit-exact, 522 bench cells min 1.53x geomean 10.22x NumPy

A probe replayed every C# argument kind of the house boundary map through the Vandermonde facades against NumPy 2.4.2,
binding each call to the overload C# overload resolution picks ({p}vander's int overload for int / short / sbyte /
byte / ushort / char degrees, the object overload for every other kind): 15,366 cases -
  x / y / z points: typed arrays of every dtype, Memory<T> / ReadOnlyMemory<T>, typed 2-D arrays, object[,], jagged
    arrays, NDArray[], List<T>, LINQ, ValueTuple / System.Tuple, every C# scalar width, char, Half, BigInteger;
  degrees (1-D) and degree containers (2-D / 3-D) of every kind, well-formed and not;
  cross-kind point stacks (typed + NDArray, list + tuple, scalar + 0-d array, ...), error-order pairs;
  the object stack of scalars (below) with every partner kind in both orders.
14,514 exact (dtype, shape, C/F/OWNDATA flags, bytes or NumPy's error type + text), 852 documented refusals (None /
str points: NumPy builds a str / object array and raises at its `+ 0.0`), and three gaps, all closed here.

1. THE OBJECT STACK OF SCALARS (NDPolyVander.ObjectScalarPoints / StackScalars / MixedDtypeVander)
   A Python int past uint64 (C# BigInteger 2**70, -2**70, 2**64, ...) among vander2d / vander3d SCALAR points makes
   NumPy's np.asarray((x, y[, z])) an OBJECT array - and NumPy still computes a numeric matrix: the `+ 0.0` runs per
   element in Python, and tuple(...) hands every dimension its OWN number:
     a Python int      -> float(int), correctly rounded (2**64 + 2**11 + 1 rounds UP to 0x43f0000000000001; .NET's
                          BigInteger -> double cast truncates), CPython's OverflowError "int too large to convert to
                          float" past the float range (the tie 2**1024 - 2**970 rounds up to 2**1024 and overflows;
                          2**1024 - 2**970 - 1 rounds down to the largest double)
     a bool / float    -> a Python float; a complex -> a complex
     an np.float16     -> np.float16 (NEP 50: + a Python float keeps float16); np.uint16 (char) -> np.float64
     a 0-d array       -> a NumPy scalar of (a + 0.0)'s dtype (float32 0-d -> np.float32, int8 0-d -> np.float64)
   NumSharp refused this. Now each element's `+ 0.0` runs through PolyNumber.Binary (Python / NumPy scalar semantics,
   left to right, so the first overflowing element raises before a later None is looked at). One dtype for every
   dimension (float64 whenever only Python numbers are involved) -> the numbers are stacked and the kernel path runs,
   bit-exact incl. NumPy's NaN priority. Mixed dtypes -> each dimension's {p}vander of its own number (a float16 matrix
   computed in float16), multiplied V_k[..., None*k, :, None*(n-1-k)] in NumPy's order through np.multiply's promotion
   (float16 x float64 multiplies in float64 after the float16 matrix; 3-D's first product rounds to its own promoted
   dtype first), each product allocated where NumPy allocates it, the (1, R) result a view (OWNDATA false) as NumPy's
   reshape leaves it. A None / str element and every non-scalar object stack (NumPy computes Python-object matrices)
   stay refused.
   Oracle: gen_polyvander section J (14,568 cases, in its own pass after A-I so every earlier case id is unchanged):
   12 huge ints (exact, rounding up, rounding down, overflowing, the tie) x 23 partner kinds (NaN, -inf, -0.0,
   complex NaN, np.float16 NaN / -0.0, 0-d float32 / int8 / bool / float16 / complex / uint64-max / NaN arrays,
   another huge int) x both orders x 3 degree pairs + 3 3-D arrangements, and the degree errors after the stack.
   _pv_refused also skips NumPy's "can only concatenate str" TypeError (a str element of an object stack).
   Planted-bug check: a truncating int -> float conversion turns 3,132 of the 31,180 cases red.

2. CHAR DEGREE CONTAINER TEXT (NDPolySeries.PythonTypeName)
   vander2d(x, y, 'c') said "object of type 'Char' has no len()"; NumPy's analog of NumSharp's char scalar is
   np.uint16: "object of type 'numpy.uint16' has no len()". PythonTypeName now maps char -> numpy.uint16 (shared by
   the other len() / iterable / subscriptable texts of the package).

3. TYPED COLLECTIONS AND NESTED FLAT LISTS IN ONE PASS (NDPolySequence - every polynomial unit)
   A List<double> of points went through the np.array coercion walk item by item (~40 ns an item: a 1000-point
   List<double> x took 40 us, 0.70-0.97x NumPy's whole {p}vander; an object[] of the same values took 2-3 us through
   the top-level TryFlatScalars fast path). The coercion now reads a TYPED COLLECTION - any non-array IEnumerable<T> of
   a NumSharp dtype: List<T>, HashSet<T>, Queue<T>, LinkedList<T>, LINQ - from its values in one pass
   (TryTypedCollection: a per-type cached materializer, one ICollection<T>.CopyTo or one enumeration;
   CollectionArray: the dtype array coercion DISCOVERS for the scalars its elements are - bool bool, every integer
   primitive int64, a ulong list uint64 past int64 and FLOAT64 for a mix of both magnitudes (NumPy's
   np.array([2**64-1, 1]) is float64), float / double float64, Complex complex128, Half / char / decimal their own; an
   EMPTY collection stays the walk's: no dtype vote, float64 (0,)). NESTED flat lists - a typed collection, or an
   object[] of Python floats / of Python ints within int64 (the top level's fast path, now also inside the 2-D / 3-D
   point stacks and N-D series rows) - become one array leaf in the walk (TryFlatLeaf / FlatLeaf): exactly the walk's
   two shape updates (the sequence's own, then ONE scalar leaf - the other n-1 are the same comparison) and the same
   values (every leaf dtype converts into the promoted dtype exactly, or with the one correctly rounded int64 / uint64
   -> float64 step the per-item conversion takes). Verified: 34 typed-vs-object[] cases (dtypes, nesting, ragged texts,
   empties) identical; the U1-U5 oracle tiers (polyeval / polyseries / polycalc / polyalgebra / polyvander) and 178
   polynomial unit tests green.
   Perf: 1000-point List<double> x 13.6-18.3x NumPy (was 0.70-0.97x); 1000-point List<double> 2-D points 8.9-11.7x;
   1000-point object[] 2-D points 7.7-10.1x.

EVERYTHING ELSE MATCHED
  typed arrays / Memory<T> / typed 2-D arrays / object[,] / jagged / NDArray[] points; short / byte / char degrees
  binding the int overload (a char degree is its code point, np.uint16's __index__); every non-integer degree's
  format(deg, '') text; 5-D / 6-D transposed and reversed points (NDIter copy) bit-exact with NumPy incl. flags; every
  result kind writeable (NumPy's are). Concurrency: 576 kernel first-compiles under 16-way contention and 16 x 400
  contended steady-state calls (one thread flipping the thread-static test hooks) all matched the single-threaded bytes.
  Found on the way, NOT fixed (library-wide, outside this family): np.multiply of two same-shape float64 arrays keeps
  the FIRST operand's NaN where NumPy keeps the second's (the oracle compares NaNs as tokens); the vander kernel imposes
  NumPy's priority itself, and the mixed-dtype object-stack path's broadcast multiplies matched NumPy's NaN bits.

TESTS
  test/NumSharp.Tests/Polynomial/PolynomialVanderArgumentKindsTests.cs (8, new): typed arrays / Memory<T> as ndarrays
  of their dtype; List<T> / LINQ / ValueTuple / System.Tuple / object[,] / jagged / NDArray[] as lists; the
  typed-collection coercion vs the object[] walk (ulong magnitude mixes, empties, nesting, ragged texts); degree kinds
  binding both overloads; non-integer degree texts; degree containers of every kind (incl. the char text); the object
  stack with NumPy 2.4.2 dumps (legvander2d(2**70, 1.5), legvander2d(2**70, np.float16(0.5)),
  chebvander2d(2**70, 0.5-1j), hermvander2d(2**64+2**11+1, 0.0) rounding up, chebvander3d(2**70, np.float16, 0-d f32),
  chebvander3d(np.float16, np.float16, 2**70) - dtype, shape, flags, bytes); its errors in NumPy's order.
  Oracle floor polyvander.jsonl 16,500 -> 31,000 (31,180 cases).

PERF - benchmark/polynomial/polyvander_{numpy.py,bench.cs,report.py}, polyvander_results.md re-rendered
  522 cells (480 + 42 argument-form cells: typed double[] x, List<double> x and 2-D points, a 0-d array degree,
  1000-point object[] 2-D points, objstack2d, objstack3d_mixed), every result SHA-256-checked before timing, both sides
  pinned to 0xF0: min 1.53x (the six page-fault-bound (100000, 121) 2-D cells, 1.53-1.67x: ~10.4 ms of each side's
  ~14-23 ms is demand-zero faults above the 64 MiB pool cap), geomean 10.22x.
    D 1-D f64 min 2.31 geo 16.1 | T dtypes 1.59 / 7.9 | L layouts 6.92 / 8.4 | N N-D 7.19 / 9.3 | V 2-D 1.53 / 11.0
    W 3-D 2.09 / 12.6 | M mixed 1.87 / 4.9 | S small 5.60 / 10.5 | A argument forms 2.98 / 8.5
  objstack2d 5.4-8.1x, objstack3d_mixed 3.1-4.2x, typed double[] 5.5-10x, 0-d array degree 6.7-12.8x.

VERIFICATION
  NumSharp.Tests (CI filter) net10.0 17,880 passed / 11 skipped, net8.0 17,879 / 11; NumSharp.Tests.Oracle net10.0 and
  net8.0 237 passed / 1 skipped each (FuzzMatrix incl. Polyvander 31,180, the leak and native-allocation gates).

DOCS: .claude/CLAUDE.md (U5 wholeness notes + two traps: a lone object[] argument IS a params object[] array; the
np.multiply NaN finding; the C# map's one-pass lists), docs/plans/numpy-polynomial.md (U5 wholeness block, perf table),
test/NumSharp.Tests.Oracle/Fuzz/README.md (section J), facade exception docs (OverflowException for the object stack).
…s; NumPy's statements over U2's arena with an int64 ramp IL kernel and a write-once zero matrix, eigvals' float32 complex result now NumPy's complex64 values; polyroots oracle 6,681 + host-pinned 1,242 bit-exact, 229 bench cells geomean 4.65x NumPy

Plan unit U7 of docs/plans/numpy-polynomial.md: polycompanion/chebcompanion/legcompanion/lagcompanion/
hermcompanion/hermecompanion and the six {p}roots - 12 names, the package now at 172 of 193. Facade members on
np.polynomial.{polynomial,chebyshev,legendre,laguerre,hermite,hermite_e} ([NDScoped], object c, NumPy's single
parameter), fully documented (errors in NumPy's order, the backend contract, the object-series divergence).

WHAT NUMPY DOES (numpy/polynomial/*.py, 2.4.2)
- {p}companion(c): `[c] = as_series([c])`; fewer than two terms -> ValueError "Series must have maximum degree of at
  least 1."; two terms -> the 1x1 matrix of the linear root (scalarmath on NumPy scalars: -c0/c1, lag 1 + c0/c1,
  herm -.5*c0/c1); otherwise mat = np.zeros((n, n), c.dtype), diagonals assigned through mat.reshape(-1)[k::n+1]
  (top = super, bot = sub, mid = main) and ONE in-place update of the last column:
    poly   bot = 1                                   mat[:, -1] -= c[:-1]/c[-1]
    cheb   scl = [1, sqrt(.5)...]; top/bot           mat[:, -1] -= (c[:-1]/c[-1])*(scl/scl[-1])*.5
    leg    scl = 1/sqrt(2*arange(n)+1); top/bot      mat[:, -1] -= (c[:-1]/c[-1])*(scl/scl[-1])*(n/(2n-1))
    lag    top/bot = -arange(1,n), mid = 2.*arange(n)+1.   mat[:, -1] += (c[:-1]/c[-1])*n
    herm   scl = cumprod([1, 1/sqrt(2.*arange(n-1,0,-1))])[::-1]; top = sqrt(.5*arange(1,n))
                                                     mat[:, -1] -= scl*c[:-1]/(2.0*c[-1])
    herme  the same without the 2.                   mat[:, -1] -= scl*c[:-1]/c[-1]
- {p}roots(c): the same as_series; < 2 terms -> np.array([], dtype=c.dtype); 2 -> the 1-element linear root;
  otherwise np.linalg.eigvals of the companion - ROTATED [::-1, ::-1] for every basis but the power series - sorted
  in place. Real roots come back as a VIEW of eigvals' complex result (w.real kept by astype(copy=False)): not
  contiguous, OWNDATA False, stride 16 bytes.

THE DTYPE RULE (the part that decides the bits)
The helper vectors (scl, top, mid) are FLOAT64 (np.sqrt of an int64 ramp, Python floats), so wherever one meets the
series the ufunc loop is result_type(c.dtype, float64): a float16/float32 series' cheb/leg/herm/herme last column is
computed in float64 and rounded ONCE into the matrix by the in-place op's same_kind output cast; the diagonals are
float64 values cast by setitem; poly/lag stay in the series dtype throughout (only Python ints meet them). Complex
series run complex128 loops (the house simd_cmul, operand order kept); NumSharp's decimal runs the same statements in
decimal. Verified BEFORE any code was written: an explicit Python model of exactly these statements matched NumPy on
9,600/9,600 random and special-value series (6 bases x float16/float32/float64/complex128, raw bytes incl. NaN
payloads).

ENGINE
- src/NumSharp.Core/Polynomial/Package/NDPolyAlgebra.Roots.cs (new partial of U2's NDPolyAlgebra, on its PolyArena):
  Companion/Roots entry points, LinearRoot (PolyNumber scalarmath), ScalarArray (np.array([[x]]): fresh owning),
  ObjectTrimLength/IsPythonZero, CompanionMatrix (the six statement lists), and one-kernel-call statement helpers:
  LeadingQuotient, Ramp, Converted (GetPolyCastKernel), SqrtInPlace (GetUnaryKernel Sqrt), CumProd
  (GetCumulativeKernel - np.multiply.accumulate's sequential scan), OpPyFloat/PyFloatOp (weak Python floats),
  Promoted/OpPromotedScalar (NEP 50 loop dtype), LastColumnUpdate (strided gather in the loop dtype, the house op, the
  output cast, the strided store), StoreDiagonal/StoreStrided (GetDiagWriteKernel; a stride-0 source broadcasts).
  No per-dtype code; every loop is an IL kernel.
- src/NumSharp.Core/Backends/Kernels/Direct/DirectILKernelGenerator.PolyRoots.cs (new): the int64 ramp kernel
  (np.arange's role: dst[i] = start + i*step as a running sum, signed down-counter so n <= 0 writes nothing).
- Rotation: np.flip(m) over both axes IS NumPy's [::-1, ::-1] view (negated strides, shifted offset) without parsing
  a slice string (the string slice cost ~0.3-0.5 us of a degree-10 root).
- WRITE-ONCE ZERO MATRIX: np.zeros via new NDArray(fillZeros: true) = calloc's lazily zeroed OS pages; the diagonal
  writes touch every 4 KB page, one demand-zero fault each (~0.9 us): a degree-200 companion took 75 us, ~70 of them
  faulting. Now np.tri's policy (np.PrefersWriteOnce): up to 64 MiB a pooled buffer cleared by np.ZeroBytes, above it
  the OS-zeroed pages (only the touched ones fault).
- Object series: as_series continues in the object dtype and companion's NEXT statement is the length check on the
  TRIMMED object array - so ObjectTrimLength applies Python's `item != 0` from the end: [None, 0], [None, 0.0, -0.0],
  [2**70, 0] raise NumPy's ValueError; past it NumPy computes with Python objects -> NotSupportedException (the
  deferred refusal). Roots of every object series are refused ([Misaligned]: NumPy returns np.array([],
  dtype=object) for one term).
- Errors in NumPy's order: as_series' empty / not-1-d / no-common-type (bool, str) texts; the degree check; for roots
  of degree >= 2 eigvals' own order - finiteness (LinAlgError) before dtype (float16/decimal TypeError) - all before
  LAPACK, so no backend is needed; with none, degree >= 2 raises MissingBackendException.

LIBRARY-WIDE FIXES THAT RODE ALONG
- np.linalg.eig / eigvals of a FLOAT32 operand with complex eigenvalues returned double-precision values; NumPy
  returns complex64 (geev in double, astype(complex64) rounds every component). CollapseEig now rounds each component
  to float32 keeping complex128 (#569) - np.linalg.eig.cs RoundComponentsToSingle: one cast + copy through a float64
  view of a contiguous result, the per-lane form for any other layout. 113 of 3,000 float32 probe roots differed
  before. The live EigLiveParityTests used [[1,-1],[1,1]] (1+-1j, float32-exact) and could not see the gap; it now
  also compares [[0,-2],[1,0]] (+-i*sqrt2) eigenvalues, eigenvectors and eigvals byte-for-byte with live NumPy.
- np.linalg AssertFinite (eig/eigvals' _assert_finite) answers isfinite(a).all() with the fused single-pass
  FiniteScan.IsAllFinite (the asarray_chkfinite kernel; same predicate per dtype and every layout, the rotated
  companion's negative strides included) instead of a bool temp + reduction: 0.45 of a 10x10 eigvals' 0.9 us wrapper.

ORACLE (bit-exact, 0 excused)
- test/oracle/gen_oracle.py gen_polyroots + mode `polyroots` (requires OPENBLAS/OMP/MKL_NUM_THREADS=1 set before
  Python starts): polyroots.jsonl 6,681 portable cases (every companion + the roots that never reach LAPACK: the
  constant/linear short cuts, float16's linalg TypeError, a non-finite companion's LinAlgError, every special-value
  pair of the linear scalarmath) and polyroots_parity.jsonl 1,242 cases (the roots that run geev) + its
  polyroots_parity.host.jsonl pin - host-pinned like linalg_parity (MatmulParityPin, threads=1; Inconclusive off the
  pinned host). Sections A dtype x length, B full-mantissa (+ a 1000x smaller leading coefficient), C trim/special
  patterns, D layouts (strided/reversed/offset/0-d/stride-0 broadcast), E Python-typed and object series
  (_pa_object_land keeps only outcomes NumPy reaches before object arithmetic), F errors in NumPy's order, G roots of
  known polynomials ({p}fromroots: distinct/repeated/clustered/wide/Wilkinson-10/conjugate pairs) + float32 series
  with complex roots (NumPy's complex64, compared by value up-cast), H result flags, I long series, J linear
  scalarmath x special values; Char = the uint16 weave in both files. The host split is RECORDED, not predicted:
  _PREigRecorder replaces np.linalg.eigvals (the modules look it up at call time) and marks a case LAPACK-bound
  exactly when eigvals got a finite float32/float64/complex128 matrix. Results whose sort order is not fixed by value
  (+0.0/-0.0 ties: NumPy's SIMD sort decides by CPU) are skipped - none occurred.
- test/NumSharp.Tests.Oracle: OpRegistry.PolySeries.cs dispatches {p}companion/{p}roots in all six modules (through
  CalcFacet: the flags facet records [C, F, OWNDATA]); FuzzCorpusTests gains Polyroots + PolyrootsParity
  ([DoNotParallelize], the LinalgParity pattern) with floors 6,600 / 1,200; Fuzz/README.md documents the tiers.
- coverage/oracle_map.json host_pins: polyroots_parity.jsonl AND the U2 polyalgebra_parity.jsonl that was missing
  (both "pinned-blas"); coverage/test_oracle_evidence.py 15/15.
- Planted bugs (built + replayed, then reverted): cheb's last column in the series dtype -> 21 portable + 5 host red;
  no [::-1, ::-1] rotation -> 695 host red; herm's linear root reassociated -.5*(c0/c1) -> 18 red; no complex64
  rounding in eigvals -> 116 host red.

TESTS
- test/NumSharp.Tests/Polynomial/PolynomialRootsTests.cs (14): NumPy's TestCompanion + test_{p}roots x 6 ported
  (polyroots' 1,000 large-root round trips included), byte dumps per basis x float32/float16/complex128/int64,
  linear-root scalarmath dumps, errors in NumPy's order without a backend (float16 NaN reports the LinAlgError
  before the TypeError), the object length check, [Misaligned] object roots, the backend-missing path, the roots'
  view flags (stride 16), decimal/char, both allocation paths (a NaN-dirtied pooled buffer and a 67 MB OS-zeroed
  matrix), and the ramp kernel's non-positive counts.
- test/NumSharp.Tests/Backends/LapackEigTests.cs: Eig_Float32ComplexResult_CarriesNumPysComplex64Values (exact
  0x3ff6a09e60000000 = (double)(float)sqrt2, eigenvector components float32-representable, a float64 control).
- test/NumSharp.Tests.Interop/EigLiveParityTests.cs: the +-i*sqrt2 float32 matrix, live NumPy, byte-exact.
- Full suites: NumSharp.Tests net10.0 17,895/0 failed, net8.0 17,894/0; Oracle net10.0 + net8.0 239 passed,
  1 skipped (RandomApiSoak, opt-in); interop Eig/PolynomialLiveParity 21/21.

PERF (NPY/NS, higher = NumSharp faster; benchmark/polynomial/polyroots_{numpy.py,bench.cs,report.py}, committed
summary polyroots_results.md; both sides pinned to 0xF0, never concurrent, best-of-rounds, every cell SHA-256-checked
before timing; NumSharp with NumSharp.Interop.OpenBLAS at one thread)
  229 cells, all bit-exact, geomean 4.65x:
  companion float64 2-2898 terms  min 0.89 geomean 7.69 | companion float32/float16/complex128/int64  1.42 / 6.77
  roots float64  0.98 / 2.65 | roots float32/complex128/int64/float16  0.99 / 1.84 | layouts 1.93 / 6.02 |
  Python lists 1.88 / 4.93
Every cell under 1.5x is bound by work both sides do identically (probed):
- roots of degree 50 and 200: LAPACK geev, the same scipy-openblas call (~1.0x);
- degree-10 complex roots: zgeev is 20.7 of NumSharp's ~22 us, NumPy's eigvals alone 23.8 us -> ceiling 1.43x
  (measured 1.35-1.5x); degree-10 float32 roots went 15.1 -> 13.1 us (1.5-1.9x) after the view-based complex64
  rounding, np.flip and FiniteScan;
- degree-200 companion: zeroing the 320 KB matrix - NumSharp's alloc+zero 4.1 us vs NumPy's np.zeros 3.9 us - plus
  0.4-0.8 us vs NumPy's ~3.3 us of Python statements: 1.27-1.63x run to run (the pinned CPUs 4-7 are two physical
  cores whose hyperthread siblings desktop load shares; the same call swung 4.47-5.48 us within one process);
- the 67 MB power-series companion: one demand-zero fault per touched page on both sides (~1.0x).
Measured on the way: a hot 312 KB buffer clears at ~86 GB/s whatever the method (Span.Clear = InitBlock = an AVX2
loop, 3.6 us; non-temporal stores 5.9 us); a fused row-by-row build (clear a row, write its nonzeros) was slower than
clearing then writing; OpenBLAS being loaded does not change the companion's time (A/B).

TRAPS
- Decide a host-pinned split by RECORDING (patch the function NumPy calls), not by predicting from shapes/dtypes.
- A conditional with an object[] arm and an NDArray arm is typed NDArray through NumSharp's implicit Array -> NDArray
  conversion (the benchmark runner crashed on its list cells until both arms were cast to object).
- cmd's %1 splits on commas: `bench_only.cmd C/poly/,T/poly/` passes only C/poly/.
- Bash heredocs lose one level of backslash escaping: a Python snippet's "\\f" became a form feed and silently kept
  the wrong filter file; write such scripts with the Write tool.
- Inside namespace NumSharp.Tests.*, `Math.Pow` binds to NumSharp.Tests.Math - spell System.Math.
- testhost PID 173468 (a long-lived earlier test host, left alive) burns ~0.8 CPU on CPU 16 - ambient load.

Docs: .claude/CLAUDE.md (U7 section + Core Source Files row + the BLAS/LAPACK table's complex64 note),
docs/plans/numpy-polynomial.md (status row + U7 delivered block), test/NumSharp.Tests.Oracle/Fuzz/README.md,
benchmark/polynomial/.gitignore.
…mPy's error, library-wide reinterpreting-alias use-after-free fixed (Forward ARC disposer), shared coercion stops throwing for Python ints past uint64; polyroots 8,195 + parity 1,354 bit-exact

A wholeness pass over plan unit U7 ({p}companion / {p}roots x 6 bases, faa1e2b9). Every C# argument kind, edge
series, flag and concurrency behaviour was validated against NumPy 2.4.2. It closed three object-series gaps, ONE
library-wide memory-safety defect the concurrency check exposed, four undisposed internal aliases the fix made
visible, and a throw/catch in the shared coercion that held the new object path under 1.5x NumPy.

==========================================================================================================
1. The probe
==========================================================================================================

A 3,096-case probe was generated against NumPy 2.4.2 and replayed through the facades. It covered:
- every C# argument kind x both functions x six bases: typed arrays, Memory<T>, ValueTuple, object[] / List, jagged
  and NDArray[] series, Half / char / decimal items, 0-d arrays, BigInteger items;
- long series up to 1,100 terms (Hermite's scl underflow, float16 rounding past 2048);
- edge numeric series: +-inf / NaN leading or interior terms, all-NaN, tiny / huge leading terms, subnormals,
  int64 / uint64 / int8 extremes, complex non-finite components, signed-zero trims.

Result after the fixes below: 2,982 exact, 0 wrong, 114 refused. Every refusal is in the documented object-dtype
classes: a one-term object series' roots, the object companion matrix, CPython TypeErrors of None / str arithmetic.

==========================================================================================================
2. Object series (a Python int past uint64, None or another object among the terms)
==========================================================================================================

NumPy's as_series goes on in the OBJECT dtype, and the items keep their kinds: Python numbers, NumPy scalars, 0-d
ARRAYS. At delivery NumSharp refused every object series past the length check. Now:

- TWO TERMS compute (ObjectLinearRoot). NumPy's linear statement (-c[0]/c[1]; lag 1 + c[0]/c[1]; herm -.5*c[0]/c[1])
  runs on the ITEMS through PolyNumber: CPython arithmetic for Python numbers (int / int correctly rounded, a Python
  complex quotient's +0.0 real part), scalarmath for NumPy scalars, ufuncs for 0-d arrays. np.array then makes a
  NUMERIC array of that one number's dtype. Examples: [2**70, np.float16(1.5)] -> float16 -inf (the Python int
  converted to float16 overflows); [2**70, np.array(1.5, float32)] -> float32. A Python int past the float range
  raises CPython's OverflowError: `integer division result too large for a float`, or `int too large to convert to
  float` for herm, whose -.5 * c[0] converts first.
- TRIMMING compares each item with 0 (Python's `item != 0`). A 0-d array item compares elementwise, so a trailing
  np.array(0.0) is trimmed (IsPythonZero).
- THREE OR MORE terms: NumPy builds an OBJECT companion matrix. {p}companion returns it (still refused: no object
  dtype). {p}roots hands it to eigvals, whose _assert_finite raises `ufunc 'isfinite' not supported for the input
  types, and the inputs could not be safely coerced to any supported types according to the casting rule ''safe''`
  (TypeError) BEFORE LAPACK, so no backend is needed. Unless the last-column arithmetic raises first: NumPy's object
  ufunc loops call Python's operator element by element in order, and with numeric items the only possible error is
  a Python int too large for a float. ObjectCompanionArithmetic replays exactly the element operations that can raise,
  in NumPy's order:
    poly / cheb / leg / lag   c[:-1] / c[-1]                 (what follows multiplies floats)
    herm                      scl * c[:-1], then 2.0 * c[-1]  (the quotient divides floats by a nonzero float)
    herme                     scl * c[:-1], then ... / c[-1]
  Two NumPy details decide the TEXT:
  - The divisor c[-1] is CAST to the object dtype before the loop (ObjectLoopValue). A NumPy scalar or 0-d array
    becomes its Python value, so `huge / np.int64(3)` inside the loop is Python int / Python int (the division text),
    where the bare expression gives the conversion text.
  - The float64 helper scl's items are Python floats, and their values cannot change whether a product raises.
  Probed over a 25-series grid x 6 bases x 2 functions: all 200 comparable outcomes exact, 88 in the documented object
  classes (np.int64 has no C# spelling; the 0-d array stands in).
- Still refused (NotSupportedException): a one-term object series' roots (NumPy's np.array([], dtype=object),
  [Misaligned]), the object companion matrix itself, and CPython's TypeErrors from None / str arithmetic (the U2
  policy: object arithmetic is not emulated).

==========================================================================================================
3. LIBRARY-WIDE: byte-reinterpreting aliases did not keep their owner's buffer alive
==========================================================================================================

The concurrency check (16 threads against single-threaded bytes; NumPy showed 0 differences under the same load)
found eigvals / {p}roots results changing. The cause was SINGLE-threaded and library-wide:
- Every NDArray holds one counted reference on its storage's slice (TryAddRef at construction, Release on Dispose,
  Abandon from the finalizer), and an ordinary view shares its owner's slice.
- A view that REINTERPRETS the bytes as another element type needs its own slice: np.real / np.imag / .real / .imag,
  view(dtype), getfield. That slice was a plain non-owning wrap (ArraySlice.Wrap, an immortal Disposer.Null), so
  disposing the owner returned the buffer to the pool UNDER a live alias. The alias then read whatever the pool handed
  out next.
- np.linalg.eig / eigvals' [NDScoped] wrapper disposes the complex result w while returning w.real. So the next
  same-size allocation (the next eigvals or {p}roots call) overwrote the first result, and under threads one call read
  another thread's eigenvalues. Reproduced single-threaded: r1 = eigvals(diag(2,3)); eigvals(diag(7,9)) turned r1 into
  [7 9]; np.real(z) after z.Dispose() read the next allocation.

FIX (UnmanagedMemoryBlock`1.cs, ArraySlice.cs, UnmanagedStorage.{Cloning,GetField}.cs):
- A new AllocationType.Forward disposer for blocks that alias another array's memory. It owns and frees nothing, and
  forwards TryAddRef / Release / Abandon / IsReleased / IsUniquelyReferenced to the OWNER's slice. An alias of an alias
  reaches the block that owns the memory.
- ArraySlice.WrapShared<T>(address, count, arcOwner) builds such a slice, used by the three alias sites:
  WrapReinterpreted (view(dtype)), AliasComplexLane (np.real / np.imag) and GetFieldAlias (getfield).
- Observable consequences, both NumPy's: the owner's refcount counts each live alias, so ndarray.resize(refcheck)
  refuses while one lives; and the buffer returns to the pool when the LAST of owner and aliases is released.
- Verified: 16 threads x (366 first calls + 6,400 steady-state calls), every stage (companion, the rotated flip copy,
  eigvals over every layout): 0 differences.

COROLLARY: an INTERNAL alias must now be disposed, because an undisposed one holds the owner's buffer out of the pool
until the finalizer. The leak gates (Catalogue_EveryEntry / EveryPropertyAndField_Read_LeavesNoUndisposedIntermediates)
caught three members, and review found a fourth of the same shape:
- the `real` / `imag` SETTERS (np.real(this) / np.imag(this) as the copy target): [NDScoped] (`this` was not created in
  the scope and is never touched);
- `setfield` (its getfield view and a value converted there): [NDScoped];
- `view<T>()` and `getfield<T>()` wrapped the alias in an untyped NDArray and called AsGeneric<T>(), which builds a
  SECOND NDArray over the storage. Both now construct the typed array directly (getfield<T> also kept losing its
  engine that way).
Internal aliases inside [NDScoped] np.* functions (eig, isreal / iscomplex, angle, frombuffer, the char reductions)
were already reclaimed by their scopes; the 200K-case corpus sweep stayed green throughout.

==========================================================================================================
4. Perf: the shared coercion threw and caught for every Python int past uint64
==========================================================================================================

The new object value path first measured 1.13-1.66x NumPy. NDPolySequence's leaf classification and
NDPolySeries.AsCoefficientArray's scalar path asked PolyNumber.DiscoveredDtype(), which THROWS for a Python int past
uint64, and caught it: ~1.4 us per leaf. Now PolyNumber.TryDiscoveredDtype classifies without throwing, and
ObjectArrayRefusal builds the same NotSupportedException unthrown, to be raised only where NumPy computes with Python
objects. DiscoveredDtype keeps its contract on top of the two. Every unit's object-series path shares this coercion.
NumPy 2.4.2 vs NumSharp, both pinned 0xF0, best-of-9 x 20,000 calls:

  polyroots([2**70, 1])                         2.344 vs 0.696 us   3.37x   (was 1.13x)
  polycompanion([2**70, 1])                     2.447 vs 0.694 us   3.53x   (was 1.18x)
  hermroots([2**70, np.float16(1.5)])           3.455 vs 0.749 us   4.61x   (was 1.66x)
  lagroots([2**70, np.array(1.5, float32)])     3.987 vs 1.364 us   2.92x   (was 1.52x)
  polyroots([2**70, 1, 0, 0.0])                 2.731 vs 0.812 us   3.36x   (was 1.27x)

==========================================================================================================
5. Oracle (gen_oracle.py polyroots; each section is its own pass after A-J, so every earlier id is unchanged)
==========================================================================================================

- polyroots.jsonl 6,681 -> 8,195 (floor 7,800 -> 8,150); polyroots_parity.jsonl 1,242 -> 1,354 (floor 1,350).
- Section K, 1,142 portable + 112 host-pinned cases:
  - two-term object series NumPy still computes (Python ints, floats, complex, bools, np.float16, 0-d arrays of several
    dtypes, tuples), trailing 0-d / -0.0 / False trims, CPython's int / int OverflowError;
  - edge numeric series for float64 / float32 / float16 / complex128 / int64 / uint64 / int8 / uint8 / int32 / uint16.
- Section L, 372 cases: object series of three or more terms. 236 OverflowErrors, both texts, from companion and roots;
  the rest eigvals' isfinite TypeError from roots. Object results and None / str TypeErrors stay filtered
  (obj3_object_land).
- Planted bugs:
  - K: object linear root refused (companion / roots), 0-d items never trimmed: 210 / 210 / 54 red;
  - L: roots refused instead of the TypeError 104; arithmetic skipped (roots / companion) 118 / 118; divisor not cast
    24; Hermite on the division path 34; 2.0 * c[-1] skipped 8 (2 before three leading-term series were added); only
    the first element checked 24.

==========================================================================================================
6. Tests (NumPy-probed literals)
==========================================================================================================

- PolynomialRootsTests (+4):
  - ObjectSeries_TwoTerms_ComputeTheLinearRootOnTheItems: bytes per basis for Python int, np.float16, 0-d float32 and
    Python complex second terms, plus the trailing 0-d zero trim;
  - ObjectSeries_TwoTerms_TooLargeForAFloat_IsCPythonsOverflowError: both texts by basis;
  - ObjectSeries_ThreeOrMoreTerms_RootsRaiseNumPysErrorInNumPysOrder: the isfinite TypeError, eleven raising series x
    both texts, Hermite's leading-term cases, None / str refusals;
  - Roots_RealRoots_SurviveLaterCallsAndAllocations.
- LapackEigTests (+1): Eigvals_RealResult_OwnsItsBuffer_SurvivesLaterCallsAndAllocations (eigvals and eig).
- ArcLifecycleTests (+4):
  - ReinterpretingAlias_SurvivesItsOwnersDispose: all 8 alias kinds, same-size thieves allocated after the owner's
    dispose, refcount -1 after the last release;
  - _CountsOnTheOwnersBlock_InEitherDisposeOrder;
  - _OfAnAlias_ForwardsToTheBlockThatOwnsTheMemory;
  - _ResizeRefcheck_SeesTheAlias (NumPy's refcheck text).
- Planted-bug check of the ARC fix: reverting any one alias site to the old wrap turned its tests red (complex lane 6
  tests, view(dtype) 2, getfield 1).

==========================================================================================================
7. Verification
==========================================================================================================

- NumSharp.Tests (CI filter): net10.0 17,904 / net8.0 17,903 passed, 0 failed.
- NumSharp.Tests.Oracle: 239 / 239 on both frameworks, including both leak gates and Polyroots / PolyrootsParity.
- Interop suites (net10.0): Interop 736, OnnxRuntime 159, MLNet 165, ParquetNet 11, System.Numerics.Tensors 54, all
  green. The ARC change touches every path that pins or aliases a buffer.

==========================================================================================================
8. Docs
==========================================================================================================

- .claude/CLAUDE.md: "Critical: View Semantics" gained the alias-ARC rule and its corollary; the U7 section has the
  object-series contract, the wholeness pass, the perf note and new counts.
- docs/plans/numpy-polynomial.md: the U7 wholeness block (probe, gaps, ARC fix, perf table, oracle, mutants) and the
  status counts.
- test/NumSharp.Tests.Oracle/Fuzz/README.md: sections K and L and the counts.
- The six facades' and the engine's XML docs: the object-series exceptions (TypeError / OverflowException /
  NotSupportedException) as they now stand.

TRAPS (for the next session):
- A non-owning wrap (Disposer.Null) is only correct for memory NOTHING disposes. Any slice carved from another
  NDArray's buffer must forward its ARC (ArraySlice.WrapShared), or the owner's Dispose frees memory a live alias
  reads. The symptom is silent: right values until the pool reuses the buffer.
- A corpus that checks each result right after producing it cannot see a use-after-free: the concurrency check and a
  same-size allocation between call and read found it.
- Making something counted surfaces every undisposed intermediate of that kind. Run the leak gates
  (UndisposedIntermediateTests) after any ARC change.
- Exceptions as control flow in a per-leaf classifier cost ~1.4 us each; classify with a Try-method and build the
  refusal object unthrown.
- NumPy's object-dtype ufunc CASTS a non-object operand to object items first (np.int64 -> Python int), so its error
  text can differ from the same expression written bare.
…s nuget.org README tab stays owner-editable for its release notes; Directory.Build.props guard fails dotnet pack if PackageReadmeFile returns

WHY
nuget.org gives a published version exactly one place that can still change: the README tab, and only
when the package does NOT embed a readme. Established 2026-10-01 against the NuGetGallery source and
the nuget.org FAQ:
- The nuspec's <releaseNotes> is immutable once published. The PackageEdits table (which held
  ReleaseNotes/Description/... edits) was dropped by migration 201801101337052_RemovePackageEdits, and
  the FAQ ("Can I edit package metadata after it's been uploaded?") says signed content, nuspec
  included, must stay immutable. The gallery also renders <releaseNotes> as plain text
  (Html.PreFormattedText), so markdown would show raw.
- PackagesController.Edit (POST) now only saves the legacy hosted README of a version (Manage page ->
  Readme), and ReadMeService.GetReadMeMdAsync(Package) ALWAYS serves the embedded file when
  HasEmbeddedReadme is set, so an embedded readme locks that tab for good.
Every 0.70.0 package (NumSharp, NumSharp.Bitmap, NumSharp.Interop.OpenBLAS, NumSharp.Interop.pythonnet,
NumSharp.Build) embeds its project-folder README.md, which is why 0.70.0 can never show its release
notes on nuget.org. 0.40.0-prerelease ... 0.60.0 embedded none, so on 2026-10-01 the owner had the README
tab of NumSharp and NumSharp.Bitmap 0.40.0-prerelease, 0.41.0-prerelease, 0.50.0-prerelease and 0.60.0
set to the GitHub release body of the same tag (done through the Manage page, read back byte-identical:
2,393 / 18,754 / 21,676 / 54,557 UTF-16 units). From the next release on, every package must arrive
readme-less so the owner can do the same by hand after publishing.

WHAT
- The nine packable projects no longer set PackageReadmeFile or pack README.md: NumSharp.Core,
  NumSharp.Bitmap, NumSharp.Interop.{MLNet, OnnxRuntime, OpenBLAS, ParquetNet,
  System.Numerics.Tensors, pythonnet} and tools/NumSharp.Build. Each csproj gets a comment at the spot
  saying the omission is deliberate. The README.md files stay in the repo as GitHub/docs documentation
  (ARCHITECTURE.md and the interop docs link them); they are simply not packed. A readme-less package
  shows its Description on the README tab until the owner submits a readme.
- Directory.Build.props: new "No embedded package README" block documenting the policy, plus the guard
  target NumSharpRejectEmbeddedPackageReadme (BeforeTargets="GenerateNuspec", so it runs on every
  dotnet pack, with or without no-build, and never on plain builds/tests). It errors when
  PackageReadmeFile is non-empty, naming the project and pointing back at the block. It is declared in
  the props, ahead of the project body, yet still sees a project's own property: a target's Condition is
  evaluated when the target runs, after the whole project has been evaluated. Without the guard, a copied
  csproj or an "add a readme" best-practice fix would silently lock the next release's README tab, and
  that only shows after an irreversible publish.
- .claude/skills/changelog/SKILL.md: new "After the release publishes - put the notes on nuget.org"
  section with the manual owner step (Manage -> Readme -> version -> Custom -> paste
  `gh release view v<version> --json body --jq .body` -> Submit; read-back via
  /packages/<id>/<version>/GetReadMeMd; an empty submit removes it).
- .claude/CLAUDE.md CI Pipeline: one-line pointer to the policy and the manual step.

VERIFIED (detached worktree at 890ff1dd + these files, never the live tree)
- dotnet pack -c Release of all nine projects: rc=0 each; every nupkg's nuspec has no <readme> and no
  README.md entry; <icon> and <license> still present.
- Guard: PackageReadmeFile=README.md passed as a global property AND written back into
  NumSharp.Bitmap.csproj's body each fail dotnet pack (rc=1) with the guard's message, and no
  package is produced.
- No test, script or workflow asserted package READMEs (git grep over test/, tools/, .github/ and the
  verify_*.sh scripts).
…instead of spinning for days - per-case CaseWatchdog over every CorpusFile replay, hang-proof constraint-bound tests; the 84-CPU-hour oracle testhost diagnosed (stale p <= 1 logseries bound)

THE BUG
-------
An oracle test host (testhost PID 173468, parent dotnet test 67924, both started 2026-09-26 21:20:41, running
test\NumSharp.Tests.Oracle\bin\Debug\net10.0) was still spinning on 2026-10-01: 84+ CPU-hours on one thread, no test
failed, no timeout fired, nothing naming what it was doing. A local `dotnet test` has no timeout (CI's test steps have
`timeout-minutes`), so one infinite loop in one corpus case holds the whole run - and a core - forever. It also locks its
build output (bin\Debug), so Debug builds of the Oracle project fail until it is gone.

DIAGNOSIS (dotnet-dump; dotnet-stack prints nothing for a thread spinning in cooperative mode)
---------
- `dotnet-dump collect -p 173468 --type Full`, `analyze`: clrthreads -> managed thread 4 (OSID 0x185a4, the one with
  303,477 CPU-s) is the only Cooperative thread. Its stack:
    MT19937.FillDouble <- DrawBufferDouble.Refill/NextDouble <- NumPyRandom.LegacyLogseries
    <- NumPyRandom.logseries(NDArray, Shape) <- OpRegistry.LegacyBroadcastDraw (OpRegistry.RandomBroadcast.cs:61)
    <- OpRegistry.Apply <- FuzzCorpusTests.CheckError <- RunCorpus("random_parity.jsonl") <- RandomParity().
- `dso` + `dumpobj` on the FuzzCorpus+Case: Id = rnd/logseries/bcast:viol:p>=1/857 - p = [0.5, 1.0], an ERROR case:
  NumPy raises `ValueError: p < 0, p >= 1 or p contains NaNs`.
- `name2ee` + `dumpil`: the loaded logseries(NDArray, Shape) DID call RandomConstraints.CheckArray(p, "p",
  CONS_BOUNDED_LT_0_1). The Debug NumSharp.dll on disk (21:20:38, three seconds before the process started - the loaded
  image) read by reflection: that branch was `AllInRange(arr, 0.0, ldc.r8 0x3FF0000000000000)` = 1.0, i.e. `p <= 1`,
  where `p < 1` is `AllInRange(arr, 0.0, Math.BitDecrement(1.0))`. (dumpil prints ldc.r8 with 6 decimals, so the exact
  literal had to come from the IL bytes.)
- So p = 1.0 passed the check, and legacy_logseries at p = 1 never accepts a draw: r = log(1 - p) = -inf makes
  q = 1 - exp(r*U) = 1, so every candidate floor(1 + log(V)/log(q)) = floor(-inf) is rejected - exactly NumPy's C
  loop, which only NumPy's parameter check keeps out.
- The in-progress build was ~40 minutes before 3a15871a; the COMMITTED code already reads BitDecrement(1.0). HEAD
  verified against NumPy 2.4.2: logseries p in {1, 1+ulp, 2, inf, +NaN, -NaN, -5e-324, -1, -inf} and zipf a in
  {1, 1-ulp, 0.5, 0, -0, -1, -inf, +NaN, -NaN}, x RandomState/Generator x (Python float, 0-d array, arrays with the bad
  value at 7 (length, position) placements across the vector body and the scalar tail): 324 calls, every outcome and
  message identical to NumPy, 0 hangs; F-ordered / strided / broadcast parameters also probed (12 calls, identical).

WHAT WAS STILL BROKEN, AND THE FIX
----------------------------------
The library is correct; the failure mode is not. For an error case whose sampler loops forever on the rejected value
(legacy logseries at p = 1; zipf at a <= 1 or NaN on both APIs), the constraint check is all that separates NumPy's
ValueError from an infinite loop - so a regressed check does not FAIL a test, it HANGS the run, with no name attached.

1. Oracle: test/NumSharp.Tests.Oracle/Fuzz/CaseWatchdog.cs (new) - a per-case deadline for every corpus replay.
   - CorpusFile's iterator (CorpusFile.Parse) opens a slot on the first MoveNext and arms it with each case's id and
     start time as it yields the case; the iterator's finally removes it (completion, early break, parse exception).
     Every streaming replay is covered with no per-loop change: RunCorpus, the Ma and RandomApi replays, the leak sweeps,
     BlasBackendDelta, the nightly soak.
   - OpRegistryRandomIsolationTests replays every rnd case of random_parity(_host) - error cases included - from
     FuzzCorpus.Load (a list), so it arms a slot explicitly per case.
   - A 5 s timer checks the live slots; a case past the limit (2 min default; NUMSHARP_ORACLE_CASE_TIMEOUT_SECONDS
     overrides, 0 disables, unreadable text keeps the default) is reported ONCE and the host is ended with
     Environment.FailFast, whose message `dotnet test` prints as the crash reason. Failing the test instead is impossible:
     the loop runs on the test thread and nothing in .NET stops a running thread. Under an attached debugger the report
     goes to the debugger output + stderr and the process lives (a breakpoint is not a hang).
   - Thread-safety: Slot.Enter stores the start time BEFORE the id (release stores) and the checker reads the id first
     (acquire loads), so a check never pairs a new case with the previous case's start. The limit and the reaction are
     swapped TOGETHER under the lock (Configure) - with two settable properties, a timer tick could pair a test's 1 ms
     limit with the process-ending default and kill the test host mid-run.
   - Contract (documented): a CorpusFile enumeration must be disposed (foreach/LINQ do it); an abandoned enumerator keeps
     its last case armed. Every current use is `using` + foreach (checked).
   - CaseWatchdogTests (4): armed-past-limit reported once naming corpus + case, the next case is a new report;
     within-limit / disposed / disabled report nothing; a CorpusFile enumeration arms exactly the case it handed out and
     an early break or completion removes the slot; ParseLimit's rules.
2. Unit tests (NumSharp.Tests):
   - RandomBroadcastTests.Constraints_LoopGuardingBounds_AreRejectedOnEveryPath (new): the 336 calls above, each must
     raise NumPy's ValueError with NumPy's text; the whole matrix runs on one background thread that publishes the call
     it is in, and the test FAILS naming that call when it does not return within 60 s.
   - [Timeout(60_000)] on both Validation_Seed42_MatchesNumPy data-row tests (LegacyRandomState, Generator.Distributions -
     their rows include logseries(1.0), zipf(1.0), zipf(0.5), zipf(nan)) and on
     Constraints_StrictBounds_AcceptTheNeighbouringDouble_Generator (zipf([1.0, 2.0])). The legacy strict-bounds test is
     not hang-prone (binomial(n=-1)'s inversion returns at once: qn = 2 > U) and is left as it was.

MUTATION PROOF (each planted, run, then the file restored from HEAD after checking the diff was only the mutant)
--------------
- A, the stale build's exact bug (CheckArray CONS_BOUNDED_LT_0_1 bound BitDecrement(1.0) -> 1.0):
  new unit test failed at 60 s: "RandomState.logseries(array n=1, 1 (0x3FF0000000000000) at 0) did not return";
  Generator strict-bounds test failed ("no exception was thrown" - see Log1p below); Oracle RandomParity AND the
  Load-based isolation test, at NUMSHARP_ORACLE_CASE_TIMEOUT_SECONDS=10: host ended after 15 s, `dotnet test`
  reported "Test host process crashed : ... corpus 'random_parity.jsonl', case 'rnd/logseries/bcast:viol:p>=1/857'".
- B, the scalar twin (Check CONS_BOUNDED_LT_0_1 `val < 1` -> `val <= 1`): legacy Validation row logseries(1.0)
  "timed out after 60000ms", Generator row failed fast, new test named "RandomState.logseries(1 (0x3FF0000000000000))";
  run finished in 2.2 min (406/409 passed). Oracle RandomApiHostLibm: host ended after 14 s naming
  'random_api/NumPyRandom.logseries(NDArray,Shape)/RandomState:-:s0:none/var3'.
- No new zombie: after every run only the old 173468 remained (the abandoned spinning threads are background threads;
  each run also had --blame-hang-timeout as a backstop, never needed).

CALIBRATION of the 2-minute default
-----------------------------------
The whole Oracle suite (CI filter, net10.0, leak sweeps included) passed with a 1 s limit checked every 100 ms (period
temporarily shortened, reverted): no committed case runs past ~1.1 s on the dev host, so 2 min is >100x headroom.

FOUND ALONG THE WAY (documented, deliberately not changed)
-----------------------------------------------------------
- Generator.Log1p is bit-exact with the CRT's log1p (np.log1p) for x > -1 - every caller's domain - but not at the edge:
  x = -1 returns NaN (CRT: -inf; u = 0 makes the correction -inf - 0/0) and x < -1 returns .NET's negative NaN (CRT:
  0x7ff8...). Probed: all 11 edge values match np.log1p bit for bit with a `u <= 0` branch, BUT that branch measured
  4.0% (best) / 5.7% (median) slower on the inversion fill's per-element loop (in-process interleaved A/B, 10M doubles,
  identical bits), and that path (standard_exponential(method="inv"), f64) is only 1.19x NumPy (58.7 vs 69.7 ms @10m) -
  so the edge is documented in the helper instead. Consequence: the Generator's logseries at an UNCHECKED p = 1 returns
  2 (r = NaN) instead of looping like NumPy's - unreachable through the checks, and the tests above still catch it.
- Pre-existing perf gap, untouched: Generator.standard_exponential(dtype=float32, method="inv") is ~0.55x NumPy
  (123.7 vs 67.9 ms @10m - double log1p + narrow per element).

THE STALE PROCESS
-----------------
Not killed here (the user's rule: no taskkill without a go-ahead). It runs a pre-3a15871a build and can never finish;
`taskkill /PID 173468 /T /F` ends it (its dotnet test parent 67924 then exits) and releases the Oracle's bin\Debug.

VERIFICATION
------------
NumSharp.Tests (CI filter): net10.0 17,905 passed / 0 failed, net8.0 17,904 / 0. NumSharp.Tests.Oracle: 243 passed /
0 failed on both (239 + the 4 watchdog tests). Release builds (the stale host locks bin\Debug).

Docs: CLAUDE.md "Keeping the suite fast" (a test that can never return must FAIL, not hang; the dotnet-dump recipe),
Fuzz/README.md Gate semantics (the watchdog). Files: test/NumSharp.Tests.Oracle/Fuzz/{CaseWatchdog.cs (new),
CaseWatchdogTests.cs (new), CorpusFile.cs, OpRegistryRandomIsolationTests.cs, README.md},
test/NumSharp.Tests/RandomSampling/{RandomBroadcast.Test.cs, LegacyRandomState.Test.cs, Generator.Distributions.Test.cs},
src/NumSharp.Core/RandomSampling/Generator.Distributions.cs (comment only), .claude/CLAUDE.md.
…n results now in NumPy's layout (NpyIter's, was C), mapdomain over N-D non-C x 0.62x -> 2.13x NumPy (transposed route + storage relabel), {p}sub NaN sign under RyuJIT's (-a)+b -> b-a rewrite; polyeval 23,866 + polyseries 19,478 bit-exact with new strides facets

A review of the numpy.polynomial package as delivered (cb64deb1 .. 890ff1dd: U1 additive family + polyutils,
U2 series algebra, U3 evaluation, U4 calculus, U5 Vandermonde, U7 companion/roots) against NumPy 2.4.2, in an
isolated worktree. It covered the API, parity beyond what the oracle tiers pin, performance on layouts the unit
matrices never measured, and the plan's own claims. Three defects and one below-NumPy performance class were found
and fixed; the plan's D9 and D12 were wrong and are corrected. The full record is
docs/plans/numpy-polynomial.md section 11.

CLEAN (worth not re-proving)
- API: every delivered name has NumPy's parameter names, order and defaults, reflected against 2.4.2's signatures.
  The 27 names NumPy has and NumSharp lacks all belong to undelivered units: U6 fit, U8 gauss/weight, U9
  chebpts/chebinterpolate, U11 format_float, the six classes (U10), and polyvalfromroots.
- Values, dtypes, flags, errors: a fresh random probe, independent of the corpora, ran 19,837 calls over
  U1/U2/U4/U5/U7. The only differences are the NaN-bit class below.
- Threads: 16 and 32 concurrent callers reproduce the single-threaded bytes.
- Decimal / Char work through every unit; decimal roots are refused with NumPy's linalg text.

F1 - {p}val RESULT LAYOUT (U3)
Defect:
- The N-D-series path allocated its result C-ordered whatever its operands were.
- A broadcast x's result was laid out by np.copy's 'K' rule, which sorts a stride-0 axis innermost (F for a row
  broadcast).
- Nothing pinned layouts: the oracle compares bytes in C order, so every value check passed.
NumPy's rule: the result is the recurrence's LAST ufunc output, and NpyIter lays out every ufunc output from that
call's operands:
- npyiter_fill_axisdata: an operand abstains (stride 0) on a broadcast extent of 1, a missing axis, its own extent
  of 1, or a real stride of 0;
- npyiter_find_best_axis_ordering: a stable insertion sort over the reversed axes in which C order wins a conflict
  between operands and an unseen pair is skipped;
- npyiter_new_temp_array: dense strides along that order.
Which copy of c the recurrence reads matters too: chebval copies (copy=True, order 'K'), the other five read c as
given (copy=None), and an integer series converts by astype(np.double) (order 'K'), except in polyval, where it is
`c + 0.0` (NpyIter).
Implementation:
- Shape.NpyIterOutputShape (View/Shape.StridePerm.cs): the three NumPy functions above, ported line for line. The
  vote table and the permutation live on the stack, and the dims array is adopted, so the strides array is the only
  allocation.
- NDPolyEval.SeriesLayout: the layout each basis's copy of c leaves.
- NDPolyEval.ResultShape: the closed forms.
  - scalar x: NpyIter's layout of c[0];
  - a 1-D series: x's layout;
  - an N-D series under tensor: the series block outside (strides x x.size), x's block inside;
  - tensor=False with an N-D series: C when a C-contiguous x or series spans the result exactly (it votes C on every
    pair, and C wins every op that meets it; every basis's last op meets x), else NDPolyEval.ReplayResultShape walks
    the basis's step table (PolySteps.Eval, the table the kernel is emitted from) over layouts until a fixpoint.
  - C fast paths: rank <= 1, C operands, and a <= 2-D series with a C / <= 1-D / scalar x take no allocation.
- Shape.DenseAxisOrder: a dense layout's axes in memory order (stable, so tied extent-1 axes keep their order).
Validation: the closed forms against NumPy on 32,740 random layouts, and the replay on 27,595, both with 0
differences. A layout probe went from 1,302 mismatches to 22, all extent-1 strides, which NumPy assigns per loop
path and no flag reads.

F2 - mapdomain LAYOUT AND SPEED (U1)
Defect:
- An N-D x that was not C-contiguous got a C-ordered result; NumPy keeps x's memory order (NpyIter over x alone).
- The same shapes were BELOW NumPy before this change: a transposed 3-D float64 x 0.62-0.65x, a float32 one 0.47x.
  The fused pass's NDIter walk gathered every element, and the 867-cell U1 matrix had measured 1-D layouts only.
Trap: writing NumPy's layout directly is 2.7x SLOWER. NumSharp's NDIter walks a pair of operands that share a
permuted layout in LOGICAL order, gathering from both; it only reverses the axes of all-F operands. Measured on a
32^3 float64 x transposed (2,0,1): np.evaluate into the matching layout 40.6 us, into a C array 14.8 us. np.add(x,
1.0, out=) shows the same, 41 vs 17.5 us. This is library-wide and not fixed here (see LEFT AS IS).
Fix:
- NDPolySeries.MapDomainWith runs every route (complex64 kernels, complex128 affine, float64/float32/complex128
  affine, fused np.evaluate) over np.transpose(x, layout.DenseAxisOrder()), which is C-contiguous for a dense
  permuted x, so the affine SIMD route takes it.
- AdoptLayout then relabels the C result's storage with NumPy's layout (UnmanagedStorage.SetShapeUnsafe, the
  ndarray.shape setter's mechanism): no element moves, and the result keeps OWNDATA and WRITEABLE as NumPy's ufunc
  output does. It checks its contract (fresh, offset 0, C-contiguous, same size) and otherwise copies through two
  same-ordered views.
- The routes moved, unchanged, into MapDomainIntoC; the earlier draft's Relayout copy and out= pass are gone.

F3 - {p}sub NaN SIGN (U1)
The fused NegateAdd combine kernel (_sub's "negate c2, then add c1" when len(c1) <= len(c2)) negates the converted
subtrahend and then adds the minuend. When the conversion was an inlined float16 -> float32/float64 helper, RyuJIT
rewrote `(-a) + b` as `b - a`, which gives a NaN subtrahend the subtraction's sign rule instead of the negated sign.
P.polysub(f32 [1, 2], f16 [nan, 3]) returned +nan where NumPy returns -nan. EmitPolyCombine now stores the negated
value and reloads it before the add. The oracle tokenizes NaN, so a byte-level unit test is the gate.

DOCS CORRECTED
- Plan D9 said "{p}val with a 2-D c returns C-contiguous c.shape[1:] + x.shape", which is true only for C operands.
  It now states NumPy's rule, plus mapdomain's.
- Plan D12 promised a NativeAOT fallback (the out= composition over NumSharp's ufuncs) that was never built: every
  polynomial kernel factory raises PlatformNotSupportedException without dynamic code. It now says so.
- The np.polynomial facade remarks listed only U3/U1/U4. They now list all six delivered families, what is still
  open, and the NativeAOT behaviour.
- Plan: new section 11 (the review), D9/D12 corrected, review notes in U1 and U3.
- CLAUDE.md polynomial sections: the layout trap, the NDIter permuted-pair trap and the relabel route, the RyuJIT
  NaN-sign trap, and updated corpus counts.
- Fuzz README: sections L and R.

GATES ADDED
- polyeval.jsonl section L: 4,650 cases (19,216 -> 23,866, floor 23,800). 3,174 are "strides" facets
  (params.facet; _result_strides records element strides with every extent <= 1 axis, and every empty result,
  zeroed, because NumPy assigns those per loop path and NumPy 2.x reports zeros for any zero-size array). The other
  1,476 are value twins of results of <= 24 elements. The matrix:
  - 15 series entries: nine float64 layouts (1-D, C, F, permuted, transposed 2-D, reversed, stepped, broadcast axis,
    broadcast series axis) plus F/permuted/broadcast int64, permuted float32 and F complex128;
  - x: every catalogue layout for a 1-D series and a 20-layout subset for an N-D one, both tensor modes, weak and
    0-d scalars;
  - val2d/val3d/grid2d/grid3d over C and F series.
- polyseries.jsonl section R: 780 mapdomain cases (18,698 -> 19,478, floor 19,400): every catalogue layout x
  float64/float32/float16/int64/complex128 x a Python tuple domain, 1-D ndarray domains and a Python complex target,
  values and strides.
- OpRegistry.Polynomial.cs PolyLayoutFacet applies the same normalization to NumSharp's result (and disposes the
  result it replaces, for the leak gate). It acts only on a string facet "strides", because the constants' facet
  param is an encoded Python str object.
- Polynomial/PolynomialResultLayoutTests.cs (13 tests): NumPy-probed strides for each rule, a 6,000-case
  closed-form-vs-replay cross-check that drives the tensor=False shortcut (asserts > 100 shortcut hits), and
  mapdomain's relabel route (strides, OWNDATA, WRITEABLE, values) across the affine, fused, complex128 and complex64
  routes.
- PolynomialSeriesTests.Sub_EqualLengths_ConvertedFloat16Subtrahend_StillNegatesFirst: F3's bytes.
- Planted bugs, each killed:
  - the layout fix reverted turns 1,337 polyeval and 23 polyseries cells red;
  - the tensor=False shortcut without its dims check is caught by the cross-check;
  - relabelling with C strides turns three mapdomain tests red;
  - F3's test fails without the store/reload.

PERF (NPY/NS; NumSharp pinned to the P-cores, best-of-9; NumPy best-of-9 on the same host; before = a7fd4cc8)
| cell | before | after | NumPy | NPY/NS after |
|---|---|---|---|---|
| chebval x C 1000, c 11 | 1,082 ns | 1,073 ns | 18,881 ns | 17.6x |
| chebval x F 32x32, c 11 | 1,136 ns | 1,182 ns | 19,445 ns | 16.5x |
| chebval x broadcast 32x32, c 11 | 2,121 ns | 1,520 ns | 20,547 ns | 13.5x |
| chebval x C 1000, c F 11x4 | 3,870 ns | 3,859 ns | 49,087 ns | 12.7x |
| hermval 0.5, c F 11x32x32 | 4,540 ns | 4,366 ns | 31,552 ns | 7.2x |
| legval tensor=False, x C, c F | 6,001 ns | 5,981 ns | 38,759 ns | 6.5x |
| legval tensor=False, x F, c C | 14,440 ns | 14,422 ns | 35,295 ns | 2.45x |
| chebval x permuted 32^3, c 11 | 21.4 us | 21.2 us | 536.6 us | 25x |
| chebval x permuted 32^3, c F 11x4 | 430 us | 423 us | 6,056 us | 14.3x |
| mapdomain x permuted 32^3 f64 | 15.5 us | 4.7 us | 10.0 us | 2.13x (was 0.65x) |
| mapdomain x permuted 32^3 f32 | 13.8 us | 2.6 us | 6.5 us | 2.48x (was 0.47x) |
| mapdomain x permuted, complex domain | 24.4 us | 18.6 us | 50.7 us | 2.72x |
| mapdomain x F 32^3 | 4.58 us | 4.87 us | 9.45 us | 1.94x |
| mapdomain x stepped (32,32,32) | 6.45 us | 6.55 us | 16.8 us | 2.55x |
| mapdomain x transposed 300x300 | 13.4 us | 10.9 us | 306 us | 28x |
| polysub f64 5 - 5 | ~289 ns | ~299 ns (noise) | 9,768 ns | 33x |
A non-C call pays 30-100 ns for its layout; C calls are unchanged. Two sources of overhead were cut on the way: the
first draft re-cloned dims and allocated the vote table on the heap, and replayed the step table for every
tensor=False call (+1.6 us on legval).

LEFT AS IS (documented in plan section 11)
- NaN bits: 91 of the 19,837 probe calls differ ONLY in a NaN's sign or payload. They come from NaNs a recurrence
  generates (inf - inf) and from NaN-vs-NaN operand priority under RyuJIT's commutative operand swaps. Library-wide;
  the oracle tokenizes NaN.
- M1: U1's mapdomain, U4 and U7 emulate NumPy's complex64 VALUES on float32 kernels, but U3's evaluation computes
  complex64 loops in complex128. That adds a value divergence on top of #569's dtype divergence.
- Library-wide, not polynomial:
  - NDIter's logical-order walk of permuted operand pairs (above);
  - ufunc outputs and NDIter ALLOCATE operands are C/F, not NpyIter's K-order (c0(2,1,1) + F(3,4) is (12,1,3) in
    NumPy and (1,2,6) here);
  - NumPy 2.x reports all-zero strides for zero-size arrays.

VERIFIED (worktree at a7fd4cc8 + this change)
- NumSharp.Tests (TestCategory!=OpenBugs&!=HighMemory): net10.0 17,672 passed / 0 failed; net8.0 17,671 / 0.
- NumSharp.Tests.Oracle (same filter): net10.0 215 passed / 0 failed; net8.0 215 / 0. Every polynomial tier is
  green: polyeval, polyseries, polycalc, polyalgebra, polyvander, polyroots.
- coverage/generate_coverage.py: 494/560 headline APIs oracle-verified; `unittest discover -s coverage` 32 OK;
  `node --test coverage/test_dashboard.cjs` 15 pass.
- The generator's pre-existing cases are byte-identical after regeneration; sections L and R are appended at the
  end, so the running id counter leaves every existing id untouched.

FILES
src/NumSharp.Core/View/Shape.StridePerm.cs (NpyIterOutputShape, DenseAxisOrder, NpyIterStackLimit)
src/NumSharp.Core/Polynomial/Package/NDPolyEval.cs (SeriesLayout, ResultShape, SameDims, ReplayResultShape,
  LayoutEnv, LayoutOf, SameLayout; Val/Evaluate pass the series layout)
src/NumSharp.Core/Polynomial/Package/NDPolySeries.cs (MapDomainWith, MapDomainIntoC, AdoptLayout)
src/NumSharp.Core/Backends/Kernels/Direct/DirectILKernelGenerator.PolySeries.cs (EmitPolyCombine store/reload)
src/NumSharp.Core/Polynomial/Package/np.polynomial.cs (facade remarks)
test/oracle/gen_oracle.py (_result_strides, facet plumbing, polyeval section L, polyseries section R)
test/NumSharp.Tests.Oracle/Fuzz/corpus/polyeval.jsonl, polyseries.jsonl (regenerated)
test/NumSharp.Tests.Oracle/Fuzz/{FuzzCorpusTests.cs (floors), OpRegistry.Polynomial.cs (PolyLayoutFacet), README.md}
test/NumSharp.Tests/Polynomial/{PolynomialResultLayoutTests.cs (new), PolynomialSeriesTests.cs}
docs/plans/numpy-polynomial.md, .claude/CLAUDE.md
…e, MLNet, System.Numerics.Tensors, ParquetNet) - merged code, but no NuGet package, no release-notes line and no public docs page until the owner releases them; CI now builds and tests all four

WHY
The owner decided on 2026-10-02 that NumSharp.Interop.OnnxRuntime, NumSharp.Interop.MLNet,
NumSharp.Interop.System.Numerics.Tensors and NumSharp.Interop.ParquetNet stay in the codebase and merge
to master with journey4, but are not released: no NuGet package is created, and no release or public
documentation mentions them, until the owner releases them.

Three of them were already on master, and public. On 2026-09-10 journey3's tip (07881ef, the code/data
split) was fast-forward-pushed straight to master, and that push carried all 38 journey3 commits made
after the v0.70.0 merge, the onnxruntime branch merge (112dba9), MLNet (bec868c) and
System.Numerics.Tensors (b3a69d3) among them. Before this commit:
- CI would pack and publish OnnxRuntime and MLNet on the next tag, and the GitHub release body hardcoded
  an install line and a badge row for each;
- docs.yml had been deploying interop/onnxruntime.md, mlnet.md and system-numerics-tensors.md to
  scisharp.github.io/NumSharp since 09-10 (all three answered HTTP 200 on 2026-10-02);
- no workflow built or tested System.Numerics.Tensors or ParquetNet at all (a 2026-10-02 probe found five
  open bugs in the Tensors bridge, recorded in the catalog).
Removing the three from master instead was considered and rejected: they are interleaved with ~30 commits
master needs, journey4 sits on master's tip, and a plain removal commit would make PR #631's merge silently
delete the 81 of their 99 package files that journey4 never edited.

WHAT
Packaging, one switch per package:
- The four csproj files set <IsPackable>false</IsPackable> beside their PackageId, with a comment.
  Probed: `dotnet pack` of such a project exits 0, writes no .nupkg and does not even compile (a project
  holding a compile error packs cleanly). `dotnet msbuild -getProperty:IsPackable` reads false, so nothing
  in Directory.Build.props overrides it, and all four pack to nothing in Release.
- build-and-release.yml keeps naming them in the signing check and in build-nuget's Build and Pack
  steps, adding System.Numerics.Tensors and ParquetNet there, so the day one is released by flipping that
  property it is built, packed, signed, published and listed with no workflow edit.
- create-release's "Compose release notes" step no longer hardcodes the package list. It reads each
  packed .nupkg's nuspec <id> and lists exactly those, NumSharp first, NumSharp.Build last, the rest in
  ordinal order, and it fails the step when it reads no id or fewer ids than there are .nupkg files (an
  unreadable package). Run under an ubuntu bash against fake packages: the five released ids are listed,
  and the unreadable and empty cases fail the step.

Docs:
- docfx.json excludes docs/interop/{onnxruntime,mlnet,system-numerics-tensors}.md. DocFX publishes every
  .md under docs/website-src whether the toc links it or not, so leaving a page out of the toc does not
  hide it. The page files stay in the repo.
- toc.yml drops their three entries (a YAML comment says why; DocFX does not render it),
  interop/index.md drops their three table rows and "See also" links, advanced/index.md drops them from
  its list of bridges, and absolute-basics.md drops the ML.NET link from the mixed-CSV note, which would
  otherwise 404. ParquetNet has no page.
- What is left is not a mention of the packages: the two website hits for "System.Numerics.Tensors" are
  code samples calling the BCL's own TensorPrimitives, and Core names the packages only in // comments,
  which the API reference does not render.

CI gate:
- interop-test gains a System.Numerics.Tensors section (build, then net8.0 and net10.0) and a ParquetNet
  section (build against the floor Parquet.Net 6.1.0, then net8.0 and net10.0) after the ML.NET one,
  each test run filtered TestCategory!=OpenBugs so a future known-bug reproduction does not fail the job.
  Local Release run: Tensors 54/54 and ParquetNet 11/11 on both TFMs. OnnxRuntime and MLNet were already
  gated there. These steps are the only thing in CI that compiles the hidden packages.
- The job's header comment, its env comment and validate-release's comment now list six suites, and the
  signing check's timing note says only pythonnet compiles there now.

.claude/CLAUDE.md: a new "Unreleased packages" section (what hidden means, the rules, how to release
one), and the OnnxRuntime and MLNet table rows are marked UNRELEASED. AGENTS.md is a symlink to it.

Master gets a matching commit for the three packages it already carries, which takes their pages off the
public site. journey4 then records that commit with `git merge -s ours`, because this commit already does
everything it does.

Verified: the workflow parses (python yaml), every steps.<id> reference resolves and step ids are
unique; the four hidden projects pack to nothing; the Tensors and ParquetNet suites pass on net8.0 and
net10.0.

Files: src/NumSharp.Interop.{OnnxRuntime,MLNet,System.Numerics.Tensors,ParquetNet}/*.csproj,
.github/workflows/build-and-release.yml, docs/website-src/{docfx.json, docs/toc.yml,
docs/interop/index.md, docs/advanced/index.md, docs/absolute-basics.md}, .claude/CLAUDE.md
@Nucs Nucs changed the title [WIP 0.71.0] numpy.ma, np.emath and 48 new np.* functions, NumPy-parity fixes, ParquetNet interop, oracle expansion [WIP 0.71.0] numpy.ma, np.emath and 48 new np.* functions, NumPy-parity fixes, oracle expansion Oct 2, 2026
Nucs added 3 commits October 2, 2026 15:55
…ensors imported the wrong elements (GetPinnedHandle points at the backing array's head), 0-d ToTensor dropped its value, empty arrays exported as an unusable rank-0 span, rank-0 BCL values read past an empty array, ToTensor read released buffers, oversize copies allocated before refusing

WHY
A probe on 2026-10-02 drove every NumSharp dtype (15) through every NDArray layout (34: contiguous to
rank 70, offset/row/column/stepped slices, transposed, permuted, Fortran, unit axes, broadcast, diagonal,
read-only flags, negative strides, complex real/imag aliases, view(dtype), NDArray<T>, frombuffer,
memmaps r/r+/c, a 3e9-element broadcast) and every kind of Tensor<T> (17) through all five verbs, on
net8.0 and net10.0. Almost everything crossed bit-exact, but five bugs did not, and checking the fixes'
premises found a sixth. The 54 existing tests missed all of them: they import only tensors that start at
offset 0, test a 0-d array only through AsTensorSpan, and test an empty array only by FlattenedLength == 0.

WHAT (each bug: effect, root cause, fix)
1. Importing a sliced tensor returned the wrong elements, and writes through the view landed outside it.
   Tensor<T>.GetPinnedHandle() pins the backing T[] and points at the ARRAY's element 0, ignoring the
   tensor's start offset (verified on System.Numerics.Tensors 10.0.0, .NET 8.0.29 and 10.0.1: for
   t.Slice(1..3, ..) of a 3x4 tensor the handle points at backing[0], the slice starts at backing[4]).
   ToNDArray's dense path and both AsNDArray paths read from that pointer, so t[1..3, ..] came back as
   rows 0..1, and AsNDArray's rows[0,0] = -1 overwrote backing[0]. Any Slice / range indexer and any
   Tensor.Create(array, start, ...) hit it. Fix: new LogicalStart(tensor) = the span's
   GetPinnableReference(), the tensor's own first element; the handle is kept only for the pin.
2. ToTensor of a 0-d array returned a rank-1 LENGTH-0 tensor: the value was silently lost.
   Tensor.Create(data, []) reads empty lengths as length 0, not as a scalar. Fix: pass [1], the shape
   AsTensorSpan already gave a 0-d array (the BCL has no rank 0), so a round trip gives (1,) with the value.
3. Every empty array exported as TensorSpan<T>.Empty: rank 0, so a (3,0,4) shape was lost (handle.Lengths
   was empty), and the BCL's own FlattenTo / Tensor.Sum throw IndexOutOfRangeException on it, as did the
   bridge's span.ToNDArray(). The documented premise ("the BCL rejects a zero-length backing") was wrong:
   it rejects one only with a NONZERO stride. Fix: export the array's own lengths with every stride 0,
   over Array.Empty<T>() (no element is addressed, so the address is unobservable); probed for (0,), (5,0),
   (0,5), (0,0), (2,0,4) and bool/decimal: FlattenTo, Fill and Sum all work.
4. (found while checking the fixes) The BCL's rank-0 values - Tensor<T>.Empty, TensorSpan<T>.Empty,
   ReadOnlyTensorSpan<T>.Empty, a default span - hold NO element but were imported as NumSharp's 0-d shape,
   which claims one: ToNDArray threw IndexOutOfRangeException, and AsNDArray(Tensor<T>.Empty) returned a
   0-d array over one element past the end of the tensor's empty managed array (an out-of-bounds read; a
   write would have corrupted the GC heap). Fix: new ImportShape(lengths, flattenedLength) maps a rank-0
   value with no element to (0,); every import verb goes through it.
5. ToTensor of a released strided array read freed memory. It densified the source with copy() BEFORE
   Pin's TryAddRef, the only disposal check, so after the pool reused the buffer it returned another
   array's data (the probe read 777s). The contiguous path always threw. Fix: take the reference on the
   source first, for every layout and size (an empty released array now throws too, as in AsTensorSpan),
   and hold it across the copy.
6. The copies refused more than int.MaxValue elements only AFTER allocating the result: span.ToNDArray
   and ToNDArray(Tensor) of a 2^40-element stride-0 value threw OutOfMemoryException trying to allocate
   4 TB, and an allocatable size was allocated and thrown away. Fix: check before allocating, and dispose
   the result if the copy fails. The message now points at AsNDArray, which has no element-count limit.

Unchanged on purpose: negative-stride views are still refused by AsTensorSpan (the BCL forbids negative
strides: "Strides cannot be less than 0"), and the copies keep their int.MaxValue limit (FlattenTo fills an
int-indexed Span<T>; a managed T[] has the same bound).

Docs: the class summary's six broken crefs (CS1574: Export./Import. prefixes, an AsNDArray(..., bool)
overload that does not exist) now resolve; every edited member is fully documented (params, returns,
exceptions incl. the new ObjectDisposedException on ToTensor); the README gains an "Edge shapes" section;
the docs page's 0-d/empty, under-the-hood, lifetime and limits text is corrected, its test count is 64, and
its claims ledger gains rows 13-17 for the six fixes.

Tests (10 new, 54 -> 64): each new test was run against the unfixed code first and failed for the bug it
names (9 red; the 10th pins already-correct behaviour: AsNDArray of a 2^40-element stride-0 tensor is a
read-only broadcast view):
- ImportTests: ToNDArray_OffsetTensors_CopyTheirOwnElements, AsNDArray_OffsetTensors_AliasTheirOwnElements
  (write-through lands on the right backing element, nothing outside changes),
  OffsetTensors_AllDtypes_ImportTheirOwnBytes (15 dtypes x dense slice / strided slice / start offset x
  both verbs), RankZeroValues_ImportAsTheEmptyVector, Copies_OverIntMaxValue_AreRefusedBeforeAllocating,
  AsNDArray_StrideZeroTensor_OverIntMaxValue_SharesZeroCopy.
- ExportTests: ToTensor_Scalar_CrossesAsSingleElementVector_KeepingItsValue,
  AsTensorSpan_Empty_KeepsItsShape_AndTheBclCanUseIt ((0,), (3,0), (0,3), (2,0,4)),
  EdgeShapes_AllDtypes_CrossIntact (15 dtypes).
- LifetimeTests: ExportVerbs_ReleasedBuffer_Throw_EveryLayout (contiguous slice, transposed, stepped,
  broadcast, empty; owner and every derived view disposed).

Verified: the suite passes 64/64 on net8.0 and net10.0, Debug and Release (CI's interop-test job runs it
in Release with TestCategory!=OpenBugs); the package builds with no warnings of its own; the full probe
matrix re-run on the fixed code shows no wrong value and no unexpected exception for any dtype/layout/kind.

Not in this commit: a parallel session is releasing this package (IsPackable, docfx.json, toc.yml, the
interop index pages, the workflow, and the docs page's ONNX/ML.NET cross-links). Its working-tree changes,
including its five hunks in system-numerics-tensors.md, are left uncommitted for it; only this commit's
seven hunks of that page were staged.

Files: src/NumSharp.Interop.System.Numerics.Tensors/{NDArrayTensorsInterop.cs, NDArrayTensorsInterop.Export.cs,
NDArrayTensorsInterop.Import.cs, TensorSpanHandle.cs, README.md},
test/NumSharp.Tests.Interop.System.Numerics.Tensors/{ExportTests.cs, ImportTests.cs, LifetimeTests.cs},
docs/website-src/docs/interop/system-numerics-tensors.md (7 of the working tree's 12 hunks).
…y staging - three ellipses committed as cp1252 mojibake and claims row 12 placed after rows 13-17

WHAT WENT WRONG
c8f1dff had to commit only its own hunks of system-numerics-tensors.md, because a parallel session was
editing the same file (removing its ONNX Runtime / ML.NET cross-links while releasing the package). The
hunk split went through a script that read `git diff -U0` with Python's subprocess(text=True), which on
Windows decodes with the locale code page (cp1252), not UTF-8, and the patch was staged with
`git apply --cached --unidiff-zero`. Two defects reached the commit:
- every U+2026 ELLIPSIS in the added lines became the mojibake "…" (UTF-8 bytes read as cp1252, written
  back as UTF-8): two on the "Under the hood" line about sliced tensors, one in claims row 13;
- claims rows 13-17 landed ABOVE row 12. They were a pure insertion with zero context, and git apply
  locates such a hunk by its NEW-side line number. That number counted the parallel session's hunks the
  split had dropped (one line fewer above), so the insertion went one row early.
The working-tree file was correct throughout (it was written by the editor, not by the patch), and the other
eight files of c8f1dff were staged directly from the working tree, so only this page was affected
(checked: no other line c8f1dff adds contains a mojibake sequence).

FIX
The page is restaged as a blob built from HEAD's bytes, with no text decoding: the three mangled sequences
become U+2026 again and row 12 moves back above row 13 (`git hash-object -w` + `git update-index --cacheinfo`).
The parallel session's five hunks on the page stay in the working tree, uncommitted; after this commit the
working tree differs from HEAD on this page by exactly those hunks.

Lesson (for any hunk-only staging): read git output as bytes or with encoding="utf-8", and never stage a
zero-context pure-insertion hunk with --unidiff-zero after dropping other hunks above it; build the blob
from HEAD and stage it with update-index instead.

Files: docs/website-src/docs/interop/system-numerics-tensors.md
….71.0 - packable again, its docs page back on the site, its page/XML docs/README no longer naming the hidden ONNX/ML.NET bridges; journey4 rebased onto master 90311b6 (old history in journey4-backup-2-10-2026, old->new map below)

WHY
Later on 2026-10-02 the owner decided that NumSharp.Interop.System.Numerics.Tensors ships in the upcoming
release (0.71.0), while NumSharp.Interop.OnnxRuntime, NumSharp.Interop.MLNet and NumSharp.Interop.ParquetNet
stay hidden. The owner also asked for journey4 to be rebased onto master, which had gained 90311b6 (the
same hiding, for the three packages master carried). This commit lands on c8f1dff and f24d0c9, a
parallel session's fixes for the six Tensors bugs found earlier that day (54 -> 64 tests).

WHAT - releasing System.Numerics.Tensors
- src/NumSharp.Interop.System.Numerics.Tensors/*.csproj: the IsPackable=false block is removed, so the
  project is packable by default again. CI already names it in the signing check and in build-nuget's
  Build and Pack steps (added in the hiding commit), and create-release lists whatever was packed, so no
  workflow logic changes; only the workflow's comments now speak of three unreleased packages.
- docs/website-src: docfx.json no longer excludes docs/interop/system-numerics-tensors.md; toc.yml lists
  it again after OpenBLAS; interop/index.md gets its bridge-table row and "See also" link back;
  advanced/index.md names it in the list of bridges again.
- Nothing that ships with it names the hidden bridges any more:
  - the docs page loses the intro's "sibling of the ONNX Runtime and ML.NET bridges", two "unlike ONNX"
    asides, the two "unlike the ONNX/ML.NET bridges" dtype notes and the two "See also" links;
  - the package's XML docs, which ship in the .nupkg and show in IntelliSense, lose all six mentions:
    NDArrayTensorsInterop.cs (class summary twice, the Layout paragraph, Pin, WrapExternal) and
    TensorSpanHandle.cs (the lease's summary);
  - its README loses two (README files are not packed, but the docs page links it).
  Only AssemblyInfo.cs's `//` comment still names them, and comments never ship.
- .claude/CLAUDE.md: "Unreleased packages" lists three packages and records that Tensors was hidden and
  released again on 2026-10-02. A new rule says a released package's XML docs and README must not name a
  hidden bridge either. The packages table gains a System.Numerics.Tensors row.

REBASE
journey4 now sits directly on master 90311b6. It was rebased with git plumbing instead of `git rebase`:
- 45c6840 (exprs) and 179b2f6 (onnxruntime) are merges with hand-made conflict resolutions. A
  linearizing rebase drops merges, and `--rebase-merges` re-merges without the original resolutions.
- Nine commit messages contain lines starting with '#', which `git rebase --continue` strips.
Method: every commit that descends from the old base 9396799 (290 of the 296 in the range) was recreated
with `git commit-tree`. The message bytes, author, committer and dates are the same, and the parents are
mapped (the old base becomes 90311b6; a rewritten parent becomes its new commit; anything else stays). Its
tree is the original tree except in the eight files 90311b6 touched. There, 90311b6's change was carried
forward by a 3-way merge (base = the old base's file, ours = the commit's own file, theirs = 90311b6's
file), with conflicts resolved to the commit's own text. The hiding commit (83738b41, now 992dbb1)
already did everything 90311b6 does, so it kept its original tree exactly. Six side-branch commits that
forked before the old base keep their SHAs (53602da, 5e65ed9, 939d063, dcbe707, 6a1d0c4, 7b3c8a5),
and so does the merge structure.
Verified before the branch moved: 290/290 messages, authors and committers byte-identical, parent counts
and mapping correct, trees differing only in those eight files; the new tip's tree identical to the old
tip's; 296 commits in 90311b6..journey4 as in 9396799..old tip; every rewritten workflow version parses as
YAML; no rewritten CLAUDE.md carries master's section twice. The `-s ours` merge 68107839 that 83738b41's
message mentions is gone, because master is now in the base. journey4-backup-2-10-2026 points at the old
tip (68107839), so every old SHA stays reachable. The parallel session's c8f1dff and f24d0c9 were made
on the rebased branch and are not in the map.

VERIFIED (this commit's tree, checked out in a clean worktree)
- `dotnet pack` of the four interop projects in Release produces exactly one package:
  NumSharp.Interop.System.Numerics.Tensors (lib/net8.0 + lib/net10.0 dll/xml/pdb, LICENSE, icon, no
  README; depends on NumSharp and System.Numerics.Tensors [10.0.0, 11.0.0)). OnnxRuntime, MLNet and
  ParquetNet pack nothing.
- .github/scripts/verify_strong_name.cs: both assemblies carry token cc7b13ffcd2ddd51 (failures: 0).
- Both shipped XML doc files: 0 mentions of ONNX, ML.NET, MLNet or Parquet.
- test/NumSharp.Tests.Interop.System.Numerics.Tensors: 64/64 on net8.0 and on net10.0.
- The workflow parses, step ids are unique, every steps.<id> reference resolves.
- No page under docs/website-src (except the two excluded ones) names or links the hidden bridges; the
  Tensors page is in toc.yml and not excluded.
This commit's tree was built with plumbing from HEAD plus this session's files, never through the shared
index, because a parallel session was committing to journey4 in the same working tree at the time.

Files: src/NumSharp.Interop.System.Numerics.Tensors/{NumSharp.Interop.System.Numerics.Tensors.csproj,
NDArrayTensorsInterop.cs, TensorSpanHandle.cs, README.md}, .github/workflows/build-and-release.yml,
docs/website-src/{docfx.json, docs/toc.yml, docs/interop/index.md, docs/interop/system-numerics-tensors.md,
docs/advanced/index.md}, .claude/CLAUDE.md

The old -> new map of the rebase (8-character SHAs, parents first):
5e3d6c3 -> ea080d8
831b664 -> 5c667b0
10a49f6 -> 76b2ab9
7192172 -> 86cc28b
d1d86c2 -> 0e9dbc4
9a368db -> 2c88b9a
1b82a7c -> 3ce81dd
ff84f7f -> aaf504e
2713086 -> dd81dda
3710d26 -> 0f16ea1
69849e1 -> d91b807
000039c -> 7630ec1
ea91017 -> df23234
355d695 -> 22d5f4d
23c4521 -> f6037bc
ab69cf8 -> dc9caff
3622a6e -> 2693d83
cb74b5f -> 4898b40
d1520ed -> c250ca8
58cd55d -> 3e7d6b3
7c3e14d -> 0e80bc8
a43dc6f -> d2c8d26
dbc0b3b -> 338d260
70a7173 -> 4463674
66f4092 -> 6e0e5ea
ea95208 -> 53b7a14
11d5b6f -> ade8d6b
c400ba2 -> 3936fd0
f247063 -> c2f129b
6180915 -> 46f5f65
7f58f01 -> 83e0e43
9af8fa3 -> f61e2d7
80dde0a -> bab51ed
6a19340 -> a15f0ee
08d5d13 -> 197aaf3
1dae35a -> c3be0b7
68081a7 -> 38f9a14
ac95ae8 -> 23ea2ad
42cffc4 -> 39e7627
4b16555 -> 30ccf0b
074d4d0 -> c2dcfa7
a0013f2 -> 4cd3b72
0be4fc4 -> d48737e
7fd70a3 -> d480418
63a0795 -> c287d70
281fece -> 3a36752
5f020a0 -> 68a84d3
45c6840 -> 0bfbc06
670afdd -> df79938
b53f5e5 -> 4460419
f3490fd -> bfc0756
4664f6b -> fb50061
1f18c44 -> 1af18f8
6373a4c -> f417827
bdd4df0 -> 7cd38a9
81f1713 -> 8af2107
bea2f60 -> ea32522
acbf571 -> 877a36b
2ba16e7 -> 36d5cae
ee366e6 -> a0c6e2f
dfd0608 -> f361da9
60df106 -> 72e4d29
179b2f6 -> c0ba8f3
fda311c -> b1937c5
1755e1b -> a3666a5
144da76 -> 686fa01
5316a61 -> 5f1476f
1b0d0d6 -> ba15bad
f7058a5 -> bf66bc0
34e6b0f -> a44ec19
01392bc -> ebfc711
60a0ea6 -> 10a5464
566fd30 -> 4406b7a
d58207e -> e740f3a
994134c -> 11f9c93
2b57d07 -> 8525d4b
1fcaea0 -> 39e5ba1
e425740 -> 8e0e48a
9071b53 -> 6db8238
fad2cbe -> ac6ec6f
d2f0467 -> de3c429
539e9c0 -> e4d73d8
6bc7d4a -> d7f1baf
36fc1ae -> c945c7b
1cb0411 -> b6f13a0
ab96ae7 -> fed9618
0c47162 -> da2dd86
0f8f8f7 -> b33b651
b3a912b -> b184183
aee243c -> 3fb1d89
1d1a226 -> 4a5235d
ea1abf7 -> 244e233
9857741 -> 832a3ca
201e4e2 -> 8f4b9e8
c1712be -> 06bb2ae
2e46e74 -> e45f1de
1b652ba -> e674512
71bf2f5 -> d870887
4368c95 -> 8b1e262
3a82d65 -> 30106bc
060809a -> 76228e4
8906619 -> d04b3c3
6e2869e -> 47c1260
6509f07 -> d07baa7
b89f714 -> bfdf174
d815adf -> 03c4ec4
bcf924f -> 76e8f78
6987944 -> 2cbe8e1
cd39ef5 -> 0370eba
89b719d -> 8558705
90b75fd -> 7d6ae94
7cabac1 -> bade986
fcdacd2 -> c70ba1f
1d82eac -> 8b86730
79e148c -> 2862d29
0c5d72c -> dc2fc07
77e419b -> e2db7be
2b67f29 -> 4b4ca2f
6847cd7 -> 2e0db80
25e8536 -> 2945841
02da9dd -> 8df5a25
d837ded -> 5e722ca
0d57514 -> 8f23550
846f53a -> 192f456
ff7283f -> 97360c8
35dd5eb -> 1dd23bb
d41eb4a -> e5afe84
cfc149d -> 989b246
db4ab83 -> 32162b6
60024b4 -> f319b65
6cf21af -> ba5c347
cf511e2 -> 392b52c
b14ff3a -> 67f8f96
ae43ba1 -> 9a1eeb7
3ad4360 -> 2e205a1
59f9932 -> 3384a0b
e7f7384 -> 7a262af
3ceb6fa -> e84fd2f
d35955a -> 55a2a6a
e3d5b77 -> 0398f30
13b7e0d -> 91d170d
2df6bf7 -> cd4a4ec
9a2d388 -> 2c1e402
abc0235 -> 07c7b94
4bfcde5 -> 1dc1d35
8a8318f -> 4406760
6a88d55 -> f500350
6c190f1 -> 4deab50
b779d95 -> 538dd00
2ea7ba6 -> e0d6a81
5635182 -> 0467720
4499d3e -> 563fc3e
b7d9093 -> 74656f1
fe756f9 -> 6d46a17
035f24f -> 29f10f2
6ee94ae -> 3bc9623
33a4d2c -> ea48e66
5f2821d -> 7a8b565
256224d -> 21d4c89
dc32aec -> 873900b
c3ebaed -> 478b531
b1a8211 -> f8b25f7
45c7913 -> f50913f
8df2598 -> f4d6b2a
4539456 -> 6f2ec47
77b967d -> 7e60327
2180151 -> 3ac37d1
2adf21d -> afe8be8
bed56ee -> e81b728
f12a03f -> a325155
3ef6b0e -> cc1be54
fe30edf -> d1b89c1
54bb26d -> c243201
cc8d56c -> 5463714
9019f9d -> 51fb9c9
cdc9088c -> 4231d2d
33d066d0 -> 6402196
cf4cba91 -> 5bc19d8
51dcf123 -> 550b2ed
e471e621 -> 034b2be
81c26c19 -> 94c477f
5658cf85 -> e405ee9
f837f20a -> 0c87022
61d1e273 -> 9de61fc
5b361bd3 -> 34f6f58
47973fc5 -> b30833b
000f45fc -> 25cc7d7
27bde580 -> b8109cb
5740fc70 -> 28fa9d7
c03f2476 -> a46259e
ddb48a00 -> 3e4a5a9
723fe758 -> 25009dd
ebc1ef06 -> 5ae3bf0
f3b9d80b -> a54450b
0dbf6da5 -> fdd728e
573a0c57 -> 801ddcb
e2b26573 -> a0d8d4e
d8d1b04a -> 5da0c78
03279983 -> 99ac94e
ac2b5430 -> 618a3ce
592b3e39 -> b298ab7
03853e5a -> 6bc242f
56e98f7a -> c5b066c
12937504 -> c6f3cd5
63104b04 -> 8da37f8
737a5573 -> 6b8e0b4
63a94fa2 -> 932b925
64dc4058 -> 9957d7a
9ff46baf -> e140097
75a799d9 -> 98bf7f9
94931b42 -> 1c8a98e
d874383c -> e877506
29ffe89c -> 30ad797
b1fdf9ad -> f3eb542
cf57b992 -> 2f02d65
143cdbb7 -> eb76e57
2fb9f901 -> dcdd656
619b443e -> e39519b
18631c84 -> 6175a01
2510bb89 -> 4462122
b5df3bbb -> bd0091f
e719db97 -> bf16cdc
f558a4fd -> ab6b6b1
02588e54 -> fc8be29
c05d3599 -> 0fb51fd
65ac6c93 -> eb1921e
704cc325 -> fbef1c4
22c56c53 -> c5422cf
ab90837a -> b53a0cc
b9736ddf -> c9de59f
5d38f1ed -> 3631fcc
724bdd9c -> d3057b3
eb64239d -> 7aba9b4
4af0a9f9 -> 4a87527
40c195c4 -> 30a35d1
63ec3406 -> ee63986
c25ce677 -> 30dbfe0
0ad77443 -> 3f22d30
ab4afb09 -> 0f78f6e
7dfc8768 -> 0f8819b
baaaaea2 -> 129fbab
02a14699 -> 78692d9
2ae46027 -> ceac36f
9363c797 -> 1d79ce6
2e8e6196 -> f8f91b3
969b0e1c -> de6333d
b2b94528 -> 268d91f
20a920e3 -> 1336e36
cb64deb1 -> 5754520
7ec2ebdd -> 7351465
bdc1da34 -> 1855ea9
ea32af61 -> fd058f0
406dd450 -> c1e0feb
2797ded2 -> fcab27e
34e14921 -> 65bdea7
34213ed1 -> aa9905b
80be9449 -> f87b4d0
5bbb7e35 -> 072eb12
f52c1441 -> cf37128
61ea12af -> 910a648
ebe18b7c -> ac89847
fa4bf379 -> c0613b2
af3a70a1 -> 2358d71
bd0f026b -> b1f282c
3a15871a -> 39d446e
099dd982 -> 5f20488
da84683f -> 85d31b5
9aac6ee9 -> eebe85a
879791da -> 20a5b60
844580a4 -> 1c7ac69
c5ef9461 -> 38ed8da
d34ff6c0 -> 8f70af7
780fe625 -> ecdab2f
b236c701 -> a72a5eb
3d4392e6 -> cdf22dc
1a367afe -> 0e39714
bea13b21 -> 3dcf999
7fb63fd8 -> 564dc69
f8ac16f7 -> ac8608d
4bc7588d -> 0896cbf
8d71581e -> 3852615
cbf272f1 -> 058f4f2
5e0d408c -> e2ff947
be1202f5 -> 37f7262
faa1e2b9 -> 0c02740
890ff1dd -> 5989050
6dd6746f -> 105efb7
a7fd4cc8 -> 3ecccf6
6314e156 -> d494e83
83738b41 -> 992dbb1
@Nucs Nucs changed the title [WIP 0.71.0] numpy.ma, np.emath and 48 new np.* functions, NumPy-parity fixes, oracle expansion [WIP 0.71.0] numpy.polynomial, np.random parity, numpy.ma, np.emath, 48 new np.* functions, System.Numerics.Tensors interop, NumPy-parity fixes Oct 2, 2026
…NoWarn - TreatAsLocalProperty="NoWarn" keeps their [Experimental] suppressions (SYSLIB5001, NUMSHARP_TENSORS) from being dropped

WHY
The first CI run of a592b7e (PR #631, run 37010896309) failed "System.Numerics.Tensors: Build" on all three OSes,
and the ubuntu `test` job's strong-naming pack, with 8 x `error NUMSHARP_TENSORS` (TensorSpanHandle.cs lines 172
and 190, both TFMs: the package's own code using its [Experimental("NUMSHARP_TENSORS")] surface). Both csproj files
suppress SYSLIB5001 and NUMSHARP_TENSORS through `<NoWarn>$(NoWarn);...</NoWarn>`, but CI passes
`-p:NoWarn=${{ env.DOTNET_NOWARN }}` to every build, test and pack. A command-line property is GLOBAL, and MSBuild
silently ignores a project's own assignment to a global property, so the suppression vanished and the
experimental-API diagnostics became errors. Local runs never pass -p:NoWarn, which is why every local build,
pack and test of the package (including a592b7e's clean-worktree verification) passed. The package was never
compiled by CI before 992dbb1 added it, so nothing had caught this.

WHAT
- src/NumSharp.Interop.System.Numerics.Tensors/*.csproj and test/NumSharp.Tests.Interop.System.Numerics.Tensors/*.csproj:
  `<Project Sdk="Microsoft.NET.Sdk" TreatAsLocalProperty="NoWarn">`, with a comment explaining it. The project's
  NoWarn assignment then wins over the global property and still starts from the command-line value, so CI's list
  (CS1570...NU5128) and the projects' two codes both apply. No other project in the repo needs this: none
  suppresses an error-severity diagnostic through NoWarn (ONNX Runtime, ML.NET and ParquetNet use no
  [Experimental] API, which is why their CI steps were green).

VERIFIED (clean worktree at a592b7e + this change, CI's exact DOTNET_NOWARN value):
- before the change: `dotnet build test/...Tensors.csproj --configuration Release -p:NoWarn=<DOTNET_NOWARN>` fails
  with the same 8 NUMSHARP_TENSORS errors CI reported;
- after: that build succeeds; `dotnet test --no-build --framework net8.0|net10.0 --filter "TestCategory!=OpenBugs"`
  64/64 on each; the signing-check form `dotnet pack <package> --configuration Release -p:NoWarn=...` produces the
  .nupkg; build-nuget's form (`dotnet build -t:Rebuild -p:NoWarn=... -p:Version=0.71.0 ...` then
  `dotnet pack --no-build ...`) produces NumSharp.Interop.System.Numerics.Tensors.0.71.0.nupkg.

NOT fixed here (pre-existing, surfaced by the same run - the first CI run of the ~114 journey4 commits made after
2026-09-23, which had never been pushed): on ubuntu and macOS, Generator/RandomState stream tests
(LongStream_Seed2024_MatchesNumPyDigest x20, Branch_Seed42_MatchesNumPy) compare against streams NumPy produced
with Windows' ucrtbase libm and are not host-pinned; on macOS (arm64) several polynomial/random tests assert x86's
negative default NaN, and the test host crashed after 3,809 tests.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant