Skip to content

v2.9.2 - #630

Closed
singaraiona wants to merge 45 commits into
masterfrom
dev
Closed

v2.9.2#630
singaraiona wants to merge 45 commits into
masterfrom
dev

Conversation

@singaraiona

Copy link
Copy Markdown
Collaborator

What & why

Checklist

  • PR targets dev (not master)
  • Commits follow Conventional Commits (feat: / fix: / perf: / docs: / …)
  • make builds cleanly (no new warnings)
  • make test passes; tests added/updated for behaviour changes

ser-vasilich and others added 30 commits September 19, 2026 19:55
The native top-N of the radix path kept every group at or beyond the
threshold.  Groups strictly beyond it number fewer than N, but the
groups AT the threshold can be nearly all of them: a count-per-group
top-10 over a near-unique key has a threshold of 1, every group ties
it, and the "kept superset" is the whole grouping, heap-sorted by
first row — a comparison sort over 100M pairs for a query whose answer
is ten rows.

The emitted prefix is the tied groups' first-seen order, so only the N
tied groups with the smallest first row can ever be taken.  The
selection now keeps the strictly-better groups plus exactly those N,
through a bounded max-heap, and sorts at most 2N entries.  The rows
emitted are the ones the full kept set produced.

Tests: the pinned kept count of the radix native top-N case becomes N;
a near-unique two-key grouping whose threshold every group ties, in
both directions and under a row selection, against the full grouping.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Rayfall had no way to stop iterating before the end of a sequence.  Every
iteration primitive — map, pmap, fold, fold-left/right, scan, scan-left/right,
prior — consumes its whole input, and there was no loop construct, so a
"repeat until done" loop had to be written as a fold over a fixed range whose
full length was paid on every call however early the work finished.  `return`
did not help: it exits the lambda, not the iteration, so the fold kept calling
the lambda for every remaining element.  Recursion was not an alternative
either, since there is no TCE and the stack tops out around 1-2k frames.

(while cond body...) evaluates cond, and while it is truthy evaluates each
body expression in order, then tests again.  It always returns null — a
statement form run for effect, and a never-taken loop has no last value to
report.  Zero body expressions is legal, so a condition with side effects can
be the whole loop; that is the shape a drain wants, where there is no sequence
to iterate over and the range was only ever scaffolding.

Implemented on both evaluators, which must agree:

  - ray_while_fn (eval.c), registered beside `if` and `do`.
  - an sf_while case in the bytecode compiler emitting JMPF over a backward
    JMP — the first backward branch the compiler produces.  The VM's op_jmp
    already checked for a pending interrupt when the displacement is negative
    (dormant until now), so a runaway loop is Ctrl-C-able for free.

Compiling the body inline, rather than letting it fall through to the generic
special-form path, is what keeps `return` working inside a loop: it reaches
the sf_return case and unwinds the lambda instead of degrading to the tree
walker's identity `return`.

Neither path pushes a scope around the body.  `let` binds only in the top
frame (env_bind_local), so a per-iteration frame would discard loop-carried
`let` state on the tree-walking path while the compiled path — whose `let`
writes a function-level slot — kept it, and the two evaluators would disagree
on the same source.  A caller wanting a fresh frame per pass writes
(while cond (do ...)), which composes.

patch_jump and emit_jump_back now share write_jump_offset; an out-of-range
displacement still sets c->error, dropping the lambda to the interpreter.

Measured on the drain shape from the issue (4-step drain under a 4096 safety
bound, release build, driver baseline subtracted): fold-left over (til 4096)
1420 us/batch, the nested 64x64 workaround 45.8 us/batch, `while` 3.8
us/batch — 374x over the original and 12x over the workaround.

Closes #588
The two remaining early-termination forms requested in #588, alongside the
`while` that landed in #590.  Neither is an unblock — `while` already covers
the reporter's case — but both complete the vocabulary he asked for.

`times`
-------
(times n body...) runs the body exactly n times and returns null.  `do` is
already progn in this language, so the bounded loop could not reuse that name;
overloading `do` on an integer head was rejected as genuinely ambiguous —
(do 5) would have to mean either "loop five times over nothing" or "return 5",
and a computed first expression that happened to be an integer would silently
change meaning.

The count is evaluated ONCE on entry, so a body mutating whatever produced it
cannot change how many passes remain.  A count of zero or less runs zero times
rather than trapping; a non-integer is a type error.

Compiled as a counted loop over a hidden local slot, reusing the backward
branch added for `while`.  ray_times_norm_fn type-checks the count and clamps
a negative bound to zero on entry, which lets the per-pass test be a bare
truthiness check on the counter — 0 is falsy, so no comparison call is needed
per pass and a negative bound cannot run away.  That helper and the decrement
are pushed as constant-pool objects rather than resolved by name, so the
loop's own arithmetic is unnamable from source and cannot be swapped out by a
`(set - ...)` override.  The counter's slot is addressed by index and its sym
carries a space, so no source token can collide with it and nested `times`
counters stay apart.  As with `while`, the body compiles inline, so `return`
unwinds the enclosing lambda from inside the loop.

fold-while
----------
(fold-while pred f init xs) offers the accumulator to `pred` before each step
and stops on a falsy answer, yielding the accumulator as it stands.  The test
precedes the first element, so a predicate false at the start returns `init`
untouched.  The predicate takes the accumulator rather than the element: that
is the form that expresses "iterate until the running result says stop", which
is the early termination actually being asked for.

One deliberate divergence from ray_fold_fn: that routes its collection through
unbox_vec_arg -> to_boxed_list, boxing every element up front.  For a
primitive whose purpose is to stop early, paying for the tail it never reaches
is the cost being removed, so elements are pulled one at a time via
collection_elem.  A plain variadic builtin — no compiler work, since it
dispatches through the normal call path.

Measured, release builds: `times` 100 ns/pass against 180 ns for `while` plus
a manual counter over 1e6 passes.  `fold-while` stopping after three elements
of a 1e6-element vector costs 3.9 us against 369 ms for the `fold-left`
equivalent, which had to box and walk all million to discover it was done.

Tests: test/rfl/lang/times.rfl and test/rfl/collection/fold_while.rfl, each
behaviour asserted on both evaluator paths where applicable — including the
count validation and negative clamp on the compiled path, which runs through
entirely different code from the tree walker's check.
feat(lang): add a `while` special form
fix(group): bound the tied groups the radix top-N selection keeps
feat(lang): add `times` and `fold-while`
ray_vec_is_null is out-of-line (no LTO) and the join called it once per key
column per row in hash_row_keys — on the build side, the probe side, and the
prefetch lookahead — plus twice per key column per hash-chain step in
join_keys_eq, across both the count and the fill pass.  A reported profile
put 13.28% of a service's samples there, on a join keyed by two SYM columns
that structurally never hold a null.

Prove once per join that no key column can hold a null and drop the call.
SYM/STR nulls are canonical empty payloads (id 0 / length 0) that HAS_NULLS
does not track, so text columns are proven by the chunked zero-scan from
#533 rather than by the flag; everything else reads the flag through slices
and takes a set bit at face value, so the proof stays O(n) and never
degrades into ray_vec_has_nulls' per-element walk.  Flag-readable columns
are settled first, so a nullable numeric key short-circuits before any text
column is scanned.  OP_CONST (atom) key slots are refused as unprovable.

Also skip the #458 null-run pre-scan under the proof.  It gates on
ray_vec_may_have_nulls, which is unconditionally true for SYM/STR, so a
SYM-keyed join ran a full per-key-per-build-row ray_vec_is_null scan before
the join proper on every execution; a null-free key set cannot contain an
all-null row, so the scan is dead.

Measured with bench/join_nullfree (release, 1:1 book join, 4M probe x 500K
build): two SYM keys 275 -> 249 ms (-9.5% median, -7.8% min); single I64 key
106 -> 99 ms (-6.3% median, -7.9% min).  A proof that fails on a text key
costs a partial scan (~8ms on a 4M-row SYM column) and gains nothing — only
nullable SYM/STR key columns pay it.

ray_join_force_null_checks forces the null-aware loops so the differential
tests and the perf gate can compare both paths in one binary;
ray_join_nullfree_keys counts the joins that took the fast path.

Closes #597
perf(join): skip per-cell key null tests on provably null-free columns
fix(system): validate launcher and timeit integer arguments
Every value-taking startup flag was gated on `i + 1 < argc` and then
consumed argv[++i] blindly, so one flag silently ate the next.  The shape
that was reported is the worst one for a service:

  rayforce -Q -p 5099 svc.rfl

`-Q` took `-p` as its value, the port was never bound, and nothing on
stdout or stderr said so.  Under a supervisor the process starts, the
script runs, the unit looks healthy, and every client gets connection
refused.  `-c` and `-t` had the identical hole.

Two neighbouring cases came out of the same code.  A trailing flag with
no value fell through to the positional-file arm and reported the
misleading `cannot open '-Q'`.  And that arm took ANY unrecognized token
as the script name, which the real positional then overwrote, so a
typo'd option was not merely ignored — it was silently swallowed:
`rayforce -x svc.rfl` ran the script as if `-x` had never been typed.

Every value-taking flag now goes through flag_value(), which refuses a
missing value and a value that is EXACTLY one of this program's flag
tokens (including the `--` app-args terminator), with a diagnostic and
exit 2 before the script is loaded.  A value that merely STARTS with '-'
stays legal: a password or a negative number is a legitimate value, and
refusing those would break working command lines for no gain.  An
unrecognized `-`-prefixed token is now an error too; a lone `-` is still
a positional, and tokens after `--` belong to the app and are never
validated.

`-Q, --querylog N` was a real flag, referenced twice in the docs, that
--help never listed beside its sibling `-t, --timeit N` — added, along
with `[-Q 0|1]` in the synopsis.

The suggestion to fail when `-p` was requested but nothing bound is
already in (#473, listen_fatal.rfl), and cannot catch this: the parser
never saw the eaten `-p`, so there was no request to check against.  The
guard on the value is what closes the class.

Also: the two pre-existing `-p` validation bail-outs now release the
runtime like every other error path, since these paths are exercised
under ASan by the new test.

Closes #600
…ionally

Every node the executor evaluates returns an owned reference — the
constant table node included, it retains its literal.  The OP_WINDOW,
OP_SORT and OP_HEAD cases released their input only when it differed
from the graph's table, taking an equal pointer for a borrowed one.  A
query's root is a constant node over that very table, so the reference
was never released: one input table per windowed, sorted or limited
query.  Invisible while a global kept the table alive; a whole table per
call when the input was built for that call, as a service that windowed
a freshly concatenated buffer on a timer found (#602).

The four cases now release the input on every path, as the join and
the plain head/tail cases already did.

Test: window, sorted and limited selects over a table built per call,
measured with the two-window bytes-allocated method of the other memory
probes, plus a shared input reused across fifty calls.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… queries

Twenty calls of each shape over one global table; the table's refcount
must be exactly what it was before, and the table still whole after.
Fails on the unfixed executor at the first window shape.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The same borrowed-query-table guard sat under the reductions; no child
evaluates to the query table there today, so nothing leaked, but the
premise is the one the sort, window and limit cases just dropped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
fix(exec): release the input table of window, sort and limit queries
fix(cli): reject a flag as another flag's value, and unknown options
ray_pool_dispatch and ray_pool_dispatch_n signalled the whole pool on
every dispatch, however narrow the window.  The main thread participates
as worker 0, so at most n_tasks-1 helpers can ever claim anything; the
surplus threads woke, raced to an already-drained window, and went
straight back to the semaphore.  On a dispatch narrower than the machine
that is pure overhead — a partition-parallel step over three partitions
woke every core on the box to run two tasks.

Signal min(n_tasks-1, n_workers) instead.  Signalling FEWER is safe:
completion is governed by the `pending` spin-wait, never by the signal
count, so an unsignalled worker simply stays asleep, and under-signalling
cannot produce the surplus-signal problem between consecutive dispatches
that the spin-wait exists to avoid.

Measured (10M-row ClickBench, 43 queries x 3, splayed): sum of per-query
hot times 5690.7 ms -> 5667.3 ms (-0.4%, within noise), user CPU 55.1 s
-> 54.8 s.  So this is waste removal rather than a speedup on workloads
whose dispatches are already wider than the pool.

Not a fix for #599.  That report's shape — a 130k-row join probe, which
splits into 16 tasks against 8 workers — is unaffected, because the clamp
lands on the same worker count it already used (measured: 5.19 s -> 5.18 s
of CPU, unchanged).  The investigation there points at the pool defaulting
to logical rather than physical cores; that is a separate change with its
own benchmark trade and is filed on its own.

test_pool.c gains pool/dispatch_narrow, which is the case the existing
coverage missed: dispatch_n_small uses 4 tasks on a 2-worker pool, so the
clamp never binds.  The new test drives windows of 1..4 tasks on a
4-worker pool, then alternates narrow and wide windows 25 times over the
same pool, asserting exact task and element counts throughout — the two
risks the change introduces are a task nobody claims and signal
accounting that drifts across dispatches and starves a later wide window.
perf(pool): wake only the workers a dispatch can keep busy
Detect Windows via $(OS) or uname (MSYS2's make hides $(OS)), link
Winsock and a statically linked winpthreads, build with 64-bit off_t
(_FILE_OFFSET_BITS=64) and an 8 MiB stack like a Linux main thread.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replace the IOCP stub with a readiness-based loop that mirrors the epoll
backend's dispatch order, so the selector state machine is identical on
every platform.  stdin (console or pipe) is not a socket, so RAY_SEL_STDIN
selectors are probed directly and the socket wait is sliced while one is
registered.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- mirror WSAGetLastError()/SO_ERROR into errno for send/recv/connect/
  accept/bind/listen: callers branch on EAGAIN/ECONNREFUSED;
- ipc_send_fn always overwrites errno, so a stale EAGAIN can no longer
  park a frame on a dead socket;
- SO_EXCLUSIVEADDRUSE instead of SO_REUSEADDR (which lets a bind steal
  a port that is in use);
- WSAStartup at load time; sock.h pulls in platform.h;
- verbose IPC capture uses a temp file that works without admin rights;
- the SIGURG out-of-band cancel stays POSIX-only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- ray_vm_unmap_file only unmaps at a view's own base: UnmapViewOfFile
  drops the whole view for an interior pointer, which freed columns that
  carry a passenger index (munmap is a no-op there);
- ray_vm_alloc_aligned returns its own allocation base, so pools are
  really released by ray_vm_free;
- ray_vm_map_fd_ro maps for real (CSV reads always failed with io);
- crash report via SetUnhandledExceptionFilter;
- heap: file-backed spill stays POSIX-only (docs/architecture/memory.md);
- domain: realpath substitute; symfile paths compare case/separator-
  insensitively so one file never gets two domains;
- aof/csr/hnsw: platform handle for fsync, portable ray_mkdir.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- term.h/profile.h include platform.h and a lean <windows.h>;
- term_write for the Windows console, errno.h and core count in the REPL;
- .sys.info reports page-size and total-mem on Windows too;
- KEY_READ no longer collides with <winreg.h>.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- runner: ';; @requires: posix' marks .rfl files whose fixtures or checks
  need a POSIX shell/filesystem; on Windows they are reported as SKIP;
- shell-free ray_test_rm_rf / ray_test_mkdir_p replace system("rm -rf")
  and "mkdir -p" in C tests; test.h maps the few POSIX helpers tests use;
- tests of POSIX-only behaviour (setrlimit, ENOTDIR, read-only dirs,
  AF_UNIX, file spill, Winsock send buffering) skip with the reason;
- fixture fixes that were latent on any platform: binary-mode CSV
  fixtures, a per-test AOF dir, journal closed before the crash rename.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
No console font Windows ships has U+2023 (checked Consolas, Cascadia Mono,
Lucida Console, Courier New) and the classic console does no font fallback,
so the prompt rendered as '?'.  Use the nearest filled triangle they all do
have, U+25BA, which is also three UTF-8 bytes.

The banner's CPU line said 'unknown': read ProcessorNameString instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…check

alloc and agg_v2 read peak RSS through getrusage; use
GetProcessMemoryInfo there.  windows_vs_linux.md records the numbers the
port was checked against.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ser-vasilich and others added 15 commits September 23, 2026 14:43
Interpolation (format/println), print and the pivot / column-name helpers
formatted an i64 with "%ld" and a (long) cast.  long is 64-bit on LP64 and
32-bit on Windows, so there 10^18 printed as -1486618624 and a pivot keyed
on values above 2^31 produced truncated column NAMES — a data defect, not
only a display one.  Use PRId64 throughout; regression test included.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…plicing

The banner was spliced from string literals and the RAYFORCE_VERSION /
RAYFORCE_GIT_COMMIT macros inside #ifdef arms; without the -D values a
static analyser reads the literal followed by the bare macro name as two
adjacent tokens and reports a syntax error.  Format it with one snprintf
at install time (the handler itself still never formats), with the
macros defaulting to empty strings.  Output unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`til` is eager, so counting a three-billion-element range built a 24 GB
vector for no extra coverage — the neighbouring literals already cross
2^31, 2^32 and 2^63.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…spatch (#621)

* fix(mem): hand workers their freed blocks back at the end of every dispatch

A parallel operator's workers allocate per-task buffers from their own
heaps and the main thread frees them once the dispatch has completed, so
each block lands on the owning worker's foreign list. The owner drained
that list only when an allocation found its freelists empty, which a
warm worker with a partly cut pool rarely does: it kept splitting fresh
pool space instead, touching new pages every round while its own freed
blocks waited. Under a steady stream of such operators (a join against a
large table per incoming batch) the process grew by a full 32 MB pool
per worker before any block was reused, and only the idle decay, which
needs the process to sit quiet, drained it earlier.

The dispatcher now drains every registered heap's foreign list at the
end of each parallel region (ray_heap_reclaim_workers), under the same
conditions the idle decay already relies on: the flag is clear, so every
worker has finished its last task and neither allocates nor frees until
the next dispatch. The blocks go back to freelists only, no pages are
released, so the next round reuses them without faulting.

Tests: heap/reclaim_workers_drains_owner (a block owned by one heap and
freed from another leaves the owner's list on reclaim and is handed out
again at the same address, no new pool) and
pool/dispatch_reclaims_worker_blocks (blocks allocated by real workers
and freed by the main thread are back with their owners by the end of
the next dispatch).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(mem): reclaim only the pool's own worker heaps after a dispatch

The heap registry holds every thread that ever allocated, not only pool
workers: a server's poll thread or an embedding's own threads are there
too, and they may be allocating while a dispatch on another thread ends.
Draining such a heap's foreign list from the dispatcher coalesces into
freelists their owner is using at that moment. Only a pool worker parked
on its semaphore is known to touch nothing of its own until the next
dispatch, so each worker now publishes its heap in the pool
(worker_heaps, set after ray_heap_init, cleared before it exits) and the
dispatcher reclaims exactly those, plus its own list on its own thread.
ray_heap_reclaim_worker takes one heap; the registry walk is gone, which
also removes its per-dispatch scan of every registry slot.

The heap test now checks the block's bytes leave the owner's books only
on reclaim, instead of expecting the next allocation at the same
address; the pool test reads the worker heaps the pool publishes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(pool,heap): make the reclaim tests independent of worker start-up and adopted heaps

The pool test read a worker's heap slot right after ray_pool_create,
which returns before the workers have started, and passed vacuously when
the main thread took every task. It now waits for each worker to publish
its heap, records which worker allocated each block and repeats the
round until a worker really allocated something, then checks that the
next dispatch left no worker with a foreign list. The heap test drains
whatever an adopted heap already held before taking its baseline.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pool): allocate the worker heap slots without an _Atomic cast

cppcheck cannot build an AST for a cast to `_Atomic(void*)*` and fails
the static-analysis run on it. The cast is not needed in C: ray_sys_alloc
returns void*, which converts to the slot pointer type on its own, and
sizeof(*pool->worker_heaps) names the element size without spelling the
type.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
… like, bulk derived key, parallel dense grouping (#626)
A merged release PR leaves a commit on master that dev lacks, so the next
release PR is blocked as out of date with the base, and after a squash it
also shows conflicts that aren't real. Add the back-merge as step 5.
* fix(system): reject arguments to gc

* docs(memory): use zero-argument gc form

* fix(system): report gc arity errors
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants