v2.9.2 - #630
Closed
singaraiona wants to merge 45 commits into
Closed
v2.9.2#630singaraiona wants to merge 45 commits into
singaraiona wants to merge 45 commits into
Conversation
The native top-N of the radix path kept every group at or beyond the threshold. Groups strictly beyond it number fewer than N, but the groups AT the threshold can be nearly all of them: a count-per-group top-10 over a near-unique key has a threshold of 1, every group ties it, and the "kept superset" is the whole grouping, heap-sorted by first row — a comparison sort over 100M pairs for a query whose answer is ten rows. The emitted prefix is the tied groups' first-seen order, so only the N tied groups with the smallest first row can ever be taken. The selection now keeps the strictly-better groups plus exactly those N, through a bounded max-heap, and sorts at most 2N entries. The rows emitted are the ones the full kept set produced. Tests: the pinned kept count of the radix native top-N case becomes N; a near-unique two-key grouping whose threshold every group ties, in both directions and under a row selection, against the full grouping. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Rayfall had no way to stop iterating before the end of a sequence. Every
iteration primitive — map, pmap, fold, fold-left/right, scan, scan-left/right,
prior — consumes its whole input, and there was no loop construct, so a
"repeat until done" loop had to be written as a fold over a fixed range whose
full length was paid on every call however early the work finished. `return`
did not help: it exits the lambda, not the iteration, so the fold kept calling
the lambda for every remaining element. Recursion was not an alternative
either, since there is no TCE and the stack tops out around 1-2k frames.
(while cond body...) evaluates cond, and while it is truthy evaluates each
body expression in order, then tests again. It always returns null — a
statement form run for effect, and a never-taken loop has no last value to
report. Zero body expressions is legal, so a condition with side effects can
be the whole loop; that is the shape a drain wants, where there is no sequence
to iterate over and the range was only ever scaffolding.
Implemented on both evaluators, which must agree:
- ray_while_fn (eval.c), registered beside `if` and `do`.
- an sf_while case in the bytecode compiler emitting JMPF over a backward
JMP — the first backward branch the compiler produces. The VM's op_jmp
already checked for a pending interrupt when the displacement is negative
(dormant until now), so a runaway loop is Ctrl-C-able for free.
Compiling the body inline, rather than letting it fall through to the generic
special-form path, is what keeps `return` working inside a loop: it reaches
the sf_return case and unwinds the lambda instead of degrading to the tree
walker's identity `return`.
Neither path pushes a scope around the body. `let` binds only in the top
frame (env_bind_local), so a per-iteration frame would discard loop-carried
`let` state on the tree-walking path while the compiled path — whose `let`
writes a function-level slot — kept it, and the two evaluators would disagree
on the same source. A caller wanting a fresh frame per pass writes
(while cond (do ...)), which composes.
patch_jump and emit_jump_back now share write_jump_offset; an out-of-range
displacement still sets c->error, dropping the lambda to the interpreter.
Measured on the drain shape from the issue (4-step drain under a 4096 safety
bound, release build, driver baseline subtracted): fold-left over (til 4096)
1420 us/batch, the nested 64x64 workaround 45.8 us/batch, `while` 3.8
us/batch — 374x over the original and 12x over the workaround.
Closes #588
The two remaining early-termination forms requested in #588, alongside the `while` that landed in #590. Neither is an unblock — `while` already covers the reporter's case — but both complete the vocabulary he asked for. `times` ------- (times n body...) runs the body exactly n times and returns null. `do` is already progn in this language, so the bounded loop could not reuse that name; overloading `do` on an integer head was rejected as genuinely ambiguous — (do 5) would have to mean either "loop five times over nothing" or "return 5", and a computed first expression that happened to be an integer would silently change meaning. The count is evaluated ONCE on entry, so a body mutating whatever produced it cannot change how many passes remain. A count of zero or less runs zero times rather than trapping; a non-integer is a type error. Compiled as a counted loop over a hidden local slot, reusing the backward branch added for `while`. ray_times_norm_fn type-checks the count and clamps a negative bound to zero on entry, which lets the per-pass test be a bare truthiness check on the counter — 0 is falsy, so no comparison call is needed per pass and a negative bound cannot run away. That helper and the decrement are pushed as constant-pool objects rather than resolved by name, so the loop's own arithmetic is unnamable from source and cannot be swapped out by a `(set - ...)` override. The counter's slot is addressed by index and its sym carries a space, so no source token can collide with it and nested `times` counters stay apart. As with `while`, the body compiles inline, so `return` unwinds the enclosing lambda from inside the loop. fold-while ---------- (fold-while pred f init xs) offers the accumulator to `pred` before each step and stops on a falsy answer, yielding the accumulator as it stands. The test precedes the first element, so a predicate false at the start returns `init` untouched. The predicate takes the accumulator rather than the element: that is the form that expresses "iterate until the running result says stop", which is the early termination actually being asked for. One deliberate divergence from ray_fold_fn: that routes its collection through unbox_vec_arg -> to_boxed_list, boxing every element up front. For a primitive whose purpose is to stop early, paying for the tail it never reaches is the cost being removed, so elements are pulled one at a time via collection_elem. A plain variadic builtin — no compiler work, since it dispatches through the normal call path. Measured, release builds: `times` 100 ns/pass against 180 ns for `while` plus a manual counter over 1e6 passes. `fold-while` stopping after three elements of a 1e6-element vector costs 3.9 us against 369 ms for the `fold-left` equivalent, which had to box and walk all million to discover it was done. Tests: test/rfl/lang/times.rfl and test/rfl/collection/fold_while.rfl, each behaviour asserted on both evaluator paths where applicable — including the count validation and negative clamp on the compiled path, which runs through entirely different code from the tree walker's check.
feat(lang): add a `while` special form
fix(group): bound the tied groups the radix top-N selection keeps
feat(lang): add `times` and `fold-while`
ray_vec_is_null is out-of-line (no LTO) and the join called it once per key column per row in hash_row_keys — on the build side, the probe side, and the prefetch lookahead — plus twice per key column per hash-chain step in join_keys_eq, across both the count and the fill pass. A reported profile put 13.28% of a service's samples there, on a join keyed by two SYM columns that structurally never hold a null. Prove once per join that no key column can hold a null and drop the call. SYM/STR nulls are canonical empty payloads (id 0 / length 0) that HAS_NULLS does not track, so text columns are proven by the chunked zero-scan from #533 rather than by the flag; everything else reads the flag through slices and takes a set bit at face value, so the proof stays O(n) and never degrades into ray_vec_has_nulls' per-element walk. Flag-readable columns are settled first, so a nullable numeric key short-circuits before any text column is scanned. OP_CONST (atom) key slots are refused as unprovable. Also skip the #458 null-run pre-scan under the proof. It gates on ray_vec_may_have_nulls, which is unconditionally true for SYM/STR, so a SYM-keyed join ran a full per-key-per-build-row ray_vec_is_null scan before the join proper on every execution; a null-free key set cannot contain an all-null row, so the scan is dead. Measured with bench/join_nullfree (release, 1:1 book join, 4M probe x 500K build): two SYM keys 275 -> 249 ms (-9.5% median, -7.8% min); single I64 key 106 -> 99 ms (-6.3% median, -7.9% min). A proof that fails on a text key costs a partial scan (~8ms on a 4M-row SYM column) and gains nothing — only nullable SYM/STR key columns pay it. ray_join_force_null_checks forces the null-aware loops so the differential tests and the perf gate can compare both paths in one binary; ray_join_nullfree_keys counts the joins that took the fast path. Closes #597
perf(join): skip per-cell key null tests on provably null-free columns
fix(system): validate launcher and timeit integer arguments
Every value-taking startup flag was gated on `i + 1 < argc` and then consumed argv[++i] blindly, so one flag silently ate the next. The shape that was reported is the worst one for a service: rayforce -Q -p 5099 svc.rfl `-Q` took `-p` as its value, the port was never bound, and nothing on stdout or stderr said so. Under a supervisor the process starts, the script runs, the unit looks healthy, and every client gets connection refused. `-c` and `-t` had the identical hole. Two neighbouring cases came out of the same code. A trailing flag with no value fell through to the positional-file arm and reported the misleading `cannot open '-Q'`. And that arm took ANY unrecognized token as the script name, which the real positional then overwrote, so a typo'd option was not merely ignored — it was silently swallowed: `rayforce -x svc.rfl` ran the script as if `-x` had never been typed. Every value-taking flag now goes through flag_value(), which refuses a missing value and a value that is EXACTLY one of this program's flag tokens (including the `--` app-args terminator), with a diagnostic and exit 2 before the script is loaded. A value that merely STARTS with '-' stays legal: a password or a negative number is a legitimate value, and refusing those would break working command lines for no gain. An unrecognized `-`-prefixed token is now an error too; a lone `-` is still a positional, and tokens after `--` belong to the app and are never validated. `-Q, --querylog N` was a real flag, referenced twice in the docs, that --help never listed beside its sibling `-t, --timeit N` — added, along with `[-Q 0|1]` in the synopsis. The suggestion to fail when `-p` was requested but nothing bound is already in (#473, listen_fatal.rfl), and cannot catch this: the parser never saw the eaten `-p`, so there was no request to check against. The guard on the value is what closes the class. Also: the two pre-existing `-p` validation bail-outs now release the runtime like every other error path, since these paths are exercised under ASan by the new test. Closes #600
…ionally Every node the executor evaluates returns an owned reference — the constant table node included, it retains its literal. The OP_WINDOW, OP_SORT and OP_HEAD cases released their input only when it differed from the graph's table, taking an equal pointer for a borrowed one. A query's root is a constant node over that very table, so the reference was never released: one input table per windowed, sorted or limited query. Invisible while a global kept the table alive; a whole table per call when the input was built for that call, as a service that windowed a freshly concatenated buffer on a timer found (#602). The four cases now release the input on every path, as the join and the plain head/tail cases already did. Test: window, sorted and limited selects over a table built per call, measured with the two-window bytes-allocated method of the other memory probes, plus a shared input reused across fifty calls. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… queries Twenty calls of each shape over one global table; the table's refcount must be exactly what it was before, and the table still whole after. Fails on the unfixed executor at the first window shape. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The same borrowed-query-table guard sat under the reductions; no child evaluates to the query table there today, so nothing leaked, but the premise is the one the sort, window and limit cases just dropped. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
fix(exec): release the input table of window, sort and limit queries
fix(cli): reject a flag as another flag's value, and unknown options
ray_pool_dispatch and ray_pool_dispatch_n signalled the whole pool on every dispatch, however narrow the window. The main thread participates as worker 0, so at most n_tasks-1 helpers can ever claim anything; the surplus threads woke, raced to an already-drained window, and went straight back to the semaphore. On a dispatch narrower than the machine that is pure overhead — a partition-parallel step over three partitions woke every core on the box to run two tasks. Signal min(n_tasks-1, n_workers) instead. Signalling FEWER is safe: completion is governed by the `pending` spin-wait, never by the signal count, so an unsignalled worker simply stays asleep, and under-signalling cannot produce the surplus-signal problem between consecutive dispatches that the spin-wait exists to avoid. Measured (10M-row ClickBench, 43 queries x 3, splayed): sum of per-query hot times 5690.7 ms -> 5667.3 ms (-0.4%, within noise), user CPU 55.1 s -> 54.8 s. So this is waste removal rather than a speedup on workloads whose dispatches are already wider than the pool. Not a fix for #599. That report's shape — a 130k-row join probe, which splits into 16 tasks against 8 workers — is unaffected, because the clamp lands on the same worker count it already used (measured: 5.19 s -> 5.18 s of CPU, unchanged). The investigation there points at the pool defaulting to logical rather than physical cores; that is a separate change with its own benchmark trade and is filed on its own. test_pool.c gains pool/dispatch_narrow, which is the case the existing coverage missed: dispatch_n_small uses 4 tasks on a 2-worker pool, so the clamp never binds. The new test drives windows of 1..4 tasks on a 4-worker pool, then alternates narrow and wide windows 25 times over the same pool, asserting exact task and element counts throughout — the two risks the change introduces are a task nobody claims and signal accounting that drifts across dispatches and starves a later wide window.
perf(pool): wake only the workers a dispatch can keep busy
Detect Windows via $(OS) or uname (MSYS2's make hides $(OS)), link Winsock and a statically linked winpthreads, build with 64-bit off_t (_FILE_OFFSET_BITS=64) and an 8 MiB stack like a Linux main thread. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replace the IOCP stub with a readiness-based loop that mirrors the epoll backend's dispatch order, so the selector state machine is identical on every platform. stdin (console or pipe) is not a socket, so RAY_SEL_STDIN selectors are probed directly and the socket wait is sliced while one is registered. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- mirror WSAGetLastError()/SO_ERROR into errno for send/recv/connect/ accept/bind/listen: callers branch on EAGAIN/ECONNREFUSED; - ipc_send_fn always overwrites errno, so a stale EAGAIN can no longer park a frame on a dead socket; - SO_EXCLUSIVEADDRUSE instead of SO_REUSEADDR (which lets a bind steal a port that is in use); - WSAStartup at load time; sock.h pulls in platform.h; - verbose IPC capture uses a temp file that works without admin rights; - the SIGURG out-of-band cancel stays POSIX-only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- ray_vm_unmap_file only unmaps at a view's own base: UnmapViewOfFile drops the whole view for an interior pointer, which freed columns that carry a passenger index (munmap is a no-op there); - ray_vm_alloc_aligned returns its own allocation base, so pools are really released by ray_vm_free; - ray_vm_map_fd_ro maps for real (CSV reads always failed with io); - crash report via SetUnhandledExceptionFilter; - heap: file-backed spill stays POSIX-only (docs/architecture/memory.md); - domain: realpath substitute; symfile paths compare case/separator- insensitively so one file never gets two domains; - aof/csr/hnsw: platform handle for fsync, portable ray_mkdir. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- term.h/profile.h include platform.h and a lean <windows.h>; - term_write for the Windows console, errno.h and core count in the REPL; - .sys.info reports page-size and total-mem on Windows too; - KEY_READ no longer collides with <winreg.h>. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
- runner: ';; @requires: posix' marks .rfl files whose fixtures or checks need a POSIX shell/filesystem; on Windows they are reported as SKIP; - shell-free ray_test_rm_rf / ray_test_mkdir_p replace system("rm -rf") and "mkdir -p" in C tests; test.h maps the few POSIX helpers tests use; - tests of POSIX-only behaviour (setrlimit, ENOTDIR, read-only dirs, AF_UNIX, file spill, Winsock send buffering) skip with the reason; - fixture fixes that were latent on any platform: binary-mode CSV fixtures, a per-test AOF dir, journal closed before the crash rename. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
No console font Windows ships has U+2023 (checked Consolas, Cascadia Mono, Lucida Console, Courier New) and the classic console does no font fallback, so the prompt rendered as '?'. Use the nearest filled triangle they all do have, U+25BA, which is also three UTF-8 bytes. The banner's CPU line said 'unknown': read ProcessorNameString instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…check alloc and agg_v2 read peak RSS through getrusage; use GetProcessMemoryInfo there. windows_vs_linux.md records the numbers the port was checked against. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Interpolation (format/println), print and the pivot / column-name helpers formatted an i64 with "%ld" and a (long) cast. long is 64-bit on LP64 and 32-bit on Windows, so there 10^18 printed as -1486618624 and a pivot keyed on values above 2^31 produced truncated column NAMES — a data defect, not only a display one. Use PRId64 throughout; regression test included. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…plicing The banner was spliced from string literals and the RAYFORCE_VERSION / RAYFORCE_GIT_COMMIT macros inside #ifdef arms; without the -D values a static analyser reads the literal followed by the bare macro name as two adjacent tokens and reports a syntax error. Format it with one snprintf at install time (the handler itself still never formats), with the macros defaulting to empty strings. Output unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
`til` is eager, so counting a three-billion-element range built a 24 GB vector for no extra coverage — the neighbouring literals already cross 2^31, 2^32 and 2^63. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…spatch (#621) * fix(mem): hand workers their freed blocks back at the end of every dispatch A parallel operator's workers allocate per-task buffers from their own heaps and the main thread frees them once the dispatch has completed, so each block lands on the owning worker's foreign list. The owner drained that list only when an allocation found its freelists empty, which a warm worker with a partly cut pool rarely does: it kept splitting fresh pool space instead, touching new pages every round while its own freed blocks waited. Under a steady stream of such operators (a join against a large table per incoming batch) the process grew by a full 32 MB pool per worker before any block was reused, and only the idle decay, which needs the process to sit quiet, drained it earlier. The dispatcher now drains every registered heap's foreign list at the end of each parallel region (ray_heap_reclaim_workers), under the same conditions the idle decay already relies on: the flag is clear, so every worker has finished its last task and neither allocates nor frees until the next dispatch. The blocks go back to freelists only, no pages are released, so the next round reuses them without faulting. Tests: heap/reclaim_workers_drains_owner (a block owned by one heap and freed from another leaves the owner's list on reclaim and is handed out again at the same address, no new pool) and pool/dispatch_reclaims_worker_blocks (blocks allocated by real workers and freed by the main thread are back with their owners by the end of the next dispatch). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(mem): reclaim only the pool's own worker heaps after a dispatch The heap registry holds every thread that ever allocated, not only pool workers: a server's poll thread or an embedding's own threads are there too, and they may be allocating while a dispatch on another thread ends. Draining such a heap's foreign list from the dispatcher coalesces into freelists their owner is using at that moment. Only a pool worker parked on its semaphore is known to touch nothing of its own until the next dispatch, so each worker now publishes its heap in the pool (worker_heaps, set after ray_heap_init, cleared before it exits) and the dispatcher reclaims exactly those, plus its own list on its own thread. ray_heap_reclaim_worker takes one heap; the registry walk is gone, which also removes its per-dispatch scan of every registry slot. The heap test now checks the block's bytes leave the owner's books only on reclaim, instead of expecting the next allocation at the same address; the pool test reads the worker heaps the pool publishes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(pool,heap): make the reclaim tests independent of worker start-up and adopted heaps The pool test read a worker's heap slot right after ray_pool_create, which returns before the workers have started, and passed vacuously when the main thread took every task. It now waits for each worker to publish its heap, records which worker allocated each block and repeats the round until a worker really allocated something, then checks that the next dispatch left no worker with a foreign list. The heap test drains whatever an adopted heap already held before taking its baseline. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix(pool): allocate the worker heap slots without an _Atomic cast cppcheck cannot build an AST for a cast to `_Atomic(void*)*` and fails the static-analysis run on it. The cast is not needed in C: ray_sys_alloc returns void*, which converts to the slot pointer type on its own, and sizeof(*pool->worker_heaps) names the element size without spelling the type. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
… like, bulk derived key, parallel dense grouping (#626)
A merged release PR leaves a commit on master that dev lacks, so the next release PR is blocked as out of date with the base, and after a squash it also shows conflicts that aren't real. Add the back-merge as step 5.
* fix(system): reject arguments to gc * docs(memory): use zero-argument gc form * fix(system): report gc arity errors
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What & why
Checklist
dev(notmaster)feat:/fix:/perf:/docs:/ …)makebuilds cleanly (no new warnings)make testpasses; tests added/updated for behaviour changes