storage: unit-test harness for the source persist sink - #38933
Draft
martykulma wants to merge 1 commit into
Draft
martykulma wants to merge 1 commit into
martykulma wants to merge 1 commit into
Conversation
Adds a timely-driven harness that scripts descriptions and data into write_batches, appends the emitted batches for real under persist's part bounds validation, and reads the shard back. Tests pin the sink's current behavior: one batch per timestamp, one batch per description, and no batch for a description that covers nothing. write_batches takes a SourceStatistics rather than the whole StorageState, so it can be built outside a dataflow that has one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
martykulma
added this pull request to stack #38937
September 19, 2026 00:30
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The sink had no unit tests, so its behavior only showed up in the shape of the
shards it wrote.
Adds a timely harness that scripts descriptions and data into
write_batches,actually appends the batches under persist's part bounds validation, and reads
the shard back. Appending for real matters here, declared bounds vs append
bounds is exactly where this operator goes wrong.
Tests pin current behavior rather than changing it. One batch per timestamp, one
per description, none for a description covering nothing.
write_batchestakesSourceStatisticsinstead of the wholeStorageStatesoit can be built outside a dataflow.
🤖 Generated with Claude Code