[python] Add temporal alignment for multimodal scans - #9536
Draft
XiaoHongbo-Hope wants to merge 3 commits into
Draft
[python] Add temporal alignment for multimodal scans#9536XiaoHongbo-Hope wants to merge 3 commits into
XiaoHongbo-Hope wants to merge 3 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Multimodal training data is often stored as independently sampled streams. A
caller currently has to collect those streams and implement timestamp matching,
episode isolation, snapshot pinning, and payload lookup itself.
This draft proposes a typed, bounded temporal-alignment API for
MultimodalTable.scan():Source names are caller-defined output namespaces, not schema or domain names.
Using an explicit mapping supports dynamically generated source lists without
coupling this API to a UI specification.
The implementation:
tolerance and deterministic earlier-frame tie breaking;
_ROW_ID, then fetches selectedpayload rows in batches;
decode only selected samples;
This is deliberately a bounded MVP. It does not add interpolation, training
windows, watermark/streaming semantics, a customer-specific JSON spec, or a
generic
ScanQuerybatch-reader API. The metadata index is currentlycoordinator-local. Since this is a draft, feedback on the API shape and whether
the alignment executor belongs in PyPaimon is especially welcome.
Tests