localExec(): keep slice results off the scheduler - #50
Closed
dan-distributive wants to merge 1 commit into
Closed
dan-distributive wants to merge 1 commit into
dan-distributive wants to merge 1 commit into
Conversation
localExec() already kept job arguments, input sets, and the work function local, but slice *results* still round-tripped through the real scheduler: the local worker POSTs each computed value to the resultSubmitter service, and the job receives it back via the scheduler's pubsub relay -- the actual value left the machine, not just a completion signal. dcp-client already ships a public API for exactly this class of problem: job.setResultStorage(url, postParams) redirects a slice's result upload to a self-hosted location instead of the scheduler's own storage (normally used for S3, Dropbox, etc. -- see "DCP Job Architecture for data routing.pptx", slides 5-8). Pointing it at a small local HTTP server keeps the real value on loopback; the server responds with a small opaque token, and that token -- not the real value -- is all that travels to the scheduler and back through the completely unmodified resultSubmitter/pubsub path, exactly how the feature already behaves for any other off-prem storage target. Results are captured via KVIN (application/x-kvin), not plain JSON: JSON would silently mangle binary/pickled results and strip a failed slice's error object down to an inert dict, breaking the existing raise_on_first_work_error behavior. Confirmed both cases explicitly: a sentinel value proven to never appear in the scheduler-relayed ResultHandle, and the full existing regression suite (simple job, pycomod with cloudpickled/numpy results via JobFS, all four error scenarios, job.wait() symmetry) passing unchanged. No public API changes -- job.localExec() behaves identically from the caller's side; this is entirely internal plumbing. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Author
|
Closing per a decision made with Ryan: the default scheduler round-trip for slice results is fine as-is. If a job needs results kept off the scheduler, that's already achievable today via the existing, public The other locality fixes (job arguments, input sets, work function source — PR #51) are unaffected by this: those mirror what Node's own |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
localExec()already kept job arguments, input sets, and the work function local (see the parent branch). Slice results were the one remaining gap: the local worker POSTs each computed value to the realresultSubmitterservice, and the job receives it back via the scheduler's pubsub relay — the actual value left the machine, not just a completion signal.This closes that gap using dcp-client's own existing, public
job.setResultStorage(url, postParams)API — normally used to redirect result storage to a self-hosted location (S3, Dropbox, etc; see "DCP Job Architecture for data routing.pptx", slides 5-8). Pointing it at a small local HTTP server keeps the real value on loopback; the server responds with a small opaque token, and that token — not the real value — is all that reaches the scheduler and comes back through the completely unmodifiedresultSubmitter/pubsub path. This is exactly how the feature already behaves for any other off-prem storage target, not a special case.Results are captured via KVIN (
application/x-kvin), not plain JSON — JSON would silently mangle binary/pickled results and strip a failed slice's error object down to an inert dict, which would have broken the existingraise_on_first_work_errorbehavior.No public API changes —
job.localExec()behaves identically from the caller's side; this is entirely internal plumbing, wired intolocalExec()itself (no new methods for callers to remember).Base branch note: this is stacked on
pythonmonkey-platform-supportsince it depends onlocalExec()'s existing local-routing machinery.Verification
ResultHandlecontents, whilejob.localExec()'s returned value is still correct.job.wait()symmetry.Test plan
_setup_local_result_storage()/_grant_send_results_origin()indcp/api/job.pylocalExec()job and confirm results are correctRuntimeErrorwith the real traceback🤖 Generated with Claude Code