What was observed
CanOpenSdoCorrectnessTests.Sdo_BlockUpload_Server_Deadline_Measures_Peer_Idle_Time_Not_The_Whole_Transfer failed once, on the ubuntu-latest leg of #234 (run 36613501052, head bce23d0): Failed: 1, Passed: 1375, Total: 1376, on net10.0.
That PR only bumps github/codeql-action/analyze in a workflow file, so no .cs change is involved. The failing assertion is the NotBe(SdoFrames.CsAbort) in the segment loop, which reads "the server must not time out a transfer whose peer keeps answering": the server sent an SDO abort mid-transfer. I only saw the log's tail, not the abort code or the segment index at which it happened, so I do not know which gap it was.
What is not known
Whether this is host perturbation or a product defect. The test's own clock analysis budgets 1.9 s of slack per gap (SdoServerTimeout 2 s against a 100 ms gap), which is about eight times the ~250 ms of actor starvation that #92 recorded. A starvation explanation therefore needs a much larger stall than anything recorded so far, and this is not the "tolerance too tight" case from #92. The failure is also in the area of #17, where the block-upload deadline was not re-armed at all. A race in the re-arm would produce the same abort.
Reproduction attempts
Not reproduced locally: 8 isolated runs and 6 runs with four busy-loop processes competing for the CPU (Linux container, net10.0), 14 of 14 passed. The full suite also passed once (1378 tests). That does not rule anything out.
Suggested next step
Attach the abort code and segment index to the assertion message (seg[0].Should().NotBe(...) currently fails without saying which of the 26 gaps it was), so the next occurrence says whether the deadline fired early in the transfer, which would point to a re-arm defect, or somewhere at random, which would point to the host. If it recurs at the same segment index, treat it as a product bug.
What was observed
CanOpenSdoCorrectnessTests.Sdo_BlockUpload_Server_Deadline_Measures_Peer_Idle_Time_Not_The_Whole_Transferfailed once, on theubuntu-latestleg of #234 (run 36613501052, headbce23d0):Failed: 1, Passed: 1375, Total: 1376, on net10.0.That PR only bumps
github/codeql-action/analyzein a workflow file, so no.cschange is involved. The failing assertion is theNotBe(SdoFrames.CsAbort)in the segment loop, which reads "the server must not time out a transfer whose peer keeps answering": the server sent an SDO abort mid-transfer. I only saw the log's tail, not the abort code or the segment index at which it happened, so I do not know which gap it was.What is not known
Whether this is host perturbation or a product defect. The test's own clock analysis budgets 1.9 s of slack per gap (
SdoServerTimeout2 s against a 100 ms gap), which is about eight times the ~250 ms of actor starvation that #92 recorded. A starvation explanation therefore needs a much larger stall than anything recorded so far, and this is not the "tolerance too tight" case from #92. The failure is also in the area of #17, where the block-upload deadline was not re-armed at all. A race in the re-arm would produce the same abort.Reproduction attempts
Not reproduced locally: 8 isolated runs and 6 runs with four busy-loop processes competing for the CPU (Linux container, net10.0), 14 of 14 passed. The full suite also passed once (1378 tests). That does not rule anything out.
Suggested next step
Attach the abort code and segment index to the assertion message (
seg[0].Should().NotBe(...)currently fails without saying which of the 26 gaps it was), so the next occurrence says whether the deadline fired early in the transfer, which would point to a re-arm defect, or somewhere at random, which would point to the host. If it recurs at the same segment index, treat it as a product bug.