Skip to content

streamable-http-client: RequestHandle::cancel hangs before response stream starts #1193

Description

@lucarlig

Describe the bug

With protocol version 2026-07-28 over Streamable HTTP, RequestHandle::cancel() can hang while the original request POST is waiting for its first response event.

The client worker awaits post_message_with_max_sse_event_size(...) inline. While that future is pending, the worker cannot process the cancellation message that should close the request's HTTP/SSE response path. A long-running tool that emits no events therefore cannot be cancelled through RMCP's public client API.

This reproduces on current main (6f8dcde, rmcp 3.1.3).

To Reproduce

  1. Start a stateless Streamable HTTP server using protocol version 2026-07-28.
  2. Make its tools/call handler wait for RequestContext::ct without returning a response or progress event.
  3. Connect with StreamableHttpClientTransport and ClientLifecycleMode::Discover.
  4. Start the call with send_cancellable_request(...).
  5. After the handler starts, call RequestHandle::cancel(...) inside a five-second timeout.

The cancel call times out:

RequestHandle::cancel should not wait for the response stream: Elapsed(())

The regression test in the proposed PR exercises this exact flow using RMCP on both sides.

Expected behavior

RequestHandle::cancel() should return promptly. RMCP should abort the pending request POST, closing that request's response path so the server cancels the handler and fires RequestContext::ct.

This is the required behavior for the modern protocol:

Logs

No transport error is logged. The cancellation future remains pending until the test timeout.

Additional context

The modern client path already converts a cancellation message into cancellation of a per-request CancellationToken, so this appears to be an ordering bug rather than an intentional limitation. The token is currently registered only after the POST produces an SSE response, which is too late to cancel a POST still waiting for its first response event.

Issues #857 and PR #967 fixed the matching server-side disconnect handling. This issue is the client-side path needed to trigger that behavior through RequestHandle::cancel().

Activity

  1. added
    P1High: significant functionality gap or spec violation
    needs confirmationBug report that needs verification from a maintainer
    T-transportTransport layer changes
    on Aug 28, 2026
  2. LizunovSergey commented on Sep 12, 2026

    @LizunovSergey

    This looks fixed on current main (b86995a) — by #1186, which merged on 2026-08-24, five days after this was filed. That would explain why it is still open with the needs confirmation label.

    The ordering the issue describes has been inverted. Worker::post_request now takes the request's cancellation token before issuing the POST and races the two:

    let cancellation = send_request
        .cancellation_token()
        .unwrap_or_else(|| transport_cancellation.child_token());
    Box::pin(async move {
        let response = tokio::select! {
            biased;
            _ = cancellation.cancelled() => None,
            _ = send_request.responder.closed() => None,
            ...
            response = client.post_message_with_max_sse_event_size(...) => Some(response),
        };

    So the token is no longer "registered only after the POST produces an SSE response" — a POST still waiting for its first response event is cancellable.

    Check I ran. Using the repo's own scripted-transport harness in tests/test_streamable_http_client_concurrency.rs, which lets the test hold a request POST pending, with ClientLifecycleMode::Discover { preferred_versions: vec![ProtocolVersion::V_2026_07_28] } to match the conditions here:

    let mut harness = Harness::with_lifecycle(
        config(),
        ClientLifecycleMode::Discover { preferred_versions: vec![ProtocolVersion::V_2026_07_28] },
    ).await?;
    let hanging = harness.cancellable("hanging").await?;
    let blocked = harness.next().await; // POST in flight, reply deliberately withheld
    
    timeout(Duration::from_secs(5), hanging.cancel(None))
        .await
        .expect("RequestHandle::cancel should not wait for the response stream")?;

    That passes — cancel returns immediately instead of timing out after five seconds.

    Two caveats, so this is confirmation rather than proof: it drives a scripted StreamableHttpClient rather than a live Axum listener, so it pins the client-worker ordering (which is where the issue located the bug) and not the server's reaction to the closed stream; and I did not test against the @webclaw/mcp-style setup in the original repro.

    Worth noting that no existing test names this scenario — the closest are cancellation_still_runs_while_recovery_waits_for_old_posts and cancellation_bypasses_queued_posts_at_capacity, both of which pin cancellation against other traffic rather than against the request's own pending POST. If you would like the case above locked in as a regression test so it cannot silently return, I am happy to send that as a small PR — say the word and I will open one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1High: significant functionality gap or spec violationT-transportTransport layer changesneeds confirmationBug report that needs verification from a maintainer

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions