Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,21 @@

## Unreleased

### Cloud Security — code-scan pushes retry when the service is busy

- `CloudSec.ingest_code_results` (and so `cloudsec code ingest` and
`cloudsec code scan --ingest`) now re-sends a push refused with HTTP 429 —
the organization already has as many pushes in progress as it may
(`error_code: "ingest_busy"`), or the request quota is spent. Nothing is
recorded for a refused push, so the re-send is safe. Each wait honours the
response's `Retry-After` (up to 120 seconds) as a floor, backs off
exponentially without one, and adds random jitter; at most 5 re-sends and
10 minutes of waiting in total. `busy_retries=0` restores the old behaviour.
- `RateLimitError.retry_after` is now filled from the response's `Retry-After`
header when the client raises on the first 429 (without `--retry`, or with
`Client.request(..., retry_quota_errors=False)`, which overrides the client's
`--retry` setting for one call).

### Cloud Security — SBOM route

- `CloudSec.get_code_sbom` / `download_code_sbom` (and `cloudsec code sbom`)
Expand Down
2 changes: 2 additions & 0 deletions doc/cli/cloud-security.md
Original file line number Diff line number Diff line change
Expand Up @@ -308,6 +308,8 @@ limacharlie cloudsec code autofix <FINDING_ID> # open the upgrade PR
limacharlie cloudsec code ingest --repo acme/api --source sarif --file report.sarif
```

`code ingest` (and `code scan --ingest`) retries a push the service answers with HTTP 429, which means the organization already has as many pushes in progress as it may, or the request quota is spent. It waits at least the response's `Retry-After`, adds random jitter so a CI fan-out that was refused together does not come back together, and gives up after 5 retries or 10 minutes of waiting, exiting with the rate-limit error. A refused push recorded nothing, so the retry is safe.

`code repos` reports `scan_status` as `scanned`, `partial` or `unknown`. `partial` means the scan tripped a limit, so the finding set is INCOMPLETE — not a clean bill. `unknown` means this view has no scan state and says so rather than guessing; `code status` is the authoritative view of the run.

`code capabilities` covers **GitHub connections only** — a GitLab or Bitbucket connection scans with its own read-only token and has no write plane to detect, so it never appears, not even as `unknown`. Use `provider manifest` for those. A capability of `available` means the control MAY be offered, not that anything fires on its own.
Expand Down
62 changes: 55 additions & 7 deletions limacharlie/client.py
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,38 @@ def _root_from_env(name: str, default: str) -> str:
HTTP_TOO_MANY_REQUESTS = 429
HTTP_GATEWAY_TIMEOUT = 504


def parse_retry_after(value: str | None, now: float | None = None) -> int | None:
"""Parse an HTTP ``Retry-After`` header into whole seconds from now.

Both forms RFC 9110 allows are read: delay-seconds (``"30"``) and an
HTTP-date. A date in the past is 0. Anything else — absent, empty,
negative, garbage — is ``None``, meaning "the server did not say", which
is different from "retry now".
"""
if value is None:
return None
value = value.strip()
if not value:
return None
if value.isascii() and value.isdigit():
# Clamped here so a hostile or broken header can neither overflow the int
# parser (Python refuses very long digit strings) nor mean "forever".
if len(value) > 9:
return 999_999_999
return int(value)
try:
from email.utils import parsedate_to_datetime

when = parsedate_to_datetime(value)
except (TypeError, ValueError, IndexError, OverflowError):
return None
if when is None or when.tzinfo is None:
return None
current = time.time() if now is None else now
return max(0, int(when.timestamp() - current + 0.999))


# Default limit for response body in debug output. Bodies longer than
# this are truncated with a "[truncated]" marker. Use debug_full_response=True
# on the Client (or --debug-full on the CLI) to disable truncation.
Expand Down Expand Up @@ -614,13 +646,16 @@ def _call_jwt_endpoint(self, auth_data: dict[str, Any]) -> str:

def _rest_call(self, url: str, verb: str, params: dict[str, Any] | None = None, alt_root: str | None = None, query_params: dict[str, Any] | list[tuple[str, str]] | None = None,
raw_body: bytes | None = None, content_type: str | None = None, is_no_auth: bool = False, timeout: int | None = None,
extra_headers: dict[str, str] | None = None, raw_response: bool = False) -> tuple[int, Any]:
extra_headers: dict[str, str] | None = None, raw_response: bool = False,
response_headers: dict[str, str] | None = None) -> tuple[int, Any]:
"""Make a single HTTP request to the API.

Args:
raw_response: Return the response body as decoded text instead of
parsed JSON (for non-JSON endpoints like CSV exports). Error
(non-2xx) bodies are still parsed as JSON when possible.
response_headers: When given, filled with the headers of an error
(non-2xx) response, so the caller can read e.g. ``Retry-After``.

Returns:
tuple: (status_code, response_data)
Expand Down Expand Up @@ -707,8 +742,11 @@ def _rest_call(self, url: str, verb: str, params: dict[str, Any] | None = None,
resp = json.loads(error_str)
except Exception:
resp = error_str
resp_headers = e.headers.items() if hasattr(e, "headers") else []
self._debug_response(e.code, list(resp_headers), error_str)
resp_headers = list(e.headers.items()) if getattr(e, "headers", None) is not None else []
if response_headers is not None:
for h_name, h_val in resp_headers:
response_headers[h_name.lower()] = h_val
self._debug_response(e.code, resp_headers, error_str)
return (e.code, resp)

except ssl.SSLError as e:
Expand All @@ -718,7 +756,7 @@ def _rest_call(self, url: str, verb: str, params: dict[str, Any] | None = None,
def request(self, verb: str, url: str, params: dict[str, Any] | None = None, alt_root: str | None = None, query_params: dict[str, Any] | list[tuple[str, str]] | None = None,
raw_body: bytes | None = None, content_type: str | None = None, is_no_auth: bool = False,
max_retries: int = 3, timeout: int | None = None, extra_headers: dict[str, str] | None = None,
raw_response: bool = False) -> Any:
raw_response: bool = False, retry_quota_errors: bool | None = None) -> Any:
"""Make an API request with retry logic and JWT management.

Args:
Expand All @@ -735,13 +773,18 @@ def request(self, verb: str, url: str, params: dict[str, Any] | None = None, alt
extra_headers: Additional HTTP headers to include.
raw_response: Return the response body as decoded text instead
of parsed JSON (for non-JSON endpoints like CSV exports).
retry_quota_errors: Override the client's ``is_retry_quota_errors``
for this call. ``False`` raises :class:`RateLimitError` on the
first 429, for a caller that runs its own backoff.

Returns:
dict: Parsed JSON response (or str when raw_response is set).

Raises:
AuthenticationError: on 401 after JWT refresh.
RateLimitError: on 429 when retry is not enabled.
RateLimitError: on 429 when retry is not enabled. Its
``retry_after`` carries the response's ``Retry-After`` in
seconds, or ``None`` when the response had none.
ApiError: on other non-200 responses after retries.
"""
has_auth_refreshed = False
Expand All @@ -753,16 +796,20 @@ def request(self, verb: str, url: str, params: dict[str, Any] | None = None, alt
else:
self.refresh_jwt()

if retry_quota_errors is None:
retry_quota_errors = self._is_retry_quota_errors

retries = 0
while retries < max_retries:
retries += 1

error_headers: dict[str, str] = {}
code, data = self._rest_call(
url, verb, params=params, alt_root=alt_root,
query_params=query_params, raw_body=raw_body,
content_type=content_type, is_no_auth=is_no_auth,
timeout=timeout, extra_headers=extra_headers,
raw_response=raw_response,
raw_response=raw_response, response_headers=error_headers,
)

if code == HTTP_OK:
Expand Down Expand Up @@ -798,13 +845,14 @@ def request(self, verb: str, url: str, params: dict[str, Any] | None = None, alt
break

if code == HTTP_TOO_MANY_REQUESTS:
if self._is_retry_quota_errors:
if retry_quota_errors:
wait = min(10 * (2 ** (retries - 1)), 60)
self._debug(f"Rate limited, waiting {wait}s before retry...")
time.sleep(wait)
continue
raise RateLimitError(
f"Rate limit exceeded: {data}",
retry_after=parse_retry_after(error_headers.get("retry-after")),
code=code,
)

Expand Down
95 changes: 92 additions & 3 deletions limacharlie/sdk/cloudsec.py
Original file line number Diff line number Diff line change
Expand Up @@ -66,12 +66,14 @@

import base64
import json
import random
import re
import time
from typing import Any, TYPE_CHECKING
from urllib.parse import quote as _quote
from urllib.request import urlopen as _urlopen

from ..errors import AuthenticationError
from ..errors import AuthenticationError, RateLimitError

if TYPE_CHECKING:
from .organization import Organization
Expand Down Expand Up @@ -124,6 +126,41 @@ def _query_pairs(**params: Any) -> list[tuple[str, str]]:
return pairs


# How a code-scan push backs off when it is told to come back later (429).
#
# A 429 on the ingest means the push was NOT processed: either the org already has as
# many pushes running and queued on the ingest service as it may ("ingest_busy", which
# carries a Retry-After), or the caller's per-identity request quota is spent. Both are
# refused before the document is processed, so nothing was recorded and re-sending the
# same push is safe. What matters is WHEN: a CI fan-out that is refused
# together must not come back together, so every wait is jittered, and the whole thing
# is bounded so a job that cannot get in fails instead of hanging.
INGEST_BUSY_RETRIES = 5
# The first backoff when the server gives no Retry-After, doubled per attempt up to the cap.
_INGEST_BUSY_BASE_S = 5.0
_INGEST_BUSY_CAP_S = 60.0
# A Retry-After past this is clamped: one header must not park a CI job for an hour.
_INGEST_RETRY_AFTER_MAX_S = 120.0
# Jitter adds up to this fraction of the wait on top of it, never below it, so a
# Retry-After is always honoured as a floor.
_INGEST_JITTER = 0.5
# The most a push waits in total across its retries. Past it, the 429 is raised.
_INGEST_BUSY_BUDGET_S = 600.0


def _ingest_busy_delay(attempt: int, retry_after: int | None, rand: Any = random) -> float:
"""Seconds to wait before retry number ``attempt + 1`` of a refused push.

The exponential backoff, raised to the server's ``Retry-After`` when it
asks for longer (clamped to :data:`_INGEST_RETRY_AFTER_MAX_S`), plus up to
:data:`_INGEST_JITTER` of that on top.
"""
delay = min(_INGEST_BUSY_CAP_S, _INGEST_BUSY_BASE_S * (2 ** attempt))
if retry_after is not None:
delay = max(delay, min(float(retry_after), _INGEST_RETRY_AFTER_MAX_S))
return delay + rand.uniform(0.0, delay * _INGEST_JITTER)


def _inventory_account_selector(
account_empty: bool | None,
account_unscoped: bool | None,
Expand Down Expand Up @@ -772,14 +809,20 @@ def _post(
query_params: list[tuple[str, str]] | None = None,
*,
raw_response: bool = False,
raw_body: bytes | None = None,
retry_quota_errors: bool | None = None,
) -> Any:
kwargs: dict[str, Any] = {}
if retry_quota_errors is not None:
kwargs["retry_quota_errors"] = retry_quota_errors
return self._org.client.request(
"POST",
f"cloudsec/{self.oid}/{path}",
query_params=query_params or None,
raw_body=json.dumps(body).encode(),
raw_body=raw_body if raw_body is not None else json.dumps(body).encode(),
content_type="application/json",
raw_response=raw_response,
**kwargs,
)

# ------------------------------------------------------------------
Expand Down Expand Up @@ -2873,6 +2916,7 @@ def ingest_code_results(
ref: str | None = None,
default_branch: str | None = None,
provider: str | None = None,
busy_retries: int = INGEST_BUSY_RETRIES,
) -> dict[str, Any]:
"""Push results your own pipeline produced for one repository.

Expand Down Expand Up @@ -2906,6 +2950,19 @@ def ingest_code_results(
sending for a repository LimaCharlie does not collect —
nothing else can state it there, and it is left unset rather
than guessed when you do not know it.
busy_retries: how many times to re-send the push when it is
refused with 429 — the organization already has as many
pushes in progress as it may (``error_code: "ingest_busy"``),
or the request quota is spent. Nothing was recorded for a
refused push, so re-sending it is safe. Each wait honours the
response's ``Retry-After`` (up to 120 seconds) as a floor, backs
off exponentially otherwise, and adds random jitter so a
refused CI fan-out does not come back as one burst; the waits
total at most 10 minutes. ``0`` raises on the first 429.

Raises:
RateLimitError: the push was still refused after ``busy_retries``
re-sends, or the next wait would pass the 10-minute budget.

Returns:
``{"result": {...}}`` — what landed: ``findings``, the SoR
Expand Down Expand Up @@ -2964,7 +3021,39 @@ def ingest_code_results(
body["default_branch"] = default_branch
if provider is not None:
body["provider"] = provider
return self._post("code/ingest", body)
# Serialized once: the document can be tens of megabytes, and a retry re-sends
# the same bytes.
raw = json.dumps(body).encode()
waited = 0.0
attempt = 0
while True:
try:
# The client's own 429 retry is off here: it would neither honour the
# Retry-After nor jitter, and stacked under this loop it would multiply
# the attempts.
return self._post("code/ingest", body, raw_body=raw, retry_quota_errors=False)
except RateLimitError as e:
delay = _ingest_busy_delay(attempt, e.retry_after)
if attempt >= busy_retries or waited + delay > _INGEST_BUSY_BUDGET_S:
if attempt == 0:
raise
# Not the default "use --retry" hint: --retry does not change this
# path, which has already backed off.
raise RateLimitError(
e.raw_message,
retry_after=e.retry_after,
suggestion=(
f"The push was refused {attempt + 1} times over {waited:.0f}s. "
"The organization has too many pushes in progress; push fewer "
"at once, or retry this one later."),
code=e.status_code,
) from e
self._org.client._debug(
f"code ingest refused (429), retrying in {delay:.1f}s "
f"({attempt + 1}/{busy_retries})")
time.sleep(delay)
waited += delay
attempt += 1

def get_code_capabilities(self, *, repo: str | None = None) -> dict[str, Any]:
"""What each connected source-control organization may actually DO
Expand Down
Loading
Loading