Skip to content

Repeated read() on builtin DCPS topics grows RSS ~70 bytes per sample, never released #343

Description

@tewbacca

Summary

Calling read() repeatedly on the builtin DCPSParticipant / DCPSPublication
/ DCPSSubscription readers grows resident memory in proportion to the number
of samples returned
, not to the number of distinct instances in the domain, and
the memory is never released. The samples are discarded by the caller
immediately.

This is native growth, not Python objects: across the same period
len(gc.get_objects()) stays flat (~38k) and tracemalloc shows no
corresponding growth, while RSS climbs linearly.

Polling the builtin topics to keep a liveness view fresh is the natural way to
write a status or topology display, and read() (rather than take()) is the
documented choice there, since builtin topics only publish on change. Doing that
at a few hundred samples per second costs ~65 MB/h, indefinitely.

Reproduction

# repro_builtin_read_leak.py — full version attached below
readers = [BuiltinDataReader(dp, t) for t in (
    BuiltinTopicDcpsParticipant, BuiltinTopicDcpsPublication,
    BuiltinTopicDcpsSubscription)]
while True:
    for r in readers:
        got = r.read(N=400)      # same live set every pass
        del got
    time.sleep(0.4)

Run python3 repro_builtin_read_leak.py plain in a domain with a few
participants. It prints RSS, cumulative samples read, and bytes per sample every
30 s.

Measured

cyclonedds-python 0.10.5, Python 3.9.25, ~170 builtin instances in the domain
(~390 samples/s re-read):

platform growth samples over
aarch64, Ubuntu 24.04, kernel 6.17 +2.2 MB 34,654 90 s
x86_64, Rocky 9.8, kernel 5.14 +2.4 MB 35,694 90 s

Linear, ~65-70 bytes per sample read, no plateau over 8 h of observation
(34 MB → 327 MB before the process was killed).

What it is not

  • Not instance accumulation — growth tracks samples returned, and the instance
    count in the domain is constant.
  • Not the QoS walk: a variant that also iterates sample.qos looking for a
    __Hostname property measured the same rate, so deserialising the QoS is not
    implicated.
  • Not the data path: the same loop shape with take() on ordinary
    (non-builtin) keyed topics is flat after warm-up, as is write() on a type
    with a nested sequence — 25,000 distinct instances cost +0.3 MB.

Workaround

Read through a NotRead ReadCondition, so each sample is returned once:

unread = SampleState.NotRead | ViewState.Any | InstanceState.Any
cond = ReadCondition(reader, unread)
got = reader.read(N=400, condition=cond)

With this, a quiet domain returns 0 samples per pass and RSS is flat
(python3 repro_builtin_read_leak.py cond: 165 samples at startup, then
nothing). The cost is that a still-alive participant no longer refreshes any
liveness timestamp, so an occasional full read() is still needed as a
backstop — which reintroduces the leak at a rate proportional to how often that
runs.

Environment

  • cyclonedds-python 0.10.5 (pip), CycloneDDS 0.10.5 built from the matching tag
  • Python 3.9.25
  • Linux aarch64 (Ubuntu 24.04.4) and x86_64 (Rocky 9.8)
  • Unicast-only discovery, AllowMulticast false, single interface configured
Full reproduction script
#!/usr/bin/env python3
"""Minimal reproduction: cyclonedds-python leaks native memory per builtin read.

Self-contained — needs only `pip install cyclonedds`, no IDL and no other
participants. Two modes, same loop, same reader:

    python3 repro_builtin_read_leak.py plain   # read() every 0.4s  -> RSS climbs linearly
    python3 repro_builtin_read_leak.py cond    # read(condition=NotRead) -> RSS flat

`plain` re-reads the same live set every pass, which is what a status/topology
view does to keep liveness fresh. Each pass returns the same samples and the
process discards them immediately, yet RSS grows in proportion to the number of
samples read — about 50 bytes per sample — and never comes back.

Measured on cyclonedds-python 0.10.5, Python 3.9.25, with ~170 builtin instances
in the domain (~390 samples/s re-read):

    aarch64, Ubuntu 24.04, kernel 6.17    +2.2 MB / 34,654 samples over 90 s
    x86_64,  Rocky 9.8,    kernel 5.14    +2.4 MB / 35,694 samples over 90 s

~65 MB/h, linear, indefinitely. A status service doing this was OOMKilled every
~6 h on a 256Mi limit.

`cond` is the workaround, not a fix: with a NotRead ReadCondition the reader
returns each sample once, so a quiet domain reads nothing and RSS stays flat.
That loses the liveness refresh, which then needs a slow periodic full read.

To see it faster, raise the instance count by starting a few dozen participants
in the same domain before running this; growth scales with samples read.
"""
import sys
import time

from cyclonedds.domain import DomainParticipant
from cyclonedds.builtin import (
    BuiltinDataReader,
    BuiltinTopicDcpsParticipant,
    BuiltinTopicDcpsPublication,
    BuiltinTopicDcpsSubscription,
)
from cyclonedds.core import InstanceState, ReadCondition, SampleState, ViewState

MODE = sys.argv[1] if len(sys.argv) > 1 else "plain"
SECONDS = int(sys.argv[2]) if len(sys.argv) > 2 else 600


def rss_mb() -> float:
    """Linux only; this is a native-allocation report, so no Python-level
    measurement (tracemalloc, gc) would show it — that is the point."""
    with open("/proc/self/status") as f:
        for line in f:
            if line.startswith("VmRSS:"):
                return int(line.split()[1]) / 1024.0
    return 0.0


def main() -> int:
    dp = DomainParticipant(0)
    readers = [BuiltinDataReader(dp, t) for t in (
        BuiltinTopicDcpsParticipant,
        BuiltinTopicDcpsPublication,
        BuiltinTopicDcpsSubscription,
    )]
    conds = None
    if MODE == "cond":
        unread = SampleState.NotRead | ViewState.Any | InstanceState.Any
        conds = [ReadCondition(r, unread) for r in readers]

    start = time.time()
    base = None
    samples = 0
    reported = -1
    while time.time() - start < SECONDS:
        for i, r in enumerate(readers):
            got = r.read(N=400, condition=conds[i]) if conds else r.read(N=400)
            samples += len(got)
            del got
        time.sleep(0.4)

        elapsed = int(time.time() - start)
        if elapsed % 30 == 0 and elapsed != reported:
            reported = elapsed
            cur = rss_mb()
            if base is None:
                base = cur
            # Bytes-per-sample is only meaningful while the sample count is
            # actually climbing. In `cond` mode it freezes after the first pass,
            # so dividing warm-up by a constant would invent a huge figure.
            if samples > 1000:
                rate = "%5.0f B/sample" % ((cur - base) * 1048576 / samples)
            else:
                rate = "(%d samples total — nothing re-read)" % samples
            print("%-5s t=%4ds rss=%7.1fMB delta=%+6.1fMB samples=%8d %s"
                  % (MODE, elapsed, cur, cur - base, samples, rate), flush=True)
    return 0


if __name__ == "__main__":
    sys.exit(main())

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions