Skip to content

Performance regression in BufferedReader.readline() after switching to PyBytesWriter #158897

Description

@leandrodamascena

Bug report

Bug description:

I have a project where I build CPython into OCI images and run them on AWS Lambda to catch performance changes, and it flagged a regression in BufferedReader.readline() after #158615, which switched it to PyBytesWriter. It shows up when lines don't fit in the buffer and readline() has to refill it:

import io

BUFFER_SIZE = 4096

medium_a = b"x" * 4080 + b"abc\n"        # 4,084 bytes
medium_b = b"y" * 4112 + b"def\n"        # 4,116 bytes
long_line = b"z" * 40960 + b"trailer\n"  # 40,968 bytes

data = (medium_a + medium_b + long_line) * 64

reader = io.BufferedReader(io.BytesIO(data), buffer_size=BUFFER_SIZE)
while reader.readline():
    pass

The regression is real and repeatable, but how big it is depends a lot on where it runs. The CPU, the allocator and the build options all seem to matter. I tested it in a few different environments, comparing #158615 with its parent:

Environment PGO+LTO Result
x86_64 VM, mixed lines yes ~43% slower (also in reverse order, and in Docker with the same images)
x86_64 VM, mixed lines no 18% slower
x86_64 VM, long lines only no 23% slower
x86_64 VM, medium lines only no 16% faster
x86_64 host no ~7% faster
arm64 VM yes / no no significant change
arm64 host yes / no faster

This also explains why the benchmark in #158615 didn't show any impact: on some machines the new code is as fast or faster. On the x86_64 host where I could run perf, the profile changes as expected: the parent spends more time in bytes_join and memmove, and the new code shows PyBytesWriter_WriteBytes, buffer growth and realloc. My guess is that the growth and realloc path is what gets more expensive in the environments where it regresses, but I couldn't prove that.

To be fair to the change, peak memory in this scenario went down from 89.7 KiB to 53.6 KiB (measured with tracemalloc), which is a nice improvement.

Since the regression depends on the environment, it can easily go unnoticed in a single benchmark, so I thought it was worth reporting. Happy to run more tests or try a patch. cc @vstinner

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux

Linked PRs

Activity

  1. vstinner commented on Oct 7, 2026

    @vstinner
    Member

    Aha, that's an interesting performance issue :-) I created a benchmark using pyperf from your script:

    # Benchmark on io.BufferedReader.readline():
    # https://github.com/python/cpython/issues/158897
    
    import pyperf
    import io
    
    BUFFER_SIZE = 4096
    
    medium_a = b"x" * 4080 + b"abc\n"        # 4,084 bytes
    medium_b = b"y" * 4112 + b"def\n"        # 4,116 bytes
    long_line = b"z" * 40960 + b"trailer\n"  # 40,968 bytes
    
    data = (medium_a + medium_b + long_line) * 64
    
    def bench(reader):
        reader.seek(0)
        while reader.readline():
            pass
    
    runner = pyperf.Runner()
    with io.BufferedReader(io.BytesIO(data), buffer_size=BUFFER_SIZE) as reader:
        runner.bench_func('bench', bench, reader)

    I compared the main branch to Python 3.15: Mean +- std dev: [py315] 1.24 ms +- 0.01 ms -> [main] 1.45 ms +- 0.01 ms: 1.18x slower.

    Oh yes, it seems like the main branch is slower on this specific benchmark.

    But when I hacked the code to try to optimize it, I get very surprising results. Benchmark results can shift from 1.45 ms to 1.10 ms (oh good!) to 1.70 ms (what's going on?).

    I looked at Linux perf:

    Samples: 4K of event 'cpu/cycles/Pu', Event count (approx.): 4722823369
    Overhead  Command  Shared Object         Symbol
      73.80%  python   python                [.] _buffered_readline
       8.82%  python   libc.so.6             [.] __memmove_avx_unaligned_erms
       1.18%  python   python                [.] _PyEval_EvalFrameDefault
       0.66%  python   python                [.] _PyObject_Malloc
       0.63%  python   python                [.] PyBuffer_FillInfo
       0.54%  python   libc.so.6             [.] _int_malloc
       0.51%  python   python                [.] PyBytesWriter_WriteBytes
       0.49%  python   python                [.] gc_collect_main
    

    Oh wow, 74% of the time is spent in _buffered_readline()!

    Moreover, the majority of the _buffered_readline() time is only spent in... just two instructions!

            │     ↓ jmp     11e
    
            │110:   addq    $0x1,%rax
      49.75 │       cmpb    $0xa,-0x1(%rax)
      49.32 │     ↓ je      270
            │11e:   cmpq    %rdx,%rax
       0.22 │     ↑ jb      110
    
    

    In C, it comes from this loop:

           while (s < end) {
               if (*s++ == '\n') {
                   ...
               }
           }

    So the majority of the time is not spent in malloc/realloc/free of PyBytesWriter or in Python <=> C layers, but just in searching for the newline character in the 4 kB buffer!

  2. maurycy commented on Oct 7, 2026

    @maurycy
    Contributor

    @vstinner

    So the majority of the time is not spent in malloc/realloc/free of PyBytesWriter or in Python <=> C layers, but just in searching for the newline character in the 4 kB buffer!

    memchr() calling? :-)

    I can take a stab if you want:

    start = self->buffer;
    const char *end = start + n;
    s = start;
    while (s < end) {
    if (*s++ == '\n') {
    if (PyBytesWriter_WriteBytes(writer, start, s - start) < 0) {
    goto error;
    }
    self->pos = s - start;
    goto found;
    }
    }

  3. vstinner commented on Oct 7, 2026

    @vstinner
    Member

    I have a project where I build CPython into OCI images and run them on AWS Lambda to catch performance changes, and it flagged a regression in BufferedReader.readline() after #158615, which switched it to PyBytesWriter. It shows up when lines don't fit in the buffer and readline() has to refill it: (...)

    That's cool to track Python performance! Thanks for tracking our performance :-)

    But I'm curious: what do you measure exactly? Your benchmark uses a buffer size of 4 kB which is way smaller than Python 3.16 default buffer size (128 kB): 32x smaller. Moreover, your script creates lines between 4,084 bytes and 40,968 bytes. What is the use case using such unusual long lines? Text files have lines closer to 80-100 bytes.

  4. vstinner commented on Oct 7, 2026

    @vstinner
    Member

    memchr() calling? :-) I can take a stab if you want

    Exactly! I already had a PR ready when I posted the comment. You're welcome to review and benchmark it!

  5. leandrodamascena commented on Oct 7, 2026

    @leandrodamascena
    Author

    But I'm curious: what do you measure exactly? Your benchmark uses a buffer size of 4 kB which is way smaller than Python 3.16 default buffer size (128 kB): 32x smaller. Moreover, your script creates lines between 4,084 bytes and 40,968 bytes. What is the use case using such unusual long lines? Text files have lines closer to 80-100 bytes.

    Thanks @vstinner, and good question! Some context: I work at AWS, specifically on AWS Lambda and its runtimes, and customers run all kinds of Python workloads there, from small API handlers to data processing jobs. Lambda is also a fairly specific environment (small Firecracker microVMs, different memory sizes, x86_64 and arm64), and since customers pay for execution time, every microsecond counts. That's why I built a project that follows merges across several CPython branches and benchmarks the ones that could have a direct impact on this environment, especially from a performance perspective. The goal is to catch regressions like this one before they reach a release and our customers, which in the end helps the whole Python community too.

    A common example is a function triggered by S3 that processes gzipped NDJSON files written by a data pipeline, record by record:

    with gzip.open("/tmp/batch.json.gz", "rb") as f:
        for line in f:
            record = json.loads(line)
            ...

    gzip wraps the stream in a BufferedReader, so this goes through readline(), and records there can easily be tens of KB (events with embedded payloads, stack traces and so on).

    The 4 KiB buffer in my benchmark was on purpose. It makes the path where a line crosses one or more buffer boundaries easy to measure. You're right that it isn't representative of the 128 KiB default in 3.16, though: with the default buffer this path runs much less often. So I'd describe this as a regression in the multi-buffer readline() path with a smaller buffer, not as a general readline() regression.

    And thanks again for the fix in #158944. I posted the Lambda results there.

  6. added a commit that references this issue on Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    extension-modulesC modules in the Modules dirperformancePerformance or resource usage

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions