Repository navigation
Performance regression in BufferedReader.readline() after switching to PyBytesWriter #158897
Description
Activity
- addedperformancePerformance or resource usagePerformance or resource usageextension-modulesC modules in the Modules dirC modules in the Modules dir
on Oct 6, 2026 Aha, that's an interesting performance issue :-) I created a benchmark using pyperf from your script:
# Benchmark on io.BufferedReader.readline(): # https://github.com/python/cpython/issues/158897 import pyperf import io BUFFER_SIZE = 4096 medium_a = b"x" * 4080 + b"abc\n" # 4,084 bytes medium_b = b"y" * 4112 + b"def\n" # 4,116 bytes long_line = b"z" * 40960 + b"trailer\n" # 40,968 bytes data = (medium_a + medium_b + long_line) * 64 def bench(reader): reader.seek(0) while reader.readline(): pass runner = pyperf.Runner() with io.BufferedReader(io.BytesIO(data), buffer_size=BUFFER_SIZE) as reader: runner.bench_func('bench', bench, reader)
I compared the main branch to Python 3.15:
Mean +- std dev: [py315] 1.24 ms +- 0.01 ms -> [main] 1.45 ms +- 0.01 ms: 1.18x slower.Oh yes, it seems like the main branch is slower on this specific benchmark.
But when I hacked the code to try to optimize it, I get very surprising results. Benchmark results can shift from 1.45 ms to 1.10 ms (oh good!) to 1.70 ms (what's going on?).
I looked at Linux perf:
Samples: 4K of event 'cpu/cycles/Pu', Event count (approx.): 4722823369 Overhead Command Shared Object Symbol 73.80% python python [.] _buffered_readline 8.82% python libc.so.6 [.] __memmove_avx_unaligned_erms 1.18% python python [.] _PyEval_EvalFrameDefault 0.66% python python [.] _PyObject_Malloc 0.63% python python [.] PyBuffer_FillInfo 0.54% python libc.so.6 [.] _int_malloc 0.51% python python [.] PyBytesWriter_WriteBytes 0.49% python python [.] gc_collect_mainOh wow, 74% of the time is spent in _buffered_readline()!
Moreover, the majority of the _buffered_readline() time is only spent in... just two instructions!
│ ↓ jmp 11e │110: addq $0x1,%rax 49.75 │ cmpb $0xa,-0x1(%rax) 49.32 │ ↓ je 270 │11e: cmpq %rdx,%rax 0.22 │ ↑ jb 110In C, it comes from this loop:
while (s < end) { if (*s++ == '\n') { ... } }
So the majority of the time is not spent in malloc/realloc/free of PyBytesWriter or in Python <=> C layers, but just in searching for the newline character in the 4 kB buffer!
Reacted by Leandro DamascenaSo the majority of the time is not spent in malloc/realloc/free of PyBytesWriter or in Python <=> C layers, but just in searching for the newline character in the 4 kB buffer!
memchr()calling? :-)I can take a stab if you want:
cpython/Modules/_io/bufferedio.c
Lines 1282 to 1293 in 1818fba
start = self->buffer; const char *end = start + n; s = start; while (s < end) { if (*s++ == '\n') { if (PyBytesWriter_WriteBytes(writer, start, s - start) < 0) { goto error; } self->pos = s - start; goto found; } } I have a project where I build CPython into OCI images and run them on AWS Lambda to catch performance changes, and it flagged a regression in BufferedReader.readline() after #158615, which switched it to PyBytesWriter. It shows up when lines don't fit in the buffer and readline() has to refill it: (...)
That's cool to track Python performance! Thanks for tracking our performance :-)
But I'm curious: what do you measure exactly? Your benchmark uses a buffer size of 4 kB which is way smaller than Python 3.16 default buffer size (128 kB): 32x smaller. Moreover, your script creates lines between 4,084 bytes and 40,968 bytes. What is the use case using such unusual long lines? Text files have lines closer to 80-100 bytes.
Reacted by Leandro Damascenamemchr() calling? :-) I can take a stab if you want
Exactly! I already had a PR ready when I posted the comment. You're welcome to review and benchmark it!
Reacted by Maurycy Pawłowski-WierońskiBut I'm curious: what do you measure exactly? Your benchmark uses a buffer size of 4 kB which is way smaller than Python 3.16 default buffer size (128 kB): 32x smaller. Moreover, your script creates lines between 4,084 bytes and 40,968 bytes. What is the use case using such unusual long lines? Text files have lines closer to 80-100 bytes.
Thanks @vstinner, and good question! Some context: I work at AWS, specifically on AWS Lambda and its runtimes, and customers run all kinds of Python workloads there, from small API handlers to data processing jobs. Lambda is also a fairly specific environment (small Firecracker microVMs, different memory sizes, x86_64 and arm64), and since customers pay for execution time, every microsecond counts. That's why I built a project that follows merges across several CPython branches and benchmarks the ones that could have a direct impact on this environment, especially from a performance perspective. The goal is to catch regressions like this one before they reach a release and our customers, which in the end helps the whole Python community too.
A common example is a function triggered by S3 that processes gzipped NDJSON files written by a data pipeline, record by record:
with gzip.open("/tmp/batch.json.gz", "rb") as f: for line in f: record = json.loads(line) ...
gzipwraps the stream in aBufferedReader, so this goes throughreadline(), and records there can easily be tens of KB (events with embedded payloads, stack traces and so on).The 4 KiB buffer in my benchmark was on purpose. It makes the path where a line crosses one or more buffer boundaries easy to measure. You're right that it isn't representative of the 128 KiB default in 3.16, though: with the default buffer this path runs much less often. So I'd describe this as a regression in the multi-buffer
readline()path with a smaller buffer, not as a generalreadline()regression.And thanks again for the fix in #158944. I posted the Lambda results there.
Reacted by Maurycy Pawłowski-Wieroński- added a commit that references this issue
on Oct 7, 2026
Bug report
Bug description:
I have a project where I build CPython into OCI images and run them on AWS Lambda to catch performance changes, and it flagged a regression in
BufferedReader.readline()after #158615, which switched it toPyBytesWriter. It shows up when lines don't fit in the buffer andreadline()has to refill it:The regression is real and repeatable, but how big it is depends a lot on where it runs. The CPU, the allocator and the build options all seem to matter. I tested it in a few different environments, comparing #158615 with its parent:
This also explains why the benchmark in #158615 didn't show any impact: on some machines the new code is as fast or faster. On the x86_64 host where I could run
perf, the profile changes as expected: the parent spends more time inbytes_joinandmemmove, and the new code showsPyBytesWriter_WriteBytes, buffer growth andrealloc. My guess is that the growth andreallocpath is what gets more expensive in the environments where it regresses, but I couldn't prove that.To be fair to the change, peak memory in this scenario went down from 89.7 KiB to 53.6 KiB (measured with
tracemalloc), which is a nice improvement.Since the regression depends on the environment, it can easily go unnoticed in a single benchmark, so I thought it was worth reporting. Happy to run more tests or try a patch. cc @vstinner
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Linked PRs