Repository navigation
Optimize UTF-8 decoding to UCS1 #158931
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Oct 6, 2026 - addedperformancePerformance or resource usagePerformance or resource usageinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Oct 7, 2026 The reasons can vary and we did see some causes due to PGO itself. I suggest that you look directly at the code to analyze what happens or open a DPO thread. We don't have anything actionable here. I suspect that decoding has other bottlenecks that hide what happens. The scanner may also be generic (that I don't remember).
Maybe @vstinner or @serhiy-storchaka do now the implementation details
- addedpendingThe issue will be closed if no feedback is providedThe issue will be closed if no feedback is provided
on Oct 7, 2026 Is it expected that decoding 2-byte UTF-8 sequences to UCS1 is no faster than decoding them to UCS2?
Objects/stringlib/codecs.hgenerates specializedutf8_decode()functions depending on the buffer kind (UCS1, UCS2, UCS4). It can make some assumptions depending onSTRINGLIB_SIZEOF_CHAR.Performance is a complex topic. I'm not sure how to answer. I don't have specific expectation on UCS1 vs UCS2 performance.
If you see an opportunity to optimize further stringlib
utf8_decode(), please propose a PR. This issue is just a question, I don't see what can be done. So I just suggest closing the issue.How about this line? It doesn’t calculate the minimal required size for the output buffer (like we do for escaping strings in the json library):
cpython/Objects/unicodeobject.c
Line 5390 in 20cb7fc
PyObject *u = PyUnicode_New(maxsize, maxchr); So, it allocates twice as much memory as necessary.
Feature or enhancement
Is it expected that decoding 2-byte UTF-8 sequences to UCS1 is no faster than decoding them to UCS2? I thought UCS1 would have an advantage because the resulting string uses half as much memory and requires half as many bytes to be written.
For comparison, the corresponding encoding operations do show a difference:
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response