gh-158567: Optimize PyFloat_Pack*() functions - #158483
Conversation
PyFloat_Pack4() and PyFloat_Pack8() avoid temporary buffer in the native byte order. Use _Py_bswapXX() functions to reverse bytes.
|
Benchmark: import pyperf, sys, _testcapi
runner = pyperf.Runner()
le = int(sys.byteorder == 'little')
for d in (1.0, float('nan')):
runner.bench_func(f'pack2({d}) native', _testcapi.float_pack, 2, d, le)
runner.bench_func(f'pack4({d}) native', _testcapi.float_pack, 4, d, le)
runner.bench_func(f'pack8({d}) native', _testcapi.float_pack, 8, d, le)
be = int(not le)
runner.bench_func(f'pack2({d}) byteswap', _testcapi.float_pack, 2, d, be)
runner.bench_func(f'pack4({d}) byteswap', _testcapi.float_pack, 4, d, be)
runner.bench_func(f'pack8({d}) byteswap', _testcapi.float_pack, 8, d, be)Result:
For example, Assembly code after: The conditional jump is replaced with more efficient |
|
cc @skirpichev |
skirpichev
left a comment
There was a problem hiding this comment.
LGTM
Though, not sure if the second case in Pack2 does make sense, see comment.
|
@skirpichev: I pushed a change to also use _Py_bswap16()+memcpy() at the end of PyFloat_Pack2(). Update benchmark results (CPU isolated, after running
Benchmark hidden because not significant (4): pack2(1.0) native, pack2(1.0) byteswap, pack4(1.0) byteswap, pack2(nan) byteswap This time it's no longer "1.11x faster", but at least, it's not slower on any benchmark: it's either faster or as fast :-)
Oh sure. It's really hard to measure the speedup, the difference is really tiny and can be lost in noise. I'm using CPU isolation on Linux to reduce the noise. Details |
|
Windows and macOS CI failed, unrelated failures. I just re-run these 2 jobs. Oh, unrelated failure: test_external_inspection failed on "Windows (free-threading) / Build and test (x64, tail-call)" with Also, the "macOS (free-threading) / build and test (macos-26)" job interrupted. I didn't see that recently, it's surprising: |
|
|
|
|
s390x Fedora Stable LTO 3.x: Oh! test_buffer fails on s390x Clang buildbot workers. Example: |
I reported the GCC crash to GCC bug tracker as: https://bugzilla.redhat.com/show_bug.cgi?id=2544649. |
|
test_buffer fails randomly on s390x with Clang. I ran Run 1:
Run 2:
Run5:
Run 6:
|
|
I generated test cases to trigger the bug on s390x with clang: import _testcapi
BIG_ENDIAN = 0
LITTLE_ENDIAN = 1
pack = _testcapi.float_pack
def format_bytes(b):
return ''.join(f'\\x{byte:02x}' for byte in b)
def test(d, expected, endian):
result = pack(2, d, endian)
if result != expected:
result = format_bytes(result)
expected = format_bytes(expected)
print(f"test failed: {d=} result='{result}' expected='{expected}' {endian=}")
test(float.fromhex('0x1.5f07c01032e34p-1'), b'\x39\x7c', BIG_ENDIAN)
test(float.fromhex('0x1.5f07c01032e34p-1'), b'\x7c\x39', LITTLE_ENDIAN)
test(float.fromhex('0x1.3ee08ca31ea80p-4'), b'\x2c\xfc', BIG_ENDIAN)
test(float.fromhex('0x1.3ee08ca31ea80p-4'), b'\xfc\x2c', LITTLE_ENDIAN)
test(float.fromhex('0x1.5ee5f0edf590cp-3'), b'\x31\x7c', BIG_ENDIAN)
test(float.fromhex('0x1.5ee5f0edf590cp-3'), b'\x7c\x31', LITTLE_ENDIAN)Output whe Python is built with If I apply #158568 fix, the clang bug goes away. |
PyFloat_Pack4() and PyFloat_Pack8() avoid temporary buffer in the native byte order.
Use _Py_bswapXX() functions to reverse bytes.
PyFloat_Pack*()functions using_Py_bswap16/32/64()functions #158567