Describe the bug, including details regarding any error messages, version, and platform.
Describe the bug
pyarrow.parquet.ParquetFile.read() returns an unexpected extra column when a literal top-level column name contains a dot and collides with a nested field path.
For example, a Parquet file contains two distinct fields:
a.b: a literal top-level column.
a: a struct containing nested field b.
When requesting columns=["a.b"], ParquetFile.read() returns both a.b and a, whereas pq.read_table() returns only a.b.
I independently reproduced this in Kaggle using PyArrow 26.0.0. The behavior was also observed locally with PyArrow 16.1.0 and 24.0.0.
Minimal reproduction
import pyarrow as pa
import pyarrow.parquet as pq
print("PyArrow version:", pa.__version__)
table = pa.table({
"a.b": [10, 20],
"a": pa.array([{"b": 1}, {"b": 2}]),
"other": [3, 4],
})
sink = pa.BufferOutputStream()
pq.write_table(table, sink)
data = sink.getvalue()
file_result = pq.ParquetFile(
pa.BufferReader(data)
).read(columns=["a.b"])
table_result = pq.read_table(
pa.BufferReader(data),
columns=["a.b"],
)
print("ParquetFile.read:", file_result.column_names, file_result.to_pydict())
print("pq.read_table:", table_result.column_names, table_result.to_pydict())
Actual behavior
PyArrow version: 26.0.0
ParquetFile.read: ['a.b', 'a'] {'a.b': [10, 20], 'a': [{'b': 1}, {'b': 2}]}
pq.read_table: ['a.b'] {'a.b': [10, 20]}
Expected behavior
Column projection should have a clear and consistent interpretation across the Parquet reading APIs.
Preferably, specifying columns=["a.b"] should select the exact top-level column named a.b when it exists, without unexpectedly including the unrelated top-level struct a.
Alternatively, if this name is considered ambiguous, a clear error or an explicit path-selection mechanism may be more appropriate.
I would appreciate maintainer guidance on the intended precedence.
Possible root cause
In python/pyarrow/parquet/core.py, _build_nested_paths() constructs lookup keys by joining path components with ".".
This makes the following distinct physical paths indistinguishable as strings:
["a.b"] -> "a.b"
["a", "b"] -> "a.b"
Consequently, _get_column_indices() can select both physical columns for a single requested name.
Impact
This may cause unexpected output schemas, unnecessary Parquet column reads, additional memory usage, or downstream failures in applications that use dotted column names alongside nested structures.
Environment
- PyArrow: 26.0.0
- Platform: Kaggle notebook (Linux)
- Reproduction: In-memory Parquet buffer; no external files or datasets required.
I can prepare a focused regression test and fix once the expected ambiguity-resolution behavior is clarified.
The investigation and initial report drafting were AI-assisted; I independently executed and reviewed the reproduction.
Component(s)
Parquet, Python
Describe the bug, including details regarding any error messages, version, and platform.
Describe the bug
pyarrow.parquet.ParquetFile.read()returns an unexpected extra column when a literal top-level column name contains a dot and collides with a nested field path.For example, a Parquet file contains two distinct fields:
a.b: a literal top-level column.a: a struct containing nested fieldb.When requesting
columns=["a.b"],ParquetFile.read()returns botha.banda, whereaspq.read_table()returns onlya.b.I independently reproduced this in Kaggle using PyArrow 26.0.0. The behavior was also observed locally with PyArrow 16.1.0 and 24.0.0.
Minimal reproduction
Actual behavior
Expected behavior
Column projection should have a clear and consistent interpretation across the Parquet reading APIs.
Preferably, specifying
columns=["a.b"]should select the exact top-level column nameda.bwhen it exists, without unexpectedly including the unrelated top-level structa.Alternatively, if this name is considered ambiguous, a clear error or an explicit path-selection mechanism may be more appropriate.
I would appreciate maintainer guidance on the intended precedence.
Possible root cause
In
python/pyarrow/parquet/core.py,_build_nested_paths()constructs lookup keys by joining path components with".".This makes the following distinct physical paths indistinguishable as strings:
Consequently,
_get_column_indices()can select both physical columns for a single requested name.Impact
This may cause unexpected output schemas, unnecessary Parquet column reads, additional memory usage, or downstream failures in applications that use dotted column names alongside nested structures.
Environment
I can prepare a focused regression test and fix once the expected ambiguity-resolution behavior is clarified.
The investigation and initial report drafting were AI-assisted; I independently executed and reviewed the reproduction.
Component(s)
Parquet, Python