Skip to content

[Python][Parquet] ParquetFile projection includes extra column when dotted field name matches nested path #52997

Description

@Mohit-25-tech

Describe the bug, including details regarding any error messages, version, and platform.

Describe the bug

pyarrow.parquet.ParquetFile.read() returns an unexpected extra column when a literal top-level column name contains a dot and collides with a nested field path.

For example, a Parquet file contains two distinct fields:

  • a.b: a literal top-level column.
  • a: a struct containing nested field b.

When requesting columns=["a.b"], ParquetFile.read() returns both a.b and a, whereas pq.read_table() returns only a.b.

I independently reproduced this in Kaggle using PyArrow 26.0.0. The behavior was also observed locally with PyArrow 16.1.0 and 24.0.0.

Minimal reproduction

import pyarrow as pa
import pyarrow.parquet as pq

print("PyArrow version:", pa.__version__)

table = pa.table({
    "a.b": [10, 20],
    "a": pa.array([{"b": 1}, {"b": 2}]),
    "other": [3, 4],
})

sink = pa.BufferOutputStream()
pq.write_table(table, sink)
data = sink.getvalue()

file_result = pq.ParquetFile(
    pa.BufferReader(data)
).read(columns=["a.b"])

table_result = pq.read_table(
    pa.BufferReader(data),
    columns=["a.b"],
)

print("ParquetFile.read:", file_result.column_names, file_result.to_pydict())
print("pq.read_table:", table_result.column_names, table_result.to_pydict())

Actual behavior

PyArrow version: 26.0.0
ParquetFile.read: ['a.b', 'a'] {'a.b': [10, 20], 'a': [{'b': 1}, {'b': 2}]}
pq.read_table:    ['a.b'] {'a.b': [10, 20]}

Expected behavior

Column projection should have a clear and consistent interpretation across the Parquet reading APIs.

Preferably, specifying columns=["a.b"] should select the exact top-level column named a.b when it exists, without unexpectedly including the unrelated top-level struct a.

Alternatively, if this name is considered ambiguous, a clear error or an explicit path-selection mechanism may be more appropriate.

I would appreciate maintainer guidance on the intended precedence.

Possible root cause

In python/pyarrow/parquet/core.py, _build_nested_paths() constructs lookup keys by joining path components with ".".

This makes the following distinct physical paths indistinguishable as strings:

["a.b"]    -> "a.b"
["a", "b"] -> "a.b"

Consequently, _get_column_indices() can select both physical columns for a single requested name.

Impact

This may cause unexpected output schemas, unnecessary Parquet column reads, additional memory usage, or downstream failures in applications that use dotted column names alongside nested structures.

Environment

  • PyArrow: 26.0.0
  • Platform: Kaggle notebook (Linux)
  • Reproduction: In-memory Parquet buffer; no external files or datasets required.

I can prepare a focused regression test and fix once the expected ambiguity-resolution behavior is clarified.

The investigation and initial report drafting were AI-assisted; I independently executed and reviewed the reproduction.

Component(s)

Parquet, Python

Activity

  1. Mohit-25-tech commented on Oct 10, 2026

    @Mohit-25-tech
    Author

    Hi maintainers,

    I've prepared a focused PR addressing this issue.

    Root cause: ParquetFile._build_nested_paths() creates identical lookup keys for a literal top-level field a.b and a nested path a → b, causing column projection to select both fields unexpectedly.

    Proposed fix:

    • Give exact top-level column names precedence over ambiguous dotted nested paths.
    • Preserve normal nested-field selection when no exact top-level match exists.
    • Add regression tests covering read(), read_row_group(), iter_batches(), deeper nested paths, duplicate selections, and pandas metadata.

    Validation:

    • Independently reproduced the original issue with PyArrow 26.0.0 in Kaggle.
    • All 8 independent Kaggle checks passed using an adapted patch.
    • 4 focused pytest tests passed using the modified Python methods.
    • Flake8, syntax checks, and git diff --check passed.
    • Full repository tests remain unverified locally.

    PR: #53010

    The precedence change is intentional, and I'd appreciate feedback on whether this is the preferred API behavior.

    Thanks for your time and review!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions