Skip to content

[Python] DataFrame interchange silently corrupts non-native-endian columns. #51548

Description

@tam3tamtam

Describe the bug, including details regarding any error messages, version, and platform.

Describe the bug, including details regarding any error messages, version, and platform.

pyarrow.interchange.from_dataframe can return incorrect values for columns
whose interchange-protocol dtype uses non-native endianness. The importer
interprets the producer's data buffer as native-endian without accounting for
the endianness field. No error is raised; the values are silently corrupted.

This was reproduced with PyArrow 26.0.0 (development build) on macOS arm64
with Python 3.13.15. The same issue can affect non-native-endian string offset
buffers.

Reproduction

import numpy as np
import pandas as pd
import pyarrow.interchange as pai

df = pd.DataFrame({"value": np.array([1, 2, 300], dtype=">i4")})
table = pai.from_dataframe(df)
print(table["value"].to_pylist())

Actual behavior

On a little-endian platform, this prints values interpreted with the wrong byte
order, for example:

[16777216, 33554432, 738263040]

Expected behavior

The imported values should preserve the producer's values:

[1, 2, 300]

Proposed fix

Convert non-native-endian data and string offset buffers to native byte order
when copying is allowed. If conversion requires a copy and allow_copy=False,
raise a RuntimeError.

Component(s)

Python

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions