tam3tamtam opened a new issue, #51548:
URL: https://github.com/apache/arrow/issues/51548
### Describe the bug, including details regarding any error messages,
version, and platform.
### Describe the bug, including details regarding any error messages,
version, and platform.
`pyarrow.interchange.from_dataframe` can return incorrect values for columns
whose interchange-protocol dtype uses non-native endianness. The importer
interprets the producer's data buffer as native-endian without accounting for
the endianness field. No error is raised; the values are silently corrupted.
This was reproduced with PyArrow `26.0.0 (development build)` on macOS arm64
with Python 3.13.15. The same issue can affect non-native-endian string
offset
buffers.
### Reproduction
```python
import numpy as np
import pandas as pd
import pyarrow.interchange as pai
df = pd.DataFrame({"value": np.array([1, 2, 300], dtype=">i4")})
table = pai.from_dataframe(df)
print(table["value"].to_pylist())
```
### Actual behavior
On a little-endian platform, this prints values interpreted with the wrong
byte
order, for example:
```text
[16777216, 33554432, 738263040]
```
### Expected behavior
The imported values should preserve the producer's values:
```text
[1, 2, 300]
```
### Proposed fix
Convert non-native-endian data and string offset buffers to native byte order
when copying is allowed. If conversion requires a copy and
`allow_copy=False`,
raise a `RuntimeError`.
### Component(s)
Python
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]