fallintoplace opened a new pull request, #2143: URL: https://github.com/apache/iceberg-go/pull/2143
**What** - Sort Puffin blob indexes when reading all blobs. **Why** - Choosing read order only needs offsets. Copying full metadata into the ordering slice adds overhead. **Implementation** - Sort `[]int` with `slices.SortFunc` and look up metadata by index. - Keep offset-order reads, footer-order results and defensive metadata copies. - Add a benchmark using real in-memory Puffin files. **Benchmark** Apple M1 Pro, darwin/arm64, Go 1.26.3. Median of 5 runs, 300ms each, `-cpu=1`. Full `ReadAllBlobs`, 64-byte payloads, writer-order offsets. | Blobs | Before µs/op | After µs/op | B/op before → after | Allocs/op before → after | | --- | ---: | ---: | ---: | ---: | | 16 | 5.0 | 5.0 | 10408 → 8576 | 85 → 82 | | 256 | 86.9 | 77.5 | 162856 → 137472 | 1285 → 1282 | | 1024 | 325.1 | 312.0 | 640424 → 550144 | 5125 → 5122 | At 1,024 blobs, the temporary ordering slice saves 90,112 bytes per call. Storage I/O can dominate overall time. ```sh go test ./puffin -run '^$' -bench '^BenchmarkReadAllBlobs$' -benchmem -benchtime=300ms -count=5 -cpu=1 ``` -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
