fallintoplace opened a new pull request, #2147: URL: https://github.com/apache/iceberg-go/pull/2147
**What** - Add opt-in raw manifest content caching across scans of a table handle. **Why** - Repeated local planning reopens immutable manifests even when the manifest list is cached. Remote FileIO adds latency for each read. **Implementation** - Byte-bounded LRU with idle expiration and a per-file size limit. - Share concurrent misses. Keep waiter cancellation separate and fall back to normal FileIO on population errors. - Read settings from saved catalog config. Disabled by default; limits are 100 MiB total, 8 MiB per manifest, 60s idle expiration. - Integrate local scan planning and transaction scans. Keep Avro decoding, projection and filtering per scan. **Benchmark** Actual repeated `Scan.PlanFiles`, real Avro manifests, warm cache, concurrency 1. Apple M1 Pro, darwin/arm64, Go 1.26.3. Median of 3 runs, 300ms each, `-cpu=1`. | Manifests | Simulated latency per open | Disabled ms/op | Warm ms/op | Manifest opens before → after | | --- | --- | ---: | ---: | ---: | | 8 | 0 | 2.21 | 2.26 | 8 → 0 | | 32 | 0 | 8.69 | 9.51 | 32 → 0 | | 8 | 1ms | 12.79 | 2.45 | 8 → 0 | | 32 | 1ms | 48.55 | 8.85 | 32 → 0 | Warm planning is about **5.2–5.5x faster** with simulated backend latency. Zero-latency controls are 2–9% slower. Each cached manifest adds one small allocation; Avro decoding remains unchanged. ```sh go test ./table -run '^$' -bench '^BenchmarkPlanFilesManifestContentCache$' -benchmem -benchtime=300ms -count=3 -cpu=1 ``` -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
