Realized I didn't share a +1 in my previous message, so adding my +1! Thanks, Andy
On Tue, Sep 22, 2026, at 2:01 PM, Francisco Guerrero wrote: > +1. Anticompaction is an area that causes a lot of pain and > seeing this proposal makes me hopeful that things will get > better soon. > > Also, I think we should consider bringing this work to 6.0 as > well assuming this lands before we release 6.0 > > Best, > - Francisco > > On 2026/09/22 13:43:37 Abe Ratnofsky wrote: >> I’m +1 on the CEP. >> >> Saving bandwidth is meaningful, particularly for those running on networked >> disks where bandwidth has a lower ceiling and a direct marginal cost. The >> project currently recommends NVMe but the cost and convenience benefits of >> newer generations of networked disks are meaningful. >> >> Regarding portability: this feels no different from supporting multiple JDK >> versions, or recommending ACCP, etc. You’ll get better performance on >> certain systems where certain features are available. This change also >> introduces performance improvements for filesystems that do not support >> FICLONERANGE like ext4. >> >> I do wish it were possible for us to version SSTables in a way that let >> users experiment with this more easily. Requiring it land in a new major >> keeps it further away from many users who would benefit. >> >> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote: >> > My biggest concern with this proposal is whether linking inodes is >> > justified. This narrows the applicability (only supported filesystems) >> > and increases the complexity, but in return saves bandwidth and allows >> > it to function when disk space is exhausted. >> > >> > I am of the opinion it is probably not justified, as bandwidth is not I >> > think typically a major constraint - we don't generally saturate the >> > disk because we are CPU-inefficient. Simply copying the files would be >> > a better starting point as much simpler and more portable. We have >> > bigger problems when we are out of space. >> > >> > Another thing to balance is whether this complexity is justified for a >> > stop-gap measure, if we expect this to be made defunct by both >> > Branimir's new file format (which permits cheaper slicing) and mutation >> > tracking (which should eliminate the need for anti-compaction). >> > >> > Separately, I wonder (if we desperately want it in the meantime) >> > whether DataStax are willing to contribute their version of this, if it >> > already exists and is already validated on real workloads. >> > >> > Finally, have we explored simply removing anti-compaction instead? This >> > would require I think a couple of components: 1) per-range repair >> > metadata; 2) either per-range sstable invalidations (so that compaction >> > may proceed on parts of the repaired/unrepaired file independently, >> > permitting it to be replaced in both sets), or rewriting the >> > uncompacted part(s) of the sstable. I may be missing some other >> > complexity, but this seems quite tractable. >> > >> > >> > >> > On 2026/09/16 12:25:28 Chris Lohfink wrote: >> >> Thanks, Branimir. The shared-index approach is an attractive option >> >> especially for transient local operations such as anticompaction. >> >> >> >> The tradeoff appears to be where the complexity lives. Reusing the >> >> original >> >> primary index makes slice creation cheaper, but its positions remain in >> >> the >> >> parent Data.db coordinate space. Readers and tools must therefore >> >> understand that the Data.db component is a slice and translate those >> >> positions before accessing the local file. >> >> >> >> The current proposal instead pays that cost once during splitting. It >> >> rebases Index.db positions, slices CompressionInfo.db, and rebuilds or >> >> apportions the child’s derived metadata. This preserves the usual idea >> >> that >> >> each SSTable is a self-contained artifact whose components share one >> >> coordinate space and lifecycle. Index reconstruction is so cheap its >> >> comparatively free relative to reading or copying Data.db, while also >> >> producing child-specific summaries, Bloom filters, and statistics (100's >> >> of >> >> ms range on *huge* sstables) >> >> >> >> For durable outputs, I prefer keeping that complexity at creation time. >> >> Sharing one physical index file among several logical SSTables would >> >> introduce additional lifecycle and bookkeeping cases around deletion, >> >> snapshots, backup and restore, import, and third-party tooling. >> >> Independent >> >> files that happen to share filesystem extents are less concerning because >> >> they retain normal component ownership semantics. >> >> >> >> That said, the DataStax approach is worth benchmarking, and it may be a >> >> better fit for BTI or another format where slicing is designed in as a >> >> first-class property. >> >> >> >> > Separately, being able to easily slice and dice files without looking >> >> inside them is a key consideration in the file format we are working on >> >> for >> >> CEP-57. >> >> >> >> Agreed,nCEP-57 seems like the right place to make sliceability a native >> >> format property. My goal here is narrower: provide the capability for BIG >> >> SSTables and current formats in the meantime without permanently >> >> introducing slice-coordinate awareness throughout BIG’s read path. BIG >> >> will >> >> remain in use for some time, and this mechanism can deliver most of the >> >> benefit with relatively contained changes. >> >> >> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]> >> >> wrote: >> >> >> >> > Hello Chris, >> >> > >> >> > A while back we implemented a similar approach for DSE's version of >> >> > zero-copy streaming. The main difference between our approach and yours >> >> > is >> >> > that we decided not to split the primary index files and instead use >> >> > them >> >> > as they are, with filtering based on the start and end key of the >> >> > section. >> >> > In the context of local operations like anticompaction, the latter may >> >> > be a >> >> > better approach as one can share the index files between all resulting >> >> > sections (and, of course, no index reconstruction is necessary). >> >> > >> >> > This code is not currently part of the Apache codebase, but DataStax's >> >> > open source fork includes support for reading these files, whose >> >> > implementation (commit >> >> > https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422) >> >> > may have some ideas you can use in designing your solution. >> >> > >> >> > Separately, being able to easily slice and dice files without looking >> >> > inside them is a key consideration in the file format we are working on >> >> > for >> >> > CEP-57. >> >> > >> >> > Regards, >> >> > Branimir >> >> > >> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella < >> >> > [email protected]> wrote: >> >> > >> >> >> Really nice idea Chris! >> >> >> >> >> >> One thing I think it’s worth flagging for the proposal: >> >> >> >> >> >> this CEP introduces a new SSTable major version. That's a relevant >> >> >> change >> >> >> for anything outside nodetool/the core read path that parses SSTables >> >> >> directly. e.g. analytics library uses the concept of per major version >> >> >> bridge to deserialize. >> >> >> >> >> >> With this approach, current bridges don't have a notion of "data >> >> >> doesn't >> >> >> start at logical offset zero” - they assume a chunk's first partition >> >> >> begins the SSTable's data - which is fine (this is a new version!). >> >> >> It'd >> >> >> help if the CEP explicitly calls out that third-party/off-node SSTable >> >> >> readers are a compatibility surface here, not just in-process Cassandra >> >> >> binaries. >> >> >> >> >> >> Bernardo >> >> >> >> >> >> *From: *Chris Lohfink <[email protected]> >> >> >> *Date: *Monday, 14 September 2026 at 21:23 >> >> >> *To: *[email protected] <[email protected]> >> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting >> >> >> >> >> >> Hi everyone, >> >> >> >> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for discussion: >> >> >> >> >> >> >> >> >> *https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting* >> >> >> <https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting> >> >> >> >> >> >> >> >> >> Anticompaction and partial-range streaming currently rewrite rows whose >> >> >> encoded representation already exists on disk. This consumes CPU, >> >> >> creates >> >> >> substantial heap churn and write amplification, and increases temporary >> >> >> disk pressure. >> >> >> >> >> >> CEP-66 proposes splitting eligible compressed SSTables by retaining >> >> >> contiguous runs of their existing compression chunks. Cassandra would >> >> >> rebuild the child SSTables' indexes and other derived components >> >> >> without >> >> >> deserializing, serializing, or recompressing their rows. >> >> >> >> >> >> "Zero-copy" here primarily means reusing the encoded bytes instead of >> >> >> rewriting rows. On filesystems that support range reflinks, Cassandra >> >> >> can >> >> >> also share the underlying extents meaning no new data written. Other >> >> >> filesystems, including ext4, would copy the already-compressed bytes >> >> >> and >> >> >> still avoid the row rewrite. >> >> >> >> >> >> The proposal is staged. It starts with an opt-in `sstablesplit >> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk. Later >> >> >> phases add BTI support, secondary indexes, anticompaction, and >> >> >> partial-range streaming. Existing implementations remain the default >> >> >> and >> >> >> provide the fallback for unsupported inputs. >> >> >> >> >> >> I'd particularly appreciate feedback on: >> >> >> >> >> >> - The retained-prefix representation and proposed Cassandra 7.0 SSTable >> >> >> format change >> >> >> - Rebuilding or conservatively deriving child metadata without decoding >> >> >> rows >> >> >> - The integrity and performance tradeoff around `Digest.crc32` >> >> >> generation >> >> >> - The staged rollout, compatibility rules, and fallback behavior >> >> >> - Any correctness, operational, or filesystem concerns the proposal has >> >> >> missed >> >> >> >> >> >> Thanks, and I look forward to the discussion. >> >> >> >> >> >> Regards, >> >> >> Chris Lohfink >> >> >> >> >> > >> >> >>
