+1 from me. In addition to the incremental repair benefits, this will make incremental cluster shrinks and expansions significantly less expensive.
On Wed, Sep 16, 2026, at 8:20 AM, Josh McKenzie wrote: >> I think AI usage should be debated as a separate topic and the CEP should be >> evaluated for the design and the approach. > +1. I'd prefer we keep this discussion here to the merits of the CEP as a > design proposal as written. > > I am broadly in favor of what you're doing here Chris and the logic of it > makes complete sense; the impact on cluster stability from reducing all the > CPU and I/O churn of anti-compaction would drastically increase the use-cases > in which it can be safely used, nevermind the efficiency gains. > > Haven't looked into the details yet but nothing smells from prior > conversations on slack about the work. ;) > > On Wed, Sep 16, 2026, at 8:25 AM, Chris Lohfink wrote: >> Thanks, Branimir. The shared-index approach is an attractive option >> especially for transient local operations such as anticompaction. >> >> The tradeoff appears to be where the complexity lives. Reusing the original >> primary index makes slice creation cheaper, but its positions remain in the >> parent Data.db coordinate space. Readers and tools must therefore understand >> that the Data.db component is a slice and translate those positions before >> accessing the local file. >> >> The current proposal instead pays that cost once during splitting. It >> rebases Index.db positions, slices CompressionInfo.db, and rebuilds or >> apportions the child’s derived metadata. This preserves the usual idea that >> each SSTable is a self-contained artifact whose components share one >> coordinate space and lifecycle. Index reconstruction is so cheap its >> comparatively free relative to reading or copying Data.db, while also >> producing child-specific summaries, Bloom filters, and statistics (100's of >> ms range on *huge* sstables) >> >> For durable outputs, I prefer keeping that complexity at creation time. >> Sharing one physical index file among several logical SSTables would >> introduce additional lifecycle and bookkeeping cases around deletion, >> snapshots, backup and restore, import, and third-party tooling. Independent >> files that happen to share filesystem extents are less concerning because >> they retain normal component ownership semantics. >> >> That said, the DataStax approach is worth benchmarking, and it may be a >> better fit for BTI or another format where slicing is designed in as a >> first-class property. >> >> > Separately, being able to easily slice and dice files without looking >> > inside them is a key consideration in the file format we are working on >> > for CEP-57. >> >> Agreed,nCEP-57 seems like the right place to make sliceability a native >> format property. My goal here is narrower: provide the capability for BIG >> SSTables and current formats in the meantime without permanently introducing >> slice-coordinate awareness throughout BIG’s read path. BIG will remain in >> use for some time, and this mechanism can deliver most of the benefit with >> relatively contained changes. >> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]> wrote: >>> Hello Chris, >>> >>> A while back we implemented a similar approach for DSE's version of >>> zero-copy streaming. The main difference between our approach and yours is >>> that we decided not to split the primary index files and instead use them >>> as they are, with filtering based on the start and end key of the section. >>> In the context of local operations like anticompaction, the latter may be a >>> better approach as one can share the index files between all resulting >>> sections (and, of course, no index reconstruction is necessary). >>> >>> This code is not currently part of the Apache codebase, but DataStax's open >>> source fork includes support for reading these files, whose implementation >>> (commit >>> https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422) >>> may have some ideas you can use in designing your solution. >>> >>> Separately, being able to easily slice and dice files without looking >>> inside them is a key consideration in the file format we are working on for >>> CEP-57. >>> >>> Regards, >>> Branimir >>> >>> On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella >>> <[email protected]> wrote: >>>> Really nice idea Chris! >>>> >>>> One thing I think it’s worth flagging for the proposal: >>>> >>>> this CEP introduces a new SSTable major version. That's a relevant change >>>> for anything outside nodetool/the core read path that parses SSTables >>>> directly. e.g. analytics library uses the concept of per major version >>>> bridge to deserialize. >>>> >>>> With this approach, current bridges don't have a notion of "data doesn't >>>> start at logical offset zero” - they assume a chunk's first partition >>>> begins the SSTable's data - which is fine (this is a new version!). It'd >>>> help if the CEP explicitly calls out that third-party/off-node SSTable >>>> readers are a compatibility surface here, not just in-process Cassandra >>>> binaries. >>>> >>>> Bernardo >>>> >>>> *From: *Chris Lohfink <[email protected]> >>>> *Date: *Monday, 14 September 2026 at 21:23 >>>> *To: *[email protected] <[email protected]> >>>> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting >>>> >>>> Hi everyone, >>>> >>>> I'd like to open CEP-66, Zero-copy SSTable splitting, for discussion: >>>> >>>> _https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting__ >>>> _ >>>> >>>> Anticompaction and partial-range streaming currently rewrite rows whose >>>> encoded representation already exists on disk. This consumes CPU, creates >>>> substantial heap churn and write amplification, and increases temporary >>>> disk pressure. >>>> >>>> CEP-66 proposes splitting eligible compressed SSTables by retaining >>>> contiguous runs of their existing compression chunks. Cassandra would >>>> rebuild the child SSTables' indexes and other derived components without >>>> deserializing, serializing, or recompressing their rows. >>>> >>>> "Zero-copy" here primarily means reusing the encoded bytes instead of >>>> rewriting rows. On filesystems that support range reflinks, Cassandra can >>>> also share the underlying extents meaning no new data written. Other >>>> filesystems, including ext4, would copy the already-compressed bytes and >>>> still avoid the row rewrite. >>>> >>>> The proposal is staged. It starts with an opt-in `sstablesplit >>>> --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk. Later >>>> phases add BTI support, secondary indexes, anticompaction, and >>>> partial-range streaming. Existing implementations remain the default and >>>> provide the fallback for unsupported inputs. >>>> >>>> I'd particularly appreciate feedback on: >>>> >>>> - The retained-prefix representation and proposed Cassandra 7.0 SSTable >>>> format change >>>> - Rebuilding or conservatively deriving child metadata without decoding >>>> rows >>>> - The integrity and performance tradeoff around `Digest.crc32` generation >>>> - The staged rollout, compatibility rules, and fallback behavior >>>> - Any correctness, operational, or filesystem concerns the proposal has >>>> missed >>>> >>>> Thanks, and I look forward to the discussion. >>>> >>>> Regards, >>>> Chris Lohfink
