+1. Anticompaction is an area that causes a lot of pain and
seeing this proposal makes me hopeful that things will get
better soon.

Also, I think we should consider bringing this work to 6.0 as
well assuming this lands before we release 6.0

Best,
- Francisco

On 2026/09/22 13:43:37 Abe Ratnofsky wrote:
> I’m +1 on the CEP.
> 
> Saving bandwidth is meaningful, particularly for those running on networked 
> disks where bandwidth has a lower ceiling and a direct marginal cost. The 
> project currently recommends NVMe but the cost and convenience benefits of 
> newer generations of networked disks are meaningful.
> 
> Regarding portability: this feels no different from supporting multiple JDK 
> versions, or recommending ACCP, etc. You’ll get better performance on certain 
> systems where certain features are available. This change also introduces 
> performance improvements for filesystems that do not support FICLONERANGE 
> like ext4.
> 
> I do wish it were possible for us to version SSTables in a way that let users 
> experiment with this more easily. Requiring it land in a new major keeps it 
> further away from many users who would benefit.
> 
> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote:
> > My biggest concern with this proposal is whether linking inodes is 
> > justified. This narrows the applicability (only supported filesystems) 
> > and increases the complexity, but in return saves bandwidth and allows 
> > it to function when disk space is exhausted.
> >
> > I am of the opinion it is probably not justified, as bandwidth is not I 
> > think typically a major constraint - we don't generally saturate the 
> > disk because we are CPU-inefficient. Simply copying the files would be 
> > a better starting point as much simpler and more portable. We have 
> > bigger problems when we are out of space.
> >
> > Another thing to balance is whether this complexity is justified for a 
> > stop-gap measure, if we expect this to be made defunct by both 
> > Branimir's new file format (which permits cheaper slicing) and mutation 
> > tracking (which should eliminate the need for anti-compaction).
> >
> > Separately, I wonder (if we desperately want it in the meantime) 
> > whether DataStax are willing to contribute their version of this, if it 
> > already exists and is already validated on real workloads.
> >
> > Finally, have we explored simply removing anti-compaction instead? This 
> > would require I think a couple of components: 1) per-range repair 
> > metadata; 2) either per-range sstable invalidations (so that compaction 
> > may proceed on parts of the repaired/unrepaired file independently, 
> > permitting it to be replaced in both sets), or rewriting the 
> > uncompacted part(s) of the sstable. I may be missing some other 
> > complexity, but this seems quite tractable.
> >
> >
> >
> > On 2026/09/16 12:25:28 Chris Lohfink wrote:
> >> Thanks, Branimir. The shared-index approach is an attractive option
> >> especially for transient local operations such as anticompaction.
> >> 
> >> The tradeoff appears to be where the complexity lives. Reusing the original
> >> primary index makes slice creation cheaper, but its positions remain in the
> >> parent Data.db coordinate space. Readers and tools must therefore
> >> understand that the Data.db component is a slice and translate those
> >> positions before accessing the local file.
> >> 
> >> The current proposal instead pays that cost once during splitting. It
> >> rebases Index.db positions, slices CompressionInfo.db, and rebuilds or
> >> apportions the child’s derived metadata. This preserves the usual idea that
> >> each SSTable is a self-contained artifact whose components share one
> >> coordinate space and lifecycle. Index reconstruction is so cheap its
> >> comparatively free relative to reading or copying Data.db, while also
> >> producing child-specific summaries, Bloom filters, and statistics (100's of
> >> ms range on *huge* sstables)
> >> 
> >> For durable outputs, I prefer keeping that complexity at creation time.
> >> Sharing one physical index file among several logical SSTables would
> >> introduce additional lifecycle and bookkeeping cases around deletion,
> >> snapshots, backup and restore, import, and third-party tooling. Independent
> >> files that happen to share filesystem extents are less concerning because
> >> they retain normal component ownership semantics.
> >> 
> >> That said, the DataStax approach is worth benchmarking, and it may be a
> >> better fit for BTI or another format where slicing is designed in as a
> >> first-class property.
> >> 
> >> > Separately, being able to easily slice and dice files without looking
> >> inside them is a key consideration in the file format we are working on for
> >> CEP-57.
> >> 
> >> Agreed,nCEP-57 seems like the right place to make sliceability a native
> >> format property. My goal here is narrower: provide the capability for BIG
> >> SSTables and current formats in the meantime without permanently
> >> introducing slice-coordinate awareness throughout BIG’s read path. BIG will
> >> remain in use for some time, and this mechanism can deliver most of the
> >> benefit with relatively contained changes.
> >> 
> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]> wrote:
> >> 
> >> > Hello Chris,
> >> >
> >> > A while back we implemented a similar approach for DSE's version of
> >> > zero-copy streaming. The main difference between our approach and yours 
> >> > is
> >> > that we decided not to split the primary index files and instead use them
> >> > as they are, with filtering based on the start and end key of the 
> >> > section.
> >> > In the context of local operations like anticompaction, the latter may 
> >> > be a
> >> > better approach as one can share the index files between all resulting
> >> > sections (and, of course, no index reconstruction is necessary).
> >> >
> >> > This code is not currently part of the Apache codebase, but DataStax's
> >> > open source fork includes support for reading these files, whose
> >> > implementation (commit
> >> > https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422)
> >> > may have some ideas you can use in designing your solution.
> >> >
> >> > Separately, being able to easily slice and dice files without looking
> >> > inside them is a key consideration in the file format we are working on 
> >> > for
> >> > CEP-57.
> >> >
> >> > Regards,
> >> > Branimir
> >> >
> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella <
> >> > [email protected]> wrote:
> >> >
> >> >> Really nice idea Chris!
> >> >>
> >> >> One thing I think it’s worth flagging for the proposal:
> >> >>
> >> >> this CEP introduces a new SSTable major version. That's a relevant 
> >> >> change
> >> >> for anything outside nodetool/the core read path that parses SSTables
> >> >> directly. e.g. analytics library uses the concept of per major version
> >> >> bridge to deserialize.
> >> >>
> >> >> With this approach, current bridges don't have a notion of "data doesn't
> >> >> start at logical offset zero” - they assume a chunk's first partition
> >> >> begins the SSTable's data - which is fine (this is a new version!). It'd
> >> >> help if the CEP explicitly calls out that third-party/off-node SSTable
> >> >> readers are a compatibility surface here, not just in-process Cassandra
> >> >> binaries.
> >> >>
> >> >> Bernardo
> >> >>
> >> >> *From: *Chris Lohfink <[email protected]>
> >> >> *Date: *Monday, 14 September 2026 at 21:23
> >> >> *To: *[email protected] <[email protected]>
> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting
> >> >>
> >> >> Hi everyone,
> >> >>
> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for discussion:
> >> >>
> >> >>
> >> >> *https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting*
> >> >> <https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting>
> >> >>
> >> >>
> >> >> Anticompaction and partial-range streaming currently rewrite rows whose
> >> >> encoded representation already exists on disk. This consumes CPU, 
> >> >> creates
> >> >> substantial heap churn and write amplification, and increases temporary
> >> >> disk pressure.
> >> >>
> >> >> CEP-66 proposes splitting eligible compressed SSTables by retaining
> >> >> contiguous runs of their existing compression chunks. Cassandra would
> >> >> rebuild the child SSTables' indexes and other derived components without
> >> >> deserializing, serializing, or recompressing their rows.
> >> >>
> >> >> "Zero-copy" here primarily means reusing the encoded bytes instead of
> >> >> rewriting rows. On filesystems that support range reflinks, Cassandra 
> >> >> can
> >> >> also share the underlying extents meaning no new data written. Other
> >> >> filesystems, including ext4, would copy the already-compressed bytes and
> >> >> still avoid the row rewrite.
> >> >>
> >> >> The proposal is staged. It starts with an opt-in `sstablesplit
> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk. Later
> >> >> phases add BTI support, secondary indexes, anticompaction, and
> >> >> partial-range streaming. Existing implementations remain the default and
> >> >> provide the fallback for unsupported inputs.
> >> >>
> >> >> I'd particularly appreciate feedback on:
> >> >>
> >> >> - The retained-prefix representation and proposed Cassandra 7.0 SSTable
> >> >> format change
> >> >> - Rebuilding or conservatively deriving child metadata without decoding
> >> >> rows
> >> >> - The integrity and performance tradeoff around `Digest.crc32` 
> >> >> generation
> >> >> - The staged rollout, compatibility rules, and fallback behavior
> >> >> - Any correctness, operational, or filesystem concerns the proposal has
> >> >> missed
> >> >>
> >> >> Thanks, and I look forward to the discussion.
> >> >>
> >> >> Regards,
> >> >> Chris Lohfink
> >> >>
> >> >
> >>
> 

Reply via email to