Realized I didn't share a +1 in my previous message, so adding my +1!

Thanks,
Andy

On Tue, Sep 22, 2026, at 2:01 PM, Francisco Guerrero wrote:
> +1. Anticompaction is an area that causes a lot of pain and
> seeing this proposal makes me hopeful that things will get
> better soon.
>
> Also, I think we should consider bringing this work to 6.0 as
> well assuming this lands before we release 6.0
>
> Best,
> - Francisco
>
> On 2026/09/22 13:43:37 Abe Ratnofsky wrote:
>> I’m +1 on the CEP.
>> 
>> Saving bandwidth is meaningful, particularly for those running on networked 
>> disks where bandwidth has a lower ceiling and a direct marginal cost. The 
>> project currently recommends NVMe but the cost and convenience benefits of 
>> newer generations of networked disks are meaningful.
>> 
>> Regarding portability: this feels no different from supporting multiple JDK 
>> versions, or recommending ACCP, etc. You’ll get better performance on 
>> certain systems where certain features are available. This change also 
>> introduces performance improvements for filesystems that do not support 
>> FICLONERANGE like ext4.
>> 
>> I do wish it were possible for us to version SSTables in a way that let 
>> users experiment with this more easily. Requiring it land in a new major 
>> keeps it further away from many users who would benefit.
>> 
>> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote:
>> > My biggest concern with this proposal is whether linking inodes is 
>> > justified. This narrows the applicability (only supported filesystems) 
>> > and increases the complexity, but in return saves bandwidth and allows 
>> > it to function when disk space is exhausted.
>> >
>> > I am of the opinion it is probably not justified, as bandwidth is not I 
>> > think typically a major constraint - we don't generally saturate the 
>> > disk because we are CPU-inefficient. Simply copying the files would be 
>> > a better starting point as much simpler and more portable. We have 
>> > bigger problems when we are out of space.
>> >
>> > Another thing to balance is whether this complexity is justified for a 
>> > stop-gap measure, if we expect this to be made defunct by both 
>> > Branimir's new file format (which permits cheaper slicing) and mutation 
>> > tracking (which should eliminate the need for anti-compaction).
>> >
>> > Separately, I wonder (if we desperately want it in the meantime) 
>> > whether DataStax are willing to contribute their version of this, if it 
>> > already exists and is already validated on real workloads.
>> >
>> > Finally, have we explored simply removing anti-compaction instead? This 
>> > would require I think a couple of components: 1) per-range repair 
>> > metadata; 2) either per-range sstable invalidations (so that compaction 
>> > may proceed on parts of the repaired/unrepaired file independently, 
>> > permitting it to be replaced in both sets), or rewriting the 
>> > uncompacted part(s) of the sstable. I may be missing some other 
>> > complexity, but this seems quite tractable.
>> >
>> >
>> >
>> > On 2026/09/16 12:25:28 Chris Lohfink wrote:
>> >> Thanks, Branimir. The shared-index approach is an attractive option
>> >> especially for transient local operations such as anticompaction.
>> >> 
>> >> The tradeoff appears to be where the complexity lives. Reusing the 
>> >> original
>> >> primary index makes slice creation cheaper, but its positions remain in 
>> >> the
>> >> parent Data.db coordinate space. Readers and tools must therefore
>> >> understand that the Data.db component is a slice and translate those
>> >> positions before accessing the local file.
>> >> 
>> >> The current proposal instead pays that cost once during splitting. It
>> >> rebases Index.db positions, slices CompressionInfo.db, and rebuilds or
>> >> apportions the child’s derived metadata. This preserves the usual idea 
>> >> that
>> >> each SSTable is a self-contained artifact whose components share one
>> >> coordinate space and lifecycle. Index reconstruction is so cheap its
>> >> comparatively free relative to reading or copying Data.db, while also
>> >> producing child-specific summaries, Bloom filters, and statistics (100's 
>> >> of
>> >> ms range on *huge* sstables)
>> >> 
>> >> For durable outputs, I prefer keeping that complexity at creation time.
>> >> Sharing one physical index file among several logical SSTables would
>> >> introduce additional lifecycle and bookkeeping cases around deletion,
>> >> snapshots, backup and restore, import, and third-party tooling. 
>> >> Independent
>> >> files that happen to share filesystem extents are less concerning because
>> >> they retain normal component ownership semantics.
>> >> 
>> >> That said, the DataStax approach is worth benchmarking, and it may be a
>> >> better fit for BTI or another format where slicing is designed in as a
>> >> first-class property.
>> >> 
>> >> > Separately, being able to easily slice and dice files without looking
>> >> inside them is a key consideration in the file format we are working on 
>> >> for
>> >> CEP-57.
>> >> 
>> >> Agreed,nCEP-57 seems like the right place to make sliceability a native
>> >> format property. My goal here is narrower: provide the capability for BIG
>> >> SSTables and current formats in the meantime without permanently
>> >> introducing slice-coordinate awareness throughout BIG’s read path. BIG 
>> >> will
>> >> remain in use for some time, and this mechanism can deliver most of the
>> >> benefit with relatively contained changes.
>> >> 
>> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]> 
>> >> wrote:
>> >> 
>> >> > Hello Chris,
>> >> >
>> >> > A while back we implemented a similar approach for DSE's version of
>> >> > zero-copy streaming. The main difference between our approach and yours 
>> >> > is
>> >> > that we decided not to split the primary index files and instead use 
>> >> > them
>> >> > as they are, with filtering based on the start and end key of the 
>> >> > section.
>> >> > In the context of local operations like anticompaction, the latter may 
>> >> > be a
>> >> > better approach as one can share the index files between all resulting
>> >> > sections (and, of course, no index reconstruction is necessary).
>> >> >
>> >> > This code is not currently part of the Apache codebase, but DataStax's
>> >> > open source fork includes support for reading these files, whose
>> >> > implementation (commit
>> >> > https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422)
>> >> > may have some ideas you can use in designing your solution.
>> >> >
>> >> > Separately, being able to easily slice and dice files without looking
>> >> > inside them is a key consideration in the file format we are working on 
>> >> > for
>> >> > CEP-57.
>> >> >
>> >> > Regards,
>> >> > Branimir
>> >> >
>> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella <
>> >> > [email protected]> wrote:
>> >> >
>> >> >> Really nice idea Chris!
>> >> >>
>> >> >> One thing I think it’s worth flagging for the proposal:
>> >> >>
>> >> >> this CEP introduces a new SSTable major version. That's a relevant 
>> >> >> change
>> >> >> for anything outside nodetool/the core read path that parses SSTables
>> >> >> directly. e.g. analytics library uses the concept of per major version
>> >> >> bridge to deserialize.
>> >> >>
>> >> >> With this approach, current bridges don't have a notion of "data 
>> >> >> doesn't
>> >> >> start at logical offset zero” - they assume a chunk's first partition
>> >> >> begins the SSTable's data - which is fine (this is a new version!). 
>> >> >> It'd
>> >> >> help if the CEP explicitly calls out that third-party/off-node SSTable
>> >> >> readers are a compatibility surface here, not just in-process Cassandra
>> >> >> binaries.
>> >> >>
>> >> >> Bernardo
>> >> >>
>> >> >> *From: *Chris Lohfink <[email protected]>
>> >> >> *Date: *Monday, 14 September 2026 at 21:23
>> >> >> *To: *[email protected] <[email protected]>
>> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting
>> >> >>
>> >> >> Hi everyone,
>> >> >>
>> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for discussion:
>> >> >>
>> >> >>
>> >> >> *https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting*
>> >> >> <https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting>
>> >> >>
>> >> >>
>> >> >> Anticompaction and partial-range streaming currently rewrite rows whose
>> >> >> encoded representation already exists on disk. This consumes CPU, 
>> >> >> creates
>> >> >> substantial heap churn and write amplification, and increases temporary
>> >> >> disk pressure.
>> >> >>
>> >> >> CEP-66 proposes splitting eligible compressed SSTables by retaining
>> >> >> contiguous runs of their existing compression chunks. Cassandra would
>> >> >> rebuild the child SSTables' indexes and other derived components 
>> >> >> without
>> >> >> deserializing, serializing, or recompressing their rows.
>> >> >>
>> >> >> "Zero-copy" here primarily means reusing the encoded bytes instead of
>> >> >> rewriting rows. On filesystems that support range reflinks, Cassandra 
>> >> >> can
>> >> >> also share the underlying extents meaning no new data written. Other
>> >> >> filesystems, including ext4, would copy the already-compressed bytes 
>> >> >> and
>> >> >> still avoid the row rewrite.
>> >> >>
>> >> >> The proposal is staged. It starts with an opt-in `sstablesplit
>> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk. Later
>> >> >> phases add BTI support, secondary indexes, anticompaction, and
>> >> >> partial-range streaming. Existing implementations remain the default 
>> >> >> and
>> >> >> provide the fallback for unsupported inputs.
>> >> >>
>> >> >> I'd particularly appreciate feedback on:
>> >> >>
>> >> >> - The retained-prefix representation and proposed Cassandra 7.0 SSTable
>> >> >> format change
>> >> >> - Rebuilding or conservatively deriving child metadata without decoding
>> >> >> rows
>> >> >> - The integrity and performance tradeoff around `Digest.crc32` 
>> >> >> generation
>> >> >> - The staged rollout, compatibility rules, and fallback behavior
>> >> >> - Any correctness, operational, or filesystem concerns the proposal has
>> >> >> missed
>> >> >>
>> >> >> Thanks, and I look forward to the discussion.
>> >> >>
>> >> >> Regards,
>> >> >> Chris Lohfink
>> >> >>
>> >> >
>> >>
>>

Reply via email to