This is such a clever way to solve a problem I wish would go away. The only
negative aspect of this proposal is that it isn't available today.

+1 from me.

On Tue, Sep 22, 2026 at 7:18 AM Andy Tolbert <[email protected]> wrote:

> Realized I didn't share a +1 in my previous message, so adding my +1!
>
> Thanks,
> Andy
>
> On Tue, Sep 22, 2026, at 2:01 PM, Francisco Guerrero wrote:
> > +1. Anticompaction is an area that causes a lot of pain and
> > seeing this proposal makes me hopeful that things will get
> > better soon.
> >
> > Also, I think we should consider bringing this work to 6.0 as
> > well assuming this lands before we release 6.0
> >
> > Best,
> > - Francisco
> >
> > On 2026/09/22 13:43:37 Abe Ratnofsky wrote:
> >> I’m +1 on the CEP.
> >>
> >> Saving bandwidth is meaningful, particularly for those running on
> networked disks where bandwidth has a lower ceiling and a direct marginal
> cost. The project currently recommends NVMe but the cost and convenience
> benefits of newer generations of networked disks are meaningful.
> >>
> >> Regarding portability: this feels no different from supporting multiple
> JDK versions, or recommending ACCP, etc. You’ll get better performance on
> certain systems where certain features are available. This change also
> introduces performance improvements for filesystems that do not support
> FICLONERANGE like ext4.
> >>
> >> I do wish it were possible for us to version SSTables in a way that let
> users experiment with this more easily. Requiring it land in a new major
> keeps it further away from many users who would benefit.
> >>
> >> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote:
> >> > My biggest concern with this proposal is whether linking inodes is
> >> > justified. This narrows the applicability (only supported
> filesystems)
> >> > and increases the complexity, but in return saves bandwidth and
> allows
> >> > it to function when disk space is exhausted.
> >> >
> >> > I am of the opinion it is probably not justified, as bandwidth is not
> I
> >> > think typically a major constraint - we don't generally saturate the
> >> > disk because we are CPU-inefficient. Simply copying the files would
> be
> >> > a better starting point as much simpler and more portable. We have
> >> > bigger problems when we are out of space.
> >> >
> >> > Another thing to balance is whether this complexity is justified for
> a
> >> > stop-gap measure, if we expect this to be made defunct by both
> >> > Branimir's new file format (which permits cheaper slicing) and
> mutation
> >> > tracking (which should eliminate the need for anti-compaction).
> >> >
> >> > Separately, I wonder (if we desperately want it in the meantime)
> >> > whether DataStax are willing to contribute their version of this, if
> it
> >> > already exists and is already validated on real workloads.
> >> >
> >> > Finally, have we explored simply removing anti-compaction instead?
> This
> >> > would require I think a couple of components: 1) per-range repair
> >> > metadata; 2) either per-range sstable invalidations (so that
> compaction
> >> > may proceed on parts of the repaired/unrepaired file independently,
> >> > permitting it to be replaced in both sets), or rewriting the
> >> > uncompacted part(s) of the sstable. I may be missing some other
> >> > complexity, but this seems quite tractable.
> >> >
> >> >
> >> >
> >> > On 2026/09/16 12:25:28 Chris Lohfink wrote:
> >> >> Thanks, Branimir. The shared-index approach is an attractive option
> >> >> especially for transient local operations such as anticompaction.
> >> >>
> >> >> The tradeoff appears to be where the complexity lives. Reusing the
> original
> >> >> primary index makes slice creation cheaper, but its positions remain
> in the
> >> >> parent Data.db coordinate space. Readers and tools must therefore
> >> >> understand that the Data.db component is a slice and translate those
> >> >> positions before accessing the local file.
> >> >>
> >> >> The current proposal instead pays that cost once during splitting. It
> >> >> rebases Index.db positions, slices CompressionInfo.db, and rebuilds
> or
> >> >> apportions the child’s derived metadata. This preserves the usual
> idea that
> >> >> each SSTable is a self-contained artifact whose components share one
> >> >> coordinate space and lifecycle. Index reconstruction is so cheap its
> >> >> comparatively free relative to reading or copying Data.db, while also
> >> >> producing child-specific summaries, Bloom filters, and statistics
> (100's of
> >> >> ms range on *huge* sstables)
> >> >>
> >> >> For durable outputs, I prefer keeping that complexity at creation
> time.
> >> >> Sharing one physical index file among several logical SSTables would
> >> >> introduce additional lifecycle and bookkeeping cases around deletion,
> >> >> snapshots, backup and restore, import, and third-party tooling.
> Independent
> >> >> files that happen to share filesystem extents are less concerning
> because
> >> >> they retain normal component ownership semantics.
> >> >>
> >> >> That said, the DataStax approach is worth benchmarking, and it may
> be a
> >> >> better fit for BTI or another format where slicing is designed in as
> a
> >> >> first-class property.
> >> >>
> >> >> > Separately, being able to easily slice and dice files without
> looking
> >> >> inside them is a key consideration in the file format we are working
> on for
> >> >> CEP-57.
> >> >>
> >> >> Agreed,nCEP-57 seems like the right place to make sliceability a
> native
> >> >> format property. My goal here is narrower: provide the capability
> for BIG
> >> >> SSTables and current formats in the meantime without permanently
> >> >> introducing slice-coordinate awareness throughout BIG’s read path.
> BIG will
> >> >> remain in use for some time, and this mechanism can deliver most of
> the
> >> >> benefit with relatively contained changes.
> >> >>
> >> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]>
> wrote:
> >> >>
> >> >> > Hello Chris,
> >> >> >
> >> >> > A while back we implemented a similar approach for DSE's version of
> >> >> > zero-copy streaming. The main difference between our approach and
> yours is
> >> >> > that we decided not to split the primary index files and instead
> use them
> >> >> > as they are, with filtering based on the start and end key of the
> section.
> >> >> > In the context of local operations like anticompaction, the latter
> may be a
> >> >> > better approach as one can share the index files between all
> resulting
> >> >> > sections (and, of course, no index reconstruction is necessary).
> >> >> >
> >> >> > This code is not currently part of the Apache codebase, but
> DataStax's
> >> >> > open source fork includes support for reading these files, whose
> >> >> > implementation (commit
> >> >> >
> https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422
> )
> >> >> > may have some ideas you can use in designing your solution.
> >> >> >
> >> >> > Separately, being able to easily slice and dice files without
> looking
> >> >> > inside them is a key consideration in the file format we are
> working on for
> >> >> > CEP-57.
> >> >> >
> >> >> > Regards,
> >> >> > Branimir
> >> >> >
> >> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella <
> >> >> > [email protected]> wrote:
> >> >> >
> >> >> >> Really nice idea Chris!
> >> >> >>
> >> >> >> One thing I think it’s worth flagging for the proposal:
> >> >> >>
> >> >> >> this CEP introduces a new SSTable major version. That's a
> relevant change
> >> >> >> for anything outside nodetool/the core read path that parses
> SSTables
> >> >> >> directly. e.g. analytics library uses the concept of per major
> version
> >> >> >> bridge to deserialize.
> >> >> >>
> >> >> >> With this approach, current bridges don't have a notion of "data
> doesn't
> >> >> >> start at logical offset zero” - they assume a chunk's first
> partition
> >> >> >> begins the SSTable's data - which is fine (this is a new
> version!). It'd
> >> >> >> help if the CEP explicitly calls out that third-party/off-node
> SSTable
> >> >> >> readers are a compatibility surface here, not just in-process
> Cassandra
> >> >> >> binaries.
> >> >> >>
> >> >> >> Bernardo
> >> >> >>
> >> >> >> *From: *Chris Lohfink <[email protected]>
> >> >> >> *Date: *Monday, 14 September 2026 at 21:23
> >> >> >> *To: *[email protected] <[email protected]>
> >> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting
> >> >> >>
> >> >> >> Hi everyone,
> >> >> >>
> >> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for
> discussion:
> >> >> >>
> >> >> >>
> >> >> >> *
> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting*
> >> >> >> <
> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting
> >
> >> >> >>
> >> >> >>
> >> >> >> Anticompaction and partial-range streaming currently rewrite rows
> whose
> >> >> >> encoded representation already exists on disk. This consumes CPU,
> creates
> >> >> >> substantial heap churn and write amplification, and increases
> temporary
> >> >> >> disk pressure.
> >> >> >>
> >> >> >> CEP-66 proposes splitting eligible compressed SSTables by
> retaining
> >> >> >> contiguous runs of their existing compression chunks. Cassandra
> would
> >> >> >> rebuild the child SSTables' indexes and other derived components
> without
> >> >> >> deserializing, serializing, or recompressing their rows.
> >> >> >>
> >> >> >> "Zero-copy" here primarily means reusing the encoded bytes
> instead of
> >> >> >> rewriting rows. On filesystems that support range reflinks,
> Cassandra can
> >> >> >> also share the underlying extents meaning no new data written.
> Other
> >> >> >> filesystems, including ext4, would copy the already-compressed
> bytes and
> >> >> >> still avoid the row rewrite.
> >> >> >>
> >> >> >> The proposal is staged. It starts with an opt-in `sstablesplit
> >> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk.
> Later
> >> >> >> phases add BTI support, secondary indexes, anticompaction, and
> >> >> >> partial-range streaming. Existing implementations remain the
> default and
> >> >> >> provide the fallback for unsupported inputs.
> >> >> >>
> >> >> >> I'd particularly appreciate feedback on:
> >> >> >>
> >> >> >> - The retained-prefix representation and proposed Cassandra 7.0
> SSTable
> >> >> >> format change
> >> >> >> - Rebuilding or conservatively deriving child metadata without
> decoding
> >> >> >> rows
> >> >> >> - The integrity and performance tradeoff around `Digest.crc32`
> generation
> >> >> >> - The staged rollout, compatibility rules, and fallback behavior
> >> >> >> - Any correctness, operational, or filesystem concerns the
> proposal has
> >> >> >> missed
> >> >> >>
> >> >> >> Thanks, and I look forward to the discussion.
> >> >> >>
> >> >> >> Regards,
> >> >> >> Chris Lohfink
> >> >> >>
> >> >> >
> >> >>
> >>
>

Reply via email to