This is such a clever way to solve a problem I wish would go away. The only negative aspect of this proposal is that it isn't available today.
+1 from me. On Tue, Sep 22, 2026 at 7:18 AM Andy Tolbert <[email protected]> wrote: > Realized I didn't share a +1 in my previous message, so adding my +1! > > Thanks, > Andy > > On Tue, Sep 22, 2026, at 2:01 PM, Francisco Guerrero wrote: > > +1. Anticompaction is an area that causes a lot of pain and > > seeing this proposal makes me hopeful that things will get > > better soon. > > > > Also, I think we should consider bringing this work to 6.0 as > > well assuming this lands before we release 6.0 > > > > Best, > > - Francisco > > > > On 2026/09/22 13:43:37 Abe Ratnofsky wrote: > >> I’m +1 on the CEP. > >> > >> Saving bandwidth is meaningful, particularly for those running on > networked disks where bandwidth has a lower ceiling and a direct marginal > cost. The project currently recommends NVMe but the cost and convenience > benefits of newer generations of networked disks are meaningful. > >> > >> Regarding portability: this feels no different from supporting multiple > JDK versions, or recommending ACCP, etc. You’ll get better performance on > certain systems where certain features are available. This change also > introduces performance improvements for filesystems that do not support > FICLONERANGE like ext4. > >> > >> I do wish it were possible for us to version SSTables in a way that let > users experiment with this more easily. Requiring it land in a new major > keeps it further away from many users who would benefit. > >> > >> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote: > >> > My biggest concern with this proposal is whether linking inodes is > >> > justified. This narrows the applicability (only supported > filesystems) > >> > and increases the complexity, but in return saves bandwidth and > allows > >> > it to function when disk space is exhausted. > >> > > >> > I am of the opinion it is probably not justified, as bandwidth is not > I > >> > think typically a major constraint - we don't generally saturate the > >> > disk because we are CPU-inefficient. Simply copying the files would > be > >> > a better starting point as much simpler and more portable. We have > >> > bigger problems when we are out of space. > >> > > >> > Another thing to balance is whether this complexity is justified for > a > >> > stop-gap measure, if we expect this to be made defunct by both > >> > Branimir's new file format (which permits cheaper slicing) and > mutation > >> > tracking (which should eliminate the need for anti-compaction). > >> > > >> > Separately, I wonder (if we desperately want it in the meantime) > >> > whether DataStax are willing to contribute their version of this, if > it > >> > already exists and is already validated on real workloads. > >> > > >> > Finally, have we explored simply removing anti-compaction instead? > This > >> > would require I think a couple of components: 1) per-range repair > >> > metadata; 2) either per-range sstable invalidations (so that > compaction > >> > may proceed on parts of the repaired/unrepaired file independently, > >> > permitting it to be replaced in both sets), or rewriting the > >> > uncompacted part(s) of the sstable. I may be missing some other > >> > complexity, but this seems quite tractable. > >> > > >> > > >> > > >> > On 2026/09/16 12:25:28 Chris Lohfink wrote: > >> >> Thanks, Branimir. The shared-index approach is an attractive option > >> >> especially for transient local operations such as anticompaction. > >> >> > >> >> The tradeoff appears to be where the complexity lives. Reusing the > original > >> >> primary index makes slice creation cheaper, but its positions remain > in the > >> >> parent Data.db coordinate space. Readers and tools must therefore > >> >> understand that the Data.db component is a slice and translate those > >> >> positions before accessing the local file. > >> >> > >> >> The current proposal instead pays that cost once during splitting. It > >> >> rebases Index.db positions, slices CompressionInfo.db, and rebuilds > or > >> >> apportions the child’s derived metadata. This preserves the usual > idea that > >> >> each SSTable is a self-contained artifact whose components share one > >> >> coordinate space and lifecycle. Index reconstruction is so cheap its > >> >> comparatively free relative to reading or copying Data.db, while also > >> >> producing child-specific summaries, Bloom filters, and statistics > (100's of > >> >> ms range on *huge* sstables) > >> >> > >> >> For durable outputs, I prefer keeping that complexity at creation > time. > >> >> Sharing one physical index file among several logical SSTables would > >> >> introduce additional lifecycle and bookkeeping cases around deletion, > >> >> snapshots, backup and restore, import, and third-party tooling. > Independent > >> >> files that happen to share filesystem extents are less concerning > because > >> >> they retain normal component ownership semantics. > >> >> > >> >> That said, the DataStax approach is worth benchmarking, and it may > be a > >> >> better fit for BTI or another format where slicing is designed in as > a > >> >> first-class property. > >> >> > >> >> > Separately, being able to easily slice and dice files without > looking > >> >> inside them is a key consideration in the file format we are working > on for > >> >> CEP-57. > >> >> > >> >> Agreed,nCEP-57 seems like the right place to make sliceability a > native > >> >> format property. My goal here is narrower: provide the capability > for BIG > >> >> SSTables and current formats in the meantime without permanently > >> >> introducing slice-coordinate awareness throughout BIG’s read path. > BIG will > >> >> remain in use for some time, and this mechanism can deliver most of > the > >> >> benefit with relatively contained changes. > >> >> > >> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]> > wrote: > >> >> > >> >> > Hello Chris, > >> >> > > >> >> > A while back we implemented a similar approach for DSE's version of > >> >> > zero-copy streaming. The main difference between our approach and > yours is > >> >> > that we decided not to split the primary index files and instead > use them > >> >> > as they are, with filtering based on the start and end key of the > section. > >> >> > In the context of local operations like anticompaction, the latter > may be a > >> >> > better approach as one can share the index files between all > resulting > >> >> > sections (and, of course, no index reconstruction is necessary). > >> >> > > >> >> > This code is not currently part of the Apache codebase, but > DataStax's > >> >> > open source fork includes support for reading these files, whose > >> >> > implementation (commit > >> >> > > https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422 > ) > >> >> > may have some ideas you can use in designing your solution. > >> >> > > >> >> > Separately, being able to easily slice and dice files without > looking > >> >> > inside them is a key consideration in the file format we are > working on for > >> >> > CEP-57. > >> >> > > >> >> > Regards, > >> >> > Branimir > >> >> > > >> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella < > >> >> > [email protected]> wrote: > >> >> > > >> >> >> Really nice idea Chris! > >> >> >> > >> >> >> One thing I think it’s worth flagging for the proposal: > >> >> >> > >> >> >> this CEP introduces a new SSTable major version. That's a > relevant change > >> >> >> for anything outside nodetool/the core read path that parses > SSTables > >> >> >> directly. e.g. analytics library uses the concept of per major > version > >> >> >> bridge to deserialize. > >> >> >> > >> >> >> With this approach, current bridges don't have a notion of "data > doesn't > >> >> >> start at logical offset zero” - they assume a chunk's first > partition > >> >> >> begins the SSTable's data - which is fine (this is a new > version!). It'd > >> >> >> help if the CEP explicitly calls out that third-party/off-node > SSTable > >> >> >> readers are a compatibility surface here, not just in-process > Cassandra > >> >> >> binaries. > >> >> >> > >> >> >> Bernardo > >> >> >> > >> >> >> *From: *Chris Lohfink <[email protected]> > >> >> >> *Date: *Monday, 14 September 2026 at 21:23 > >> >> >> *To: *[email protected] <[email protected]> > >> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting > >> >> >> > >> >> >> Hi everyone, > >> >> >> > >> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for > discussion: > >> >> >> > >> >> >> > >> >> >> * > https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting* > >> >> >> < > https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting > > > >> >> >> > >> >> >> > >> >> >> Anticompaction and partial-range streaming currently rewrite rows > whose > >> >> >> encoded representation already exists on disk. This consumes CPU, > creates > >> >> >> substantial heap churn and write amplification, and increases > temporary > >> >> >> disk pressure. > >> >> >> > >> >> >> CEP-66 proposes splitting eligible compressed SSTables by > retaining > >> >> >> contiguous runs of their existing compression chunks. Cassandra > would > >> >> >> rebuild the child SSTables' indexes and other derived components > without > >> >> >> deserializing, serializing, or recompressing their rows. > >> >> >> > >> >> >> "Zero-copy" here primarily means reusing the encoded bytes > instead of > >> >> >> rewriting rows. On filesystems that support range reflinks, > Cassandra can > >> >> >> also share the underlying extents meaning no new data written. > Other > >> >> >> filesystems, including ext4, would copy the already-compressed > bytes and > >> >> >> still avoid the row rewrite. > >> >> >> > >> >> >> The proposal is staged. It starts with an opt-in `sstablesplit > >> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk. > Later > >> >> >> phases add BTI support, secondary indexes, anticompaction, and > >> >> >> partial-range streaming. Existing implementations remain the > default and > >> >> >> provide the fallback for unsupported inputs. > >> >> >> > >> >> >> I'd particularly appreciate feedback on: > >> >> >> > >> >> >> - The retained-prefix representation and proposed Cassandra 7.0 > SSTable > >> >> >> format change > >> >> >> - Rebuilding or conservatively deriving child metadata without > decoding > >> >> >> rows > >> >> >> - The integrity and performance tradeoff around `Digest.crc32` > generation > >> >> >> - The staged rollout, compatibility rules, and fallback behavior > >> >> >> - Any correctness, operational, or filesystem concerns the > proposal has > >> >> >> missed > >> >> >> > >> >> >> Thanks, and I look forward to the discussion. > >> >> >> > >> >> >> Regards, > >> >> >> Chris Lohfink > >> >> >> > >> >> > > >> >> > >> >
