+1. Anticompaction is an area that causes a lot of pain and seeing this proposal makes me hopeful that things will get better soon.
Also, I think we should consider bringing this work to 6.0 as well assuming this lands before we release 6.0 Best, - Francisco On 2026/09/22 13:43:37 Abe Ratnofsky wrote: > I’m +1 on the CEP. > > Saving bandwidth is meaningful, particularly for those running on networked > disks where bandwidth has a lower ceiling and a direct marginal cost. The > project currently recommends NVMe but the cost and convenience benefits of > newer generations of networked disks are meaningful. > > Regarding portability: this feels no different from supporting multiple JDK > versions, or recommending ACCP, etc. You’ll get better performance on certain > systems where certain features are available. This change also introduces > performance improvements for filesystems that do not support FICLONERANGE > like ext4. > > I do wish it were possible for us to version SSTables in a way that let users > experiment with this more easily. Requiring it land in a new major keeps it > further away from many users who would benefit. > > On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote: > > My biggest concern with this proposal is whether linking inodes is > > justified. This narrows the applicability (only supported filesystems) > > and increases the complexity, but in return saves bandwidth and allows > > it to function when disk space is exhausted. > > > > I am of the opinion it is probably not justified, as bandwidth is not I > > think typically a major constraint - we don't generally saturate the > > disk because we are CPU-inefficient. Simply copying the files would be > > a better starting point as much simpler and more portable. We have > > bigger problems when we are out of space. > > > > Another thing to balance is whether this complexity is justified for a > > stop-gap measure, if we expect this to be made defunct by both > > Branimir's new file format (which permits cheaper slicing) and mutation > > tracking (which should eliminate the need for anti-compaction). > > > > Separately, I wonder (if we desperately want it in the meantime) > > whether DataStax are willing to contribute their version of this, if it > > already exists and is already validated on real workloads. > > > > Finally, have we explored simply removing anti-compaction instead? This > > would require I think a couple of components: 1) per-range repair > > metadata; 2) either per-range sstable invalidations (so that compaction > > may proceed on parts of the repaired/unrepaired file independently, > > permitting it to be replaced in both sets), or rewriting the > > uncompacted part(s) of the sstable. I may be missing some other > > complexity, but this seems quite tractable. > > > > > > > > On 2026/09/16 12:25:28 Chris Lohfink wrote: > >> Thanks, Branimir. The shared-index approach is an attractive option > >> especially for transient local operations such as anticompaction. > >> > >> The tradeoff appears to be where the complexity lives. Reusing the original > >> primary index makes slice creation cheaper, but its positions remain in the > >> parent Data.db coordinate space. Readers and tools must therefore > >> understand that the Data.db component is a slice and translate those > >> positions before accessing the local file. > >> > >> The current proposal instead pays that cost once during splitting. It > >> rebases Index.db positions, slices CompressionInfo.db, and rebuilds or > >> apportions the child’s derived metadata. This preserves the usual idea that > >> each SSTable is a self-contained artifact whose components share one > >> coordinate space and lifecycle. Index reconstruction is so cheap its > >> comparatively free relative to reading or copying Data.db, while also > >> producing child-specific summaries, Bloom filters, and statistics (100's of > >> ms range on *huge* sstables) > >> > >> For durable outputs, I prefer keeping that complexity at creation time. > >> Sharing one physical index file among several logical SSTables would > >> introduce additional lifecycle and bookkeeping cases around deletion, > >> snapshots, backup and restore, import, and third-party tooling. Independent > >> files that happen to share filesystem extents are less concerning because > >> they retain normal component ownership semantics. > >> > >> That said, the DataStax approach is worth benchmarking, and it may be a > >> better fit for BTI or another format where slicing is designed in as a > >> first-class property. > >> > >> > Separately, being able to easily slice and dice files without looking > >> inside them is a key consideration in the file format we are working on for > >> CEP-57. > >> > >> Agreed,nCEP-57 seems like the right place to make sliceability a native > >> format property. My goal here is narrower: provide the capability for BIG > >> SSTables and current formats in the meantime without permanently > >> introducing slice-coordinate awareness throughout BIG’s read path. BIG will > >> remain in use for some time, and this mechanism can deliver most of the > >> benefit with relatively contained changes. > >> > >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov <[email protected]> wrote: > >> > >> > Hello Chris, > >> > > >> > A while back we implemented a similar approach for DSE's version of > >> > zero-copy streaming. The main difference between our approach and yours > >> > is > >> > that we decided not to split the primary index files and instead use them > >> > as they are, with filtering based on the start and end key of the > >> > section. > >> > In the context of local operations like anticompaction, the latter may > >> > be a > >> > better approach as one can share the index files between all resulting > >> > sections (and, of course, no index reconstruction is necessary). > >> > > >> > This code is not currently part of the Apache codebase, but DataStax's > >> > open source fork includes support for reading these files, whose > >> > implementation (commit > >> > https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422) > >> > may have some ideas you can use in designing your solution. > >> > > >> > Separately, being able to easily slice and dice files without looking > >> > inside them is a key consideration in the file format we are working on > >> > for > >> > CEP-57. > >> > > >> > Regards, > >> > Branimir > >> > > >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella < > >> > [email protected]> wrote: > >> > > >> >> Really nice idea Chris! > >> >> > >> >> One thing I think it’s worth flagging for the proposal: > >> >> > >> >> this CEP introduces a new SSTable major version. That's a relevant > >> >> change > >> >> for anything outside nodetool/the core read path that parses SSTables > >> >> directly. e.g. analytics library uses the concept of per major version > >> >> bridge to deserialize. > >> >> > >> >> With this approach, current bridges don't have a notion of "data doesn't > >> >> start at logical offset zero” - they assume a chunk's first partition > >> >> begins the SSTable's data - which is fine (this is a new version!). It'd > >> >> help if the CEP explicitly calls out that third-party/off-node SSTable > >> >> readers are a compatibility surface here, not just in-process Cassandra > >> >> binaries. > >> >> > >> >> Bernardo > >> >> > >> >> *From: *Chris Lohfink <[email protected]> > >> >> *Date: *Monday, 14 September 2026 at 21:23 > >> >> *To: *[email protected] <[email protected]> > >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting > >> >> > >> >> Hi everyone, > >> >> > >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for discussion: > >> >> > >> >> > >> >> *https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting* > >> >> <https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting> > >> >> > >> >> > >> >> Anticompaction and partial-range streaming currently rewrite rows whose > >> >> encoded representation already exists on disk. This consumes CPU, > >> >> creates > >> >> substantial heap churn and write amplification, and increases temporary > >> >> disk pressure. > >> >> > >> >> CEP-66 proposes splitting eligible compressed SSTables by retaining > >> >> contiguous runs of their existing compression chunks. Cassandra would > >> >> rebuild the child SSTables' indexes and other derived components without > >> >> deserializing, serializing, or recompressing their rows. > >> >> > >> >> "Zero-copy" here primarily means reusing the encoded bytes instead of > >> >> rewriting rows. On filesystems that support range reflinks, Cassandra > >> >> can > >> >> also share the underlying extents meaning no new data written. Other > >> >> filesystems, including ext4, would copy the already-compressed bytes and > >> >> still avoid the row rewrite. > >> >> > >> >> The proposal is staged. It starts with an opt-in `sstablesplit > >> >> --zero-copy` mode for BIG-format SSTables in Cassandra 7.0/trunk. Later > >> >> phases add BTI support, secondary indexes, anticompaction, and > >> >> partial-range streaming. Existing implementations remain the default and > >> >> provide the fallback for unsupported inputs. > >> >> > >> >> I'd particularly appreciate feedback on: > >> >> > >> >> - The retained-prefix representation and proposed Cassandra 7.0 SSTable > >> >> format change > >> >> - Rebuilding or conservatively deriving child metadata without decoding > >> >> rows > >> >> - The integrity and performance tradeoff around `Digest.crc32` > >> >> generation > >> >> - The staged rollout, compatibility rules, and fallback behavior > >> >> - Any correctness, operational, or filesystem concerns the proposal has > >> >> missed > >> >> > >> >> Thanks, and I look forward to the discussion. > >> >> > >> >> Regards, > >> >> Chris Lohfink > >> >> > >> > > >> >
