I actually believe you can still do it in BTI; it just requires reading some of the Data component. We wont need to deserialize everything like current but will still have IO costs. I think when we get to the BTI part of implementation we can experiment with a few different approaches more thoroughly including possibly the DSE implementation but I'm hesitant to have sstables be dependent on each other's components for operational simplicity. There's also the option of reusing the parents bloom filter at the cost of increased false positives.
Chris On Wed, Sep 30, 2026 at 5:11 PM Dmitry Konstantinov <[email protected]> wrote: > Regarding the BTI scenario, if I am not mistaken for narrow partitions we > do not have a full partition key within a primary index, it is only in Data > file, so to rebuild things like bloom filters we will have to read the data > file itself... I suppose it can be only one of the reasons why DSE > implementation does not split primary indexes. > > On Tue, 22 Sept 2026 at 20:03, Patrick McFadin <[email protected]> wrote: > >> This is such a clever way to solve a problem I wish would go away. The >> only negative aspect of this proposal is that it isn't available today. >> >> +1 from me. >> >> On Tue, Sep 22, 2026 at 7:18 AM Andy Tolbert <[email protected]> >> wrote: >> >>> Realized I didn't share a +1 in my previous message, so adding my +1! >>> >>> Thanks, >>> Andy >>> >>> On Tue, Sep 22, 2026, at 2:01 PM, Francisco Guerrero wrote: >>> > +1. Anticompaction is an area that causes a lot of pain and >>> > seeing this proposal makes me hopeful that things will get >>> > better soon. >>> > >>> > Also, I think we should consider bringing this work to 6.0 as >>> > well assuming this lands before we release 6.0 >>> > >>> > Best, >>> > - Francisco >>> > >>> > On 2026/09/22 13:43:37 Abe Ratnofsky wrote: >>> >> I’m +1 on the CEP. >>> >> >>> >> Saving bandwidth is meaningful, particularly for those running on >>> networked disks where bandwidth has a lower ceiling and a direct marginal >>> cost. The project currently recommends NVMe but the cost and convenience >>> benefits of newer generations of networked disks are meaningful. >>> >> >>> >> Regarding portability: this feels no different from supporting >>> multiple JDK versions, or recommending ACCP, etc. You’ll get better >>> performance on certain systems where certain features are available. This >>> change also introduces performance improvements for filesystems that do not >>> support FICLONERANGE like ext4. >>> >> >>> >> I do wish it were possible for us to version SSTables in a way that >>> let users experiment with this more easily. Requiring it land in a new >>> major keeps it further away from many users who would benefit. >>> >> >>> >> On Tue, Sep 22, 2026, at 5:38 AM, Benedict Elliott Smith wrote: >>> >> > My biggest concern with this proposal is whether linking inodes is >>> >> > justified. This narrows the applicability (only supported >>> filesystems) >>> >> > and increases the complexity, but in return saves bandwidth and >>> allows >>> >> > it to function when disk space is exhausted. >>> >> > >>> >> > I am of the opinion it is probably not justified, as bandwidth is >>> not I >>> >> > think typically a major constraint - we don't generally saturate >>> the >>> >> > disk because we are CPU-inefficient. Simply copying the files would >>> be >>> >> > a better starting point as much simpler and more portable. We have >>> >> > bigger problems when we are out of space. >>> >> > >>> >> > Another thing to balance is whether this complexity is justified >>> for a >>> >> > stop-gap measure, if we expect this to be made defunct by both >>> >> > Branimir's new file format (which permits cheaper slicing) and >>> mutation >>> >> > tracking (which should eliminate the need for anti-compaction). >>> >> > >>> >> > Separately, I wonder (if we desperately want it in the meantime) >>> >> > whether DataStax are willing to contribute their version of this, >>> if it >>> >> > already exists and is already validated on real workloads. >>> >> > >>> >> > Finally, have we explored simply removing anti-compaction instead? >>> This >>> >> > would require I think a couple of components: 1) per-range repair >>> >> > metadata; 2) either per-range sstable invalidations (so that >>> compaction >>> >> > may proceed on parts of the repaired/unrepaired file independently, >>> >> > permitting it to be replaced in both sets), or rewriting the >>> >> > uncompacted part(s) of the sstable. I may be missing some other >>> >> > complexity, but this seems quite tractable. >>> >> > >>> >> > >>> >> > >>> >> > On 2026/09/16 12:25:28 Chris Lohfink wrote: >>> >> >> Thanks, Branimir. The shared-index approach is an attractive option >>> >> >> especially for transient local operations such as anticompaction. >>> >> >> >>> >> >> The tradeoff appears to be where the complexity lives. Reusing the >>> original >>> >> >> primary index makes slice creation cheaper, but its positions >>> remain in the >>> >> >> parent Data.db coordinate space. Readers and tools must therefore >>> >> >> understand that the Data.db component is a slice and translate >>> those >>> >> >> positions before accessing the local file. >>> >> >> >>> >> >> The current proposal instead pays that cost once during splitting. >>> It >>> >> >> rebases Index.db positions, slices CompressionInfo.db, and >>> rebuilds or >>> >> >> apportions the child’s derived metadata. This preserves the usual >>> idea that >>> >> >> each SSTable is a self-contained artifact whose components share >>> one >>> >> >> coordinate space and lifecycle. Index reconstruction is so cheap >>> its >>> >> >> comparatively free relative to reading or copying Data.db, while >>> also >>> >> >> producing child-specific summaries, Bloom filters, and statistics >>> (100's of >>> >> >> ms range on *huge* sstables) >>> >> >> >>> >> >> For durable outputs, I prefer keeping that complexity at creation >>> time. >>> >> >> Sharing one physical index file among several logical SSTables >>> would >>> >> >> introduce additional lifecycle and bookkeeping cases around >>> deletion, >>> >> >> snapshots, backup and restore, import, and third-party tooling. >>> Independent >>> >> >> files that happen to share filesystem extents are less concerning >>> because >>> >> >> they retain normal component ownership semantics. >>> >> >> >>> >> >> That said, the DataStax approach is worth benchmarking, and it may >>> be a >>> >> >> better fit for BTI or another format where slicing is designed in >>> as a >>> >> >> first-class property. >>> >> >> >>> >> >> > Separately, being able to easily slice and dice files without >>> looking >>> >> >> inside them is a key consideration in the file format we are >>> working on for >>> >> >> CEP-57. >>> >> >> >>> >> >> Agreed,nCEP-57 seems like the right place to make sliceability a >>> native >>> >> >> format property. My goal here is narrower: provide the capability >>> for BIG >>> >> >> SSTables and current formats in the meantime without permanently >>> >> >> introducing slice-coordinate awareness throughout BIG’s read path. >>> BIG will >>> >> >> remain in use for some time, and this mechanism can deliver most >>> of the >>> >> >> benefit with relatively contained changes. >>> >> >> >>> >> >> On Wed, Sep 16, 2026 at 3:04 AM Branimir Lambov < >>> [email protected]> wrote: >>> >> >> >>> >> >> > Hello Chris, >>> >> >> > >>> >> >> > A while back we implemented a similar approach for DSE's version >>> of >>> >> >> > zero-copy streaming. The main difference between our approach >>> and yours is >>> >> >> > that we decided not to split the primary index files and instead >>> use them >>> >> >> > as they are, with filtering based on the start and end key of >>> the section. >>> >> >> > In the context of local operations like anticompaction, the >>> latter may be a >>> >> >> > better approach as one can share the index files between all >>> resulting >>> >> >> > sections (and, of course, no index reconstruction is necessary). >>> >> >> > >>> >> >> > This code is not currently part of the Apache codebase, but >>> DataStax's >>> >> >> > open source fork includes support for reading these files, whose >>> >> >> > implementation (commit >>> >> >> > >>> https://github.com/datastax/cassandra/commit/38c44d1abcf2793337b7e954fc98517c5691f422 >>> ) >>> >> >> > may have some ideas you can use in designing your solution. >>> >> >> > >>> >> >> > Separately, being able to easily slice and dice files without >>> looking >>> >> >> > inside them is a key consideration in the file format we are >>> working on for >>> >> >> > CEP-57. >>> >> >> > >>> >> >> > Regards, >>> >> >> > Branimir >>> >> >> > >>> >> >> > On Tue, Sep 15, 2026 at 1:11 PM Bernardo Botella < >>> >> >> > [email protected]> wrote: >>> >> >> > >>> >> >> >> Really nice idea Chris! >>> >> >> >> >>> >> >> >> One thing I think it’s worth flagging for the proposal: >>> >> >> >> >>> >> >> >> this CEP introduces a new SSTable major version. That's a >>> relevant change >>> >> >> >> for anything outside nodetool/the core read path that parses >>> SSTables >>> >> >> >> directly. e.g. analytics library uses the concept of per major >>> version >>> >> >> >> bridge to deserialize. >>> >> >> >> >>> >> >> >> With this approach, current bridges don't have a notion of >>> "data doesn't >>> >> >> >> start at logical offset zero” - they assume a chunk's first >>> partition >>> >> >> >> begins the SSTable's data - which is fine (this is a new >>> version!). It'd >>> >> >> >> help if the CEP explicitly calls out that third-party/off-node >>> SSTable >>> >> >> >> readers are a compatibility surface here, not just in-process >>> Cassandra >>> >> >> >> binaries. >>> >> >> >> >>> >> >> >> Bernardo >>> >> >> >> >>> >> >> >> *From: *Chris Lohfink <[email protected]> >>> >> >> >> *Date: *Monday, 14 September 2026 at 21:23 >>> >> >> >> *To: *[email protected] <[email protected]> >>> >> >> >> *Subject: *[DISCUSS] CEP-66: Zero-copy SSTable splitting >>> >> >> >> >>> >> >> >> Hi everyone, >>> >> >> >> >>> >> >> >> I'd like to open CEP-66, Zero-copy SSTable splitting, for >>> discussion: >>> >> >> >> >>> >> >> >> >>> >> >> >> * >>> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting* >>> >> >> >> < >>> https://cwiki.apache.org/confluence/spaces/CASSANDRA/pages/451972773/draft+CEP-66+Zero-copy+SSTable+splitting >>> > >>> >> >> >> >>> >> >> >> >>> >> >> >> Anticompaction and partial-range streaming currently rewrite >>> rows whose >>> >> >> >> encoded representation already exists on disk. This consumes >>> CPU, creates >>> >> >> >> substantial heap churn and write amplification, and increases >>> temporary >>> >> >> >> disk pressure. >>> >> >> >> >>> >> >> >> CEP-66 proposes splitting eligible compressed SSTables by >>> retaining >>> >> >> >> contiguous runs of their existing compression chunks. Cassandra >>> would >>> >> >> >> rebuild the child SSTables' indexes and other derived >>> components without >>> >> >> >> deserializing, serializing, or recompressing their rows. >>> >> >> >> >>> >> >> >> "Zero-copy" here primarily means reusing the encoded bytes >>> instead of >>> >> >> >> rewriting rows. On filesystems that support range reflinks, >>> Cassandra can >>> >> >> >> also share the underlying extents meaning no new data written. >>> Other >>> >> >> >> filesystems, including ext4, would copy the already-compressed >>> bytes and >>> >> >> >> still avoid the row rewrite. >>> >> >> >> >>> >> >> >> The proposal is staged. It starts with an opt-in `sstablesplit >>> >> >> >> --zero-copy` mode for BIG-format SSTables in Cassandra >>> 7.0/trunk. Later >>> >> >> >> phases add BTI support, secondary indexes, anticompaction, and >>> >> >> >> partial-range streaming. Existing implementations remain the >>> default and >>> >> >> >> provide the fallback for unsupported inputs. >>> >> >> >> >>> >> >> >> I'd particularly appreciate feedback on: >>> >> >> >> >>> >> >> >> - The retained-prefix representation and proposed Cassandra 7.0 >>> SSTable >>> >> >> >> format change >>> >> >> >> - Rebuilding or conservatively deriving child metadata without >>> decoding >>> >> >> >> rows >>> >> >> >> - The integrity and performance tradeoff around `Digest.crc32` >>> generation >>> >> >> >> - The staged rollout, compatibility rules, and fallback behavior >>> >> >> >> - Any correctness, operational, or filesystem concerns the >>> proposal has >>> >> >> >> missed >>> >> >> >> >>> >> >> >> Thanks, and I look forward to the discussion. >>> >> >> >> >>> >> >> >> Regards, >>> >> >> >> Chris Lohfink >>> >> >> >> >>> >> >> > >>> >> >> >>> >> >>> >> > > -- > Dmitry Konstantinov >
