For the record Josh, I am including pre-3.0. I was involved in customer support 
escalations for 1.2, 2.0 and 2.1: the software was shoddy, to put it mildly. 
8099 was a symptom, not the disease.

The project had a culture of doing stuff without sufficient care or 
consideration. This is by no means a simple matter of needing more testing or 
less timeline pressure, nor is it easy to define a "bar" to ensure it does not 
regress. If we had simple metrics, we would have settled it a long time ago.

You can still see customer scars crop up on forum discussions periodically. 


On 2026/09/24 19:01:02 Josh McKenzie wrote:
> > Apache Cassandra was fundamentally undeployable for four years between Nov 
> > 2015 - 2019.
> CASSANDRA-8099 was a maximal manifestation of a specific approach to 
> engineering and calendar constraints we've seen time and again on the 
> project; I don't want us to conflate things here. That was a herculean 
> monolithic body of work performed in inhuman conditions (in vim!) that was 
> ultimately so invasive, all the unit tests in the code-base were commented 
> out and Jake and I spent a grueling 1.5-2 months hand-rewriting basically all 
> the unit tests in that code-base to get things to even build and run, much 
> less pass.
> 
> Massive blast radius changes that are un-sustainably complex, under-tested, 
> where we don't property, fuzz, check coverage, check complexity, or A/B 
> compare against a known good system (i.e. pre/post) correctness testing are 
> going to destabilize the database at any time under any regime of tooling. We 
> certainly could speed-run our way back into destabilization with LLM's as 
> they are a force multiplier for both the good and the bad of one's 
> engineering practices, but that's a solvable problem by better defining what 
> our bars of quality are (Definition of Done anyone?) and holding ourselves 
> accountable to delivering at that bar. 
> 
> On Thu, Sep 24, 2026, at 2:40 PM, David Capwell via dev wrote:
> >> 
> >> 
> >> if you really want to pursue it I would ask that we do it offline to avoid 
> >> polluting an already busy conversation
> >> 
> > People are directly responding saying that they feel discrimination 
> > currently and that the policy tries to codify that discrimination, so I 
> > feel its 100% on topic. This thread has presented 0 evidence that LLM usage 
> > has lowered the quality of contributions merged and has so far been vibes 
> > and feeling; I have yet to see any evidence to justify such discrimination 
> > so I will keep pushing back until such evidence is presented so we can have 
> > a informed debate.
> > 
> >> namely how we handle shallow and localised bug fixes. I would be happy 
> >> adding a clear entry to the “Permitted” section for this. No doubt there 
> >> are many refinements needed to the Restricted text as well, that might 
> >> also capture some of your concerns.
> >> 
> > I will not sign off on cherry picking areas where "safe" to use a tool; so 
> > no it does not capture my concerns.
> > 
> > 
> > 
> >> I don't know if everyone remembers, but ten years ago Cassandra was full 
> >> of serious correctness and stability issues. Despite developing it, I 
> >> would not have run it myself or recommend that anyone use it. We have dug 
> >> ourselves out of that hole, but it took years of discipline and effort, 
> >> and we're still (deservedly) recovering our reputation. 
> >> 
> >> Let's use this new technology to improve the quality of our contributions, 
> >> not squander our hard-earned gains in the name of speed. It will be hard 
> >> to recover our reputation a second time.
> >> 
> > I do recall the 3.x line and put in a significant amount of effort to 
> > harden it. There were behaviors I noticed after joining Cassandra that I 
> > feel directly contributed to 3.0 and the decade of catch up; behaviors that 
> > still linger in parts of the community today.
> > 
> > As I look on trunk and look at committed code and trace back to PRs and 
> > JIRA I see the following:
> > 
> >  • large patches approved without comments
> >  • 0 evidence that tests were run
> > I then look at our CI and see tests failing for months. As you start to 
> > triage you start to see some of them show real issues; yet they linger for 
> > months not being addressed... when CI is unstable it takes a lot of effort 
> > to triage "did my patch break the test", and I have seen time and time 
> > again people do not put in that effort, and shrug off as "its just a flakey 
> > test"; then our CI failure rate grows.
> > 
> > Non of this has anything to do with LLMs but LLMs running in this 
> > environment is far more dangerous as there are not checks in place to "hold 
> > the bar". I am all for raising the bar universally; expecting both humans 
> > and LLMs to match that bar.
> > 
> >> `I’d support something that boils down to roughly this:
> >> 
> >> 1.) 2 committers must understand an LLM-assisted change before it commits. 
> >> (Perhaps separately we can explore the question of why we haven’t added 
> >> any new committers to the core project for about a year. I’m also still 
> >> not entirely sure if it’s acceptable within our guidelines for a committer 
> >> to +1 a patch after delegating review.)
> >> 
> >> 2.) Patch authors must demonstrate enough understanding to discuss their 
> >> own patch, whether or not parts of it are generated by an LLM.
> >> 
> >> 3.) The “Assisted-by” tag should be used to indicate any non-trivial LLM 
> >> usage in the generation of a patch, just like we have used Co-authored-by 
> >> historically.
> >> 
> >> 4.) Comments and other things that aren't the actual code (but could sow 
> >> confusion) should be held to the same standard we'd expect from a human 
> >> writer. If we don’t yet agree on that standard, we can formalize enough of 
> >> it to guide both humans and LLMs.
> `
> > I can get behind this proposal but i would tweak it as 1/2 i don't think 
> > really need to special case LLM usage
> > 
> >  1. 2 committers must understand the change before it commits.
> >  2. Patch authors must demonstrate enough understanding to discuss their 
> > own patch
> > Nothing about those 2 need to be scoped to LLM usage and honestly matches 
> > most PMCs I have talked to understanding of our bar (as Benedict pointed 
> > out, the actual wording could be interpreted to allow rubber stamping from 
> > committer)
> > 
> > As for 3 I am cool with this. ASF recommends the same (it says 
> > `Generated-by` but that discount's the human's effort) as its useful for 
> > audits and tooling. Having `Assisted-by` tag should not imply anything 
> > about the committed patch as it should have gone through the same bar we 
> > all expect; its just for tools auditing.
> > 
> > 
> >> On Sep 24, 2026, at 10:37 AM, Aleksey Yeshchenko via dev 
> >> <[email protected]> wrote:
> >> 
> >> Meant "doing away with", sorry. Non-native speaker with a headache here. 
> >> Thanks Caleb for spotting.
> >> 
> >>> On 24 Sep 2026, at 17:35, Aleksey Yeshchenko via dev 
> >>> <[email protected]> wrote:
> >>> 
> >>> P.S. I assume it's obvious from the text above that I don't believe that 
> >>> getting away with human code review is a viable option.
> >> 
> >> 
> >>> On 24 Sep 2026, at 17:37, Štefan Miklošovič <[email protected]> 
> >>> wrote:
> >>> 
> >>> Good call on checkerframework, we even have a patch for it. Work of
> >>> Jacek Lewandowski. We might just drive it to completion. Using AI for
> >>> finishing it would be quite ironic.
> >>> 
> >>> (1) https://github.com/apache/cassandra/pull/2370
> >>> 
> >>> On Thu, Sep 24, 2026 at 6:19 PM Jon Haddad <[email protected]> 
> >>> wrote:
> >>>> 
> >>>> There are some really good points being brought up about stability of 
> >>>> the codebase, maintainability, quality of reviews, correctness bugs, and 
> >>>> I agree with all of them.  I think it would be helpful to take a step 
> >>>> back and consider how those bugs got there in the first place, how they 
> >>>> were fixed, and what we could do to further advance the codebase so they 
> >>>> don't creep back.  LLMs can be used either with very tight guardrails, 
> >>>> or in what's effectively YOLO mode, and there's a big difference in the 
> >>>> quality of the results you get.
> >>>> 
> >>>> One thing to keep in mind, a lot of the initial code in C* was added 
> >>>> without comprehensive testing.  I hope we can all agree that it's a lot 
> >>>> easier to break code that doesn't have high quality tests.  During the 
> >>>> code freeze, a lot of people people spent several years relentlessly 
> >>>> finding and fixing bugs. This was probably a pretty frustrating time for 
> >>>> anyone who was focused on fixing other people's bugs when they wanted to 
> >>>> build features.  I think we should recognize the effort here and 
> >>>> appreciate the foundation that the project stands on now. I can 
> >>>> understand how anyone involved with this effort would be apprehensive 
> >>>> about seeing years of their life swept away by an agent that was driven 
> >>>> by goal seeking to remove all the tests that it broke instead of fixing 
> >>>> them.
> >>>> 
> >>>> When I picked up the work to improve cursor compaction, the first thing 
> >>>> I asked myself was how can I make sure I don't break this?  How do I 
> >>>> even know it works properly?  There were some tricky parts to the code, 
> >>>> and I really didn't want to come in and immediately break stuff.  That's 
> >>>> why I started with an entire patch dedicated to adding test infra to it. 
> >>>>  90% of the patch was tests, and in my other cursor patches, it remains 
> >>>> *at least* 80% of my patches.  It was a *lot* faster to add almost 10K 
> >>>> lines of tests that handled a byte for byte differential testing paired 
> >>>> with harry to find over 30 bugs that caused cursor to corrupt results.  
> >>>> Range tombstones alone were at least a dozen bugs, but I also found 
> >>>> issues with static columns, reverse ordering, etc.  Randomizing schemas 
> >>>> and data in burn tests to generate different shapes of data, to ensure 
> >>>> they all result in the same output at the end. JMH tests to ensure there 
> >>>> weren't performance regressions, hours of profiling. These were all a 
> >>>> *lot* easier to do with the LLM helping me out.  In the process I've 
> >>>> found bugs that have been lingering in the codebase for years.
> >>>> 
> >>>> That's a long story, but hopefully we all agree that having 
> >>>> comprehensive tests is a great way to ensure that both humans and LLMs 
> >>>> don't break things that are working.
> >>>> 
> >>>> The lesson: we need to keep improving our testing.  Everything that we 
> >>>> touch, should be left in a better state than how we found it with regard 
> >>>> to test coverage.
> >>>> 
> >>>> Test coverage isn't everything though, there's always little subtle bugs 
> >>>> that don't get found in testing, that can slip in despite our best 
> >>>> efforts.  It's debatable if humans will be as good as agents for coding 
> >>>> in the long term, for spotting small defects.  I sincerely doubt it.  
> >>>> For the time being though, we still have people involved. It's probably 
> >>>> a good time to start using more static analysis tools to identify 
> >>>> problematic code and to add this to CI.  Dmitry had a suggestion 
> >>>> recently for checkerframework to detect leaking contexts, a problem he 
> >>>> spotted when reviewing my branch.  It would be great to have that 
> >>>> integrated into our CI and dev workflow so we can simply avoid an entire 
> >>>> class of bugs.
> >>>> 
> >>>> There's also PMD, which is excellent for finding code that can be hard 
> >>>> to understand.  I *highly* suggest you all run PMD to analyze for 
> >>>> cognitive complexity and high npath scores.  This was made popular by 
> >>>> the folks at Sonar and I've found it to be an excellent feedback 
> >>>> mechanism for structuring code.  The default max they set is 15, which 
> >>>> is the point where it starts to become difficult to verify something 
> >>>> works without making a massive investment.  We've got areas in the 
> >>>> codebase that are in the hundreds, and some parts even higher.  These 
> >>>> have been contributed by humans, and are all high risk points for both 
> >>>> humans and agents to start messing around with.  They're also in some 
> >>>> fairly critical areas that are very likely to break, so I understand why 
> >>>> people would not want an agent anywhere near it.
> >>>> 
> >>>> Unfortunately, it's not an easy problem to address.  There's so many 
> >>>> places where the code is structured in a way that has so many branches, 
> >>>> so many conditions, that it's effectively impossible for a human to 
> >>>> understand, creating a fear of messing around in it.  There's plenty of 
> >>>> areas that deserve extreme scrutiny, and we should be careful of what we 
> >>>> add, whether it's human or agent.
> >>>> 
> >>>> The codebase today requires a high degree of internal knowledge to 
> >>>> navigate.  There's land mines everywhere. We should be looking to make 
> >>>> conscious improvements by moving the code forward, so it's easier to 
> >>>> make changes to small, well tested components with minimal side effects. 
> >>>>  Not making it harder for people to use the tools that aid in that 
> >>>> process.
> >>>> 
> >>>> Here's what we could do to achieve the underlying goal of not breaking 
> >>>> the DB:
> >>>> 
> >>>> Add cognitive complexlity and npath via PMD as a feedback mechanism.
> >>>> 
> >>>> Code that's hard to understand is hard to review.  It's also hard to 
> >>>> test. Let's break down the complex code so more people can contribute, 
> >>>> safely.
> >>>> 
> >>>> Add checkerframework to our tooling,
> >>>> 
> >>>> Properly annotate the codebase for it and reduce the surface area that 
> >>>> things can break.  Less brittle codebase = we can move faster.
> >>>> 
> >>>> Use jacoco to find areas of the codebase with poor testing.
> >>>> 
> >>>> Let's improve the test coverage there, LLMs are great for this.  We have 
> >>>> a ton of static tests, these can become more dynamic, parameterized, and 
> >>>> leverage harry.
> >>>> 
> >>>> Refactor parts of the codebase that have high cognitive complexlity and 
> >>>> NPath scores.
> >>>> 
> >>>> This should be lowered over time to meet some high watermark, say 25 
> >>>> maximum, although I'd prefer 15 which is where the Sonar folks settled.
> >>>> 
> >>>> Move forward moving the codebase to a more modular structure
> >>>> 
> >>>> We've talked about Gradle on and off - but it can really be a huge help 
> >>>> with incremental, modular builds. This is pretty easy to do with an 
> >>>> agent and we could have it done in a couple days.
> >>>> 
> >>>> Enforce boundaries with ArchUnit
> >>>> 
> >>>> If we want to enforce certain code boundaries, this is the way to do it. 
> >>>> Should not be part of manual review.
> >>>> 
> >>>> Add LLM review for all incoming PRs before a human
> >>>> 
> >>>> The goal here is to automate the initial part of the review process that 
> >>>> reviewers should spot, and raise the bar for the initial contribution.  
> >>>> When the code gets reviewed by a human, it should already have passed a 
> >>>> large variety of initial checks.  This should shorten the review cycle 
> >>>> and result in higher quality patches.  I've had Claude reviewing all my 
> >>>> PRs in my personal projects for a while now and it consistently gives 
> >>>> great feedback that I almost always incorporate.
> >>>> 
> >>>> In my ideal world, we'd also auto-format all code
> >>>> 
> >>>> Consistent formatting throughout the codebase would be amazing, but 
> >>>> that's just one man's dream.
> >>>> 
> >>>> Hopefully there's at least a couple things in this list we could move 
> >>>> forward with in the short term, as it'll help improve the code quality 
> >>>> regardless of how it's created.
> >>>> 
> >>>> Jon
> >>>> 
> >>>> https://checkerframework.org/manual/#aliasing-leaking-contexts
> >>>> https://www.sonarsource.com/docs/CognitiveComplexity.pdf
> >>>> https://pmd.github.io/pmd/pmd_rules_java_design.html
> >>>> 
> >>>> 
> >>>> 
> >>>> 
> >>>> 
> >>>> On Thu, Sep 24, 2026 at 7:38 AM C. Scott Andreas <[email protected]> 
> >>>> wrote:
> >>>>> 
> >>>>> From Benedict:
> >>>>> 
> >>>>> “I don't know if everyone remembers, but ten years ago Cassandra was 
> >>>>> full of serious correctness and stability issues. Despite developing 
> >>>>> it, I would not have run it myself or recommend that anyone use it. We 
> >>>>> have dug ourselves out of that hole, but it took years of discipline 
> >>>>> and effort, and we're still (deservedly) recovering our reputation.”
> >>>>> 
> >>>>> Expanding on this point for those who may not have been active in the 
> >>>>> project at this time —
> >>>>> 
> >>>>> Apache Cassandra was fundamentally undeployable for four years between 
> >>>>> Nov 2015 - 2019. The database literally lost data if you ran a 
> >>>>> read-only SELECT query ordered descending (C-14513, C-14515). If you 
> >>>>> haven’t read these tickets before, please take a moment to do so.
> >>>>> 
> >>>>> It took years of careful work via property-based testing, fuzzing, and 
> >>>>> deterministic simulation to restore Cassandra’s status as a usable 
> >>>>> system of record. Once 14513 and 14515 were identified, nearly 30 
> >>>>> additional critical data loss and incorrect response bugs were 
> >>>>> identified.
> >>>>> 
> >>>>> It is essential for the project’s future that we don’t regress to this 
> >>>>> state chasing AI-generated features motivated by fear. The fact that 
> >>>>> examples cited in this thread which boast shiny features but have 
> >>>>> critical shortcomings unknown to their author supports this argument.
> >>>>> 
> >>>>> The most common path for large corpuses of AI-generated software is 
> >>>>> elation and reveling in a feature matrix, followed by abandonment.
> >>>>> 
> >>>>> I endorse this point:
> >>>>> 
> >>>>> “Let's use this new technology to improve the quality of our 
> >>>>> contributions, not squander our hard-earned gains in the name of speed. 
> >>>>> It will be hard to recover our reputation a second time.”
> >>>>> 
> >>>>> Patrick, I don’t want your note regarding a TCM issue to go 
> >>>>> unaddressed. Please file a Jira ticket and the patch if you like. I 
> >>>>> can’t comment on the patch as I haven’t seen it, but together we will 
> >>>>> solve the problem.
> >>>>> 
> >>>>> – Scott
> >>>>> 
> >>>>>> On Sep 24, 2026, at 4:01 AM, Benedict Elliott Smith 
> >>>>>> <[email protected]> wrote:
> >>>>>> 
> >>>>>> Hi Patrick,
> >>>>>> 
> >>>>>> As I mentioned in my reply to David, I would be happy to create a 
> >>>>>> carve out for shallow and localised bug fixes in the "Permitted" 
> >>>>>> section. Would this alleviate some of your concerns regarding your 
> >>>>>> ability to contribute to the project?
> >>>>>> 
> >>>>>> I appreciate your pointing out Ferrosa's Accord implementation 
> >>>>>> however, as it is a *great* example of the problems we're leaping 
> >>>>>> into. I took a look, and within about 30s found that the protocol is 
> >>>>>> fundamentally incorrect, having failed to address CASSANDRA-18365. 
> >>>>>> This is despite claiming to be tested with Jepsen that should in 
> >>>>>> principle find this fault. I followed up by using Claude to 
> >>>>>> interrogate the implementation further, and immediately found other 
> >>>>>> serious correctness issues.
> >>>>>> 
> >>>>>> I use LLMs daily now to help facilitate Accord development, and while 
> >>>>>> they are powerful they are NOT able to author the code themselves, 
> >>>>>> even when building upon a strong human-authored foundation.
> >>>>>> 
> >>>>>> I don't know if everyone remembers, but ten years ago Cassandra was 
> >>>>>> full of serious correctness and stability issues. Despite developing 
> >>>>>> it, I would not have run it myself or recommend that anyone use it. We 
> >>>>>> have dug ourselves out of that hole, but it took years of discipline 
> >>>>>> and effort, and we're still (deservedly) recovering our reputation.
> >>>>>> 
> >>>>>> Let's use this new technology to improve the quality of our 
> >>>>>> contributions, not squander our hard-earned gains in the name of 
> >>>>>> speed. It will be hard to recover our reputation a second time.
> >>>>>> 
> >>>>>> 
> >>>>>>> On 2026/09/23 19:16:17 Patrick McFadin wrote:
> >>>>>>> I was waiting for this moment to hit our project and I'm glad we're 
> >>>>>>> here. I
> >>>>>>> am deeply concerned for our project and its future, as we have 
> >>>>>>> increasingly
> >>>>>>> made it difficult to contribute. I had hoped that this new era of
> >>>>>>> software tools powered by AI would expand the project's reach and 
> >>>>>>> bring
> >>>>>>> more diverse thoughts and ideas. This policy proposal is the exact 
> >>>>>>> opposite
> >>>>>>> of what we need. We have been sitting on a Cassandra 6 release alpha 
> >>>>>>> for
> >>>>>>> months. We need to accelerate and embrace new ways of being or be left
> >>>>>>> behind. As I read that policy, my first and gut level reactions:
> >>>>>>> - It comes across as elitist and class protectionism. Committer 
> >>>>>>> should not
> >>>>>>> be special but this proposal makes that designation even more sacred.
> >>>>>>> - It signals that our project is so fragile that only a few people 
> >>>>>>> "Really
> >>>>>>> understand it" That's some SQLite vibes right there.
> >>>>>>> - Trying to fix a problem that doesn't exist
> >>>>>>> Sadly, i think this policy change would also exclude a lot of 
> >>>>>>> comitters.
> >>>>>>> We aren't alone in this moment. The Linux project just went through
> >>>>>>> this. You can find the thread with a simple Google, but similar hard
> >>>>>>> feelings were being expressed "AI is going to ruin our project!", "The
> >>>>>>> unwashed masses are going to contribute terrible code!", "We have to
> >>>>>>> protect our precious status as Linux maintainers!"  Linus being 
> >>>>>>> Linus, was
> >>>>>>> deeply invloved and they adopted a super simple statement that covers 
> >>>>>>> all
> >>>>>>> bases. Human or Human using AI. “You are expected to understand and 
> >>>>>>> to be
> >>>>>>> able to defend everything you submit.”  Love that.
> >>>>>>> In the larger picture, I'll restate. I'm worried for our project. In 
> >>>>>>> late
> >>>>>>> 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real corner 
> >>>>>>> and
> >>>>>>> in the hands of somebody that knows how to build software, this tool 
> >>>>>>> is
> >>>>>>> like jet fuel. Here's some examples of new projects being hyper 
> >>>>>>> fueled by
> >>>>>>> AI coding tools.
> >>>>>>> Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity
> >>>>>>> Turso - Rust re-write of SQLite
> >>>>>>> Bun - Rust re-write of itself from Zig.
> >>>>>>> Think this couldn't happen to us? Already has:
> >>>>>>> https://github.com/ferrosadb/ferrosa. Ben is using it to power his own
> >>>>>>> startup, but it was him alone using a ton of local AI coding agents. 
> >>>>>>> He
> >>>>>>> even implemented Accord. Yeah...
> >>>>>>> The cracks are already starting to show. There is a black market 
> >>>>>>> economy of
> >>>>>>> Cassandra patches happening now. Not going to name names or call 
> >>>>>>> people
> >>>>>>> out,  but there are fixes and optimizations living in branches 
> >>>>>>> outside of
> >>>>>>> the Cassandra project. Why? I'll use myself as an example. I fixed a 
> >>>>>>> nasty
> >>>>>>> bug I ran into with TCM a few weeks ago. Wrote the tests. It passes 
> >>>>>>> CI and
> >>>>>>> lives in my personal branch. I'm sitting here really wondering if I 
> >>>>>>> want to
> >>>>>>> go through the ritual humiliation of being roasted for using AI to 
> >>>>>>> fix it.
> >>>>>>> Me. I am worried about contrinuting code the Cassandra. What the hell 
> >>>>>>> does
> >>>>>>> that say?
> >>>>>>> I have my CQLite project that I've been doing a release around once a
> >>>>>>> month. I would love to donate that to the Cassandra project but I 
> >>>>>>> wouldn't
> >>>>>>> if it essentially killed any progress.
> >>>>>>> My larger counter proposal would be to:
> >>>>>>> - Adopt the “You are expected to understand and to be able to defend
> >>>>>>> everything you submit.” approach the Linux project has adopted.
> >>>>>>> - Loosen up the contributor process and our worry on trunk. Let 1000
> >>>>>>> flowers bloom and bring it in.
> >>>>>>> - And finally, to give some people more peace of mind and open more 
> >>>>>>> doors,
> >>>>>>> adopt what other projects have done and provide more pluggability. 
> >>>>>>> Let new
> >>>>>>> ideas have an easy place to connect.
> >>>>>>> We are at a fork in the road. What are we going to do? And then I 
> >>>>>>> have to
> >>>>>>> ask myself, what am I going to do as a contributor?
> >>>>>>> Patrick
> >>>>>>> On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston <[email protected]>
> >>>>>>> wrote:
> >>>>>>>> I’m not necessarily opposed to having a policy, but so far we have 
> >>>>>>>> some
> >>>>>>>> specific proposals addressing a problem statement that’s very 
> >>>>>>>> nebulous.
> >>>>>>>> What is the community failing to do on its own that we’re trying to 
> >>>>>>>> correct
> >>>>>>>> with policy? What outcomes are we trying to create or prevent? 
> >>>>>>>> Having some
> >>>>>>>> examples and specific problems to discuss would help focus the 
> >>>>>>>> conversation.
> >>>>>>>>> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote:
> >>>>>>>> Benedict,
> >>>>>>>> Thanks for clarifying. My concern still remains. This criteria would 
> >>>>>>>> be
> >>>>>>>> difficult to define and apply consistently. What counts as “similar” 
> >>>>>>>> scope
> >>>>>>>> or area, “mostly correct,” or sufficiently independent work? More
> >>>>>>>> importantly, how do we prevent such vague criteria from creating an
> >>>>>>>> informal hierarchy where some contributors work is routinely 
> >>>>>>>> accepted while
> >>>>>>>> others is routinely rejected?
> >>>>>>>> If the intent is to limit AI-assisted code changes to Cassandra
> >>>>>>>> contributors, or to contributors who have previously worked in that
> >>>>>>>> component without AI, that would at least be clear and enforceable.
> >>>>>>>>> On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith <
> >>>>>>>> [email protected]> wrote:
> >>>>>>>>> Core code changes
> >>>>>>>>> Chris: Do you object to the first or second line you quote? Because 
> >>>>>>>>> the
> >>>>>>>> first line is effectively motivation for the second line, and can be
> >>>>>>>> removed (or more clearly combined). If it’s the second line, then I 
> >>>>>>>> do not
> >>>>>>>> think this is an unreasonable expectation, and we can get into a 
> >>>>>>>> proper
> >>>>>>>> debate about it.
> >>>>>>>>> Shailaja, since you only snipped the first sentence, your concerns 
> >>>>>>>>> might
> >>>>>>>> also be mostly answered by this clarification? “Minimal third-party
> >>>>>>>> guidance” implies you have some concerns about the second line, but 
> >>>>>>>> all of
> >>>>>>>> our policies have some ambiguity because legalese is even worse. I 
> >>>>>>>> don’t
> >>>>>>>> think the ambiguity here would be challenging to navigate though we 
> >>>>>>>> can
> >>>>>>>> certainly refine it. This specific snippet is meant to convey an
> >>>>>>>> expectation that a contributor has autonomously produced patches of 
> >>>>>>>> similar
> >>>>>>>> scope that were mostly correct, so that they have demonstrated the 
> >>>>>>>> level of
> >>>>>>>> understanding necessary to guide another party to a successful patch 
> >>>>>>>> (i.e.
> >>>>>>>> an LLM in this case).
> >>>>>>>>> On 2026/09/23 10:54:16 Benedict Elliott Smith wrote:
> >>>>>>>>>> Thanks everyone for your input so far. I’ll respond in brief to the
> >>>>>>>> main themes, in (mostly) separate emails so they can each have their 
> >>>>>>>> own
> >>>>>>>> debate chain.
> >>>>>>>>>> Should we have a policy (Blake/Josh*/Jon/Dinesh)
> >>>>>>>>>> I think we would all agree that LLMs represent the biggest change 
> >>>>>>>>>> to
> >>>>>>>> this community (and software more generally) since its inception, 
> >>>>>>>> and we
> >>>>>>>> all now have enough experience with the technology to have formed 
> >>>>>>>> opinions
> >>>>>>>> about how it is best managed. We also evidently have not all arrived 
> >>>>>>>> at the
> >>>>>>>> same conclusions. In this situation, it would be an abdication of our
> >>>>>>>> responsibilities as a management committee to not agree *some* 
> >>>>>>>> policy.
> >>>>>>>>>> I intend to conduct straw polls as the discussion evolves, so if 
> >>>>>>>>>> you
> >>>>>>>> prefer an alternative policy - or modifications to this policy - I 
> >>>>>>>> would
> >>>>>>>> encourage you to make those alternative proposals.
> >>>>>>>>>> *Veto/Consensus (Josh)
> >>>>>>>>>> It was fair to call out my poor use of language on this topic, so 
> >>>>>>>>>> let
> >>>>>>>> me rephrase a little. The community is built on consensus, and work 
> >>>>>>>> should
> >>>>>>>> not be merged when there are outstanding concerns to address. The 
> >>>>>>>> explicit
> >>>>>>>> -1 should only be used rarely, because the prior expectation should 
> >>>>>>>> prevent
> >>>>>>>> it ever being needed. I (and others) have outstanding concerns on LLM
> >>>>>>>> generated work that can only be addressed through this process right 
> >>>>>>>> here,
> >>>>>>>> so to merge such work while maintaining the community’s consensus we 
> >>>>>>>> must
> >>>>>>>> agree some policy.
> >>>>>>>>>> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote:
> >>>>>>>>>>> I am strongly -1 on this
> >>>>>>>>>>> - Core code changes made by LLM may only be proposed by 
> >>>>>>>>>>> contributors
> >>>>>>>> with demonstrated expertise
> >>>>>>>>>>> That creates a new, subjective privileged class of contributors 
> >>>>>>>>>>> and
> >>>>>>>> turns a tool choice into an eligibility test. Who decides whether 
> >>>>>>>> expertise
> >>>>>>>> has been “demonstrated,” what counts as “minimal third-party 
> >>>>>>>> guidance,” and
> >>>>>>>> how could those judgments be applied consistently or fairly?
> >>>>>>>>>>> Apache already has a better model, anyone may contribute, trust 
> >>>>>>>>>>> and
> >>>>>>>> additional repository privileges are earned transparently over time. 
> >>>>>>>> The
> >>>>>>>> ASF describes its communities as flat, and says that newcomer ideas 
> >>>>>>>> have as
> >>>>>>>> much input as those from original creators. We should not add a 
> >>>>>>>> separate,
> >>>>>>>> informal hierarchy in which certain people may use common 
> >>>>>>>> development tools
> >>>>>>>> while others may not.
> >>>>>>>>>>>> On Sep 23, 2026, at 6:33 AM, Chris Lohfink <[email protected]>
> >>>>>>>> wrote:
> >>>>>>>>>>>> - Core code changes made by LLM may only be proposed by 
> >>>>>>>>>>>> contributors
> >>>>>>>> with demonstrated expertise
> >>>>>>>>>>>> - Must have produced similar patches in size, scope and area
> >>>>>>>> unassisted and with minimal third-party guidance
> >>>>>>>>>>>> I really don't like this one or its wording. Definitely too "the
> >>>>>>>> peasants are getting uppity lets build a wall". Lets not let a 
> >>>>>>>> subjective
> >>>>>>>> thing like demonstrated expertise (who decides that?) be if it's ok 
> >>>>>>>> or not.
> >>>>>>>> Hold the same standards for code quality and process for it all. I 
> >>>>>>>> don't
> >>>>>>>> want this to be: only people on the storage team in Apple can use AI.
> >>>>>>>>>>>> Chris
> >>>>>>>>>>>> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto:
> >>>>>>>> [email protected]>> wrote:
> >>>>>>>>>>>>> I agree with Stefan and think this is both a reasonable and
> >>>>>>>> thoughtful proposal.
> >>>>>>>>>>>>> Here are some things I like about it:
> >>>>>>>>>>>>> – It outlines areas where LLM usage is unambiguously useful to 
> >>>>>>>>>>>>> the
> >>>>>>>> project’s developers and users.
> >>>>>>>>>>>>> – It defines a spectrum of recommendations and cautions.
> >>>>>>>>>>>>> – The only prohibited areas are extremely narrow and say nothing
> >>>>>>>> about code at all.
> >>>>>>>>>>>>> Some in this thread are responding as if this proposal seeks to
> >>>>>>>> prohibit or sharply limit use of LLMs. In fact, it’s one of the most 
> >>>>>>>> open
> >>>>>>>> and welcoming I’ve seen for an OSS project of our size where many are
> >>>>>>>> adopting policies that simply ban them entirely. I’ve re-appended the
> >>>>>>>> proposal below my message as it seems to have been lost in threaded
> >>>>>>>> replies, and would encourage folks to give it a second read.
> >>>>>>>>>>>>> Some brief thoughts based on my own use of LLMs:
> >>>>>>>>>>>>> – I find them fantastically useful for reviewing and identifying
> >>>>>>>> problems that have slipped through review - primarily via Alex 
> >>>>>>>> Petrov’s
> >>>>>>>> /deep-review skill, which I have running in a VM in a loop executing 
> >>>>>>>> over
> >>>>>>>> every new commit in the project as of a few days ago. I will be 
> >>>>>>>> posting a
> >>>>>>>> few hand-authored Jira tickets based on findings that appear 
> >>>>>>>> legitimate to
> >>>>>>>> me. For now, the loop is posting them as issue drafts for my own 
> >>>>>>>> review on
> >>>>>>>> my personal fork which you can find here:
> >>>>>>>> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug
> >>>>>>>>>>>>> – They’re great for enabling use of model checkers and formal
> >>>>>>>> methods where such work would have previously been prohibitively 
> >>>>>>>> expensive,
> >>>>>>>> such as Blake’s work on a TLA+ proof of aspects of Mutation Tracking 
> >>>>>>>> and
> >>>>>>>> Benedict/Fedor’s work on a machine-checkable proof of the Accord 
> >>>>>>>> protocol
> >>>>>>>> in Lean.
> >>>>>>>>>>>>> – They are stunning for allowing me to experiment with ideas 
> >>>>>>>>>>>>> that
> >>>>>>>> would have otherwise been a summer internship’s scope of work. Some
> >>>>>>>> examples include an io_uring prototype, exploring the impact of
> >>>>>>>> page-aligned compressed chunk sizes, an API shim bridging the 3.x 
> >>>>>>>> and 4.x
> >>>>>>>> Java Drivers, and potential enhancements to Zstandard.
> >>>>>>>>>>>>> – And they shine when given grunt-work that is critical to the
> >>>>>>>> project but a miserable labor for humans, such as triaging, 
> >>>>>>>> reproducing,
> >>>>>>>> and root-causing flaky tests, which David Capwell now has running in 
> >>>>>>>> a loop
> >>>>>>>> to help us improve CI stability in the project.
> >>>>>>>>>>>>> I never thought I’d be so positive on what’s possible via 
> >>>>>>>>>>>>> language
> >>>>>>>> models a year ago. At the same time, I also agree that they present
> >>>>>>>> challenges and risks that can be managed through thoughtful 
> >>>>>>>> discussion and
> >>>>>>>> policy. Some of the concerns that I think are important to guard 
> >>>>>>>> against
> >>>>>>>> include:
> >>>>>>>>>>>>> – Asymmetry of effort between author and reviewers: As
> >>>>>>>> token-generating machines, LLMs can generate diffs of extraordinary 
> >>>>>>>> size
> >>>>>>>> very rapidly. /deep-review is great for chewing through diffs and
> >>>>>>>> identifying defects. But it should be used by the contributor 
> >>>>>>>> themselves to
> >>>>>>>> identify issues – not to replace the role of the reviewer with more
> >>>>>>>> electricity. The role of the reviewers extends beyond identifying and
> >>>>>>>> highlighting defects. It encompasses architecture, harmony with the
> >>>>>>>> existing codebase, thinking ahead to future evolution of the 
> >>>>>>>> project, and
> >>>>>>>> replicates context on the project as new code is committed. These 
> >>>>>>>> functions
> >>>>>>>> cannot be automated away.
> >>>>>>>>>>>>> – Hesitancy of authors to engage manually with code they have
> >>>>>>>> generated: This is not specific to Cassandra, but it is a behavior 
> >>>>>>>> that I
> >>>>>>>> have seen in several “highly-electric” projects. There’s a bimodal 
> >>>>>>>> tendency
> >>>>>>>> toward code that is entirely generated or entirely human-authored - 
> >>>>>>>> but it
> >>>>>>>> is rare for someone to prepare an AI-authored patch to take an 
> >>>>>>>> offramp and
> >>>>>>>> spend a significant amount of time refining the work by hand in an 
> >>>>>>>> IDE.
> >>>>>>>> This hesitancy toward human participation in authorship of 
> >>>>>>>> LLM-generated
> >>>>>>>> code is very concerning to me.
> >>>>>>>>>>>>> – Harmony with the existing codebase: Due to the tunnel-vision 
> >>>>>>>>>>>>> of
> >>>>>>>> context windows, LLMs are generally unaware of conventions and norms
> >>>>>>>> present in codebases and very frequently reinvent concepts in a 
> >>>>>>>> generation
> >>>>>>>> turn to suit a goal without view of the project’s overall 
> >>>>>>>> architecture.
> >>>>>>>> This results in a profusion of messy and duplicated concepts that 
> >>>>>>>> gradually
> >>>>>>>> sprawl about a codebase.
> >>>>>>>>>>>>> Again, none of these are grounds for prohibition of usage of
> >>>>>>>> language models in developing the project. They’re just problems we 
> >>>>>>>> need to
> >>>>>>>> bear in mind and guard against – and I think the proposal is 
> >>>>>>>> designed to do
> >>>>>>>> just that.
> >>>>>>>>>>>>> I’m thrilled by the potential of LLMs to improve Apache 
> >>>>>>>>>>>>> Cassandra
> >>>>>>>> and we already see it happening through a vast number of issues that 
> >>>>>>>> are
> >>>>>>>> being reported and fixed. But there’s also danger in taking ATVs 
> >>>>>>>> down a
> >>>>>>>> hiking trail full of people.
> >>>>>>>>>>>>> Regarding the prohibition on prose, I’ll simply say: I recently
> >>>>>>>> found myself in a scenario where I found a Claude-authored document 
> >>>>>>>> so
> >>>>>>>> inscrutable that I piped it back into a model, directed it to 
> >>>>>>>> rewrite it in
> >>>>>>>> ASD-STE100, read it myself, and responded based on the 
> >>>>>>>> summarization. As a
> >>>>>>>> humanities grad, this is probably the worst language crime I have
> >>>>>>>> committed. But it was in response to language that was itself so
> >>>>>>>> idiosyncratic that it was unreadable to me in its original form. I 
> >>>>>>>> hope
> >>>>>>>> this never happens in the Apache Cassandra project.
> >>>>>>>>>>>>> I’ll close with a quote from an excellent article written by 
> >>>>>>>>>>>>> Colin
> >>>>>>>> Breck, an engineer who works on large-scale data systems:
> >>>>>>>> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/
> >>>>>>>>>>>>> Colin wrote:
> >>>>>>>>>>>>>> I don’t want to live in a world where you use AI to summarize
> >>>>>>>> something important into unreadable text, and then I use AI in an 
> >>>>>>>> attempt
> >>>>>>>> to decipher it. I want to hear you, imperfections and all. I want 
> >>>>>>>> your
> >>>>>>>> interpretation of aesthetics, beauty, quality, relationship, time. I 
> >>>>>>>> want
> >>>>>>>> to know how you feel. I want you to cut through and tell me what 
> >>>>>>>> really
> >>>>>>>> matters.
> >>>>>>>>>>>>>> Intentional writing will likely become more valuable. People 
> >>>>>>>>>>>>>> who
> >>>>>>>> write, and write to think, to think deeply and carefully, or to 
> >>>>>>>> create, to
> >>>>>>>> share, or to capture something important without explicitly 
> >>>>>>>> expressing it
> >>>>>>>> will continue to write and produce original work. The people who 
> >>>>>>>> never were
> >>>>>>>> writers will use AI to produce lots of text.
> >>>>>>>>>>>>> I hope that our culture can remain one of intentional writing 
> >>>>>>>>>>>>> and
> >>>>>>>> intentional engineering. I enjoy reading the voice of the author in
> >>>>>>>> comments, code, and tickets in Cassandra – the different ways we use
> >>>>>>>> language based on where we grew up and how we learned English, the
> >>>>>>>> translated idioms from our various backgrounds, and terse comments 
> >>>>>>>> that
> >>>>>>>> recognize the difference between code whose function is obvious and 
> >>>>>>>> what
> >>>>>>>> warrants genuine exposition. When I read code in Cassandra, it’s a 
> >>>>>>>> delight
> >>>>>>>> to recognize the author based on their writing style before flipping 
> >>>>>>>> on
> >>>>>>>> `git annotate` to reveal the origin.
> >>>>>>>>>>>>> I’d encourage folks to re-read the original proposal below. It 
> >>>>>>>>>>>>> is
> >>>>>>>> very permissive. The guidance strikes me not just as reasonable, but
> >>>>>>>> genuinely important to maintaining the health of the project.
> >>>>>>>>>>>>> – Scott
> >>>>>>>>>>>>> =====
> >>>>>>>>>>>>> Encouraged:
> >>>>>>>>>>>>> - Reviewing and otherwise validating human-authored patches 
> >>>>>>>>>>>>> before
> >>>>>>>> submission
> >>>>>>>>>>>>> - Debugging, diagnosing etc
> >>>>>>>>>>>>> Permitted:
> >>>>>>>>>>>>> - Generating or modifying tests, scripts, tooling or any other
> >>>>>>>> non-user facing changes
> >>>>>>>>>>>>> - Minor changes to human-authored patches that are carefully
> >>>>>>>> reviewed by the author
> >>>>>>>>>>>>> Restricted:
> >>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by 
> >>>>>>>>>>>>> contributors
> >>>>>>>> with demonstrated expertise
> >>>>>>>>>>>>> - Must have produced similar patches in size, scope and area
> >>>>>>>> unassisted and with minimal third-party guidance
> >>>>>>>>>>>>> - Core code changes made by LLM require an additional reviewer
> >>>>>>>>>>>>> - LLM review is not a substitute for human review, and must be 
> >>>>>>>>>>>>> used
> >>>>>>>> only to augment a complete and independent human understanding of 
> >>>>>>>> the patch.
> >>>>>>>>>>>>> Prohibited:
> >>>>>>>>>>>>> - All public prose must be human authored. This includes inline
> >>>>>>>> comments, docs, posts to Jira etc.
> >>>>>>>>>>>>> All LLM generated changes MUST be disclosed:
> >>>>>>>>>>>>> - Outlined to any reviewer;
> >>>>>>>>>>>>> - Summarised in the commit message;
> >>>>>>>>>>>>> - Large blocks or files must be individually marked with some 
> >>>>>>>>>>>>> agreed
> >>>>>>>> message like "created by <some AI>"
> >>>>>>>>>>>>> =====
> >>>>>>>>>>>>>> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected]
> >>>>>>>> <mailto:[email protected]>> wrote:
> >>>>>>>>>>>>>> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected]
> >>>>>>>> <mailto:[email protected]>> wrote:
> >>>>>>>>>>>>>>> Restricted:
> >>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by
> >>>>>>>> contributors with demonstrated expertise
> >>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area
> >>>>>>>> unassisted and with minimal third-party guidance
> >>>>>>>>>>>>>> I am -1 on this. This sounds like gate keeping attempt. It 
> >>>>>>>>>>>>>> narrowly
> >>>>>>>> limits the pool to a few people on the project that have historically
> >>>>>>>> contributed to certain parts of the codebase. This policy will 
> >>>>>>>> prohibit
> >>>>>>>> skilled software engineers with domain expertise from proposing LLM
> >>>>>>>> assisted changes simply because they have not contributed to the 
> >>>>>>>> project.
> >>>>>>>> This is unrealistic and a net negative for the project to attract 
> >>>>>>>> talent
> >>>>>>>> and grow our community.
> >>>>>>>>>>>>>>> - Core code changes made by LLM require an additional reviewer
> >>>>>>>>>>>>>> Can you be more precise what is this in addition to? How many 
> >>>>>>>>>>>>>> total
> >>>>>>>> reviewers do you expect and what is the purpose of additional 
> >>>>>>>> reviewer? and
> >>>>>>>> why?
> >>>>>>>>>>>>>> Taking a step back - what are you trying to solve here?
> >>>>>>>>>>>>>> Dinesh

Reply via email to