To my eye, 2 and 3 are redundant. I’m all about KISS here. 

Patrick

> On Sep 24, 2026, at 5:40 PM, Caleb Rackliffe <[email protected]> wrote:
> Blake, I’d support those 3 guidelines.
> 
> One hostile reaction to them would, I’m guessing, be that guideline number 1 
> is too weak, and that the burden of quality would continue to rest entirely 
> on committer reviewers. I suppose my reply there would be that anyone who 
> continually spams the project with patches they do not fully understand wound 
> incur a pretty heavy reputational penalty and the problem would more or less 
> sort itself out.
> 
> 
>> On Sep 24, 2026, at 6:07 PM, Benedict Elliott Smith <[email protected]> 
>> wrote:
>> 
>> For the record Josh, I am including pre-3.0. I was involved in customer 
>> support escalations for 1.2, 2.0 and 2.1: the software was shoddy, to put it 
>> mildly. 8099 was a symptom, not the disease.
>> 
>> The project had a culture of doing stuff without sufficient care or 
>> consideration. This is by no means a simple matter of needing more testing 
>> or less timeline pressure, nor is it easy to define a "bar" to ensure it 
>> does not regress. If we had simple metrics, we would have settled it a long 
>> time ago.
>> 
>> You can still see customer scars crop up on forum discussions periodically.
>> 
>> 
>> On 2026/09/24 19:01:02 Josh McKenzie wrote:
>>>> Apache Cassandra was fundamentally undeployable for four years between Nov 
>>>> 2015 - 2019.
>>> CASSANDRA-8099 was a maximal manifestation of a specific approach to 
>>> engineering and calendar constraints we've seen time and again on the 
>>> project; I don't want us to conflate things here. That was a herculean 
>>> monolithic body of work performed in inhuman conditions (in vim!) that was 
>>> ultimately so invasive, all the unit tests in the code-base were commented 
>>> out and Jake and I spent a grueling 1.5-2 months hand-rewriting basically 
>>> all the unit tests in that code-base to get things to even build and run, 
>>> much less pass.
>>> Massive blast radius changes that are un-sustainably complex, under-tested, 
>>> where we don't property, fuzz, check coverage, check complexity, or A/B 
>>> compare against a known good system (i.e. pre/post) correctness testing are 
>>> going to destabilize the database at any time under any regime of tooling. 
>>> We certainly could speed-run our way back into destabilization with LLM's 
>>> as they are a force multiplier for both the good and the bad of one's 
>>> engineering practices, but that's a solvable problem by better defining 
>>> what our bars of quality are (Definition of Done anyone?) and holding 
>>> ourselves accountable to delivering at that bar.
>>> On Thu, Sep 24, 2026, at 2:40 PM, David Capwell via dev wrote:
>>>>> if you really want to pursue it I would ask that we do it offline to 
>>>>> avoid polluting an already busy conversation
>>>> People are directly responding saying that they feel discrimination 
>>>> currently and that the policy tries to codify that discrimination, so I 
>>>> feel its 100% on topic. This thread has presented 0 evidence that LLM 
>>>> usage has lowered the quality of contributions merged and has so far been 
>>>> vibes and feeling; I have yet to see any evidence to justify such 
>>>> discrimination so I will keep pushing back until such evidence is 
>>>> presented so we can have a informed debate.
>>>>> namely how we handle shallow and localised bug fixes. I would be happy 
>>>>> adding a clear entry to the “Permitted” section for this. No doubt there 
>>>>> are many refinements needed to the Restricted text as well, that might 
>>>>> also capture some of your concerns.
>>>> I will not sign off on cherry picking areas where "safe" to use a tool; so 
>>>> no it does not capture my concerns.
>>>>> I don't know if everyone remembers, but ten years ago Cassandra was full 
>>>>> of serious correctness and stability issues. Despite developing it, I 
>>>>> would not have run it myself or recommend that anyone use it. We have dug 
>>>>> ourselves out of that hole, but it took years of discipline and effort, 
>>>>> and we're still (deservedly) recovering our reputation.
>>>>> Let's use this new technology to improve the quality of our 
>>>>> contributions, not squander our hard-earned gains in the name of speed. 
>>>>> It will be hard to recover our reputation a second time.
>>>> I do recall the 3.x line and put in a significant amount of effort to 
>>>> harden it. There were behaviors I noticed after joining Cassandra that I 
>>>> feel directly contributed to 3.0 and the decade of catch up; behaviors 
>>>> that still linger in parts of the community today.
>>>> As I look on trunk and look at committed code and trace back to PRs and 
>>>> JIRA I see the following:
>>>> • large patches approved without comments
>>>> • 0 evidence that tests were run
>>>> I then look at our CI and see tests failing for months. As you start to 
>>>> triage you start to see some of them show real issues; yet they linger for 
>>>> months not being addressed... when CI is unstable it takes a lot of effort 
>>>> to triage "did my patch break the test", and I have seen time and time 
>>>> again people do not put in that effort, and shrug off as "its just a 
>>>> flakey test"; then our CI failure rate grows.
>>>> Non of this has anything to do with LLMs but LLMs running in this 
>>>> environment is far more dangerous as there are not checks in place to 
>>>> "hold the bar". I am all for raising the bar universally; expecting both 
>>>> humans and LLMs to match that bar.
>>>>> `I’d support something that boils down to roughly this:
>>>>> 1.) 2 committers must understand an LLM-assisted change before it 
>>>>> commits. (Perhaps separately we can explore the question of why we 
>>>>> haven’t added any new committers to the core project for about a year. 
>>>>> I’m also still not entirely sure if it’s acceptable within our guidelines 
>>>>> for a committer to +1 a patch after delegating review.)
>>>>> 2.) Patch authors must demonstrate enough understanding to discuss their 
>>>>> own patch, whether or not parts of it are generated by an LLM.
>>>>> 3.) The “Assisted-by” tag should be used to indicate any non-trivial LLM 
>>>>> usage in the generation of a patch, just like we have used Co-authored-by 
>>>>> historically.
>>>>> 4.) Comments and other things that aren't the actual code (but could sow 
>>>>> confusion) should be held to the same standard we'd expect from a human 
>>>>> writer. If we don’t yet agree on that standard, we can formalize enough 
>>>>> of it to guide both humans and LLMs.
>>> `
>>>> I can get behind this proposal but i would tweak it as 1/2 i don't think 
>>>> really need to special case LLM usage
>>>> 1. 2 committers must understand the change before it commits.
>>>> 2. Patch authors must demonstrate enough understanding to discuss their 
>>>> own patch
>>>> Nothing about those 2 need to be scoped to LLM usage and honestly matches 
>>>> most PMCs I have talked to understanding of our bar (as Benedict pointed 
>>>> out, the actual wording could be interpreted to allow rubber stamping from 
>>>> committer)
>>>> As for 3 I am cool with this. ASF recommends the same (it says 
>>>> `Generated-by` but that discount's the human's effort) as its useful for 
>>>> audits and tooling. Having `Assisted-by` tag should not imply anything 
>>>> about the committed patch as it should have gone through the same bar we 
>>>> all expect; its just for tools auditing.
>>>>> On Sep 24, 2026, at 10:37 AM, Aleksey Yeshchenko via dev 
>>>>> <[email protected]> wrote:
>>>>> Meant "doing away with", sorry. Non-native speaker with a headache here. 
>>>>> Thanks Caleb for spotting.
>>>>>> On 24 Sep 2026, at 17:35, Aleksey Yeshchenko via dev 
>>>>>> <[email protected]> wrote:
>>>>>> P.S. I assume it's obvious from the text above that I don't believe that 
>>>>>> getting away with human code review is a viable option.
>>>>>> On 24 Sep 2026, at 17:37, Štefan Miklošovič <[email protected]> 
>>>>>> wrote:
>>>>>> Good call on checkerframework, we even have a patch for it. Work of
>>>>>> Jacek Lewandowski. We might just drive it to completion. Using AI for
>>>>>> finishing it would be quite ironic.
>>>>>> (1) https://github.com/apache/cassandra/pull/2370
>>>>>> On Thu, Sep 24, 2026 at 6:19 PM Jon Haddad <[email protected]> 
>>>>>> wrote:
>>>>>>> There are some really good points being brought up about stability of 
>>>>>>> the codebase, maintainability, quality of reviews, correctness bugs, 
>>>>>>> and I agree with all of them.  I think it would be helpful to take a 
>>>>>>> step back and consider how those bugs got there in the first place, how 
>>>>>>> they were fixed, and what we could do to further advance the codebase 
>>>>>>> so they don't creep back.  LLMs can be used either with very tight 
>>>>>>> guardrails, or in what's effectively YOLO mode, and there's a big 
>>>>>>> difference in the quality of the results you get.
>>>>>>> One thing to keep in mind, a lot of the initial code in C* was added 
>>>>>>> without comprehensive testing.  I hope we can all agree that it's a lot 
>>>>>>> easier to break code that doesn't have high quality tests.  During the 
>>>>>>> code freeze, a lot of people people spent several years relentlessly 
>>>>>>> finding and fixing bugs. This was probably a pretty frustrating time 
>>>>>>> for anyone who was focused on fixing other people's bugs when they 
>>>>>>> wanted to build features.  I think we should recognize the effort here 
>>>>>>> and appreciate the foundation that the project stands on now. I can 
>>>>>>> understand how anyone involved with this effort would be apprehensive 
>>>>>>> about seeing years of their life swept away by an agent that was driven 
>>>>>>> by goal seeking to remove all the tests that it broke instead of fixing 
>>>>>>> them.
>>>>>>> When I picked up the work to improve cursor compaction, the first thing 
>>>>>>> I asked myself was how can I make sure I don't break this?  How do I 
>>>>>>> even know it works properly?  There were some tricky parts to the code, 
>>>>>>> and I really didn't want to come in and immediately break stuff.  
>>>>>>> That's why I started with an entire patch dedicated to adding test 
>>>>>>> infra to it.  90% of the patch was tests, and in my other cursor 
>>>>>>> patches, it remains *at least* 80% of my patches.  It was a *lot* 
>>>>>>> faster to add almost 10K lines of tests that handled a byte for byte 
>>>>>>> differential testing paired with harry to find over 30 bugs that caused 
>>>>>>> cursor to corrupt results.  Range tombstones alone were at least a 
>>>>>>> dozen bugs, but I also found issues with static columns, reverse 
>>>>>>> ordering, etc.  Randomizing schemas and data in burn tests to generate 
>>>>>>> different shapes of data, to ensure they all result in the same output 
>>>>>>> at the end. JMH tests to ensure there weren't performance regressions, 
>>>>>>> hours of profiling. These were all a *lot* easier to do with the LLM 
>>>>>>> helping me out.  In the process I've found bugs that have been 
>>>>>>> lingering in the codebase for years.
>>>>>>> That's a long story, but hopefully we all agree that having 
>>>>>>> comprehensive tests is a great way to ensure that both humans and LLMs 
>>>>>>> don't break things that are working.
>>>>>>> The lesson: we need to keep improving our testing.  Everything that we 
>>>>>>> touch, should be left in a better state than how we found it with 
>>>>>>> regard to test coverage.
>>>>>>> Test coverage isn't everything though, there's always little subtle 
>>>>>>> bugs that don't get found in testing, that can slip in despite our best 
>>>>>>> efforts.  It's debatable if humans will be as good as agents for coding 
>>>>>>> in the long term, for spotting small defects.  I sincerely doubt it.  
>>>>>>> For the time being though, we still have people involved. It's probably 
>>>>>>> a good time to start using more static analysis tools to identify 
>>>>>>> problematic code and to add this to CI.  Dmitry had a suggestion 
>>>>>>> recently for checkerframework to detect leaking contexts, a problem he 
>>>>>>> spotted when reviewing my branch.  It would be great to have that 
>>>>>>> integrated into our CI and dev workflow so we can simply avoid an 
>>>>>>> entire class of bugs.
>>>>>>> There's also PMD, which is excellent for finding code that can be hard 
>>>>>>> to understand.  I *highly* suggest you all run PMD to analyze for 
>>>>>>> cognitive complexity and high npath scores.  This was made popular by 
>>>>>>> the folks at Sonar and I've found it to be an excellent feedback 
>>>>>>> mechanism for structuring code.  The default max they set is 15, which 
>>>>>>> is the point where it starts to become difficult to verify something 
>>>>>>> works without making a massive investment.  We've got areas in the 
>>>>>>> codebase that are in the hundreds, and some parts even higher.  These 
>>>>>>> have been contributed by humans, and are all high risk points for both 
>>>>>>> humans and agents to start messing around with.  They're also in some 
>>>>>>> fairly critical areas that are very likely to break, so I understand 
>>>>>>> why people would not want an agent anywhere near it.
>>>>>>> Unfortunately, it's not an easy problem to address.  There's so many 
>>>>>>> places where the code is structured in a way that has so many branches, 
>>>>>>> so many conditions, that it's effectively impossible for a human to 
>>>>>>> understand, creating a fear of messing around in it.  There's plenty of 
>>>>>>> areas that deserve extreme scrutiny, and we should be careful of what 
>>>>>>> we add, whether it's human or agent.
>>>>>>> The codebase today requires a high degree of internal knowledge to 
>>>>>>> navigate.  There's land mines everywhere. We should be looking to make 
>>>>>>> conscious improvements by moving the code forward, so it's easier to 
>>>>>>> make changes to small, well tested components with minimal side 
>>>>>>> effects.  Not making it harder for people to use the tools that aid in 
>>>>>>> that process.
>>>>>>> Here's what we could do to achieve the underlying goal of not breaking 
>>>>>>> the DB:
>>>>>>> Add cognitive complexlity and npath via PMD as a feedback mechanism.
>>>>>>> Code that's hard to understand is hard to review.  It's also hard to 
>>>>>>> test. Let's break down the complex code so more people can contribute, 
>>>>>>> safely.
>>>>>>> Add checkerframework to our tooling,
>>>>>>> Properly annotate the codebase for it and reduce the surface area that 
>>>>>>> things can break.  Less brittle codebase = we can move faster.
>>>>>>> Use jacoco to find areas of the codebase with poor testing.
>>>>>>> Let's improve the test coverage there, LLMs are great for this.  We 
>>>>>>> have a ton of static tests, these can become more dynamic, 
>>>>>>> parameterized, and leverage harry.
>>>>>>> Refactor parts of the codebase that have high cognitive complexlity and 
>>>>>>> NPath scores.
>>>>>>> This should be lowered over time to meet some high watermark, say 25 
>>>>>>> maximum, although I'd prefer 15 which is where the Sonar folks settled.
>>>>>>> Move forward moving the codebase to a more modular structure
>>>>>>> We've talked about Gradle on and off - but it can really be a huge help 
>>>>>>> with incremental, modular builds. This is pretty easy to do with an 
>>>>>>> agent and we could have it done in a couple days.
>>>>>>> Enforce boundaries with ArchUnit
>>>>>>> If we want to enforce certain code boundaries, this is the way to do 
>>>>>>> it. Should not be part of manual review.
>>>>>>> Add LLM review for all incoming PRs before a human
>>>>>>> The goal here is to automate the initial part of the review process 
>>>>>>> that reviewers should spot, and raise the bar for the initial 
>>>>>>> contribution.  When the code gets reviewed by a human, it should 
>>>>>>> already have passed a large variety of initial checks.  This should 
>>>>>>> shorten the review cycle and result in higher quality patches.  I've 
>>>>>>> had Claude reviewing all my PRs in my personal projects for a while now 
>>>>>>> and it consistently gives great feedback that I almost always 
>>>>>>> incorporate.
>>>>>>> In my ideal world, we'd also auto-format all code
>>>>>>> Consistent formatting throughout the codebase would be amazing, but 
>>>>>>> that's just one man's dream.
>>>>>>> Hopefully there's at least a couple things in this list we could move 
>>>>>>> forward with in the short term, as it'll help improve the code quality 
>>>>>>> regardless of how it's created.
>>>>>>> Jon
>>>>>>> https://checkerframework.org/manual/#aliasing-leaking-contexts
>>>>>>> https://www.sonarsource.com/docs/CognitiveComplexity.pdf
>>>>>>> https://pmd.github.io/pmd/pmd_rules_java_design.html
>>>>>>> On Thu, Sep 24, 2026 at 7:38 AM C. Scott Andreas <[email protected]> 
>>>>>>> wrote:
>>>>>>>> From Benedict:
>>>>>>>> “I don't know if everyone remembers, but ten years ago Cassandra was 
>>>>>>>> full of serious correctness and stability issues. Despite developing 
>>>>>>>> it, I would not have run it myself or recommend that anyone use it. We 
>>>>>>>> have dug ourselves out of that hole, but it took years of discipline 
>>>>>>>> and effort, and we're still (deservedly) recovering our reputation.”
>>>>>>>> Expanding on this point for those who may not have been active in the 
>>>>>>>> project at this time —
>>>>>>>> Apache Cassandra was fundamentally undeployable for four years between 
>>>>>>>> Nov 2015 - 2019. The database literally lost data if you ran a 
>>>>>>>> read-only SELECT query ordered descending (C-14513, C-14515). If you 
>>>>>>>> haven’t read these tickets before, please take a moment to do so.
>>>>>>>> It took years of careful work via property-based testing, fuzzing, and 
>>>>>>>> deterministic simulation to restore Cassandra’s status as a usable 
>>>>>>>> system of record. Once 14513 and 14515 were identified, nearly 30 
>>>>>>>> additional critical data loss and incorrect response bugs were 
>>>>>>>> identified.
>>>>>>>> It is essential for the project’s future that we don’t regress to this 
>>>>>>>> state chasing AI-generated features motivated by fear. The fact that 
>>>>>>>> examples cited in this thread which boast shiny features but have 
>>>>>>>> critical shortcomings unknown to their author supports this argument.
>>>>>>>> The most common path for large corpuses of AI-generated software is 
>>>>>>>> elation and reveling in a feature matrix, followed by abandonment.
>>>>>>>> I endorse this point:
>>>>>>>> “Let's use this new technology to improve the quality of our 
>>>>>>>> contributions, not squander our hard-earned gains in the name of 
>>>>>>>> speed. It will be hard to recover our reputation a second time.”
>>>>>>>> Patrick, I don’t want your note regarding a TCM issue to go 
>>>>>>>> unaddressed. Please file a Jira ticket and the patch if you like. I 
>>>>>>>> can’t comment on the patch as I haven’t seen it, but together we will 
>>>>>>>> solve the problem.
>>>>>>>> – Scott
>>>>>>>>> On Sep 24, 2026, at 4:01 AM, Benedict Elliott Smith 
>>>>>>>>> <[email protected]> wrote:
>>>>>>>>> Hi Patrick,
>>>>>>>>> As I mentioned in my reply to David, I would be happy to create a 
>>>>>>>>> carve out for shallow and localised bug fixes in the "Permitted" 
>>>>>>>>> section. Would this alleviate some of your concerns regarding your 
>>>>>>>>> ability to contribute to the project?
>>>>>>>>> I appreciate your pointing out Ferrosa's Accord implementation 
>>>>>>>>> however, as it is a *great* example of the problems we're leaping 
>>>>>>>>> into. I took a look, and within about 30s found that the protocol is 
>>>>>>>>> fundamentally incorrect, having failed to address CASSANDRA-18365. 
>>>>>>>>> This is despite claiming to be tested with Jepsen that should in 
>>>>>>>>> principle find this fault. I followed up by using Claude to 
>>>>>>>>> interrogate the implementation further, and immediately found other 
>>>>>>>>> serious correctness issues.
>>>>>>>>> I use LLMs daily now to help facilitate Accord development, and while 
>>>>>>>>> they are powerful they are NOT able to author the code themselves, 
>>>>>>>>> even when building upon a strong human-authored foundation.
>>>>>>>>> I don't know if everyone remembers, but ten years ago Cassandra was 
>>>>>>>>> full of serious correctness and stability issues. Despite developing 
>>>>>>>>> it, I would not have run it myself or recommend that anyone use it. 
>>>>>>>>> We have dug ourselves out of that hole, but it took years of 
>>>>>>>>> discipline and effort, and we're still (deservedly) recovering our 
>>>>>>>>> reputation.
>>>>>>>>> Let's use this new technology to improve the quality of our 
>>>>>>>>> contributions, not squander our hard-earned gains in the name of 
>>>>>>>>> speed. It will be hard to recover our reputation a second time.
>>>>>>>>>> On 2026/09/23 19:16:17 Patrick McFadin wrote:
>>>>>>>>>> I was waiting for this moment to hit our project and I'm glad we're 
>>>>>>>>>> here. I
>>>>>>>>>> am deeply concerned for our project and its future, as we have 
>>>>>>>>>> increasingly
>>>>>>>>>> made it difficult to contribute. I had hoped that this new era of
>>>>>>>>>> software tools powered by AI would expand the project's reach and 
>>>>>>>>>> bring
>>>>>>>>>> more diverse thoughts and ideas. This policy proposal is the exact 
>>>>>>>>>> opposite
>>>>>>>>>> of what we need. We have been sitting on a Cassandra 6 release alpha 
>>>>>>>>>> for
>>>>>>>>>> months. We need to accelerate and embrace new ways of being or be 
>>>>>>>>>> left
>>>>>>>>>> behind. As I read that policy, my first and gut level reactions:
>>>>>>>>>> - It comes across as elitist and class protectionism. Committer 
>>>>>>>>>> should not
>>>>>>>>>> be special but this proposal makes that designation even more sacred.
>>>>>>>>>> - It signals that our project is so fragile that only a few people 
>>>>>>>>>> "Really
>>>>>>>>>> understand it" That's some SQLite vibes right there.
>>>>>>>>>> - Trying to fix a problem that doesn't exist
>>>>>>>>>> Sadly, i think this policy change would also exclude a lot of 
>>>>>>>>>> comitters.
>>>>>>>>>> We aren't alone in this moment. The Linux project just went through
>>>>>>>>>> this. You can find the thread with a simple Google, but similar hard
>>>>>>>>>> feelings were being expressed "AI is going to ruin our project!", 
>>>>>>>>>> "The
>>>>>>>>>> unwashed masses are going to contribute terrible code!", "We have to
>>>>>>>>>> protect our precious status as Linux maintainers!"  Linus being 
>>>>>>>>>> Linus, was
>>>>>>>>>> deeply invloved and they adopted a super simple statement that 
>>>>>>>>>> covers all
>>>>>>>>>> bases. Human or Human using AI. “You are expected to understand and 
>>>>>>>>>> to be
>>>>>>>>>> able to defend everything you submit.”  Love that.
>>>>>>>>>> In the larger picture, I'll restate. I'm worried for our project. In 
>>>>>>>>>> late
>>>>>>>>>> 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real 
>>>>>>>>>> corner and
>>>>>>>>>> in the hands of somebody that knows how to build software, this tool 
>>>>>>>>>> is
>>>>>>>>>> like jet fuel. Here's some examples of new projects being hyper 
>>>>>>>>>> fueled by
>>>>>>>>>> AI coding tools.
>>>>>>>>>> Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity
>>>>>>>>>> Turso - Rust re-write of SQLite
>>>>>>>>>> Bun - Rust re-write of itself from Zig.
>>>>>>>>>> Think this couldn't happen to us? Already has:
>>>>>>>>>> https://github.com/ferrosadb/ferrosa. Ben is using it to power his 
>>>>>>>>>> own
>>>>>>>>>> startup, but it was him alone using a ton of local AI coding agents. 
>>>>>>>>>> He
>>>>>>>>>> even implemented Accord. Yeah...
>>>>>>>>>> The cracks are already starting to show. There is a black market 
>>>>>>>>>> economy of
>>>>>>>>>> Cassandra patches happening now. Not going to name names or call 
>>>>>>>>>> people
>>>>>>>>>> out,  but there are fixes and optimizations living in branches 
>>>>>>>>>> outside of
>>>>>>>>>> the Cassandra project. Why? I'll use myself as an example. I fixed a 
>>>>>>>>>> nasty
>>>>>>>>>> bug I ran into with TCM a few weeks ago. Wrote the tests. It passes 
>>>>>>>>>> CI and
>>>>>>>>>> lives in my personal branch. I'm sitting here really wondering if I 
>>>>>>>>>> want to
>>>>>>>>>> go through the ritual humiliation of being roasted for using AI to 
>>>>>>>>>> fix it.
>>>>>>>>>> Me. I am worried about contrinuting code the Cassandra. What the 
>>>>>>>>>> hell does
>>>>>>>>>> that say?
>>>>>>>>>> I have my CQLite project that I've been doing a release around once a
>>>>>>>>>> month. I would love to donate that to the Cassandra project but I 
>>>>>>>>>> wouldn't
>>>>>>>>>> if it essentially killed any progress.
>>>>>>>>>> My larger counter proposal would be to:
>>>>>>>>>> - Adopt the “You are expected to understand and to be able to defend
>>>>>>>>>> everything you submit.” approach the Linux project has adopted.
>>>>>>>>>> - Loosen up the contributor process and our worry on trunk. Let 1000
>>>>>>>>>> flowers bloom and bring it in.
>>>>>>>>>> - And finally, to give some people more peace of mind and open more 
>>>>>>>>>> doors,
>>>>>>>>>> adopt what other projects have done and provide more pluggability. 
>>>>>>>>>> Let new
>>>>>>>>>> ideas have an easy place to connect.
>>>>>>>>>> We are at a fork in the road. What are we going to do? And then I 
>>>>>>>>>> have to
>>>>>>>>>> ask myself, what am I going to do as a contributor?
>>>>>>>>>> Patrick
>>>>>>>>>> On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston 
>>>>>>>>>> <[email protected]>
>>>>>>>>>> wrote:
>>>>>>>>>>> I’m not necessarily opposed to having a policy, but so far we have 
>>>>>>>>>>> some
>>>>>>>>>>> specific proposals addressing a problem statement that’s very 
>>>>>>>>>>> nebulous.
>>>>>>>>>>> What is the community failing to do on its own that we’re trying to 
>>>>>>>>>>> correct
>>>>>>>>>>> with policy? What outcomes are we trying to create or prevent? 
>>>>>>>>>>> Having some
>>>>>>>>>>> examples and specific problems to discuss would help focus the 
>>>>>>>>>>> conversation.
>>>>>>>>>>>> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote:
>>>>>>>>>>> Benedict,
>>>>>>>>>>> Thanks for clarifying. My concern still remains. This criteria 
>>>>>>>>>>> would be
>>>>>>>>>>> difficult to define and apply consistently. What counts as 
>>>>>>>>>>> “similar” scope
>>>>>>>>>>> or area, “mostly correct,” or sufficiently independent work? More
>>>>>>>>>>> importantly, how do we prevent such vague criteria from creating an
>>>>>>>>>>> informal hierarchy where some contributors work is routinely 
>>>>>>>>>>> accepted while
>>>>>>>>>>> others is routinely rejected?
>>>>>>>>>>> If the intent is to limit AI-assisted code changes to Cassandra
>>>>>>>>>>> contributors, or to contributors who have previously worked in that
>>>>>>>>>>> component without AI, that would at least be clear and enforceable.
>>>>>>>>>>>> On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith <
>>>>>>>>>>> [email protected]> wrote:
>>>>>>>>>>>> Core code changes
>>>>>>>>>>>> Chris: Do you object to the first or second line you quote? 
>>>>>>>>>>>> Because the
>>>>>>>>>>> first line is effectively motivation for the second line, and can be
>>>>>>>>>>> removed (or more clearly combined). If it’s the second line, then I 
>>>>>>>>>>> do not
>>>>>>>>>>> think this is an unreasonable expectation, and we can get into a 
>>>>>>>>>>> proper
>>>>>>>>>>> debate about it.
>>>>>>>>>>>> Shailaja, since you only snipped the first sentence, your concerns 
>>>>>>>>>>>> might
>>>>>>>>>>> also be mostly answered by this clarification? “Minimal third-party
>>>>>>>>>>> guidance” implies you have some concerns about the second line, but 
>>>>>>>>>>> all of
>>>>>>>>>>> our policies have some ambiguity because legalese is even worse. I 
>>>>>>>>>>> don’t
>>>>>>>>>>> think the ambiguity here would be challenging to navigate though we 
>>>>>>>>>>> can
>>>>>>>>>>> certainly refine it. This specific snippet is meant to convey an
>>>>>>>>>>> expectation that a contributor has autonomously produced patches of 
>>>>>>>>>>> similar
>>>>>>>>>>> scope that were mostly correct, so that they have demonstrated the 
>>>>>>>>>>> level of
>>>>>>>>>>> understanding necessary to guide another party to a successful 
>>>>>>>>>>> patch (i.e.
>>>>>>>>>>> an LLM in this case).
>>>>>>>>>>>> On 2026/09/23 10:54:16 Benedict Elliott Smith wrote:
>>>>>>>>>>>>> Thanks everyone for your input so far. I’ll respond in brief to 
>>>>>>>>>>>>> the
>>>>>>>>>>> main themes, in (mostly) separate emails so they can each have 
>>>>>>>>>>> their own
>>>>>>>>>>> debate chain.
>>>>>>>>>>>>> Should we have a policy (Blake/Josh*/Jon/Dinesh)
>>>>>>>>>>>>> I think we would all agree that LLMs represent the biggest change 
>>>>>>>>>>>>> to
>>>>>>>>>>> this community (and software more generally) since its inception, 
>>>>>>>>>>> and we
>>>>>>>>>>> all now have enough experience with the technology to have formed 
>>>>>>>>>>> opinions
>>>>>>>>>>> about how it is best managed. We also evidently have not all 
>>>>>>>>>>> arrived at the
>>>>>>>>>>> same conclusions. In this situation, it would be an abdication of 
>>>>>>>>>>> our
>>>>>>>>>>> responsibilities as a management committee to not agree *some* 
>>>>>>>>>>> policy.
>>>>>>>>>>>>> I intend to conduct straw polls as the discussion evolves, so if 
>>>>>>>>>>>>> you
>>>>>>>>>>> prefer an alternative policy - or modifications to this policy - I 
>>>>>>>>>>> would
>>>>>>>>>>> encourage you to make those alternative proposals.
>>>>>>>>>>>>> *Veto/Consensus (Josh)
>>>>>>>>>>>>> It was fair to call out my poor use of language on this topic, so 
>>>>>>>>>>>>> let
>>>>>>>>>>> me rephrase a little. The community is built on consensus, and work 
>>>>>>>>>>> should
>>>>>>>>>>> not be merged when there are outstanding concerns to address. The 
>>>>>>>>>>> explicit
>>>>>>>>>>> -1 should only be used rarely, because the prior expectation should 
>>>>>>>>>>> prevent
>>>>>>>>>>> it ever being needed. I (and others) have outstanding concerns on 
>>>>>>>>>>> LLM
>>>>>>>>>>> generated work that can only be addressed through this process 
>>>>>>>>>>> right here,
>>>>>>>>>>> so to merge such work while maintaining the community’s consensus 
>>>>>>>>>>> we must
>>>>>>>>>>> agree some policy.
>>>>>>>>>>>>> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote:
>>>>>>>>>>>>>> I am strongly -1 on this
>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by 
>>>>>>>>>>>>>> contributors
>>>>>>>>>>> with demonstrated expertise
>>>>>>>>>>>>>> That creates a new, subjective privileged class of contributors 
>>>>>>>>>>>>>> and
>>>>>>>>>>> turns a tool choice into an eligibility test. Who decides whether 
>>>>>>>>>>> expertise
>>>>>>>>>>> has been “demonstrated,” what counts as “minimal third-party 
>>>>>>>>>>> guidance,” and
>>>>>>>>>>> how could those judgments be applied consistently or fairly?
>>>>>>>>>>>>>> Apache already has a better model, anyone may contribute, trust 
>>>>>>>>>>>>>> and
>>>>>>>>>>> additional repository privileges are earned transparently over 
>>>>>>>>>>> time. The
>>>>>>>>>>> ASF describes its communities as flat, and says that newcomer ideas 
>>>>>>>>>>> have as
>>>>>>>>>>> much input as those from original creators. We should not add a 
>>>>>>>>>>> separate,
>>>>>>>>>>> informal hierarchy in which certain people may use common 
>>>>>>>>>>> development tools
>>>>>>>>>>> while others may not.
>>>>>>>>>>>>>>> On Sep 23, 2026, at 6:33 AM, Chris Lohfink 
>>>>>>>>>>>>>>> <[email protected]>
>>>>>>>>>>> wrote:
>>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by 
>>>>>>>>>>>>>>> contributors
>>>>>>>>>>> with demonstrated expertise
>>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area
>>>>>>>>>>> unassisted and with minimal third-party guidance
>>>>>>>>>>>>>>> I really don't like this one or its wording. Definitely too "the
>>>>>>>>>>> peasants are getting uppity lets build a wall". Lets not let a 
>>>>>>>>>>> subjective
>>>>>>>>>>> thing like demonstrated expertise (who decides that?) be if it's ok 
>>>>>>>>>>> or not.
>>>>>>>>>>> Hold the same standards for code quality and process for it all. I 
>>>>>>>>>>> don't
>>>>>>>>>>> want this to be: only people on the storage team in Apple can use 
>>>>>>>>>>> AI.
>>>>>>>>>>>>>>> Chris
>>>>>>>>>>>>>>> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto:
>>>>>>>>>>> [email protected]>> wrote:
>>>>>>>>>>>>>>>> I agree with Stefan and think this is both a reasonable and
>>>>>>>>>>> thoughtful proposal.
>>>>>>>>>>>>>>>> Here are some things I like about it:
>>>>>>>>>>>>>>>> – It outlines areas where LLM usage is unambiguously useful to 
>>>>>>>>>>>>>>>> the
>>>>>>>>>>> project’s developers and users.
>>>>>>>>>>>>>>>> – It defines a spectrum of recommendations and cautions.
>>>>>>>>>>>>>>>> – The only prohibited areas are extremely narrow and say 
>>>>>>>>>>>>>>>> nothing
>>>>>>>>>>> about code at all.
>>>>>>>>>>>>>>>> Some in this thread are responding as if this proposal seeks to
>>>>>>>>>>> prohibit or sharply limit use of LLMs. In fact, it’s one of the 
>>>>>>>>>>> most open
>>>>>>>>>>> and welcoming I’ve seen for an OSS project of our size where many 
>>>>>>>>>>> are
>>>>>>>>>>> adopting policies that simply ban them entirely. I’ve re-appended 
>>>>>>>>>>> the
>>>>>>>>>>> proposal below my message as it seems to have been lost in threaded
>>>>>>>>>>> replies, and would encourage folks to give it a second read.
>>>>>>>>>>>>>>>> Some brief thoughts based on my own use of LLMs:
>>>>>>>>>>>>>>>> – I find them fantastically useful for reviewing and 
>>>>>>>>>>>>>>>> identifying
>>>>>>>>>>> problems that have slipped through review - primarily via Alex 
>>>>>>>>>>> Petrov’s
>>>>>>>>>>> /deep-review skill, which I have running in a VM in a loop 
>>>>>>>>>>> executing over
>>>>>>>>>>> every new commit in the project as of a few days ago. I will be 
>>>>>>>>>>> posting a
>>>>>>>>>>> few hand-authored Jira tickets based on findings that appear 
>>>>>>>>>>> legitimate to
>>>>>>>>>>> me. For now, the loop is posting them as issue drafts for my own 
>>>>>>>>>>> review on
>>>>>>>>>>> my personal fork which you can find here:
>>>>>>>>>>> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug
>>>>>>>>>>>>>>>> – They’re great for enabling use of model checkers and formal
>>>>>>>>>>> methods where such work would have previously been prohibitively 
>>>>>>>>>>> expensive,
>>>>>>>>>>> such as Blake’s work on a TLA+ proof of aspects of Mutation 
>>>>>>>>>>> Tracking and
>>>>>>>>>>> Benedict/Fedor’s work on a machine-checkable proof of the Accord 
>>>>>>>>>>> protocol
>>>>>>>>>>> in Lean.
>>>>>>>>>>>>>>>> – They are stunning for allowing me to experiment with ideas 
>>>>>>>>>>>>>>>> that
>>>>>>>>>>> would have otherwise been a summer internship’s scope of work. Some
>>>>>>>>>>> examples include an io_uring prototype, exploring the impact of
>>>>>>>>>>> page-aligned compressed chunk sizes, an API shim bridging the 3.x 
>>>>>>>>>>> and 4.x
>>>>>>>>>>> Java Drivers, and potential enhancements to Zstandard.
>>>>>>>>>>>>>>>> – And they shine when given grunt-work that is critical to the
>>>>>>>>>>> project but a miserable labor for humans, such as triaging, 
>>>>>>>>>>> reproducing,
>>>>>>>>>>> and root-causing flaky tests, which David Capwell now has running 
>>>>>>>>>>> in a loop
>>>>>>>>>>> to help us improve CI stability in the project.
>>>>>>>>>>>>>>>> I never thought I’d be so positive on what’s possible via 
>>>>>>>>>>>>>>>> language
>>>>>>>>>>> models a year ago. At the same time, I also agree that they present
>>>>>>>>>>> challenges and risks that can be managed through thoughtful 
>>>>>>>>>>> discussion and
>>>>>>>>>>> policy. Some of the concerns that I think are important to guard 
>>>>>>>>>>> against
>>>>>>>>>>> include:
>>>>>>>>>>>>>>>> – Asymmetry of effort between author and reviewers: As
>>>>>>>>>>> token-generating machines, LLMs can generate diffs of extraordinary 
>>>>>>>>>>> size
>>>>>>>>>>> very rapidly. /deep-review is great for chewing through diffs and
>>>>>>>>>>> identifying defects. But it should be used by the contributor 
>>>>>>>>>>> themselves to
>>>>>>>>>>> identify issues – not to replace the role of the reviewer with more
>>>>>>>>>>> electricity. The role of the reviewers extends beyond identifying 
>>>>>>>>>>> and
>>>>>>>>>>> highlighting defects. It encompasses architecture, harmony with the
>>>>>>>>>>> existing codebase, thinking ahead to future evolution of the 
>>>>>>>>>>> project, and
>>>>>>>>>>> replicates context on the project as new code is committed. These 
>>>>>>>>>>> functions
>>>>>>>>>>> cannot be automated away.
>>>>>>>>>>>>>>>> – Hesitancy of authors to engage manually with code they have
>>>>>>>>>>> generated: This is not specific to Cassandra, but it is a behavior 
>>>>>>>>>>> that I
>>>>>>>>>>> have seen in several “highly-electric” projects. There’s a bimodal 
>>>>>>>>>>> tendency
>>>>>>>>>>> toward code that is entirely generated or entirely human-authored - 
>>>>>>>>>>> but it
>>>>>>>>>>> is rare for someone to prepare an AI-authored patch to take an 
>>>>>>>>>>> offramp and
>>>>>>>>>>> spend a significant amount of time refining the work by hand in an 
>>>>>>>>>>> IDE.
>>>>>>>>>>> This hesitancy toward human participation in authorship of 
>>>>>>>>>>> LLM-generated
>>>>>>>>>>> code is very concerning to me.
>>>>>>>>>>>>>>>> – Harmony with the existing codebase: Due to the tunnel-vision 
>>>>>>>>>>>>>>>> of
>>>>>>>>>>> context windows, LLMs are generally unaware of conventions and norms
>>>>>>>>>>> present in codebases and very frequently reinvent concepts in a 
>>>>>>>>>>> generation
>>>>>>>>>>> turn to suit a goal without view of the project’s overall 
>>>>>>>>>>> architecture.
>>>>>>>>>>> This results in a profusion of messy and duplicated concepts that 
>>>>>>>>>>> gradually
>>>>>>>>>>> sprawl about a codebase.
>>>>>>>>>>>>>>>> Again, none of these are grounds for prohibition of usage of
>>>>>>>>>>> language models in developing the project. They’re just problems we 
>>>>>>>>>>> need to
>>>>>>>>>>> bear in mind and guard against – and I think the proposal is 
>>>>>>>>>>> designed to do
>>>>>>>>>>> just that.
>>>>>>>>>>>>>>>> I’m thrilled by the potential of LLMs to improve Apache 
>>>>>>>>>>>>>>>> Cassandra
>>>>>>>>>>> and we already see it happening through a vast number of issues 
>>>>>>>>>>> that are
>>>>>>>>>>> being reported and fixed. But there’s also danger in taking ATVs 
>>>>>>>>>>> down a
>>>>>>>>>>> hiking trail full of people.
>>>>>>>>>>>>>>>> Regarding the prohibition on prose, I’ll simply say: I recently
>>>>>>>>>>> found myself in a scenario where I found a Claude-authored document 
>>>>>>>>>>> so
>>>>>>>>>>> inscrutable that I piped it back into a model, directed it to 
>>>>>>>>>>> rewrite it in
>>>>>>>>>>> ASD-STE100, read it myself, and responded based on the 
>>>>>>>>>>> summarization. As a
>>>>>>>>>>> humanities grad, this is probably the worst language crime I have
>>>>>>>>>>> committed. But it was in response to language that was itself so
>>>>>>>>>>> idiosyncratic that it was unreadable to me in its original form. I 
>>>>>>>>>>> hope
>>>>>>>>>>> this never happens in the Apache Cassandra project.
>>>>>>>>>>>>>>>> I’ll close with a quote from an excellent article written by 
>>>>>>>>>>>>>>>> Colin
>>>>>>>>>>> Breck, an engineer who works on large-scale data systems:
>>>>>>>>>>> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/
>>>>>>>>>>>>>>>> Colin wrote:
>>>>>>>>>>>>>>>>> I don’t want to live in a world where you use AI to summarize
>>>>>>>>>>> something important into unreadable text, and then I use AI in an 
>>>>>>>>>>> attempt
>>>>>>>>>>> to decipher it. I want to hear you, imperfections and all. I want 
>>>>>>>>>>> your
>>>>>>>>>>> interpretation of aesthetics, beauty, quality, relationship, time. 
>>>>>>>>>>> I want
>>>>>>>>>>> to know how you feel. I want you to cut through and tell me what 
>>>>>>>>>>> really
>>>>>>>>>>> matters.
>>>>>>>>>>>>>>>>> Intentional writing will likely become more valuable. People 
>>>>>>>>>>>>>>>>> who
>>>>>>>>>>> write, and write to think, to think deeply and carefully, or to 
>>>>>>>>>>> create, to
>>>>>>>>>>> share, or to capture something important without explicitly 
>>>>>>>>>>> expressing it
>>>>>>>>>>> will continue to write and produce original work. The people who 
>>>>>>>>>>> never were
>>>>>>>>>>> writers will use AI to produce lots of text.
>>>>>>>>>>>>>>>> I hope that our culture can remain one of intentional writing 
>>>>>>>>>>>>>>>> and
>>>>>>>>>>> intentional engineering. I enjoy reading the voice of the author in
>>>>>>>>>>> comments, code, and tickets in Cassandra – the different ways we use
>>>>>>>>>>> language based on where we grew up and how we learned English, the
>>>>>>>>>>> translated idioms from our various backgrounds, and terse comments 
>>>>>>>>>>> that
>>>>>>>>>>> recognize the difference between code whose function is obvious and 
>>>>>>>>>>> what
>>>>>>>>>>> warrants genuine exposition. When I read code in Cassandra, it’s a 
>>>>>>>>>>> delight
>>>>>>>>>>> to recognize the author based on their writing style before 
>>>>>>>>>>> flipping on
>>>>>>>>>>> `git annotate` to reveal the origin.
>>>>>>>>>>>>>>>> I’d encourage folks to re-read the original proposal below. It 
>>>>>>>>>>>>>>>> is
>>>>>>>>>>> very permissive. The guidance strikes me not just as reasonable, but
>>>>>>>>>>> genuinely important to maintaining the health of the project.
>>>>>>>>>>>>>>>> – Scott
>>>>>>>>>>>>>>>> =====
>>>>>>>>>>>>>>>> Encouraged:
>>>>>>>>>>>>>>>> - Reviewing and otherwise validating human-authored patches 
>>>>>>>>>>>>>>>> before
>>>>>>>>>>> submission
>>>>>>>>>>>>>>>> - Debugging, diagnosing etc
>>>>>>>>>>>>>>>> Permitted:
>>>>>>>>>>>>>>>> - Generating or modifying tests, scripts, tooling or any other
>>>>>>>>>>> non-user facing changes
>>>>>>>>>>>>>>>> - Minor changes to human-authored patches that are carefully
>>>>>>>>>>> reviewed by the author
>>>>>>>>>>>>>>>> Restricted:
>>>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by 
>>>>>>>>>>>>>>>> contributors
>>>>>>>>>>> with demonstrated expertise
>>>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area
>>>>>>>>>>> unassisted and with minimal third-party guidance
>>>>>>>>>>>>>>>> - Core code changes made by LLM require an additional reviewer
>>>>>>>>>>>>>>>> - LLM review is not a substitute for human review, and must be 
>>>>>>>>>>>>>>>> used
>>>>>>>>>>> only to augment a complete and independent human understanding of 
>>>>>>>>>>> the patch.
>>>>>>>>>>>>>>>> Prohibited:
>>>>>>>>>>>>>>>> - All public prose must be human authored. This includes inline
>>>>>>>>>>> comments, docs, posts to Jira etc.
>>>>>>>>>>>>>>>> All LLM generated changes MUST be disclosed:
>>>>>>>>>>>>>>>> - Outlined to any reviewer;
>>>>>>>>>>>>>>>> - Summarised in the commit message;
>>>>>>>>>>>>>>>> - Large blocks or files must be individually marked with some 
>>>>>>>>>>>>>>>> agreed
>>>>>>>>>>> message like "created by <some AI>"
>>>>>>>>>>>>>>>> =====
>>>>>>>>>>>>>>>>> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected]
>>>>>>>>>>> <mailto:[email protected]>> wrote:
>>>>>>>>>>>>>>>>> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected]
>>>>>>>>>>> <mailto:[email protected]>> wrote:
>>>>>>>>>>>>>>>>>> Restricted:
>>>>>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by
>>>>>>>>>>> contributors with demonstrated expertise
>>>>>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area
>>>>>>>>>>> unassisted and with minimal third-party guidance
>>>>>>>>>>>>>>>>> I am -1 on this. This sounds like gate keeping attempt. It 
>>>>>>>>>>>>>>>>> narrowly
>>>>>>>>>>> limits the pool to a few people on the project that have 
>>>>>>>>>>> historically
>>>>>>>>>>> contributed to certain parts of the codebase. This policy will 
>>>>>>>>>>> prohibit
>>>>>>>>>>> skilled software engineers with domain expertise from proposing LLM
>>>>>>>>>>> assisted changes simply because they have not contributed to the 
>>>>>>>>>>> project.
>>>>>>>>>>> This is unrealistic and a net negative for the project to attract 
>>>>>>>>>>> talent
>>>>>>>>>>> and grow our community.
>>>>>>>>>>>>>>>>>> - Core code changes made by LLM require an additional 
>>>>>>>>>>>>>>>>>> reviewer
>>>>>>>>>>>>>>>>> Can you be more precise what is this in addition to? How many 
>>>>>>>>>>>>>>>>> total
>>>>>>>>>>> reviewers do you expect and what is the purpose of additional 
>>>>>>>>>>> reviewer? and
>>>>>>>>>>> why?
>>>>>>>>>>>>>>>>> Taking a step back - what are you trying to solve here?
>>>>>>>>>>>>>>>>> Dinesh

Reply via email to