For the record Josh, I am including pre-3.0. I was involved in customer support escalations for 1.2, 2.0 and 2.1: the software was shoddy, to put it mildly. 8099 was a symptom, not the disease.
The project had a culture of doing stuff without sufficient care or consideration. This is by no means a simple matter of needing more testing or less timeline pressure, nor is it easy to define a "bar" to ensure it does not regress. If we had simple metrics, we would have settled it a long time ago. You can still see customer scars crop up on forum discussions periodically. On 2026/09/24 19:01:02 Josh McKenzie wrote: > > Apache Cassandra was fundamentally undeployable for four years between Nov > > 2015 - 2019. > CASSANDRA-8099 was a maximal manifestation of a specific approach to > engineering and calendar constraints we've seen time and again on the > project; I don't want us to conflate things here. That was a herculean > monolithic body of work performed in inhuman conditions (in vim!) that was > ultimately so invasive, all the unit tests in the code-base were commented > out and Jake and I spent a grueling 1.5-2 months hand-rewriting basically all > the unit tests in that code-base to get things to even build and run, much > less pass. > > Massive blast radius changes that are un-sustainably complex, under-tested, > where we don't property, fuzz, check coverage, check complexity, or A/B > compare against a known good system (i.e. pre/post) correctness testing are > going to destabilize the database at any time under any regime of tooling. We > certainly could speed-run our way back into destabilization with LLM's as > they are a force multiplier for both the good and the bad of one's > engineering practices, but that's a solvable problem by better defining what > our bars of quality are (Definition of Done anyone?) and holding ourselves > accountable to delivering at that bar. > > On Thu, Sep 24, 2026, at 2:40 PM, David Capwell via dev wrote: > >> > >> > >> if you really want to pursue it I would ask that we do it offline to avoid > >> polluting an already busy conversation > >> > > People are directly responding saying that they feel discrimination > > currently and that the policy tries to codify that discrimination, so I > > feel its 100% on topic. This thread has presented 0 evidence that LLM usage > > has lowered the quality of contributions merged and has so far been vibes > > and feeling; I have yet to see any evidence to justify such discrimination > > so I will keep pushing back until such evidence is presented so we can have > > a informed debate. > > > >> namely how we handle shallow and localised bug fixes. I would be happy > >> adding a clear entry to the “Permitted” section for this. No doubt there > >> are many refinements needed to the Restricted text as well, that might > >> also capture some of your concerns. > >> > > I will not sign off on cherry picking areas where "safe" to use a tool; so > > no it does not capture my concerns. > > > > > > > >> I don't know if everyone remembers, but ten years ago Cassandra was full > >> of serious correctness and stability issues. Despite developing it, I > >> would not have run it myself or recommend that anyone use it. We have dug > >> ourselves out of that hole, but it took years of discipline and effort, > >> and we're still (deservedly) recovering our reputation. > >> > >> Let's use this new technology to improve the quality of our contributions, > >> not squander our hard-earned gains in the name of speed. It will be hard > >> to recover our reputation a second time. > >> > > I do recall the 3.x line and put in a significant amount of effort to > > harden it. There were behaviors I noticed after joining Cassandra that I > > feel directly contributed to 3.0 and the decade of catch up; behaviors that > > still linger in parts of the community today. > > > > As I look on trunk and look at committed code and trace back to PRs and > > JIRA I see the following: > > > > • large patches approved without comments > > • 0 evidence that tests were run > > I then look at our CI and see tests failing for months. As you start to > > triage you start to see some of them show real issues; yet they linger for > > months not being addressed... when CI is unstable it takes a lot of effort > > to triage "did my patch break the test", and I have seen time and time > > again people do not put in that effort, and shrug off as "its just a flakey > > test"; then our CI failure rate grows. > > > > Non of this has anything to do with LLMs but LLMs running in this > > environment is far more dangerous as there are not checks in place to "hold > > the bar". I am all for raising the bar universally; expecting both humans > > and LLMs to match that bar. > > > >> `I’d support something that boils down to roughly this: > >> > >> 1.) 2 committers must understand an LLM-assisted change before it commits. > >> (Perhaps separately we can explore the question of why we haven’t added > >> any new committers to the core project for about a year. I’m also still > >> not entirely sure if it’s acceptable within our guidelines for a committer > >> to +1 a patch after delegating review.) > >> > >> 2.) Patch authors must demonstrate enough understanding to discuss their > >> own patch, whether or not parts of it are generated by an LLM. > >> > >> 3.) The “Assisted-by” tag should be used to indicate any non-trivial LLM > >> usage in the generation of a patch, just like we have used Co-authored-by > >> historically. > >> > >> 4.) Comments and other things that aren't the actual code (but could sow > >> confusion) should be held to the same standard we'd expect from a human > >> writer. If we don’t yet agree on that standard, we can formalize enough of > >> it to guide both humans and LLMs. > ` > > I can get behind this proposal but i would tweak it as 1/2 i don't think > > really need to special case LLM usage > > > > 1. 2 committers must understand the change before it commits. > > 2. Patch authors must demonstrate enough understanding to discuss their > > own patch > > Nothing about those 2 need to be scoped to LLM usage and honestly matches > > most PMCs I have talked to understanding of our bar (as Benedict pointed > > out, the actual wording could be interpreted to allow rubber stamping from > > committer) > > > > As for 3 I am cool with this. ASF recommends the same (it says > > `Generated-by` but that discount's the human's effort) as its useful for > > audits and tooling. Having `Assisted-by` tag should not imply anything > > about the committed patch as it should have gone through the same bar we > > all expect; its just for tools auditing. > > > > > >> On Sep 24, 2026, at 10:37 AM, Aleksey Yeshchenko via dev > >> <[email protected]> wrote: > >> > >> Meant "doing away with", sorry. Non-native speaker with a headache here. > >> Thanks Caleb for spotting. > >> > >>> On 24 Sep 2026, at 17:35, Aleksey Yeshchenko via dev > >>> <[email protected]> wrote: > >>> > >>> P.S. I assume it's obvious from the text above that I don't believe that > >>> getting away with human code review is a viable option. > >> > >> > >>> On 24 Sep 2026, at 17:37, Štefan Miklošovič <[email protected]> > >>> wrote: > >>> > >>> Good call on checkerframework, we even have a patch for it. Work of > >>> Jacek Lewandowski. We might just drive it to completion. Using AI for > >>> finishing it would be quite ironic. > >>> > >>> (1) https://github.com/apache/cassandra/pull/2370 > >>> > >>> On Thu, Sep 24, 2026 at 6:19 PM Jon Haddad <[email protected]> > >>> wrote: > >>>> > >>>> There are some really good points being brought up about stability of > >>>> the codebase, maintainability, quality of reviews, correctness bugs, and > >>>> I agree with all of them. I think it would be helpful to take a step > >>>> back and consider how those bugs got there in the first place, how they > >>>> were fixed, and what we could do to further advance the codebase so they > >>>> don't creep back. LLMs can be used either with very tight guardrails, > >>>> or in what's effectively YOLO mode, and there's a big difference in the > >>>> quality of the results you get. > >>>> > >>>> One thing to keep in mind, a lot of the initial code in C* was added > >>>> without comprehensive testing. I hope we can all agree that it's a lot > >>>> easier to break code that doesn't have high quality tests. During the > >>>> code freeze, a lot of people people spent several years relentlessly > >>>> finding and fixing bugs. This was probably a pretty frustrating time for > >>>> anyone who was focused on fixing other people's bugs when they wanted to > >>>> build features. I think we should recognize the effort here and > >>>> appreciate the foundation that the project stands on now. I can > >>>> understand how anyone involved with this effort would be apprehensive > >>>> about seeing years of their life swept away by an agent that was driven > >>>> by goal seeking to remove all the tests that it broke instead of fixing > >>>> them. > >>>> > >>>> When I picked up the work to improve cursor compaction, the first thing > >>>> I asked myself was how can I make sure I don't break this? How do I > >>>> even know it works properly? There were some tricky parts to the code, > >>>> and I really didn't want to come in and immediately break stuff. That's > >>>> why I started with an entire patch dedicated to adding test infra to it. > >>>> 90% of the patch was tests, and in my other cursor patches, it remains > >>>> *at least* 80% of my patches. It was a *lot* faster to add almost 10K > >>>> lines of tests that handled a byte for byte differential testing paired > >>>> with harry to find over 30 bugs that caused cursor to corrupt results. > >>>> Range tombstones alone were at least a dozen bugs, but I also found > >>>> issues with static columns, reverse ordering, etc. Randomizing schemas > >>>> and data in burn tests to generate different shapes of data, to ensure > >>>> they all result in the same output at the end. JMH tests to ensure there > >>>> weren't performance regressions, hours of profiling. These were all a > >>>> *lot* easier to do with the LLM helping me out. In the process I've > >>>> found bugs that have been lingering in the codebase for years. > >>>> > >>>> That's a long story, but hopefully we all agree that having > >>>> comprehensive tests is a great way to ensure that both humans and LLMs > >>>> don't break things that are working. > >>>> > >>>> The lesson: we need to keep improving our testing. Everything that we > >>>> touch, should be left in a better state than how we found it with regard > >>>> to test coverage. > >>>> > >>>> Test coverage isn't everything though, there's always little subtle bugs > >>>> that don't get found in testing, that can slip in despite our best > >>>> efforts. It's debatable if humans will be as good as agents for coding > >>>> in the long term, for spotting small defects. I sincerely doubt it. > >>>> For the time being though, we still have people involved. It's probably > >>>> a good time to start using more static analysis tools to identify > >>>> problematic code and to add this to CI. Dmitry had a suggestion > >>>> recently for checkerframework to detect leaking contexts, a problem he > >>>> spotted when reviewing my branch. It would be great to have that > >>>> integrated into our CI and dev workflow so we can simply avoid an entire > >>>> class of bugs. > >>>> > >>>> There's also PMD, which is excellent for finding code that can be hard > >>>> to understand. I *highly* suggest you all run PMD to analyze for > >>>> cognitive complexity and high npath scores. This was made popular by > >>>> the folks at Sonar and I've found it to be an excellent feedback > >>>> mechanism for structuring code. The default max they set is 15, which > >>>> is the point where it starts to become difficult to verify something > >>>> works without making a massive investment. We've got areas in the > >>>> codebase that are in the hundreds, and some parts even higher. These > >>>> have been contributed by humans, and are all high risk points for both > >>>> humans and agents to start messing around with. They're also in some > >>>> fairly critical areas that are very likely to break, so I understand why > >>>> people would not want an agent anywhere near it. > >>>> > >>>> Unfortunately, it's not an easy problem to address. There's so many > >>>> places where the code is structured in a way that has so many branches, > >>>> so many conditions, that it's effectively impossible for a human to > >>>> understand, creating a fear of messing around in it. There's plenty of > >>>> areas that deserve extreme scrutiny, and we should be careful of what we > >>>> add, whether it's human or agent. > >>>> > >>>> The codebase today requires a high degree of internal knowledge to > >>>> navigate. There's land mines everywhere. We should be looking to make > >>>> conscious improvements by moving the code forward, so it's easier to > >>>> make changes to small, well tested components with minimal side effects. > >>>> Not making it harder for people to use the tools that aid in that > >>>> process. > >>>> > >>>> Here's what we could do to achieve the underlying goal of not breaking > >>>> the DB: > >>>> > >>>> Add cognitive complexlity and npath via PMD as a feedback mechanism. > >>>> > >>>> Code that's hard to understand is hard to review. It's also hard to > >>>> test. Let's break down the complex code so more people can contribute, > >>>> safely. > >>>> > >>>> Add checkerframework to our tooling, > >>>> > >>>> Properly annotate the codebase for it and reduce the surface area that > >>>> things can break. Less brittle codebase = we can move faster. > >>>> > >>>> Use jacoco to find areas of the codebase with poor testing. > >>>> > >>>> Let's improve the test coverage there, LLMs are great for this. We have > >>>> a ton of static tests, these can become more dynamic, parameterized, and > >>>> leverage harry. > >>>> > >>>> Refactor parts of the codebase that have high cognitive complexlity and > >>>> NPath scores. > >>>> > >>>> This should be lowered over time to meet some high watermark, say 25 > >>>> maximum, although I'd prefer 15 which is where the Sonar folks settled. > >>>> > >>>> Move forward moving the codebase to a more modular structure > >>>> > >>>> We've talked about Gradle on and off - but it can really be a huge help > >>>> with incremental, modular builds. This is pretty easy to do with an > >>>> agent and we could have it done in a couple days. > >>>> > >>>> Enforce boundaries with ArchUnit > >>>> > >>>> If we want to enforce certain code boundaries, this is the way to do it. > >>>> Should not be part of manual review. > >>>> > >>>> Add LLM review for all incoming PRs before a human > >>>> > >>>> The goal here is to automate the initial part of the review process that > >>>> reviewers should spot, and raise the bar for the initial contribution. > >>>> When the code gets reviewed by a human, it should already have passed a > >>>> large variety of initial checks. This should shorten the review cycle > >>>> and result in higher quality patches. I've had Claude reviewing all my > >>>> PRs in my personal projects for a while now and it consistently gives > >>>> great feedback that I almost always incorporate. > >>>> > >>>> In my ideal world, we'd also auto-format all code > >>>> > >>>> Consistent formatting throughout the codebase would be amazing, but > >>>> that's just one man's dream. > >>>> > >>>> Hopefully there's at least a couple things in this list we could move > >>>> forward with in the short term, as it'll help improve the code quality > >>>> regardless of how it's created. > >>>> > >>>> Jon > >>>> > >>>> https://checkerframework.org/manual/#aliasing-leaking-contexts > >>>> https://www.sonarsource.com/docs/CognitiveComplexity.pdf > >>>> https://pmd.github.io/pmd/pmd_rules_java_design.html > >>>> > >>>> > >>>> > >>>> > >>>> > >>>> On Thu, Sep 24, 2026 at 7:38 AM C. Scott Andreas <[email protected]> > >>>> wrote: > >>>>> > >>>>> From Benedict: > >>>>> > >>>>> “I don't know if everyone remembers, but ten years ago Cassandra was > >>>>> full of serious correctness and stability issues. Despite developing > >>>>> it, I would not have run it myself or recommend that anyone use it. We > >>>>> have dug ourselves out of that hole, but it took years of discipline > >>>>> and effort, and we're still (deservedly) recovering our reputation.” > >>>>> > >>>>> Expanding on this point for those who may not have been active in the > >>>>> project at this time — > >>>>> > >>>>> Apache Cassandra was fundamentally undeployable for four years between > >>>>> Nov 2015 - 2019. The database literally lost data if you ran a > >>>>> read-only SELECT query ordered descending (C-14513, C-14515). If you > >>>>> haven’t read these tickets before, please take a moment to do so. > >>>>> > >>>>> It took years of careful work via property-based testing, fuzzing, and > >>>>> deterministic simulation to restore Cassandra’s status as a usable > >>>>> system of record. Once 14513 and 14515 were identified, nearly 30 > >>>>> additional critical data loss and incorrect response bugs were > >>>>> identified. > >>>>> > >>>>> It is essential for the project’s future that we don’t regress to this > >>>>> state chasing AI-generated features motivated by fear. The fact that > >>>>> examples cited in this thread which boast shiny features but have > >>>>> critical shortcomings unknown to their author supports this argument. > >>>>> > >>>>> The most common path for large corpuses of AI-generated software is > >>>>> elation and reveling in a feature matrix, followed by abandonment. > >>>>> > >>>>> I endorse this point: > >>>>> > >>>>> “Let's use this new technology to improve the quality of our > >>>>> contributions, not squander our hard-earned gains in the name of speed. > >>>>> It will be hard to recover our reputation a second time.” > >>>>> > >>>>> Patrick, I don’t want your note regarding a TCM issue to go > >>>>> unaddressed. Please file a Jira ticket and the patch if you like. I > >>>>> can’t comment on the patch as I haven’t seen it, but together we will > >>>>> solve the problem. > >>>>> > >>>>> – Scott > >>>>> > >>>>>> On Sep 24, 2026, at 4:01 AM, Benedict Elliott Smith > >>>>>> <[email protected]> wrote: > >>>>>> > >>>>>> Hi Patrick, > >>>>>> > >>>>>> As I mentioned in my reply to David, I would be happy to create a > >>>>>> carve out for shallow and localised bug fixes in the "Permitted" > >>>>>> section. Would this alleviate some of your concerns regarding your > >>>>>> ability to contribute to the project? > >>>>>> > >>>>>> I appreciate your pointing out Ferrosa's Accord implementation > >>>>>> however, as it is a *great* example of the problems we're leaping > >>>>>> into. I took a look, and within about 30s found that the protocol is > >>>>>> fundamentally incorrect, having failed to address CASSANDRA-18365. > >>>>>> This is despite claiming to be tested with Jepsen that should in > >>>>>> principle find this fault. I followed up by using Claude to > >>>>>> interrogate the implementation further, and immediately found other > >>>>>> serious correctness issues. > >>>>>> > >>>>>> I use LLMs daily now to help facilitate Accord development, and while > >>>>>> they are powerful they are NOT able to author the code themselves, > >>>>>> even when building upon a strong human-authored foundation. > >>>>>> > >>>>>> I don't know if everyone remembers, but ten years ago Cassandra was > >>>>>> full of serious correctness and stability issues. Despite developing > >>>>>> it, I would not have run it myself or recommend that anyone use it. We > >>>>>> have dug ourselves out of that hole, but it took years of discipline > >>>>>> and effort, and we're still (deservedly) recovering our reputation. > >>>>>> > >>>>>> Let's use this new technology to improve the quality of our > >>>>>> contributions, not squander our hard-earned gains in the name of > >>>>>> speed. It will be hard to recover our reputation a second time. > >>>>>> > >>>>>> > >>>>>>> On 2026/09/23 19:16:17 Patrick McFadin wrote: > >>>>>>> I was waiting for this moment to hit our project and I'm glad we're > >>>>>>> here. I > >>>>>>> am deeply concerned for our project and its future, as we have > >>>>>>> increasingly > >>>>>>> made it difficult to contribute. I had hoped that this new era of > >>>>>>> software tools powered by AI would expand the project's reach and > >>>>>>> bring > >>>>>>> more diverse thoughts and ideas. This policy proposal is the exact > >>>>>>> opposite > >>>>>>> of what we need. We have been sitting on a Cassandra 6 release alpha > >>>>>>> for > >>>>>>> months. We need to accelerate and embrace new ways of being or be left > >>>>>>> behind. As I read that policy, my first and gut level reactions: > >>>>>>> - It comes across as elitist and class protectionism. Committer > >>>>>>> should not > >>>>>>> be special but this proposal makes that designation even more sacred. > >>>>>>> - It signals that our project is so fragile that only a few people > >>>>>>> "Really > >>>>>>> understand it" That's some SQLite vibes right there. > >>>>>>> - Trying to fix a problem that doesn't exist > >>>>>>> Sadly, i think this policy change would also exclude a lot of > >>>>>>> comitters. > >>>>>>> We aren't alone in this moment. The Linux project just went through > >>>>>>> this. You can find the thread with a simple Google, but similar hard > >>>>>>> feelings were being expressed "AI is going to ruin our project!", "The > >>>>>>> unwashed masses are going to contribute terrible code!", "We have to > >>>>>>> protect our precious status as Linux maintainers!" Linus being > >>>>>>> Linus, was > >>>>>>> deeply invloved and they adopted a super simple statement that covers > >>>>>>> all > >>>>>>> bases. Human or Human using AI. “You are expected to understand and > >>>>>>> to be > >>>>>>> able to defend everything you submit.” Love that. > >>>>>>> In the larger picture, I'll restate. I'm worried for our project. In > >>>>>>> late > >>>>>>> 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real corner > >>>>>>> and > >>>>>>> in the hands of somebody that knows how to build software, this tool > >>>>>>> is > >>>>>>> like jet fuel. Here's some examples of new projects being hyper > >>>>>>> fueled by > >>>>>>> AI coding tools. > >>>>>>> Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity > >>>>>>> Turso - Rust re-write of SQLite > >>>>>>> Bun - Rust re-write of itself from Zig. > >>>>>>> Think this couldn't happen to us? Already has: > >>>>>>> https://github.com/ferrosadb/ferrosa. Ben is using it to power his own > >>>>>>> startup, but it was him alone using a ton of local AI coding agents. > >>>>>>> He > >>>>>>> even implemented Accord. Yeah... > >>>>>>> The cracks are already starting to show. There is a black market > >>>>>>> economy of > >>>>>>> Cassandra patches happening now. Not going to name names or call > >>>>>>> people > >>>>>>> out, but there are fixes and optimizations living in branches > >>>>>>> outside of > >>>>>>> the Cassandra project. Why? I'll use myself as an example. I fixed a > >>>>>>> nasty > >>>>>>> bug I ran into with TCM a few weeks ago. Wrote the tests. It passes > >>>>>>> CI and > >>>>>>> lives in my personal branch. I'm sitting here really wondering if I > >>>>>>> want to > >>>>>>> go through the ritual humiliation of being roasted for using AI to > >>>>>>> fix it. > >>>>>>> Me. I am worried about contrinuting code the Cassandra. What the hell > >>>>>>> does > >>>>>>> that say? > >>>>>>> I have my CQLite project that I've been doing a release around once a > >>>>>>> month. I would love to donate that to the Cassandra project but I > >>>>>>> wouldn't > >>>>>>> if it essentially killed any progress. > >>>>>>> My larger counter proposal would be to: > >>>>>>> - Adopt the “You are expected to understand and to be able to defend > >>>>>>> everything you submit.” approach the Linux project has adopted. > >>>>>>> - Loosen up the contributor process and our worry on trunk. Let 1000 > >>>>>>> flowers bloom and bring it in. > >>>>>>> - And finally, to give some people more peace of mind and open more > >>>>>>> doors, > >>>>>>> adopt what other projects have done and provide more pluggability. > >>>>>>> Let new > >>>>>>> ideas have an easy place to connect. > >>>>>>> We are at a fork in the road. What are we going to do? And then I > >>>>>>> have to > >>>>>>> ask myself, what am I going to do as a contributor? > >>>>>>> Patrick > >>>>>>> On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston <[email protected]> > >>>>>>> wrote: > >>>>>>>> I’m not necessarily opposed to having a policy, but so far we have > >>>>>>>> some > >>>>>>>> specific proposals addressing a problem statement that’s very > >>>>>>>> nebulous. > >>>>>>>> What is the community failing to do on its own that we’re trying to > >>>>>>>> correct > >>>>>>>> with policy? What outcomes are we trying to create or prevent? > >>>>>>>> Having some > >>>>>>>> examples and specific problems to discuss would help focus the > >>>>>>>> conversation. > >>>>>>>>> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote: > >>>>>>>> Benedict, > >>>>>>>> Thanks for clarifying. My concern still remains. This criteria would > >>>>>>>> be > >>>>>>>> difficult to define and apply consistently. What counts as “similar” > >>>>>>>> scope > >>>>>>>> or area, “mostly correct,” or sufficiently independent work? More > >>>>>>>> importantly, how do we prevent such vague criteria from creating an > >>>>>>>> informal hierarchy where some contributors work is routinely > >>>>>>>> accepted while > >>>>>>>> others is routinely rejected? > >>>>>>>> If the intent is to limit AI-assisted code changes to Cassandra > >>>>>>>> contributors, or to contributors who have previously worked in that > >>>>>>>> component without AI, that would at least be clear and enforceable. > >>>>>>>>> On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith < > >>>>>>>> [email protected]> wrote: > >>>>>>>>> Core code changes > >>>>>>>>> Chris: Do you object to the first or second line you quote? Because > >>>>>>>>> the > >>>>>>>> first line is effectively motivation for the second line, and can be > >>>>>>>> removed (or more clearly combined). If it’s the second line, then I > >>>>>>>> do not > >>>>>>>> think this is an unreasonable expectation, and we can get into a > >>>>>>>> proper > >>>>>>>> debate about it. > >>>>>>>>> Shailaja, since you only snipped the first sentence, your concerns > >>>>>>>>> might > >>>>>>>> also be mostly answered by this clarification? “Minimal third-party > >>>>>>>> guidance” implies you have some concerns about the second line, but > >>>>>>>> all of > >>>>>>>> our policies have some ambiguity because legalese is even worse. I > >>>>>>>> don’t > >>>>>>>> think the ambiguity here would be challenging to navigate though we > >>>>>>>> can > >>>>>>>> certainly refine it. This specific snippet is meant to convey an > >>>>>>>> expectation that a contributor has autonomously produced patches of > >>>>>>>> similar > >>>>>>>> scope that were mostly correct, so that they have demonstrated the > >>>>>>>> level of > >>>>>>>> understanding necessary to guide another party to a successful patch > >>>>>>>> (i.e. > >>>>>>>> an LLM in this case). > >>>>>>>>> On 2026/09/23 10:54:16 Benedict Elliott Smith wrote: > >>>>>>>>>> Thanks everyone for your input so far. I’ll respond in brief to the > >>>>>>>> main themes, in (mostly) separate emails so they can each have their > >>>>>>>> own > >>>>>>>> debate chain. > >>>>>>>>>> Should we have a policy (Blake/Josh*/Jon/Dinesh) > >>>>>>>>>> I think we would all agree that LLMs represent the biggest change > >>>>>>>>>> to > >>>>>>>> this community (and software more generally) since its inception, > >>>>>>>> and we > >>>>>>>> all now have enough experience with the technology to have formed > >>>>>>>> opinions > >>>>>>>> about how it is best managed. We also evidently have not all arrived > >>>>>>>> at the > >>>>>>>> same conclusions. In this situation, it would be an abdication of our > >>>>>>>> responsibilities as a management committee to not agree *some* > >>>>>>>> policy. > >>>>>>>>>> I intend to conduct straw polls as the discussion evolves, so if > >>>>>>>>>> you > >>>>>>>> prefer an alternative policy - or modifications to this policy - I > >>>>>>>> would > >>>>>>>> encourage you to make those alternative proposals. > >>>>>>>>>> *Veto/Consensus (Josh) > >>>>>>>>>> It was fair to call out my poor use of language on this topic, so > >>>>>>>>>> let > >>>>>>>> me rephrase a little. The community is built on consensus, and work > >>>>>>>> should > >>>>>>>> not be merged when there are outstanding concerns to address. The > >>>>>>>> explicit > >>>>>>>> -1 should only be used rarely, because the prior expectation should > >>>>>>>> prevent > >>>>>>>> it ever being needed. I (and others) have outstanding concerns on LLM > >>>>>>>> generated work that can only be addressed through this process right > >>>>>>>> here, > >>>>>>>> so to merge such work while maintaining the community’s consensus we > >>>>>>>> must > >>>>>>>> agree some policy. > >>>>>>>>>> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote: > >>>>>>>>>>> I am strongly -1 on this > >>>>>>>>>>> - Core code changes made by LLM may only be proposed by > >>>>>>>>>>> contributors > >>>>>>>> with demonstrated expertise > >>>>>>>>>>> That creates a new, subjective privileged class of contributors > >>>>>>>>>>> and > >>>>>>>> turns a tool choice into an eligibility test. Who decides whether > >>>>>>>> expertise > >>>>>>>> has been “demonstrated,” what counts as “minimal third-party > >>>>>>>> guidance,” and > >>>>>>>> how could those judgments be applied consistently or fairly? > >>>>>>>>>>> Apache already has a better model, anyone may contribute, trust > >>>>>>>>>>> and > >>>>>>>> additional repository privileges are earned transparently over time. > >>>>>>>> The > >>>>>>>> ASF describes its communities as flat, and says that newcomer ideas > >>>>>>>> have as > >>>>>>>> much input as those from original creators. We should not add a > >>>>>>>> separate, > >>>>>>>> informal hierarchy in which certain people may use common > >>>>>>>> development tools > >>>>>>>> while others may not. > >>>>>>>>>>>> On Sep 23, 2026, at 6:33 AM, Chris Lohfink <[email protected]> > >>>>>>>> wrote: > >>>>>>>>>>>> - Core code changes made by LLM may only be proposed by > >>>>>>>>>>>> contributors > >>>>>>>> with demonstrated expertise > >>>>>>>>>>>> - Must have produced similar patches in size, scope and area > >>>>>>>> unassisted and with minimal third-party guidance > >>>>>>>>>>>> I really don't like this one or its wording. Definitely too "the > >>>>>>>> peasants are getting uppity lets build a wall". Lets not let a > >>>>>>>> subjective > >>>>>>>> thing like demonstrated expertise (who decides that?) be if it's ok > >>>>>>>> or not. > >>>>>>>> Hold the same standards for code quality and process for it all. I > >>>>>>>> don't > >>>>>>>> want this to be: only people on the storage team in Apple can use AI. > >>>>>>>>>>>> Chris > >>>>>>>>>>>> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto: > >>>>>>>> [email protected]>> wrote: > >>>>>>>>>>>>> I agree with Stefan and think this is both a reasonable and > >>>>>>>> thoughtful proposal. > >>>>>>>>>>>>> Here are some things I like about it: > >>>>>>>>>>>>> – It outlines areas where LLM usage is unambiguously useful to > >>>>>>>>>>>>> the > >>>>>>>> project’s developers and users. > >>>>>>>>>>>>> – It defines a spectrum of recommendations and cautions. > >>>>>>>>>>>>> – The only prohibited areas are extremely narrow and say nothing > >>>>>>>> about code at all. > >>>>>>>>>>>>> Some in this thread are responding as if this proposal seeks to > >>>>>>>> prohibit or sharply limit use of LLMs. In fact, it’s one of the most > >>>>>>>> open > >>>>>>>> and welcoming I’ve seen for an OSS project of our size where many are > >>>>>>>> adopting policies that simply ban them entirely. I’ve re-appended the > >>>>>>>> proposal below my message as it seems to have been lost in threaded > >>>>>>>> replies, and would encourage folks to give it a second read. > >>>>>>>>>>>>> Some brief thoughts based on my own use of LLMs: > >>>>>>>>>>>>> – I find them fantastically useful for reviewing and identifying > >>>>>>>> problems that have slipped through review - primarily via Alex > >>>>>>>> Petrov’s > >>>>>>>> /deep-review skill, which I have running in a VM in a loop executing > >>>>>>>> over > >>>>>>>> every new commit in the project as of a few days ago. I will be > >>>>>>>> posting a > >>>>>>>> few hand-authored Jira tickets based on findings that appear > >>>>>>>> legitimate to > >>>>>>>> me. For now, the loop is posting them as issue drafts for my own > >>>>>>>> review on > >>>>>>>> my personal fork which you can find here: > >>>>>>>> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug > >>>>>>>>>>>>> – They’re great for enabling use of model checkers and formal > >>>>>>>> methods where such work would have previously been prohibitively > >>>>>>>> expensive, > >>>>>>>> such as Blake’s work on a TLA+ proof of aspects of Mutation Tracking > >>>>>>>> and > >>>>>>>> Benedict/Fedor’s work on a machine-checkable proof of the Accord > >>>>>>>> protocol > >>>>>>>> in Lean. > >>>>>>>>>>>>> – They are stunning for allowing me to experiment with ideas > >>>>>>>>>>>>> that > >>>>>>>> would have otherwise been a summer internship’s scope of work. Some > >>>>>>>> examples include an io_uring prototype, exploring the impact of > >>>>>>>> page-aligned compressed chunk sizes, an API shim bridging the 3.x > >>>>>>>> and 4.x > >>>>>>>> Java Drivers, and potential enhancements to Zstandard. > >>>>>>>>>>>>> – And they shine when given grunt-work that is critical to the > >>>>>>>> project but a miserable labor for humans, such as triaging, > >>>>>>>> reproducing, > >>>>>>>> and root-causing flaky tests, which David Capwell now has running in > >>>>>>>> a loop > >>>>>>>> to help us improve CI stability in the project. > >>>>>>>>>>>>> I never thought I’d be so positive on what’s possible via > >>>>>>>>>>>>> language > >>>>>>>> models a year ago. At the same time, I also agree that they present > >>>>>>>> challenges and risks that can be managed through thoughtful > >>>>>>>> discussion and > >>>>>>>> policy. Some of the concerns that I think are important to guard > >>>>>>>> against > >>>>>>>> include: > >>>>>>>>>>>>> – Asymmetry of effort between author and reviewers: As > >>>>>>>> token-generating machines, LLMs can generate diffs of extraordinary > >>>>>>>> size > >>>>>>>> very rapidly. /deep-review is great for chewing through diffs and > >>>>>>>> identifying defects. But it should be used by the contributor > >>>>>>>> themselves to > >>>>>>>> identify issues – not to replace the role of the reviewer with more > >>>>>>>> electricity. The role of the reviewers extends beyond identifying and > >>>>>>>> highlighting defects. It encompasses architecture, harmony with the > >>>>>>>> existing codebase, thinking ahead to future evolution of the > >>>>>>>> project, and > >>>>>>>> replicates context on the project as new code is committed. These > >>>>>>>> functions > >>>>>>>> cannot be automated away. > >>>>>>>>>>>>> – Hesitancy of authors to engage manually with code they have > >>>>>>>> generated: This is not specific to Cassandra, but it is a behavior > >>>>>>>> that I > >>>>>>>> have seen in several “highly-electric” projects. There’s a bimodal > >>>>>>>> tendency > >>>>>>>> toward code that is entirely generated or entirely human-authored - > >>>>>>>> but it > >>>>>>>> is rare for someone to prepare an AI-authored patch to take an > >>>>>>>> offramp and > >>>>>>>> spend a significant amount of time refining the work by hand in an > >>>>>>>> IDE. > >>>>>>>> This hesitancy toward human participation in authorship of > >>>>>>>> LLM-generated > >>>>>>>> code is very concerning to me. > >>>>>>>>>>>>> – Harmony with the existing codebase: Due to the tunnel-vision > >>>>>>>>>>>>> of > >>>>>>>> context windows, LLMs are generally unaware of conventions and norms > >>>>>>>> present in codebases and very frequently reinvent concepts in a > >>>>>>>> generation > >>>>>>>> turn to suit a goal without view of the project’s overall > >>>>>>>> architecture. > >>>>>>>> This results in a profusion of messy and duplicated concepts that > >>>>>>>> gradually > >>>>>>>> sprawl about a codebase. > >>>>>>>>>>>>> Again, none of these are grounds for prohibition of usage of > >>>>>>>> language models in developing the project. They’re just problems we > >>>>>>>> need to > >>>>>>>> bear in mind and guard against – and I think the proposal is > >>>>>>>> designed to do > >>>>>>>> just that. > >>>>>>>>>>>>> I’m thrilled by the potential of LLMs to improve Apache > >>>>>>>>>>>>> Cassandra > >>>>>>>> and we already see it happening through a vast number of issues that > >>>>>>>> are > >>>>>>>> being reported and fixed. But there’s also danger in taking ATVs > >>>>>>>> down a > >>>>>>>> hiking trail full of people. > >>>>>>>>>>>>> Regarding the prohibition on prose, I’ll simply say: I recently > >>>>>>>> found myself in a scenario where I found a Claude-authored document > >>>>>>>> so > >>>>>>>> inscrutable that I piped it back into a model, directed it to > >>>>>>>> rewrite it in > >>>>>>>> ASD-STE100, read it myself, and responded based on the > >>>>>>>> summarization. As a > >>>>>>>> humanities grad, this is probably the worst language crime I have > >>>>>>>> committed. But it was in response to language that was itself so > >>>>>>>> idiosyncratic that it was unreadable to me in its original form. I > >>>>>>>> hope > >>>>>>>> this never happens in the Apache Cassandra project. > >>>>>>>>>>>>> I’ll close with a quote from an excellent article written by > >>>>>>>>>>>>> Colin > >>>>>>>> Breck, an engineer who works on large-scale data systems: > >>>>>>>> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/ > >>>>>>>>>>>>> Colin wrote: > >>>>>>>>>>>>>> I don’t want to live in a world where you use AI to summarize > >>>>>>>> something important into unreadable text, and then I use AI in an > >>>>>>>> attempt > >>>>>>>> to decipher it. I want to hear you, imperfections and all. I want > >>>>>>>> your > >>>>>>>> interpretation of aesthetics, beauty, quality, relationship, time. I > >>>>>>>> want > >>>>>>>> to know how you feel. I want you to cut through and tell me what > >>>>>>>> really > >>>>>>>> matters. > >>>>>>>>>>>>>> Intentional writing will likely become more valuable. People > >>>>>>>>>>>>>> who > >>>>>>>> write, and write to think, to think deeply and carefully, or to > >>>>>>>> create, to > >>>>>>>> share, or to capture something important without explicitly > >>>>>>>> expressing it > >>>>>>>> will continue to write and produce original work. The people who > >>>>>>>> never were > >>>>>>>> writers will use AI to produce lots of text. > >>>>>>>>>>>>> I hope that our culture can remain one of intentional writing > >>>>>>>>>>>>> and > >>>>>>>> intentional engineering. I enjoy reading the voice of the author in > >>>>>>>> comments, code, and tickets in Cassandra – the different ways we use > >>>>>>>> language based on where we grew up and how we learned English, the > >>>>>>>> translated idioms from our various backgrounds, and terse comments > >>>>>>>> that > >>>>>>>> recognize the difference between code whose function is obvious and > >>>>>>>> what > >>>>>>>> warrants genuine exposition. When I read code in Cassandra, it’s a > >>>>>>>> delight > >>>>>>>> to recognize the author based on their writing style before flipping > >>>>>>>> on > >>>>>>>> `git annotate` to reveal the origin. > >>>>>>>>>>>>> I’d encourage folks to re-read the original proposal below. It > >>>>>>>>>>>>> is > >>>>>>>> very permissive. The guidance strikes me not just as reasonable, but > >>>>>>>> genuinely important to maintaining the health of the project. > >>>>>>>>>>>>> – Scott > >>>>>>>>>>>>> ===== > >>>>>>>>>>>>> Encouraged: > >>>>>>>>>>>>> - Reviewing and otherwise validating human-authored patches > >>>>>>>>>>>>> before > >>>>>>>> submission > >>>>>>>>>>>>> - Debugging, diagnosing etc > >>>>>>>>>>>>> Permitted: > >>>>>>>>>>>>> - Generating or modifying tests, scripts, tooling or any other > >>>>>>>> non-user facing changes > >>>>>>>>>>>>> - Minor changes to human-authored patches that are carefully > >>>>>>>> reviewed by the author > >>>>>>>>>>>>> Restricted: > >>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by > >>>>>>>>>>>>> contributors > >>>>>>>> with demonstrated expertise > >>>>>>>>>>>>> - Must have produced similar patches in size, scope and area > >>>>>>>> unassisted and with minimal third-party guidance > >>>>>>>>>>>>> - Core code changes made by LLM require an additional reviewer > >>>>>>>>>>>>> - LLM review is not a substitute for human review, and must be > >>>>>>>>>>>>> used > >>>>>>>> only to augment a complete and independent human understanding of > >>>>>>>> the patch. > >>>>>>>>>>>>> Prohibited: > >>>>>>>>>>>>> - All public prose must be human authored. This includes inline > >>>>>>>> comments, docs, posts to Jira etc. > >>>>>>>>>>>>> All LLM generated changes MUST be disclosed: > >>>>>>>>>>>>> - Outlined to any reviewer; > >>>>>>>>>>>>> - Summarised in the commit message; > >>>>>>>>>>>>> - Large blocks or files must be individually marked with some > >>>>>>>>>>>>> agreed > >>>>>>>> message like "created by <some AI>" > >>>>>>>>>>>>> ===== > >>>>>>>>>>>>>> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected] > >>>>>>>> <mailto:[email protected]>> wrote: > >>>>>>>>>>>>>> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected] > >>>>>>>> <mailto:[email protected]>> wrote: > >>>>>>>>>>>>>>> Restricted: > >>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by > >>>>>>>> contributors with demonstrated expertise > >>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area > >>>>>>>> unassisted and with minimal third-party guidance > >>>>>>>>>>>>>> I am -1 on this. This sounds like gate keeping attempt. It > >>>>>>>>>>>>>> narrowly > >>>>>>>> limits the pool to a few people on the project that have historically > >>>>>>>> contributed to certain parts of the codebase. This policy will > >>>>>>>> prohibit > >>>>>>>> skilled software engineers with domain expertise from proposing LLM > >>>>>>>> assisted changes simply because they have not contributed to the > >>>>>>>> project. > >>>>>>>> This is unrealistic and a net negative for the project to attract > >>>>>>>> talent > >>>>>>>> and grow our community. > >>>>>>>>>>>>>>> - Core code changes made by LLM require an additional reviewer > >>>>>>>>>>>>>> Can you be more precise what is this in addition to? How many > >>>>>>>>>>>>>> total > >>>>>>>> reviewers do you expect and what is the purpose of additional > >>>>>>>> reviewer? and > >>>>>>>> why? > >>>>>>>>>>>>>> Taking a step back - what are you trying to solve here? > >>>>>>>>>>>>>> Dinesh
