Blake, I’d support those 3 guidelines. One hostile reaction to them would, I’m guessing, be that guideline number 1 is too weak, and that the burden of quality would continue to rest entirely on committer reviewers. I suppose my reply there would be that anyone who continually spams the project with patches they do not fully understand wound incur a pretty heavy reputational penalty and the problem would more or less sort itself out.
> On Sep 24, 2026, at 6:07 PM, Benedict Elliott Smith <[email protected]> > wrote: > > For the record Josh, I am including pre-3.0. I was involved in customer > support escalations for 1.2, 2.0 and 2.1: the software was shoddy, to put it > mildly. 8099 was a symptom, not the disease. > > The project had a culture of doing stuff without sufficient care or > consideration. This is by no means a simple matter of needing more testing or > less timeline pressure, nor is it easy to define a "bar" to ensure it does > not regress. If we had simple metrics, we would have settled it a long time > ago. > > You can still see customer scars crop up on forum discussions periodically. > > > On 2026/09/24 19:01:02 Josh McKenzie wrote: >>> Apache Cassandra was fundamentally undeployable for four years between Nov >>> 2015 - 2019. >> CASSANDRA-8099 was a maximal manifestation of a specific approach to >> engineering and calendar constraints we've seen time and again on the >> project; I don't want us to conflate things here. That was a herculean >> monolithic body of work performed in inhuman conditions (in vim!) that was >> ultimately so invasive, all the unit tests in the code-base were commented >> out and Jake and I spent a grueling 1.5-2 months hand-rewriting basically >> all the unit tests in that code-base to get things to even build and run, >> much less pass. >> >> Massive blast radius changes that are un-sustainably complex, under-tested, >> where we don't property, fuzz, check coverage, check complexity, or A/B >> compare against a known good system (i.e. pre/post) correctness testing are >> going to destabilize the database at any time under any regime of tooling. >> We certainly could speed-run our way back into destabilization with LLM's as >> they are a force multiplier for both the good and the bad of one's >> engineering practices, but that's a solvable problem by better defining what >> our bars of quality are (Definition of Done anyone?) and holding ourselves >> accountable to delivering at that bar. >> >> On Thu, Sep 24, 2026, at 2:40 PM, David Capwell via dev wrote: >>>> >>>> >>>> if you really want to pursue it I would ask that we do it offline to avoid >>>> polluting an already busy conversation >>>> >>> People are directly responding saying that they feel discrimination >>> currently and that the policy tries to codify that discrimination, so I >>> feel its 100% on topic. This thread has presented 0 evidence that LLM usage >>> has lowered the quality of contributions merged and has so far been vibes >>> and feeling; I have yet to see any evidence to justify such discrimination >>> so I will keep pushing back until such evidence is presented so we can have >>> a informed debate. >>> >>>> namely how we handle shallow and localised bug fixes. I would be happy >>>> adding a clear entry to the “Permitted” section for this. No doubt there >>>> are many refinements needed to the Restricted text as well, that might >>>> also capture some of your concerns. >>>> >>> I will not sign off on cherry picking areas where "safe" to use a tool; so >>> no it does not capture my concerns. >>> >>> >>> >>>> I don't know if everyone remembers, but ten years ago Cassandra was full >>>> of serious correctness and stability issues. Despite developing it, I >>>> would not have run it myself or recommend that anyone use it. We have dug >>>> ourselves out of that hole, but it took years of discipline and effort, >>>> and we're still (deservedly) recovering our reputation. >>>> >>>> Let's use this new technology to improve the quality of our contributions, >>>> not squander our hard-earned gains in the name of speed. It will be hard >>>> to recover our reputation a second time. >>>> >>> I do recall the 3.x line and put in a significant amount of effort to >>> harden it. There were behaviors I noticed after joining Cassandra that I >>> feel directly contributed to 3.0 and the decade of catch up; behaviors that >>> still linger in parts of the community today. >>> >>> As I look on trunk and look at committed code and trace back to PRs and >>> JIRA I see the following: >>> >>> • large patches approved without comments >>> • 0 evidence that tests were run >>> I then look at our CI and see tests failing for months. As you start to >>> triage you start to see some of them show real issues; yet they linger for >>> months not being addressed... when CI is unstable it takes a lot of effort >>> to triage "did my patch break the test", and I have seen time and time >>> again people do not put in that effort, and shrug off as "its just a flakey >>> test"; then our CI failure rate grows. >>> >>> Non of this has anything to do with LLMs but LLMs running in this >>> environment is far more dangerous as there are not checks in place to "hold >>> the bar". I am all for raising the bar universally; expecting both humans >>> and LLMs to match that bar. >>> >>>> `I’d support something that boils down to roughly this: >>>> >>>> 1.) 2 committers must understand an LLM-assisted change before it commits. >>>> (Perhaps separately we can explore the question of why we haven’t added >>>> any new committers to the core project for about a year. I’m also still >>>> not entirely sure if it’s acceptable within our guidelines for a committer >>>> to +1 a patch after delegating review.) >>>> >>>> 2.) Patch authors must demonstrate enough understanding to discuss their >>>> own patch, whether or not parts of it are generated by an LLM. >>>> >>>> 3.) The “Assisted-by” tag should be used to indicate any non-trivial LLM >>>> usage in the generation of a patch, just like we have used Co-authored-by >>>> historically. >>>> >>>> 4.) Comments and other things that aren't the actual code (but could sow >>>> confusion) should be held to the same standard we'd expect from a human >>>> writer. If we don’t yet agree on that standard, we can formalize enough of >>>> it to guide both humans and LLMs. >> ` >>> I can get behind this proposal but i would tweak it as 1/2 i don't think >>> really need to special case LLM usage >>> >>> 1. 2 committers must understand the change before it commits. >>> 2. Patch authors must demonstrate enough understanding to discuss their own >>> patch >>> Nothing about those 2 need to be scoped to LLM usage and honestly matches >>> most PMCs I have talked to understanding of our bar (as Benedict pointed >>> out, the actual wording could be interpreted to allow rubber stamping from >>> committer) >>> >>> As for 3 I am cool with this. ASF recommends the same (it says >>> `Generated-by` but that discount's the human's effort) as its useful for >>> audits and tooling. Having `Assisted-by` tag should not imply anything >>> about the committed patch as it should have gone through the same bar we >>> all expect; its just for tools auditing. >>> >>> >>>> On Sep 24, 2026, at 10:37 AM, Aleksey Yeshchenko via dev >>>> <[email protected]> wrote: >>>> >>>> Meant "doing away with", sorry. Non-native speaker with a headache here. >>>> Thanks Caleb for spotting. >>>> >>>>> On 24 Sep 2026, at 17:35, Aleksey Yeshchenko via dev >>>>> <[email protected]> wrote: >>>>> >>>>> P.S. I assume it's obvious from the text above that I don't believe that >>>>> getting away with human code review is a viable option. >>>> >>>> >>>>> On 24 Sep 2026, at 17:37, Štefan Miklošovič <[email protected]> >>>>> wrote: >>>>> >>>>> Good call on checkerframework, we even have a patch for it. Work of >>>>> Jacek Lewandowski. We might just drive it to completion. Using AI for >>>>> finishing it would be quite ironic. >>>>> >>>>> (1) https://github.com/apache/cassandra/pull/2370 >>>>> >>>>> On Thu, Sep 24, 2026 at 6:19 PM Jon Haddad <[email protected]> >>>>> wrote: >>>>>> >>>>>> There are some really good points being brought up about stability of >>>>>> the codebase, maintainability, quality of reviews, correctness bugs, and >>>>>> I agree with all of them. I think it would be helpful to take a step >>>>>> back and consider how those bugs got there in the first place, how they >>>>>> were fixed, and what we could do to further advance the codebase so they >>>>>> don't creep back. LLMs can be used either with very tight guardrails, >>>>>> or in what's effectively YOLO mode, and there's a big difference in the >>>>>> quality of the results you get. >>>>>> >>>>>> One thing to keep in mind, a lot of the initial code in C* was added >>>>>> without comprehensive testing. I hope we can all agree that it's a lot >>>>>> easier to break code that doesn't have high quality tests. During the >>>>>> code freeze, a lot of people people spent several years relentlessly >>>>>> finding and fixing bugs. This was probably a pretty frustrating time for >>>>>> anyone who was focused on fixing other people's bugs when they wanted to >>>>>> build features. I think we should recognize the effort here and >>>>>> appreciate the foundation that the project stands on now. I can >>>>>> understand how anyone involved with this effort would be apprehensive >>>>>> about seeing years of their life swept away by an agent that was driven >>>>>> by goal seeking to remove all the tests that it broke instead of fixing >>>>>> them. >>>>>> >>>>>> When I picked up the work to improve cursor compaction, the first thing >>>>>> I asked myself was how can I make sure I don't break this? How do I >>>>>> even know it works properly? There were some tricky parts to the code, >>>>>> and I really didn't want to come in and immediately break stuff. That's >>>>>> why I started with an entire patch dedicated to adding test infra to it. >>>>>> 90% of the patch was tests, and in my other cursor patches, it remains >>>>>> *at least* 80% of my patches. It was a *lot* faster to add almost 10K >>>>>> lines of tests that handled a byte for byte differential testing paired >>>>>> with harry to find over 30 bugs that caused cursor to corrupt results. >>>>>> Range tombstones alone were at least a dozen bugs, but I also found >>>>>> issues with static columns, reverse ordering, etc. Randomizing schemas >>>>>> and data in burn tests to generate different shapes of data, to ensure >>>>>> they all result in the same output at the end. JMH tests to ensure there >>>>>> weren't performance regressions, hours of profiling. These were all a >>>>>> *lot* easier to do with the LLM helping me out. In the process I've >>>>>> found bugs that have been lingering in the codebase for years. >>>>>> >>>>>> That's a long story, but hopefully we all agree that having >>>>>> comprehensive tests is a great way to ensure that both humans and LLMs >>>>>> don't break things that are working. >>>>>> >>>>>> The lesson: we need to keep improving our testing. Everything that we >>>>>> touch, should be left in a better state than how we found it with regard >>>>>> to test coverage. >>>>>> >>>>>> Test coverage isn't everything though, there's always little subtle bugs >>>>>> that don't get found in testing, that can slip in despite our best >>>>>> efforts. It's debatable if humans will be as good as agents for coding >>>>>> in the long term, for spotting small defects. I sincerely doubt it. >>>>>> For the time being though, we still have people involved. It's probably >>>>>> a good time to start using more static analysis tools to identify >>>>>> problematic code and to add this to CI. Dmitry had a suggestion >>>>>> recently for checkerframework to detect leaking contexts, a problem he >>>>>> spotted when reviewing my branch. It would be great to have that >>>>>> integrated into our CI and dev workflow so we can simply avoid an entire >>>>>> class of bugs. >>>>>> >>>>>> There's also PMD, which is excellent for finding code that can be hard >>>>>> to understand. I *highly* suggest you all run PMD to analyze for >>>>>> cognitive complexity and high npath scores. This was made popular by >>>>>> the folks at Sonar and I've found it to be an excellent feedback >>>>>> mechanism for structuring code. The default max they set is 15, which >>>>>> is the point where it starts to become difficult to verify something >>>>>> works without making a massive investment. We've got areas in the >>>>>> codebase that are in the hundreds, and some parts even higher. These >>>>>> have been contributed by humans, and are all high risk points for both >>>>>> humans and agents to start messing around with. They're also in some >>>>>> fairly critical areas that are very likely to break, so I understand why >>>>>> people would not want an agent anywhere near it. >>>>>> >>>>>> Unfortunately, it's not an easy problem to address. There's so many >>>>>> places where the code is structured in a way that has so many branches, >>>>>> so many conditions, that it's effectively impossible for a human to >>>>>> understand, creating a fear of messing around in it. There's plenty of >>>>>> areas that deserve extreme scrutiny, and we should be careful of what we >>>>>> add, whether it's human or agent. >>>>>> >>>>>> The codebase today requires a high degree of internal knowledge to >>>>>> navigate. There's land mines everywhere. We should be looking to make >>>>>> conscious improvements by moving the code forward, so it's easier to >>>>>> make changes to small, well tested components with minimal side effects. >>>>>> Not making it harder for people to use the tools that aid in that >>>>>> process. >>>>>> >>>>>> Here's what we could do to achieve the underlying goal of not breaking >>>>>> the DB: >>>>>> >>>>>> Add cognitive complexlity and npath via PMD as a feedback mechanism. >>>>>> >>>>>> Code that's hard to understand is hard to review. It's also hard to >>>>>> test. Let's break down the complex code so more people can contribute, >>>>>> safely. >>>>>> >>>>>> Add checkerframework to our tooling, >>>>>> >>>>>> Properly annotate the codebase for it and reduce the surface area that >>>>>> things can break. Less brittle codebase = we can move faster. >>>>>> >>>>>> Use jacoco to find areas of the codebase with poor testing. >>>>>> >>>>>> Let's improve the test coverage there, LLMs are great for this. We have >>>>>> a ton of static tests, these can become more dynamic, parameterized, and >>>>>> leverage harry. >>>>>> >>>>>> Refactor parts of the codebase that have high cognitive complexlity and >>>>>> NPath scores. >>>>>> >>>>>> This should be lowered over time to meet some high watermark, say 25 >>>>>> maximum, although I'd prefer 15 which is where the Sonar folks settled. >>>>>> >>>>>> Move forward moving the codebase to a more modular structure >>>>>> >>>>>> We've talked about Gradle on and off - but it can really be a huge help >>>>>> with incremental, modular builds. This is pretty easy to do with an >>>>>> agent and we could have it done in a couple days. >>>>>> >>>>>> Enforce boundaries with ArchUnit >>>>>> >>>>>> If we want to enforce certain code boundaries, this is the way to do it. >>>>>> Should not be part of manual review. >>>>>> >>>>>> Add LLM review for all incoming PRs before a human >>>>>> >>>>>> The goal here is to automate the initial part of the review process that >>>>>> reviewers should spot, and raise the bar for the initial contribution. >>>>>> When the code gets reviewed by a human, it should already have passed a >>>>>> large variety of initial checks. This should shorten the review cycle >>>>>> and result in higher quality patches. I've had Claude reviewing all my >>>>>> PRs in my personal projects for a while now and it consistently gives >>>>>> great feedback that I almost always incorporate. >>>>>> >>>>>> In my ideal world, we'd also auto-format all code >>>>>> >>>>>> Consistent formatting throughout the codebase would be amazing, but >>>>>> that's just one man's dream. >>>>>> >>>>>> Hopefully there's at least a couple things in this list we could move >>>>>> forward with in the short term, as it'll help improve the code quality >>>>>> regardless of how it's created. >>>>>> >>>>>> Jon >>>>>> >>>>>> https://checkerframework.org/manual/#aliasing-leaking-contexts >>>>>> https://www.sonarsource.com/docs/CognitiveComplexity.pdf >>>>>> https://pmd.github.io/pmd/pmd_rules_java_design.html >>>>>> >>>>>> >>>>>> >>>>>> >>>>>> >>>>>> On Thu, Sep 24, 2026 at 7:38 AM C. Scott Andreas <[email protected]> >>>>>> wrote: >>>>>>> >>>>>>> From Benedict: >>>>>>> >>>>>>> “I don't know if everyone remembers, but ten years ago Cassandra was >>>>>>> full of serious correctness and stability issues. Despite developing >>>>>>> it, I would not have run it myself or recommend that anyone use it. We >>>>>>> have dug ourselves out of that hole, but it took years of discipline >>>>>>> and effort, and we're still (deservedly) recovering our reputation.” >>>>>>> >>>>>>> Expanding on this point for those who may not have been active in the >>>>>>> project at this time — >>>>>>> >>>>>>> Apache Cassandra was fundamentally undeployable for four years between >>>>>>> Nov 2015 - 2019. The database literally lost data if you ran a >>>>>>> read-only SELECT query ordered descending (C-14513, C-14515). If you >>>>>>> haven’t read these tickets before, please take a moment to do so. >>>>>>> >>>>>>> It took years of careful work via property-based testing, fuzzing, and >>>>>>> deterministic simulation to restore Cassandra’s status as a usable >>>>>>> system of record. Once 14513 and 14515 were identified, nearly 30 >>>>>>> additional critical data loss and incorrect response bugs were >>>>>>> identified. >>>>>>> >>>>>>> It is essential for the project’s future that we don’t regress to this >>>>>>> state chasing AI-generated features motivated by fear. The fact that >>>>>>> examples cited in this thread which boast shiny features but have >>>>>>> critical shortcomings unknown to their author supports this argument. >>>>>>> >>>>>>> The most common path for large corpuses of AI-generated software is >>>>>>> elation and reveling in a feature matrix, followed by abandonment. >>>>>>> >>>>>>> I endorse this point: >>>>>>> >>>>>>> “Let's use this new technology to improve the quality of our >>>>>>> contributions, not squander our hard-earned gains in the name of speed. >>>>>>> It will be hard to recover our reputation a second time.” >>>>>>> >>>>>>> Patrick, I don’t want your note regarding a TCM issue to go >>>>>>> unaddressed. Please file a Jira ticket and the patch if you like. I >>>>>>> can’t comment on the patch as I haven’t seen it, but together we will >>>>>>> solve the problem. >>>>>>> >>>>>>> – Scott >>>>>>> >>>>>>>> On Sep 24, 2026, at 4:01 AM, Benedict Elliott Smith >>>>>>>> <[email protected]> wrote: >>>>>>>> >>>>>>>> Hi Patrick, >>>>>>>> >>>>>>>> As I mentioned in my reply to David, I would be happy to create a >>>>>>>> carve out for shallow and localised bug fixes in the "Permitted" >>>>>>>> section. Would this alleviate some of your concerns regarding your >>>>>>>> ability to contribute to the project? >>>>>>>> >>>>>>>> I appreciate your pointing out Ferrosa's Accord implementation >>>>>>>> however, as it is a *great* example of the problems we're leaping >>>>>>>> into. I took a look, and within about 30s found that the protocol is >>>>>>>> fundamentally incorrect, having failed to address CASSANDRA-18365. >>>>>>>> This is despite claiming to be tested with Jepsen that should in >>>>>>>> principle find this fault. I followed up by using Claude to >>>>>>>> interrogate the implementation further, and immediately found other >>>>>>>> serious correctness issues. >>>>>>>> >>>>>>>> I use LLMs daily now to help facilitate Accord development, and while >>>>>>>> they are powerful they are NOT able to author the code themselves, >>>>>>>> even when building upon a strong human-authored foundation. >>>>>>>> >>>>>>>> I don't know if everyone remembers, but ten years ago Cassandra was >>>>>>>> full of serious correctness and stability issues. Despite developing >>>>>>>> it, I would not have run it myself or recommend that anyone use it. We >>>>>>>> have dug ourselves out of that hole, but it took years of discipline >>>>>>>> and effort, and we're still (deservedly) recovering our reputation. >>>>>>>> >>>>>>>> Let's use this new technology to improve the quality of our >>>>>>>> contributions, not squander our hard-earned gains in the name of >>>>>>>> speed. It will be hard to recover our reputation a second time. >>>>>>>> >>>>>>>> >>>>>>>>> On 2026/09/23 19:16:17 Patrick McFadin wrote: >>>>>>>>> I was waiting for this moment to hit our project and I'm glad we're >>>>>>>>> here. I >>>>>>>>> am deeply concerned for our project and its future, as we have >>>>>>>>> increasingly >>>>>>>>> made it difficult to contribute. I had hoped that this new era of >>>>>>>>> software tools powered by AI would expand the project's reach and >>>>>>>>> bring >>>>>>>>> more diverse thoughts and ideas. This policy proposal is the exact >>>>>>>>> opposite >>>>>>>>> of what we need. We have been sitting on a Cassandra 6 release alpha >>>>>>>>> for >>>>>>>>> months. We need to accelerate and embrace new ways of being or be left >>>>>>>>> behind. As I read that policy, my first and gut level reactions: >>>>>>>>> - It comes across as elitist and class protectionism. Committer >>>>>>>>> should not >>>>>>>>> be special but this proposal makes that designation even more sacred. >>>>>>>>> - It signals that our project is so fragile that only a few people >>>>>>>>> "Really >>>>>>>>> understand it" That's some SQLite vibes right there. >>>>>>>>> - Trying to fix a problem that doesn't exist >>>>>>>>> Sadly, i think this policy change would also exclude a lot of >>>>>>>>> comitters. >>>>>>>>> We aren't alone in this moment. The Linux project just went through >>>>>>>>> this. You can find the thread with a simple Google, but similar hard >>>>>>>>> feelings were being expressed "AI is going to ruin our project!", "The >>>>>>>>> unwashed masses are going to contribute terrible code!", "We have to >>>>>>>>> protect our precious status as Linux maintainers!" Linus being >>>>>>>>> Linus, was >>>>>>>>> deeply invloved and they adopted a super simple statement that covers >>>>>>>>> all >>>>>>>>> bases. Human or Human using AI. “You are expected to understand and >>>>>>>>> to be >>>>>>>>> able to defend everything you submit.” Love that. >>>>>>>>> In the larger picture, I'll restate. I'm worried for our project. In >>>>>>>>> late >>>>>>>>> 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real corner >>>>>>>>> and >>>>>>>>> in the hands of somebody that knows how to build software, this tool >>>>>>>>> is >>>>>>>>> like jet fuel. Here's some examples of new projects being hyper >>>>>>>>> fueled by >>>>>>>>> AI coding tools. >>>>>>>>> Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity >>>>>>>>> Turso - Rust re-write of SQLite >>>>>>>>> Bun - Rust re-write of itself from Zig. >>>>>>>>> Think this couldn't happen to us? Already has: >>>>>>>>> https://github.com/ferrosadb/ferrosa. Ben is using it to power his own >>>>>>>>> startup, but it was him alone using a ton of local AI coding agents. >>>>>>>>> He >>>>>>>>> even implemented Accord. Yeah... >>>>>>>>> The cracks are already starting to show. There is a black market >>>>>>>>> economy of >>>>>>>>> Cassandra patches happening now. Not going to name names or call >>>>>>>>> people >>>>>>>>> out, but there are fixes and optimizations living in branches >>>>>>>>> outside of >>>>>>>>> the Cassandra project. Why? I'll use myself as an example. I fixed a >>>>>>>>> nasty >>>>>>>>> bug I ran into with TCM a few weeks ago. Wrote the tests. It passes >>>>>>>>> CI and >>>>>>>>> lives in my personal branch. I'm sitting here really wondering if I >>>>>>>>> want to >>>>>>>>> go through the ritual humiliation of being roasted for using AI to >>>>>>>>> fix it. >>>>>>>>> Me. I am worried about contrinuting code the Cassandra. What the hell >>>>>>>>> does >>>>>>>>> that say? >>>>>>>>> I have my CQLite project that I've been doing a release around once a >>>>>>>>> month. I would love to donate that to the Cassandra project but I >>>>>>>>> wouldn't >>>>>>>>> if it essentially killed any progress. >>>>>>>>> My larger counter proposal would be to: >>>>>>>>> - Adopt the “You are expected to understand and to be able to defend >>>>>>>>> everything you submit.” approach the Linux project has adopted. >>>>>>>>> - Loosen up the contributor process and our worry on trunk. Let 1000 >>>>>>>>> flowers bloom and bring it in. >>>>>>>>> - And finally, to give some people more peace of mind and open more >>>>>>>>> doors, >>>>>>>>> adopt what other projects have done and provide more pluggability. >>>>>>>>> Let new >>>>>>>>> ideas have an easy place to connect. >>>>>>>>> We are at a fork in the road. What are we going to do? And then I >>>>>>>>> have to >>>>>>>>> ask myself, what am I going to do as a contributor? >>>>>>>>> Patrick >>>>>>>>> On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston <[email protected]> >>>>>>>>> wrote: >>>>>>>>>> I’m not necessarily opposed to having a policy, but so far we have >>>>>>>>>> some >>>>>>>>>> specific proposals addressing a problem statement that’s very >>>>>>>>>> nebulous. >>>>>>>>>> What is the community failing to do on its own that we’re trying to >>>>>>>>>> correct >>>>>>>>>> with policy? What outcomes are we trying to create or prevent? >>>>>>>>>> Having some >>>>>>>>>> examples and specific problems to discuss would help focus the >>>>>>>>>> conversation. >>>>>>>>>>> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote: >>>>>>>>>> Benedict, >>>>>>>>>> Thanks for clarifying. My concern still remains. This criteria would >>>>>>>>>> be >>>>>>>>>> difficult to define and apply consistently. What counts as “similar” >>>>>>>>>> scope >>>>>>>>>> or area, “mostly correct,” or sufficiently independent work? More >>>>>>>>>> importantly, how do we prevent such vague criteria from creating an >>>>>>>>>> informal hierarchy where some contributors work is routinely >>>>>>>>>> accepted while >>>>>>>>>> others is routinely rejected? >>>>>>>>>> If the intent is to limit AI-assisted code changes to Cassandra >>>>>>>>>> contributors, or to contributors who have previously worked in that >>>>>>>>>> component without AI, that would at least be clear and enforceable. >>>>>>>>>>> On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith < >>>>>>>>>> [email protected]> wrote: >>>>>>>>>>> Core code changes >>>>>>>>>>> Chris: Do you object to the first or second line you quote? Because >>>>>>>>>>> the >>>>>>>>>> first line is effectively motivation for the second line, and can be >>>>>>>>>> removed (or more clearly combined). If it’s the second line, then I >>>>>>>>>> do not >>>>>>>>>> think this is an unreasonable expectation, and we can get into a >>>>>>>>>> proper >>>>>>>>>> debate about it. >>>>>>>>>>> Shailaja, since you only snipped the first sentence, your concerns >>>>>>>>>>> might >>>>>>>>>> also be mostly answered by this clarification? “Minimal third-party >>>>>>>>>> guidance” implies you have some concerns about the second line, but >>>>>>>>>> all of >>>>>>>>>> our policies have some ambiguity because legalese is even worse. I >>>>>>>>>> don’t >>>>>>>>>> think the ambiguity here would be challenging to navigate though we >>>>>>>>>> can >>>>>>>>>> certainly refine it. This specific snippet is meant to convey an >>>>>>>>>> expectation that a contributor has autonomously produced patches of >>>>>>>>>> similar >>>>>>>>>> scope that were mostly correct, so that they have demonstrated the >>>>>>>>>> level of >>>>>>>>>> understanding necessary to guide another party to a successful patch >>>>>>>>>> (i.e. >>>>>>>>>> an LLM in this case). >>>>>>>>>>> On 2026/09/23 10:54:16 Benedict Elliott Smith wrote: >>>>>>>>>>>> Thanks everyone for your input so far. I’ll respond in brief to the >>>>>>>>>> main themes, in (mostly) separate emails so they can each have their >>>>>>>>>> own >>>>>>>>>> debate chain. >>>>>>>>>>>> Should we have a policy (Blake/Josh*/Jon/Dinesh) >>>>>>>>>>>> I think we would all agree that LLMs represent the biggest change >>>>>>>>>>>> to >>>>>>>>>> this community (and software more generally) since its inception, >>>>>>>>>> and we >>>>>>>>>> all now have enough experience with the technology to have formed >>>>>>>>>> opinions >>>>>>>>>> about how it is best managed. We also evidently have not all arrived >>>>>>>>>> at the >>>>>>>>>> same conclusions. In this situation, it would be an abdication of our >>>>>>>>>> responsibilities as a management committee to not agree *some* >>>>>>>>>> policy. >>>>>>>>>>>> I intend to conduct straw polls as the discussion evolves, so if >>>>>>>>>>>> you >>>>>>>>>> prefer an alternative policy - or modifications to this policy - I >>>>>>>>>> would >>>>>>>>>> encourage you to make those alternative proposals. >>>>>>>>>>>> *Veto/Consensus (Josh) >>>>>>>>>>>> It was fair to call out my poor use of language on this topic, so >>>>>>>>>>>> let >>>>>>>>>> me rephrase a little. The community is built on consensus, and work >>>>>>>>>> should >>>>>>>>>> not be merged when there are outstanding concerns to address. The >>>>>>>>>> explicit >>>>>>>>>> -1 should only be used rarely, because the prior expectation should >>>>>>>>>> prevent >>>>>>>>>> it ever being needed. I (and others) have outstanding concerns on LLM >>>>>>>>>> generated work that can only be addressed through this process right >>>>>>>>>> here, >>>>>>>>>> so to merge such work while maintaining the community’s consensus we >>>>>>>>>> must >>>>>>>>>> agree some policy. >>>>>>>>>>>> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote: >>>>>>>>>>>>> I am strongly -1 on this >>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by >>>>>>>>>>>>> contributors >>>>>>>>>> with demonstrated expertise >>>>>>>>>>>>> That creates a new, subjective privileged class of contributors >>>>>>>>>>>>> and >>>>>>>>>> turns a tool choice into an eligibility test. Who decides whether >>>>>>>>>> expertise >>>>>>>>>> has been “demonstrated,” what counts as “minimal third-party >>>>>>>>>> guidance,” and >>>>>>>>>> how could those judgments be applied consistently or fairly? >>>>>>>>>>>>> Apache already has a better model, anyone may contribute, trust >>>>>>>>>>>>> and >>>>>>>>>> additional repository privileges are earned transparently over time. >>>>>>>>>> The >>>>>>>>>> ASF describes its communities as flat, and says that newcomer ideas >>>>>>>>>> have as >>>>>>>>>> much input as those from original creators. We should not add a >>>>>>>>>> separate, >>>>>>>>>> informal hierarchy in which certain people may use common >>>>>>>>>> development tools >>>>>>>>>> while others may not. >>>>>>>>>>>>>> On Sep 23, 2026, at 6:33 AM, Chris Lohfink <[email protected]> >>>>>>>>>> wrote: >>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by >>>>>>>>>>>>>> contributors >>>>>>>>>> with demonstrated expertise >>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area >>>>>>>>>> unassisted and with minimal third-party guidance >>>>>>>>>>>>>> I really don't like this one or its wording. Definitely too "the >>>>>>>>>> peasants are getting uppity lets build a wall". Lets not let a >>>>>>>>>> subjective >>>>>>>>>> thing like demonstrated expertise (who decides that?) be if it's ok >>>>>>>>>> or not. >>>>>>>>>> Hold the same standards for code quality and process for it all. I >>>>>>>>>> don't >>>>>>>>>> want this to be: only people on the storage team in Apple can use AI. >>>>>>>>>>>>>> Chris >>>>>>>>>>>>>> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto: >>>>>>>>>> [email protected]>> wrote: >>>>>>>>>>>>>>> I agree with Stefan and think this is both a reasonable and >>>>>>>>>> thoughtful proposal. >>>>>>>>>>>>>>> Here are some things I like about it: >>>>>>>>>>>>>>> – It outlines areas where LLM usage is unambiguously useful to >>>>>>>>>>>>>>> the >>>>>>>>>> project’s developers and users. >>>>>>>>>>>>>>> – It defines a spectrum of recommendations and cautions. >>>>>>>>>>>>>>> – The only prohibited areas are extremely narrow and say nothing >>>>>>>>>> about code at all. >>>>>>>>>>>>>>> Some in this thread are responding as if this proposal seeks to >>>>>>>>>> prohibit or sharply limit use of LLMs. In fact, it’s one of the most >>>>>>>>>> open >>>>>>>>>> and welcoming I’ve seen for an OSS project of our size where many are >>>>>>>>>> adopting policies that simply ban them entirely. I’ve re-appended the >>>>>>>>>> proposal below my message as it seems to have been lost in threaded >>>>>>>>>> replies, and would encourage folks to give it a second read. >>>>>>>>>>>>>>> Some brief thoughts based on my own use of LLMs: >>>>>>>>>>>>>>> – I find them fantastically useful for reviewing and identifying >>>>>>>>>> problems that have slipped through review - primarily via Alex >>>>>>>>>> Petrov’s >>>>>>>>>> /deep-review skill, which I have running in a VM in a loop executing >>>>>>>>>> over >>>>>>>>>> every new commit in the project as of a few days ago. I will be >>>>>>>>>> posting a >>>>>>>>>> few hand-authored Jira tickets based on findings that appear >>>>>>>>>> legitimate to >>>>>>>>>> me. For now, the loop is posting them as issue drafts for my own >>>>>>>>>> review on >>>>>>>>>> my personal fork which you can find here: >>>>>>>>>> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug >>>>>>>>>>>>>>> – They’re great for enabling use of model checkers and formal >>>>>>>>>> methods where such work would have previously been prohibitively >>>>>>>>>> expensive, >>>>>>>>>> such as Blake’s work on a TLA+ proof of aspects of Mutation Tracking >>>>>>>>>> and >>>>>>>>>> Benedict/Fedor’s work on a machine-checkable proof of the Accord >>>>>>>>>> protocol >>>>>>>>>> in Lean. >>>>>>>>>>>>>>> – They are stunning for allowing me to experiment with ideas >>>>>>>>>>>>>>> that >>>>>>>>>> would have otherwise been a summer internship’s scope of work. Some >>>>>>>>>> examples include an io_uring prototype, exploring the impact of >>>>>>>>>> page-aligned compressed chunk sizes, an API shim bridging the 3.x >>>>>>>>>> and 4.x >>>>>>>>>> Java Drivers, and potential enhancements to Zstandard. >>>>>>>>>>>>>>> – And they shine when given grunt-work that is critical to the >>>>>>>>>> project but a miserable labor for humans, such as triaging, >>>>>>>>>> reproducing, >>>>>>>>>> and root-causing flaky tests, which David Capwell now has running in >>>>>>>>>> a loop >>>>>>>>>> to help us improve CI stability in the project. >>>>>>>>>>>>>>> I never thought I’d be so positive on what’s possible via >>>>>>>>>>>>>>> language >>>>>>>>>> models a year ago. At the same time, I also agree that they present >>>>>>>>>> challenges and risks that can be managed through thoughtful >>>>>>>>>> discussion and >>>>>>>>>> policy. Some of the concerns that I think are important to guard >>>>>>>>>> against >>>>>>>>>> include: >>>>>>>>>>>>>>> – Asymmetry of effort between author and reviewers: As >>>>>>>>>> token-generating machines, LLMs can generate diffs of extraordinary >>>>>>>>>> size >>>>>>>>>> very rapidly. /deep-review is great for chewing through diffs and >>>>>>>>>> identifying defects. But it should be used by the contributor >>>>>>>>>> themselves to >>>>>>>>>> identify issues – not to replace the role of the reviewer with more >>>>>>>>>> electricity. The role of the reviewers extends beyond identifying and >>>>>>>>>> highlighting defects. It encompasses architecture, harmony with the >>>>>>>>>> existing codebase, thinking ahead to future evolution of the >>>>>>>>>> project, and >>>>>>>>>> replicates context on the project as new code is committed. These >>>>>>>>>> functions >>>>>>>>>> cannot be automated away. >>>>>>>>>>>>>>> – Hesitancy of authors to engage manually with code they have >>>>>>>>>> generated: This is not specific to Cassandra, but it is a behavior >>>>>>>>>> that I >>>>>>>>>> have seen in several “highly-electric” projects. There’s a bimodal >>>>>>>>>> tendency >>>>>>>>>> toward code that is entirely generated or entirely human-authored - >>>>>>>>>> but it >>>>>>>>>> is rare for someone to prepare an AI-authored patch to take an >>>>>>>>>> offramp and >>>>>>>>>> spend a significant amount of time refining the work by hand in an >>>>>>>>>> IDE. >>>>>>>>>> This hesitancy toward human participation in authorship of >>>>>>>>>> LLM-generated >>>>>>>>>> code is very concerning to me. >>>>>>>>>>>>>>> – Harmony with the existing codebase: Due to the tunnel-vision >>>>>>>>>>>>>>> of >>>>>>>>>> context windows, LLMs are generally unaware of conventions and norms >>>>>>>>>> present in codebases and very frequently reinvent concepts in a >>>>>>>>>> generation >>>>>>>>>> turn to suit a goal without view of the project’s overall >>>>>>>>>> architecture. >>>>>>>>>> This results in a profusion of messy and duplicated concepts that >>>>>>>>>> gradually >>>>>>>>>> sprawl about a codebase. >>>>>>>>>>>>>>> Again, none of these are grounds for prohibition of usage of >>>>>>>>>> language models in developing the project. They’re just problems we >>>>>>>>>> need to >>>>>>>>>> bear in mind and guard against – and I think the proposal is >>>>>>>>>> designed to do >>>>>>>>>> just that. >>>>>>>>>>>>>>> I’m thrilled by the potential of LLMs to improve Apache >>>>>>>>>>>>>>> Cassandra >>>>>>>>>> and we already see it happening through a vast number of issues that >>>>>>>>>> are >>>>>>>>>> being reported and fixed. But there’s also danger in taking ATVs >>>>>>>>>> down a >>>>>>>>>> hiking trail full of people. >>>>>>>>>>>>>>> Regarding the prohibition on prose, I’ll simply say: I recently >>>>>>>>>> found myself in a scenario where I found a Claude-authored document >>>>>>>>>> so >>>>>>>>>> inscrutable that I piped it back into a model, directed it to >>>>>>>>>> rewrite it in >>>>>>>>>> ASD-STE100, read it myself, and responded based on the >>>>>>>>>> summarization. As a >>>>>>>>>> humanities grad, this is probably the worst language crime I have >>>>>>>>>> committed. But it was in response to language that was itself so >>>>>>>>>> idiosyncratic that it was unreadable to me in its original form. I >>>>>>>>>> hope >>>>>>>>>> this never happens in the Apache Cassandra project. >>>>>>>>>>>>>>> I’ll close with a quote from an excellent article written by >>>>>>>>>>>>>>> Colin >>>>>>>>>> Breck, an engineer who works on large-scale data systems: >>>>>>>>>> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/ >>>>>>>>>>>>>>> Colin wrote: >>>>>>>>>>>>>>>> I don’t want to live in a world where you use AI to summarize >>>>>>>>>> something important into unreadable text, and then I use AI in an >>>>>>>>>> attempt >>>>>>>>>> to decipher it. I want to hear you, imperfections and all. I want >>>>>>>>>> your >>>>>>>>>> interpretation of aesthetics, beauty, quality, relationship, time. I >>>>>>>>>> want >>>>>>>>>> to know how you feel. I want you to cut through and tell me what >>>>>>>>>> really >>>>>>>>>> matters. >>>>>>>>>>>>>>>> Intentional writing will likely become more valuable. People >>>>>>>>>>>>>>>> who >>>>>>>>>> write, and write to think, to think deeply and carefully, or to >>>>>>>>>> create, to >>>>>>>>>> share, or to capture something important without explicitly >>>>>>>>>> expressing it >>>>>>>>>> will continue to write and produce original work. The people who >>>>>>>>>> never were >>>>>>>>>> writers will use AI to produce lots of text. >>>>>>>>>>>>>>> I hope that our culture can remain one of intentional writing >>>>>>>>>>>>>>> and >>>>>>>>>> intentional engineering. I enjoy reading the voice of the author in >>>>>>>>>> comments, code, and tickets in Cassandra – the different ways we use >>>>>>>>>> language based on where we grew up and how we learned English, the >>>>>>>>>> translated idioms from our various backgrounds, and terse comments >>>>>>>>>> that >>>>>>>>>> recognize the difference between code whose function is obvious and >>>>>>>>>> what >>>>>>>>>> warrants genuine exposition. When I read code in Cassandra, it’s a >>>>>>>>>> delight >>>>>>>>>> to recognize the author based on their writing style before flipping >>>>>>>>>> on >>>>>>>>>> `git annotate` to reveal the origin. >>>>>>>>>>>>>>> I’d encourage folks to re-read the original proposal below. It >>>>>>>>>>>>>>> is >>>>>>>>>> very permissive. The guidance strikes me not just as reasonable, but >>>>>>>>>> genuinely important to maintaining the health of the project. >>>>>>>>>>>>>>> – Scott >>>>>>>>>>>>>>> ===== >>>>>>>>>>>>>>> Encouraged: >>>>>>>>>>>>>>> - Reviewing and otherwise validating human-authored patches >>>>>>>>>>>>>>> before >>>>>>>>>> submission >>>>>>>>>>>>>>> - Debugging, diagnosing etc >>>>>>>>>>>>>>> Permitted: >>>>>>>>>>>>>>> - Generating or modifying tests, scripts, tooling or any other >>>>>>>>>> non-user facing changes >>>>>>>>>>>>>>> - Minor changes to human-authored patches that are carefully >>>>>>>>>> reviewed by the author >>>>>>>>>>>>>>> Restricted: >>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by >>>>>>>>>>>>>>> contributors >>>>>>>>>> with demonstrated expertise >>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area >>>>>>>>>> unassisted and with minimal third-party guidance >>>>>>>>>>>>>>> - Core code changes made by LLM require an additional reviewer >>>>>>>>>>>>>>> - LLM review is not a substitute for human review, and must be >>>>>>>>>>>>>>> used >>>>>>>>>> only to augment a complete and independent human understanding of >>>>>>>>>> the patch. >>>>>>>>>>>>>>> Prohibited: >>>>>>>>>>>>>>> - All public prose must be human authored. This includes inline >>>>>>>>>> comments, docs, posts to Jira etc. >>>>>>>>>>>>>>> All LLM generated changes MUST be disclosed: >>>>>>>>>>>>>>> - Outlined to any reviewer; >>>>>>>>>>>>>>> - Summarised in the commit message; >>>>>>>>>>>>>>> - Large blocks or files must be individually marked with some >>>>>>>>>>>>>>> agreed >>>>>>>>>> message like "created by <some AI>" >>>>>>>>>>>>>>> ===== >>>>>>>>>>>>>>>> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected] >>>>>>>>>> <mailto:[email protected]>> wrote: >>>>>>>>>>>>>>>> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected] >>>>>>>>>> <mailto:[email protected]>> wrote: >>>>>>>>>>>>>>>>> Restricted: >>>>>>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by >>>>>>>>>> contributors with demonstrated expertise >>>>>>>>>>>>>>>>> - Must have produced similar patches in size, scope and area >>>>>>>>>> unassisted and with minimal third-party guidance >>>>>>>>>>>>>>>> I am -1 on this. This sounds like gate keeping attempt. It >>>>>>>>>>>>>>>> narrowly >>>>>>>>>> limits the pool to a few people on the project that have historically >>>>>>>>>> contributed to certain parts of the codebase. This policy will >>>>>>>>>> prohibit >>>>>>>>>> skilled software engineers with domain expertise from proposing LLM >>>>>>>>>> assisted changes simply because they have not contributed to the >>>>>>>>>> project. >>>>>>>>>> This is unrealistic and a net negative for the project to attract >>>>>>>>>> talent >>>>>>>>>> and grow our community. >>>>>>>>>>>>>>>>> - Core code changes made by LLM require an additional reviewer >>>>>>>>>>>>>>>> Can you be more precise what is this in addition to? How many >>>>>>>>>>>>>>>> total >>>>>>>>>> reviewers do you expect and what is the purpose of additional >>>>>>>>>> reviewer? and >>>>>>>>>> why? >>>>>>>>>>>>>>>> Taking a step back - what are you trying to solve here? >>>>>>>>>>>>>>>> Dinesh
