Meant "doing away with", sorry. Non-native speaker with a headache here. Thanks Caleb for spotting.
> On 24 Sep 2026, at 17:35, Aleksey Yeshchenko via dev > <[email protected]> wrote: > > P.S. I assume it's obvious from the text above that I don't believe that > getting away with human code review is a viable option. > On 24 Sep 2026, at 17:37, Štefan Miklošovič <[email protected]> wrote: > > Good call on checkerframework, we even have a patch for it. Work of > Jacek Lewandowski. We might just drive it to completion. Using AI for > finishing it would be quite ironic. > > (1) https://github.com/apache/cassandra/pull/2370 > > On Thu, Sep 24, 2026 at 6:19 PM Jon Haddad <[email protected]> wrote: >> >> There are some really good points being brought up about stability of the >> codebase, maintainability, quality of reviews, correctness bugs, and I agree >> with all of them. I think it would be helpful to take a step back and >> consider how those bugs got there in the first place, how they were fixed, >> and what we could do to further advance the codebase so they don't creep >> back. LLMs can be used either with very tight guardrails, or in what's >> effectively YOLO mode, and there's a big difference in the quality of the >> results you get. >> >> One thing to keep in mind, a lot of the initial code in C* was added without >> comprehensive testing. I hope we can all agree that it's a lot easier to >> break code that doesn't have high quality tests. During the code freeze, a >> lot of people people spent several years relentlessly finding and fixing >> bugs. This was probably a pretty frustrating time for anyone who was focused >> on fixing other people's bugs when they wanted to build features. I think >> we should recognize the effort here and appreciate the foundation that the >> project stands on now. I can understand how anyone involved with this effort >> would be apprehensive about seeing years of their life swept away by an >> agent that was driven by goal seeking to remove all the tests that it broke >> instead of fixing them. >> >> When I picked up the work to improve cursor compaction, the first thing I >> asked myself was how can I make sure I don't break this? How do I even know >> it works properly? There were some tricky parts to the code, and I really >> didn't want to come in and immediately break stuff. That's why I started >> with an entire patch dedicated to adding test infra to it. 90% of the patch >> was tests, and in my other cursor patches, it remains *at least* 80% of my >> patches. It was a *lot* faster to add almost 10K lines of tests that >> handled a byte for byte differential testing paired with harry to find over >> 30 bugs that caused cursor to corrupt results. Range tombstones alone were >> at least a dozen bugs, but I also found issues with static columns, reverse >> ordering, etc. Randomizing schemas and data in burn tests to generate >> different shapes of data, to ensure they all result in the same output at >> the end. JMH tests to ensure there weren't performance regressions, hours of >> profiling. These were all a *lot* easier to do with the LLM helping me out. >> In the process I've found bugs that have been lingering in the codebase for >> years. >> >> That's a long story, but hopefully we all agree that having comprehensive >> tests is a great way to ensure that both humans and LLMs don't break things >> that are working. >> >> The lesson: we need to keep improving our testing. Everything that we >> touch, should be left in a better state than how we found it with regard to >> test coverage. >> >> Test coverage isn't everything though, there's always little subtle bugs >> that don't get found in testing, that can slip in despite our best efforts. >> It's debatable if humans will be as good as agents for coding in the long >> term, for spotting small defects. I sincerely doubt it. For the time being >> though, we still have people involved. It's probably a good time to start >> using more static analysis tools to identify problematic code and to add >> this to CI. Dmitry had a suggestion recently for checkerframework to detect >> leaking contexts, a problem he spotted when reviewing my branch. It would >> be great to have that integrated into our CI and dev workflow so we can >> simply avoid an entire class of bugs. >> >> There's also PMD, which is excellent for finding code that can be hard to >> understand. I *highly* suggest you all run PMD to analyze for cognitive >> complexity and high npath scores. This was made popular by the folks at >> Sonar and I've found it to be an excellent feedback mechanism for >> structuring code. The default max they set is 15, which is the point where >> it starts to become difficult to verify something works without making a >> massive investment. We've got areas in the codebase that are in the >> hundreds, and some parts even higher. These have been contributed by >> humans, and are all high risk points for both humans and agents to start >> messing around with. They're also in some fairly critical areas that are >> very likely to break, so I understand why people would not want an agent >> anywhere near it. >> >> Unfortunately, it's not an easy problem to address. There's so many places >> where the code is structured in a way that has so many branches, so many >> conditions, that it's effectively impossible for a human to understand, >> creating a fear of messing around in it. There's plenty of areas that >> deserve extreme scrutiny, and we should be careful of what we add, whether >> it's human or agent. >> >> The codebase today requires a high degree of internal knowledge to navigate. >> There's land mines everywhere. We should be looking to make conscious >> improvements by moving the code forward, so it's easier to make changes to >> small, well tested components with minimal side effects. Not making it >> harder for people to use the tools that aid in that process. >> >> Here's what we could do to achieve the underlying goal of not breaking the >> DB: >> >> Add cognitive complexlity and npath via PMD as a feedback mechanism. >> >> Code that's hard to understand is hard to review. It's also hard to test. >> Let's break down the complex code so more people can contribute, safely. >> >> Add checkerframework to our tooling, >> >> Properly annotate the codebase for it and reduce the surface area that >> things can break. Less brittle codebase = we can move faster. >> >> Use jacoco to find areas of the codebase with poor testing. >> >> Let's improve the test coverage there, LLMs are great for this. We have a >> ton of static tests, these can become more dynamic, parameterized, and >> leverage harry. >> >> Refactor parts of the codebase that have high cognitive complexlity and >> NPath scores. >> >> This should be lowered over time to meet some high watermark, say 25 >> maximum, although I'd prefer 15 which is where the Sonar folks settled. >> >> Move forward moving the codebase to a more modular structure >> >> We've talked about Gradle on and off - but it can really be a huge help with >> incremental, modular builds. This is pretty easy to do with an agent and we >> could have it done in a couple days. >> >> Enforce boundaries with ArchUnit >> >> If we want to enforce certain code boundaries, this is the way to do it. >> Should not be part of manual review. >> >> Add LLM review for all incoming PRs before a human >> >> The goal here is to automate the initial part of the review process that >> reviewers should spot, and raise the bar for the initial contribution. When >> the code gets reviewed by a human, it should already have passed a large >> variety of initial checks. This should shorten the review cycle and result >> in higher quality patches. I've had Claude reviewing all my PRs in my >> personal projects for a while now and it consistently gives great feedback >> that I almost always incorporate. >> >> In my ideal world, we'd also auto-format all code >> >> Consistent formatting throughout the codebase would be amazing, but that's >> just one man's dream. >> >> Hopefully there's at least a couple things in this list we could move >> forward with in the short term, as it'll help improve the code quality >> regardless of how it's created. >> >> Jon >> >> https://checkerframework.org/manual/#aliasing-leaking-contexts >> https://www.sonarsource.com/docs/CognitiveComplexity.pdf >> https://pmd.github.io/pmd/pmd_rules_java_design.html >> >> >> >> >> >> On Thu, Sep 24, 2026 at 7:38 AM C. Scott Andreas <[email protected]> >> wrote: >>> >>> From Benedict: >>> >>> “I don't know if everyone remembers, but ten years ago Cassandra was full >>> of serious correctness and stability issues. Despite developing it, I would >>> not have run it myself or recommend that anyone use it. We have dug >>> ourselves out of that hole, but it took years of discipline and effort, and >>> we're still (deservedly) recovering our reputation.” >>> >>> Expanding on this point for those who may not have been active in the >>> project at this time — >>> >>> Apache Cassandra was fundamentally undeployable for four years between Nov >>> 2015 - 2019. The database literally lost data if you ran a read-only SELECT >>> query ordered descending (C-14513, C-14515). If you haven’t read these >>> tickets before, please take a moment to do so. >>> >>> It took years of careful work via property-based testing, fuzzing, and >>> deterministic simulation to restore Cassandra’s status as a usable system >>> of record. Once 14513 and 14515 were identified, nearly 30 additional >>> critical data loss and incorrect response bugs were identified. >>> >>> It is essential for the project’s future that we don’t regress to this >>> state chasing AI-generated features motivated by fear. The fact that >>> examples cited in this thread which boast shiny features but have critical >>> shortcomings unknown to their author supports this argument. >>> >>> The most common path for large corpuses of AI-generated software is elation >>> and reveling in a feature matrix, followed by abandonment. >>> >>> I endorse this point: >>> >>> “Let's use this new technology to improve the quality of our contributions, >>> not squander our hard-earned gains in the name of speed. It will be hard to >>> recover our reputation a second time.” >>> >>> Patrick, I don’t want your note regarding a TCM issue to go unaddressed. >>> Please file a Jira ticket and the patch if you like. I can’t comment on the >>> patch as I haven’t seen it, but together we will solve the problem. >>> >>> – Scott >>> >>>> On Sep 24, 2026, at 4:01 AM, Benedict Elliott Smith <[email protected]> >>>> wrote: >>>> >>>> Hi Patrick, >>>> >>>> As I mentioned in my reply to David, I would be happy to create a carve >>>> out for shallow and localised bug fixes in the "Permitted" section. Would >>>> this alleviate some of your concerns regarding your ability to contribute >>>> to the project? >>>> >>>> I appreciate your pointing out Ferrosa's Accord implementation however, as >>>> it is a *great* example of the problems we're leaping into. I took a look, >>>> and within about 30s found that the protocol is fundamentally incorrect, >>>> having failed to address CASSANDRA-18365. This is despite claiming to be >>>> tested with Jepsen that should in principle find this fault. I followed up >>>> by using Claude to interrogate the implementation further, and immediately >>>> found other serious correctness issues. >>>> >>>> I use LLMs daily now to help facilitate Accord development, and while they >>>> are powerful they are NOT able to author the code themselves, even when >>>> building upon a strong human-authored foundation. >>>> >>>> I don't know if everyone remembers, but ten years ago Cassandra was full >>>> of serious correctness and stability issues. Despite developing it, I >>>> would not have run it myself or recommend that anyone use it. We have dug >>>> ourselves out of that hole, but it took years of discipline and effort, >>>> and we're still (deservedly) recovering our reputation. >>>> >>>> Let's use this new technology to improve the quality of our contributions, >>>> not squander our hard-earned gains in the name of speed. It will be hard >>>> to recover our reputation a second time. >>>> >>>> >>>>> On 2026/09/23 19:16:17 Patrick McFadin wrote: >>>>> I was waiting for this moment to hit our project and I'm glad we're here. >>>>> I >>>>> am deeply concerned for our project and its future, as we have >>>>> increasingly >>>>> made it difficult to contribute. I had hoped that this new era of >>>>> software tools powered by AI would expand the project's reach and bring >>>>> more diverse thoughts and ideas. This policy proposal is the exact >>>>> opposite >>>>> of what we need. We have been sitting on a Cassandra 6 release alpha for >>>>> months. We need to accelerate and embrace new ways of being or be left >>>>> behind. As I read that policy, my first and gut level reactions: >>>>> - It comes across as elitist and class protectionism. Committer should not >>>>> be special but this proposal makes that designation even more sacred. >>>>> - It signals that our project is so fragile that only a few people "Really >>>>> understand it" That's some SQLite vibes right there. >>>>> - Trying to fix a problem that doesn't exist >>>>> Sadly, i think this policy change would also exclude a lot of comitters. >>>>> We aren't alone in this moment. The Linux project just went through >>>>> this. You can find the thread with a simple Google, but similar hard >>>>> feelings were being expressed "AI is going to ruin our project!", "The >>>>> unwashed masses are going to contribute terrible code!", "We have to >>>>> protect our precious status as Linux maintainers!" Linus being Linus, was >>>>> deeply invloved and they adopted a super simple statement that covers all >>>>> bases. Human or Human using AI. “You are expected to understand and to be >>>>> able to defend everything you submit.” Love that. >>>>> In the larger picture, I'll restate. I'm worried for our project. In late >>>>> 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real corner and >>>>> in the hands of somebody that knows how to build software, this tool is >>>>> like jet fuel. Here's some examples of new projects being hyper fueled by >>>>> AI coding tools. >>>>> Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity >>>>> Turso - Rust re-write of SQLite >>>>> Bun - Rust re-write of itself from Zig. >>>>> Think this couldn't happen to us? Already has: >>>>> https://github.com/ferrosadb/ferrosa. Ben is using it to power his own >>>>> startup, but it was him alone using a ton of local AI coding agents. He >>>>> even implemented Accord. Yeah... >>>>> The cracks are already starting to show. There is a black market economy >>>>> of >>>>> Cassandra patches happening now. Not going to name names or call people >>>>> out, but there are fixes and optimizations living in branches outside of >>>>> the Cassandra project. Why? I'll use myself as an example. I fixed a nasty >>>>> bug I ran into with TCM a few weeks ago. Wrote the tests. It passes CI and >>>>> lives in my personal branch. I'm sitting here really wondering if I want >>>>> to >>>>> go through the ritual humiliation of being roasted for using AI to fix it. >>>>> Me. I am worried about contrinuting code the Cassandra. What the hell does >>>>> that say? >>>>> I have my CQLite project that I've been doing a release around once a >>>>> month. I would love to donate that to the Cassandra project but I wouldn't >>>>> if it essentially killed any progress. >>>>> My larger counter proposal would be to: >>>>> - Adopt the “You are expected to understand and to be able to defend >>>>> everything you submit.” approach the Linux project has adopted. >>>>> - Loosen up the contributor process and our worry on trunk. Let 1000 >>>>> flowers bloom and bring it in. >>>>> - And finally, to give some people more peace of mind and open more doors, >>>>> adopt what other projects have done and provide more pluggability. Let new >>>>> ideas have an easy place to connect. >>>>> We are at a fork in the road. What are we going to do? And then I have to >>>>> ask myself, what am I going to do as a contributor? >>>>> Patrick >>>>> On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston <[email protected]> >>>>> wrote: >>>>>> I’m not necessarily opposed to having a policy, but so far we have some >>>>>> specific proposals addressing a problem statement that’s very nebulous. >>>>>> What is the community failing to do on its own that we’re trying to >>>>>> correct >>>>>> with policy? What outcomes are we trying to create or prevent? Having >>>>>> some >>>>>> examples and specific problems to discuss would help focus the >>>>>> conversation. >>>>>>> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote: >>>>>> Benedict, >>>>>> Thanks for clarifying. My concern still remains. This criteria would be >>>>>> difficult to define and apply consistently. What counts as “similar” >>>>>> scope >>>>>> or area, “mostly correct,” or sufficiently independent work? More >>>>>> importantly, how do we prevent such vague criteria from creating an >>>>>> informal hierarchy where some contributors work is routinely accepted >>>>>> while >>>>>> others is routinely rejected? >>>>>> If the intent is to limit AI-assisted code changes to Cassandra >>>>>> contributors, or to contributors who have previously worked in that >>>>>> component without AI, that would at least be clear and enforceable. >>>>>>> On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith < >>>>>> [email protected]> wrote: >>>>>>> Core code changes >>>>>>> Chris: Do you object to the first or second line you quote? Because the >>>>>> first line is effectively motivation for the second line, and can be >>>>>> removed (or more clearly combined). If it’s the second line, then I do >>>>>> not >>>>>> think this is an unreasonable expectation, and we can get into a proper >>>>>> debate about it. >>>>>>> Shailaja, since you only snipped the first sentence, your concerns might >>>>>> also be mostly answered by this clarification? “Minimal third-party >>>>>> guidance” implies you have some concerns about the second line, but all >>>>>> of >>>>>> our policies have some ambiguity because legalese is even worse. I don’t >>>>>> think the ambiguity here would be challenging to navigate though we can >>>>>> certainly refine it. This specific snippet is meant to convey an >>>>>> expectation that a contributor has autonomously produced patches of >>>>>> similar >>>>>> scope that were mostly correct, so that they have demonstrated the level >>>>>> of >>>>>> understanding necessary to guide another party to a successful patch >>>>>> (i.e. >>>>>> an LLM in this case). >>>>>>> On 2026/09/23 10:54:16 Benedict Elliott Smith wrote: >>>>>>>> Thanks everyone for your input so far. I’ll respond in brief to the >>>>>> main themes, in (mostly) separate emails so they can each have their own >>>>>> debate chain. >>>>>>>> Should we have a policy (Blake/Josh*/Jon/Dinesh) >>>>>>>> I think we would all agree that LLMs represent the biggest change to >>>>>> this community (and software more generally) since its inception, and we >>>>>> all now have enough experience with the technology to have formed >>>>>> opinions >>>>>> about how it is best managed. We also evidently have not all arrived at >>>>>> the >>>>>> same conclusions. In this situation, it would be an abdication of our >>>>>> responsibilities as a management committee to not agree *some* policy. >>>>>>>> I intend to conduct straw polls as the discussion evolves, so if you >>>>>> prefer an alternative policy - or modifications to this policy - I would >>>>>> encourage you to make those alternative proposals. >>>>>>>> *Veto/Consensus (Josh) >>>>>>>> It was fair to call out my poor use of language on this topic, so let >>>>>> me rephrase a little. The community is built on consensus, and work >>>>>> should >>>>>> not be merged when there are outstanding concerns to address. The >>>>>> explicit >>>>>> -1 should only be used rarely, because the prior expectation should >>>>>> prevent >>>>>> it ever being needed. I (and others) have outstanding concerns on LLM >>>>>> generated work that can only be addressed through this process right >>>>>> here, >>>>>> so to merge such work while maintaining the community’s consensus we must >>>>>> agree some policy. >>>>>>>> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote: >>>>>>>>> I am strongly -1 on this >>>>>>>>> - Core code changes made by LLM may only be proposed by contributors >>>>>> with demonstrated expertise >>>>>>>>> That creates a new, subjective privileged class of contributors and >>>>>> turns a tool choice into an eligibility test. Who decides whether >>>>>> expertise >>>>>> has been “demonstrated,” what counts as “minimal third-party guidance,” >>>>>> and >>>>>> how could those judgments be applied consistently or fairly? >>>>>>>>> Apache already has a better model, anyone may contribute, trust and >>>>>> additional repository privileges are earned transparently over time. The >>>>>> ASF describes its communities as flat, and says that newcomer ideas have >>>>>> as >>>>>> much input as those from original creators. We should not add a separate, >>>>>> informal hierarchy in which certain people may use common development >>>>>> tools >>>>>> while others may not. >>>>>>>>>> On Sep 23, 2026, at 6:33 AM, Chris Lohfink <[email protected]> >>>>>> wrote: >>>>>>>>>> - Core code changes made by LLM may only be proposed by contributors >>>>>> with demonstrated expertise >>>>>>>>>> - Must have produced similar patches in size, scope and area >>>>>> unassisted and with minimal third-party guidance >>>>>>>>>> I really don't like this one or its wording. Definitely too "the >>>>>> peasants are getting uppity lets build a wall". Lets not let a subjective >>>>>> thing like demonstrated expertise (who decides that?) be if it's ok or >>>>>> not. >>>>>> Hold the same standards for code quality and process for it all. I don't >>>>>> want this to be: only people on the storage team in Apple can use AI. >>>>>>>>>> Chris >>>>>>>>>> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto: >>>>>> [email protected]>> wrote: >>>>>>>>>>> I agree with Stefan and think this is both a reasonable and >>>>>> thoughtful proposal. >>>>>>>>>>> Here are some things I like about it: >>>>>>>>>>> – It outlines areas where LLM usage is unambiguously useful to the >>>>>> project’s developers and users. >>>>>>>>>>> – It defines a spectrum of recommendations and cautions. >>>>>>>>>>> – The only prohibited areas are extremely narrow and say nothing >>>>>> about code at all. >>>>>>>>>>> Some in this thread are responding as if this proposal seeks to >>>>>> prohibit or sharply limit use of LLMs. In fact, it’s one of the most open >>>>>> and welcoming I’ve seen for an OSS project of our size where many are >>>>>> adopting policies that simply ban them entirely. I’ve re-appended the >>>>>> proposal below my message as it seems to have been lost in threaded >>>>>> replies, and would encourage folks to give it a second read. >>>>>>>>>>> Some brief thoughts based on my own use of LLMs: >>>>>>>>>>> – I find them fantastically useful for reviewing and identifying >>>>>> problems that have slipped through review - primarily via Alex Petrov’s >>>>>> /deep-review skill, which I have running in a VM in a loop executing over >>>>>> every new commit in the project as of a few days ago. I will be posting a >>>>>> few hand-authored Jira tickets based on findings that appear legitimate >>>>>> to >>>>>> me. For now, the loop is posting them as issue drafts for my own review >>>>>> on >>>>>> my personal fork which you can find here: >>>>>> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug >>>>>>>>>>> – They’re great for enabling use of model checkers and formal >>>>>> methods where such work would have previously been prohibitively >>>>>> expensive, >>>>>> such as Blake’s work on a TLA+ proof of aspects of Mutation Tracking and >>>>>> Benedict/Fedor’s work on a machine-checkable proof of the Accord protocol >>>>>> in Lean. >>>>>>>>>>> – They are stunning for allowing me to experiment with ideas that >>>>>> would have otherwise been a summer internship’s scope of work. Some >>>>>> examples include an io_uring prototype, exploring the impact of >>>>>> page-aligned compressed chunk sizes, an API shim bridging the 3.x and 4.x >>>>>> Java Drivers, and potential enhancements to Zstandard. >>>>>>>>>>> – And they shine when given grunt-work that is critical to the >>>>>> project but a miserable labor for humans, such as triaging, reproducing, >>>>>> and root-causing flaky tests, which David Capwell now has running in a >>>>>> loop >>>>>> to help us improve CI stability in the project. >>>>>>>>>>> I never thought I’d be so positive on what’s possible via language >>>>>> models a year ago. At the same time, I also agree that they present >>>>>> challenges and risks that can be managed through thoughtful discussion >>>>>> and >>>>>> policy. Some of the concerns that I think are important to guard against >>>>>> include: >>>>>>>>>>> – Asymmetry of effort between author and reviewers: As >>>>>> token-generating machines, LLMs can generate diffs of extraordinary size >>>>>> very rapidly. /deep-review is great for chewing through diffs and >>>>>> identifying defects. But it should be used by the contributor themselves >>>>>> to >>>>>> identify issues – not to replace the role of the reviewer with more >>>>>> electricity. The role of the reviewers extends beyond identifying and >>>>>> highlighting defects. It encompasses architecture, harmony with the >>>>>> existing codebase, thinking ahead to future evolution of the project, and >>>>>> replicates context on the project as new code is committed. These >>>>>> functions >>>>>> cannot be automated away. >>>>>>>>>>> – Hesitancy of authors to engage manually with code they have >>>>>> generated: This is not specific to Cassandra, but it is a behavior that I >>>>>> have seen in several “highly-electric” projects. There’s a bimodal >>>>>> tendency >>>>>> toward code that is entirely generated or entirely human-authored - but >>>>>> it >>>>>> is rare for someone to prepare an AI-authored patch to take an offramp >>>>>> and >>>>>> spend a significant amount of time refining the work by hand in an IDE. >>>>>> This hesitancy toward human participation in authorship of LLM-generated >>>>>> code is very concerning to me. >>>>>>>>>>> – Harmony with the existing codebase: Due to the tunnel-vision of >>>>>> context windows, LLMs are generally unaware of conventions and norms >>>>>> present in codebases and very frequently reinvent concepts in a >>>>>> generation >>>>>> turn to suit a goal without view of the project’s overall architecture. >>>>>> This results in a profusion of messy and duplicated concepts that >>>>>> gradually >>>>>> sprawl about a codebase. >>>>>>>>>>> Again, none of these are grounds for prohibition of usage of >>>>>> language models in developing the project. They’re just problems we need >>>>>> to >>>>>> bear in mind and guard against – and I think the proposal is designed to >>>>>> do >>>>>> just that. >>>>>>>>>>> I’m thrilled by the potential of LLMs to improve Apache Cassandra >>>>>> and we already see it happening through a vast number of issues that are >>>>>> being reported and fixed. But there’s also danger in taking ATVs down a >>>>>> hiking trail full of people. >>>>>>>>>>> Regarding the prohibition on prose, I’ll simply say: I recently >>>>>> found myself in a scenario where I found a Claude-authored document so >>>>>> inscrutable that I piped it back into a model, directed it to rewrite it >>>>>> in >>>>>> ASD-STE100, read it myself, and responded based on the summarization. As >>>>>> a >>>>>> humanities grad, this is probably the worst language crime I have >>>>>> committed. But it was in response to language that was itself so >>>>>> idiosyncratic that it was unreadable to me in its original form. I hope >>>>>> this never happens in the Apache Cassandra project. >>>>>>>>>>> I’ll close with a quote from an excellent article written by Colin >>>>>> Breck, an engineer who works on large-scale data systems: >>>>>> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/ >>>>>>>>>>> Colin wrote: >>>>>>>>>>>> I don’t want to live in a world where you use AI to summarize >>>>>> something important into unreadable text, and then I use AI in an attempt >>>>>> to decipher it. I want to hear you, imperfections and all. I want your >>>>>> interpretation of aesthetics, beauty, quality, relationship, time. I want >>>>>> to know how you feel. I want you to cut through and tell me what really >>>>>> matters. >>>>>>>>>>>> Intentional writing will likely become more valuable. People who >>>>>> write, and write to think, to think deeply and carefully, or to create, >>>>>> to >>>>>> share, or to capture something important without explicitly expressing it >>>>>> will continue to write and produce original work. The people who never >>>>>> were >>>>>> writers will use AI to produce lots of text. >>>>>>>>>>> I hope that our culture can remain one of intentional writing and >>>>>> intentional engineering. I enjoy reading the voice of the author in >>>>>> comments, code, and tickets in Cassandra – the different ways we use >>>>>> language based on where we grew up and how we learned English, the >>>>>> translated idioms from our various backgrounds, and terse comments that >>>>>> recognize the difference between code whose function is obvious and what >>>>>> warrants genuine exposition. When I read code in Cassandra, it’s a >>>>>> delight >>>>>> to recognize the author based on their writing style before flipping on >>>>>> `git annotate` to reveal the origin. >>>>>>>>>>> I’d encourage folks to re-read the original proposal below. It is >>>>>> very permissive. The guidance strikes me not just as reasonable, but >>>>>> genuinely important to maintaining the health of the project. >>>>>>>>>>> – Scott >>>>>>>>>>> ===== >>>>>>>>>>> Encouraged: >>>>>>>>>>> - Reviewing and otherwise validating human-authored patches before >>>>>> submission >>>>>>>>>>> - Debugging, diagnosing etc >>>>>>>>>>> Permitted: >>>>>>>>>>> - Generating or modifying tests, scripts, tooling or any other >>>>>> non-user facing changes >>>>>>>>>>> - Minor changes to human-authored patches that are carefully >>>>>> reviewed by the author >>>>>>>>>>> Restricted: >>>>>>>>>>> - Core code changes made by LLM may only be proposed by contributors >>>>>> with demonstrated expertise >>>>>>>>>>> - Must have produced similar patches in size, scope and area >>>>>> unassisted and with minimal third-party guidance >>>>>>>>>>> - Core code changes made by LLM require an additional reviewer >>>>>>>>>>> - LLM review is not a substitute for human review, and must be used >>>>>> only to augment a complete and independent human understanding of the >>>>>> patch. >>>>>>>>>>> Prohibited: >>>>>>>>>>> - All public prose must be human authored. This includes inline >>>>>> comments, docs, posts to Jira etc. >>>>>>>>>>> All LLM generated changes MUST be disclosed: >>>>>>>>>>> - Outlined to any reviewer; >>>>>>>>>>> - Summarised in the commit message; >>>>>>>>>>> - Large blocks or files must be individually marked with some agreed >>>>>> message like "created by <some AI>" >>>>>>>>>>> ===== >>>>>>>>>>>> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected] >>>>>> <mailto:[email protected]>> wrote: >>>>>>>>>>>> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected] >>>>>> <mailto:[email protected]>> wrote: >>>>>>>>>>>>> Restricted: >>>>>>>>>>>>> - Core code changes made by LLM may only be proposed by >>>>>> contributors with demonstrated expertise >>>>>>>>>>>>> - Must have produced similar patches in size, scope and area >>>>>> unassisted and with minimal third-party guidance >>>>>>>>>>>> I am -1 on this. This sounds like gate keeping attempt. It narrowly >>>>>> limits the pool to a few people on the project that have historically >>>>>> contributed to certain parts of the codebase. This policy will prohibit >>>>>> skilled software engineers with domain expertise from proposing LLM >>>>>> assisted changes simply because they have not contributed to the project. >>>>>> This is unrealistic and a net negative for the project to attract talent >>>>>> and grow our community. >>>>>>>>>>>>> - Core code changes made by LLM require an additional reviewer >>>>>>>>>>>> Can you be more precise what is this in addition to? How many total >>>>>> reviewers do you expect and what is the purpose of additional reviewer? >>>>>> and >>>>>> why? >>>>>>>>>>>> Taking a step back - what are you trying to solve here? >>>>>>>>>>>> Dinesh
