Good call on checkerframework, we even have a patch for it. Work of
Jacek Lewandowski. We might just drive it to completion. Using AI for
finishing it would be quite ironic.

(1) https://github.com/apache/cassandra/pull/2370

On Thu, Sep 24, 2026 at 6:19 PM Jon Haddad <[email protected]> wrote:
>
> There are some really good points being brought up about stability of the 
> codebase, maintainability, quality of reviews, correctness bugs, and I agree 
> with all of them.  I think it would be helpful to take a step back and 
> consider how those bugs got there in the first place, how they were fixed, 
> and what we could do to further advance the codebase so they don't creep 
> back.  LLMs can be used either with very tight guardrails, or in what's 
> effectively YOLO mode, and there's a big difference in the quality of the 
> results you get.
>
> One thing to keep in mind, a lot of the initial code in C* was added without 
> comprehensive testing.  I hope we can all agree that it's a lot easier to 
> break code that doesn't have high quality tests.  During the code freeze, a 
> lot of people people spent several years relentlessly finding and fixing 
> bugs. This was probably a pretty frustrating time for anyone who was focused 
> on fixing other people's bugs when they wanted to build features.  I think we 
> should recognize the effort here and appreciate the foundation that the 
> project stands on now. I can understand how anyone involved with this effort 
> would be apprehensive about seeing years of their life swept away by an agent 
> that was driven by goal seeking to remove all the tests that it broke instead 
> of fixing them.
>
> When I picked up the work to improve cursor compaction, the first thing I 
> asked myself was how can I make sure I don't break this?  How do I even know 
> it works properly?  There were some tricky parts to the code, and I really 
> didn't want to come in and immediately break stuff.  That's why I started 
> with an entire patch dedicated to adding test infra to it.  90% of the patch 
> was tests, and in my other cursor patches, it remains *at least* 80% of my 
> patches.  It was a *lot* faster to add almost 10K lines of tests that handled 
> a byte for byte differential testing paired with harry to find over 30 bugs 
> that caused cursor to corrupt results.  Range tombstones alone were at least 
> a dozen bugs, but I also found issues with static columns, reverse ordering, 
> etc.  Randomizing schemas and data in burn tests to generate different shapes 
> of data, to ensure they all result in the same output at the end. JMH tests 
> to ensure there weren't performance regressions, hours of profiling. These 
> were all a *lot* easier to do with the LLM helping me out.  In the process 
> I've found bugs that have been lingering in the codebase for years.
>
> That's a long story, but hopefully we all agree that having comprehensive 
> tests is a great way to ensure that both humans and LLMs don't break things 
> that are working.
>
> The lesson: we need to keep improving our testing.  Everything that we touch, 
> should be left in a better state than how we found it with regard to test 
> coverage.
>
> Test coverage isn't everything though, there's always little subtle bugs that 
> don't get found in testing, that can slip in despite our best efforts.  It's 
> debatable if humans will be as good as agents for coding in the long term, 
> for spotting small defects.  I sincerely doubt it.  For the time being 
> though, we still have people involved. It's probably a good time to start 
> using more static analysis tools to identify problematic code and to add this 
> to CI.  Dmitry had a suggestion recently for checkerframework to detect 
> leaking contexts, a problem he spotted when reviewing my branch.  It would be 
> great to have that integrated into our CI and dev workflow so we can simply 
> avoid an entire class of bugs.
>
> There's also PMD, which is excellent for finding code that can be hard to 
> understand.  I *highly* suggest you all run PMD to analyze for cognitive 
> complexity and high npath scores.  This was made popular by the folks at 
> Sonar and I've found it to be an excellent feedback mechanism for structuring 
> code.  The default max they set is 15, which is the point where it starts to 
> become difficult to verify something works without making a massive 
> investment.  We've got areas in the codebase that are in the hundreds, and 
> some parts even higher.  These have been contributed by humans, and are all 
> high risk points for both humans and agents to start messing around with.  
> They're also in some fairly critical areas that are very likely to break, so 
> I understand why people would not want an agent anywhere near it.
>
> Unfortunately, it's not an easy problem to address.  There's so many places 
> where the code is structured in a way that has so many branches, so many 
> conditions, that it's effectively impossible for a human to understand, 
> creating a fear of messing around in it.  There's plenty of areas that 
> deserve extreme scrutiny, and we should be careful of what we add, whether 
> it's human or agent.
>
> The codebase today requires a high degree of internal knowledge to navigate.  
> There's land mines everywhere. We should be looking to make conscious 
> improvements by moving the code forward, so it's easier to make changes to 
> small, well tested components with minimal side effects.  Not making it 
> harder for people to use the tools that aid in that process.
>
> Here's what we could do to achieve the underlying goal of not breaking the DB:
>
> Add cognitive complexlity and npath via PMD as a feedback mechanism.
>
> Code that's hard to understand is hard to review.  It's also hard to test. 
> Let's break down the complex code so more people can contribute, safely.
>
> Add checkerframework to our tooling,
>
> Properly annotate the codebase for it and reduce the surface area that things 
> can break.  Less brittle codebase = we can move faster.
>
> Use jacoco to find areas of the codebase with poor testing.
>
> Let's improve the test coverage there, LLMs are great for this.  We have a 
> ton of static tests, these can become more dynamic, parameterized, and 
> leverage harry.
>
> Refactor parts of the codebase that have high cognitive complexlity and NPath 
> scores.
>
> This should be lowered over time to meet some high watermark, say 25 maximum, 
> although I'd prefer 15 which is where the Sonar folks settled.
>
> Move forward moving the codebase to a more modular structure
>
> We've talked about Gradle on and off - but it can really be a huge help with 
> incremental, modular builds. This is pretty easy to do with an agent and we 
> could have it done in a couple days.
>
> Enforce boundaries with ArchUnit
>
> If we want to enforce certain code boundaries, this is the way to do it. 
> Should not be part of manual review.
>
> Add LLM review for all incoming PRs before a human
>
> The goal here is to automate the initial part of the review process that 
> reviewers should spot, and raise the bar for the initial contribution.  When 
> the code gets reviewed by a human, it should already have passed a large 
> variety of initial checks.  This should shorten the review cycle and result 
> in higher quality patches.  I've had Claude reviewing all my PRs in my 
> personal projects for a while now and it consistently gives great feedback 
> that I almost always incorporate.
>
> In my ideal world, we'd also auto-format all code
>
> Consistent formatting throughout the codebase would be amazing, but that's 
> just one man's dream.
>
> Hopefully there's at least a couple things in this list we could move forward 
> with in the short term, as it'll help improve the code quality regardless of 
> how it's created.
>
> Jon
>
> https://checkerframework.org/manual/#aliasing-leaking-contexts
> https://www.sonarsource.com/docs/CognitiveComplexity.pdf
> https://pmd.github.io/pmd/pmd_rules_java_design.html
>
>
>
>
>
> On Thu, Sep 24, 2026 at 7:38 AM C. Scott Andreas <[email protected]> wrote:
>>
>> From Benedict:
>>
>> “I don't know if everyone remembers, but ten years ago Cassandra was full of 
>> serious correctness and stability issues. Despite developing it, I would not 
>> have run it myself or recommend that anyone use it. We have dug ourselves 
>> out of that hole, but it took years of discipline and effort, and we're 
>> still (deservedly) recovering our reputation.”
>>
>> Expanding on this point for those who may not have been active in the 
>> project at this time —
>>
>> Apache Cassandra was fundamentally undeployable for four years between Nov 
>> 2015 - 2019. The database literally lost data if you ran a read-only SELECT 
>> query ordered descending (C-14513, C-14515). If you haven’t read these 
>> tickets before, please take a moment to do so.
>>
>> It took years of careful work via property-based testing, fuzzing, and 
>> deterministic simulation to restore Cassandra’s status as a usable system of 
>> record. Once 14513 and 14515 were identified, nearly 30 additional critical 
>> data loss and incorrect response bugs were identified.
>>
>> It is essential for the project’s future that we don’t regress to this state 
>> chasing AI-generated features motivated by fear. The fact that examples 
>> cited in this thread which boast shiny features but have critical 
>> shortcomings unknown to their author supports this argument.
>>
>> The most common path for large corpuses of AI-generated software is elation 
>> and reveling in a feature matrix, followed by abandonment.
>>
>> I endorse this point:
>>
>> “Let's use this new technology to improve the quality of our contributions, 
>> not squander our hard-earned gains in the name of speed. It will be hard to 
>> recover our reputation a second time.”
>>
>> Patrick, I don’t want your note regarding a TCM issue to go unaddressed. 
>> Please file a Jira ticket and the patch if you like. I can’t comment on the 
>> patch as I haven’t seen it, but together we will solve the problem.
>>
>> – Scott
>>
>> > On Sep 24, 2026, at 4:01 AM, Benedict Elliott Smith <[email protected]> 
>> > wrote:
>> >
>> > Hi Patrick,
>> >
>> > As I mentioned in my reply to David, I would be happy to create a carve 
>> > out for shallow and localised bug fixes in the "Permitted" section. Would 
>> > this alleviate some of your concerns regarding your ability to contribute 
>> > to the project?
>> >
>> > I appreciate your pointing out Ferrosa's Accord implementation however, as 
>> > it is a *great* example of the problems we're leaping into. I took a look, 
>> > and within about 30s found that the protocol is fundamentally incorrect, 
>> > having failed to address CASSANDRA-18365. This is despite claiming to be 
>> > tested with Jepsen that should in principle find this fault. I followed up 
>> > by using Claude to interrogate the implementation further, and immediately 
>> > found other serious correctness issues.
>> >
>> > I use LLMs daily now to help facilitate Accord development, and while they 
>> > are powerful they are NOT able to author the code themselves, even when 
>> > building upon a strong human-authored foundation.
>> >
>> > I don't know if everyone remembers, but ten years ago Cassandra was full 
>> > of serious correctness and stability issues. Despite developing it, I 
>> > would not have run it myself or recommend that anyone use it. We have dug 
>> > ourselves out of that hole, but it took years of discipline and effort, 
>> > and we're still (deservedly) recovering our reputation.
>> >
>> > Let's use this new technology to improve the quality of our contributions, 
>> > not squander our hard-earned gains in the name of speed. It will be hard 
>> > to recover our reputation a second time.
>> >
>> >
>> >> On 2026/09/23 19:16:17 Patrick McFadin wrote:
>> >> I was waiting for this moment to hit our project and I'm glad we're here. 
>> >> I
>> >> am deeply concerned for our project and its future, as we have 
>> >> increasingly
>> >> made it difficult to contribute. I had hoped that this new era of
>> >> software tools powered by AI would expand the project's reach and bring
>> >> more diverse thoughts and ideas. This policy proposal is the exact 
>> >> opposite
>> >> of what we need. We have been sitting on a Cassandra 6 release alpha for
>> >> months. We need to accelerate and embrace new ways of being or be left
>> >> behind. As I read that policy, my first and gut level reactions:
>> >> - It comes across as elitist and class protectionism. Committer should not
>> >> be special but this proposal makes that designation even more sacred.
>> >> - It signals that our project is so fragile that only a few people "Really
>> >> understand it" That's some SQLite vibes right there.
>> >> - Trying to fix a problem that doesn't exist
>> >> Sadly, i think this policy change would also exclude a lot of comitters.
>> >> We aren't alone in this moment. The Linux project just went through
>> >> this. You can find the thread with a simple Google, but similar hard
>> >> feelings were being expressed "AI is going to ruin our project!", "The
>> >> unwashed masses are going to contribute terrible code!", "We have to
>> >> protect our precious status as Linux maintainers!"  Linus being Linus, was
>> >> deeply invloved and they adopted a super simple statement that covers all
>> >> bases. Human or Human using AI. “You are expected to understand and to be
>> >> able to defend everything you submit.”  Love that.
>> >> In the larger picture, I'll restate. I'm worried for our project. In late
>> >> 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real corner and
>> >> in the hands of somebody that knows how to build software, this tool is
>> >> like jet fuel. Here's some examples of new projects being hyper fueled by
>> >> AI coding tools.
>> >> Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity
>> >> Turso - Rust re-write of SQLite
>> >> Bun - Rust re-write of itself from Zig.
>> >> Think this couldn't happen to us? Already has:
>> >> https://github.com/ferrosadb/ferrosa. Ben is using it to power his own
>> >> startup, but it was him alone using a ton of local AI coding agents. He
>> >> even implemented Accord. Yeah...
>> >> The cracks are already starting to show. There is a black market economy 
>> >> of
>> >> Cassandra patches happening now. Not going to name names or call people
>> >> out,  but there are fixes and optimizations living in branches outside of
>> >> the Cassandra project. Why? I'll use myself as an example. I fixed a nasty
>> >> bug I ran into with TCM a few weeks ago. Wrote the tests. It passes CI and
>> >> lives in my personal branch. I'm sitting here really wondering if I want 
>> >> to
>> >> go through the ritual humiliation of being roasted for using AI to fix it.
>> >> Me. I am worried about contrinuting code the Cassandra. What the hell does
>> >> that say?
>> >> I have my CQLite project that I've been doing a release around once a
>> >> month. I would love to donate that to the Cassandra project but I wouldn't
>> >> if it essentially killed any progress.
>> >> My larger counter proposal would be to:
>> >> - Adopt the “You are expected to understand and to be able to defend
>> >> everything you submit.” approach the Linux project has adopted.
>> >> - Loosen up the contributor process and our worry on trunk. Let 1000
>> >> flowers bloom and bring it in.
>> >> - And finally, to give some people more peace of mind and open more doors,
>> >> adopt what other projects have done and provide more pluggability. Let new
>> >> ideas have an easy place to connect.
>> >> We are at a fork in the road. What are we going to do? And then I have to
>> >> ask myself, what am I going to do as a contributor?
>> >> Patrick
>> >> On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston <[email protected]>
>> >> wrote:
>> >>> I’m not necessarily opposed to having a policy, but so far we have some
>> >>> specific proposals addressing a problem statement that’s very nebulous.
>> >>> What is the community failing to do on its own that we’re trying to 
>> >>> correct
>> >>> with policy? What outcomes are we trying to create or prevent? Having 
>> >>> some
>> >>> examples and specific problems to discuss would help focus the 
>> >>> conversation.
>> >>>> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote:
>> >>> Benedict,
>> >>> Thanks for clarifying. My concern still remains. This criteria would be
>> >>> difficult to define and apply consistently. What counts as “similar” 
>> >>> scope
>> >>> or area, “mostly correct,” or sufficiently independent work? More
>> >>> importantly, how do we prevent such vague criteria from creating an
>> >>> informal hierarchy where some contributors work is routinely accepted 
>> >>> while
>> >>> others is routinely rejected?
>> >>> If the intent is to limit AI-assisted code changes to Cassandra
>> >>> contributors, or to contributors who have previously worked in that
>> >>> component without AI, that would at least be clear and enforceable.
>> >>>> On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith <
>> >>> [email protected]> wrote:
>> >>>> Core code changes
>> >>>> Chris: Do you object to the first or second line you quote? Because the
>> >>> first line is effectively motivation for the second line, and can be
>> >>> removed (or more clearly combined). If it’s the second line, then I do 
>> >>> not
>> >>> think this is an unreasonable expectation, and we can get into a proper
>> >>> debate about it.
>> >>>> Shailaja, since you only snipped the first sentence, your concerns might
>> >>> also be mostly answered by this clarification? “Minimal third-party
>> >>> guidance” implies you have some concerns about the second line, but all 
>> >>> of
>> >>> our policies have some ambiguity because legalese is even worse. I don’t
>> >>> think the ambiguity here would be challenging to navigate though we can
>> >>> certainly refine it. This specific snippet is meant to convey an
>> >>> expectation that a contributor has autonomously produced patches of 
>> >>> similar
>> >>> scope that were mostly correct, so that they have demonstrated the level 
>> >>> of
>> >>> understanding necessary to guide another party to a successful patch 
>> >>> (i.e.
>> >>> an LLM in this case).
>> >>>> On 2026/09/23 10:54:16 Benedict Elliott Smith wrote:
>> >>>>> Thanks everyone for your input so far. I’ll respond in brief to the
>> >>> main themes, in (mostly) separate emails so they can each have their own
>> >>> debate chain.
>> >>>>> Should we have a policy (Blake/Josh*/Jon/Dinesh)
>> >>>>> I think we would all agree that LLMs represent the biggest change to
>> >>> this community (and software more generally) since its inception, and we
>> >>> all now have enough experience with the technology to have formed 
>> >>> opinions
>> >>> about how it is best managed. We also evidently have not all arrived at 
>> >>> the
>> >>> same conclusions. In this situation, it would be an abdication of our
>> >>> responsibilities as a management committee to not agree *some* policy.
>> >>>>> I intend to conduct straw polls as the discussion evolves, so if you
>> >>> prefer an alternative policy - or modifications to this policy - I would
>> >>> encourage you to make those alternative proposals.
>> >>>>> *Veto/Consensus (Josh)
>> >>>>> It was fair to call out my poor use of language on this topic, so let
>> >>> me rephrase a little. The community is built on consensus, and work 
>> >>> should
>> >>> not be merged when there are outstanding concerns to address. The 
>> >>> explicit
>> >>> -1 should only be used rarely, because the prior expectation should 
>> >>> prevent
>> >>> it ever being needed. I (and others) have outstanding concerns on LLM
>> >>> generated work that can only be addressed through this process right 
>> >>> here,
>> >>> so to merge such work while maintaining the community’s consensus we must
>> >>> agree some policy.
>> >>>>> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote:
>> >>>>>> I am strongly -1 on this
>> >>>>>> - Core code changes made by LLM may only be proposed by contributors
>> >>> with demonstrated expertise
>> >>>>>> That creates a new, subjective privileged class of contributors and
>> >>> turns a tool choice into an eligibility test. Who decides whether 
>> >>> expertise
>> >>> has been “demonstrated,” what counts as “minimal third-party guidance,” 
>> >>> and
>> >>> how could those judgments be applied consistently or fairly?
>> >>>>>> Apache already has a better model, anyone may contribute, trust and
>> >>> additional repository privileges are earned transparently over time. The
>> >>> ASF describes its communities as flat, and says that newcomer ideas have 
>> >>> as
>> >>> much input as those from original creators. We should not add a separate,
>> >>> informal hierarchy in which certain people may use common development 
>> >>> tools
>> >>> while others may not.
>> >>>>>>> On Sep 23, 2026, at 6:33 AM, Chris Lohfink <[email protected]>
>> >>> wrote:
>> >>>>>>> - Core code changes made by LLM may only be proposed by contributors
>> >>> with demonstrated expertise
>> >>>>>>> - Must have produced similar patches in size, scope and area
>> >>> unassisted and with minimal third-party guidance
>> >>>>>>> I really don't like this one or its wording. Definitely too "the
>> >>> peasants are getting uppity lets build a wall". Lets not let a subjective
>> >>> thing like demonstrated expertise (who decides that?) be if it's ok or 
>> >>> not.
>> >>> Hold the same standards for code quality and process for it all. I don't
>> >>> want this to be: only people on the storage team in Apple can use AI.
>> >>>>>>> Chris
>> >>>>>>> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto:
>> >>> [email protected]>> wrote:
>> >>>>>>>> I agree with Stefan and think this is both a reasonable and
>> >>> thoughtful proposal.
>> >>>>>>>> Here are some things I like about it:
>> >>>>>>>> – It outlines areas where LLM usage is unambiguously useful to the
>> >>> project’s developers and users.
>> >>>>>>>> – It defines a spectrum of recommendations and cautions.
>> >>>>>>>> – The only prohibited areas are extremely narrow and say nothing
>> >>> about code at all.
>> >>>>>>>> Some in this thread are responding as if this proposal seeks to
>> >>> prohibit or sharply limit use of LLMs. In fact, it’s one of the most open
>> >>> and welcoming I’ve seen for an OSS project of our size where many are
>> >>> adopting policies that simply ban them entirely. I’ve re-appended the
>> >>> proposal below my message as it seems to have been lost in threaded
>> >>> replies, and would encourage folks to give it a second read.
>> >>>>>>>> Some brief thoughts based on my own use of LLMs:
>> >>>>>>>> – I find them fantastically useful for reviewing and identifying
>> >>> problems that have slipped through review - primarily via Alex Petrov’s
>> >>> /deep-review skill, which I have running in a VM in a loop executing over
>> >>> every new commit in the project as of a few days ago. I will be posting a
>> >>> few hand-authored Jira tickets based on findings that appear legitimate 
>> >>> to
>> >>> me. For now, the loop is posting them as issue drafts for my own review 
>> >>> on
>> >>> my personal fork which you can find here:
>> >>> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug
>> >>>>>>>> – They’re great for enabling use of model checkers and formal
>> >>> methods where such work would have previously been prohibitively 
>> >>> expensive,
>> >>> such as Blake’s work on a TLA+ proof of aspects of Mutation Tracking and
>> >>> Benedict/Fedor’s work on a machine-checkable proof of the Accord protocol
>> >>> in Lean.
>> >>>>>>>> – They are stunning for allowing me to experiment with ideas that
>> >>> would have otherwise been a summer internship’s scope of work. Some
>> >>> examples include an io_uring prototype, exploring the impact of
>> >>> page-aligned compressed chunk sizes, an API shim bridging the 3.x and 4.x
>> >>> Java Drivers, and potential enhancements to Zstandard.
>> >>>>>>>> – And they shine when given grunt-work that is critical to the
>> >>> project but a miserable labor for humans, such as triaging, reproducing,
>> >>> and root-causing flaky tests, which David Capwell now has running in a 
>> >>> loop
>> >>> to help us improve CI stability in the project.
>> >>>>>>>> I never thought I’d be so positive on what’s possible via language
>> >>> models a year ago. At the same time, I also agree that they present
>> >>> challenges and risks that can be managed through thoughtful discussion 
>> >>> and
>> >>> policy. Some of the concerns that I think are important to guard against
>> >>> include:
>> >>>>>>>> – Asymmetry of effort between author and reviewers: As
>> >>> token-generating machines, LLMs can generate diffs of extraordinary size
>> >>> very rapidly. /deep-review is great for chewing through diffs and
>> >>> identifying defects. But it should be used by the contributor themselves 
>> >>> to
>> >>> identify issues – not to replace the role of the reviewer with more
>> >>> electricity. The role of the reviewers extends beyond identifying and
>> >>> highlighting defects. It encompasses architecture, harmony with the
>> >>> existing codebase, thinking ahead to future evolution of the project, and
>> >>> replicates context on the project as new code is committed. These 
>> >>> functions
>> >>> cannot be automated away.
>> >>>>>>>> – Hesitancy of authors to engage manually with code they have
>> >>> generated: This is not specific to Cassandra, but it is a behavior that I
>> >>> have seen in several “highly-electric” projects. There’s a bimodal 
>> >>> tendency
>> >>> toward code that is entirely generated or entirely human-authored - but 
>> >>> it
>> >>> is rare for someone to prepare an AI-authored patch to take an offramp 
>> >>> and
>> >>> spend a significant amount of time refining the work by hand in an IDE.
>> >>> This hesitancy toward human participation in authorship of LLM-generated
>> >>> code is very concerning to me.
>> >>>>>>>> – Harmony with the existing codebase: Due to the tunnel-vision of
>> >>> context windows, LLMs are generally unaware of conventions and norms
>> >>> present in codebases and very frequently reinvent concepts in a 
>> >>> generation
>> >>> turn to suit a goal without view of the project’s overall architecture.
>> >>> This results in a profusion of messy and duplicated concepts that 
>> >>> gradually
>> >>> sprawl about a codebase.
>> >>>>>>>> Again, none of these are grounds for prohibition of usage of
>> >>> language models in developing the project. They’re just problems we need 
>> >>> to
>> >>> bear in mind and guard against – and I think the proposal is designed to 
>> >>> do
>> >>> just that.
>> >>>>>>>> I’m thrilled by the potential of LLMs to improve Apache Cassandra
>> >>> and we already see it happening through a vast number of issues that are
>> >>> being reported and fixed. But there’s also danger in taking ATVs down a
>> >>> hiking trail full of people.
>> >>>>>>>> Regarding the prohibition on prose, I’ll simply say: I recently
>> >>> found myself in a scenario where I found a Claude-authored document so
>> >>> inscrutable that I piped it back into a model, directed it to rewrite it 
>> >>> in
>> >>> ASD-STE100, read it myself, and responded based on the summarization. As 
>> >>> a
>> >>> humanities grad, this is probably the worst language crime I have
>> >>> committed. But it was in response to language that was itself so
>> >>> idiosyncratic that it was unreadable to me in its original form. I hope
>> >>> this never happens in the Apache Cassandra project.
>> >>>>>>>> I’ll close with a quote from an excellent article written by Colin
>> >>> Breck, an engineer who works on large-scale data systems:
>> >>> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/
>> >>>>>>>> Colin wrote:
>> >>>>>>>>> I don’t want to live in a world where you use AI to summarize
>> >>> something important into unreadable text, and then I use AI in an attempt
>> >>> to decipher it. I want to hear you, imperfections and all. I want your
>> >>> interpretation of aesthetics, beauty, quality, relationship, time. I want
>> >>> to know how you feel. I want you to cut through and tell me what really
>> >>> matters.
>> >>>>>>>>> Intentional writing will likely become more valuable. People who
>> >>> write, and write to think, to think deeply and carefully, or to create, 
>> >>> to
>> >>> share, or to capture something important without explicitly expressing it
>> >>> will continue to write and produce original work. The people who never 
>> >>> were
>> >>> writers will use AI to produce lots of text.
>> >>>>>>>> I hope that our culture can remain one of intentional writing and
>> >>> intentional engineering. I enjoy reading the voice of the author in
>> >>> comments, code, and tickets in Cassandra – the different ways we use
>> >>> language based on where we grew up and how we learned English, the
>> >>> translated idioms from our various backgrounds, and terse comments that
>> >>> recognize the difference between code whose function is obvious and what
>> >>> warrants genuine exposition. When I read code in Cassandra, it’s a 
>> >>> delight
>> >>> to recognize the author based on their writing style before flipping on
>> >>> `git annotate` to reveal the origin.
>> >>>>>>>> I’d encourage folks to re-read the original proposal below. It is
>> >>> very permissive. The guidance strikes me not just as reasonable, but
>> >>> genuinely important to maintaining the health of the project.
>> >>>>>>>> – Scott
>> >>>>>>>> =====
>> >>>>>>>> Encouraged:
>> >>>>>>>> - Reviewing and otherwise validating human-authored patches before
>> >>> submission
>> >>>>>>>> - Debugging, diagnosing etc
>> >>>>>>>> Permitted:
>> >>>>>>>> - Generating or modifying tests, scripts, tooling or any other
>> >>> non-user facing changes
>> >>>>>>>> - Minor changes to human-authored patches that are carefully
>> >>> reviewed by the author
>> >>>>>>>> Restricted:
>> >>>>>>>> - Core code changes made by LLM may only be proposed by contributors
>> >>> with demonstrated expertise
>> >>>>>>>> - Must have produced similar patches in size, scope and area
>> >>> unassisted and with minimal third-party guidance
>> >>>>>>>> - Core code changes made by LLM require an additional reviewer
>> >>>>>>>> - LLM review is not a substitute for human review, and must be used
>> >>> only to augment a complete and independent human understanding of the 
>> >>> patch.
>> >>>>>>>> Prohibited:
>> >>>>>>>> - All public prose must be human authored. This includes inline
>> >>> comments, docs, posts to Jira etc.
>> >>>>>>>> All LLM generated changes MUST be disclosed:
>> >>>>>>>> - Outlined to any reviewer;
>> >>>>>>>> - Summarised in the commit message;
>> >>>>>>>> - Large blocks or files must be individually marked with some agreed
>> >>> message like "created by <some AI>"
>> >>>>>>>> =====
>> >>>>>>>>> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected]
>> >>> <mailto:[email protected]>> wrote:
>> >>>>>>>>> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected]
>> >>> <mailto:[email protected]>> wrote:
>> >>>>>>>>>> Restricted:
>> >>>>>>>>>> - Core code changes made by LLM may only be proposed by
>> >>> contributors with demonstrated expertise
>> >>>>>>>>>> - Must have produced similar patches in size, scope and area
>> >>> unassisted and with minimal third-party guidance
>> >>>>>>>>> I am -1 on this. This sounds like gate keeping attempt. It narrowly
>> >>> limits the pool to a few people on the project that have historically
>> >>> contributed to certain parts of the codebase. This policy will prohibit
>> >>> skilled software engineers with domain expertise from proposing LLM
>> >>> assisted changes simply because they have not contributed to the project.
>> >>> This is unrealistic and a net negative for the project to attract talent
>> >>> and grow our community.
>> >>>>>>>>>> - Core code changes made by LLM require an additional reviewer
>> >>>>>>>>> Can you be more precise what is this in addition to? How many total
>> >>> reviewers do you expect and what is the purpose of additional reviewer? 
>> >>> and
>> >>> why?
>> >>>>>>>>> Taking a step back - what are you trying to solve here?
>> >>>>>>>>> Dinesh

Reply via email to