I absolutely remember and we solved it with new test frameworks, tooling,
ensuring rigorous testing and pausing all contributions to stabilize the
codebase. The critical point is that we did not selectively slow down
patches only from some contributors with a highly subjective criteria.

IIRC - we need substantial resources to run said tooling. Over the years
there have been resources committed to the project but I don’t think
they’re sufficient.

Here’s the bottom line - we have asymmetry in contributions. We’ve had more
contributors than reviewers historically. This was always the case.
LLMs have amplified this gap.

IMHO the solution is multi-faceted -

1. Find ways to leverage LLMs to improve our testing and verification story.

2. Actively mentor committers on the project I help them grow into
maintainers for complex systems.

3. Large users of Cassandra contribute resources to the project for testing
at scale (either committing resources or validating patches via their
internal test infrastructure and publishing reports).

In the meanwhile, I will support a sensible policy to treat all incoming
patches the same and expect the author to fully understand and explain
their decisions and prove the correctness of their code via rigorous
testing.

Thanks,

Dinesh

PS: This is a 100% human generated response.

On Thu, Sep 24, 2026 at 8:36 AM Caleb Rackliffe <[email protected]>
wrote:

> I think most of us DO remember, but I’m worried that requiring even more
> committer involvement than we already do and/or gatekeeping contributions
> around amorphous criteria could be just as dangerous to the project.
> (Clearly not alone in this worry.)
>
> Again, this is worth a separate thread, possibly, but there are some
> things that appear to be scaring people from trying to make contributions
> that have nothing to do with AI or our policy around it.
>
> > On Sep 24, 2026, at 9:38 AM, C. Scott Andreas <[email protected]>
> wrote:
> >
> > From Benedict:
> >
> > “I don't know if everyone remembers, but ten years ago Cassandra was
> full of serious correctness and stability issues. Despite developing it, I
> would not have run it myself or recommend that anyone use it. We have dug
> ourselves out of that hole, but it took years of discipline and effort, and
> we're still (deservedly) recovering our reputation.”
> >
> > Expanding on this point for those who may not have been active in the
> project at this time —
> >
> > Apache Cassandra was fundamentally undeployable for four years between
> Nov 2015 - 2019. The database literally lost data if you ran a read-only
> SELECT query ordered descending (C-14513, C-14515). If you haven’t read
> these tickets before, please take a moment to do so.
> >
> > It took years of careful work via property-based testing, fuzzing, and
> deterministic simulation to restore Cassandra’s status as a usable system
> of record. Once 14513 and 14515 were identified, nearly 30 additional
> critical data loss and incorrect response bugs were identified.
> >
> > It is essential for the project’s future that we don’t regress to this
> state chasing AI-generated features motivated by fear. The fact that
> examples cited in this thread which boast shiny features but have critical
> shortcomings unknown to their author supports this argument.
> >
> > The most common path for large corpuses of AI-generated software is
> elation and reveling in a feature matrix, followed by abandonment.
> >
> > I endorse this point:
> >
> > “Let's use this new technology to improve the quality of our
> contributions, not squander our hard-earned gains in the name of speed. It
> will be hard to recover our reputation a second time.”
> >
> > Patrick, I don’t want your note regarding a TCM issue to go unaddressed.
> Please file a Jira ticket and the patch if you like. I can’t comment on the
> patch as I haven’t seen it, but together we will solve the problem.
> >
> > – Scott
> >
> >> On Sep 24, 2026, at 4:01 AM, Benedict Elliott Smith <
> [email protected]> wrote:
> >>
> >> Hi Patrick,
> >>
> >> As I mentioned in my reply to David, I would be happy to create a carve
> out for shallow and localised bug fixes in the "Permitted" section. Would
> this alleviate some of your concerns regarding your ability to contribute
> to the project?
> >>
> >> I appreciate your pointing out Ferrosa's Accord implementation however,
> as it is a *great* example of the problems we're leaping into. I took a
> look, and within about 30s found that the protocol is fundamentally
> incorrect, having failed to address CASSANDRA-18365. This is despite
> claiming to be tested with Jepsen that should in principle find this fault.
> I followed up by using Claude to interrogate the implementation further,
> and immediately found other serious correctness issues.
> >>
> >> I use LLMs daily now to help facilitate Accord development, and while
> they are powerful they are NOT able to author the code themselves, even
> when building upon a strong human-authored foundation.
> >>
> >> I don't know if everyone remembers, but ten years ago Cassandra was
> full of serious correctness and stability issues. Despite developing it, I
> would not have run it myself or recommend that anyone use it. We have dug
> ourselves out of that hole, but it took years of discipline and effort, and
> we're still (deservedly) recovering our reputation.
> >>
> >> Let's use this new technology to improve the quality of our
> contributions, not squander our hard-earned gains in the name of speed. It
> will be hard to recover our reputation a second time.
> >>
> >>
> >>>> On 2026/09/23 19:16:17 Patrick McFadin wrote:
> >>> I was waiting for this moment to hit our project and I'm glad we're
> here. I
> >>> am deeply concerned for our project and its future, as we have
> increasingly
> >>> made it difficult to contribute. I had hoped that this new era of
> >>> software tools powered by AI would expand the project's reach and bring
> >>> more diverse thoughts and ideas. This policy proposal is the exact
> opposite
> >>> of what we need. We have been sitting on a Cassandra 6 release alpha
> for
> >>> months. We need to accelerate and embrace new ways of being or be left
> >>> behind. As I read that policy, my first and gut level reactions:
> >>> - It comes across as elitist and class protectionism. Committer should
> not
> >>> be special but this proposal makes that designation even more sacred.
> >>> - It signals that our project is so fragile that only a few people
> "Really
> >>> understand it" That's some SQLite vibes right there.
> >>> - Trying to fix a problem that doesn't exist
> >>> Sadly, i think this policy change would also exclude a lot of
> comitters.
> >>> We aren't alone in this moment. The Linux project just went through
> >>> this. You can find the thread with a simple Google, but similar hard
> >>> feelings were being expressed "AI is going to ruin our project!", "The
> >>> unwashed masses are going to contribute terrible code!", "We have to
> >>> protect our precious status as Linux maintainers!"  Linus being Linus,
> was
> >>> deeply invloved and they adopted a super simple statement that covers
> all
> >>> bases. Human or Human using AI. “You are expected to understand and to
> be
> >>> able to defend everything you submit.”  Love that.
> >>> In the larger picture, I'll restate. I'm worried for our project. In
> late
> >>> 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real corner
> and
> >>> in the hands of somebody that knows how to build software, this tool is
> >>> like jet fuel. Here's some examples of new projects being hyper fueled
> by
> >>> AI coding tools.
> >>> Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity
> >>> Turso - Rust re-write of SQLite
> >>> Bun - Rust re-write of itself from Zig.
> >>> Think this couldn't happen to us? Already has:
> >>> https://github.com/ferrosadb/ferrosa. Ben is using it to power his own
> >>> startup, but it was him alone using a ton of local AI coding agents. He
> >>> even implemented Accord. Yeah...
> >>> The cracks are already starting to show. There is a black market
> economy of
> >>> Cassandra patches happening now. Not going to name names or call people
> >>> out,  but there are fixes and optimizations living in branches outside
> of
> >>> the Cassandra project. Why? I'll use myself as an example. I fixed a
> nasty
> >>> bug I ran into with TCM a few weeks ago. Wrote the tests. It passes CI
> and
> >>> lives in my personal branch. I'm sitting here really wondering if I
> want to
> >>> go through the ritual humiliation of being roasted for using AI to fix
> it.
> >>> Me. I am worried about contrinuting code the Cassandra. What the hell
> does
> >>> that say?
> >>> I have my CQLite project that I've been doing a release around once a
> >>> month. I would love to donate that to the Cassandra project but I
> wouldn't
> >>> if it essentially killed any progress.
> >>> My larger counter proposal would be to:
> >>> - Adopt the “You are expected to understand and to be able to defend
> >>> everything you submit.” approach the Linux project has adopted.
> >>> - Loosen up the contributor process and our worry on trunk. Let 1000
> >>> flowers bloom and bring it in.
> >>> - And finally, to give some people more peace of mind and open more
> doors,
> >>> adopt what other projects have done and provide more pluggability. Let
> new
> >>> ideas have an easy place to connect.
> >>> We are at a fork in the road. What are we going to do? And then I have
> to
> >>> ask myself, what am I going to do as a contributor?
> >>> Patrick
> >>> On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston <[email protected]>
> >>> wrote:
> >>>> I’m not necessarily opposed to having a policy, but so far we have
> some
> >>>> specific proposals addressing a problem statement that’s very
> nebulous.
> >>>> What is the community failing to do on its own that we’re trying to
> correct
> >>>> with policy? What outcomes are we trying to create or prevent? Having
> some
> >>>> examples and specific problems to discuss would help focus the
> conversation.
> >>>>> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote:
> >>>> Benedict,
> >>>> Thanks for clarifying. My concern still remains. This criteria would
> be
> >>>> difficult to define and apply consistently. What counts as “similar”
> scope
> >>>> or area, “mostly correct,” or sufficiently independent work? More
> >>>> importantly, how do we prevent such vague criteria from creating an
> >>>> informal hierarchy where some contributors work is routinely accepted
> while
> >>>> others is routinely rejected?
> >>>> If the intent is to limit AI-assisted code changes to Cassandra
> >>>> contributors, or to contributors who have previously worked in that
> >>>> component without AI, that would at least be clear and enforceable.
> >>>>> On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith <
> >>>> [email protected]> wrote:
> >>>>> Core code changes
> >>>>> Chris: Do you object to the first or second line you quote? Because
> the
> >>>> first line is effectively motivation for the second line, and can be
> >>>> removed (or more clearly combined). If it’s the second line, then I
> do not
> >>>> think this is an unreasonable expectation, and we can get into a
> proper
> >>>> debate about it.
> >>>>> Shailaja, since you only snipped the first sentence, your concerns
> might
> >>>> also be mostly answered by this clarification? “Minimal third-party
> >>>> guidance” implies you have some concerns about the second line, but
> all of
> >>>> our policies have some ambiguity because legalese is even worse. I
> don’t
> >>>> think the ambiguity here would be challenging to navigate though we
> can
> >>>> certainly refine it. This specific snippet is meant to convey an
> >>>> expectation that a contributor has autonomously produced patches of
> similar
> >>>> scope that were mostly correct, so that they have demonstrated the
> level of
> >>>> understanding necessary to guide another party to a successful patch
> (i.e.
> >>>> an LLM in this case).
> >>>>> On 2026/09/23 10:54:16 Benedict Elliott Smith wrote:
> >>>>>> Thanks everyone for your input so far. I’ll respond in brief to the
> >>>> main themes, in (mostly) separate emails so they can each have their
> own
> >>>> debate chain.
> >>>>>> Should we have a policy (Blake/Josh*/Jon/Dinesh)
> >>>>>> I think we would all agree that LLMs represent the biggest change to
> >>>> this community (and software more generally) since its inception, and
> we
> >>>> all now have enough experience with the technology to have formed
> opinions
> >>>> about how it is best managed. We also evidently have not all arrived
> at the
> >>>> same conclusions. In this situation, it would be an abdication of our
> >>>> responsibilities as a management committee to not agree *some* policy.
> >>>>>> I intend to conduct straw polls as the discussion evolves, so if you
> >>>> prefer an alternative policy - or modifications to this policy - I
> would
> >>>> encourage you to make those alternative proposals.
> >>>>>> *Veto/Consensus (Josh)
> >>>>>> It was fair to call out my poor use of language on this topic, so
> let
> >>>> me rephrase a little. The community is built on consensus, and work
> should
> >>>> not be merged when there are outstanding concerns to address. The
> explicit
> >>>> -1 should only be used rarely, because the prior expectation should
> prevent
> >>>> it ever being needed. I (and others) have outstanding concerns on LLM
> >>>> generated work that can only be addressed through this process right
> here,
> >>>> so to merge such work while maintaining the community’s consensus we
> must
> >>>> agree some policy.
> >>>>>> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote:
> >>>>>>> I am strongly -1 on this
> >>>>>>> - Core code changes made by LLM may only be proposed by
> contributors
> >>>> with demonstrated expertise
> >>>>>>> That creates a new, subjective privileged class of contributors and
> >>>> turns a tool choice into an eligibility test. Who decides whether
> expertise
> >>>> has been “demonstrated,” what counts as “minimal third-party
> guidance,” and
> >>>> how could those judgments be applied consistently or fairly?
> >>>>>>> Apache already has a better model, anyone may contribute, trust and
> >>>> additional repository privileges are earned transparently over time.
> The
> >>>> ASF describes its communities as flat, and says that newcomer ideas
> have as
> >>>> much input as those from original creators. We should not add a
> separate,
> >>>> informal hierarchy in which certain people may use common development
> tools
> >>>> while others may not.
> >>>>>>>> On Sep 23, 2026, at 6:33 AM, Chris Lohfink <[email protected]>
> >>>> wrote:
> >>>>>>>> - Core code changes made by LLM may only be proposed by
> contributors
> >>>> with demonstrated expertise
> >>>>>>>> - Must have produced similar patches in size, scope and area
> >>>> unassisted and with minimal third-party guidance
> >>>>>>>> I really don't like this one or its wording. Definitely too "the
> >>>> peasants are getting uppity lets build a wall". Lets not let a
> subjective
> >>>> thing like demonstrated expertise (who decides that?) be if it's ok
> or not.
> >>>> Hold the same standards for code quality and process for it all. I
> don't
> >>>> want this to be: only people on the storage team in Apple can use AI.
> >>>>>>>> Chris
> >>>>>>>> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto:
> >>>> [email protected]>> wrote:
> >>>>>>>>> I agree with Stefan and think this is both a reasonable and
> >>>> thoughtful proposal.
> >>>>>>>>> Here are some things I like about it:
> >>>>>>>>> – It outlines areas where LLM usage is unambiguously useful to
> the
> >>>> project’s developers and users.
> >>>>>>>>> – It defines a spectrum of recommendations and cautions.
> >>>>>>>>> – The only prohibited areas are extremely narrow and say nothing
> >>>> about code at all.
> >>>>>>>>> Some in this thread are responding as if this proposal seeks to
> >>>> prohibit or sharply limit use of LLMs. In fact, it’s one of the most
> open
> >>>> and welcoming I’ve seen for an OSS project of our size where many are
> >>>> adopting policies that simply ban them entirely. I’ve re-appended the
> >>>> proposal below my message as it seems to have been lost in threaded
> >>>> replies, and would encourage folks to give it a second read.
> >>>>>>>>> Some brief thoughts based on my own use of LLMs:
> >>>>>>>>> – I find them fantastically useful for reviewing and identifying
> >>>> problems that have slipped through review - primarily via Alex
> Petrov’s
> >>>> /deep-review skill, which I have running in a VM in a loop executing
> over
> >>>> every new commit in the project as of a few days ago. I will be
> posting a
> >>>> few hand-authored Jira tickets based on findings that appear
> legitimate to
> >>>> me. For now, the loop is posting them as issue drafts for my own
> review on
> >>>> my personal fork which you can find here:
> >>>>
> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug
> >>>>>>>>> – They’re great for enabling use of model checkers and formal
> >>>> methods where such work would have previously been prohibitively
> expensive,
> >>>> such as Blake’s work on a TLA+ proof of aspects of Mutation Tracking
> and
> >>>> Benedict/Fedor’s work on a machine-checkable proof of the Accord
> protocol
> >>>> in Lean.
> >>>>>>>>> – They are stunning for allowing me to experiment with ideas that
> >>>> would have otherwise been a summer internship’s scope of work. Some
> >>>> examples include an io_uring prototype, exploring the impact of
> >>>> page-aligned compressed chunk sizes, an API shim bridging the 3.x and
> 4.x
> >>>> Java Drivers, and potential enhancements to Zstandard.
> >>>>>>>>> – And they shine when given grunt-work that is critical to the
> >>>> project but a miserable labor for humans, such as triaging,
> reproducing,
> >>>> and root-causing flaky tests, which David Capwell now has running in
> a loop
> >>>> to help us improve CI stability in the project.
> >>>>>>>>> I never thought I’d be so positive on what’s possible via
> language
> >>>> models a year ago. At the same time, I also agree that they present
> >>>> challenges and risks that can be managed through thoughtful
> discussion and
> >>>> policy. Some of the concerns that I think are important to guard
> against
> >>>> include:
> >>>>>>>>> – Asymmetry of effort between author and reviewers: As
> >>>> token-generating machines, LLMs can generate diffs of extraordinary
> size
> >>>> very rapidly. /deep-review is great for chewing through diffs and
> >>>> identifying defects. But it should be used by the contributor
> themselves to
> >>>> identify issues – not to replace the role of the reviewer with more
> >>>> electricity. The role of the reviewers extends beyond identifying and
> >>>> highlighting defects. It encompasses architecture, harmony with the
> >>>> existing codebase, thinking ahead to future evolution of the project,
> and
> >>>> replicates context on the project as new code is committed. These
> functions
> >>>> cannot be automated away.
> >>>>>>>>> – Hesitancy of authors to engage manually with code they have
> >>>> generated: This is not specific to Cassandra, but it is a behavior
> that I
> >>>> have seen in several “highly-electric” projects. There’s a bimodal
> tendency
> >>>> toward code that is entirely generated or entirely human-authored -
> but it
> >>>> is rare for someone to prepare an AI-authored patch to take an
> offramp and
> >>>> spend a significant amount of time refining the work by hand in an
> IDE.
> >>>> This hesitancy toward human participation in authorship of
> LLM-generated
> >>>> code is very concerning to me.
> >>>>>>>>> – Harmony with the existing codebase: Due to the tunnel-vision of
> >>>> context windows, LLMs are generally unaware of conventions and norms
> >>>> present in codebases and very frequently reinvent concepts in a
> generation
> >>>> turn to suit a goal without view of the project’s overall
> architecture.
> >>>> This results in a profusion of messy and duplicated concepts that
> gradually
> >>>> sprawl about a codebase.
> >>>>>>>>> Again, none of these are grounds for prohibition of usage of
> >>>> language models in developing the project. They’re just problems we
> need to
> >>>> bear in mind and guard against – and I think the proposal is designed
> to do
> >>>> just that.
> >>>>>>>>> I’m thrilled by the potential of LLMs to improve Apache Cassandra
> >>>> and we already see it happening through a vast number of issues that
> are
> >>>> being reported and fixed. But there’s also danger in taking ATVs down
> a
> >>>> hiking trail full of people.
> >>>>>>>>> Regarding the prohibition on prose, I’ll simply say: I recently
> >>>> found myself in a scenario where I found a Claude-authored document so
> >>>> inscrutable that I piped it back into a model, directed it to rewrite
> it in
> >>>> ASD-STE100, read it myself, and responded based on the summarization.
> As a
> >>>> humanities grad, this is probably the worst language crime I have
> >>>> committed. But it was in response to language that was itself so
> >>>> idiosyncratic that it was unreadable to me in its original form. I
> hope
> >>>> this never happens in the Apache Cassandra project.
> >>>>>>>>> I’ll close with a quote from an excellent article written by
> Colin
> >>>> Breck, an engineer who works on large-scale data systems:
> >>>> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/
> >>>>>>>>> Colin wrote:
> >>>>>>>>>> I don’t want to live in a world where you use AI to summarize
> >>>> something important into unreadable text, and then I use AI in an
> attempt
> >>>> to decipher it. I want to hear you, imperfections and all. I want your
> >>>> interpretation of aesthetics, beauty, quality, relationship, time. I
> want
> >>>> to know how you feel. I want you to cut through and tell me what
> really
> >>>> matters.
> >>>>>>>>>> Intentional writing will likely become more valuable. People who
> >>>> write, and write to think, to think deeply and carefully, or to
> create, to
> >>>> share, or to capture something important without explicitly
> expressing it
> >>>> will continue to write and produce original work. The people who
> never were
> >>>> writers will use AI to produce lots of text.
> >>>>>>>>> I hope that our culture can remain one of intentional writing and
> >>>> intentional engineering. I enjoy reading the voice of the author in
> >>>> comments, code, and tickets in Cassandra – the different ways we use
> >>>> language based on where we grew up and how we learned English, the
> >>>> translated idioms from our various backgrounds, and terse comments
> that
> >>>> recognize the difference between code whose function is obvious and
> what
> >>>> warrants genuine exposition. When I read code in Cassandra, it’s a
> delight
> >>>> to recognize the author based on their writing style before flipping
> on
> >>>> `git annotate` to reveal the origin.
> >>>>>>>>> I’d encourage folks to re-read the original proposal below. It is
> >>>> very permissive. The guidance strikes me not just as reasonable, but
> >>>> genuinely important to maintaining the health of the project.
> >>>>>>>>> – Scott
> >>>>>>>>> =====
> >>>>>>>>> Encouraged:
> >>>>>>>>> - Reviewing and otherwise validating human-authored patches
> before
> >>>> submission
> >>>>>>>>> - Debugging, diagnosing etc
> >>>>>>>>> Permitted:
> >>>>>>>>> - Generating or modifying tests, scripts, tooling or any other
> >>>> non-user facing changes
> >>>>>>>>> - Minor changes to human-authored patches that are carefully
> >>>> reviewed by the author
> >>>>>>>>> Restricted:
> >>>>>>>>> - Core code changes made by LLM may only be proposed by
> contributors
> >>>> with demonstrated expertise
> >>>>>>>>> - Must have produced similar patches in size, scope and area
> >>>> unassisted and with minimal third-party guidance
> >>>>>>>>> - Core code changes made by LLM require an additional reviewer
> >>>>>>>>> - LLM review is not a substitute for human review, and must be
> used
> >>>> only to augment a complete and independent human understanding of the
> patch.
> >>>>>>>>> Prohibited:
> >>>>>>>>> - All public prose must be human authored. This includes inline
> >>>> comments, docs, posts to Jira etc.
> >>>>>>>>> All LLM generated changes MUST be disclosed:
> >>>>>>>>> - Outlined to any reviewer;
> >>>>>>>>> - Summarised in the commit message;
> >>>>>>>>> - Large blocks or files must be individually marked with some
> agreed
> >>>> message like "created by <some AI>"
> >>>>>>>>> =====
> >>>>>>>>>> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected]
> >>>> <mailto:[email protected]>> wrote:
> >>>>>>>>>> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected]
> >>>> <mailto:[email protected]>> wrote:
> >>>>>>>>>>> Restricted:
> >>>>>>>>>>> - Core code changes made by LLM may only be proposed by
> >>>> contributors with demonstrated expertise
> >>>>>>>>>>> - Must have produced similar patches in size, scope and area
> >>>> unassisted and with minimal third-party guidance
> >>>>>>>>>> I am -1 on this. This sounds like gate keeping attempt. It
> narrowly
> >>>> limits the pool to a few people on the project that have historically
> >>>> contributed to certain parts of the codebase. This policy will
> prohibit
> >>>> skilled software engineers with domain expertise from proposing LLM
> >>>> assisted changes simply because they have not contributed to the
> project.
> >>>> This is unrealistic and a net negative for the project to attract
> talent
> >>>> and grow our community.
> >>>>>>>>>>> - Core code changes made by LLM require an additional reviewer
> >>>>>>>>>> Can you be more precise what is this in addition to? How many
> total
> >>>> reviewers do you expect and what is the purpose of additional
> reviewer? and
> >>>> why?
> >>>>>>>>>> Taking a step back - what are you trying to solve here?
> >>>>>>>>>> Dinesh
>

Reply via email to