And something I forgot to add. The ASF has some new tooling which might actually help us here:
https://magpie.apache.org/ On Wed, Sep 23, 2026 at 12:16 PM Patrick McFadin <[email protected]> wrote: > I was waiting for this moment to hit our project and I'm glad we're > here. I am deeply concerned for our project and its future, as we have > increasingly made it difficult to contribute. I had hoped that this new era > of software tools powered by AI would expand the project's reach and bring > more diverse thoughts and ideas. This policy proposal is the exact opposite > of what we need. We have been sitting on a Cassandra 6 release alpha for > months. We need to accelerate and embrace new ways of being or be left > behind. As I read that policy, my first and gut level reactions: > > - It comes across as elitist and class protectionism. Committer should > not be special but this proposal makes that designation even more sacred. > - It signals that our project is so fragile that only a few > people "Really understand it" That's some SQLite vibes right there. > - Trying to fix a problem that doesn't exist > > Sadly, i think this policy change would also exclude a lot of comitters. > > We aren't alone in this moment. The Linux project just went through > this. You can find the thread with a simple Google, but similar hard > feelings were being expressed "AI is going to ruin our project!", "The > unwashed masses are going to contribute terrible code!", "We have to > protect our precious status as Linux maintainers!" Linus being Linus, was > deeply invloved and they adopted a super simple statement that covers all > bases. Human or Human using AI. “You are expected to understand and to be > able to defend everything you submit.” Love that. > > In the larger picture, I'll restate. I'm worried for our project. In late > 2025(Opus 4.5 IYKYK), early 2026, AI coding LLMs turned a real corner and > in the hands of somebody that knows how to build software, this tool is > like jet fuel. Here's some examples of new projects being hyper fueled by > AI coding tools. > > Apache Iggy - Complete rust replacement of kafka. Crazy fast velocity > Turso - Rust re-write of SQLite > Bun - Rust re-write of itself from Zig. > > Think this couldn't happen to us? Already has: > https://github.com/ferrosadb/ferrosa. Ben is using it to power his own > startup, but it was him alone using a ton of local AI coding agents. He > even implemented Accord. Yeah... > > The cracks are already starting to show. There is a black market economy > of Cassandra patches happening now. Not going to name names or call people > out, but there are fixes and optimizations living in branches outside of > the Cassandra project. Why? I'll use myself as an example. I fixed a nasty > bug I ran into with TCM a few weeks ago. Wrote the tests. It passes CI and > lives in my personal branch. I'm sitting here really wondering if I want to > go through the ritual humiliation of being roasted for using AI to fix it. > Me. I am worried about contrinuting code the Cassandra. What the hell does > that say? > > I have my CQLite project that I've been doing a release around once a > month. I would love to donate that to the Cassandra project but I wouldn't > if it essentially killed any progress. > > My larger counter proposal would be to: > - Adopt the “You are expected to understand and to be able to defend > everything you submit.” approach the Linux project has adopted. > - Loosen up the contributor process and our worry on trunk. Let 1000 > flowers bloom and bring it in. > - And finally, to give some people more peace of mind and open more > doors, adopt what other projects have done and provide more pluggability. > Let new ideas have an easy place to connect. > > We are at a fork in the road. What are we going to do? And then I have to > ask myself, what am I going to do as a contributor? > > Patrick > > On Wed, Sep 23, 2026 at 6:16 AM Blake Eggleston <[email protected]> > wrote: > >> I’m not necessarily opposed to having a policy, but so far we have some >> specific proposals addressing a problem statement that’s very nebulous. >> What is the community failing to do on its own that we’re trying to correct >> with policy? What outcomes are we trying to create or prevent? Having some >> examples and specific problems to discuss would help focus the conversation. >> >> On Wed, Sep 23, 2026, at 4:34 AM, Shailaja Koppu via dev wrote: >> >> Benedict, >> >> Thanks for clarifying. My concern still remains. This criteria would be >> difficult to define and apply consistently. What counts as “similar” scope >> or area, “mostly correct,” or sufficiently independent work? More >> importantly, how do we prevent such vague criteria from creating an >> informal hierarchy where some contributors work is routinely accepted while >> others is routinely rejected? >> >> If the intent is to limit AI-assisted code changes to Cassandra >> contributors, or to contributors who have previously worked in that >> component without AI, that would at least be clear and enforceable. >> >> >> >> >> > On Sep 23, 2026, at 12:01 PM, Benedict Elliott Smith < >> [email protected]> wrote: >> > >> > Core code changes >> > Chris: Do you object to the first or second line you quote? Because the >> first line is effectively motivation for the second line, and can be >> removed (or more clearly combined). If it’s the second line, then I do not >> think this is an unreasonable expectation, and we can get into a proper >> debate about it. >> > >> > Shailaja, since you only snipped the first sentence, your concerns >> might also be mostly answered by this clarification? “Minimal third-party >> guidance” implies you have some concerns about the second line, but all of >> our policies have some ambiguity because legalese is even worse. I don’t >> think the ambiguity here would be challenging to navigate though we can >> certainly refine it. This specific snippet is meant to convey an >> expectation that a contributor has autonomously produced patches of similar >> scope that were mostly correct, so that they have demonstrated the level of >> understanding necessary to guide another party to a successful patch (i.e. >> an LLM in this case). >> > >> > >> > On 2026/09/23 10:54:16 Benedict Elliott Smith wrote: >> >> Thanks everyone for your input so far. I’ll respond in brief to the >> main themes, in (mostly) separate emails so they can each have their own >> debate chain. >> >> >> >> Should we have a policy (Blake/Josh*/Jon/Dinesh) >> >> I think we would all agree that LLMs represent the biggest change to >> this community (and software more generally) since its inception, and we >> all now have enough experience with the technology to have formed opinions >> about how it is best managed. We also evidently have not all arrived at the >> same conclusions. In this situation, it would be an abdication of our >> responsibilities as a management committee to not agree *some* policy. >> >> >> >> I intend to conduct straw polls as the discussion evolves, so if you >> prefer an alternative policy - or modifications to this policy - I would >> encourage you to make those alternative proposals. >> >> >> >> *Veto/Consensus (Josh) >> >> It was fair to call out my poor use of language on this topic, so let >> me rephrase a little. The community is built on consensus, and work should >> not be merged when there are outstanding concerns to address. The explicit >> -1 should only be used rarely, because the prior expectation should prevent >> it ever being needed. I (and others) have outstanding concerns on LLM >> generated work that can only be addressed through this process right here, >> so to merge such work while maintaining the community’s consensus we must >> agree some policy. >> >> >> >> >> >> >> >> On 2026/09/23 09:58:27 Shailaja Koppu via dev wrote: >> >>> I am strongly -1 on this >> >>> - Core code changes made by LLM may only be proposed by contributors >> with demonstrated expertise >> >>> That creates a new, subjective privileged class of contributors and >> turns a tool choice into an eligibility test. Who decides whether expertise >> has been “demonstrated,” what counts as “minimal third-party guidance,” and >> how could those judgments be applied consistently or fairly? >> >>> >> >>> Apache already has a better model, anyone may contribute, trust and >> additional repository privileges are earned transparently over time. The >> ASF describes its communities as flat, and says that newcomer ideas have as >> much input as those from original creators. We should not add a separate, >> informal hierarchy in which certain people may use common development tools >> while others may not. >> >>> >> >>> >> >>> >> >>> >> >>>> On Sep 23, 2026, at 6:33 AM, Chris Lohfink <[email protected]> >> wrote: >> >>>> >> >>>> >> >>>> - Core code changes made by LLM may only be proposed by contributors >> with demonstrated expertise >> >>>> - Must have produced similar patches in size, scope and area >> unassisted and with minimal third-party guidance >> >>>> >> >>>> I really don't like this one or its wording. Definitely too "the >> peasants are getting uppity lets build a wall". Lets not let a subjective >> thing like demonstrated expertise (who decides that?) be if it's ok or not. >> Hold the same standards for code quality and process for it all. I don't >> want this to be: only people on the storage team in Apple can use AI. >> >>>> >> >>>> Chris >> >>>> >> >>>> On Wed, Sep 23, 2026 at 12:16 AM <[email protected] <mailto: >> [email protected]>> wrote: >> >>>>> I agree with Stefan and think this is both a reasonable and >> thoughtful proposal. >> >>>>> >> >>>>> Here are some things I like about it: >> >>>>> >> >>>>> – It outlines areas where LLM usage is unambiguously useful to the >> project’s developers and users. >> >>>>> – It defines a spectrum of recommendations and cautions. >> >>>>> – The only prohibited areas are extremely narrow and say nothing >> about code at all. >> >>>>> >> >>>>> Some in this thread are responding as if this proposal seeks to >> prohibit or sharply limit use of LLMs. In fact, it’s one of the most open >> and welcoming I’ve seen for an OSS project of our size where many are >> adopting policies that simply ban them entirely. I’ve re-appended the >> proposal below my message as it seems to have been lost in threaded >> replies, and would encourage folks to give it a second read. >> >>>>> >> >>>>> Some brief thoughts based on my own use of LLMs: >> >>>>> >> >>>>> – I find them fantastically useful for reviewing and identifying >> problems that have slipped through review - primarily via Alex Petrov’s >> /deep-review skill, which I have running in a VM in a loop executing over >> every new commit in the project as of a few days ago. I will be posting a >> few hand-authored Jira tickets based on findings that appear legitimate to >> me. For now, the loop is posting them as issue drafts for my own review on >> my personal fork which you can find here: >> https://github.com/cscotta/cassandra/issues?q=is%3Aissue%20state%3Aopen%20label%3Abug >> >>>>> – They’re great for enabling use of model checkers and formal >> methods where such work would have previously been prohibitively expensive, >> such as Blake’s work on a TLA+ proof of aspects of Mutation Tracking and >> Benedict/Fedor’s work on a machine-checkable proof of the Accord protocol >> in Lean. >> >>>>> – They are stunning for allowing me to experiment with ideas that >> would have otherwise been a summer internship’s scope of work. Some >> examples include an io_uring prototype, exploring the impact of >> page-aligned compressed chunk sizes, an API shim bridging the 3.x and 4.x >> Java Drivers, and potential enhancements to Zstandard. >> >>>>> – And they shine when given grunt-work that is critical to the >> project but a miserable labor for humans, such as triaging, reproducing, >> and root-causing flaky tests, which David Capwell now has running in a loop >> to help us improve CI stability in the project. >> >>>>> >> >>>>> I never thought I’d be so positive on what’s possible via language >> models a year ago. At the same time, I also agree that they present >> challenges and risks that can be managed through thoughtful discussion and >> policy. Some of the concerns that I think are important to guard against >> include: >> >>>>> >> >>>>> – Asymmetry of effort between author and reviewers: As >> token-generating machines, LLMs can generate diffs of extraordinary size >> very rapidly. /deep-review is great for chewing through diffs and >> identifying defects. But it should be used by the contributor themselves to >> identify issues – not to replace the role of the reviewer with more >> electricity. The role of the reviewers extends beyond identifying and >> highlighting defects. It encompasses architecture, harmony with the >> existing codebase, thinking ahead to future evolution of the project, and >> replicates context on the project as new code is committed. These functions >> cannot be automated away. >> >>>>> – Hesitancy of authors to engage manually with code they have >> generated: This is not specific to Cassandra, but it is a behavior that I >> have seen in several “highly-electric” projects. There’s a bimodal tendency >> toward code that is entirely generated or entirely human-authored - but it >> is rare for someone to prepare an AI-authored patch to take an offramp and >> spend a significant amount of time refining the work by hand in an IDE. >> This hesitancy toward human participation in authorship of LLM-generated >> code is very concerning to me. >> >>>>> – Harmony with the existing codebase: Due to the tunnel-vision of >> context windows, LLMs are generally unaware of conventions and norms >> present in codebases and very frequently reinvent concepts in a generation >> turn to suit a goal without view of the project’s overall architecture. >> This results in a profusion of messy and duplicated concepts that gradually >> sprawl about a codebase. >> >>>>> >> >>>>> Again, none of these are grounds for prohibition of usage of >> language models in developing the project. They’re just problems we need to >> bear in mind and guard against – and I think the proposal is designed to do >> just that. >> >>>>> >> >>>>> I’m thrilled by the potential of LLMs to improve Apache Cassandra >> and we already see it happening through a vast number of issues that are >> being reported and fixed. But there’s also danger in taking ATVs down a >> hiking trail full of people. >> >>>>> >> >>>>> Regarding the prohibition on prose, I’ll simply say: I recently >> found myself in a scenario where I found a Claude-authored document so >> inscrutable that I piped it back into a model, directed it to rewrite it in >> ASD-STE100, read it myself, and responded based on the summarization. As a >> humanities grad, this is probably the worst language crime I have >> committed. But it was in response to language that was itself so >> idiosyncratic that it was unreadable to me in its original form. I hope >> this never happens in the Apache Cassandra project. >> >>>>> >> >>>>> I’ll close with a quote from an excellent article written by Colin >> Breck, an engineer who works on large-scale data systems: >> https://blog.colinbreck.com/i-dont-want-to-read-what-you-didnt-write/ >> >>>>> >> >>>>> Colin wrote: >> >>>>> >> >>>>>> I don’t want to live in a world where you use AI to summarize >> something important into unreadable text, and then I use AI in an attempt >> to decipher it. I want to hear you, imperfections and all. I want your >> interpretation of aesthetics, beauty, quality, relationship, time. I want >> to know how you feel. I want you to cut through and tell me what really >> matters. >> >>>>> >> >>>>>> Intentional writing will likely become more valuable. People who >> write, and write to think, to think deeply and carefully, or to create, to >> share, or to capture something important without explicitly expressing it >> will continue to write and produce original work. The people who never were >> writers will use AI to produce lots of text. >> >>>>> >> >>>>> I hope that our culture can remain one of intentional writing and >> intentional engineering. I enjoy reading the voice of the author in >> comments, code, and tickets in Cassandra – the different ways we use >> language based on where we grew up and how we learned English, the >> translated idioms from our various backgrounds, and terse comments that >> recognize the difference between code whose function is obvious and what >> warrants genuine exposition. When I read code in Cassandra, it’s a delight >> to recognize the author based on their writing style before flipping on >> `git annotate` to reveal the origin. >> >>>>> >> >>>>> I’d encourage folks to re-read the original proposal below. It is >> very permissive. The guidance strikes me not just as reasonable, but >> genuinely important to maintaining the health of the project. >> >>>>> >> >>>>> – Scott >> >>>>> >> >>>>> ===== >> >>>>> Encouraged: >> >>>>> - Reviewing and otherwise validating human-authored patches before >> submission >> >>>>> - Debugging, diagnosing etc >> >>>>> >> >>>>> Permitted: >> >>>>> - Generating or modifying tests, scripts, tooling or any other >> non-user facing changes >> >>>>> - Minor changes to human-authored patches that are carefully >> reviewed by the author >> >>>>> >> >>>>> Restricted: >> >>>>> - Core code changes made by LLM may only be proposed by >> contributors with demonstrated expertise >> >>>>> - Must have produced similar patches in size, scope and area >> unassisted and with minimal third-party guidance >> >>>>> - Core code changes made by LLM require an additional reviewer >> >>>>> - LLM review is not a substitute for human review, and must be used >> only to augment a complete and independent human understanding of the patch. >> >>>>> >> >>>>> Prohibited: >> >>>>> - All public prose must be human authored. This includes inline >> comments, docs, posts to Jira etc. >> >>>>> >> >>>>> All LLM generated changes MUST be disclosed: >> >>>>> - Outlined to any reviewer; >> >>>>> - Summarised in the commit message; >> >>>>> - Large blocks or files must be individually marked with some >> agreed message like "created by <some AI>" >> >>>>> ===== >> >>>>> >> >>>>>> On Sep 22, 2026, at 9:13 PM, Dinesh Joshi <[email protected] >> <mailto:[email protected]>> wrote: >> >>>>>> >> >>>>>> On Tue, Sep 22, 2026 at 3:28 AM Benedict <[email protected] >> <mailto:[email protected]>> wrote: >> >>>>>>> >> >>>>>>> Restricted: >> >>>>>>> - Core code changes made by LLM may only be proposed by >> contributors with demonstrated expertise >> >>>>>>> - Must have produced similar patches in size, scope and area >> unassisted and with minimal third-party guidance >> >>>>>> >> >>>>>> I am -1 on this. This sounds like gate keeping attempt. It >> narrowly limits the pool to a few people on the project that have >> historically contributed to certain parts of the codebase. This policy will >> prohibit skilled software engineers with domain expertise from proposing >> LLM assisted changes simply because they have not contributed to the >> project. This is unrealistic and a net negative for the project to attract >> talent and grow our community. >> >>>>>> >> >>>>>>> - Core code changes made by LLM require an additional reviewer >> >>>>>> >> >>>>>> Can you be more precise what is this in addition to? How many >> total reviewers do you expect and what is the purpose of additional >> reviewer? and why? >> >>>>>> >> >>>>>> Taking a step back - what are you trying to solve here? >> >>>>>> >> >>>>>> Dinesh >> >>>>>> >> >>>>> >> >>> >> >>> >> >> >> >> >>
