While I too want to lower the bar for contributors to join us, I am happy to see something like the proposed policy and I don’t think reading a short document is too much to ask when contributing to our large and critical code base. We’ve had and have much larger barriers to entry than that. One reason I’m excited about AI use in the project is I think it can help us lower more of those.
While I agree with the spirit of the proposed policy and some of what’s in it I think as written it will have us back here often to re-discuss this topic for a couple reasons. First, “strongly discouraged” will have different meanings to different community members and as the policy acknowledges it’s unenforceable so this will lead to debates based on the ambiguities in the text. Second, all of us are on various points on a spectrum on where we see AIs abilities today and where we see them going. The policy is written to be a sort of blend of our current opinions on where it is today so it seems likely we are back here today as our opinions shift (I know my beliefs change daily to weekly these days in both directions), capabilities change, and new consensus is needed. I propose we do have a document but one more I like of the “motivating and guiding principles section” and that we rely on our other existing policies and assumption of positive intent of contributors. I have proposed some below. I am sure several here will find these too lenient given what I presume to be where they fall on the AI adoption spectrum compared to me and I respect that. Some days I am likely right there with you, others I’m more bullish. You’ll find my proposal below does not encourage trying to delineate parts of the codebase that can and cannot be contributed to with an LLM but holds standards regardless. I encourage us to find ways to set the bar for quality when LLMs are used now and in the future vs trying to limit them based on today’s opinions and capabilities that are rapidly changing. Some proposed ideas for some guiding principles with that in mind. It’s likely not a complete list. * LLMs do not change accountability. You are ultimately accountable for code produced or reviewed in your name. Whether hand written or by an LLM. If you own an agentic process performing coding or review tasks you are still responsible and accountable for what it produces or what actions it takes. We as a community are accountable for understanding and being knowledgeable about the software we provide to others. LLMs do not change this. * Humans must be involved in the merging of code either by producing or reviewing code and is subject to the accountability requirement above. Nothing about using LLMs changes existing policies regarding committers required to merge code, vetoes, or other voting procedures. * All LLM use must be attributed to both you and the harness, provider, and model being used. The project makes no specific recommendations or requirements regarding the toolchain used as long as the user has legal access and provides attribution. * LLM generated code has a higher standard of automated testing than human code, for which we have already adopted an incredibly high standard. Proposed fixes or performance improvements must include runnable demonstrations. Use of LLMs does not absolve the accountable contributor of existing requirements such as providing a test plan in JIRA. LLM generated code must be automatically linted to meet the projects code standards. LLM generated code is not an excuse to ignore the projects existing policies on code style. * It is strongly preferred that documentation and comments are not LLM generated. We have found these to be of low quality. However, if done, the aforementioned accountability remains with the contributor. “The LLM wrote it” is not an acceptable dismissal of a review comment in any context. It is recommended that documentation and comments continue to be human written and optional LLM reviewed or edited with human supervision. On Tue, Sep 29, 2026 at 08:57 Caleb Rackliffe <[email protected]> wrote: > @Josh I tried to define “directly generated” at the bottom, although there > isn’t a proper footnote/link. I don’t think it matters at this point. There > doesn’t appear to be any appetite for something in the middle, i.e. what I > was attempting to do there. > > On Sep 29, 2026, at 10:00 AM, Josh McKenzie <[email protected]> wrote: > > > Some questions that are still unclear to me after reading through this > thread and the PR - and I assume a new contributor would be confused as > well: > > Re: what qualifies as "Directly Generated" by an LLM: > > - If someone generates a full implementation and testing for something > via an LLM then goes through line by line and cleans things up and makes > changes, does that qualify as Directly Generated or not? > - If they have fine-tuned a local model to comments in their own > verbal style, is that strongly discouraged because an LLM generated it? > What if they review it line-by-line? What if they write things by hand then > have an LLM rephrase things and leave the LLM's final directly generated > text in place? > - What happens if 15% of the comments generated by the LLM are > concise, clear, and only explain non-obvious "why's" of the code? Should a > contributor go through and rephrase those lines in order to keep them from > being directly generated? > > If I was a contributor looking for a project to start getting involved > with baroque and bespoke rules would be incredibly off-putting to me. > Honestly, the set of rules we have and social norms today are incredibly > off-putting to many long-term contributors already who have a deep vested > social and professional interest in the project succeeding. Who still would > love to work technically on the project but are driven away by this culture. > > We're trying to hit a middle ground of not being too prescriptive but not > leaving everything open to the interpretation of the reader which is just > breeding more confusion. All in a space where the progress of the > underlying tools is faster than anything I can recall in our field. > Whatever policy we come up with now will probably be slightly outdated even > by the time we ratify it unless it's incredibly high level and instead > tries to codify our *values* and trust people to live up to them. > > Which I'd argue is exactly what Blake's simple proposal does. It's durable > in the face of change and focuses on what's important to us and the > community instead of engaging in pedantry and policing that just ends up > confusing everyone further. > > On Tue, Sep 29, 2026, at 4:42 AM, Aleksey Yeshchenko via dev wrote: > > If anyone wants to follow along and/or add comments, we've created > https://github.com/apache/cassandra/pull/5220 > > This is now quite qualitatively different from "Rust policy but without > the committer exception for critical sections", I'm afraid. > > Watered down beyond what we discussed here and offline, and not what I and > most folks who endorsed a Rust-like policy voted for. > > I'll make some edits to restore it to the shape we discussed last night. > > -- > AY > > On 29 Sep 2026, at 00:17, Caleb Rackliffe <[email protected]> > wrote: > > If anyone wants to follow along and/or add comments, we've created > https://github.com/apache/cassandra/pull/5220 > > On Mon, Sep 28, 2026 at 2:45 PM Aleksey Yeshchenko via dev < > [email protected]> wrote: > > It is a little confusing to keep track of the proposed diffs to Rust's > policy. I think David is preparing a version with all the changes applied > to it, so there is no ambiguity. > > How would we handle the “Non-critical” part of the experimental section? > The policy exempts rust-lang members from that… does this mean we’d exempt > committers but not non-committers. What’s the Cassandra analog of the > non-critical section? > > > The modified proposal removes that paragraph (about exemptions) > altogether, thus treating all C* developers equally and allowing code > generation for non-critical parts only for everyone. > > If you are still confused (which would be understandable - it is > confusing), perhaps wait for David's doc, to make sure we are all on the > same page wrt what's being proposed first. > > -- > AY > > On 28 Sep 2026, at 20:16, Blake Eggleston <[email protected]> wrote: > > This is something I could support as well. > > 2 things: > > Our docs tend to be neglected. While ideally our docs would be 100% human > generated, they’re mostly just not generated at the moment. While not > ideal, I think relaxing the rust LLM policy as it relates to docs would be > a net positive for users, provided they’re human reviewed and edited. > > How would we handle the “Non-critical” part of the experimental section? > The policy exempts rust-lang members from that… does this mean we’d exempt > committers but not non-committers. What’s the Cassandra analog of the > non-critical section? > > On Mon, Sep 28, 2026, at 12:10 PM, Francisco Guerrero wrote: > > I've gone over the the Rust policy. I am in support of the Rust > version with the tweaks proposed by Caleb. > > Best, > - Francisco > > On 2026/09/28 18:45:03 Aleksey Yeshchenko via dev wrote: > > Some last minute amends to the suggested policy's TL;DR, with Caleb's > approval: > > > > - It’s fine to use LLMs to answer questions, analyze, distill, refine, > check, suggest, review. > > - LLMs work best when used as a tool to write better, not faster. > > > > "But not to create." bit is covered in detail by the full policy and is > impossible to summarise well in two words. > > > > With other changes as outlined by Caleb in the quoted email, I would be > happy to support this fine-tuned version of Rust's policy. > > > > -- > > AY > > > > > On 28 Sep 2026, at 19:27, Caleb Rackliffe <[email protected]> > wrote: > > > > > > To clarify, I would remove the "Experimental" tag and make that > section apply to all contributors. (In other words, encourage attribution, > quality, and human decision-making for all of us.) > > > > > > The spirit of this is really a one line change to the Rust policy: > > > > > > > It’s fine to use LLMs to answer questions, analyze, distill, refine, > check, suggest, review. But not to create. > > > > > > ...becomes... > > > > > > It’s fine to use LLMs to answer questions, analyze, distill, refine, > check, suggest, review. But not to decide. > > > > > > > > > > > > On Mon, Sep 28, 2026 at 12:54 PM Caleb Rackliffe < > [email protected] <mailto:[email protected]>> wrote: > > >> I finally read the Rust and Lucene policy docs in more detail... > > >> > > >> > https://flagged.apple.com:443/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11 > <https://flagged.apple.com/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11> > < > https://flagged.apple.com/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11 > > > > >> https://github.com/apache/lucene/blob/main/AI_POLICY.md > > >> > > >> I think I agree with a lot of what's written in both, and they > overlap quite a lot, especially around communication (docs, issue comments, > etc.) that should be primarily human-to-human. Everything useful in the > Lucene policy is already included in the Rust policy though. If we could > take the Rust policy, generalize away the Rust-specific things, simplify > it, and remove the "Experimental" tag (and probably the "non-critical" > qualifier) on the "LLM-created code changes intended of review" section, I > think that's something a large majority of us would be able to live with. > > >> > > >> If we can get this right, it's simply clarifying the set of things > contributors (including existing committers) can do to have the best chance > at getting engagement from reviewers. > > >> > > >> I don't know how much appetite there is out there for a formal draft > of this, and we already have 3-4 proposals, but I could attempt it if that > would be useful... > > >> > > >> > > >> On Mon, Sep 28, 2026 at 10:50 AM Štefan Miklošovič < > [email protected] <mailto:[email protected]>> wrote: > > >>> A clarification from my side, I asked "what is wrong with this" in my > > >>> latest email: > > >>> > > >>> "For these reasons, it should be expected that the person producing > > >>> the patch has already demonstrated their expertise and commitment by > > >>> producing and maintaining similar patches without the use of AI" > > >>> > > >>> It is "almost fine", the part of "similar patches without the use of > > >>> AI" should not be there. It should stop before that. > > >>> > > >>> Otherwise this is going to exclude people who have a decade of > > >>> experience with Cassandra and contributed countless patches of > various > > >>> size and complexity while according to that exact wording, they would > > >>> not be eligible to contribute an AI patch. That is silly. I think > this > > >>> is wrong. It does not matter how it was produced. What is important > is > > >>> established trust and if a patch is correct. What does even the size > > >>> of a patch have in common with that? Expertise and commitment! Not > > >>> "the series of patches this committer ever produced was not complex > > >>> enough so we can't take that code in". > > >>> > > >>> Also, who is exactly going to measure that anyway? What are the > > >>> _objective_ criteria who qualifies? Somebody might come and say > "while > > >>> based on my criteria, (because I do not like this person), I do not > > >>> think that the patches of this person qualify, because ...". We need > > >>> hard data on whether it can be merged or not, performance > improvement, > > >>> stability ... > > >>> > > >>> I think this particular wording would need to be refined further. > > >>> > > >>> On Mon, Sep 28, 2026 at 4:34 PM Štefan Miklošovič > > >>> <[email protected] <mailto:[email protected]>> wrote: > > >>> > > > >>> > Right ... for that reason I don't think we should restrict anybody > to > > >>> > create a PR or anything like that, putting some artificial > constraints > > >>> > people will eventually bypass anyway. We don't have that under > > >>> > control. What we have under control is the review part of that. A > > >>> > patch not merged will not be released. The review itself is the > > >>> > "gate". > > >>> > > > >>> > If a PR, even done by AI, is up to standards, has everything it > should > > >>> > have and it is technically correct, then I can not reject to merge > > >>> > that only on the basis it was AI-generated. A patch like a patch. > The > > >>> > code speaks. The ultimate gate is if a patch is correct or not, not > > >>> > how it was produced. > > >>> > > > >>> > Do I gravitate with my trust more towards established members of > the > > >>> > community? Definitely. The trust is earned over the years. > Implicitly, > > >>> > I am trusting a newcomer less. Sorry but not sorry. If somebody > calls > > >>> > this "gating", I don't think they see the nuances enough. Yeah, > call > > >>> > it a gate if you want ... > > >>> > > > >>> > That is why I agree with Benedict here, he said: > > >>> > > > >>> > "For these reasons, it should be expected that the person producing > > >>> > the patch has already demonstrated their expertise and commitment > by > > >>> > producing and maintaining similar patches without the use of AI". > > >>> > > > >>> > What is wrong about this? > > >>> > > > >>> > Look at this contributor (1). This is an excellent example. 10 > patches > > >>> > in fast cadence three weeks ago. We never heard about this person > > >>> > before nor after the patches were created. What about hitting a ML > > >>> > saying "hey, guys, I have a set of patches which scratch my itches, > > >>> > can you take a look, please?". I don't know ... just be a bit ... > > >>> > human about all of this? The maintainers are people too. I am not > > >>> > obliged to take in and cooperate with whoever comes by, dumps their > > >>> > stuff and then they ... wait. Well, so wait. See where you got > three > > >>> > weeks after? Nowhere. > > >>> > > > >>> > Caleb put it nicely, we are "only" humans. > > >>> > > > >>> > (1) > https://github.com/apache/cassandra/pulls?q=is%3Apr+state%3Aopen+author%3Acheeeee > > >>> > > > >>> > On Mon, Sep 28, 2026 at 2:51 PM Shailaja Koppu via dev > > >>> > <[email protected] <mailto:[email protected]>> > wrote: > > >>> > > > > >>> > > Hi Stefan, > > >>> > > > > >>> > > From personal side, I completely agree with you. I am giving > potential options only to address concerns like - new contributors > overwhelming the community with AI generated PRs just to show as add-on in > their profile and vanish after that, or purely AI opened PRs without > developer review or understanding. But the later can happen with anyone > including committers due to workload/deadlines or misled by AI etc. Also, > someone can copy a AI generated patch line by line skipping comments, which > looks like a handwritten code. > > >>> > > > > >>> > > > > >>> > > Thanks, > > >>> > > Shailaja > > >>> > > > > >>> > > > > >>> > > > > >>> > > > On Sep 28, 2026, at 12:23 PM, Štefan Miklošovič < > [email protected] <mailto:[email protected]>> wrote: > > >>> > > > > > >>> > > >> - Only Cassandra committers may submit AI-assisted PRs. This > would mean new contributors first write and understand code without AI > before becoming committers; or > > >>> > > >> - Contributors may submit AI-assisted changes in a > component/subcomponent only after they have submitted at least one non-AI > PR in that component/subcomponent. > > >>> > > > > > >>> > > > I am not sure if I am missing something but can you all > explain in > > >>> > > > simple terms how is this actually enforceable in practice? > > >>> > > > > > >>> > > > "Only Cassandra committers may submit AI-assisted PRs" - there > is no > > >>> > > > restriction who can create a PR and how. It is not like we see > that a > > >>> > > > PR is created with heavy AI usage, then we check if a > contributor is a > > >>> > > > committer and when they are not we comment on that PR saying - > "hold > > >>> > > > your horses mate, we checked the list and you are not a > committer, > > >>> > > > sorry, we have to close this". > > >>> > > > > > >>> > > > If a PR is crafted "carefuly" then it might look like a > completely > > >>> > > > legitimate piece of work while it is still 100% prompted and > the > > >>> > > > author does not have a clue what they did. I mean ... how do > you make > > >>> > > > the difference between what is "real" and what is AI-driven > 100%? I > > >>> > > > think that even if we "guessed" which one is which, the > possibility to > > >>> > > > see this is being progressively erased as this tech is > evolving and we > > >>> > > > will eventually not have a clue. > > >>> > > > > > > > >
