I like your wording Josh. If we move forward with some form of my proposal I would +1 a version with your wording on comments and docs instead of mine.
Jordan On Wed, Sep 30, 2026 at 12:44 Josh McKenzie <[email protected]> wrote: > as someone with both plenty of AI use experience, and more C* experience > than most folks arguing in the opposite direction here. > > Let's try and avoid getting into measuring contests on the thread here > please. :) > > We all have different experiences both inside and outside C* and that > diversity of expertise is a strength. We should be leaning on each other > and sharing our experience working with *any* new technology (be it > LLM's, build systems, reviewing systems, IDE's, etc) to make Cassandra > better. And ourselves. > > @Jordan - I re-read your email and your clause on docs and testing is less > binary than I'd recalled. I'd advocate for changing it a bit though: > > * Pay special attention to documentation and comments; they serve a unique > role in directing attention for future users and maintainers and as such, > blindly generating them with an LLM harms the project and makes things > harder for both immediate reviewers and future maintainers. Contributors > often gloss over the importance of clear comments and documentation, making > them more likely to blindly use LLM generated content here and not suitably > edit or take ownership of them. Prefer writing these by hand, and if you do > use an LLM during generation of any such documentation, the aforementioned > accountability remains with the contributor and the output should be > closely reviewed and edited. > > Maybe nits, but primary things I don't like about the original text: > - "We have found these to be of low quality." <- this is true of both code > *and* docs and comments at various stages of the evolution of this tech > and of harnesses. And the problem is really people generating things and > not reviewing them or taking ownership, and comments and docs are just more > likely to be glossed over. > - "'The LLM wrote it' is not an acceptable dismissal of a review comment > in any context": This is universal too, not just a doc / comment issue. > > On Wed, Sep 30, 2026, at 3:10 PM, Aleksey Yeshchenko via dev wrote: > > What if we break out the discussion on attribution in another thread? It > seems like one of the only things most of us agree should happen in some > capacity. Vote on that and then leave all the more amorphous stuff in this > thread... > > > Please don't. We can talk about attribution if and when we agree on the > crux of it first. > > As someone who is relatively far on the AI adoption journey > > not everyone is at the same stage of AI adoption > > > AI adoption is not a straight line. There is more than one way to hold an > agent. I use Pi+Opus *extensively*, just not to generate *production* code. > This is the case for Benedict as well. > > What I'm advocating for is a way of holding it appropriately - in context > of development of a distributed DBMS like ours, but specifically ours - as > someone with both plenty of AI use experience, and more C* experience than > most folks arguing in the opposite direction here. > > -- > AY > > On 30 Sep 2026, at 19:38, Caleb Rackliffe <[email protected]> > wrote: > > What if we break out the discussion on attribution in another thread? It > seems like one of the only things most of us agree should happen in some > capacity. Vote on that and then leave all the more amorphous stuff in this > thread... > > On Wed, Sep 30, 2026 at 1:29 PM Jon Haddad <[email protected]> > wrote: > > I didn't say it's too hard, not sure where you got that. It's just mildly > annoying that we'd have a list of models used on every commit. I don't see > the value. > > On Wed, Sep 30, 2026 at 11:21 AM Jordan West <[email protected]> wrote: > > I’m not gonna die on the attribution hill past “some kind of attribution / > acknowledgement of LLM use is important” but I think it’s a bit farfetched > to say “we have this tool that can write all these amazing tests and code > but tracking what models are used and writing them down is too hard”. > Skills that can automate development and review can be easily extended to > track models being used. But again, it seems like there is enough concern > for specific attribution that a more general attribution like originally > proposed would be better. I don’t personally agree with the concerns about > specific attribution raised so far but see many possible compromises. I > would encourage us to debate specific language on a proposal (which again > I’m happy to create). > > On testing, a few of you are implying / saying that given LLMs we can > raise the testing bar higher than it is today. I actually agree with that. > But it implies almost a need to use LLMs which would in turn make it harder > for contributors who aren’t at your level of adoption of AI. As someone who > is relatively far on the AI adoption journey I don’t personally see this as > a bad thing but it doesn’t feel like our community is ready for that and > many of the same folks proposing testing can be done better with LLMs are > the same group advocating to not make the bar higher for contributors or > have special classes of contributors. Again personally, I like raising the > testing bar higher but I would encourage us to drive consensus on that > separately so we can reach some form of consensus here. IMO consensus here > will require accepting not everyone is at the same stage of AI adoption and > viewpoints as others might be. If we are trying to get everyone to see AI > the same and then make a policy i think we will have a hard time. > > Jordan > > On Wed, Sep 30, 2026 at 10:43 Jon Haddad <[email protected]> wrote: > > I think the main reason is that it's far less effort to create a > comprehensive system for testing with an LLM than it is by hand. For > example, with cursor compaction, we have parameterized, differential, > fuzzed tests. I was able to go through several iterations and different > ideas quite cheaply, in order to arrive where it is now. Being able to > experiment with multiple paths forward is a huge general advantage, and on > the side of testing it makes it a no brainer. > > LLMs also don't complain when they need to make big revisions, or plumb > things through that would otherwise be an annoyance. They don't mind the > grunt work, and are damn good at it. Add a reviewer that runs PMD and > Jacoco and tells you where the code is either too complex or untested AND > can quickly refactor it, or put together a few hundred lines of test code > is remarkable. > > Jon > > > > On Tue, Sep 29, 2026 at 1:56 PM Dinesh Joshi <[email protected]> wrote: > > Jordan, thanks for the guiding principles. I think this is a short-enough > prose that I expect contributors to realistic read, digest and apply. > > One clarification though - why is the bar for LLM generated code higher > than human generated code? Why isn't the bar the same for both? Why does it > have to be a function of the thing that generated the code? > > Realistically, most developers are using LLM assistance is writing tests. > Like Caleb said, the cost of writing tests has fallen to zero. If anything, > I would say that the bar for tests should be higher regardless of who > generated the code. > > An unintended side effect of having a higher test bar for LLM generated > code is that it may incentivize contributors to hide / underplay the use of > LLMs. > > > > On Tue, Sep 29, 2026 at 1:04 PM Jordan West <[email protected]> wrote: > > On Tue, Sep 29, 2026 at 12:54 Caleb Rackliffe <[email protected]> > wrote: > > @Jordan I agree with essentially everything you’ve said here. The only > exception is that I wouldn’t want to lower the testing bar (assuming we > agree on what that means) for non-LLM-authored patches. The cost of writing > tests has esssntially fallen to zero. > > > > > > We’re absolutely aligned on that. If I implied otherwise that was an error > on my part. My only intent is the bar should be even higher > for LLM generated code than whatever the bar is for human generated code > not that we lower the human generated bar we have or agree to in the > future. Maybe we have to wordsmith that some to be more clear? > > > > On Sep 29, 2026, at 2:30 PM, Jordan West <[email protected]> wrote: > > > While I too want to lower the bar for contributors to join us, I am happy > to see something like the proposed policy and I don’t think reading a short > document is too much to ask when contributing to our large and critical > code base. We’ve had and have much larger barriers to entry than that. One > reason I’m excited about AI use in the project is I think it can help us > lower more of those. > > While I agree with the spirit of the proposed policy and some of what’s in > it I think as written it will have us back here often to re-discuss this > topic for a couple reasons. First, “strongly discouraged” will have > different meanings to different community members and as the policy > acknowledges it’s unenforceable so this will lead to debates based on the > ambiguities in the text. Second, all of us are on various points on a > spectrum on where we see AIs abilities today and where we see them going. > The policy is written to be a sort of blend of our current opinions on > where it is today so it seems likely we are back here today as our opinions > shift (I know my beliefs change daily to weekly these days in both > directions), capabilities change, and new consensus is needed. > > I propose we do have a document but one more I like of the “motivating and > guiding principles section” and that we rely on our other existing policies > and assumption of positive intent of contributors. I have proposed some > below. I am sure several here will find these too lenient given what I > presume to be where they fall on the AI adoption spectrum compared to me > and I respect that. Some days I am likely right there with you, others I’m > more bullish. You’ll find my proposal below does not encourage trying to > delineate parts of the codebase that can and cannot be contributed to with > an LLM but holds standards regardless. I encourage us to find ways to set > the bar for quality when LLMs are used now and in the future vs trying to > limit them based on today’s opinions and capabilities that are rapidly > changing. > > Some proposed ideas for some guiding principles with that in mind. It’s > likely not a complete list. > > * LLMs do not change accountability. You are ultimately accountable for > code produced or reviewed in your name. Whether hand written or by an LLM. > If you own an agentic process performing coding or review tasks you are > still responsible and accountable for what it produces or what actions it > takes. We as a community are accountable for understanding and being > knowledgeable about the software we provide to others. LLMs do not change > this. > > * Humans must be involved in the merging of code either by producing or > reviewing code and is subject to the accountability requirement above. > Nothing about using LLMs changes existing policies regarding committers > required to merge code, vetoes, or other voting procedures. > > * All LLM use must be attributed to both you and the harness, provider, > and model being used. The project makes no specific recommendations or > requirements regarding the toolchain used as long as the user has legal > access and provides attribution. > > * LLM generated code has a higher standard of automated testing than human > code, for which we have already adopted an incredibly high standard. > Proposed fixes or performance improvements must include runnable > demonstrations. Use of LLMs does not absolve the accountable contributor of > existing requirements such as providing a test plan in JIRA. LLM generated > code must be automatically linted to meet the projects code standards. LLM > generated code is not an excuse to ignore the projects existing policies on > code style. > > * It is strongly preferred that documentation and comments are not LLM > generated. We have found these to be of low quality. However, if done, the > aforementioned accountability remains with the contributor. “The LLM wrote > it” is not an acceptable dismissal of a review comment in any context. It > is recommended that documentation and comments continue to be human written > and optional LLM reviewed or edited with human supervision. > > On Tue, Sep 29, 2026 at 08:57 Caleb Rackliffe <[email protected]> > wrote: > > @Josh I tried to define “directly generated” at the bottom, although there > isn’t a proper footnote/link. I don’t think it matters at this point. There > doesn’t appear to be any appetite for something in the middle, i.e. what I > was attempting to do there. > > On Sep 29, 2026, at 10:00 AM, Josh McKenzie <[email protected]> wrote: > > > Some questions that are still unclear to me after reading through this > thread and the PR - and I assume a new contributor would be confused as > well: > > Re: what qualifies as "Directly Generated" by an LLM: > > - If someone generates a full implementation and testing for something > via an LLM then goes through line by line and cleans things up and makes > changes, does that qualify as Directly Generated or not? > - If they have fine-tuned a local model to comments in their own > verbal style, is that strongly discouraged because an LLM generated it? > What if they review it line-by-line? What if they write things by hand then > have an LLM rephrase things and leave the LLM's final directly generated > text in place? > - What happens if 15% of the comments generated by the LLM are > concise, clear, and only explain non-obvious "why's" of the code? Should a > contributor go through and rephrase those lines in order to keep them from > being directly generated? > > If I was a contributor looking for a project to start getting involved > with baroque and bespoke rules would be incredibly off-putting to me. > Honestly, the set of rules we have and social norms today are incredibly > off-putting to many long-term contributors already who have a deep vested > social and professional interest in the project succeeding. Who still would > love to work technically on the project but are driven away by this culture. > > We're trying to hit a middle ground of not being too prescriptive but not > leaving everything open to the interpretation of the reader which is just > breeding more confusion. All in a space where the progress of the > underlying tools is faster than anything I can recall in our field. > Whatever policy we come up with now will probably be slightly outdated even > by the time we ratify it unless it's incredibly high level and instead > tries to codify our *values* and trust people to live up to them. > > Which I'd argue is exactly what Blake's simple proposal does. It's durable > in the face of change and focuses on what's important to us and the > community instead of engaging in pedantry and policing that just ends up > confusing everyone further. > > On Tue, Sep 29, 2026, at 4:42 AM, Aleksey Yeshchenko via dev wrote: > > If anyone wants to follow along and/or add comments, we've created > https://github.com/apache/cassandra/pull/5220 > > This is now quite qualitatively different from "Rust policy but without > the committer exception for critical sections", I'm afraid. > > Watered down beyond what we discussed here and offline, and not what I and > most folks who endorsed a Rust-like policy voted for. > > I'll make some edits to restore it to the shape we discussed last night. > > -- > AY > > On 29 Sep 2026, at 00:17, Caleb Rackliffe <[email protected]> > wrote: > > If anyone wants to follow along and/or add comments, we've created > https://github.com/apache/cassandra/pull/5220 > > On Mon, Sep 28, 2026 at 2:45 PM Aleksey Yeshchenko via dev < > [email protected]> wrote: > > It is a little confusing to keep track of the proposed diffs to Rust's > policy. I think David is preparing a version with all the changes applied > to it, so there is no ambiguity. > > How would we handle the “Non-critical” part of the experimental section? > The policy exempts rust-lang members from that… does this mean we’d exempt > committers but not non-committers. What’s the Cassandra analog of the > non-critical section? > > > The modified proposal removes that paragraph (about exemptions) > altogether, thus treating all C* developers equally and allowing code > generation for non-critical parts only for everyone. > > If you are still confused (which would be understandable - it is > confusing), perhaps wait for David's doc, to make sure we are all on the > same page wrt what's being proposed first. > > -- > AY > > On 28 Sep 2026, at 20:16, Blake Eggleston <[email protected]> wrote: > > This is something I could support as well. > > 2 things: > > Our docs tend to be neglected. While ideally our docs would be 100% human > generated, they’re mostly just not generated at the moment. While not > ideal, I think relaxing the rust LLM policy as it relates to docs would be > a net positive for users, provided they’re human reviewed and edited. > > How would we handle the “Non-critical” part of the experimental section? > The policy exempts rust-lang members from that… does this mean we’d exempt > committers but not non-committers. What’s the Cassandra analog of the > non-critical section? > > On Mon, Sep 28, 2026, at 12:10 PM, Francisco Guerrero wrote: > > I've gone over the the Rust policy. I am in support of the Rust > version with the tweaks proposed by Caleb. > > Best, > - Francisco > > On 2026/09/28 18:45:03 Aleksey Yeshchenko via dev wrote: > > Some last minute amends to the suggested policy's TL;DR, with Caleb's > approval: > > > > - It’s fine to use LLMs to answer questions, analyze, distill, refine, > check, suggest, review. > > - LLMs work best when used as a tool to write better, not faster. > > > > "But not to create." bit is covered in detail by the full policy and is > impossible to summarise well in two words. > > > > With other changes as outlined by Caleb in the quoted email, I would be > happy to support this fine-tuned version of Rust's policy. > > > > -- > > AY > > > > > On 28 Sep 2026, at 19:27, Caleb Rackliffe <[email protected]> > wrote: > > > > > > To clarify, I would remove the "Experimental" tag and make that > section apply to all contributors. (In other words, encourage attribution, > quality, and human decision-making for all of us.) > > > > > > The spirit of this is really a one line change to the Rust policy: > > > > > > > It’s fine to use LLMs to answer questions, analyze, distill, refine, > check, suggest, review. But not to create. > > > > > > ...becomes... > > > > > > It’s fine to use LLMs to answer questions, analyze, distill, refine, > check, suggest, review. But not to decide. > > > > > > > > > > > > On Mon, Sep 28, 2026 at 12:54 PM Caleb Rackliffe < > [email protected] <mailto:[email protected]>> wrote: > > >> I finally read the Rust and Lucene policy docs in more detail... > > >> > > >> > https://flagged.apple.com:443/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11 > <https://flagged.apple.com/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11> > < > https://flagged.apple.com/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11 > > > > >> https://github.com/apache/lucene/blob/main/AI_POLICY.md > > >> > > >> I think I agree with a lot of what's written in both, and they > overlap quite a lot, especially around communication (docs, issue comments, > etc.) that should be primarily human-to-human. Everything useful in the > Lucene policy is already included in the Rust policy though. If we could > take the Rust policy, generalize away the Rust-specific things, simplify > it, and remove the "Experimental" tag (and probably the "non-critical" > qualifier) on the "LLM-created code changes intended of review" section, I > think that's something a large majority of us would be able to live with. > > >> > > >> If we can get this right, it's simply clarifying the set of things > contributors (including existing committers) can do to have the best chance > at getting engagement from reviewers. > > >> > > >> I don't know how much appetite there is out there for a formal draft > of this, and we already have 3-4 proposals, but I could attempt it if that > would be useful... > > >> > > >> > > >> On Mon, Sep 28, 2026 at 10:50 AM Štefan Miklošovič < > [email protected] <mailto:[email protected]>> wrote: > > >>> A clarification from my side, I asked "what is wrong with this" in my > > >>> latest email: > > >>> > > >>> "For these reasons, it should be expected that the person producing > > >>> the patch has already demonstrated their expertise and commitment by > > >>> producing and maintaining similar patches without the use of AI" > > >>> > > >>> It is "almost fine", the part of "similar patches without the use of > > >>> AI" should not be there. It should stop before that. > > >>> > > >>> Otherwise this is going to exclude people who have a decade of > > >>> experience with Cassandra and contributed countless patches of > various > > >>> size and complexity while according to that exact wording, they would > > >>> not be eligible to contribute an AI patch. That is silly. I think > this > > >>> is wrong. It does not matter how it was produced. What is important > is > > >>> established trust and if a patch is correct. What does even the size > > >>> of a patch have in common with that? Expertise and commitment! Not > > >>> "the series of patches this committer ever produced was not complex > > >>> enough so we can't take that code in". > > >>> > > >>> Also, who is exactly going to measure that anyway? What are the > > >>> _objective_ criteria who qualifies? Somebody might come and say > "while > > >>> based on my criteria, (because I do not like this person), I do not > > >>> think that the patches of this person qualify, because ...". We need > > >>> hard data on whether it can be merged or not, performance > improvement, > > >>> stability ... > > >>> > > >>> I think this particular wording would need to be refined further. > > >>> > > >>> On Mon, Sep 28, 2026 at 4:34 PM Štefan Miklošovič > > >>> <[email protected] <mailto:[email protected]>> wrote: > > >>> > > > >>> > Right ... for that reason I don't think we should restrict anybody > to > > >>> > create a PR or anything like that, putting some artificial > constraints > > >>> > people will eventually bypass anyway. We don't have that under > > >>> > control. What we have under control is the review part of that. A > > >>> > patch not merged will not be released. The review itself is the > > >>> > "gate". > > >>> > > > >>> > If a PR, even done by AI, is up to standards, has everything it > should > > >>> > have and it is technically correct, then I can not reject to merge > > >>> > that only on the basis it was AI-generated. A patch like a patch. > The > > >>> > code speaks. The ultimate gate is if a patch is correct or not, not > > >>> > how it was produced. > > >>> > > > >>> > Do I gravitate with my trust more towards established members of > the > > >>> > community? Definitely. The trust is earned over the years. > Implicitly, > > >>> > I am trusting a newcomer less. Sorry but not sorry. If somebody > calls > > >>> > this "gating", I don't think they see the nuances enough. Yeah, > call > > >>> > it a gate if you want ... > > >>> > > > >>> > That is why I agree with Benedict here, he said: > > >>> > > > >>> > "For these reasons, it should be expected that the person producing > > >>> > the patch has already demonstrated their expertise and commitment > by > > >>> > producing and maintaining similar patches without the use of AI". > > >>> > > > >>> > What is wrong about this? > > >>> > > > >>> > Look at this contributor (1). This is an excellent example. 10 > patches > > >>> > in fast cadence three weeks ago. We never heard about this person > > >>> > before nor after the patches were created. What about hitting a ML > > >>> > saying "hey, guys, I have a set of patches which scratch my itches, > > >>> > can you take a look, please?". I don't know ... just be a bit ... > > >>> > human about all of this? The maintainers are people too. I am not > > >>> > obliged to take in and cooperate with whoever comes by, dumps their > > >>> > stuff and then they ... wait. Well, so wait. See where you got > three > > >>> > weeks after? Nowhere. > > >>> > > > >>> > Caleb put it nicely, we are "only" humans. > > >>> > > > >>> > (1) > https://github.com/apache/cassandra/pulls?q=is%3Apr+state%3Aopen+author%3Acheeeee > > >>> > > > >>> > On Mon, Sep 28, 2026 at 2:51 PM Shailaja Koppu via dev > > >>> > <[email protected] <mailto:[email protected]>> > wrote: > > >>> > > > > >>> > > Hi Stefan, > > >>> > > > > >>> > > From personal side, I completely agree with you. I am giving > potential options only to address concerns like - new contributors > overwhelming the community with AI generated PRs just to show as add-on in > their profile and vanish after that, or purely AI opened PRs without > developer review or understanding. But the later can happen with anyone > including committers due to workload/deadlines or misled by AI etc. Also, > someone can copy a AI generated patch line by line skipping comments, which > looks like a handwritten code. > > >>> > > > > >>> > > > > >>> > > Thanks, > > >>> > > Shailaja > > >>> > > > > >>> > > > > >>> > > > > >>> > > > On Sep 28, 2026, at 12:23 PM, Štefan Miklošovič < > [email protected] <mailto:[email protected]>> wrote: > > >>> > > > > > >>> > > >> - Only Cassandra committers may submit AI-assisted PRs. This > would mean new contributors first write and understand code without AI > before becoming committers; or > > >>> > > >> - Contributors may submit AI-assisted changes in a > component/subcomponent only after they have submitted at least one non-AI > PR in that component/subcomponent. > > >>> > > > > > >>> > > > I am not sure if I am missing something but can you all > explain in > > >>> > > > simple terms how is this actually enforceable in practice? > > >>> > > > > > >>> > > > "Only Cassandra committers may submit AI-assisted PRs" - there > is no > > >>> > > > restriction who can create a PR and how. It is not like we see > that a > > >>> > > > PR is created with heavy AI usage, then we check if a > contributor is a > > >>> > > > committer and when they are not we comment on that PR saying - > "hold > > >>> > > > your horses mate, we checked the list and you are not a > committer, > > >>> > > > sorry, we have to close this". > > >>> > > > > > >>> > > > If a PR is crafted "carefuly" then it might look like a > completely > > >>> > > > legitimate piece of work while it is still 100% prompted and > the > > >>> > > > author does not have a clue what they did. I mean ... how do > you make > > >>> > > > the difference between what is "real" and what is AI-driven > 100%? I > > >>> > > > think that even if we "guessed" which one is which, the > possibility to > > >>> > > > see this is being progressively erased as this tech is > evolving and we > > >>> > > > will eventually not have a clue. > > >>> > > > > > > > >
