> What if we break out the discussion on attribution in another thread? It > seems like one of the only things most of us agree should happen in some > capacity. Vote on that and then leave all the more amorphous stuff in this > thread...
Please don't. We can talk about attribution if and when we agree on the crux of it first. >>> As someone who is relatively far on the AI adoption journey >>> not everyone is at the same stage of AI adoption AI adoption is not a straight line. There is more than one way to hold an agent. I use Pi+Opus *extensively*, just not to generate *production* code. This is the case for Benedict as well. What I'm advocating for is a way of holding it appropriately - in context of development of a distributed DBMS like ours, but specifically ours - as someone with both plenty of AI use experience, and more C* experience than most folks arguing in the opposite direction here. -- AY > On 30 Sep 2026, at 19:38, Caleb Rackliffe <[email protected]> wrote: > > What if we break out the discussion on attribution in another thread? It > seems like one of the only things most of us agree should happen in some > capacity. Vote on that and then leave all the more amorphous stuff in this > thread... > > On Wed, Sep 30, 2026 at 1:29 PM Jon Haddad <[email protected] > <mailto:[email protected]>> wrote: >> I didn't say it's too hard, not sure where you got that. It's just mildly >> annoying that we'd have a list of models used on every commit. I don't see >> the value. >> >> On Wed, Sep 30, 2026 at 11:21 AM Jordan West <[email protected] >> <mailto:[email protected]>> wrote: >>> I’m not gonna die on the attribution hill past “some kind of attribution / >>> acknowledgement of LLM use is important” but I think it’s a bit farfetched >>> to say “we have this tool that can write all these amazing tests and code >>> but tracking what models are used and writing them down is too hard”. >>> Skills that can automate development and review can be easily extended to >>> track models being used. But again, it seems like there is enough concern >>> for specific attribution that a more general attribution like originally >>> proposed would be better. I don’t personally agree with the concerns about >>> specific attribution raised so far but see many possible compromises. I >>> would encourage us to debate specific language on a proposal (which again >>> I’m happy to create). >>> >>> On testing, a few of you are implying / saying that given LLMs we can raise >>> the testing bar higher than it is today. I actually agree with that. But it >>> implies almost a need to use LLMs which would in turn make it harder for >>> contributors who aren’t at your level of adoption of AI. As someone who is >>> relatively far on the AI adoption journey I don’t personally see this as a >>> bad thing but it doesn’t feel like our community is ready for that and many >>> of the same folks proposing testing can be done better with LLMs are the >>> same group advocating to not make the bar higher for contributors or have >>> special classes of contributors. Again personally, I like raising the >>> testing bar higher but I would encourage us to drive consensus on that >>> separately so we can reach some form of consensus here. IMO consensus here >>> will require accepting not everyone is at the same stage of AI adoption and >>> viewpoints as others might be. If we are trying to get everyone to see AI >>> the same and then make a policy i think we will have a hard time. >>> >>> Jordan >>> >>> On Wed, Sep 30, 2026 at 10:43 Jon Haddad <[email protected] >>> <mailto:[email protected]>> wrote: >>>> I think the main reason is that it's far less effort to create a >>>> comprehensive system for testing with an LLM than it is by hand. For >>>> example, with cursor compaction, we have parameterized, differential, >>>> fuzzed tests. I was able to go through several iterations and different >>>> ideas quite cheaply, in order to arrive where it is now. Being able to >>>> experiment with multiple paths forward is a huge general advantage, and on >>>> the side of testing it makes it a no brainer. >>>> >>>> LLMs also don't complain when they need to make big revisions, or plumb >>>> things through that would otherwise be an annoyance. They don't mind the >>>> grunt work, and are damn good at it. Add a reviewer that runs PMD and >>>> Jacoco and tells you where the code is either too complex or untested AND >>>> can quickly refactor it, or put together a few hundred lines of test code >>>> is remarkable. >>>> >>>> Jon >>>> >>>> >>>> >>>> On Tue, Sep 29, 2026 at 1:56 PM Dinesh Joshi <[email protected] >>>> <mailto:[email protected]>> wrote: >>>>> Jordan, thanks for the guiding principles. I think this is a short-enough >>>>> prose that I expect contributors to realistic read, digest and apply. >>>>> >>>>> One clarification though - why is the bar for LLM generated code higher >>>>> than human generated code? Why isn't the bar the same for both? Why does >>>>> it have to be a function of the thing that generated the code? >>>>> >>>>> Realistically, most developers are using LLM assistance is writing tests. >>>>> Like Caleb said, the cost of writing tests has fallen to zero. If >>>>> anything, I would say that the bar for tests should be higher regardless >>>>> of who generated the code. >>>>> >>>>> An unintended side effect of having a higher test bar for LLM generated >>>>> code is that it may incentivize contributors to hide / underplay the use >>>>> of LLMs. >>>>> >>>>> >>>>> >>>>> On Tue, Sep 29, 2026 at 1:04 PM Jordan West <[email protected] >>>>> <mailto:[email protected]>> wrote: >>>>>> On Tue, Sep 29, 2026 at 12:54 Caleb Rackliffe <[email protected] >>>>>> <mailto:[email protected]>> wrote: >>>>>>> @Jordan I agree with essentially everything you’ve said here. The only >>>>>>> exception is that I wouldn’t want to lower the testing bar (assuming we >>>>>>> agree on what that means) for non-LLM-authored patches. The cost of >>>>>>> writing tests has esssntially fallen to zero. >>>>>> >>>>>> >>>>>> We’re absolutely aligned on that. If I implied otherwise that was an >>>>>> error on my part. My only intent is the bar should be even higher >>>>>> for LLM generated code than whatever the bar is for human generated code >>>>>> not that we lower the human generated bar we have or agree to in the >>>>>> future. Maybe we have to wordsmith that some to be more clear? >>>>>> >>>>>>> >>>>>>>> On Sep 29, 2026, at 2:30 PM, Jordan West <[email protected] >>>>>>>> <mailto:[email protected]>> wrote: >>>>>>>> >>>>>>>> >>>>>>>> While I too want to lower the bar for contributors to join us, I am >>>>>>>> happy to see something like the proposed policy and I don’t think >>>>>>>> reading a short document is too much to ask when contributing to our >>>>>>>> large and critical code base. We’ve had and have much larger barriers >>>>>>>> to entry than that. One reason I’m excited about AI use in the project >>>>>>>> is I think it can help us lower more of those. >>>>>>>> >>>>>>>> While I agree with the spirit of the proposed policy and some of >>>>>>>> what’s in it I think as written it will have us back here often to >>>>>>>> re-discuss this topic for a couple reasons. First, “strongly >>>>>>>> discouraged” will have different meanings to different community >>>>>>>> members and as the policy acknowledges it’s unenforceable so this will >>>>>>>> lead to debates based on the ambiguities in the text. Second, all of >>>>>>>> us are on various points on a spectrum on where we see AIs abilities >>>>>>>> today and where we see them going. The policy is written to be a sort >>>>>>>> of blend of our current opinions on where it is today so it seems >>>>>>>> likely we are back here today as our opinions shift (I know my beliefs >>>>>>>> change daily to weekly these days in both directions), capabilities >>>>>>>> change, and new consensus is needed. >>>>>>>> >>>>>>>> I propose we do have a document but one more I like of the “motivating >>>>>>>> and guiding principles section” and that we rely on our other existing >>>>>>>> policies and assumption of positive intent of contributors. I have >>>>>>>> proposed some below. I am sure several here will find these too >>>>>>>> lenient given what I presume to be where they fall on the AI adoption >>>>>>>> spectrum compared to me and I respect that. Some days I am likely >>>>>>>> right there with you, others I’m more bullish. You’ll find my proposal >>>>>>>> below does not encourage trying to delineate parts of the codebase >>>>>>>> that can and cannot be contributed to with an LLM but holds standards >>>>>>>> regardless. I encourage us to find ways to set the bar for quality >>>>>>>> when LLMs are used now and in the future vs trying to limit them based >>>>>>>> on today’s opinions and capabilities that are rapidly changing. >>>>>>>> >>>>>>>> Some proposed ideas for some guiding principles with that in mind. >>>>>>>> It’s likely not a complete list. >>>>>>>> >>>>>>>> * LLMs do not change accountability. You are ultimately accountable >>>>>>>> for code produced or reviewed in your name. Whether hand written or by >>>>>>>> an LLM. If you own an agentic process performing coding or review >>>>>>>> tasks you are still responsible and accountable for what it produces >>>>>>>> or what actions it takes. We as a community are accountable for >>>>>>>> understanding and being knowledgeable about the software we provide to >>>>>>>> others. LLMs do not change this. >>>>>>>> >>>>>>>> * Humans must be involved in the merging of code either by producing >>>>>>>> or reviewing code and is subject to the accountability requirement >>>>>>>> above. Nothing about using LLMs changes existing policies regarding >>>>>>>> committers required to merge code, vetoes, or other voting procedures. >>>>>>>> >>>>>>>> * All LLM use must be attributed to both you and the harness, >>>>>>>> provider, and model being used. The project makes no specific >>>>>>>> recommendations or requirements regarding the toolchain used as long >>>>>>>> as the user has legal access and provides attribution. >>>>>>>> >>>>>>>> * LLM generated code has a higher standard of automated testing than >>>>>>>> human code, for which we have already adopted an incredibly high >>>>>>>> standard. Proposed fixes or performance improvements must include >>>>>>>> runnable demonstrations. Use of LLMs does not absolve the accountable >>>>>>>> contributor of existing requirements such as providing a test plan in >>>>>>>> JIRA. LLM generated code must be automatically linted to meet the >>>>>>>> projects code standards. LLM generated code is not an excuse to ignore >>>>>>>> the projects existing policies on code style. >>>>>>>> >>>>>>>> * It is strongly preferred that documentation and comments are not LLM >>>>>>>> generated. We have found these to be of low quality. However, if done, >>>>>>>> the aforementioned accountability remains with the contributor. “The >>>>>>>> LLM wrote it” is not an acceptable dismissal of a review comment in >>>>>>>> any context. It is recommended that documentation and comments >>>>>>>> continue to be human written and optional LLM reviewed or edited with >>>>>>>> human supervision. >>>>>>>> >>>>>>>> On Tue, Sep 29, 2026 at 08:57 Caleb Rackliffe >>>>>>>> <[email protected] <mailto:[email protected]>> wrote: >>>>>>>>> @Josh I tried to define “directly generated” at the bottom, although >>>>>>>>> there isn’t a proper footnote/link. I don’t think it matters at this >>>>>>>>> point. There doesn’t appear to be any appetite for something in the >>>>>>>>> middle, i.e. what I was attempting to do there. >>>>>>>>> >>>>>>>>>> On Sep 29, 2026, at 10:00 AM, Josh McKenzie <[email protected] >>>>>>>>>> <mailto:[email protected]>> wrote: >>>>>>>>>> >>>>>>>>>> >>>>>>>>>> Some questions that are still unclear to me after reading through >>>>>>>>>> this thread and the PR - and I assume a new contributor would be >>>>>>>>>> confused as well: >>>>>>>>>> >>>>>>>>>> Re: what qualifies as "Directly Generated" by an LLM: >>>>>>>>>> If someone generates a full implementation and testing for something >>>>>>>>>> via an LLM then goes through line by line and cleans things up and >>>>>>>>>> makes changes, does that qualify as Directly Generated or not? >>>>>>>>>> If they have fine-tuned a local model to comments in their own >>>>>>>>>> verbal style, is that strongly discouraged because an LLM generated >>>>>>>>>> it? What if they review it line-by-line? What if they write things >>>>>>>>>> by hand then have an LLM rephrase things and leave the LLM's final >>>>>>>>>> directly generated text in place? >>>>>>>>>> What happens if 15% of the comments generated by the LLM are >>>>>>>>>> concise, clear, and only explain non-obvious "why's" of the code? >>>>>>>>>> Should a contributor go through and rephrase those lines in order to >>>>>>>>>> keep them from being directly generated? >>>>>>>>>> If I was a contributor looking for a project to start getting >>>>>>>>>> involved with baroque and bespoke rules would be incredibly >>>>>>>>>> off-putting to me. Honestly, the set of rules we have and social >>>>>>>>>> norms today are incredibly off-putting to many long-term >>>>>>>>>> contributors already who have a deep vested social and professional >>>>>>>>>> interest in the project succeeding. Who still would love to work >>>>>>>>>> technically on the project but are driven away by this culture. >>>>>>>>>> >>>>>>>>>> We're trying to hit a middle ground of not being too prescriptive >>>>>>>>>> but not leaving everything open to the interpretation of the reader >>>>>>>>>> which is just breeding more confusion. All in a space where the >>>>>>>>>> progress of the underlying tools is faster than anything I can >>>>>>>>>> recall in our field. Whatever policy we come up with now will >>>>>>>>>> probably be slightly outdated even by the time we ratify it unless >>>>>>>>>> it's incredibly high level and instead tries to codify our values >>>>>>>>>> and trust people to live up to them. >>>>>>>>>> >>>>>>>>>> Which I'd argue is exactly what Blake's simple proposal does. It's >>>>>>>>>> durable in the face of change and focuses on what's important to us >>>>>>>>>> and the community instead of engaging in pedantry and policing that >>>>>>>>>> just ends up confusing everyone further. >>>>>>>>>> >>>>>>>>>> On Tue, Sep 29, 2026, at 4:42 AM, Aleksey Yeshchenko via dev wrote: >>>>>>>>>>>> If anyone wants to follow along and/or add comments, we've created >>>>>>>>>>>> https://github.com/apache/cassandra/pull/5220 >>>>>>>>>>> >>>>>>>>>>> This is now quite qualitatively different from "Rust policy but >>>>>>>>>>> without the committer exception for critical sections", I'm afraid. >>>>>>>>>>> >>>>>>>>>>> Watered down beyond what we discussed here and offline, and not >>>>>>>>>>> what I and most folks who endorsed a Rust-like policy voted for. >>>>>>>>>>> >>>>>>>>>>> I'll make some edits to restore it to the shape we discussed last >>>>>>>>>>> night. >>>>>>>>>>> >>>>>>>>>>> -- >>>>>>>>>>> AY >>>>>>>>>>> >>>>>>>>>>>> On 29 Sep 2026, at 00:17, Caleb Rackliffe >>>>>>>>>>>> <[email protected] <mailto:[email protected]>> wrote: >>>>>>>>>>>> >>>>>>>>>>>> If anyone wants to follow along and/or add comments, we've created >>>>>>>>>>>> https://github.com/apache/cassandra/pull/5220 >>>>>>>>>>>> >>>>>>>>>>>> On Mon, Sep 28, 2026 at 2:45 PM Aleksey Yeshchenko via dev >>>>>>>>>>>> <[email protected] <mailto:[email protected]>> wrote: >>>>>>>>>>>> It is a little confusing to keep track of the proposed diffs to >>>>>>>>>>>> Rust's policy. I think David is preparing a version with all the >>>>>>>>>>>> changes applied to it, so there is no ambiguity. >>>>>>>>>>>> >>>>>>>>>>>>> How would we handle the “Non-critical” part of the experimental >>>>>>>>>>>>> section? The policy exempts rust-lang members from that… does >>>>>>>>>>>>> this mean we’d exempt committers but not non-committers. What’s >>>>>>>>>>>>> the Cassandra analog of the non-critical section? >>>>>>>>>>>> >>>>>>>>>>>> The modified proposal removes that paragraph (about exemptions) >>>>>>>>>>>> altogether, thus treating all C* developers equally and allowing >>>>>>>>>>>> code generation for non-critical parts only for everyone. >>>>>>>>>>>> >>>>>>>>>>>> If you are still confused (which would be understandable - it is >>>>>>>>>>>> confusing), perhaps wait for David's doc, to make sure we are all >>>>>>>>>>>> on the same page wrt what's being proposed first. >>>>>>>>>>>> >>>>>>>>>>>> -- >>>>>>>>>>>> AY >>>>>>>>>>>> >>>>>>>>>>>>> On 28 Sep 2026, at 20:16, Blake Eggleston <[email protected] >>>>>>>>>>>>> <mailto:[email protected]>> wrote: >>>>>>>>>>>>> >>>>>>>>>>>>> This is something I could support as well. >>>>>>>>>>>>> >>>>>>>>>>>>> 2 things: >>>>>>>>>>>>> >>>>>>>>>>>>> Our docs tend to be neglected. While ideally our docs would be >>>>>>>>>>>>> 100% human generated, they’re mostly just not generated at the >>>>>>>>>>>>> moment. While not ideal, I think relaxing the rust LLM policy as >>>>>>>>>>>>> it relates to docs would be a net positive for users, provided >>>>>>>>>>>>> they’re human reviewed and edited. >>>>>>>>>>>>> >>>>>>>>>>>>> How would we handle the “Non-critical” part of the experimental >>>>>>>>>>>>> section? The policy exempts rust-lang members from that… does >>>>>>>>>>>>> this mean we’d exempt committers but not non-committers. What’s >>>>>>>>>>>>> the Cassandra analog of the non-critical section? >>>>>>>>>>>>> >>>>>>>>>>>>> On Mon, Sep 28, 2026, at 12:10 PM, Francisco Guerrero wrote: >>>>>>>>>>>>>> I've gone over the the Rust policy. I am in support of the Rust >>>>>>>>>>>>>> version with the tweaks proposed by Caleb. >>>>>>>>>>>>>> >>>>>>>>>>>>>> Best, >>>>>>>>>>>>>> - Francisco >>>>>>>>>>>>>> >>>>>>>>>>>>>> On 2026/09/28 18:45:03 Aleksey Yeshchenko via dev wrote: >>>>>>>>>>>>>> > Some last minute amends to the suggested policy's TL;DR, with >>>>>>>>>>>>>> > Caleb's approval: >>>>>>>>>>>>>> > >>>>>>>>>>>>>> > - It’s fine to use LLMs to answer questions, analyze, distill, >>>>>>>>>>>>>> > refine, check, suggest, review. >>>>>>>>>>>>>> > - LLMs work best when used as a tool to write better, not >>>>>>>>>>>>>> > faster. >>>>>>>>>>>>>> > >>>>>>>>>>>>>> > "But not to create." bit is covered in detail by the full >>>>>>>>>>>>>> > policy and is impossible to summarise well in two words. >>>>>>>>>>>>>> > >>>>>>>>>>>>>> > With other changes as outlined by Caleb in the quoted email, I >>>>>>>>>>>>>> > would be happy to support this fine-tuned version of Rust's >>>>>>>>>>>>>> > policy. >>>>>>>>>>>>>> > >>>>>>>>>>>>>> > -- >>>>>>>>>>>>>> > AY >>>>>>>>>>>>>> > >>>>>>>>>>>>>> > > On 28 Sep 2026, at 19:27, Caleb Rackliffe >>>>>>>>>>>>>> > > <[email protected] <mailto:[email protected]>> >>>>>>>>>>>>>> > > wrote: >>>>>>>>>>>>>> > > >>>>>>>>>>>>>> > > To clarify, I would remove the "Experimental" tag and make >>>>>>>>>>>>>> > > that section apply to all contributors. (In other words, >>>>>>>>>>>>>> > > encourage attribution, quality, and human decision-making >>>>>>>>>>>>>> > > for all of us.) >>>>>>>>>>>>>> > > >>>>>>>>>>>>>> > > The spirit of this is really a one line change to the Rust >>>>>>>>>>>>>> > > policy: >>>>>>>>>>>>>> > > >>>>>>>>>>>>>> > > > It’s fine to use LLMs to answer questions, analyze, >>>>>>>>>>>>>> > > > distill, refine, check, suggest, review. But not to create. >>>>>>>>>>>>>> > > >>>>>>>>>>>>>> > > ...becomes... >>>>>>>>>>>>>> > > >>>>>>>>>>>>>> > > It’s fine to use LLMs to answer questions, analyze, distill, >>>>>>>>>>>>>> > > refine, check, suggest, review. But not to decide. >>>>>>>>>>>>>> > > >>>>>>>>>>>>>> > > >>>>>>>>>>>>>> > > >>>>>>>>>>>>>> > > On Mon, Sep 28, 2026 at 12:54 PM Caleb Rackliffe >>>>>>>>>>>>>> > > <[email protected] <mailto:[email protected]> >>>>>>>>>>>>>> > > <mailto:[email protected] >>>>>>>>>>>>>> > > <mailto:[email protected]>>> wrote: >>>>>>>>>>>>>> > >> I finally read the Rust and Lucene policy docs in more >>>>>>>>>>>>>> > >> detail... >>>>>>>>>>>>>> > >> >>>>>>>>>>>>>> > >> https://flagged.apple.com:443/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11 >>>>>>>>>>>>>> > >> >>>>>>>>>>>>>> > >> <https://flagged.apple.com/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11> >>>>>>>>>>>>>> > >> >>>>>>>>>>>>>> > >> <https://flagged.apple.com/proxy?t2=Dx1j3o8bN8&o=aHR0cHM6Ly9mb3JnZS5ydXN0LWxhbmcub3JnL3BvbGljaWVzL2xsbS11c2FnZS5odG1s&emid=10aadc4f-b52a-4662-9281-ce00061a230c&c=11> >>>>>>>>>>>>>> > >> https://github.com/apache/lucene/blob/main/AI_POLICY.md >>>>>>>>>>>>>> > >> >>>>>>>>>>>>>> > >> I think I agree with a lot of what's written in both, and >>>>>>>>>>>>>> > >> they overlap quite a lot, especially around communication >>>>>>>>>>>>>> > >> (docs, issue comments, etc.) that should be primarily >>>>>>>>>>>>>> > >> human-to-human. Everything useful in the Lucene policy is >>>>>>>>>>>>>> > >> already included in the Rust policy though. If we could >>>>>>>>>>>>>> > >> take the Rust policy, generalize away the Rust-specific >>>>>>>>>>>>>> > >> things, simplify it, and remove the "Experimental" tag (and >>>>>>>>>>>>>> > >> probably the "non-critical" qualifier) on the "LLM-created >>>>>>>>>>>>>> > >> code changes intended of review" section, I think that's >>>>>>>>>>>>>> > >> something a large majority of us would be able to live with. >>>>>>>>>>>>>> > >> >>>>>>>>>>>>>> > >> If we can get this right, it's simply clarifying the set of >>>>>>>>>>>>>> > >> things contributors (including existing committers) can do >>>>>>>>>>>>>> > >> to have the best chance at getting engagement from >>>>>>>>>>>>>> > >> reviewers. >>>>>>>>>>>>>> > >> >>>>>>>>>>>>>> > >> I don't know how much appetite there is out there for a >>>>>>>>>>>>>> > >> formal draft of this, and we already have 3-4 proposals, >>>>>>>>>>>>>> > >> but I could attempt it if that would be useful... >>>>>>>>>>>>>> > >> >>>>>>>>>>>>>> > >> >>>>>>>>>>>>>> > >> On Mon, Sep 28, 2026 at 10:50 AM Štefan Miklošovič >>>>>>>>>>>>>> > >> <[email protected] <mailto:[email protected]> >>>>>>>>>>>>>> > >> <mailto:[email protected] >>>>>>>>>>>>>> > >> <mailto:[email protected]>>> wrote: >>>>>>>>>>>>>> > >>> A clarification from my side, I asked "what is wrong with >>>>>>>>>>>>>> > >>> this" in my >>>>>>>>>>>>>> > >>> latest email: >>>>>>>>>>>>>> > >>> >>>>>>>>>>>>>> > >>> "For these reasons, it should be expected that the person >>>>>>>>>>>>>> > >>> producing >>>>>>>>>>>>>> > >>> the patch has already demonstrated their expertise and >>>>>>>>>>>>>> > >>> commitment by >>>>>>>>>>>>>> > >>> producing and maintaining similar patches without the use >>>>>>>>>>>>>> > >>> of AI" >>>>>>>>>>>>>> > >>> >>>>>>>>>>>>>> > >>> It is "almost fine", the part of "similar patches without >>>>>>>>>>>>>> > >>> the use of >>>>>>>>>>>>>> > >>> AI" should not be there. It should stop before that. >>>>>>>>>>>>>> > >>> >>>>>>>>>>>>>> > >>> Otherwise this is going to exclude people who have a >>>>>>>>>>>>>> > >>> decade of >>>>>>>>>>>>>> > >>> experience with Cassandra and contributed countless >>>>>>>>>>>>>> > >>> patches of various >>>>>>>>>>>>>> > >>> size and complexity while according to that exact wording, >>>>>>>>>>>>>> > >>> they would >>>>>>>>>>>>>> > >>> not be eligible to contribute an AI patch. That is silly. >>>>>>>>>>>>>> > >>> I think this >>>>>>>>>>>>>> > >>> is wrong. It does not matter how it was produced. What is >>>>>>>>>>>>>> > >>> important is >>>>>>>>>>>>>> > >>> established trust and if a patch is correct. What does >>>>>>>>>>>>>> > >>> even the size >>>>>>>>>>>>>> > >>> of a patch have in common with that? Expertise and >>>>>>>>>>>>>> > >>> commitment! Not >>>>>>>>>>>>>> > >>> "the series of patches this committer ever produced was >>>>>>>>>>>>>> > >>> not complex >>>>>>>>>>>>>> > >>> enough so we can't take that code in". >>>>>>>>>>>>>> > >>> >>>>>>>>>>>>>> > >>> Also, who is exactly going to measure that anyway? What >>>>>>>>>>>>>> > >>> are the >>>>>>>>>>>>>> > >>> _objective_ criteria who qualifies? Somebody might come >>>>>>>>>>>>>> > >>> and say "while >>>>>>>>>>>>>> > >>> based on my criteria, (because I do not like this person), >>>>>>>>>>>>>> > >>> I do not >>>>>>>>>>>>>> > >>> think that the patches of this person qualify, because >>>>>>>>>>>>>> > >>> ...". We need >>>>>>>>>>>>>> > >>> hard data on whether it can be merged or not, performance >>>>>>>>>>>>>> > >>> improvement, >>>>>>>>>>>>>> > >>> stability ... >>>>>>>>>>>>>> > >>> >>>>>>>>>>>>>> > >>> I think this particular wording would need to be refined >>>>>>>>>>>>>> > >>> further. >>>>>>>>>>>>>> > >>> >>>>>>>>>>>>>> > >>> On Mon, Sep 28, 2026 at 4:34 PM Štefan Miklošovič >>>>>>>>>>>>>> > >>> <[email protected] <mailto:[email protected]> >>>>>>>>>>>>>> > >>> <mailto:[email protected] >>>>>>>>>>>>>> > >>> <mailto:[email protected]>>> wrote: >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > Right ... for that reason I don't think we should >>>>>>>>>>>>>> > >>> > restrict anybody to >>>>>>>>>>>>>> > >>> > create a PR or anything like that, putting some >>>>>>>>>>>>>> > >>> > artificial constraints >>>>>>>>>>>>>> > >>> > people will eventually bypass anyway. We don't have that >>>>>>>>>>>>>> > >>> > under >>>>>>>>>>>>>> > >>> > control. What we have under control is the review part >>>>>>>>>>>>>> > >>> > of that. A >>>>>>>>>>>>>> > >>> > patch not merged will not be released. The review itself >>>>>>>>>>>>>> > >>> > is the >>>>>>>>>>>>>> > >>> > "gate". >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > If a PR, even done by AI, is up to standards, has >>>>>>>>>>>>>> > >>> > everything it should >>>>>>>>>>>>>> > >>> > have and it is technically correct, then I can not >>>>>>>>>>>>>> > >>> > reject to merge >>>>>>>>>>>>>> > >>> > that only on the basis it was AI-generated. A patch like >>>>>>>>>>>>>> > >>> > a patch. The >>>>>>>>>>>>>> > >>> > code speaks. The ultimate gate is if a patch is correct >>>>>>>>>>>>>> > >>> > or not, not >>>>>>>>>>>>>> > >>> > how it was produced. >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > Do I gravitate with my trust more towards established >>>>>>>>>>>>>> > >>> > members of the >>>>>>>>>>>>>> > >>> > community? Definitely. The trust is earned over the >>>>>>>>>>>>>> > >>> > years. Implicitly, >>>>>>>>>>>>>> > >>> > I am trusting a newcomer less. Sorry but not sorry. If >>>>>>>>>>>>>> > >>> > somebody calls >>>>>>>>>>>>>> > >>> > this "gating", I don't think they see the nuances >>>>>>>>>>>>>> > >>> > enough. Yeah, call >>>>>>>>>>>>>> > >>> > it a gate if you want ... >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > That is why I agree with Benedict here, he said: >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > "For these reasons, it should be expected that the >>>>>>>>>>>>>> > >>> > person producing >>>>>>>>>>>>>> > >>> > the patch has already demonstrated their expertise and >>>>>>>>>>>>>> > >>> > commitment by >>>>>>>>>>>>>> > >>> > producing and maintaining similar patches without the >>>>>>>>>>>>>> > >>> > use of AI". >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > What is wrong about this? >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > Look at this contributor (1). This is an excellent >>>>>>>>>>>>>> > >>> > example. 10 patches >>>>>>>>>>>>>> > >>> > in fast cadence three weeks ago. We never heard about >>>>>>>>>>>>>> > >>> > this person >>>>>>>>>>>>>> > >>> > before nor after the patches were created. What about >>>>>>>>>>>>>> > >>> > hitting a ML >>>>>>>>>>>>>> > >>> > saying "hey, guys, I have a set of patches which scratch >>>>>>>>>>>>>> > >>> > my itches, >>>>>>>>>>>>>> > >>> > can you take a look, please?". I don't know ... just be >>>>>>>>>>>>>> > >>> > a bit ... >>>>>>>>>>>>>> > >>> > human about all of this? The maintainers are people too. >>>>>>>>>>>>>> > >>> > I am not >>>>>>>>>>>>>> > >>> > obliged to take in and cooperate with whoever comes by, >>>>>>>>>>>>>> > >>> > dumps their >>>>>>>>>>>>>> > >>> > stuff and then they ... wait. Well, so wait. See where >>>>>>>>>>>>>> > >>> > you got three >>>>>>>>>>>>>> > >>> > weeks after? Nowhere. >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > Caleb put it nicely, we are "only" humans. >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > (1) >>>>>>>>>>>>>> > >>> > https://github.com/apache/cassandra/pulls?q=is%3Apr+state%3Aopen+author%3Acheeeee >>>>>>>>>>>>>> > >>> > >>>>>>>>>>>>>> > >>> > On Mon, Sep 28, 2026 at 2:51 PM Shailaja Koppu via dev >>>>>>>>>>>>>> > >>> > <[email protected] >>>>>>>>>>>>>> > >>> > <mailto:[email protected]> >>>>>>>>>>>>>> > >>> > <mailto:[email protected] >>>>>>>>>>>>>> > >>> > <mailto:[email protected]>>> wrote: >>>>>>>>>>>>>> > >>> > > >>>>>>>>>>>>>> > >>> > > Hi Stefan, >>>>>>>>>>>>>> > >>> > > >>>>>>>>>>>>>> > >>> > > From personal side, I completely agree with you. I am >>>>>>>>>>>>>> > >>> > > giving potential options only to address concerns like >>>>>>>>>>>>>> > >>> > > - new contributors overwhelming the community with AI >>>>>>>>>>>>>> > >>> > > generated PRs just to show as add-on in their profile >>>>>>>>>>>>>> > >>> > > and vanish after that, or purely AI opened PRs without >>>>>>>>>>>>>> > >>> > > developer review or understanding. But the later can >>>>>>>>>>>>>> > >>> > > happen with anyone including committers due to >>>>>>>>>>>>>> > >>> > > workload/deadlines or misled by AI etc. Also, someone >>>>>>>>>>>>>> > >>> > > can copy a AI generated patch line by line skipping >>>>>>>>>>>>>> > >>> > > comments, which looks like a handwritten code. >>>>>>>>>>>>>> > >>> > > >>>>>>>>>>>>>> > >>> > > >>>>>>>>>>>>>> > >>> > > Thanks, >>>>>>>>>>>>>> > >>> > > Shailaja >>>>>>>>>>>>>> > >>> > > >>>>>>>>>>>>>> > >>> > > >>>>>>>>>>>>>> > >>> > > >>>>>>>>>>>>>> > >>> > > > On Sep 28, 2026, at 12:23 PM, Štefan Miklošovič >>>>>>>>>>>>>> > >>> > > > <[email protected] >>>>>>>>>>>>>> > >>> > > > <mailto:[email protected]> >>>>>>>>>>>>>> > >>> > > > <mailto:[email protected] >>>>>>>>>>>>>> > >>> > > > <mailto:[email protected]>>> wrote: >>>>>>>>>>>>>> > >>> > > > >>>>>>>>>>>>>> > >>> > > >> - Only Cassandra committers may submit AI-assisted >>>>>>>>>>>>>> > >>> > > >> PRs. This would mean new contributors first write >>>>>>>>>>>>>> > >>> > > >> and understand code without AI before becoming >>>>>>>>>>>>>> > >>> > > >> committers; or >>>>>>>>>>>>>> > >>> > > >> - Contributors may submit AI-assisted changes in a >>>>>>>>>>>>>> > >>> > > >> component/subcomponent only after they have >>>>>>>>>>>>>> > >>> > > >> submitted at least one non-AI PR in that >>>>>>>>>>>>>> > >>> > > >> component/subcomponent. >>>>>>>>>>>>>> > >>> > > > >>>>>>>>>>>>>> > >>> > > > I am not sure if I am missing something but can you >>>>>>>>>>>>>> > >>> > > > all explain in >>>>>>>>>>>>>> > >>> > > > simple terms how is this actually enforceable in >>>>>>>>>>>>>> > >>> > > > practice? >>>>>>>>>>>>>> > >>> > > > >>>>>>>>>>>>>> > >>> > > > "Only Cassandra committers may submit AI-assisted >>>>>>>>>>>>>> > >>> > > > PRs" - there is no >>>>>>>>>>>>>> > >>> > > > restriction who can create a PR and how. It is not >>>>>>>>>>>>>> > >>> > > > like we see that a >>>>>>>>>>>>>> > >>> > > > PR is created with heavy AI usage, then we check if >>>>>>>>>>>>>> > >>> > > > a contributor is a >>>>>>>>>>>>>> > >>> > > > committer and when they are not we comment on that >>>>>>>>>>>>>> > >>> > > > PR saying - "hold >>>>>>>>>>>>>> > >>> > > > your horses mate, we checked the list and you are >>>>>>>>>>>>>> > >>> > > > not a committer, >>>>>>>>>>>>>> > >>> > > > sorry, we have to close this". >>>>>>>>>>>>>> > >>> > > > >>>>>>>>>>>>>> > >>> > > > If a PR is crafted "carefuly" then it might look >>>>>>>>>>>>>> > >>> > > > like a completely >>>>>>>>>>>>>> > >>> > > > legitimate piece of work while it is still 100% >>>>>>>>>>>>>> > >>> > > > prompted and the >>>>>>>>>>>>>> > >>> > > > author does not have a clue what they did. I mean >>>>>>>>>>>>>> > >>> > > > ... how do you make >>>>>>>>>>>>>> > >>> > > > the difference between what is "real" and what is >>>>>>>>>>>>>> > >>> > > > AI-driven 100%? I >>>>>>>>>>>>>> > >>> > > > think that even if we "guessed" which one is which, >>>>>>>>>>>>>> > >>> > > > the possibility to >>>>>>>>>>>>>> > >>> > > > see this is being progressively erased as this tech >>>>>>>>>>>>>> > >>> > > > is evolving and we >>>>>>>>>>>>>> > >>> > > > will eventually not have a clue. >>>>>>>>>>>>>> > >>> > > >>>>>>>>>>>>>> > >>>>>>>>>>>>>> > >>>>>>>>>>>>>>
