There is no point to LLM generated documentation or comments: next year, they
will generate better ones, and can always do it on demand. If the docs are to
be LLM generated, a reader is better off generating them as needed so they get
them from the latest models.
On 2026/09/30 13:12:20 Josh McKenzie wrote:
> > * It is strongly preferred that documentation and comments are not LLM
> > generated. We have found these to be of low quality. However, if done, the
> > aforementioned accountability remains with the contributor. “The LLM wrote
> > it” is not an acceptable dismissal of a review comment in any context. It
> > is recommended that documentation and comments continue to be human written
> > and optional LLM reviewed or edited with human supervision.
> I'd prefer we move from discouraging the use of certain tooling in
> documentation and comments to being more prescriptive about what quality
> documentation and comments looks like. Since this tooling is evolving very
> rapidly (and different peoples' local setups differs greatly in capability
> and quality of output even today), this feels like a rule that's going to
> rapidly age out of relevance.
>
> ---
>
> Taking a crack at an alternative:
> • Code should be clearly and concisely documented. Classes should contain at
> least minimal javadoc headers explaining why they exist and their role in the
> architecture, and any non-trivial methods should likewise document their
> "why" and any non-obvious contextual assumptions at the time of their
> authoring. Include {@link} pointers to other methods they have behavioral
> dependencies on. Comments in the code should not be redundant with what the
> code is doing; if the code is opaque enough that you have to write comments
> to explain what it's doing that's probably a sign the code should be
> refactored and simplified.
>
> ---
> The reason I'd advocate for this over the blanket discouragement of LLMs is
> that we have a lot of really poorly documented code and a lack of alignment
> on this front as a project is hurting us; these guidelines would help set
> expectations for what we collectively think good comments look like and leave
> it up to a contributor on how they want to get to that bar of quality.
>
> Also: thanks for the measured and well thought out proposal Jordan. I agree
> w/others that we really should just be more declarative about what the bar of
> quality is and apply it evenly for code regardless of where it comes from
> (human, LLM) and lift that bar up to where we largely all agree it should be.
> If we have such agreement. :)
>
> On Tue, Sep 29, 2026, at 11:54 PM, C. Scott Andreas wrote:
> > Dinesh,
> >
> > I follow the point that you are making, but your reply dismisses Jane’s
> > central concern without addressing it: the review burden.
> >
> > “Software testing, certification, and verification” are orthogonal concerns
> > to the burden on reviewers and the activity of review: ensuring harmony
> > with the overall codebase and architecture, replicating knowledge of the
> > implementation among maintainers, and refining software together.
> >
> > These are functions that aren’t and can’t be addressed by software. They
> > are the “community” part of code.
> >
> > – Scott
> >
> >> On Sep 29, 2026, at 8:39 PM, Dinesh Joshi <[email protected]> wrote:
> >>
> >> Jane,
> >>
> >> Everything that you just said is true about human generated code as well.
> >> Believe it or not every engineer that hand writes code also believes that
> >> their code is perfect and they understand it completely. Yet reviewers
> >> find issues with it (and even they miss things that later get caught in
> >> testing or production).
> >>
> >> The issue with AI is that we can produce a lot of code very quickly and
> >> create a lot of work for reviewers.
> >>
> >> The problem we should be addressing is that of software testing,
> >> verification and certification of the database.
> >>
> >> Thanks,
> >>
> >> Dinesh
> >>
> >> On Tue, Sep 29, 2026 at 4:58 PM Jane H <[email protected]> wrote:
> >>> I'd support detailed and explicit policies, like Benedict's, Rust's, or
> >>> Caleb's, at least they can actually help the problem of AI assisted code
> >>> increasing reviewers' burden, though they might introduce other problems.
> >>>
> >>> Blake, why I think your 3 point simple rules won't work as you wish:
> >>> people have very different understandings in what's considered
> >>> responsible enough or comprehensive enough understanding.
> >>>
> >>> Everyone who has brought AI-coded PRs for me to review, new contributors
> >>> or committers, who costs around 4x time from me to review their PRs
> >>> compared to human-written ones (already more time than implementing on my
> >>> own), believes that they are fully responsible for their code and fully
> >>> understand their code. Every one of them believes so.
> >>>
> >>> I tried reminding them that they have to be fully responsible for the
> >>> verification of their code. Didn't help in my own experience. (I don't
> >>> blame them. We are all figuring out how to use this new tool. And I
> >>> voluntarily pick up the tasks of reviewing those PRs. Not their fault.
> >>> But that's why we need explicit guidance/policies.)
> >>>
> >>> I've already met people who believe one of the following is considered
> >>> responsible enough:
> >>> - Read every line of code and think it makes sense.
> >>> - Pass other AI's code reviews
> >>> - Pass all existing tests and AI-coded new unit tests and integration
> >>> tests
> >>>
> >>> But I think none of them is enough. Therefore, when someone says they are
> >>> fully responsible for/fully understanding their code, it doesn't really
> >>> reduce my reviewer's load. Those detailed and explicit policies like
> >>> Benedict's or Rust's, are largely just defining what's considered
> >>> responsible enough. So they will help.