On Tue, Aug 18, 2026 at 09:02:17AM -0700, Sean Christopherson wrote:
> On Tue, Aug 18, 2026, Pratyush Yadav wrote:
> > Hi Sean,
> > 
> > On Mon, Aug 17 2026, Sean Christopherson wrote:
> > 
> > > On Sat, Aug 15, 2026, Pratyush Yadav wrote:
> > >> On Wed, Aug 12 2026, Sean Christopherson wrote:
> > >> > On Wed, Aug 12, 2026, Pratyush Yadav wrote:
> > >> So with live update, we don't need to keep backwards compatibility in
> > >> the ABI forever. Of course, it is good to minimize changes, but we have
> > >> more freedom to change it.
> > >
> > > Ah.  So this is the heart of the disconnect.  I very strongly disagree 
> > > with the
> > > statement that live update doesn't need to support backwards 
> > > compatibility.  I
> > > can totally believe that the folks working on live update are ok breaking 
> > > backwards
> > > compatibility because their use cases are "fine" with such breakage.  And 
> > > I can
> > > also believe live update as an upstream kernel feature being developed by 
> > > those
> > > same folks is also ok with breaking backwards compatibility.
> > >
> > > But with my upstream KVM maintainer hat on, I am not ok with that.  I did 
> > > not agree
> > > to support a world where KVM is allowed to break backwards compatibility, 
> > > so long
> > > as it's done "carefully" or whatever.  If y'all want to deal with the 
> > > resulting
> > > complexity, that's fine by me, but you'll be doing it without KVM.
> > 
> > Let's step back a bit. I don't think backwards compatibility in the
> > _ABI_ is all that important in live update's context. What is important
> > is that our users are able to live update from kernel version X to Y.
> > The line format the kernel uses to describe its state is an
> > implementation detail. Our users never see it.
> 
> The rule is thou shalt not break userspace.  Whether or not the
> breakage is the result of an explicit ABI change is irrelevant.

I argued strongly for this simpler approach, I will write here the
reasoning I gave.

Firstly, "live update" is not uABI per-say. It is not "userspace
breaking", and it is not a "regression". It is an internal kernel
mechanism to allow the kernel to self-upgrade. In the industry it is
typical any of these hitless update/patch/etc schemes to only work
between a few tightly controlled version pairs.

Asking upstream to carry a full matrix of every single version pair
ever released is massive over-engineering and cost on upstream
maintainers. Refusing to do this is not a uABI breakage, it is not a
regression, it is simply a lack of a feature.

I think upstream should start by agreeing to only support forward
going upgrades within a single stable branch. If live update becomes a
success then maybe we could do from a stable branch to the immediate
next stable branch too. I don't know.

This alone is already much more powerful than live patching.

The balancing issue here is the impact on the kernel subsystems
adopting luo. Every time you change anything about the internal
function of the subystem you now have to go test a massive version
pair matrix to make sure everything works? No thanks.

Luo is very invasive and a huge PITA at the best of times. It doesn't
need to be even worse.

> It's probably fine for Google and other large companies that tightly
> control their kernels and use cases, and have the resources to
> juggle the resulting complexity, e.g. have kernel engineers on staff
> to track feature and dependencies, coordinate and plan kernel
> upgrades, etc.

Every downstream that wants to support luo is going to necessarily
severely restrict the version pairs that can work. Like a RH type
distro may only support it for 9.1.x -> 9.1.x+1 - they can carry the
cost of figuring out how to manage their patching and compatability
matrix on their own with the tools upstream provides.

> engineers on staff to help them thread the needle you describe
> above.  And if supporting live update as a general feature for all
> users of the kernel isn't being factored into design considerations,
> then that needs to change, otherwise this is all dead in the water.

General uses can use upstream and do CVE upgrades within a single
stable branch. That's good enough, lets start there. We don't need to
boil the ocean.

> I also don't see the point.  Maintaining a rigid save/restore ABI is
> annoying, but it's not _hard_ (or at least, not _that_ hard),
> especially if there's a set

I was told KVM had the smallest luo footprint of everything, so
perhaps your perspective is different.

In other places luo becomes coupled to the internal datastructures of
the kernel, we literally have to preserve lots of kernel memory
utterly unchanged without any kind of serialization at all. This
inevitably creates restrictions on the evolution of the kernel in
general if we can no longer change in-kernel architecture because it
would make some data structure incompatible with a kernel 10 years
old.

There is also a need for *downgrade*. Meaning the N+1 kernel has to
strictly emulate the limitations and capabilities of the N kernel so
it can be rolled back. Can you imagine the nightmare of trying to
codify the feature progression for every single kernel release forever
in some impossible scheme to enable this? Again, no thanks.

So, what upstream can reasonably do, that still provides real value
without severely burdening every subsystem, is upgrades within a
stable kernel branch only.

If some distro or CSP wants to do more, they can deal with it. They
have a far, far simpler problem because they can exactly lock down
their version pairs to a very small universe. They don't need a single
kernel that can emulate 10 different releases of behaviors
concurrently.

Jason

Reply via email to