On Wed, 2026-07-22 at 17:24 -0700, Ackerley Tng wrote:
> David Woodhouse <[email protected]> writes:
> 
> > For dedicated hosting environments, there are a few reasons to prefer
> > using external PFNMAP memory for guests, rather than kernel-managed
> > memory:
> > 
> >  • No struct page overheads
> > 
> >  • Easily managed in 1GiB pages (or larger chunks); no fragmentation/THP
> > 
> >  • Faster kexec/KHO live update (every millisecond spent faffing around
> >    with memory management is an extra millisecond of steal time taken
> >    from guests)
> > 
> 
> Also, no need to split/merge pages when converting from private to
> shared if the host mappings can't depend on the refcounts in struct
> pages :).
> 
> > Today, there is only one implementation of guest_memfd: the internal
> > shmem-based page cache in virt/kvm/guest_memfd.c. This series allows
> > for other code to provide "a guest_memfd". This is an early path-
> > finding proof of concept tested with AMD SEV-SNP as well as non-CoCo
> > guests on x86 and arm64. I'm planning to build a memory-based file
> > system which gives files for both guest_memfd as well as memory-based
> > storage to allow things like userspace VM and orchestrator state
> > to be passed over KHO/kexec.
> > 
> 
> Thanks for raising this, good to know that there is more interest!
> 
> Vishal and I have been working on guest_memfd with HugeTLB for a while
> now. I'll be restarting upstream work on HugeTLB after the conversions
> series merges.
> 
> Vishal, Frank and I brought up having guest_memfd with HugeTLB with
> VM_PFNMAP mappings at a guest_memfd biweekly meeting a while ago. The
> consensus was that guest_memfd HugeTLB must support GUP and so VM_PFNMAP
> can't be the only way that guest_memfd supports HugeTLB.
> 
> Your proposal is true non-page-struct memory, which has been on the
> discussion guest_memfd biweekly discussion backlog for a while
> now. Google is definitely also interested in this.
> 
> Frank also described some use cases [1] other than HugeTLB and the
> regular PAGE_SIZE pages from the buddy that is currently supported.
> 
> I submitted a proposal for LPC 2026 to discuss the interface. What I had
> in mind was to have the KVM_GUEST_MEMFD_CREATE ioctl take a "template"
> fd. guest_memfd will then read off what it needs from the "template" fd
> (ops, like you have here), and the "template" fd could be VFIO [2] or
> HugeTLB.
> 
> You have a provider_fd = open("/dev/gmem_provider"), how about passing
> that provider_fd to guest_memfd, where /dev/gmem_provider could
> represent the page-less memory, like:
> 
>   ioctl(kvm_fd, KVM_GUEST_MEMFD_CREATE, provider_fd);
> 
> It would be cool if the fallback could be set up flexibly, like:
> 
>   ioctl(kvm_fd, KVM_GUEST_MEMFD_CREATE,
>         [provider_fd, fallback_provider_fd, ...]);
> 
> I think guest_memfd should be the wrapper around other memory providers
> because then the CoCo shared/private logic and special page table
> handling for both host and stage 2 page tables is centralized in
> guest_memfd, and also, KVM doesn't have to understand other, different
> kinds of fds other than guest_memfd.

Yeah. My approach was to make those "other kinds of fds" essentially
the same — they are "a guest_memfd", and KVM didn't have to understand
anything else (after a little bit of rework to have the generic ops so
that other code can create its own guest_memfds). As opposed to the
approach in Xu Yilun's TDISP series, which did teach KVM about a new
kind of memslot.

Likewise, I used dma-buf so that IOMMUFD didn't have to understand
anything new either.

I like your approach, inverting the wrapping of guest_memfd/dma-buf.
The provider then gives a plain dma-buf (not associated with any KVM as
it is in my current code). Userspace can use that *directly* with the
IOMMU, and passes it to your KVM_GUEST_MEMFD_CREATE in order to get a
guest_memfd for use with a given KVM. I think a lot of the warts in my
series fall away if we do it that way round.

I'll have a go at reworking it that way round... and see what different
warts it ends up with. The interesting part will be getting actual PFNs
out of the dma-buf (which is already under active discussion).

I'm not sure I understand the 'fallback provider' part. If one provider
doesn't have a page for a given address, KVM would iterate over them
all until it finds one that does? Is that about falling back from one
THP pool to another? Is this something that the generic KVM code needs
to know about, or could it be handled *within* a provider, which can
chain to the next pool if it can't find a page? Wouldn't we want that
done transparently for both KVM and IOMMU consumers?



> 
> [1] 
> https://lore.kernel.org/all/CAPTztWajm_JLpp9BjRcX=h72r25ELrXeGkOXVachybBxLJGS=g...@mail.gmail.com/
> [2] 
> https://lore.kernel.org/all/CAEvNRgFpJWQ5M5sQhGpQUV3GbBq9N+MQhhaxdxa=d8ky94s...@mail.gmail.com/
> 
> > This series lifts guest_memfd's operations into a small ops struct and
> > converts KVM's own gmem to use it, so it becomes one implementation
> > among peers rather than the only game in town.  It also adapts the
> > SEV-SNP fast paths to page-less operation, plumbs iommufd to map a
> > guest_memfd-backing dma-buf as CPU RAM (not MMIO), and adds the
> > provider->KVM revocation primitive that makes overcommit-style reclaim
> > possible.
> > 
> > It's largely AI-built at the moment. My main design concerns are around
> > how we tell that a given file is "a guest_memfd", and how we plug into
> > IOMMUFD. Presenting a dma-buf was the less intrusive choice because that
> > path already has support for taking pages away, but I'm less convinced
> > that it's the cleaner choice in the long term.
> > 
> 
> IIUC from this series, you're creating a dmabuf fd out of a guest_memfd,
> what does "taking pages away" mean, in this case, from where/who?

Taken away from the guest. Ballooning pages out, either for traditional
memory overcommit, or for the models where a guest donates some of its
pages to the system to run a "child" which is properly isolated from
the guest.

The physical page is reclaimed and needs to be taken out of the KVM and
IOMMU stage2 mappings. If KVM wants to touch it again it must call the
provider to get a new PFN for that range. This is roughly equivalent to
the MMU notifier path on userspace VMAs in the normal memslot case.
IOMMUFD has no equivalent of that for userspace addresses —
IOMMU_IOAS_MAP pins the pages so they *can't* go away, which is why I
used dma-buf for the IOMMUFD interface.

I've actually been fleshing out the overcommit path a bit more,
experimenting with adding async-pf support to guest_memfd in 
https://git.infradead.org/?p=users/dwmw2/linux.git;a=commitdiff;h=894dd0fcb34
(the sample provider just pretends its pages are slow to populate, so
the guest gets to run something else while a "reclaimed" page comes
back).

Attachment: smime.p7s
Description: S/MIME cryptographic signature

Reply via email to