From: Naman Jain <[email protected]> Sent: Tuesday, September 22, 
2026 2:09 AM
> 
> On 9/22/2026 2:39 AM, Michael Kelley wrote:
> > From: Naman Jain <[email protected]> Sent: Sunday, September 6, 
> > 2026 10:48 PM
> >>
> >> On Hyper-V guests each virtual PCI bus is enumerated by its own
> >> hv_pci_probe() call. The probe performs several synchronous host
> >> request/response exchanges while negotiating the protocol, querying bus
> >> relations, entering D0, and reporting allocated resources. These waits
> >> are latency-bound rather than CPU-bound.
> >>
> >> hv_pci registers as an ordinary VMBus driver, so driver_register() walks
> >> matching vPCI buses and probes them sequentially while the driver's
> >> initcall runs. On guests that expose several devices, each through its
> >> own vPCI bus, this serialization adds the host round-trip latencies to
> >> device initialization.
> >>
> >> Each bus is described by its own struct hv_pcibus_device, so
> >> independent buses can be probed concurrently. Request asynchronous
> >> probing via PROBE_PREFER_ASYNCHRONOUS, causing the driver core to
> >> schedule matching buses for asynchronous probe work.
> >>
> >> On an Azure Standard_L32s_v3 guest with five vPCI targets (four NVMe
> >> controllers and one Mellanox VF), Linux 7.2.3 was tested with one warm-up
> >> and three measured boots per variant. The median interval from the first
> >> hv_pci_probe() entry to the last return decreased from 2847.968 ms to
> >> 2786.709 ms, a 61.259 ms (2.15%) improvement.
> >
> > The elapsed time improvement is rather disappointing given the
> > complexity of the probing sequence and the number of interactions
> > with the Hyper-V host. Do you have any insight into why there isn't a
> > larger reduction? Is something mostly serializing the work even though
> > PROBE_PREFER_ASYNCHRONOUS is specified?
> >
> > Michael
> >
> 
> I can see these reasons for not seeing great improvements:
> 1. Timing of device offers from the host is beyond the control of guest
> and the Hyper-V host may also be serializing the requests from the host.

Ah, right. This is probably the key factor.

> 2. Shared locks that needs to be handled separately:
>     * hyperv_mmio_lock during VMBus MMIO allocation.

This probably has minimal impact. While there's a decent amount
of code protected by the lock, I don't think there's any interaction
with the host (even via traps), so it should run quickly.

>     * pci_rescan_remove_lock during PCI resource assignment and device
> addition.

OK. I don’t know about this one.

> 
> 
> I digged more into it, and it is indeed because of late offers from
> Hyper-V. I was considering the start of first probe to the last return,
> for time calculations.
> 
> For the 4 PCI devices on my setup whose offers were delivered together,
> the performance improvement was about 24%. However with the last offer
> coming late for MLX PCI device, overall improvement in time was lesser
> in terms of percentage.

For traditional configs where the VF NIC is paired with a netvsc
instance, the host doesn't offer the VF NIC to the guest until the
corresponding netvsc instance has been probed and the netvsc
driver has told the host it will accept a VF NIC. In these cases, the
Mellanox or MANA VF NIC device is always "late" and probably
shouldn't be included when determining the speed-up of doing
hv_pci probing asynchronously.

> 
> Dexuan had removed pci_rescan_remove_lock in his previous upstream
> attempt, but I ommitted it intentionally this time because from AI
> review, I saw a potential race condition that we would introduce if we
> remove it. Secondly, I did not observe any benefits of removing this
> lock. But I am going to revisit it again.

Thanks for the discussion. Using PROBE_PREFER_ASYNCHRONOUS
is still a good thing to use, but the benefit will accrue the most when
there are a significant number of NVMe devices and when Hyper-V
is quick about offering them. As I described above, it will be harder
to get parallelism with the VF NIC.

Michael

Reply via email to