On Fri, 28 Aug 2026, Priyank Rathod wrote:
> Per PCIe Base Specification r6.0, sec 8.4.4 ("Lane Margining at
> Receiver"), PCIe devices operating at 16.0 GT/s (Gen 4) or higher data
> rates support the Lane Margining at Receiver Extended Capability
> (ID 0x27), and it is mandatory for receivers operating at 64.0 GT/s
> (Gen 6) or higher data rates. Lane Margining allows software to
> evaluate high-speed link margins by measuring timing and voltage steps
> for each individual physical lane and receiver.
>
> Add driver and debugfs support for PCIe Lane Margining at Receiver:
>
> - Add Lane Margining at Receiver Extended Capability register
> definitions (PCI_EXT_CAP_ID_LMR, PCI_LMR_PORT_CAP, PCI_LMR_PORT_STS,
> PCI_LMR_LANE_CTRL, PCI_LMR_LANE_STS) to <uapi/linux/pci_regs.h>.
> - Add Kconfig option CONFIG_PCIE_LMR (under drivers/pci/pcie/Kconfig)
> dependent on DEBUG_FS.
> - Implement drivers/pci/pcie/margin.c to probe the capability on Gen4+
> links and expose per-device debugfs entries under:
> /sys/kernel/debug/pci/pcie_lmr_<pci_dev_name>/
> providing control over margining enablement, receiver selection, and
> execution of timing/voltage margin step commands. Distinguish
> between missing mandatory LMR capability on Gen6+ vs optional on
> Gen4/Gen5.
> - Hook pci_lmr_init() into pci_init_capabilities() during device probe
> in drivers/pci/probe.c and pci_lmr_exit() into drivers/pci/remove.c.
> - Add kselftest script under tools/testing/selftests/pcie_lmt/pcie_lmt.sh
> to test debugfs capability reads, enablement, and stepping.
> - Add MAINTAINERS entry for PCIe Lane Margining at Receiver (LMR).
>
> Signed-off-by: Priyank Rathod <[email protected]>
> ---
> Changes in v7:
> - Implemented dedicated ASPM Inhibit API (pci_lmr_aspm_inhibit) strictly
> enforcing PCIe Base Specification Revision 7.0 sec 7.5.3.7 sequencing
> (Downstream Component first on disable, Upstream Component first on restore)
> with 2-3ms L0 settling.
> - Added pci_lmr_ensure_aspm_inhibited() before executing physical margin
> steps to protect against out-of-band ASPM re-enabling.
Hi,
Unfortunately, this is ASPM thing is still wrong solution and you even
admit it yourself in a pci_info message that ASPM got enabled
"unexpectedly".
Please handle ASPM disabling (and reenabling) through the aspm driver, do
not write to ASPMC in this driver at all!! (For better disable/enable ASPM
driver API, you may have to look in one of the pending ASPM series and
build on top of that.)
If/when you need "inhibit", that should be done by the ASPM driver. Note
though that the way aspm driver handles state disabling is not entirely
satisfactory, It should probably keep counter for each ASPM state and only
allow enabling ASPM state X when nothing has disabled it, because a driver
could also want to disable ASPM and possibly re-enabled it later (e.g.
over a duration of a reset or fw update). Saving LNKCTL here would mess
that up so it should be ASPM driver's sole responsability to prevent
enabled ASPM prematurely if anything has asked a state to be kept off.
> - Eliminated pci_bus_sem lock inversion by resolving link partners once
> before acquiring device locks and referencing them via pci_lmr_get_ports().
> - Switched pci_reset_lmr() to asynchronous pm_runtime_put() for remote
> partner to prevent stranding it in D0 without blocking local device reset.
> - Aligned all specification citations, table numbers (Table 4-77 / Table
> 4-73), and register bitfields to PCIe Base Spec r7.0 and r6.0.
> - Wrapped comments to stay within 100-column checkpatch limit and added
> @partner kerneldoc description.
Some code comments are much longer than the code lines, and they are hard
to read because of that. The comments would be better to still be limited
to 80-chars, even if for code the 80 chars limit can be exceeded where it
makes sense.
I see you also added HASW/HAWD handling which is good!
> - Link to v6:
> https://lore.kernel.org/r/[email protected]
>
> Changes in v6:
> - Added kernel documentation under Documentation/PCI/pcie-lmr.rst and
> indexed in Documentation/PCI/index.rst (Ilpo Järvinen).
> - Updated MAINTAINERS with Documentation/PCI/pcie-lmr.rst (Ilpo Järvinen).
> - Aligned capability bit naming and comments with PCIe Base Specification
> r6.0 sec 8.4.4 Table "Report Margining Capabilities Payload" (Ilpo Järvinen).
> - Clarified Sample Multiple Receivers concurrency verification and rules
> across physical lanes in kerneldoc and documentation (Ilpo Järvinen).
> - Refactored pci_lmr_run_cmd() to pass struct pci_margin_dev *mdev
> directly, eliminating redundant NULL checks and using mdev->num_lanes (Ilpo
> Järvinen).
> - Converted PCI config read/write return checking across all helpers to
> pcibios_err_to_errno() (Ilpo Järvinen).
> - Reversed return logic in pci_lmr_demargin_lane() to return early on error
> (Ilpo Järvinen).
> - Refactored margin_lane_step_write() to eliminate bool is_voltage
> parameter, using command type (LMR_TYPE_TIMING / LMR_TYPE_VOLTAGE) and
> switch/case with consolidated bounds checks (Ilpo Järvinen).
> - Renamed __pci_suspend_lmr_locked() to pci_lmr_disable_locked() to avoid
> PM terminology confusion and added lockdep_assert_held(&mdev->lock) (Ilpo
> Järvinen).
> - Replaced -EACCES with -EBUSY across debugfs show/write callbacks when
> margining is inactive (Ilpo Järvinen).
> - Clarified comment for active operating link speed check (Gen4+ capability
> vs dynamically operating speed) in margin_enable_write() (Ilpo Järvinen).
> - Added WARN_ON_ONCE(!dev) check in pci_lmr_init() (Ilpo Järvinen).
> - Demoted capability detection log message from pci_info to pci_dbg to
> prevent boot log noise (Ilpo Järvinen).
> - Fixed timing step mask extraction in pci_lmr_cache_rx_info() to use 6-bit
> LMR_TIMING_STEP_MASK (sashiko-bot).
> - Resumed runtime PM via pm_runtime_resume_and_get() before performing
> config space reads in margin_enable_write() (sashiko-bot).
> - Switched to pm_runtime_put_sync() during margining teardown (sashiko-bot).
> - Link to v5:
> https://lore.kernel.org/r/[email protected]
>
> PCI/pcie: Add PCIe Lane Margining at Receiver (LMR) support
>
> Per PCIe Base Specification r6.0, section 8.4.4 ("Lane Margining at
> Receiver"),
> PCIe devices operating at 16.0 GT/s (Gen 4) or higher data rates support the
> Lane Margining at Receiver Extended Capability (ID 0x27), and it is mandatory
> for receivers operating at 64.0 GT/s (Gen 6) or higher data rates.
>
> Lane Margining allows system software to evaluate high-speed link signal
> integrity and margins by measuring timing and voltage steps for each physical
> lane and receiver independently.
>
> This series introduces kernel driver support, debugfs controls, and a
> kselftest automation script for PCIe Lane Margining at Receiver (LMR/LMT).
>
> ==============================================================================
> 1. How to Enable & Configure
> ==============================================================================
> Enable the Kconfig option under PCI support:
> CONFIG_PCIE_LMR=y (or =m)
> (Depends on CONFIG_PCI and CONFIG_DEBUG_FS)
>
> Upon boot or device hotplug on Gen4+ links (>= 16.0 GT/s), the driver probes
> Extended Capability ID 0x27 and exposes per-device debugfs interfaces:
> /sys/kernel/debug/pci/pcie_lmr_<domain>:<bus>:<dev>.<func>/
>
> ==============================================================================
> 2. How to Use the Debugfs Interface (Manual Margining)
> ==============================================================================
> Inspect device-wide margining capabilities and port status:
> # Inspect root device LMR capabilities & status
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/capabilities
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/port_status
>
> Enable active Lane Margining on the device:
> # Enable Lane Margining state machine
> echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
>
> Inspect and step individual lanes (e.g. lane0):
> # Select target receiver (0 = local receiver, 1..6 = retimers/link partners)
> echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/receiver
>
> # Check available timing and voltage steps for this receiver
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/caps
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_timing_steps
> cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_voltage_steps
>
> # Step timing margin or voltage margin offset
> echo 2 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
> echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
>
> # Reset margin offset back to nominal (0)
> echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
> echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
>
> Disable Lane Margining when finished:
> echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
>
> ==============================================================================
> 3. How to Run Automated Kselftests Using the Test Script
> ==============================================================================
> An automated kselftest script is included to test capability reads, receiver
> selection, and margining commands across all enumerated LMR devices:
>
> # Run directly as root
> sudo ./tools/testing/selftests/pcie_lmt/pcie_lmt.sh
>
> Or run via the kselftest Makefile harness:
> make -C tools/testing/selftests TARGETS=pcie_lmt run_tests
>
> Sample script output on an LMR-capable device:
> pcie_lmt: testing PCIe LMR debugfs entries
> pcie_lmt: probing device pcie_lmr_0000:01:00.0
> pcie_lmr_0000:01:00.0: capabilities read OK
> pcie_lmr_0000:01:00.0: port_status read OK
> pcie_lmr_0000:01:00.0: margining enabled OK
> pcie_lmr_0000:01:00.0: testing lane0
> pcie_lmr_0000:01:00.0: testing lane1
> pcie_lmr_0000:01:00.0: margining disabled OK
> pcie_lmt [PASS]
>
> To: Bjorn Helgaas <[email protected]>
> To: Shuah Khan <[email protected]>
> Cc: [email protected]
> Cc: [email protected]
> Cc: [email protected]
> Cc: Ilpo Järvinen <[email protected]>
>
> Changes in v5:
> - Sorted #include directives alphabetically and added missing includes for
> bits.h, bitfield.h, cleanup.h, overflow.h, and slab.h (Ilpo Järvinen).
> - Converted bitmasks to GENMASK() and BIT() macros and used FIELD_PREP()
> and FIELD_GET() instead of manual bit shifts (Ilpo Järvinen).
> - Added pci_lmr_sts_payload() helper to cleanly extract the status payload
> byte before applying step and capability masks (Ilpo Järvinen).
> - Replaced manual mutex locking sequences with guard(mutex)(&mdev->lock)
> across show and write callbacks to simplify control flow (Ilpo Järvinen).
> - Documented mutex lock protection scope in kerneldoc for struct
> pci_margin_dev (Ilpo Järvinen).
> - Used standard PCI_POSSIBLE_ERROR(), str_yes_no(), and scnprintf() helpers
> throughout the driver (Ilpo Järvinen).
> - Clarified receiver range (0..6 per PCIe r6.0 sec 8.4.4; 7 reserved) in
> comments and validation checks (Ilpo Järvinen).
> - Deduplicated timing and voltage show/write handlers using
> margin_lane_steps_show() and margin_lane_step_write() (Ilpo Järvinen).
> - Placed speed check immediately following pcie_get_speed_cap() and handled
> PCI_SPEED_UNKNOWN (Ilpo Järvinen).
> - Converted lanes in struct pci_margin_dev to a flexible array member with
> __counted_by(num_lanes) allocated via struct_size() (Ilpo Järvinen).
>
> Changes in v4:
> - Added Sample Multiple Receivers (Bit 5) concurrency verification in
> margin_lane_timing_write() and margin_lane_voltage_write() per PCIe r6.0 sec
> 8.4.4, returning -EBUSY if another lane on the same receiver is already
> margined when simultaneous lane margining is not supported.
> - Added active operating link speed verification (PCI_EXP_LNKSTA_CLS >=
> 16.0 GT/s) in margin_enable_write() before enabling LMR, as LMR commands are
> physically undefined on links operating at Gen1/Gen2/Gen3 speeds.
> - Added fast-path hardware NAK detection in pci_lmr_run_cmd() to return
> -EOPNOTSUPP immediately if a receiver echoes MTYPE == NO_CMD (0x7) after
> command issuance rather than waiting 150ms for a timeout.
> - Added pci_reset_lmr() hooked into __pci_reset_function_locked() to
> synchronize software state and demargin on FLR or Secondary Bus Reset.
> - Comprehensive NULL pointer checks and array/lane/receiver bounds checks
> added across all internal helpers and debugfs write handlers.
> - Added MAINTAINERS entry for PCIe Lane Margining at Receiver (LMR).
>
> Changes in v2:
> - Fixed NO_CMD (0x7) clearing in pci_lmr_run_cmd() before issuing new
> commands per PCIe r6.0 sec 8.4.4.
> - Protected plane->rx updates with mdev->lock in
> margin_lane_receiver_write().
> - Corrected Margining Port Capabilities bit definition to
> PCI_LMR_PORT_CAP_USES_SW_READY (0x0001) in <uapi/linux/pci_regs.h>.
> - Updated kselftest script (pcie_lmt.sh) to locate LMR debugfs entries.
> - Validated integer bounds against LMR_MAX_TIMING_STEP /
> LMR_MAX_VOLTAGE_STEP before narrowing u8 cast.
> - Moved mdev->enabled checks inside mutex_lock(&mdev->lock) to eliminate
> TOCTOU races.
> - Checked return values of all pci_read_config_word() calls, propagating
> -EIO on failure.
> - Eliminated dead store of cap in margin_enable_write().
> - Explicitly checked speed == PCIE_SPEED_64_0GT in pci_lmr_init() to avoid
> misidentifying PCI_SPEED_UNKNOWN (0xFF) as Gen6.
> ---
> Documentation/PCI/index.rst | 1 +
> Documentation/PCI/pcie-lmr.rst | 174 +++
> MAINTAINERS | 8 +
> drivers/pci/pci-driver.c | 2 +
> drivers/pci/pci.c | 1 +
> drivers/pci/pci.h | 17 +
> drivers/pci/pcie/Kconfig | 12 +
> drivers/pci/pcie/Makefile | 1 +
> drivers/pci/pcie/margin.c | 1618
> ++++++++++++++++++++++++++
> drivers/pci/probe.c | 1 +
> drivers/pci/remove.c | 2 +-
> include/linux/pci.h | 6 +
> include/uapi/linux/pci_regs.h | 18 +
> tools/testing/selftests/Makefile | 1 +
> tools/testing/selftests/pcie_lmt/Makefile | 3 +
> tools/testing/selftests/pcie_lmt/pcie_lmt.sh | 105 ++
> 16 files changed, 1969 insertions(+), 1 deletion(-)
>
> diff --git a/Documentation/PCI/index.rst b/Documentation/PCI/index.rst
> index 5d720d2a415e..9170c98cbf3f 100644
> --- a/Documentation/PCI/index.rst
> +++ b/Documentation/PCI/index.rst
> @@ -20,3 +20,4 @@ PCI Bus Subsystem
> controller/index
> boot-interrupts
> tph
> + pcie-lmr
> diff --git a/Documentation/PCI/pcie-lmr.rst b/Documentation/PCI/pcie-lmr.rst
> new file mode 100644
> index 000000000000..ffcd4fcf72cc
> --- /dev/null
> +++ b/Documentation/PCI/pcie-lmr.rst
> @@ -0,0 +1,174 @@
> +.. SPDX-License-Identifier: GPL-2.0
> +
> +=======================================================
> +PCI Express Lane Margining at Receiver (LMR) Subsystem
> +=======================================================
> +
> +:Author: Priyank Rathod <[email protected]>
> +:Copyright: 2026 Google LLC
> +
> +Overview
> +========
> +
> +Lane Margining at Receiver (LMR), specified in the PCI Express Base
> +Specification (Revision 7.0 sec 7.7.11 & sec 8.4.4; r6.0 sec 7.7.10 & sec
> 8.4.4),
> +allows system software to evaluate high-speed link physical signal integrity
> and
> +eye margins. LMR measures available timing (jitter/phase) and voltage margin
> +offsets for each physical lane and receiver independently while the link is
> +operating in active L0 state.
> +
> +Lane Margining Extended Capability (ID 0x27) is optional for links operating
> at
> +16.0 GT/s (PCIe Gen 4) and 32.0 GT/s (Gen 5), and is mandatory for receivers
> +operating at 64.0 GT/s (Gen 6) and higher.
> +
> +Target Receivers
> +================
> +
> +Each physical lane can margin up to 7 distinct receivers per PCIe link:
> +
> +* **Receiver 0 (Local Receiver)**: The receiver in the immediate link
> partner.
> +* **Receivers 1 to 6 (Retimers)**: Retimer pseudo-ports along the physical
> link
> + (up to 3 retimers, each with upstream and downstream pseudo-ports).
> +* **Receiver 7**: Reserved per PCIe Base Specification.
> +
> +Kernel Configuration
> +====================
> +
> +Enable the kernel configuration option under PCI support:
> +
> +.. code-block:: none
> +
> + CONFIG_PCIE_LMR=y (or =m)
> +
> +Dependencies:
> +* ``CONFIG_PCI``
> +* ``CONFIG_DEBUG_FS``
> +
> +Debugfs Interface Guide
> +=======================
> +
> +When an LMR-capable device is enumerated on a Gen4+ link, the kernel exposes
> +per-device control and status files under debugfs:
> +
> +.. code-block:: none
> +
> + /sys/kernel/debug/pci/pcie_lmr_<domain>:<bus>:<dev>.<func>/
> +
> +Device-Level Attributes
> +-----------------------
> +
> +* ``capabilities`` (read-only):
> + Displays the 16-bit Margining Port Capabilities register and whether the
> + device uses the Software Ready handshake bit.
> +
> +* ``port_status`` (read-only):
> + Displays the Margining Port Status register, indicating Margining Ready and
> + SW Ready states.
> +
> +* ``enable`` (read-write):
> + Enables (``1``) or disables (``0``) Lane Margining on the device.
> + Enabling margining locks the link into D0, prevents runtime PM suspend,
> + disables ASPM L0s/L1, disables hardware autonomous link width/speed
> changes,
> + and verifies that the link is operating at >= 16.0 GT/s.
> + Disabling margining restores ASPM, hardware autonomous width/speed
> settings,
> + and runtime PM, and returns all lanes to nominal (normal) operating
> settings.
> +
> +Lane-Level Attributes
> +---------------------
> +
> +For each physical lane (``lane0``, ``lane1``, ...):
> +
> +* ``receiver`` (read-write):
> + Gets or sets the active target receiver number (``0`` for local receiver,
> + ``1..6`` for retimers). Switching receivers automatically clears previous
> + offsets back to normal settings per PCIe single-receiver margining
> requirements.
> +
> +* ``caps`` (read-only):
> + Reports the target receiver's margining capabilities (PCIe Base
> Specification
> + Revision 7.0 Table 4-77 and Table 8-13; r6.0 Table 4-73 & Table 8-11):
> + - Voltage Margining support (supported vs unsupported)
> + - Independent Left/Right Timing Margining support (independent vs
> symmetric)
> + - Independent Up/Down Voltage Margining support (independent vs symmetric)
> + - Error Sampler vs Main Sampler (independent error sampler vs intrusive
> main sampler)
> + - Sample Reporting Method (sampling rate vs sample count)
> +
> +* ``num_timing_steps`` (read-only):
> + Maximum timing margin steps supported by the receiver (0..63).
> +
> +* ``num_voltage_steps`` (read-only):
> + Maximum voltage margin steps supported by the receiver (0..127).
> +
> +* ``margin_timing`` (read-write):
> + Applies timing margin step offset (+/-). Writing ``0`` clears timing margin
> + back to nominal.
> +
> +* ``margin_voltage`` (read-write):
> + Applies voltage margin step offset (+/-). Writing ``0`` clears voltage
> margin
> + back to nominal.
> +
> +Manual Margining Example
> +========================
> +
> +1. Inspect device capabilities and status:
> +
> +.. code-block:: sh
> +
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/capabilities
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/port_status
> +
> +2. Enable Lane Margining mode:
> +
> +.. code-block:: sh
> +
> + echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
> +
> +3. Configure target receiver and inspect step limits on lane 0:
> +
> +.. code-block:: sh
> +
> + echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/receiver
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/caps
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_timing_steps
> + cat /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/num_voltage_steps
> +
> +4. Apply timing and voltage margin steps:
> +
> +.. code-block:: sh
> +
> + # Step timing margin +2 steps
> + echo 2 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
> +
> + # Step voltage margin +1 step
> + echo 1 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
> +
> +5. Reset margins back to nominal:
> +
> +.. code-block:: sh
> +
> + echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_timing
> + echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/lane0/margin_voltage
> +
> +6. Disable Lane Margining when complete:
> +
> +.. code-block:: sh
> +
> + echo 0 > /sys/kernel/debug/pci/pcie_lmr_0000:01:00.0/enable
> +
> +Automated Testing via Kselftest
> +===============================
> +
> +The kernel includes an automated kselftest script under
> +``tools/testing/selftests/pcie_lmt/pcie_lmt.sh`` to probe, validate, and
> exercise
> +debugfs controls across all enumerated LMR devices.
> +
> +Run directly as root:
> +
> +.. code-block:: sh
> +
> + sudo ./tools/testing/selftests/pcie_lmt/pcie_lmt.sh
> +
> +Or run via the kselftest test harness:
> +
> +.. code-block:: sh
> +
> + make -C tools/testing/selftests TARGETS=pcie_lmt run_tests
> diff --git a/MAINTAINERS b/MAINTAINERS
> index b7094a616afd..b5deaae11bfe 100644
> --- a/MAINTAINERS
> +++ b/MAINTAINERS
> @@ -21059,6 +21059,14 @@ F:
> Documentation/devicetree/bindings/pci/qcom,sa8255p-pcie-ep.yaml
> F: drivers/pci/controller/dwc/pcie-qcom-common.c
> F: drivers/pci/controller/dwc/pcie-qcom-ep.c
>
> +PCIE LANE MARGINING AT RECEIVER (LMR)
> +M: Priyank Rathod <[email protected]>
> +L: [email protected]
> +S: Maintained
> +F: Documentation/PCI/pcie-lmr.rst
> +F: drivers/pci/pcie/margin.c
> +F: tools/testing/selftests/pcie_lmt/
> +
> PCMCIA SUBSYSTEM
> M: Dominik Brodowski <[email protected]>
> S: Odd Fixes
> diff --git a/drivers/pci/pci-driver.c b/drivers/pci/pci-driver.c
> index f36778e62ac1..ded3925aab1d 100644
> --- a/drivers/pci/pci-driver.c
> +++ b/drivers/pci/pci-driver.c
> @@ -743,6 +743,8 @@ static int pci_pm_prepare(struct device *dev)
>
> dev_pm_set_strict_midlayer(dev, true);
>
> + pci_suspend_lmr(pci_dev);
> +
> if (pm && pm->prepare) {
> int error = pm->prepare(dev);
> if (error < 0)
> diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c
> index 77b17b13ee61..e6d8cd0094f1 100644
> --- a/drivers/pci/pci.c
> +++ b/drivers/pci/pci.c
> @@ -5058,6 +5058,7 @@ static void pci_dev_save_and_disable(struct pci_dev
> *dev)
> */
> pci_set_power_state(dev, PCI_D0);
>
> + pci_reset_lmr(dev);
> pci_save_state(dev);
> /*
> * Disable the device by clearing the Command register, except for
> diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h
> index 4469e1a77f3c..ea66fb65a848 100644
> --- a/drivers/pci/pci.h
> +++ b/drivers/pci/pci.h
> @@ -796,6 +796,11 @@ static inline bool pci_dev_test_and_set_removed(struct
> pci_dev *dev)
> return test_and_set_bit(PCI_DEV_REMOVED, &dev->priv_flags);
> }
>
> +static inline bool pci_dev_is_removed(struct pci_dev *dev)
> +{
> + return test_bit(PCI_DEV_REMOVED, &dev->priv_flags);
> +}
> +
> static inline void pci_dev_allow_binding(struct pci_dev *dev)
> {
> set_bit(PCI_DEV_ALLOW_BINDING, &dev->priv_flags);
> @@ -1023,6 +1028,18 @@ static inline void pci_no_tph(void) { }
> static inline void pci_tph_init(struct pci_dev *dev) { }
> #endif
>
> +#ifdef CONFIG_PCIE_LMR
> +void pci_lmr_init(struct pci_dev *dev);
> +void pci_lmr_exit(struct pci_dev *dev);
> +void pci_suspend_lmr(struct pci_dev *dev);
> +void pci_reset_lmr(struct pci_dev *dev);
> +#else
> +static inline void pci_lmr_init(struct pci_dev *dev) { }
> +static inline void pci_lmr_exit(struct pci_dev *dev) { }
> +static inline void pci_suspend_lmr(struct pci_dev *dev) { }
> +static inline void pci_reset_lmr(struct pci_dev *dev) { }
> +#endif
> +
> #ifdef CONFIG_PCIE_PTM
> void pci_ptm_init(struct pci_dev *dev);
> void pci_save_ptm_state(struct pci_dev *dev);
> diff --git a/drivers/pci/pcie/Kconfig b/drivers/pci/pcie/Kconfig
> index 207c2deae35f..3b021ca2fe84 100644
> --- a/drivers/pci/pcie/Kconfig
> +++ b/drivers/pci/pcie/Kconfig
> @@ -137,6 +137,18 @@ config PCIE_PTM
> This is only useful if you have devices that support PTM, but it
> is safe to enable even if you don't.
>
> +config PCIE_LMR
> + bool "PCI Express Lane Margining at Receiver Support"
> + depends on DEBUG_FS
> + help
> + This enables the PCI Express Lane Margining at Receiver support.
> + Lane Margining allows software to determine the voltage and
> + timing margin of each lane on a PCIe link (16.0 GT/s and above).
> + The margining data is exposed via debugfs.
> +
> + This is only useful if you have devices that support lane
> + margining, but it is safe to enable even if you don't.
> +
> config PCIE_EDR
> bool "PCI Express Error Disconnect Recover support"
> depends on PCIE_DPC && ACPI
> diff --git a/drivers/pci/pcie/Makefile b/drivers/pci/pcie/Makefile
> index b0b43a18c304..aac45ae0402e 100644
> --- a/drivers/pci/pcie/Makefile
> +++ b/drivers/pci/pcie/Makefile
> @@ -13,4 +13,5 @@ obj-$(CONFIG_PCIEAER_INJECT) += aer_inject.o
> obj-$(CONFIG_PCIE_PME) += pme.o
> obj-$(CONFIG_PCIE_DPC) += dpc.o
> obj-$(CONFIG_PCIE_PTM) += ptm.o
> +obj-$(CONFIG_PCIE_LMR) += margin.o
> obj-$(CONFIG_PCIE_EDR) += edr.o
> diff --git a/drivers/pci/pcie/margin.c b/drivers/pci/pcie/margin.c
> new file mode 100644
> index 000000000000..a428726c500d
> --- /dev/null
> +++ b/drivers/pci/pcie/margin.c
> @@ -0,0 +1,1618 @@
> +// SPDX-License-Identifier: GPL-2.0
> +/*
> + * PCI Express Lane Margining at Receiver
> + *
> + * Copyright (C) 2026 Google LLC
> + * Author: Priyank Rathod <[email protected]>
> + *
> + * Lane Margining at Receiver (PCIe Base Specification Revision 7.0, sec
> 7.7.11 &
> + * sec 8.4.4; r6.0 sec 7.7.10 & sec 8.4.4) allows system software to
> determine
> + * the voltage and timing margins of each physical lane on a PCIe link. The
> + * Extended Capability (ID 0x27) is available for receivers operating at
> 16.0 GT/s
> + * (Gen4) or higher data rates, and is mandatory for receivers operating at
> 64.0 GT/s
> + * (Gen6) or higher data rates.
> + *
> + * This driver implements:
> + * - Probing Extended Capability ID 0x27 and Margining Port Capabilities.
> + * - Managing ASPM L0s/L1 link states during active margining with
> restoration.
> + * - PCIe Base Specification NO_CMD (0x7) clearing handshake per receiver
> and lane.
> + * - Caching receiver capabilities & step counts to avoid side-effects
> + * when setting to normal settings.
> + * - Handling Symmetric vs Independent Left/Right & Up/Down margin steps.
> + * - Runtime PM protection (D0 enforcement) during active margining.
> + * - Exposing per-device debugfs interfaces under /sys/kernel/debug/pci/.
> + */
> +
> +#include <linux/bitfield.h>
> +#include <linux/bits.h>
> +#include <linux/cleanup.h>
> +#include <linux/debugfs.h>
> +#include <linux/delay.h>
> +#include <linux/err.h>
> +#include <linux/errno.h>
> +#include <linux/jiffies.h>
> +#include <linux/kstrtox.h>
> +#include <linux/minmax.h>
> +#include <linux/mutex.h>
> +#include <linux/overflow.h>
> +#include <linux/pci.h>
> +#include <linux/pm_runtime.h>
> +#include <linux/seq_file.h>
> +#include <linux/slab.h>
> +#include <linux/sprintf.h>
> +#include <linux/string_choices.h>
> +#include <linux/types.h>
> +
> +#include "../pci.h"
> +
> +/*
> + * Margining Type (MTYPE) field encodings (bits 5:3) in Margining Lane
> Control
> + * and Margining Lane Status registers per PCIe Base Specification Revision
> 7.0:
> + * - Section 7.7.11 "Lane Margining at the Receiver Extended Capability (ID
> 0x27)"
> + * (Margining Lane Control Register & Margining Lane Status Register)
> + * [r6.0 Section 7.7.10]
> + * - Section 4.2.18.2 "Margin Command and Response Flow"
> + * (Table 4-77 "Margin Commands and Corresponding Responses")
> + * [r6.0 Table 4-73]
> + *
> + * Encodings:
> + * 001b (0x1) - Report Margin Control Capabilities
> + * 010b (0x2) - Set Margining Parameters (Go to Normal Settings, Clear
> Error Log)
> + * 011b (0x3) - Step Margin Timing
> + * 100b (0x4) - Step Margin Voltage
> + * 111b (0x7) - No Command
> + * (000b, 101b-110b are Reserved)
> + */
> +#define LMR_TYPE_REPORT_CAPS 0x1 /* Report Margin Control Capabilities */
> +#define LMR_TYPE_SET_PARAMS 0x2 /* Set Margining Parameters */
> +#define LMR_TYPE_TIMING 0x3 /* Step Margin Timing */
> +#define LMR_TYPE_VOLTAGE 0x4 /* Step Margin Voltage */
> +#define LMR_TYPE_NO_CMD 0x7 /* No Command */
Align values for better readability in define blocks like this.
> +
> +/* Command Payloads per PCIe Base Specification Revision 7.0 Table 4-77
> (r6.0 Table 4-73) */
> +#define LMR_PAYLOAD_REPORT_CAPS 0x88 /* Report Margin Control Capabilities */
> +#define LMR_PAYLOAD_REPORT_VOLT_STEPS 0x89 /* Report Margining Voltage Steps
> */
> +#define LMR_PAYLOAD_REPORT_TIM_STEPS 0x8A /* Report Margining Timing Steps */
> +#define LMR_PAYLOAD_GO_TO_NORMAL 0x0F /* Go to Normal Settings */
> +#define LMR_PAYLOAD_CLEAR_ERROR_LOG 0x55 /* Clear Error Log */
> +#define LMR_PAYLOAD_NO_CMD 0x9C /* No Command */
Ditto.
> +
> +/* LMR command timing parameters */
> +#define LMR_CMD_TIMEOUT_MS 150
> +#define LMR_CMD_SLEEP_MIN_US 100
> +#define LMR_CMD_SLEEP_MAX_US 250
> +#define LMR_ENABLE_TIMEOUT_MS 150
> +#define LMR_ENABLE_SLEEP_MIN_US 1000
> +#define LMR_ENABLE_SLEEP_MAX_US 2000
> +
> +/*
> + * LMR parameter limits per PCIe Base Specification Revision 7.0:
> + * - Max lanes (32): sec 7.7.11 & Table 8-13 (MMaxLanes max 31)
> + * - Receiver numbers 0..6: Table 4-76 (assignment) & Table 4-77 (valid for
> commands)
> + * - Max timing step (63): sec 4.2.18.1.2, Table 4-77 (8Ah), & Table 8-13
> + * - Max voltage step (127): sec 4.2.18.1.2, Table 4-77 (89h), & Table 8-13
> + */
> +#define LMR_MAX_LANES 32
> +#define LMR_MAX_RX_NUM 6
> +#define LMR_MAX_TIMING_STEP 63
> +#define LMR_MAX_VOLTAGE_STEP 127
> +
> +/* LMR PCIe generation numbers and helper */
> +#define LMR_GEN6 6
> +#define LMR_GEN5 5
> +#define LMR_GEN4 4
> +
> +#define LMR_SPEED_TO_GEN(speed) \
> + ((speed) >= PCIE_SPEED_64_0GT ? LMR_GEN6 : \
> + (speed) >= PCIE_SPEED_32_0GT ? LMR_GEN5 : \
> + LMR_GEN4)
> +
> +/* LMR lane register stride */
> +#define LMR_LANE_REG_STRIDE 4
> +
> +/* LMR receivers and directions */
> +#define LMR_RX_LOCAL 0
> +#define LMR_STEP_DIR_INCREASE 1
> +#define LMR_STEP_DIR_DECREASE 0
> +
> +/*
> + * Margining Payload field masks for Step Margin Timing and Step Margin
> Voltage
> + * per PCIe Base Specification Revision 7.0 sec 4.2.18.1.2
> + * ("Margin Payload for Step Margin Commands"):
> + *
> + * Step Margin Timing Payload:
> + * Bit 7: Reserved (must be 0b)
> + * Bit 6: Direction (0 = Left/Decrease, 1 = Right/Increase)
> + * Bits 5:0: Margin Step (0..63)
> + *
> + * Step Margin Voltage Payload:
> + * Bit 7: Direction (0 = Down/Decrease, 1 = Up/Increase)
> + * Bits 6:0: Margin Step (0..127)
> + */
> +#define LMR_TIMING_STEP_MASK GENMASK(5, 0)
> +#define LMR_TIMING_DIR_MASK BIT(6)
> +#define LMR_VOLTAGE_STEP_MASK GENMASK(6, 0)
> +#define LMR_VOLTAGE_DIR_MASK BIT(7)
> +
> +/*
> + * Margin Payload step direction field encodings per PCIe Base Specification
> + * Revision 7.0 sec 4.2.18.1.2 ("Margin Payload for Step Margin Commands"):
> + *
> + * For timing:
> + * Bit 6: 0b = Right of normal setting (also 0b Reserved for symmetric)
> + * 1b = Left of normal setting (when MIndLeftRightTiming is Set)
> + * For voltage:
> + * Bit 7: 0b = Up from normal setting (also 0b Reserved for symmetric)
> + * 1b = Down from normal setting (when MIndUpDownVoltage is Set)
> + */
> +#define LMR_STEP_DIR_RIGHT_OR_UP 0
> +#define LMR_STEP_DIR_LEFT_OR_DOWN 1
> +
> +/*
> + * Report Margin Control Capabilities (Command 88h) response payload bit
> fields
> + * per PCIe Base Specification Revision 7.0 Table 4-77 & Table 8-13 (r6.0
> Table 4-73 & Table 8-11):
Having the reference numbers for the latest spec version is enough.
> + * Bit 0: MVoltageSupported (1 = Voltage margining supported; 0 = Not
> supported)
> + * Bit 1: MIndUpDownVoltage (1 = Independent Up/Down voltage supported;
> 0 = Symmetric)
> + * Bit 2: MIndLeftRightTiming (1 = Independent Left/Right timing
> supported; 0 = Symmetric)
> + * Bit 3: MSampleReportingMethod (1 = Sampling rate supported; 0 =
> Sample count supported)
> + * Bit 4: MIndErrorSampler (1 = Independent error sampler; 0 = Main data
> sampler)
Thanks, this is better now. I have to admit it's not entirely your fault
things are as confusing as they are (the spec could have been clearer when
it comes to defining these).
> + * Bits 7:5: Reserved
> + */
> +#define LMR_CAP_VOLTAGE_SUPPORTED BIT(0)
> +#define LMR_CAP_IND_UP_DOWN_VOLTAGE BIT(1)
> +#define LMR_CAP_IND_LEFT_RIGHT_TIMING BIT(2)
> +#define LMR_CAP_SAMPLE_REPORT_METHOD BIT(3)
> +#define LMR_CAP_IND_ERROR_SAMPLER BIT(4)
Align all BIT()s.
> +
> +/*
> + * Step Margin Execution Status (Bits 7:6 of response payload per PCIe Base
> + * Specification Revision 7.0 sec 4.2.18.1.1 "Step Margin Execution Status"):
> + * 00b: Too many errors - Receiver autonomously went back to default settings
> + * 01b: Set up for margin in progress
> + * 10b: Margining in progress
> + * 11b: NAK - Unsupported Lane Margining command was issued
> + */
> +#define LMR_STS_EXEC_MASK GENMASK(7, 6)
> +#define LMR_STS_EXEC_TOO_MANY_ERR 0x0
> +#define LMR_STS_EXEC_SETUP_IN_PROGRESS 0x1
> +#define LMR_STS_EXEC_IN_PROGRESS 0x2
> +#define LMR_STS_EXEC_NAK 0x3
> +#define LMR_STS_ERR_CNT_MASK GENMASK(5, 0)
> +
> +/**
> + * struct pci_margin_rx_info - Cached Lane Margining receiver capabilities
> + * @caps_cached: True if receiver capabilities and step limits are cached
> + * @caps: Margining capabilities byte reported by receiver
> + * @num_timing_steps: Maximum timing margin steps supported by receiver
> + * @num_voltage_steps: Maximum voltage margin steps supported by receiver
> + */
> +struct pci_margin_rx_info {
> + bool caps_cached;
> + u8 caps;
> + u8 num_timing_steps;
> + u8 num_voltage_steps;
> +};
> +
> +/**
> + * struct pci_margin_lane - Per-lane margining state
> + * @mdev: Parent LMR margin device
> + * @lane: Physical lane index (0..num_lanes - 1)
> + * @rx: Selected target receiver number (0 = local, 1..6 = retimers)
> + * @timing_val: Current applied timing margin step offset (+/-)
> + * @voltage_val: Current applied voltage margin step offset (+/-)
> + * @rx_info: Cached receiver capabilities per receiver number
> + */
> +struct pci_margin_lane {
> + struct pci_margin_dev *mdev;
> + int lane;
> + u8 rx;
> + int timing_val;
> + int voltage_val;
> + struct pci_margin_rx_info rx_info[LMR_MAX_RX_NUM + 1];
> +};
> +
> +/**
> + * struct pci_margin_dev - PCIe Lane Margining device instance
> + * @dev: Underlying PCI device
> + * @partner: Connected link partner device across the PCIe link
> + * @cap: Extended capability offset (PCI_EXT_CAP_ID_LMR)
> + * @debugfs: Root debugfs dentry for this device
> + * @lock: Mutex protecting LMR hardware access, active margining enablement,
> + * target receiver selection, lane margining steps, and ASPM state
> + * @enabled: True if Lane Margining is currently enabled
> + * @aspm_saved: True if original ASPM configuration has been saved
> + * @saved_dsp_aspm: Saved ASPM control register bits for Downstream Port
> + * @saved_usp_aspm: Saved ASPM control register bits for Upstream Port
> + * @autonomous_saved: True if original autonomous width/speed configuration
> has been saved
> + * @saved_dsp_lnkctl: Saved Link Control register bits for Downstream Port
> + * @saved_dsp_lnkctl2: Saved Link Control 2 register bits for Downstream Port
> + * @saved_usp_lnkctl: Saved Link Control register bits for Upstream Port
> + * @saved_usp_lnkctl2: Saved Link Control 2 register bits for Upstream Port
> + * @num_lanes: Number of lanes on the link
> + * @lanes: Flexible array of per-lane state structures
> + */
> +struct pci_margin_dev {
> + struct pci_dev *dev;
> + struct pci_dev *partner;
> + u16 cap;
> + struct dentry *debugfs;
> + struct mutex lock;
> + bool enabled;
> + bool aspm_saved;
> + u16 saved_dsp_aspm;
> + u16 saved_usp_aspm;
> + bool autonomous_saved;
> + u16 saved_dsp_lnkctl;
> + u16 saved_dsp_lnkctl2;
> + u16 saved_usp_lnkctl;
> + u16 saved_usp_lnkctl2;
> + int num_lanes;
> + struct pci_margin_lane lanes[] __counted_by(num_lanes);
> +};
> +
> +#if IS_ENABLED(CONFIG_DEBUG_FS)
> +static DEFINE_MUTEX(pci_debugfs_root_lock);
> +static struct dentry *pci_debugfs_root_dir;
> +
> +static struct dentry *get_pci_debugfs_root(void)
> +{
> + mutex_lock(&pci_debugfs_root_lock);
> + if (!pci_debugfs_root_dir)
> + pci_debugfs_root_dir = debugfs_lookup("pci", NULL);
> + if (!pci_debugfs_root_dir)
> + pci_debugfs_root_dir = debugfs_create_dir("pci", NULL);
> + mutex_unlock(&pci_debugfs_root_lock);
> + return pci_debugfs_root_dir;
> +}
> +#endif
> +
> +/*
> + * pci_lmr_get_link_partners() - Identify Downstream and Upstream Port link
> partners.
> + *
> + * For Root Ports and Switch Downstream Ports, @dev is the Downstream Port,
> and the
> + * connected device on the secondary bus is the Upstream Port.
> + * For Endpoints and Switch Upstream Ports, @dev is the Upstream Port, and
> the
> + * upstream bridge is the Downstream Port.
> + */
> +static void pci_lmr_get_link_partners(struct pci_dev *dev,
> + struct pci_dev **downstream_port,
> + struct pci_dev **upstream_port)
> +{
> + if (pci_pcie_type(dev) == PCI_EXP_TYPE_ROOT_PORT ||
> + pci_pcie_type(dev) == PCI_EXP_TYPE_DOWNSTREAM) {
> + *downstream_port = dev;
> + down_read(&pci_bus_sem);
> + *upstream_port = dev->subordinate ?
> +
> pci_dev_get(list_first_entry_or_null(&dev->subordinate->devices,
> + struct pci_dev,
> bus_list)) : NULL;
> + up_read(&pci_bus_sem);
> + } else {
> + *downstream_port = pci_upstream_bridge(dev);
> + *upstream_port = dev;
> + }
> +}
> +
> +static void pci_lmr_put_link_partners(struct pci_dev *dev,
> + struct pci_dev *downstream_port,
> + struct pci_dev *upstream_port)
> +{
> + if (pci_pcie_type(dev) == PCI_EXP_TYPE_ROOT_PORT ||
> + pci_pcie_type(dev) == PCI_EXP_TYPE_DOWNSTREAM) {
> + if (upstream_port)
> + pci_dev_put(upstream_port);
> + }
> +}
> +
> +/*
> + * pci_lmr_get_ports() - Identify Downstream and Upstream Port link partners
> + * using the already tracked mdev->dev and mdev->partner devices.
> + *
> + * For Root Ports and Switch Downstream Ports, @dev is the Downstream Port
> and
> + * @partner is the Upstream Port. For Endpoints and Switch Upstream Ports,
> + * @partner is the Downstream Port and @dev is the Upstream Port.
> + *
> + * Context: Called with mdev->lock held and partner already established.
> + * Does NOT acquire pci_bus_sem, preventing lock inversion deadlocks with
> + * device_lock.
> + */
> +static void pci_lmr_get_ports(struct pci_margin_dev *mdev,
> + struct pci_dev **downstream_port,
> + struct pci_dev **upstream_port)
> +{
> + struct pci_dev *dev = mdev->dev;
> + struct pci_dev *partner = mdev->partner;
> +
> + if (pci_pcie_type(dev) == PCI_EXP_TYPE_ROOT_PORT ||
> + pci_pcie_type(dev) == PCI_EXP_TYPE_DOWNSTREAM) {
> + *downstream_port = dev;
> + *upstream_port = partner;
> + } else {
> + *downstream_port = partner;
> + *upstream_port = dev;
> + }
> +}
This feels like duplicating similar functionality with the aspm driver
that also wants to infer ends of the link when giving a pci_dev in. The
aspm driver currently does that within, but it kind of duplicating
pci_bus.
It would be nice to avoid the duplication and have something similar for
this in PCI core.
I'd have already tried to move it out of the aspm driver into pci_bus but
I highly suspect pci_bus is allocated too late for it to be trivial to
just embed link information into the struct pci_bus.
--
i.
> +
> +/*
> + * pci_lmr_aspm_inhibit() - Inhibit or restore ASPM L0s/L1 during active
> margining.
> + * PCIe Base Specification Revision 7.0 sec 7.5.3.7 ("Link Control
> Register"):
> + * - To disable/inhibit ASPM, software on Downstream Component (Endpoint /
> Upstream Port)
> + * must disable ASPM prior to disabling ASPM on Upstream Component (Root
> Port / Downstream Port).
> + * - To enable/restore ASPM, software on Upstream Component (Root Port /
> Downstream Port)
> + * must enable ASPM prior to enabling ASPM on Downstream Component
> (Endpoint / Upstream Port).
> + */
> +static void pci_lmr_aspm_inhibit(struct pci_margin_dev *mdev, bool inhibit)
> +{
> + struct pci_dev *downstream_port, *upstream_port;
> + struct pci_dev *partner = mdev->partner;
> + u16 ctl;
> +
> + pci_lmr_get_ports(mdev, &downstream_port, &upstream_port);
> +
> + if (inhibit) {
> + if (mdev->aspm_saved)
> + return;
> +
> + /*
> + * If link partner already saved ASPM state, inherit it to
> + * prevent overwriting with 0.
> + */
> + if (partner && partner->lmr && partner->lmr->aspm_saved) {
> + mdev->saved_dsp_aspm = partner->lmr->saved_dsp_aspm;
> + mdev->saved_usp_aspm = partner->lmr->saved_usp_aspm;
> + mdev->aspm_saved = true;
> + return;
> + }
> +
> + /*
> + * 1. Downstream Component (upstream_port) must be disabled
> + * FIRST per sec 7.5.3.7.
> + */
> + if (upstream_port && pci_is_pcie(upstream_port) &&
> + upstream_port->current_state == PCI_D0) {
> + if (!pcie_capability_read_word(upstream_port,
> PCI_EXP_LNKCTL, &ctl)) {
> + mdev->saved_usp_aspm = ctl &
> PCI_EXP_LNKCTL_ASPMC;
> + pcie_capability_clear_word(upstream_port,
> PCI_EXP_LNKCTL,
> +
> PCI_EXP_LNKCTL_ASPMC);
> + }
> + }
> +
> + /*
> + * 2. Upstream Component (downstream_port) must be disabled
> + * SECOND per sec 7.5.3.7.
> + */
> + if (downstream_port && pci_is_pcie(downstream_port) &&
> + downstream_port->current_state == PCI_D0) {
> + if (!pcie_capability_read_word(downstream_port,
> PCI_EXP_LNKCTL, &ctl)) {
> + mdev->saved_dsp_aspm = ctl &
> PCI_EXP_LNKCTL_ASPMC;
> + pcie_capability_clear_word(downstream_port,
> PCI_EXP_LNKCTL,
> +
> PCI_EXP_LNKCTL_ASPMC);
> + }
> + }
> +
> + mdev->aspm_saved = true;
> +
> + /*
> + * Ensure link is settled in L0 mode per PCIe Base
> + * Specification Revision 7.0 sec 8.4.4.
> + */
> + usleep_range(2000, 3000);
> + } else {
> + if (!mdev->aspm_saved)
> + return;
> +
> + /*
> + * 1. Upstream Component (downstream_port) MUST be restored
> + * FIRST per sec 7.5.3.7.
> + */
> + if (downstream_port && pci_is_pcie(downstream_port) &&
> + downstream_port->current_state == PCI_D0) {
> + pcie_capability_clear_and_set_word(
> + downstream_port, PCI_EXP_LNKCTL,
> PCI_EXP_LNKCTL_ASPMC,
> + mdev->saved_dsp_aspm);
> + }
> +
> + /*
> + * 2. Downstream Component (upstream_port) MUST be restored
> + * SECOND per sec 7.5.3.7.
> + */
> + if (upstream_port && pci_is_pcie(upstream_port) &&
> + upstream_port->current_state == PCI_D0) {
> + pcie_capability_clear_and_set_word(
> + upstream_port, PCI_EXP_LNKCTL,
> PCI_EXP_LNKCTL_ASPMC,
> + mdev->saved_usp_aspm);
> + }
> +
> + mdev->aspm_saved = false;
> + }
> +}
> +
> +/*
> + * pci_lmr_ensure_aspm_inhibited() - Verify and re-enforce ASPM inhibit
> state.
> + *
> + * Checks both Downstream Port and Upstream Port to guarantee that
> out-of-band
> + * OS events (e.g. background power transitions or sysfs modifications) have
> + * not unexpectedly re-enabled ASPM on either component. If ASPM was
> re-enabled,
> + * re-inhibits it in the spec-mandated order and waits for the link to settle
> + * in L0 before physical lane margin steps are executed.
> + */
> +static void pci_lmr_ensure_aspm_inhibited(struct pci_margin_dev *mdev)
> +{
> + struct pci_dev *downstream_port, *upstream_port;
> + u16 dsp_ctl = 0, usp_ctl = 0;
> + bool re_inhibit = false;
> +
> + pci_lmr_get_ports(mdev, &downstream_port, &upstream_port);
> +
> + if (upstream_port && pci_is_pcie(upstream_port)) {
> + if (!pcie_capability_read_word(upstream_port, PCI_EXP_LNKCTL,
> &usp_ctl) &&
> + (usp_ctl & PCI_EXP_LNKCTL_ASPMC))
> + re_inhibit = true;
> + }
> +
> + if (downstream_port && pci_is_pcie(downstream_port)) {
> + if (!pcie_capability_read_word(downstream_port, PCI_EXP_LNKCTL,
> &dsp_ctl) &&
> + (dsp_ctl & PCI_EXP_LNKCTL_ASPMC))
> + re_inhibit = true;
> + }
> +
> + if (re_inhibit) {
> + pci_info_ratelimited(mdev->dev,
> + "ASPM re-enabled unexpectedly;
> re-enforcing ASPM inhibit for LMR\n");
> + /* Disable Downstream Component first, Upstream Component
> second per sec 7.5.3.7 */
> + if (upstream_port && pci_is_pcie(upstream_port))
> + pcie_capability_clear_word(upstream_port,
> PCI_EXP_LNKCTL,
> + PCI_EXP_LNKCTL_ASPMC);
> + if (downstream_port && pci_is_pcie(downstream_port))
> + pcie_capability_clear_word(downstream_port,
> PCI_EXP_LNKCTL,
> + PCI_EXP_LNKCTL_ASPMC);
> + /* Ensure link returns to and settles in L0 mode before
> proceeding */
> + usleep_range(2000, 3000);
> + }
> +}
> +
> +/*
> + * Helpers to manage Autonomous Width/Speed transitions per PCIe Base
> Specification Revision 7.0:
> + * - Section 7.5.3.7 "Link Control Register" (Hardware Autonomous Width
> Disable, bit 9)
> + * - Section 7.5.3.17 "Link Control 2 Register" (Hardware Autonomous Speed
> Disable, bit 5)
> + * - Section 4.2.18.4 "Receiver Margin Testing Requirements"
> + * - Section 8.4.4 "Lane Margining at the Receiver - Electrical Requirements"
> + *
> + * Both Downstream Port and Upstream Port must save and set Hardware
> Autonomous
> + * Width Disable and Hardware Autonomous Speed Disable bits during margining
> to
> + * guarantee that the link remains in a stable active L0 state.
> + */
> +static void pci_lmr_disable_autonomous(struct pci_margin_dev *mdev)
> +{
> + struct pci_dev *downstream_port, *upstream_port;
> + struct pci_dev *partner = mdev->partner;
> + u16 lnkctl, lnkctl2;
> +
> + if (mdev->autonomous_saved)
> + return;
> +
> + /* If link partner already saved autonomous settings, inherit them */
> + if (partner && partner->lmr && partner->lmr->autonomous_saved) {
> + mdev->saved_dsp_lnkctl = partner->lmr->saved_dsp_lnkctl;
> + mdev->saved_dsp_lnkctl2 = partner->lmr->saved_dsp_lnkctl2;
> + mdev->saved_usp_lnkctl = partner->lmr->saved_usp_lnkctl;
> + mdev->saved_usp_lnkctl2 = partner->lmr->saved_usp_lnkctl2;
> + mdev->autonomous_saved = true;
> + return;
> + }
> +
> + pci_lmr_get_ports(mdev, &downstream_port, &upstream_port);
> +
> + /* 1. Downstream Component (upstream_port): Save and Disable FIRST */
> + if (upstream_port && pci_is_pcie(upstream_port) &&
> + upstream_port->current_state == PCI_D0) {
> + if (!pcie_capability_read_word(upstream_port, PCI_EXP_LNKCTL,
> &lnkctl)) {
> + mdev->saved_usp_lnkctl = lnkctl;
> + pcie_capability_set_word(upstream_port, PCI_EXP_LNKCTL,
> + PCI_EXP_LNKCTL_HAWD);
> + }
> +
> + if (!pcie_capability_read_word(upstream_port, PCI_EXP_LNKCTL2,
> &lnkctl2)) {
> + mdev->saved_usp_lnkctl2 = lnkctl2;
> + pcie_capability_set_word(upstream_port, PCI_EXP_LNKCTL2,
> + PCI_EXP_LNKCTL2_HASD);
> + }
> + }
> +
> + /* 2. Upstream Component (downstream_port): Save and Disable SECOND */
> + if (downstream_port && pci_is_pcie(downstream_port) &&
> + downstream_port->current_state == PCI_D0) {
> + if (!pcie_capability_read_word(downstream_port, PCI_EXP_LNKCTL,
> &lnkctl)) {
> + mdev->saved_dsp_lnkctl = lnkctl;
> + pcie_capability_set_word(downstream_port,
> PCI_EXP_LNKCTL,
> + PCI_EXP_LNKCTL_HAWD);
> + }
> +
> + if (!pcie_capability_read_word(downstream_port,
> PCI_EXP_LNKCTL2, &lnkctl2)) {
> + mdev->saved_dsp_lnkctl2 = lnkctl2;
> + pcie_capability_set_word(downstream_port,
> PCI_EXP_LNKCTL2,
> + PCI_EXP_LNKCTL2_HASD);
> + }
> + }
> +
> + mdev->autonomous_saved = true;
> +}
> +
> +static void pci_lmr_restore_autonomous(struct pci_margin_dev *mdev)
> +{
> + struct pci_dev *downstream_port, *upstream_port;
> +
> + if (!mdev->autonomous_saved)
> + return;
> +
> + pci_lmr_get_ports(mdev, &downstream_port, &upstream_port);
> +
> + /* 1. Upstream Component (downstream_port) restored FIRST */
> + if (downstream_port && pci_is_pcie(downstream_port) &&
> + downstream_port->current_state == PCI_D0) {
> + pcie_capability_clear_and_set_word(
> + downstream_port, PCI_EXP_LNKCTL, PCI_EXP_LNKCTL_HAWD,
> + mdev->saved_dsp_lnkctl & PCI_EXP_LNKCTL_HAWD);
> + pcie_capability_clear_and_set_word(
> + downstream_port, PCI_EXP_LNKCTL2, PCI_EXP_LNKCTL2_HASD,
> + mdev->saved_dsp_lnkctl2 & PCI_EXP_LNKCTL2_HASD);
> + }
> +
> + /* 2. Downstream Component (upstream_port) restored SECOND */
> + if (upstream_port && pci_is_pcie(upstream_port) &&
> + upstream_port->current_state == PCI_D0) {
> + pcie_capability_clear_and_set_word(
> + upstream_port, PCI_EXP_LNKCTL, PCI_EXP_LNKCTL_HAWD,
> + mdev->saved_usp_lnkctl & PCI_EXP_LNKCTL_HAWD);
> + pcie_capability_clear_and_set_word(
> + upstream_port, PCI_EXP_LNKCTL2, PCI_EXP_LNKCTL2_HASD,
> + mdev->saved_usp_lnkctl2 & PCI_EXP_LNKCTL2_HASD);
> + }
> +
> + mdev->autonomous_saved = false;
> +}
> +
> +static inline u8 pci_lmr_sts_payload(u16 sts)
> +{
> + return FIELD_GET(PCI_LMR_LANE_STS_PAYLOAD, sts);
> +}
> +
> +/*
> + * pci_lmr_run_cmd() - Issue LMR command to Lane Control and wait for Status.
> + * Must be called with mdev->lock held.
> + */
> +static int pci_lmr_run_cmd(struct pci_margin_dev *mdev, int lane, u8 rx, u8
> type,
> + u8 usage, u8 payload, u16 *status_val)
> +{
> + struct pci_dev *dev;
> + u16 lmr, ctrl_offset, sts_offset;
> + u16 ctrl, sts;
> + unsigned long timeout;
> + int ret;
> +
> + if (!mdev || lane < 0 || lane >= mdev->num_lanes || rx > LMR_MAX_RX_NUM)
> + return -EINVAL;
> +
> + dev = mdev->dev;
> + lmr = mdev->cap;
> + ctrl_offset = lmr + PCI_LMR_LANE_CTRL + LMR_LANE_REG_STRIDE * lane;
> + sts_offset = lmr + PCI_LMR_LANE_STS + LMR_LANE_REG_STRIDE * lane;
> +
> + /*
> + * Per PCIe Base Specification Revision 7.0 sec 4.2.18.2 & Table 4-77
> + * (r6.0 Table 4-73), software must issue NO_CMD (0x7) with payload
> + * 0x9C targeting the specific receiver (rx) to clear MTYPE in Lane
> + * Status before issuing a subsequent command.
> + */
> + if (type != LMR_TYPE_NO_CMD) {
> + ctrl = FIELD_PREP(PCI_LMR_LANE_CTRL_RX_NUM, rx) |
> + FIELD_PREP(PCI_LMR_LANE_CTRL_MTYPE, LMR_TYPE_NO_CMD) |
> + FIELD_PREP(PCI_LMR_LANE_CTRL_USAGE, 0) |
> + FIELD_PREP(PCI_LMR_LANE_CTRL_PAYLOAD,
> + LMR_PAYLOAD_NO_CMD);
> +
> + ret = pci_write_config_word(dev, ctrl_offset, ctrl);
> + if (ret != PCIBIOS_SUCCESSFUL)
> + return pcibios_err_to_errno(ret);
> +
> + timeout = jiffies + msecs_to_jiffies(LMR_CMD_TIMEOUT_MS);
> + while (1) {
> + ret = pci_read_config_word(dev, sts_offset, &sts);
> + if (ret != PCIBIOS_SUCCESSFUL)
> + return pcibios_err_to_errno(ret);
> + if (PCI_POSSIBLE_ERROR(sts))
> + return -ENODEV;
> + if (FIELD_GET(PCI_LMR_LANE_STS_MTYPE, sts) ==
> LMR_TYPE_NO_CMD &&
> + FIELD_GET(PCI_LMR_LANE_STS_RX_NUM, sts) == rx)
> + break;
> + if (time_after(jiffies, timeout))
> + return -ETIMEDOUT;
> + usleep_range(LMR_CMD_SLEEP_MIN_US,
> LMR_CMD_SLEEP_MAX_US);
> + }
> + }
> +
> + ctrl = FIELD_PREP(PCI_LMR_LANE_CTRL_RX_NUM, rx) |
> + FIELD_PREP(PCI_LMR_LANE_CTRL_MTYPE, type) |
> + FIELD_PREP(PCI_LMR_LANE_CTRL_USAGE, usage) |
> + FIELD_PREP(PCI_LMR_LANE_CTRL_PAYLOAD, payload);
> +
> + ret = pci_write_config_word(dev, ctrl_offset, ctrl);
> + if (ret != PCIBIOS_SUCCESSFUL)
> + return pcibios_err_to_errno(ret);
> +
> + timeout = jiffies + msecs_to_jiffies(LMR_CMD_TIMEOUT_MS);
> + while (1) {
> + ret = pci_read_config_word(dev, sts_offset, &sts);
> + if (ret != PCIBIOS_SUCCESSFUL)
> + return pcibios_err_to_errno(ret);
> + if (PCI_POSSIBLE_ERROR(sts))
> + return -ENODEV;
> +
> + if (FIELD_GET(PCI_LMR_LANE_STS_MTYPE, sts) == type &&
> + FIELD_GET(PCI_LMR_LANE_STS_RX_NUM, sts) == rx) {
> + if (status_val)
> + *status_val = sts;
> + return 0;
> + }
> +
> + if (time_after(jiffies, timeout)) {
> + /*
> + * Per PCIe Base Specification Revision 7.0 sec 4.2.18.2
> + * & Table 4-77 (r6.0 Table 4-73), if receiver echoes
> + * NO_CMD (0x7) after command issuance, it indicates
> NAK.
> + */
> + if (FIELD_GET(PCI_LMR_LANE_STS_MTYPE, sts) ==
> LMR_TYPE_NO_CMD &&
> + FIELD_GET(PCI_LMR_LANE_STS_RX_NUM, sts) == rx)
> + return -EOPNOTSUPP;
> + break;
> + }
> +
> + usleep_range(LMR_CMD_SLEEP_MIN_US, LMR_CMD_SLEEP_MAX_US);
> + }
> +
> + return -ETIMEDOUT;
> +}
> +
> +/*
> + * pci_lmr_clear_to_normal_lane() - Clear lane margin back to normal settings
> + * per PCIe Base Specification Revision 7.0 sec 4.2.18.2 & Table 4-77 (r6.0
> Table 4-73).
> + * Issues Set Margining Parameters (MTYPE 010b) with "Go to Normal Settings"
> (Payload 0x0F).
> + */
> +static int pci_lmr_clear_to_normal_lane(struct pci_margin_lane *plane)
> +{
> + u16 sts;
> + int ret;
> +
> + if (!plane || !plane->mdev)
> + return -EINVAL;
> +
> + ret = pci_lmr_run_cmd(plane->mdev, plane->lane, plane->rx,
> + LMR_TYPE_SET_PARAMS, 0, LMR_PAYLOAD_GO_TO_NORMAL,
> + &sts);
> + if (ret)
> + return ret;
> +
> + plane->timing_val = 0;
> + plane->voltage_val = 0;
> + return 0;
> +}
> +
> +static int pci_lmr_cache_rx_info(struct pci_margin_lane *plane, u8 rx)
> +{
> + struct pci_margin_rx_info *info;
> + u16 sts;
> + int ret;
> +
> + if (!plane || rx > LMR_MAX_RX_NUM)
> + return -EINVAL;
> +
> + info = &plane->rx_info[rx];
> +
> + if (info->caps_cached)
> + return 0;
> +
> + /* Issuing REPORT_CAPS aborts active margin; clear to normal settings */
> + ret = pci_lmr_clear_to_normal_lane(plane);
> + if (ret)
> + return ret;
> +
> + /* Report Capabilities: MTYPE 001b, Payload 0x88 */
> + ret = pci_lmr_run_cmd(plane->mdev, plane->lane, rx,
> + LMR_TYPE_REPORT_CAPS, 0, LMR_PAYLOAD_REPORT_CAPS,
> + &sts);
> + if (ret)
> + return ret;
> + info->caps = pci_lmr_sts_payload(sts);
> +
> + /* Report Timing Steps: MTYPE 001b, Payload 0x8A */
> + ret = pci_lmr_run_cmd(plane->mdev, plane->lane, rx,
> + LMR_TYPE_REPORT_CAPS, 0,
> + LMR_PAYLOAD_REPORT_TIM_STEPS, &sts);
> + if (ret)
> + return ret;
> + info->num_timing_steps = FIELD_GET(LMR_TIMING_STEP_MASK,
> pci_lmr_sts_payload(sts));
> +
> + /* Report Voltage Steps: MTYPE 001b, Payload 0x89 */
> + ret = pci_lmr_run_cmd(plane->mdev, plane->lane, rx,
> + LMR_TYPE_REPORT_CAPS, 0,
> + LMR_PAYLOAD_REPORT_VOLT_STEPS, &sts);
> + if (ret)
> + return ret;
> + info->num_voltage_steps = FIELD_GET(LMR_VOLTAGE_STEP_MASK,
> pci_lmr_sts_payload(sts));
> +
> + info->caps_cached = true;
> + return 0;
> +}
> +
> +#if IS_ENABLED(CONFIG_DEBUG_FS)
> +
> +static int margin_caps_show(struct seq_file *s, void *v)
> +{
> + struct pci_margin_dev *mdev = s->private;
> + struct pci_dev *dev = mdev->dev;
> + u16 cap;
> + int ret;
> +
> + /* Wake the hardware and hold the PM reference before accessing
> registers */
> + ret = pm_runtime_resume_and_get(&dev->dev);
> + if (ret < 0)
> + return ret;
> +
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_CAP, &cap);
> + pm_runtime_put_sync(&dev->dev);
> +
> + if (ret != PCIBIOS_SUCCESSFUL)
> + return pcibios_err_to_errno(ret);
> +
> + seq_printf(s, "Port Capabilities: %#06x\n", cap);
> + seq_printf(s, " Uses SW Ready: %s\n",
> + str_yes_no(cap & PCI_LMR_PORT_CAP_USES_SW_READY));
> + return 0;
> +}
> +DEFINE_SHOW_ATTRIBUTE(margin_caps);
> +
> +static int margin_port_status_show(struct seq_file *s, void *v)
> +{
> + struct pci_margin_dev *mdev = s->private;
> + struct pci_dev *dev = mdev->dev;
> + u16 sts;
> + int ret;
> +
> + /* Wake the hardware and hold the PM reference before accessing
> registers */
> + ret = pm_runtime_resume_and_get(&dev->dev);
> + if (ret < 0)
> + return ret;
> +
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
> + pm_runtime_put_sync(&dev->dev);
> +
> + if (ret != PCIBIOS_SUCCESSFUL)
> + return pcibios_err_to_errno(ret);
> +
> + seq_printf(s, "Port Status: %#06x\n", sts);
> + seq_printf(s, " Margining Ready: %s\n",
> + str_yes_no(sts & PCI_LMR_PORT_STS_MARGIN_READY));
> + seq_printf(s, " SW Ready: %s\n",
> + str_yes_no(sts & PCI_LMR_PORT_STS_SW_READY));
> + return 0;
> +}
> +DEFINE_SHOW_ATTRIBUTE(margin_port_status);
> +
> +static int margin_enable_show(struct seq_file *s, void *v)
> +{
> + struct pci_margin_dev *mdev = s->private;
> +
> + guard(mutex)(&mdev->lock);
> + seq_printf(s, "%d\n", mdev->enabled);
> + return 0;
> +}
> +
> +static void pci_lmr_disable_locked(struct pci_margin_dev *mdev)
> +{
> + struct pci_dev *dev;
> + int i, ret;
> + u16 sts;
> +
> + if (!mdev)
> + return;
> +
> + lockdep_assert_held(&mdev->lock);
> +
> + if (!mdev->enabled)
> + return;
> +
> + dev = mdev->dev;
> +
> + for (i = 0; i < mdev->num_lanes; i++)
> + pci_lmr_clear_to_normal_lane(&mdev->lanes[i]);
> +
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, &sts);
> + if (ret == PCIBIOS_SUCCESSFUL) {
> + sts &= ~PCI_LMR_PORT_STS_SW_READY;
> + pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
> + }
> + pci_lmr_aspm_inhibit(mdev, false);
> + pci_lmr_restore_autonomous(mdev);
> +
> + if (mdev->partner) {
> + pm_runtime_put_sync(&mdev->partner->dev);
> + pci_dev_put(mdev->partner);
> + mdev->partner = NULL;
> + }
> +
> + pm_runtime_put_sync(&dev->dev);
> + mdev->enabled = false;
> +}
> +
> +static int pci_lmr_enable_locked(struct pci_margin_dev *mdev,
> + struct pci_dev *downstream_port,
> + struct pci_dev *upstream_port)
> +{
> + struct pci_dev *dev = mdev->dev;
> + struct pci_dev *partner = NULL;
> + unsigned long timeout;
> + u16 sts, cap, lnksta;
> + int ret, i;
> +
> + lockdep_assert_held(&mdev->lock);
> +
> + /* Ensure device is powered (D0) before reading configuration registers
> */
> + ret = pm_runtime_resume_and_get(&dev->dev);
> + if (ret < 0)
> + return ret;
> +
> + partner = (dev == downstream_port) ? upstream_port : downstream_port;
> +
> + /* Prevent concurrent LMR on both ends of the same link */
> + if (partner && partner->lmr && partner->lmr->enabled) {
> + ret = -EBUSY;
> + goto err_rpm;
> + }
> +
> + if (partner) {
> + ret = pm_runtime_resume_and_get(&partner->dev);
> + if (ret < 0)
> + goto err_rpm;
> + mdev->partner = pci_dev_get(partner);
> + }
> +
> + /*
> + * PCIe Base Specification Revision 7.0 sec 8.4.4: LMR is physically
> + * undefined below 16.0 GT/s. Even if a device supports Gen4+, if the
> + * link is currently trained and operating at Gen1..Gen3 speeds
> + * (< 16.0 GT/s) in Link Status Register (sec 7.5.3.8, Current Link
> + * Speed), reject margining.
> + */
> + pcie_capability_read_word(dev, PCI_EXP_LNKSTA, &lnksta);
> + if ((lnksta & PCI_EXP_LNKSTA_CLS) < PCI_EXP_LNKSTA_CLS_16_0GB) {
> + ret = -EOPNOTSUPP;
> + goto err_partner_rpm;
> + }
> +
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_CAP, &cap);
> + if (ret != PCIBIOS_SUCCESSFUL) {
> + ret = pcibios_err_to_errno(ret);
> + goto err_partner_rpm;
> + }
> +
> + /* Disable Autonomous Width and Speed transitions */
> + pci_lmr_disable_autonomous(mdev);
> +
> + /* Inhibit ASPM L0s/L1 during margining with restoration path */
> + pci_lmr_aspm_inhibit(mdev, true);
> +
> + if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS,
> &sts);
> + if (ret != PCIBIOS_SUCCESSFUL) {
> + ret = pcibios_err_to_errno(ret);
> + goto err_aspm;
> + }
> + sts |= PCI_LMR_PORT_STS_SW_READY;
> + pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS, sts);
> + }
> +
> + timeout = jiffies + msecs_to_jiffies(LMR_ENABLE_TIMEOUT_MS);
> + while (1) {
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS,
> &sts);
> + if (ret != PCIBIOS_SUCCESSFUL) {
> + ret = pcibios_err_to_errno(ret);
> + goto err_sw_ready;
> + }
> + if (PCI_POSSIBLE_ERROR(sts)) {
> + ret = -ENODEV;
> + goto err_sw_ready;
> + }
> + if (sts & PCI_LMR_PORT_STS_MARGIN_READY)
> + break;
> + if (time_after(jiffies, timeout)) {
> + ret = -ETIMEDOUT;
> + goto err_sw_ready;
> + }
> + usleep_range(LMR_ENABLE_SLEEP_MIN_US, LMR_ENABLE_SLEEP_MAX_US);
> + }
> +
> + /* Cache capabilities for configured receiver on all lanes */
> + for (i = 0; i < mdev->num_lanes; i++) {
> + ret = pci_lmr_cache_rx_info(&mdev->lanes[i], mdev->lanes[i].rx);
> + if (ret)
> + goto err_sw_ready;
> + }
> + mdev->enabled = true;
> + return 0;
> +
> +err_sw_ready:
> + if (cap & PCI_LMR_PORT_CAP_USES_SW_READY) {
> + u16 clean_sts;
> + int clean_ret;
> +
> + clean_ret = pci_read_config_word(
> + dev, mdev->cap + PCI_LMR_PORT_STS, &clean_sts);
> + if (clean_ret == PCIBIOS_SUCCESSFUL) {
> + clean_sts &= ~PCI_LMR_PORT_STS_SW_READY;
> + pci_write_config_word(dev, mdev->cap + PCI_LMR_PORT_STS,
> + clean_sts);
> + }
> + }
> +err_aspm:
> + pci_lmr_aspm_inhibit(mdev, false);
> + pci_lmr_restore_autonomous(mdev);
> +err_partner_rpm:
> + if (mdev->partner) {
> + pm_runtime_put_sync(&mdev->partner->dev);
> + pci_dev_put(mdev->partner);
> + mdev->partner = NULL;
> + }
> +err_rpm:
> + pm_runtime_put_sync(&dev->dev);
> + return ret;
> +}
> +
> +static ssize_t margin_enable_write(struct file *file,
> + const char __user *user_buf, size_t count,
> + loff_t *ppos)
> +{
> + struct seq_file *s = file->private_data;
> + struct pci_margin_dev *mdev = s->private;
> + struct pci_dev *dev = mdev->dev;
> + struct pci_dev *downstream_port, *upstream_port;
> + bool enable;
> + int ret;
> +
> + ret = kstrtobool_from_user(user_buf, count, &enable);
> + if (ret)
> + return ret;
> +
> + pci_lmr_get_link_partners(dev, &downstream_port, &upstream_port);
> +
> + /* Strict hierarchical lock order: Downstream Port (parent) before
> Upstream Port (child) */
> + if (downstream_port)
> + pci_dev_lock(downstream_port);
> + if (upstream_port && upstream_port != downstream_port)
> + pci_dev_lock(upstream_port);
> +
> + mutex_lock(&mdev->lock);
> +
> + if (mdev->enabled == enable) {
> + ret = count;
> + } else if (!enable) {
> + pci_lmr_disable_locked(mdev);
> + ret = count;
> + } else {
> + ret = pci_lmr_enable_locked(mdev, downstream_port,
> upstream_port);
> + if (!ret)
> + ret = count;
> + }
> +
> + mutex_unlock(&mdev->lock);
> +
> + if (upstream_port && upstream_port != downstream_port)
> + pci_dev_unlock(upstream_port);
> + if (downstream_port)
> + pci_dev_unlock(downstream_port);
> +
> + pci_lmr_put_link_partners(dev, downstream_port, upstream_port);
> +
> + return ret;
> +}
> +
> +static int margin_enable_open(struct inode *inode, struct file *file)
> +{
> + return single_open(file, margin_enable_show, inode->i_private);
> +}
> +
> +static const struct file_operations margin_enable_fops = {
> + .open = margin_enable_open,
> + .read = seq_read,
> + .write = margin_enable_write,
> + .llseek = seq_lseek,
> + .release = single_release,
> +};
> +
> +static int margin_lane_receiver_show(struct seq_file *s, void *v)
> +{
> + struct pci_margin_lane *plane = s->private;
> +
> + guard(mutex)(&plane->mdev->lock);
> + seq_printf(s, "%d\n", plane->rx);
> + return 0;
> +}
> +
> +static ssize_t margin_lane_receiver_write(struct file *file, const char
> __user *user_buf,
> + size_t count, loff_t *ppos)
> +{
> + struct seq_file *s = file->private_data;
> + struct pci_margin_lane *plane = s->private;
> + struct pci_margin_dev *mdev = plane->mdev;
> + int ret;
> + u8 rx;
> +
> + ret = kstrtou8_from_user(user_buf, count, 0, &rx);
> + if (ret)
> + return ret;
> +
> + /*
> + * Valid receiver numbers are 0..6 per PCIe Base Specification
> + * Revision 7.0 sec 4.2.18.1 & Table 4-76 (r6.0 Table 4-72);
> + * 7 is reserved.
> + */
> + if (rx > LMR_MAX_RX_NUM)
> + return -EINVAL;
> +
> + guard(mutex)(&mdev->lock);
> + if (plane->rx == rx)
> + return count;
> +
> + if (mdev->enabled) {
> + /* Clear previous receiver to normal settings per
> single-receiver rule */
> + ret = pci_lmr_clear_to_normal_lane(plane);
> + if (ret)
> + return ret;
> + ret = pci_lmr_cache_rx_info(plane, rx);
> + if (ret)
> + return ret;
> + }
> +
> + plane->rx = rx;
> + return count;
> +}
> +
> +static int margin_lane_receiver_open(struct inode *inode, struct file *file)
> +{
> + return single_open(file, margin_lane_receiver_show, inode->i_private);
> +}
> +
> +static const struct file_operations margin_lane_receiver_fops = {
> + .open = margin_lane_receiver_open,
> + .read = seq_read,
> + .write = margin_lane_receiver_write,
> + .llseek = seq_lseek,
> + .release = single_release,
> +};
> +
> +static int margin_lane_caps_show(struct seq_file *s, void *v)
> +{
> + struct pci_margin_lane *plane = s->private;
> + struct pci_margin_dev *mdev = plane->mdev;
> + struct pci_margin_rx_info *info;
> + int ret;
> + u8 val;
> +
> + guard(mutex)(&mdev->lock);
> + if (!mdev->enabled)
> + return -EBUSY;
> +
> + ret = pci_lmr_cache_rx_info(plane, plane->rx);
> + if (ret)
> + return ret;
> +
> + info = &plane->rx_info[plane->rx];
> + val = info->caps;
> + seq_printf(s, "Lane %d Rx %d Capabilities: %#02x\n", plane->lane,
> plane->rx, val);
> + seq_printf(s, " Voltage Supported: %s\n",
> + str_yes_no(val & LMR_CAP_VOLTAGE_SUPPORTED));
> + seq_printf(s, " Left/Right: %s\n",
> + (val & LMR_CAP_IND_LEFT_RIGHT_TIMING) ? "independent" :
> "symmetric");
> + seq_printf(s, " Up/Down: %s\n",
> + (val & LMR_CAP_IND_UP_DOWN_VOLTAGE) ? "independent" :
> "symmetric");
> + seq_printf(s, " Error Sampler: %s\n",
> + (val & LMR_CAP_IND_ERROR_SAMPLER) ? "independent" :
> + "main sampler");
> + seq_printf(s, " Sample Reporting: %s\n",
> + (val & LMR_CAP_SAMPLE_REPORT_METHOD) ? "rate" : "count");
> + return 0;
> +}
> +DEFINE_SHOW_ATTRIBUTE(margin_lane_caps);
> +
> +static int margin_lane_steps_show(struct seq_file *s, u8 type)
> +{
> + struct pci_margin_lane *plane = s->private;
> + struct pci_margin_dev *mdev = plane->mdev;
> + struct pci_margin_rx_info *info;
> + int ret;
> +
> + guard(mutex)(&mdev->lock);
> + if (!mdev->enabled)
> + return -EBUSY;
> +
> + ret = pci_lmr_cache_rx_info(plane, plane->rx);
> + if (ret)
> + return ret;
> +
> + info = &plane->rx_info[plane->rx];
> + seq_printf(s, "%d\n", (type == LMR_TYPE_VOLTAGE) ?
> + info->num_voltage_steps : info->num_timing_steps);
> + return 0;
> +}
> +
> +static int margin_lane_timing_steps_show(struct seq_file *s, void *v)
> +{
> + return margin_lane_steps_show(s, LMR_TYPE_TIMING);
> +}
> +DEFINE_SHOW_ATTRIBUTE(margin_lane_timing_steps);
> +
> +static int margin_lane_voltage_steps_show(struct seq_file *s, void *v)
> +{
> + return margin_lane_steps_show(s, LMR_TYPE_VOLTAGE);
> +}
> +DEFINE_SHOW_ATTRIBUTE(margin_lane_voltage_steps);
> +
> +/*
> + * pci_lmr_check_sample_multiple_rx() - Check multi-receiver concurrency.
> + * Per PCIe Base Specification Revision 7.0 sec 4.2.18.2 & sec 8.4.4:
> + * "For Receivers where MIndErrorSampler is 0b, at most one such Receiver is
> + * permitted to be margined at a time. However, margining may be performed on
> + * multiple Lanes simultaneously, as long as it is within the maximum number
> of
> + * Lanes the device supports."
> + *
> + * If the target receiver uses an independent error sampler
> (MIndErrorSampler == 1b),
> + * margining will not produce errors in the live data stream, and multiple
> receivers
> + * may be margined concurrently. If MIndErrorSampler is 0b (main data
> sampler),
> + * software must ensure that no other receiver on any lane is currently
> margined.
> + */
> +static bool pci_lmr_check_sample_multiple_rx(struct pci_margin_dev *mdev,
> + struct pci_margin_lane *plane)
> +{
> + struct pci_margin_rx_info *info = &plane->rx_info[plane->rx];
> + int i;
> +
> + /* If receiver has an independent error sampler, concurrent margining
> is permitted */
> + if (info->caps & LMR_CAP_IND_ERROR_SAMPLER)
> + return true;
> +
> + for (i = 0; i < mdev->num_lanes; i++) {
> + struct pci_margin_lane *other = &mdev->lanes[i];
> + struct pci_margin_rx_info *other_info;
> +
> + if (i == plane->lane)
> + continue;
> +
> + other_info = &other->rx_info[other->rx];
> + /*
> + * For receivers using the main data sampler, reject only if
> another lane
> + * is actively margining a DIFFERENT receiver that ALSO uses
> the main data sampler.
> + */
> + if (other->rx != plane->rx &&
> + !(other_info->caps & LMR_CAP_IND_ERROR_SAMPLER) &&
> + (other->timing_val != 0 || other->voltage_val != 0))
> + return false;
> + }
> + return true;
> +}
> +
> +static ssize_t margin_lane_step_write(struct file *file, const char __user
> *user_buf,
> + size_t count, u8 type)
> +{
> + struct seq_file *s = file->private_data;
> + struct pci_margin_lane *plane = s->private;
> + struct pci_margin_dev *mdev = plane->mdev;
> + struct pci_margin_rx_info *info;
> + u8 step, dir, payload;
> + int max_step, val, ret;
> + u16 sts;
> + u8 caps;
> +
> + ret = kstrtoint_from_user(user_buf, count, 0, &val);
> + if (ret)
> + return ret;
> +
> + guard(mutex)(&mdev->lock);
> + if (!mdev->enabled)
> + return -EBUSY;
> +
> + /*
> + * Ensure ASPM remains inhibited on both link partners before issuing
> + * margin steps. PCIe Base Specification Revision 7.0 sec 8.4.4 requires
> + * the link to stay in L0.
> + */
> + pci_lmr_ensure_aspm_inhibited(mdev);
> +
> + if (val == 0) {
> + /* Step this specific axis to 0 without resetting the
> orthogonal axis */
> + if (type == LMR_TYPE_TIMING) {
> + if (plane->voltage_val == 0) {
> + ret = pci_lmr_clear_to_normal_lane(plane);
> + } else {
> + ret = pci_lmr_run_cmd(mdev, plane->lane,
> plane->rx,
> + LMR_TYPE_TIMING, 0, 0,
> &sts);
> + if (!ret)
> + plane->timing_val = 0;
> + }
> + } else {
> + if (plane->timing_val == 0) {
> + ret = pci_lmr_clear_to_normal_lane(plane);
> + } else {
> + ret = pci_lmr_run_cmd(mdev, plane->lane,
> plane->rx,
> + LMR_TYPE_VOLTAGE, 0, 0,
> &sts);
> + if (!ret)
> + plane->voltage_val = 0;
> + }
> + }
> + return ret ? ret : count;
> + }
> +
> + ret = pci_lmr_cache_rx_info(plane, plane->rx);
> + if (ret)
> + return ret;
> +
> + if (!pci_lmr_check_sample_multiple_rx(mdev, plane))
> + return -EBUSY;
> +
> + info = &plane->rx_info[plane->rx];
> + caps = info->caps;
> +
> + switch (type) {
> + case LMR_TYPE_TIMING:
> + if (val < -LMR_MAX_TIMING_STEP || val > LMR_MAX_TIMING_STEP)
> + return -EINVAL;
> + if (val < 0) {
> + if (!(caps & LMR_CAP_IND_LEFT_RIGHT_TIMING))
> + return -EINVAL;
> + step = -val;
> + dir = LMR_STEP_DIR_LEFT_OR_DOWN; /* 1b: Left */
> + } else {
> + step = val;
> + dir = LMR_STEP_DIR_RIGHT_OR_UP; /* 0b: Right or
> Symmetric (Reserved 0b) */
> + }
> + max_step = info->num_timing_steps;
> + if (step > max_step)
> + return -EINVAL;
> +
> + payload = FIELD_PREP(LMR_TIMING_DIR_MASK, dir) |
> + FIELD_PREP(LMR_TIMING_STEP_MASK, step);
> + ret = pci_lmr_run_cmd(mdev, plane->lane, plane->rx,
> + LMR_TYPE_TIMING, 0, payload, &sts);
> + if (ret)
> + return ret;
> + if (FIELD_GET(LMR_STS_EXEC_MASK, pci_lmr_sts_payload(sts)) ==
> + LMR_STS_EXEC_NAK)
> + return -EOPNOTSUPP;
> + if (FIELD_GET(LMR_STS_EXEC_MASK, pci_lmr_sts_payload(sts)) ==
> + LMR_STS_EXEC_TOO_MANY_ERR) {
> + plane->timing_val = 0;
> + plane->voltage_val = 0;
> + return -EIO;
> + }
> + plane->timing_val = val;
> + break;
> +
> + case LMR_TYPE_VOLTAGE:
> + if (!(caps & LMR_CAP_VOLTAGE_SUPPORTED))
> + return -EOPNOTSUPP;
> + if (val < -LMR_MAX_VOLTAGE_STEP || val > LMR_MAX_VOLTAGE_STEP)
> + return -EINVAL;
> + if (val < 0) {
> + if (!(caps & LMR_CAP_IND_UP_DOWN_VOLTAGE))
> + return -EINVAL;
> + step = -val;
> + dir = LMR_STEP_DIR_LEFT_OR_DOWN; /* 1b: Down */
> + } else {
> + step = val;
> + dir = LMR_STEP_DIR_RIGHT_OR_UP; /* 0b: Up or Symmetric
> (Reserved 0b) */
> + }
> + max_step = info->num_voltage_steps;
> + if (step > max_step)
> + return -EINVAL;
> +
> + payload = FIELD_PREP(LMR_VOLTAGE_DIR_MASK, dir) |
> + FIELD_PREP(LMR_VOLTAGE_STEP_MASK, step);
> + ret = pci_lmr_run_cmd(mdev, plane->lane, plane->rx,
> + LMR_TYPE_VOLTAGE, 0, payload, &sts);
> + if (ret)
> + return ret;
> + if (FIELD_GET(LMR_STS_EXEC_MASK, pci_lmr_sts_payload(sts)) ==
> + LMR_STS_EXEC_NAK)
> + return -EOPNOTSUPP;
> + if (FIELD_GET(LMR_STS_EXEC_MASK, pci_lmr_sts_payload(sts)) ==
> + LMR_STS_EXEC_TOO_MANY_ERR) {
> + plane->timing_val = 0;
> + plane->voltage_val = 0;
> + return -EIO;
> + }
> + plane->voltage_val = val;
> + break;
> +
> + default:
> + return -EINVAL;
> + }
> +
> + return count;
> +}
> +
> +static ssize_t margin_lane_timing_write(struct file *file, const char __user
> *user_buf,
> + size_t count, loff_t *ppos)
> +{
> + return margin_lane_step_write(file, user_buf, count, LMR_TYPE_TIMING);
> +}
> +
> +static int margin_lane_step_show(struct seq_file *s, u8 type)
> +{
> + struct pci_margin_lane *plane = s->private;
> +
> + guard(mutex)(&plane->mdev->lock);
> + seq_printf(s, "%d\n", (type == LMR_TYPE_VOLTAGE) ?
> + plane->voltage_val : plane->timing_val);
> + return 0;
> +}
> +
> +static int margin_lane_timing_show(struct seq_file *s, void *v)
> +{
> + return margin_lane_step_show(s, LMR_TYPE_TIMING);
> +}
> +
> +static int margin_lane_timing_open(struct inode *inode, struct file *file)
> +{
> + return single_open(file, margin_lane_timing_show, inode->i_private);
> +}
> +
> +static const struct file_operations margin_lane_timing_fops = {
> + .open = margin_lane_timing_open,
> + .read = seq_read,
> + .write = margin_lane_timing_write,
> + .llseek = seq_lseek,
> + .release = single_release,
> +};
> +
> +static ssize_t margin_lane_voltage_write(struct file *file, const char
> __user *user_buf,
> + size_t count, loff_t *ppos)
> +{
> + return margin_lane_step_write(file, user_buf, count, LMR_TYPE_VOLTAGE);
> +}
> +
> +static int margin_lane_voltage_show(struct seq_file *s, void *v)
> +{
> + return margin_lane_step_show(s, LMR_TYPE_VOLTAGE);
> +}
> +
> +static int margin_lane_voltage_open(struct inode *inode, struct file *file)
> +{
> + return single_open(file, margin_lane_voltage_show, inode->i_private);
> +}
> +
> +static const struct file_operations margin_lane_voltage_fops = {
> + .open = margin_lane_voltage_open,
> + .read = seq_read,
> + .write = margin_lane_voltage_write,
> + .llseek = seq_lseek,
> + .release = single_release,
> +};
> +
> +static void pci_margin_debugfs_init(struct pci_margin_dev *mdev)
> +{
> + struct pci_dev *dev = mdev->dev;
> + struct dentry *parent;
> + char dirname[64];
> + int i;
> +
> + parent = get_pci_debugfs_root();
> + scnprintf(dirname, sizeof(dirname), "pcie_lmr_%s", dev_name(&dev->dev));
> + mdev->debugfs = debugfs_create_dir(dirname, parent);
> +
> + debugfs_create_file("capabilities", 0444, mdev->debugfs, mdev,
> &margin_caps_fops);
> + debugfs_create_file("port_status", 0444, mdev->debugfs, mdev,
> &margin_port_status_fops);
> + debugfs_create_file("enable", 0644, mdev->debugfs, mdev,
> &margin_enable_fops);
> +
> + for (i = 0; i < mdev->num_lanes; i++) {
> + struct pci_margin_lane *plane = &mdev->lanes[i];
> + struct dentry *lane_dir;
> + char lane_name[16];
> +
> + scnprintf(lane_name, sizeof(lane_name), "lane%d", i);
> + lane_dir = debugfs_create_dir(lane_name, mdev->debugfs);
> +
> + debugfs_create_file("receiver", 0644, lane_dir, plane,
> &margin_lane_receiver_fops);
> + debugfs_create_file("caps", 0444, lane_dir, plane,
> &margin_lane_caps_fops);
> + debugfs_create_file("num_timing_steps", 0444, lane_dir, plane,
> + &margin_lane_timing_steps_fops);
> + debugfs_create_file("num_voltage_steps", 0444, lane_dir, plane,
> + &margin_lane_voltage_steps_fops);
> + debugfs_create_file("margin_timing", 0644, lane_dir, plane,
> + &margin_lane_timing_fops);
> + debugfs_create_file("margin_voltage", 0644, lane_dir, plane,
> + &margin_lane_voltage_fops);
> + }
> +}
> +
> +static void pci_margin_debugfs_remove(struct pci_margin_dev *mdev)
> +{
> + debugfs_remove_recursive(mdev->debugfs);
> +}
> +
> +#else
> +static inline void pci_margin_debugfs_init(struct pci_margin_dev *mdev) { }
> +static inline void pci_margin_debugfs_remove(struct pci_margin_dev *mdev) { }
> +#endif
> +
> +void pci_lmr_init(struct pci_dev *dev)
> +{
> + struct pci_margin_dev *mdev;
> + enum pci_bus_speed speed;
> + u32 lnkcap;
> + u16 lmr;
> + int num_lanes, ret, i;
> +
> + if (WARN_ON_ONCE(!dev) || !pci_is_pcie(dev))
> + return;
> +
> + /*
> + * Per PCIe Base Specification Revision 7.0 sec 7.7.11:
> + * For devices associated with an Upstream Port (Endpoints),
> + * the Lane Margining Extended Capability must be implemented in
> + * Function 0 (and only Function 0).
> + */
> + if (pci_is_pcie(dev) && pci_pcie_type(dev) == PCI_EXP_TYPE_ENDPOINT &&
> + PCI_FUNC(dev->devfn) != 0)
> + return;
> +
> + speed = pcie_get_speed_cap(dev);
> + if (speed < PCIE_SPEED_16_0GT || speed == PCI_SPEED_UNKNOWN)
> + return;
> +
> + lmr = pci_find_ext_capability(dev, PCI_EXT_CAP_ID_LMR);
> + if (!lmr) {
> + if (speed >= PCIE_SPEED_64_0GT)
> + pci_warn(dev,
> + "Missing Lane Margining at Receiver Capability
> (mandatory for Gen6+)\n");
> + else
> + pci_dbg(dev,
> + "Optional Lane Margining at Receiver Capability
> not found\n");
> + return;
> + }
> +
> + /*
> + * Maximum Link Width (MLW) per PCIe Base Specification Revision 7.0
> sec 7.5.3.6
> + * ("Link Capabilities Register", bits 9:4).
> + */
> + ret = pcie_capability_read_dword(dev, PCI_EXP_LNKCAP, &lnkcap);
> + if (ret != PCIBIOS_SUCCESSFUL)
> + return;
> + num_lanes = FIELD_GET(PCI_EXP_LNKCAP_MLW, lnkcap);
> + if (num_lanes == 0 || num_lanes > LMR_MAX_LANES) {
> + pci_warn(dev, "Invalid link width %d for LMR\n", num_lanes);
> + return;
> + }
> +
> + dev->lmr_cap = lmr;
> +
> + mdev = kzalloc(struct_size(mdev, lanes, num_lanes), GFP_KERNEL);
> + if (!mdev)
> + return;
> +
> + mdev->num_lanes = num_lanes;
> + mdev->dev = dev;
> + mdev->cap = lmr;
> + mutex_init(&mdev->lock);
> +
> + for (i = 0; i < num_lanes; i++) {
> + mdev->lanes[i].mdev = mdev;
> + mdev->lanes[i].lane = i;
> + mdev->lanes[i].rx = LMR_RX_LOCAL;
> + }
> +
> + pci_margin_debugfs_init(mdev);
> +
> + dev->lmr = mdev;
> +
> + pci_dbg(dev, "Lane Margining at Receiver (Gen%u) Capability detected\n",
> + LMR_SPEED_TO_GEN(speed));
> +}
> +
> +void pci_lmr_exit(struct pci_dev *dev)
> +{
> + struct pci_margin_dev *mdev;
> +
> + if (!dev || !dev->lmr)
> + return;
> +
> + mdev = dev->lmr;
> +
> + /* 1. Tear down user-facing debugfs files FIRST to prevent concurrent
> access */
> + pci_margin_debugfs_remove(mdev);
> +
> + /* 2. Disarm dev->lmr under device_lock to serialize with pci_reset_lmr
> */
> + pci_dev_lock(dev);
> + mdev = dev->lmr;
> + if (!mdev) {
> + pci_dev_unlock(dev);
> + return;
> + }
> + scoped_guard(mutex, &mdev->lock) {
> + dev->lmr = NULL;
> + pci_lmr_disable_locked(mdev);
> + }
> + pci_dev_unlock(dev);
> +
> + /* 3. Safe to destroy structures */
> + mutex_destroy(&mdev->lock);
> + kfree(mdev);
> +}
> +
> +void pci_suspend_lmr(struct pci_dev *dev)
> +{
> + struct pci_margin_dev *mdev = dev->lmr;
> +
> + if (!dev || !mdev)
> + return;
> +
> + guard(mutex)(&mdev->lock);
> + pci_lmr_disable_locked(mdev);
> +}
> +
> +void pci_reset_lmr(struct pci_dev *dev)
> +{
> + struct pci_margin_dev *mdev;
> + u16 sts;
> + int ret, i;
> +
> + if (!dev || pci_dev_is_removed(dev))
> + return;
> +
> + device_lock_assert(&dev->dev);
> +
> + mdev = dev->lmr;
> + if (!mdev)
> + return;
> +
> + guard(mutex)
> + (&mdev->lock);
> + if (mdev->enabled) {
> + for (i = 0; i < mdev->num_lanes; i++) {
> + /*
> + * FLR does not reset Physical Layer registers like LMR.
> + * Return physical samplers to nominal in hardware to
> prevent
> + * persistent receiver eye skew.
> + */
> + pci_lmr_clear_to_normal_lane(&mdev->lanes[i]);
> + mdev->lanes[i].timing_val = 0;
> + mdev->lanes[i].voltage_val = 0;
> + }
> +
> + /* Clear SW_READY in hardware to reset margining state machine
> */
> + ret = pci_read_config_word(dev, mdev->cap + PCI_LMR_PORT_STS,
> &sts);
> + if (ret == PCIBIOS_SUCCESSFUL) {
> + sts &= ~PCI_LMR_PORT_STS_SW_READY;
> + pci_write_config_word(dev, mdev->cap +
> PCI_LMR_PORT_STS, sts);
> + }
> +
> + /* Restore original hardware ASPM before saved states can seal
> the leak */
> + pci_lmr_aspm_inhibit(mdev, false);
> + pci_lmr_restore_autonomous(mdev);
> +
> + if (mdev->partner) {
> + /*
> + * Drop remote partner's PM reference and schedule idle
> check
> + * asynchronously so the partner does not remain
> stranded in
> + * RPM_ACTIVE (D0) indefinitely.
> + */
> + pm_runtime_put(&mdev->partner->dev);
> + pci_dev_put(mdev->partner);
> + mdev->partner = NULL;
> + }
> +
> + /*
> + * Decrement runtime PM usage counter without triggering
> synchronous
> + * suspend, ensuring the device remains in D0 during
> pci_save_state()
> + * and the subsequent reset sequence.
> + */
> + pm_runtime_put_noidle(&dev->dev);
> + mdev->enabled = false;
> + }
> +}
> diff --git a/drivers/pci/probe.c b/drivers/pci/probe.c
> index dd0abbc63e18..352b95568ebf 100644
> --- a/drivers/pci/probe.c
> +++ b/drivers/pci/probe.c
> @@ -2666,6 +2666,7 @@ static void pci_init_capabilities(struct pci_dev *dev)
> pci_pasid_init(dev); /* Process Address Space ID */
> pci_acs_init(dev); /* Access Control Services */
> pci_ptm_init(dev); /* Precision Time Measurement */
> + pci_lmr_init(dev); /* Lane Margining at Receiver */
> pci_aer_init(dev); /* Advanced Error Reporting */
> pci_dpc_init(dev); /* Downstream Port Containment */
> pci_rcec_init(dev); /* Root Complex Event Collector */
> diff --git a/drivers/pci/remove.c b/drivers/pci/remove.c
> index d8bffa21498a..11f8d129d8ea 100644
> --- a/drivers/pci/remove.c
> +++ b/drivers/pci/remove.c
> @@ -36,7 +36,7 @@ static void pci_destroy_dev(struct pci_dev *dev)
>
> pci_doe_sysfs_teardown(dev);
> pci_npem_remove(dev);
> -
> + pci_lmr_exit(dev);
> /*
> * While device is in D0 drop the device from TSM link operations
> * including unbind and disconnect (IDE + SPDM teardown).
> diff --git a/include/linux/pci.h b/include/linux/pci.h
> index 64b308b6e61c..ef1275f4c5b6 100644
> --- a/include/linux/pci.h
> +++ b/include/linux/pci.h
> @@ -349,6 +349,8 @@ struct rcec_ea;
> * number resources to allow for hierarchy expansion.
> * @is_pciehp: PCIe Hot-Plug Capable bridge.
> */
> +struct pci_margin_dev;
> +
> struct pci_dev {
> struct list_head bus_list; /* Node in per-bus list */
> struct pci_bus *bus; /* Bus this device is on */
> @@ -528,6 +530,10 @@ struct pci_dev {
> atomic_t ptm_enable_cnt;
> u8 ptm_granularity;
> #endif
> +#ifdef CONFIG_PCIE_LMR
> + u16 lmr_cap; /* Lane Margining Capability */
> + struct pci_margin_dev *lmr;
> +#endif
> #ifdef CONFIG_PCI_MSI
> void __iomem *msix_base;
> raw_spinlock_t msi_lock;
> diff --git a/include/uapi/linux/pci_regs.h b/include/uapi/linux/pci_regs.h
> index facaa324bd86..90cbe310e62f 100644
> --- a/include/uapi/linux/pci_regs.h
> +++ b/include/uapi/linux/pci_regs.h
> @@ -757,6 +757,7 @@
> #define PCI_EXT_CAP_ID_VF_REBAR 0x24 /* VF Resizable BAR */
> #define PCI_EXT_CAP_ID_DLF 0x25 /* Data Link Feature */
> #define PCI_EXT_CAP_ID_PL_16GT 0x26 /* Physical Layer 16.0 GT/s */
> +#define PCI_EXT_CAP_ID_LMR 0x27 /* Lane Margining at Receiver */
> #define PCI_EXT_CAP_ID_NPEM 0x29 /* Native PCIe Enclosure Management */
> #define PCI_EXT_CAP_ID_PL_32GT 0x2A /* Physical Layer 32.0 GT/s */
> #define PCI_EXT_CAP_ID_DOE 0x2E /* Data Object Exchange */
> @@ -1181,6 +1182,23 @@
> #define PCI_PL_16GT_LE_CTRL_USP_TX_PRESET_MASK 0x000000F0
> #define PCI_PL_16GT_LE_CTRL_USP_TX_PRESET_SHIFT 4
>
> +/* Lane Margining at Receiver */
> +#define PCI_LMR_PORT_CAP 0x04 /* Margining Port Capabilities
> */
> +#define PCI_LMR_PORT_CAP_USES_SW_READY 0x0001 /* Margining Uses
> Software Ready */
> +#define PCI_LMR_PORT_STS 0x06 /* Margining Port Status */
> +#define PCI_LMR_PORT_STS_MARGIN_READY 0x0001 /* Margining Ready */
> +#define PCI_LMR_PORT_STS_SW_READY 0x0002 /* Margining SW Ready */
> +#define PCI_LMR_LANE_CTRL 0x08 /* Margining Lane Control */
> +#define PCI_LMR_LANE_CTRL_RX_NUM 0x0007 /* Receiver Number */
> +#define PCI_LMR_LANE_CTRL_MTYPE 0x0038 /* Margining Type */
> +#define PCI_LMR_LANE_CTRL_USAGE 0x0040 /* Margining Usage Model */
> +#define PCI_LMR_LANE_CTRL_PAYLOAD 0xFF00 /* Margining Payload */
> +#define PCI_LMR_LANE_STS 0x0A /* Margining Lane Status */
> +#define PCI_LMR_LANE_STS_RX_NUM 0x0007 /* Receiver Number */
> +#define PCI_LMR_LANE_STS_MTYPE 0x0038 /* Margining Type */
> +#define PCI_LMR_LANE_STS_USAGE 0x0040 /* Margining Usage
> Model */
> +#define PCI_LMR_LANE_STS_PAYLOAD 0xFF00 /* Margining Payload */
> +
> /* Physical Layer 32.0 GT/s */
> #define PCI_PL_32GT_LE_CTRL 0x20 /* Lane Equalization Control Register */
>
> diff --git a/tools/testing/selftests/Makefile
> b/tools/testing/selftests/Makefile
> index 8a4b6ddc68df..6990d999388a 100644
> --- a/tools/testing/selftests/Makefile
> +++ b/tools/testing/selftests/Makefile
> @@ -91,6 +91,7 @@ TARGETS += net/tcp_ao
> TARGETS += nolibc
> TARGETS += pci_endpoint
> TARGETS += pcie_bwctrl
> +TARGETS += pcie_lmt
> TARGETS += perf_events
> TARGETS += pidfd
> TARGETS += pid_namespace
> diff --git a/tools/testing/selftests/pcie_lmt/Makefile
> b/tools/testing/selftests/pcie_lmt/Makefile
> new file mode 100644
> index 000000000000..36ac85937d78
> --- /dev/null
> +++ b/tools/testing/selftests/pcie_lmt/Makefile
> @@ -0,0 +1,3 @@
> +# SPDX-License-Identifier: GPL-2.0
> +TEST_PROGS = pcie_lmt.sh
> +include ../lib.mk
> diff --git a/tools/testing/selftests/pcie_lmt/pcie_lmt.sh
> b/tools/testing/selftests/pcie_lmt/pcie_lmt.sh
> new file mode 100755
> index 000000000000..22c00c2b8956
> --- /dev/null
> +++ b/tools/testing/selftests/pcie_lmt/pcie_lmt.sh
> @@ -0,0 +1,105 @@
> +#!/bin/bash
> +# SPDX-License-Identifier: GPL-2.0
> +#
> +# Copyright (C) 2026 Google LLC
> +# Author: Priyank Rathod <[email protected]>
> +#
> +# Kselftest for PCIe Lane Margining at Receiver (LMR / LMT)
> +# Tests the debugfs interface exposed by drivers/pci/pcie/margin.c
> +# (/sys/kernel/debug/pci/pcie_lmr_<pci_dev_name>/)
> +
> +set -e
> +
> +TESTNAME="pcie_lmt"
> +
> +# Kselftest framework requirement - SKIP code is 4.
> +ksft_skip=4
> +retval=0
> +skipmsg="skip all tests:"
> +
> +if [ $UID != 0 ]; then
> + echo "$skipmsg must be run as root" >&2
> + exit $ksft_skip
> +fi
> +
> +DEBUGFS=$(mount -t debugfs | head -1 | awk '{ print $3 }')
> +if [ -z "$DEBUGFS" ]; then
> + if [ -d "/sys/kernel/debug" ]; then
> + DEBUGFS="/sys/kernel/debug"
> + else
> + echo "$skipmsg debugfs is not mounted" >&2
> + exit $ksft_skip
> + fi
> +fi
> +
> +if [ ! -d "$DEBUGFS/pci" ]; then
> + # Allow searching debugfs root or pci directory
> + :
> +fi
> +
> +LMR_DEVS=$(ls -d $DEBUGFS/pci/pcie_lmr_* $DEBUGFS/pcie_lmr_* 2>/dev/null ||
> true)
> +if [ -z "$LMR_DEVS" ]; then
> + echo "$skipmsg no PCIe LMR devices found in $DEBUGFS/" >&2
> + exit $ksft_skip
> +fi
> +
> +cleanup_dev()
> +{
> + local dev="$1"
> + echo 0 > "$dev/enable" 2>/dev/null || true
> +}
> +
> +echo "$TESTNAME: testing PCIe LMR debugfs entries"
> +
> +for dev in $LMR_DEVS; do
> + dev_name=$(basename "$dev")
> + echo "$TESTNAME: probing device $dev_name"
> +
> + if [ ! -r "$dev/capabilities" ] || [ ! -r "$dev/port_status" ] ||
> + [ ! -r "$dev/enable" ] || [ ! -w "$dev/enable" ]; then
> + echo "$TESTNAME: $dev_name missing mandatory root attributes"
> + retval=1
> + continue
> + fi
> +
> + caps=$(cat "$dev/capabilities")
> + status=$(cat "$dev/port_status")
> + echo " $dev_name: capabilities read OK"
> + echo " $dev_name: port_status read OK"
> +
> + trap 'cleanup_dev "$dev"' EXIT
> +
> + if ! echo 1 > "$dev/enable" 2>/dev/null; then
> + echo " $dev_name: margining not ready by hardware (skipping
> active lanes)"
> + continue
> + fi
> +
> + echo " $dev_name: margining enabled OK"
> +
> + for lane_dir in $(ls -d "$dev"/lane* 2>/dev/null || true); do
> + lane=$(basename "$lane_dir")
> + echo " $dev_name: testing $lane"
> +
> + # Test setting receiver (Rx 0 is always local receiver)
> + echo 0 > "$lane_dir/receiver"
> + cat "$lane_dir/caps" > /dev/null
> + cat "$lane_dir/num_timing_steps" > /dev/null
> + cat "$lane_dir/num_voltage_steps" > /dev/null
> +
> + # Test resetting timing and voltage margin
> + echo 0 > "$lane_dir/margin_timing"
> + echo 0 > "$lane_dir/margin_voltage"
> + done
> +
> + echo 0 > "$dev/enable"
> + trap - EXIT
> + echo " $dev_name: margining disabled OK"
> +done
> +
> +if [ $retval -eq 0 ]; then
> + echo "$TESTNAME [PASS]"
> +else
> + echo "$TESTNAME [FAIL]"
> +fi
> +
> +exit $retval
>
> ---
> base-commit: 0f23d56f17fdfc7db69d51f64c8b91bbab947aa9
> change-id: 20260818-pcie-lmt-3044d586aaec
>
> Best regards,
>
--
i.