From: Mohammad Shuab Siddique <[email protected]> Hardware already reports an invalid/bad DMA address on a Tx BD via the TX_CMPL_ERRORS_DMA_ERROR bit in the Tx completion record, but the driver never inspected it, so a bad mbuf->buf_iova on Tx completed silently with no visibility.
Check the bit in bnxt_handle_tx_cp() and in the AVX2/SSE/NEON vector Tx-completion handlers, and count occurrences in a new per-queue tx_dma_err counter. The counter is folded into the standard oerrors stat and also exposed as a named xstat (tx_dma_err_cmpl) for finer-grained visibility. The xstat name reflects what is actually counted: one increment per Tx completion record that carries the error bit, not one per packet -- a coalesced Tx completion can cover several packets. The xstat itself is also a port-level sum across queues, not a per-queue value, even though the underlying counter field is per-queue. tx_dma_err has a single writer (the lcore polling that Tx queue's completions) and is read from other threads via xstats, so the increment stores the new value with rte_atomic_store_explicit() after a relaxed load, rather than rte_atomic_fetch_add_explicit(): with only one writer there is nothing to race against for the read-modify-write itself, so the stronger fetch_add (a locked RMW on most architectures) is unneeded on this per-packet path; the atomic store still ensures other threads never observe a torn value. Signed-off-by: Mohammad Shuab Siddique <[email protected]> --- v3: * Renamed the xstat from tx_dma_err_pkts to tx_dma_err_cmpl, and reworded the release note from "per-queue" to "port-level" to match what bnxt_dev_xstats_get_op() actually returns. Stephen Hemminger pointed out both: the xstat counts completions, not packets (a coalesced completion covers several packets, so "_pkts" overstates granularity), and the value summed across queues is exposed as one port-wide xstat, not one per queue. * Added the DMA-error check to the NEON vector Tx-completion handler (bnxt_handle_tx_cp_vec() in bnxt_rxtx_vec_neon.c), matching the AVX2/SSE handlers already covered -- also per Stephen Hemminger. * Switched the increment from rte_atomic_fetch_add_explicit() to a relaxed load + rte_atomic_store_explicit(): tx_dma_err has a single writer (the lcore polling that queue's completions), so the stronger fetch_add (a locked read-modify-write on most architectures) was unneeded on this per-packet path. * Changed bnxt_stats_reset_op()'s tx_dma_err reset from a direct `= 0` assignment (v2) to rte_atomic_store_explicit(..., 0, ...), matching the atomic store now used for the increment and for the other reset site (bnxt_dev_xstats_reset_op()). v2: * Added a release notes entry documenting the new xstat, per reviewer request. * NOTE: apply this patch before "net/bnxt: add support for queue size of 16384" -- both add a bullet under the same "Updated bnxt driver" release-notes heading, and the latter's context assumes this one's bullet is already present. doc/guides/rel_notes/release_26_11.rst | 6 ++++++ drivers/net/bnxt/bnxt_rxtx_vec_avx2.c | 8 ++++++++ drivers/net/bnxt/bnxt_rxtx_vec_neon.c | 8 ++++++++ drivers/net/bnxt/bnxt_rxtx_vec_sse.c | 8 ++++++++ drivers/net/bnxt/bnxt_stats.c | 26 ++++++++++++++++++++++++++ drivers/net/bnxt/bnxt_stats.h | 3 +++ drivers/net/bnxt/bnxt_txq.h | 1 + drivers/net/bnxt/bnxt_txr.c | 8 ++++++++ 8 files changed, 68 insertions(+) diff --git a/doc/guides/rel_notes/release_26_11.rst b/doc/guides/rel_notes/release_26_11.rst index dec96ccbc7..7ab289adf1 100644 --- a/doc/guides/rel_notes/release_26_11.rst +++ b/doc/guides/rel_notes/release_26_11.rst @@ -74,6 +74,12 @@ New Features ``xdp_meta_rx_ts_valid_mask``. * Added ``read_clock`` operation to query the PTP hardware clock. +* **Updated bnxt driver.** + + * Added a ``tx_dma_err_cmpl`` xstat to report Tx completions that the + device flagged with a DMA error. This is a port-level counter, and + is also folded into the standard ``oerrors`` counter. + * **Updated Intel iavf driver.** * Runtime Rx/Tx queue setup is now automatically disabled diff --git a/drivers/net/bnxt/bnxt_rxtx_vec_avx2.c b/drivers/net/bnxt/bnxt_rxtx_vec_avx2.c index 50b3602839..b22bb16fa0 100644 --- a/drivers/net/bnxt/bnxt_rxtx_vec_avx2.c +++ b/drivers/net/bnxt/bnxt_rxtx_vec_avx2.c @@ -743,6 +743,14 @@ bnxt_handle_tx_cp_vec(struct bnxt_tx_queue *txq) if (!bnxt_cpr_cmp_valid(txcmp, raw_cons, ring_mask + 1)) break; + uint16_t errors_v = rte_le_to_cpu_16(txcmp->errors_v); + + if (unlikely(errors_v & TX_CMPL_ERRORS_DMA_ERROR)) + rte_atomic_store_explicit(&txq->tx_dma_err, + rte_atomic_load_explicit(&txq->tx_dma_err, + rte_memory_order_relaxed) + 1, + rte_memory_order_relaxed); + nb_tx_pkts += txcmp->opaque; raw_cons = NEXT_RAW_CMP(raw_cons); } while (nb_tx_pkts < ring_mask); diff --git a/drivers/net/bnxt/bnxt_rxtx_vec_neon.c b/drivers/net/bnxt/bnxt_rxtx_vec_neon.c index 03f39280e5..086ba43363 100644 --- a/drivers/net/bnxt/bnxt_rxtx_vec_neon.c +++ b/drivers/net/bnxt/bnxt_rxtx_vec_neon.c @@ -355,6 +355,14 @@ bnxt_handle_tx_cp_vec(struct bnxt_tx_queue *txq) if (!bnxt_cpr_cmp_valid(txcmp, raw_cons, ring_mask + 1)) break; + uint16_t errors_v = rte_le_to_cpu_16(txcmp->errors_v); + + if (unlikely(errors_v & TX_CMPL_ERRORS_DMA_ERROR)) + rte_atomic_store_explicit(&txq->tx_dma_err, + rte_atomic_load_explicit(&txq->tx_dma_err, + rte_memory_order_relaxed) + 1, + rte_memory_order_relaxed); + if (likely(CMP_TYPE(txcmp) == TX_CMPL_TYPE_TX_L2)) nb_tx_pkts += txcmp->opaque; else diff --git a/drivers/net/bnxt/bnxt_rxtx_vec_sse.c b/drivers/net/bnxt/bnxt_rxtx_vec_sse.c index 7d455b6f56..4024a80b51 100644 --- a/drivers/net/bnxt/bnxt_rxtx_vec_sse.c +++ b/drivers/net/bnxt/bnxt_rxtx_vec_sse.c @@ -577,6 +577,14 @@ bnxt_handle_tx_cp_vec(struct bnxt_tx_queue *txq) if (!bnxt_cpr_cmp_valid(txcmp, raw_cons, ring_mask + 1)) break; + uint16_t errors_v = rte_le_to_cpu_16(txcmp->errors_v); + + if (unlikely(errors_v & TX_CMPL_ERRORS_DMA_ERROR)) + rte_atomic_store_explicit(&txq->tx_dma_err, + rte_atomic_load_explicit(&txq->tx_dma_err, + rte_memory_order_relaxed) + 1, + rte_memory_order_relaxed); + if (likely(CMP_TYPE(txcmp) == TX_CMPL_TYPE_TX_L2)) nb_tx_pkts += txcmp->opaque; else diff --git a/drivers/net/bnxt/bnxt_stats.c b/drivers/net/bnxt/bnxt_stats.c index 37b33f0505..8f7c867ff5 100644 --- a/drivers/net/bnxt/bnxt_stats.c +++ b/drivers/net/bnxt/bnxt_stats.c @@ -685,6 +685,8 @@ static int bnxt_stats_get_ext(struct rte_eth_dev *eth_dev, bnxt_stats->oerrors += rte_atomic_load_explicit(&txq->tx_mbuf_drop, rte_memory_order_relaxed); + bnxt_stats->oerrors += rte_atomic_load_explicit(&txq->tx_dma_err, + rte_memory_order_relaxed); if (!txq->tx_started) continue; @@ -758,6 +760,9 @@ int bnxt_stats_get_op(struct rte_eth_dev *eth_dev, bnxt_stats->oerrors += rte_atomic_load_explicit(&txq->tx_mbuf_drop, rte_memory_order_relaxed); + bnxt_stats->oerrors += + rte_atomic_load_explicit(&txq->tx_dma_err, + rte_memory_order_relaxed); } return rc; @@ -808,6 +813,8 @@ int bnxt_stats_reset_op(struct rte_eth_dev *eth_dev) struct bnxt_tx_queue *txq = bp->tx_queues[i]; txq->tx_mbuf_drop = 0; + rte_atomic_store_explicit(&txq->tx_dma_err, 0, + rte_memory_order_relaxed); } bnxt_clear_prev_stat(bp); @@ -911,6 +918,7 @@ int bnxt_dev_xstats_get_op(struct rte_eth_dev *eth_dev, RTE_DIM(bnxt_tx_stats_strings) + sz + RTE_DIM(bnxt_rx_ext_stats_strings) + RTE_DIM(bnxt_tx_ext_stats_strings) + + BNXT_NUM_SW_XSTATS + bnxt_flow_stats_cnt(bp); if (n < stat_count || xstats == NULL) @@ -1033,6 +1041,14 @@ int bnxt_dev_xstats_get_op(struct rte_eth_dev *eth_dev, count++; } + xstats[count].id = count; + xstats[count].value = 0; + for (i = 0; i < bp->tx_cp_nr_rings; i++) + xstats[count].value += + rte_atomic_load_explicit(&bp->tx_queues[i]->tx_dma_err, + rte_memory_order_relaxed); + count++; + if (bp->fw_cap & BNXT_FW_CAP_ADV_FLOW_COUNTERS && bp->fw_cap & BNXT_FW_CAP_ADV_FLOW_MGMT && BNXT_FLOW_XSTATS_EN(bp)) { @@ -1112,6 +1128,7 @@ int bnxt_dev_xstats_get_names_op(struct rte_eth_dev *eth_dev, sz + RTE_DIM(bnxt_rx_ext_stats_strings) + RTE_DIM(bnxt_tx_ext_stats_strings) + + BNXT_NUM_SW_XSTATS + bnxt_flow_stats_cnt(bp); if (xstats_names == NULL || size < stat_cnt) @@ -1165,6 +1182,10 @@ int bnxt_dev_xstats_get_names_op(struct rte_eth_dev *eth_dev, count++; } + strlcpy(xstats_names[count].name, "tx_dma_err_cmpl", + sizeof(xstats_names[count].name)); + count++; + if (bp->fw_cap & BNXT_FW_CAP_ADV_FLOW_COUNTERS && bp->fw_cap & BNXT_FW_CAP_ADV_FLOW_MGMT && BNXT_FLOW_XSTATS_EN(bp)) { @@ -1190,6 +1211,7 @@ int bnxt_dev_xstats_get_names_op(struct rte_eth_dev *eth_dev, int bnxt_dev_xstats_reset_op(struct rte_eth_dev *eth_dev) { struct bnxt *bp = eth_dev->data->dev_private; + unsigned int i; int ret; ret = is_bnxt_in_error(bp); @@ -1202,6 +1224,10 @@ int bnxt_dev_xstats_reset_op(struct rte_eth_dev *eth_dev) return -ENOTSUP; } + for (i = 0; i < bp->tx_cp_nr_rings; i++) + rte_atomic_store_explicit(&bp->tx_queues[i]->tx_dma_err, 0, + rte_memory_order_relaxed); + ret = bnxt_hwrm_port_clr_stats(bp); if (ret != 0) PMD_DRV_LOG_LINE(ERR, "Failed to reset xstats: %s", diff --git a/drivers/net/bnxt/bnxt_stats.h b/drivers/net/bnxt/bnxt_stats.h index c0508e773a..534b9af5e8 100644 --- a/drivers/net/bnxt/bnxt_stats.h +++ b/drivers/net/bnxt/bnxt_stats.h @@ -8,6 +8,9 @@ #include <ethdev_driver.h> +/* Number of software (non-HWRM) xstats appended after the FW-reported ones. */ +#define BNXT_NUM_SW_XSTATS 1 + void bnxt_free_stats(struct bnxt *bp); int bnxt_stats_get_op(struct rte_eth_dev *eth_dev, struct rte_eth_stats *bnxt_stats, struct eth_queue_stats *qstats); diff --git a/drivers/net/bnxt/bnxt_txq.h b/drivers/net/bnxt/bnxt_txq.h index ac8af91c57..525f841789 100644 --- a/drivers/net/bnxt/bnxt_txq.h +++ b/drivers/net/bnxt/bnxt_txq.h @@ -36,6 +36,7 @@ struct bnxt_tx_queue { struct rte_mbuf **free; uint64_t offloads; RTE_ATOMIC(uint64_t) tx_mbuf_drop; + RTE_ATOMIC(uint64_t) tx_dma_err; }; void bnxt_free_txq_stats(struct bnxt_tx_queue *txq); diff --git a/drivers/net/bnxt/bnxt_txr.c b/drivers/net/bnxt/bnxt_txr.c index 36188346f1..64ff42b38e 100644 --- a/drivers/net/bnxt/bnxt_txr.c +++ b/drivers/net/bnxt/bnxt_txr.c @@ -782,6 +782,14 @@ static int bnxt_handle_tx_cp(struct bnxt_tx_queue *txq) if (!bnxt_cpr_cmp_valid(txcmp, raw_cons, ring_mask + 1)) break; + uint16_t errors_v = rte_le_to_cpu_16(txcmp->errors_v); + + if (unlikely(errors_v & TX_CMPL_ERRORS_DMA_ERROR)) + rte_atomic_store_explicit(&txq->tx_dma_err, + rte_atomic_load_explicit(&txq->tx_dma_err, + rte_memory_order_relaxed) + 1, + rte_memory_order_relaxed); + if (CMP_TYPE(txcmp) == CMPL_BASE_TYPE_TX_L2_COAL) { struct tx_cmpl_coal *txcmp_c = (struct tx_cmpl_coal *)txcmp; -- 2.47.3

