The HPASS audio subsystem on the Qualcomm Nord SoC has three QDSP6
instances, ADSP0/1/2. They are not independent: ADSP0 owns the resources
the other two need (PLLs, AG_NOC, RSCp, CESTA, THROTTLE, QTMR) and
initializes them as part of its own boot, so ADSP1/ADSP2 can only cold
boot once ADSP0 is up. Their firmware also cannot tolerate one instance
going down on its own -- stopping or restarting a single member leaves
sysmon/glink_ssr notify requests to the siblings timing out, and the
whole cluster unable to recover without manually stopping all three and
restarting them in dependency order.

This series models that constraint. A new "qcom,cluster-root" phandle
groups the instances; every member carries it, the resource-owning
instance included, whose phandle references itself, so both cluster
membership and root-ness are answerable from the node being probed alone.
Instances without the property -- which is every PAS instance on every
other SoC -- keep qcom_pas.cluster NULL and are completely unaffected.

Two invariants fall out of the hardware, and are implemented
independently of each other:

 - Root-first bring-up. A dependent member's start blocks until the root
   has booted.

 - Coupled, two-phase teardown. Stopping or crashing any member takes
   the whole cluster down. Hardware testing showed that whichever
   member's PAS/TZ power-off happens chronologically last always fails
   its own graceful shutdown ack -- both the sysmon SSCTL QMI ack and
   the SMEM-bit/IRQ ack -- regardless of whether that member is the
   root. So phase 1 fans the stop out to every participant concurrently
   and waits at a barrier for all of them to finish requesting and
   awaiting their own ack while the cluster is still fully powered; only
   past the barrier does phase 2 let each member actually power its
   hardware off. That way no member is ever asking firmware to ack a
   shutdown after a sibling's hardware is already gone. On a crash the
   cluster is then auto-restarted root-first; an explicit stop leaves
   the ordered restart to higher-level policy.

The remaining wrinkle is that a cluster member must not send SSR notify
to its siblings, whose firmware cannot handle it, while still notifying
non-cluster peers such as mpss and cdsp. Both notify paths therefore
suppress receiver-side on a cluster id: sysmon (in remoteproc) and
glink_ssr (in rpmsg). Since RPMSG_QCOM_GLINK does not and should not
depend on REMOTEPROC, the id is a plain opaque "void *" cookie on struct
rproc, plumbed down into the glink edge rather than looked up through
struct rproc from the rpmsg side.

Shawn Guo (8):
  dt-bindings: remoteproc: qcom,nord-pas: Add qcom,cluster-root property
  remoteproc: Add cluster field to struct rproc
  remoteproc: qcom: sysmon: Suppress notify between cluster siblings
  rpmsg: qcom_glink_ssr: Suppress cleanup notify between cluster
    siblings
  remoteproc: qcom: pas: Track HPASS ADSP cluster membership
  remoteproc: qcom: pas: Sequence HPASS ADSP cluster boot root-first
  remoteproc: qcom: pas: Enforce coupled stop/crash for HPASS ADSP
    clusters
  remoteproc: qcom: pas: Add Nord ADSP1/2 support

 .../bindings/remoteproc/qcom,nord-pas.yaml    |  17 +
 drivers/remoteproc/qcom_common.c              |   7 +-
 drivers/remoteproc/qcom_common.h              |   2 +-
 drivers/remoteproc/qcom_q6v5_pas.c            | 524 ++++++++++++++++++
 drivers/remoteproc/qcom_sysmon.c              |  18 +-
 drivers/rpmsg/qcom_glink_native.c             |  20 +-
 drivers/rpmsg/qcom_glink_native.h             |   5 +-
 drivers/rpmsg/qcom_glink_rpm.c                |   2 +-
 drivers/rpmsg/qcom_glink_smem.c               |   5 +-
 drivers/rpmsg/qcom_glink_ssr.c                |  36 +-
 include/linux/remoteproc.h                    |   3 +
 include/linux/rpmsg/qcom_glink.h              |  10 +-
 12 files changed, 627 insertions(+), 22 deletions(-)

-- 
2.43.0


Reply via email to