AMD General

Hi

> -----Original Message-----
> From: Jiang, Haochen <[email protected]>
> Sent: Monday, July 6, 2026 9:01 AM
> To: Kumar, Venkataramanan <[email protected]>; gcc-
> [email protected]
> Cc: Liu, Hongtao <[email protected]>; [email protected]; Sharma,
> Dipesh <[email protected]>
> Subject: RE: [PATCH 2/7] Add tmm as fake register
>
> Caution: This message originated from an External Source. Use proper caution
> when opening attachments, clicking links, or responding.
>
>
> > From: Kumar, Venkataramanan <[email protected]>
> > Sent: Monday, July 6, 2026 1:54 AM
> >
> > > From: Haochen Jiang <[email protected]>
> > > Sent: Thursday, July 2, 2026 12:19 PM
> > > To: [email protected]
> > > Cc: [email protected]; [email protected]; Sharma, Dipesh
> > > <[email protected]>; Kumar, Venkataramanan
> > > <[email protected]>
> > > Subject: [PATCH 2/7] Add tmm as fake register
> > >
> > > Caution: This message originated from an External Source. Use proper
> > caution
> > > when opening attachments, clicking links, or responding.
> > >
> > >
> > > In legacy AMX, we use inline assembly for intrins since we are not
> > > doing tile register allocation. For ACE, we will try a different way.
> > >
> > > Although we are not doing tile register allocation for now since
> > > there is no convenient way for register spill, which is a must for
> > > register allocation, we
> > are
> > > going to make it convenient for future implementation for potential
> > > register allocation. Thus, we will introduce a fake tmm register.
> >
> > I understand deferring full tile RA for now.  The tile data can be
> > spilled via repeated TILEMOVROW/TILEMOVCOL extract to ZMM and stores
> > to an aligned stack slot (reverse on reload), or via
> > TILELOADD/TILESTORED on the AMX-TILE path.  Please mention in the
> > commit message to avoid implying spill is fundamentally not possible.
>
> The spill is exactly what you said. But I did not get what do you mean by " 
> avoid
> implying spill is fundamentally not possible ". Spill is doable but just not
> convenient. Since we should never mis-use legacy AMX and ACE, the only way
> to do the spill is using tilemovrow/tilemovcol for 16 times, which might even
> cause spill on zmms.
>
> If there is a strong request for RA, that is the way we could do but user 
> still
> need to be careful on register usage, which is the same as now for sers to do
> the register allocation on their own. The spill is too costly.

ACE specification mentions typed tile-pointer intrinsics (e.g. 
_tile_zero(__tile1024i *tdst)); LLVM supports them, and GCC should too for 
spec/user compatibility. This can be added later.

>
> > > diff --git a/gcc/config/i386/i386.cc b/gcc/config/i386/i386.cc index
> > > e66958db7ac..9cdd0138104 100644
> > > --- a/gcc/config/i386/i386.cc
> > > +++ b/gcc/config/i386/i386.cc
> > > @@ -177,6 +177,8 @@ enum reg_class const
> > > regclass_map[FIRST_PSEUDO_REGISTER] =
> > >    GENERAL_REGS, GENERAL_REGS, GENERAL_REGS, GENERAL_REGS,
> > >    GENERAL_REGS, GENERAL_REGS, GENERAL_REGS, GENERAL_REGS,
> > >    GENERAL_REGS, GENERAL_REGS, GENERAL_REGS, GENERAL_REGS,
> >
> > Why is only one fake register although we have 8 tile registers ??
> >
>
> Several reasons:
> 1. It is an easier implementation method to cooperate with using reg number in
> intrin.
> 2. Under the current situation of ACE, we do not need to complicate the user
> case. The user case it to split the matrix into several part, do outer product
> simultaneously, then accumulate. So, all the tmms will be used at the same
> time. No need to distinguish them at this phase for now.

OK.

>
> > >  /* ADJUST_REG_ALLOC_ORDER is a macro which permits reg_alloc_order
> > >     to be rearranged based on a particular function.  When using sse math,
> > > @@ -2107,7 +2111,8 @@ do {                                                
> > >       \
> > >   "xmm28", "xmm29", "xmm30", "xmm31",                                   \
> > >  "k0", "k1", "k2", "k3", "k4", "k5", "k6", "k7",                       \
> > >  "r16", "r17", "r18", "r19", "r20", "r21", "r22", "r23",               \
> > > - "r24", "r25", "r26", "r27", "r28", "r29", "r30", "r31" }
> > > + "r24", "r25", "r26", "r27", "r28", "r29", "r30", "r31",               \
> >
> > Can we add "tmm" here ??
> >
>
> It is never used so we did not add it. But I suppose the answer is yes for
> clearance.

Thanks, the patch look good.

Regards,
Venkat.

>
> Thx,
> Haochen

Reply via email to