LLVM 24.0.0git
AMDGPURegisterBankInfo.cpp
Go to the documentation of this file.
1//===- AMDGPURegisterBankInfo.cpp -------------------------------*- C++ -*-==//
2//
3// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
4// See https://llvm.org/LICENSE.txt for license information.
5// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
6//
7//===----------------------------------------------------------------------===//
8/// \file
9/// This file implements the targeting of the RegisterBankInfo class for
10/// AMDGPU.
11///
12/// \par
13///
14/// AMDGPU has unique register bank constraints that require special high level
15/// strategies to deal with. There are two main true physical register banks
16/// VGPR (vector), and SGPR (scalar). Additionally the VCC register bank is a
17/// sort of pseudo-register bank needed to represent SGPRs used in a vector
18/// boolean context. There is also the AGPR bank, which is a special purpose
19/// physical register bank present on some subtargets.
20///
21/// Copying from VGPR to SGPR is generally illegal, unless the value is known to
22/// be uniform. It is generally not valid to legalize operands by inserting
23/// copies as on other targets. Operations which require uniform, SGPR operands
24/// generally require scalarization by repeatedly executing the instruction,
25/// activating each set of lanes using a unique set of input values. This is
26/// referred to as a waterfall loop.
27///
28/// \par Booleans
29///
30/// Booleans (s1 values) requires special consideration. A vector compare result
31/// is naturally a bitmask with one bit per lane, in a 32 or 64-bit
32/// register. These are represented with the VCC bank. During selection, we need
33/// to be able to unambiguously go back from a register class to a register
34/// bank. To distinguish whether an SGPR should use the SGPR or VCC register
35/// bank, we need to know the use context type. An SGPR s1 value always means a
36/// VCC bank value, otherwise it will be the SGPR bank. A scalar compare sets
37/// SCC, which is a 1-bit unaddressable register. This will need to be copied to
38/// a 32-bit virtual register. Taken together, this means we need to adjust the
39/// type of boolean operations to be regbank legal. All SALU booleans need to be
40/// widened to 32-bits, and all VALU booleans need to be s1 values.
41///
42/// A noteworthy exception to the s1-means-vcc rule is for legalization artifact
43/// casts. G_TRUNC s1 results, and G_SEXT/G_ZEXT/G_ANYEXT sources are never vcc
44/// bank. A non-boolean source (such as a truncate from a 1-bit load from
45/// memory) will require a copy to the VCC bank which will require clearing the
46/// high bits and inserting a compare.
47///
48/// \par Constant bus restriction
49///
50/// VALU instructions have a limitation known as the constant bus
51/// restriction. Most VALU instructions can use SGPR operands, but may read at
52/// most 1 SGPR or constant literal value (this to 2 in gfx10 for most
53/// instructions). This is one unique SGPR, so the same SGPR may be used for
54/// multiple operands. From a register bank perspective, any combination of
55/// operands should be legal as an SGPR, but this is contextually dependent on
56/// the SGPR operands all being the same register. There is therefore optimal to
57/// choose the SGPR with the most uses to minimize the number of copies.
58///
59/// We avoid trying to solve this problem in RegBankSelect. Any VALU G_*
60/// operation should have its source operands all mapped to VGPRs (except for
61/// VCC), inserting copies from any SGPR operands. This the most trivial legal
62/// mapping. Anything beyond the simplest 1:1 instruction selection would be too
63/// complicated to solve here. Every optimization pattern or instruction
64/// selected to multiple outputs would have to enforce this rule, and there
65/// would be additional complexity in tracking this rule for every G_*
66/// operation. By forcing all inputs to VGPRs, it also simplifies the task of
67/// picking the optimal operand combination from a post-isel optimization pass.
68///
69//===----------------------------------------------------------------------===//
70
72
73#include "AMDGPU.h"
75#include "AMDGPUInstrInfo.h"
76#include "AMDGPULaneMaskUtils.h"
77#include "GCNSubtarget.h"
79#include "SIRegisterInfo.h"
85#include "llvm/IR/IntrinsicsAMDGPU.h"
86
87#define GET_TARGET_REGBANK_IMPL
88#include "AMDGPUGenRegisterBank.inc"
89
90// This file will be TableGen'ed at some point.
91#include "AMDGPUGenRegisterBankInfo.def"
92
93using namespace llvm;
94using namespace MIPatternMatch;
95
96namespace {
97
98// Observer to apply a register bank to new registers created by LegalizerHelper.
99class ApplyRegBankMapping final : public GISelChangeObserver {
100private:
102 const AMDGPURegisterBankInfo &RBI;
104 const RegisterBank *NewBank;
106
107public:
108 ApplyRegBankMapping(MachineIRBuilder &B, const AMDGPURegisterBankInfo &RBI_,
109 MachineRegisterInfo &MRI_, const RegisterBank *RB)
110 : B(B), RBI(RBI_), MRI(MRI_), NewBank(RB) {
111 assert(!B.isObservingChanges());
112 B.setChangeObserver(*this);
113 }
114
115 ~ApplyRegBankMapping() override {
116 for (MachineInstr *MI : NewInsts)
117 applyBank(*MI);
118
119 B.stopObservingChanges();
120 }
121
122 /// Set any registers that don't have a set register class or bank to SALU.
123 void applyBank(MachineInstr &MI) {
124 const unsigned Opc = MI.getOpcode();
125 if (Opc == AMDGPU::G_ANYEXT || Opc == AMDGPU::G_ZEXT ||
126 Opc == AMDGPU::G_SEXT) {
127 // LegalizerHelper wants to use the basic legalization artifacts when
128 // widening etc. We don't handle selection with vcc in artifact sources,
129 // so we need to use a select instead to handle these properly.
130 Register DstReg = MI.getOperand(0).getReg();
131 Register SrcReg = MI.getOperand(1).getReg();
132 const RegisterBank *SrcBank = RBI.getRegBank(SrcReg, MRI, *RBI.TRI);
133 if (SrcBank == &AMDGPU::VCCRegBank) {
134 const LLT S32 = LLT::scalar(32);
135 assert(MRI.getType(SrcReg) == LLT::scalar(1));
136 assert(MRI.getType(DstReg) == S32);
137 assert(NewBank == &AMDGPU::VGPRRegBank);
138
139 // Replace the extension with a select, which really uses the boolean
140 // source.
141 B.setInsertPt(*MI.getParent(), MI);
142
143 auto True = B.buildConstant(S32, Opc == AMDGPU::G_SEXT ? -1 : 1);
144 auto False = B.buildConstant(S32, 0);
145 B.buildSelect(DstReg, SrcReg, True, False);
146 MRI.setRegBank(True.getReg(0), *NewBank);
147 MRI.setRegBank(False.getReg(0), *NewBank);
148 MI.eraseFromParent();
149 }
150
151 assert(!MRI.getRegClassOrRegBank(DstReg));
152 MRI.setRegBank(DstReg, *NewBank);
153 return;
154 }
155
156#ifndef NDEBUG
157 if (Opc == AMDGPU::G_TRUNC) {
158 Register DstReg = MI.getOperand(0).getReg();
159 const RegisterBank *DstBank = RBI.getRegBank(DstReg, MRI, *RBI.TRI);
160 assert(DstBank != &AMDGPU::VCCRegBank);
161 }
162#endif
163
164 for (MachineOperand &Op : MI.operands()) {
165 if (!Op.isReg())
166 continue;
167
168 // We may see physical registers if building a real MI
169 Register Reg = Op.getReg();
170 if (Reg.isPhysical() || MRI.getRegClassOrRegBank(Reg))
171 continue;
172
173 const RegisterBank *RB = NewBank;
174 if (MRI.getType(Reg) == LLT::scalar(1)) {
175 assert(NewBank == &AMDGPU::VGPRRegBank &&
176 "s1 operands should only be used for vector bools");
177 assert((MI.getOpcode() != AMDGPU::G_TRUNC &&
178 MI.getOpcode() != AMDGPU::G_ANYEXT) &&
179 "not expecting legalization artifacts here");
180 RB = &AMDGPU::VCCRegBank;
181 }
182
183 MRI.setRegBank(Reg, *RB);
184 }
185 }
186
187 void erasingInstr(MachineInstr &MI) override {}
188
189 void createdInstr(MachineInstr &MI) override {
190 // At this point, the instruction was just inserted and has no operands.
191 NewInsts.push_back(&MI);
192 }
193
194 void changingInstr(MachineInstr &MI) override {}
195 void changedInstr(MachineInstr &MI) override {
196 // FIXME: In principle we should probably add the instruction to NewInsts,
197 // but the way the LegalizerHelper uses the observer, we will always see the
198 // registers we need to set the regbank on also referenced in a new
199 // instruction.
200 }
201};
202
203} // anonymous namespace
204
206 : Subtarget(ST), TRI(Subtarget.getRegisterInfo()),
207 TII(Subtarget.getInstrInfo()) {
208
209 // HACK: Until this is fully tablegen'd.
210 static llvm::once_flag InitializeRegisterBankFlag;
211
212 static auto InitializeRegisterBankOnce = [this]() {
213 assert(&getRegBank(AMDGPU::SGPRRegBankID) == &AMDGPU::SGPRRegBank &&
214 &getRegBank(AMDGPU::VGPRRegBankID) == &AMDGPU::VGPRRegBank &&
215 &getRegBank(AMDGPU::AGPRRegBankID) == &AMDGPU::AGPRRegBank);
216 (void)this;
217 };
218
219 llvm::call_once(InitializeRegisterBankFlag, InitializeRegisterBankOnce);
220}
221
222static bool isVectorRegisterBank(const RegisterBank &Bank) {
223 unsigned BankID = Bank.getID();
224 return BankID == AMDGPU::VGPRRegBankID || BankID == AMDGPU::AGPRRegBankID;
225}
226
228 return RB != &AMDGPU::SGPRRegBank;
229}
230
232 const RegisterBank &Src,
233 TypeSize Size) const {
234 // TODO: Should there be a UniformVGPRRegBank which can use readfirstlane?
235 if (Dst.getID() == AMDGPU::SGPRRegBankID &&
236 (isVectorRegisterBank(Src) || Src.getID() == AMDGPU::VCCRegBankID)) {
237 return std::numeric_limits<unsigned>::max();
238 }
239
240 // Bool values are tricky, because the meaning is based on context. The SCC
241 // and VCC banks are for the natural scalar and vector conditions produced by
242 // a compare.
243 //
244 // Legalization doesn't know about the necessary context, so an s1 use may
245 // have been a truncate from an arbitrary value, in which case a copy (lowered
246 // as a compare with 0) needs to be inserted.
247 if (Size == 1 &&
248 (Dst.getID() == AMDGPU::SGPRRegBankID) &&
249 (isVectorRegisterBank(Src) ||
250 Src.getID() == AMDGPU::SGPRRegBankID ||
251 Src.getID() == AMDGPU::VCCRegBankID))
252 return std::numeric_limits<unsigned>::max();
253
254 // There is no direct copy between AGPRs.
255 if (Dst.getID() == AMDGPU::AGPRRegBankID &&
256 Src.getID() == AMDGPU::AGPRRegBankID)
257 return 4;
258
259 return RegisterBankInfo::copyCost(Dst, Src, Size);
260}
261
263 const ValueMapping &ValMapping,
264 const RegisterBank *CurBank) const {
265 // Check if this is a breakdown for G_LOAD to move the pointer from SGPR to
266 // VGPR.
267 // FIXME: Is there a better way to do this?
268 if (ValMapping.NumBreakDowns >= 2 || ValMapping.BreakDown[0].Length >= 64)
269 return 10; // This is expensive.
270
271 assert(ValMapping.NumBreakDowns == 2 &&
272 ValMapping.BreakDown[0].Length == 32 &&
273 ValMapping.BreakDown[0].StartIdx == 0 &&
274 ValMapping.BreakDown[1].Length == 32 &&
275 ValMapping.BreakDown[1].StartIdx == 32 &&
276 ValMapping.BreakDown[0].RegBank == ValMapping.BreakDown[1].RegBank);
277
278 // 32-bit extract of a 64-bit value is just access of a subregister, so free.
279 // TODO: Cost of 0 hits assert, though it's not clear it's what we really
280 // want.
281
282 // TODO: 32-bit insert to a 64-bit SGPR may incur a non-free copy due to SGPR
283 // alignment restrictions, but this probably isn't important.
284 return 1;
285}
286
287const RegisterBank &
289 LLT Ty) const {
290 // We promote real scalar booleans to SReg_32. Any SGPR using s1 is really a
291 // VCC-like use.
292 if (TRI->isSGPRClass(&RC)) {
293 // FIXME: This probably came from a copy from a physical register, which
294 // should be inferable from the copied to-type. We don't have many boolean
295 // physical register constraints so just assume a normal SGPR for now.
296 if (!Ty.isValid())
297 return AMDGPU::SGPRRegBank;
298
299 return Ty == LLT::scalar(1) ? AMDGPU::VCCRegBank : AMDGPU::SGPRRegBank;
300 }
301
302 return TRI->isAGPRClass(&RC) ? AMDGPU::AGPRRegBank : AMDGPU::VGPRRegBank;
303}
304
305template <unsigned NumOps>
308 const MachineInstr &MI, const MachineRegisterInfo &MRI,
309 const std::array<unsigned, NumOps> RegSrcOpIdx,
311
312 InstructionMappings AltMappings;
313
314 SmallVector<const ValueMapping *, 10> Operands(MI.getNumOperands());
315
316 unsigned Sizes[NumOps];
317 for (unsigned I = 0; I < NumOps; ++I) {
318 Register Reg = MI.getOperand(RegSrcOpIdx[I]).getReg();
319 Sizes[I] = getSizeInBits(Reg, MRI, *TRI);
320 }
321
322 for (unsigned I = 0, E = MI.getNumExplicitDefs(); I != E; ++I) {
323 unsigned SizeI = getSizeInBits(MI.getOperand(I).getReg(), MRI, *TRI);
324 Operands[I] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SizeI);
325 }
326
327 // getInstrMapping's default mapping uses ID 1, so start at 2.
328 unsigned MappingID = 2;
329 for (const auto &Entry : Table) {
330 for (unsigned I = 0; I < NumOps; ++I) {
331 int OpIdx = RegSrcOpIdx[I];
332 Operands[OpIdx] = AMDGPU::getValueMapping(Entry.RegBanks[I], Sizes[I]);
333 }
334
335 AltMappings.push_back(&getInstructionMapping(MappingID++, Entry.Cost,
336 getOperandsMapping(Operands),
337 Operands.size()));
338 }
339
340 return AltMappings;
341}
342
345 const MachineInstr &MI, const MachineRegisterInfo &MRI) const {
347 case Intrinsic::amdgcn_readlane: {
348 static const OpRegBankEntry<3> Table[2] = {
349 // Perfectly legal.
350 { { AMDGPU::SGPRRegBankID, AMDGPU::VGPRRegBankID, AMDGPU::SGPRRegBankID }, 1 },
351
352 // Need a readfirstlane for the index.
353 { { AMDGPU::SGPRRegBankID, AMDGPU::VGPRRegBankID, AMDGPU::VGPRRegBankID }, 2 }
354 };
355
356 const std::array<unsigned, 3> RegSrcOpIdx = { { 0, 2, 3 } };
357 return addMappingFromTable<3>(MI, MRI, RegSrcOpIdx, Table);
358 }
359 case Intrinsic::amdgcn_writelane: {
360 static const OpRegBankEntry<4> Table[4] = {
361 // Perfectly legal.
362 { { AMDGPU::VGPRRegBankID, AMDGPU::SGPRRegBankID, AMDGPU::SGPRRegBankID, AMDGPU::VGPRRegBankID }, 1 },
363
364 // Need readfirstlane of first op
365 { { AMDGPU::VGPRRegBankID, AMDGPU::VGPRRegBankID, AMDGPU::SGPRRegBankID, AMDGPU::VGPRRegBankID }, 2 },
366
367 // Need readfirstlane of second op
368 { { AMDGPU::VGPRRegBankID, AMDGPU::SGPRRegBankID, AMDGPU::VGPRRegBankID, AMDGPU::VGPRRegBankID }, 2 },
369
370 // Need readfirstlane of both ops
371 { { AMDGPU::VGPRRegBankID, AMDGPU::VGPRRegBankID, AMDGPU::VGPRRegBankID, AMDGPU::VGPRRegBankID }, 3 }
372 };
373
374 // rsrc, voffset, offset
375 const std::array<unsigned, 4> RegSrcOpIdx = { { 0, 2, 3, 4 } };
376 return addMappingFromTable<4>(MI, MRI, RegSrcOpIdx, Table);
377 }
378 default:
380 }
381}
382
385 const MachineInstr &MI, const MachineRegisterInfo &MRI) const {
386
388 case Intrinsic::amdgcn_s_buffer_load:
389 case Intrinsic::amdgcn_ptr_s_buffer_load: {
390 static const OpRegBankEntry<2> Table[4] = {
391 // Perfectly legal.
392 { { AMDGPU::SGPRRegBankID, AMDGPU::SGPRRegBankID }, 1 },
393
394 // Only need 1 register in loop
395 { { AMDGPU::SGPRRegBankID, AMDGPU::VGPRRegBankID }, 300 },
396
397 // Have to waterfall the resource.
398 { { AMDGPU::VGPRRegBankID, AMDGPU::SGPRRegBankID }, 1000 },
399
400 // Have to waterfall the resource, and the offset.
401 { { AMDGPU::VGPRRegBankID, AMDGPU::VGPRRegBankID }, 1500 }
402 };
403
404 // rsrc, offset
405 const std::array<unsigned, 2> RegSrcOpIdx = { { 2, 3 } };
406 return addMappingFromTable<2>(MI, MRI, RegSrcOpIdx, Table);
407 }
408 case Intrinsic::amdgcn_ds_ordered_add:
409 case Intrinsic::amdgcn_ds_ordered_swap: {
410 // VGPR = M0, VGPR
411 static const OpRegBankEntry<3> Table[2] = {
412 // Perfectly legal.
413 { { AMDGPU::VGPRRegBankID, AMDGPU::SGPRRegBankID, AMDGPU::VGPRRegBankID }, 1 },
414
415 // Need a readfirstlane for m0
416 { { AMDGPU::VGPRRegBankID, AMDGPU::VGPRRegBankID, AMDGPU::VGPRRegBankID }, 2 }
417 };
418
419 const std::array<unsigned, 3> RegSrcOpIdx = { { 0, 2, 3 } };
420 return addMappingFromTable<3>(MI, MRI, RegSrcOpIdx, Table);
421 }
422 case Intrinsic::amdgcn_s_sendmsg:
423 case Intrinsic::amdgcn_s_sendmsghalt: {
424 // FIXME: Should have no register for immediate
425 static const OpRegBankEntry<1> Table[2] = {
426 // Perfectly legal.
427 { { AMDGPU::SGPRRegBankID }, 1 },
428
429 // Need readlane
430 { { AMDGPU::VGPRRegBankID }, 3 }
431 };
432
433 const std::array<unsigned, 1> RegSrcOpIdx = { { 2 } };
434 return addMappingFromTable<1>(MI, MRI, RegSrcOpIdx, Table);
435 }
436 default:
438 }
439}
440
441// FIXME: Returns uniform if there's no source value information. This is
442// probably wrong.
444 if (!MI.hasOneMemOperand())
445 return false;
446
447 const MachineMemOperand *MMO = *MI.memoperands_begin();
448 const unsigned AS = MMO->getAddrSpace();
449 const bool IsConst = AS == AMDGPUAS::CONSTANT_ADDRESS ||
451 const unsigned MemSize = 8 * MMO->getSize().getValue();
452
453 // Require 4-byte alignment.
454 return (MMO->getAlign() >= Align(4) ||
455 (Subtarget.hasScalarSubwordLoads() &&
456 ((MemSize == 16 && MMO->getAlign() >= Align(2)) ||
457 (MemSize == 8 && MMO->getAlign() >= Align(1))))) &&
458 // Can't do a scalar atomic load.
459 !MMO->isAtomic() &&
460 // Don't use scalar loads for volatile accesses to non-constant address
461 // spaces.
462 (IsConst || !MMO->isVolatile()) &&
463 // Memory must be known constant, or not written before this load.
464 (IsConst || MMO->isInvariant() || (MMO->getFlags() & MONoClobber)) &&
466}
467
470 const MachineInstr &MI) const {
471
472 const MachineFunction &MF = *MI.getMF();
473 const MachineRegisterInfo &MRI = MF.getRegInfo();
474
475
476 InstructionMappings AltMappings;
477 switch (MI.getOpcode()) {
478 case TargetOpcode::G_CONSTANT:
479 case TargetOpcode::G_IMPLICIT_DEF: {
480 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
481 if (Size == 1) {
482 static const OpRegBankEntry<1> Table[3] = {
483 { { AMDGPU::VGPRRegBankID }, 1 },
484 { { AMDGPU::SGPRRegBankID }, 1 },
485 { { AMDGPU::VCCRegBankID }, 1 }
486 };
487
488 return addMappingFromTable<1>(MI, MRI, {{ 0 }}, Table);
489 }
490
491 [[fallthrough]];
492 }
493 case TargetOpcode::G_FCONSTANT:
494 case TargetOpcode::G_FRAME_INDEX:
495 case TargetOpcode::G_GLOBAL_VALUE: {
496 static const OpRegBankEntry<1> Table[2] = {
497 { { AMDGPU::VGPRRegBankID }, 1 },
498 { { AMDGPU::SGPRRegBankID }, 1 }
499 };
500
501 return addMappingFromTable<1>(MI, MRI, {{ 0 }}, Table);
502 }
503 case TargetOpcode::G_AND:
504 case TargetOpcode::G_OR:
505 case TargetOpcode::G_XOR: {
506 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
507
508 if (Size == 1) {
509 // s_{and|or|xor}_b32 set scc when the result of the 32-bit op is not 0.
510 const InstructionMapping &SCCMapping = getInstructionMapping(
511 1, 1, getOperandsMapping(
512 {AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 32),
513 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 32),
514 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 32)}),
515 3); // Num Operands
516 AltMappings.push_back(&SCCMapping);
517
518 const InstructionMapping &VCCMapping0 = getInstructionMapping(
519 2, 1, getOperandsMapping(
520 {AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, Size),
521 AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, Size),
522 AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, Size)}),
523 3); // Num Operands
524 AltMappings.push_back(&VCCMapping0);
525 return AltMappings;
526 }
527
528 if (Size != 64)
529 break;
530
531 const InstructionMapping &SSMapping = getInstructionMapping(
532 1, 1, getOperandsMapping(
533 {AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size),
534 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size),
535 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size)}),
536 3); // Num Operands
537 AltMappings.push_back(&SSMapping);
538
539 const InstructionMapping &VVMapping = getInstructionMapping(
540 2, 2, getOperandsMapping(
541 {AMDGPU::getValueMappingSGPR64Only(AMDGPU::VGPRRegBankID, Size),
542 AMDGPU::getValueMappingSGPR64Only(AMDGPU::VGPRRegBankID, Size),
543 AMDGPU::getValueMappingSGPR64Only(AMDGPU::VGPRRegBankID, Size)}),
544 3); // Num Operands
545 AltMappings.push_back(&VVMapping);
546 break;
547 }
548 case TargetOpcode::G_LOAD:
549 case TargetOpcode::G_ZEXTLOAD:
550 case TargetOpcode::G_SEXTLOAD: {
551 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
552 LLT PtrTy = MRI.getType(MI.getOperand(1).getReg());
553 unsigned PtrSize = PtrTy.getSizeInBits();
554 unsigned AS = PtrTy.getAddressSpace();
555
559 const InstructionMapping &SSMapping = getInstructionMapping(
560 1, 1, getOperandsMapping(
561 {AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size),
562 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, PtrSize)}),
563 2); // Num Operands
564 AltMappings.push_back(&SSMapping);
565 }
566
567 const InstructionMapping &VVMapping = getInstructionMapping(
568 2, 1,
570 {AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size),
571 AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, PtrSize)}),
572 2); // Num Operands
573 AltMappings.push_back(&VVMapping);
574
575 // It may be possible to have a vgpr = load sgpr mapping here, because
576 // the mubuf instructions support this kind of load, but probably for only
577 // gfx7 and older. However, the addressing mode matching in the instruction
578 // selector should be able to do a better job of detecting and selecting
579 // these kinds of loads from the vgpr = load vgpr mapping.
580
581 return AltMappings;
582
583 }
584 case TargetOpcode::G_SELECT: {
585 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
586 const InstructionMapping &SSMapping = getInstructionMapping(1, 1,
587 getOperandsMapping({AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size),
588 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 1),
589 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size),
590 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size)}),
591 4); // Num Operands
592 AltMappings.push_back(&SSMapping);
593
594 const InstructionMapping &VVMapping = getInstructionMapping(2, 1,
595 getOperandsMapping({AMDGPU::getValueMappingSGPR64Only(AMDGPU::VGPRRegBankID, Size),
596 AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1),
597 AMDGPU::getValueMappingSGPR64Only(AMDGPU::VGPRRegBankID, Size),
598 AMDGPU::getValueMappingSGPR64Only(AMDGPU::VGPRRegBankID, Size)}),
599 4); // Num Operands
600 AltMappings.push_back(&VVMapping);
601
602 return AltMappings;
603 }
604 case TargetOpcode::G_UADDE:
605 case TargetOpcode::G_USUBE:
606 case TargetOpcode::G_SADDE:
607 case TargetOpcode::G_SSUBE: {
608 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
609 const InstructionMapping &SSMapping = getInstructionMapping(1, 1,
611 {AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size),
612 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 1),
613 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size),
614 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size),
615 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 1)}),
616 5); // Num Operands
617 AltMappings.push_back(&SSMapping);
618
619 const InstructionMapping &VVMapping = getInstructionMapping(2, 1,
620 getOperandsMapping({AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size),
621 AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1),
622 AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size),
623 AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size),
624 AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1)}),
625 5); // Num Operands
626 AltMappings.push_back(&VVMapping);
627 return AltMappings;
628 }
629 case AMDGPU::G_BRCOND: {
630 assert(MRI.getType(MI.getOperand(0).getReg()).getSizeInBits() == 1);
631
632 // TODO: Change type to 32 for scalar
634 1, 1, getOperandsMapping(
635 {AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 1), nullptr}),
636 2); // Num Operands
637 AltMappings.push_back(&SMapping);
638
640 1, 1, getOperandsMapping(
641 {AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1), nullptr }),
642 2); // Num Operands
643 AltMappings.push_back(&VMapping);
644 return AltMappings;
645 }
646 case AMDGPU::G_INTRINSIC:
647 case AMDGPU::G_INTRINSIC_CONVERGENT:
649 case AMDGPU::G_INTRINSIC_W_SIDE_EFFECTS:
650 case AMDGPU::G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS:
652 default:
653 break;
654 }
656}
657
661 LLT HalfTy,
662 Register Reg) const {
663 assert(HalfTy.getSizeInBits() == 32);
664 MachineRegisterInfo *MRI = B.getMRI();
665 Register LoLHS = MRI->createGenericVirtualRegister(HalfTy);
666 Register HiLHS = MRI->createGenericVirtualRegister(HalfTy);
667 const RegisterBank *Bank = getRegBank(Reg, *MRI, *TRI);
668 MRI->setRegBank(LoLHS, *Bank);
669 MRI->setRegBank(HiLHS, *Bank);
670
671 Regs.push_back(LoLHS);
672 Regs.push_back(HiLHS);
673
674 B.buildInstr(AMDGPU::G_UNMERGE_VALUES)
675 .addDef(LoLHS)
676 .addDef(HiLHS)
677 .addUse(Reg);
678}
679
680/// Replace the current type each register in \p Regs has with \p NewTy
682 LLT NewTy) {
683 for (Register Reg : Regs) {
684 assert(MRI.getType(Reg).getSizeInBits() == NewTy.getSizeInBits());
685 MRI.setType(Reg, NewTy);
686 }
687}
688
690 if (Ty.isVector()) {
691 assert(Ty.getElementCount().isKnownMultipleOf(2));
692 return LLT::scalarOrVector(Ty.getElementCount().divideCoefficientBy(2),
693 Ty.getElementType());
694 }
695
696 assert(Ty.getScalarSizeInBits() % 2 == 0);
697 return LLT::scalar(Ty.getScalarSizeInBits() / 2);
698}
699
700// Build one or more V_READFIRSTLANE_B32 instructions to move the given vector
701// source value into a scalar register.
704 Register Src) const {
705 LLT Ty = MRI.getType(Src);
706 const RegisterBank *Bank = getRegBank(Src, MRI, *TRI);
707
708 if (Bank == &AMDGPU::SGPRRegBank)
709 return Src;
710
711 unsigned Bits = Ty.getSizeInBits();
712 assert(Bits % 32 == 0);
713
714 if (Bank != &AMDGPU::VGPRRegBank) {
715 // We need to copy from AGPR to VGPR
716 Src = B.buildCopy(Ty, Src).getReg(0);
717 MRI.setRegBank(Src, AMDGPU::VGPRRegBank);
718 }
719
720 LLT S32 = LLT::scalar(32);
721 unsigned NumParts = Bits / 32;
724
725 if (Bits == 32) {
726 SrcParts.push_back(Src);
727 } else {
728 auto Unmerge = B.buildUnmerge(S32, Src);
729 for (unsigned i = 0; i < NumParts; ++i)
730 SrcParts.push_back(Unmerge.getReg(i));
731 }
732
733 for (unsigned i = 0; i < NumParts; ++i) {
734 Register SrcPart = SrcParts[i];
735 Register DstPart = MRI.createVirtualRegister(&AMDGPU::SReg_32_XM0RegClass);
736 MRI.setType(DstPart, NumParts == 1 ? Ty : S32);
737
738 const TargetRegisterClass *Constrained =
739 constrainGenericRegister(SrcPart, AMDGPU::VGPR_32RegClass, MRI);
740 (void)Constrained;
741 assert(Constrained && "Failed to constrain readfirstlane src reg");
742
743 B.buildInstr(AMDGPU::V_READFIRSTLANE_B32, {DstPart}, {SrcPart});
744
745 DstParts.push_back(DstPart);
746 }
747
748 if (Bits == 32)
749 return DstParts[0];
750
751 Register Dst = B.buildMergeLikeInstr(Ty, DstParts).getReg(0);
752 MRI.setRegBank(Dst, AMDGPU::SGPRRegBank);
753 return Dst;
754}
755
756/// Legalize instruction \p MI where operands in \p OpIndices must be SGPRs. If
757/// any of the required SGPR operands are VGPRs, perform a waterfall loop to
758/// execute the instruction for each unique combination of values in all lanes
759/// in the wave. The block will be split such that rest of the instructions are
760/// moved to a new block.
761///
762/// Essentially performs this loop:
763//
764/// Save Execution Mask
765/// For (Lane : Wavefront) {
766/// Enable Lane, Disable all other lanes
767/// SGPR = read SGPR value for current lane from VGPR
768/// VGPRResult[Lane] = use_op SGPR
769/// }
770/// Restore Execution Mask
771///
772/// There is additional complexity to try for compare values to identify the
773/// unique values used.
776 SmallSet<Register, 4> &SGPROperandRegs) const {
777 // Track use registers which have already been expanded with a readfirstlane
778 // sequence. This may have multiple uses if moving a sequence.
779 DenseMap<Register, Register> WaterfalledRegMap;
780
781 MachineBasicBlock &MBB = B.getMBB();
782 MachineFunction *MF = &B.getMF();
783
784 const TargetRegisterClass *WaveRC = TRI->getWaveMaskRegClass();
785 const AMDGPU::LaneMaskConstants &LMC =
787
788#ifndef NDEBUG
789 const int OrigRangeSize = std::distance(Range.begin(), Range.end());
790#endif
791
792 MachineRegisterInfo &MRI = *B.getMRI();
793 Register SaveExecReg = MRI.createVirtualRegister(WaveRC);
794 Register InitSaveExecReg = MRI.createVirtualRegister(WaveRC);
795
796 // Don't bother using generic instructions/registers for the exec mask.
797 B.buildInstr(TargetOpcode::IMPLICIT_DEF)
798 .addDef(InitSaveExecReg);
799
800 Register PhiExec = MRI.createVirtualRegister(WaveRC);
801 Register NewExec = MRI.createVirtualRegister(WaveRC);
802
803 // To insert the loop we need to split the block. Move everything before this
804 // point to a new block, and insert a new empty block before this instruction.
807 MachineBasicBlock *RemainderBB = MF->CreateMachineBasicBlock();
808 MachineBasicBlock *RestoreExecBB = MF->CreateMachineBasicBlock();
810 ++MBBI;
811 MF->insert(MBBI, LoopBB);
812 MF->insert(MBBI, BodyBB);
813 MF->insert(MBBI, RestoreExecBB);
814 MF->insert(MBBI, RemainderBB);
815
816 LoopBB->addSuccessor(BodyBB);
817 BodyBB->addSuccessor(RestoreExecBB);
818 BodyBB->addSuccessor(LoopBB);
819
820 // Move the rest of the block into a new block.
822 RemainderBB->splice(RemainderBB->begin(), &MBB, Range.end(), MBB.end());
823
824 MBB.addSuccessor(LoopBB);
825 RestoreExecBB->addSuccessor(RemainderBB);
826
827 B.setInsertPt(*LoopBB, LoopBB->end());
828
829 B.buildInstr(TargetOpcode::PHI)
830 .addDef(PhiExec)
831 .addReg(InitSaveExecReg)
832 .addMBB(&MBB)
833 .addReg(NewExec)
834 .addMBB(BodyBB);
835
836 const DebugLoc &DL = B.getDL();
837
838 MachineInstr &FirstInst = *Range.begin();
839
840 // Move the instruction into the loop body. Note we moved everything after
841 // Range.end() already into a new block, so Range.end() is no longer valid.
842 BodyBB->splice(BodyBB->end(), &MBB, Range.begin(), MBB.end());
843
844 // Figure out the iterator range after splicing the instructions.
845 MachineBasicBlock::iterator NewBegin = FirstInst.getIterator();
846 auto NewEnd = BodyBB->end();
847
848 B.setMBB(*LoopBB);
849
850 LLT S1 = LLT::scalar(1);
851 Register CondReg;
852
853 assert(std::distance(NewBegin, NewEnd) == OrigRangeSize);
854
855 for (MachineInstr &MI : make_range(NewBegin, NewEnd)) {
856 for (MachineOperand &Op : MI.all_uses()) {
857 Register OldReg = Op.getReg();
858 if (!SGPROperandRegs.count(OldReg))
859 continue;
860
861 // See if we already processed this register in another instruction in the
862 // sequence.
863 auto OldVal = WaterfalledRegMap.find(OldReg);
864 if (OldVal != WaterfalledRegMap.end()) {
865 Op.setReg(OldVal->second);
866 continue;
867 }
868
869 Register OpReg = Op.getReg();
870 LLT OpTy = MRI.getType(OpReg);
871
872 const RegisterBank *OpBank = getRegBank(OpReg, MRI, *TRI);
873 if (OpBank != &AMDGPU::VGPRRegBank) {
874 // Insert copy from AGPR to VGPR before the loop.
875 B.setMBB(MBB);
876 OpReg = B.buildCopy(OpTy, OpReg).getReg(0);
877 MRI.setRegBank(OpReg, AMDGPU::VGPRRegBank);
878 B.setMBB(*LoopBB);
879 }
880
881 Register CurrentLaneReg = buildReadFirstLane(B, MRI, OpReg);
882
883 // Build the comparison(s).
884 unsigned OpSize = OpTy.getSizeInBits();
885 bool Is64 = OpSize % 64 == 0;
886 unsigned PartSize = Is64 ? 64 : 32;
887 LLT PartTy = LLT::scalar(PartSize);
888 unsigned NumParts = OpSize / PartSize;
890 SmallVector<Register, 8> CurrentLaneParts;
891
892 if (NumParts == 1) {
893 OpParts.push_back(OpReg);
894 CurrentLaneParts.push_back(CurrentLaneReg);
895 } else {
896 auto UnmergeOp = B.buildUnmerge(PartTy, OpReg);
897 auto UnmergeCurrentLane = B.buildUnmerge(PartTy, CurrentLaneReg);
898 for (unsigned i = 0; i < NumParts; ++i) {
899 OpParts.push_back(UnmergeOp.getReg(i));
900 CurrentLaneParts.push_back(UnmergeCurrentLane.getReg(i));
901 MRI.setRegBank(OpParts[i], AMDGPU::VGPRRegBank);
902 MRI.setRegBank(CurrentLaneParts[i], AMDGPU::SGPRRegBank);
903 }
904 }
905
906 for (unsigned i = 0; i < NumParts; ++i) {
907 auto CmpReg = B.buildICmp(CmpInst::ICMP_EQ, S1, CurrentLaneParts[i],
908 OpParts[i]).getReg(0);
909 MRI.setRegBank(CmpReg, AMDGPU::VCCRegBank);
910
911 if (!CondReg) {
912 CondReg = CmpReg;
913 } else {
914 CondReg = B.buildAnd(S1, CondReg, CmpReg).getReg(0);
915 MRI.setRegBank(CondReg, AMDGPU::VCCRegBank);
916 }
917 }
918
919 Op.setReg(CurrentLaneReg);
920
921 // Make sure we don't re-process this register again.
922 WaterfalledRegMap.insert(std::pair(OldReg, Op.getReg()));
923 }
924 }
925
926 // The ballot becomes a no-op during instruction selection.
927 CondReg = B.buildIntrinsic(Intrinsic::amdgcn_ballot,
928 {LLT::scalar(Subtarget.isWave32() ? 32 : 64)})
929 .addReg(CondReg)
930 .getReg(0);
931 MRI.setRegClass(CondReg, WaveRC);
932
933 // Update EXEC, save the original EXEC value to VCC.
934 B.buildInstr(LMC.AndSaveExecOpc)
935 .addDef(NewExec)
936 .addReg(CondReg, RegState::Kill);
937
938 MRI.setSimpleHint(NewExec, CondReg);
939
940 B.setInsertPt(*BodyBB, BodyBB->end());
941
942 // Update EXEC, switch all done bits to 0 and all todo bits to 1.
943 B.buildInstr(LMC.XorTermOpc)
944 .addDef(LMC.ExecReg)
945 .addReg(LMC.ExecReg)
946 .addReg(NewExec);
947
948 // XXX - s_xor_b64 sets scc to 1 if the result is nonzero, so can we use
949 // s_cbranch_scc0?
950
951 // Loop back to V_READFIRSTLANE_B32 if there are still variants to cover.
952 B.buildInstr(AMDGPU::SI_WATERFALL_LOOP).addMBB(LoopBB);
953
954 // Save the EXEC mask before the loop.
955 BuildMI(MBB, MBB.end(), DL, TII->get(LMC.MovOpc), SaveExecReg)
956 .addReg(LMC.ExecReg);
957
958 // Restore the EXEC mask after the loop.
959 B.setMBB(*RestoreExecBB);
960 B.buildInstr(LMC.MovTermOpc).addDef(LMC.ExecReg).addReg(SaveExecReg);
961
962 // Set the insert point after the original instruction, so any new
963 // instructions will be in the remainder.
964 B.setInsertPt(*RemainderBB, RemainderBB->begin());
965
966 return true;
967}
968
969// Return any unique registers used by \p MI at \p OpIndices that need to be
970// handled in a waterfall loop. Returns these registers in \p
971// SGPROperandRegs. Returns true if there are any operands to handle and a
972// waterfall loop is necessary.
974 SmallSet<Register, 4> &SGPROperandRegs, MachineInstr &MI,
975 MachineRegisterInfo &MRI, ArrayRef<unsigned> OpIndices) const {
976 for (unsigned Op : OpIndices) {
977 assert(MI.getOperand(Op).isUse());
978 Register Reg = MI.getOperand(Op).getReg();
979 const RegisterBank *OpBank = getRegBank(Reg, MRI, *TRI);
980 if (OpBank->getID() != AMDGPU::SGPRRegBankID)
981 SGPROperandRegs.insert(Reg);
982 }
983
984 // No operands need to be replaced, so no need to loop.
985 return !SGPROperandRegs.empty();
986}
987
990 // Use a set to avoid extra readfirstlanes in the case where multiple operands
991 // are the same register.
992 SmallSet<Register, 4> SGPROperandRegs;
993
994 if (!collectWaterfallOperands(SGPROperandRegs, MI, *B.getMRI(), OpIndices))
995 return false;
996
997 MachineBasicBlock::iterator I = MI.getIterator();
998 return executeInWaterfallLoop(B, make_range(I, std::next(I)),
999 SGPROperandRegs);
1000}
1001
1002// Legalize an operand that must be an SGPR by inserting a readfirstlane.
1004 MachineIRBuilder &B, MachineInstr &MI, unsigned OpIdx) const {
1005 Register Reg = MI.getOperand(OpIdx).getReg();
1006 MachineRegisterInfo &MRI = *B.getMRI();
1007 const RegisterBank *Bank = getRegBank(Reg, MRI, *TRI);
1008 if (Bank == &AMDGPU::SGPRRegBank)
1009 return;
1010
1011 Reg = buildReadFirstLane(B, MRI, Reg);
1012 MI.getOperand(OpIdx).setReg(Reg);
1013}
1014
1015/// Split \p Ty into 2 pieces. The first will have \p FirstSize bits, and the
1016/// rest will be in the remainder.
1017static std::pair<LLT, LLT> splitUnequalType(LLT Ty, unsigned FirstSize) {
1018 unsigned TotalSize = Ty.getSizeInBits();
1019 if (!Ty.isVector())
1020 return {LLT::scalar(FirstSize), LLT::scalar(TotalSize - FirstSize)};
1021
1022 LLT EltTy = Ty.getElementType();
1023 unsigned EltSize = EltTy.getSizeInBits();
1024 assert(FirstSize % EltSize == 0);
1025
1026 unsigned FirstPartNumElts = FirstSize / EltSize;
1027 unsigned RemainderElts = (TotalSize - FirstSize) / EltSize;
1028
1029 return {LLT::scalarOrVector(ElementCount::getFixed(FirstPartNumElts), EltTy),
1030 LLT::scalarOrVector(ElementCount::getFixed(RemainderElts), EltTy)};
1031}
1032
1034 if (!Ty.isVector())
1035 return LLT::scalar(128);
1036
1037 LLT EltTy = Ty.getElementType();
1038 assert(128 % EltTy.getSizeInBits() == 0);
1039 return LLT::fixed_vector(128 / EltTy.getSizeInBits(), EltTy);
1040}
1041
1045 MachineInstr &MI) const {
1046 MachineRegisterInfo &MRI = *B.getMRI();
1047 Register DstReg = MI.getOperand(0).getReg();
1048 const LLT LoadTy = MRI.getType(DstReg);
1049 unsigned LoadSize = LoadTy.getSizeInBits();
1050 MachineMemOperand *MMO = *MI.memoperands_begin();
1051 const unsigned MaxNonSmrdLoadSize = 128;
1052
1053 const RegisterBank *DstBank =
1054 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
1055 if (DstBank == &AMDGPU::SGPRRegBank) {
1056 // There are some special cases that we need to look at for 32 bit and 96
1057 // bit SGPR loads otherwise we have nothing to do.
1058 if (LoadSize != 32 && (LoadSize != 96 || Subtarget.hasScalarDwordx3Loads()))
1059 return false;
1060
1061 const unsigned MemSize = 8 * MMO->getSize().getValue();
1062 // Scalar loads of size 8 or 16 bit with proper alignment may be widened to
1063 // 32 bit. Check to see if we need to widen the memory access, 8 or 16 bit
1064 // scalar loads should have a load size of 32 but memory access size of less
1065 // than 32.
1066 if (LoadSize == 32 &&
1067 (MemSize == 32 || LoadTy.isVector() || !isScalarLoadLegal(MI)))
1068 return false;
1069
1070 if (LoadSize == 32 &&
1071 ((MemSize == 8 && MMO->getAlign() >= Align(1)) ||
1072 (MemSize == 16 && MMO->getAlign() >= Align(2))) &&
1074 Subtarget.getGeneration() >= AMDGPUSubtarget::GFX12)
1075 return false;
1076
1077 Register PtrReg = MI.getOperand(1).getReg();
1078
1079 ApplyRegBankMapping ApplyBank(B, *this, MRI, DstBank);
1080
1081 if (LoadSize == 32) {
1082 // This is an extending load from a sub-dword size. Widen the memory
1083 // access size to 4 bytes and clear the extra high bits appropriately
1084 const LLT S32 = LLT::scalar(32);
1085 if (MI.getOpcode() == AMDGPU::G_SEXTLOAD) {
1086 // Must extend the sign bit into higher bits for a G_SEXTLOAD
1087 auto WideLoad = B.buildLoadFromOffset(S32, PtrReg, *MMO, 0);
1088 B.buildSExtInReg(MI.getOperand(0), WideLoad, MemSize);
1089 } else if (MI.getOpcode() == AMDGPU::G_ZEXTLOAD) {
1090 // Must extend zero into higher bits with an AND for a G_ZEXTLOAD
1091 auto WideLoad = B.buildLoadFromOffset(S32, PtrReg, *MMO, 0);
1092 B.buildZExtInReg(MI.getOperand(0), WideLoad, MemSize);
1093 } else
1094 // We do not need to touch the higher bits for regular loads.
1095 B.buildLoadFromOffset(MI.getOperand(0), PtrReg, *MMO, 0);
1096 } else {
1097 // 96-bit loads are only available for vector loads. We need to split this
1098 // into a 64-bit part, and 32 (unless we can widen to a 128-bit load).
1099 if (MMO->getAlign() < Align(16)) {
1100 LegalizerHelper Helper(B.getMF(), ApplyBank, B);
1101 LLT Part64, Part32;
1102 std::tie(Part64, Part32) = splitUnequalType(LoadTy, 64);
1103 if (Helper.reduceLoadStoreWidth(cast<GAnyLoad>(MI), 0, Part64) !=
1105 return false;
1106 return true;
1107 }
1108 LLT WiderTy = widen96To128(LoadTy);
1109 auto WideLoad = B.buildLoadFromOffset(WiderTy, PtrReg, *MMO, 0);
1110 if (WiderTy.isScalar()) {
1111 B.buildTrunc(MI.getOperand(0), WideLoad);
1112 } else {
1113 B.buildDeleteTrailingVectorElements(MI.getOperand(0).getReg(),
1114 WideLoad);
1115 }
1116 }
1117
1118 MI.eraseFromParent();
1119 return true;
1120 }
1121
1122 // 128-bit loads are supported for all instruction types.
1123 if (LoadSize <= MaxNonSmrdLoadSize)
1124 return false;
1125
1126 SmallVector<Register, 1> SrcRegs(OpdMapper.getVRegs(1));
1127
1128 if (SrcRegs.empty())
1129 SrcRegs.push_back(MI.getOperand(1).getReg());
1130
1131 // RegBankSelect only emits scalar types, so we need to reset the pointer
1132 // operand to a pointer type.
1133 Register BasePtrReg = SrcRegs[0];
1134 LLT PtrTy = MRI.getType(MI.getOperand(1).getReg());
1135 MRI.setType(BasePtrReg, PtrTy);
1136
1137 // The following are the loads not splitted enough during legalization
1138 // because it was not clear they are smem-load or vmem-load
1141 assert(LoadSize % MaxNonSmrdLoadSize == 0);
1142 unsigned NumSplitParts = LoadTy.getSizeInBits() / MaxNonSmrdLoadSize;
1143 const LLT LoadSplitTy = LoadTy.divide(NumSplitParts);
1144 ApplyRegBankMapping O(B, *this, MRI, &AMDGPU::VGPRRegBank);
1145 LegalizerHelper Helper(B.getMF(), O, B);
1146 if (LoadTy.isVector()) {
1147 if (Helper.fewerElementsVector(MI, 0, LoadSplitTy) !=
1149 return false;
1150 } else {
1151 if (Helper.narrowScalar(MI, 0, LoadSplitTy) != LegalizerHelper::Legalized)
1152 return false;
1153 }
1154 }
1155
1156 MRI.setRegBank(DstReg, AMDGPU::VGPRRegBank);
1157 return true;
1158}
1159
1163 MachineInstr &MI) const {
1164 MachineRegisterInfo &MRI = *B.getMRI();
1165 const MachineFunction &MF = B.getMF();
1166 const GCNSubtarget &ST = MF.getSubtarget<GCNSubtarget>();
1167 const auto &TFI = *ST.getFrameLowering();
1168
1169 // Guard in case the stack growth direction ever changes with scratch
1170 // instructions.
1171 assert(TFI.getStackGrowthDirection() == TargetFrameLowering::StackGrowsUp &&
1172 "Stack grows upwards for AMDGPU");
1173
1174 Register Dst = MI.getOperand(0).getReg();
1175 Register AllocSize = MI.getOperand(1).getReg();
1176 Align Alignment = assumeAligned(MI.getOperand(2).getImm());
1177
1178 // When using flat-scratch, the stack offset is unscaled.
1179 const bool HasFlatScratch = ST.hasFlatScratchEnabled();
1180 const unsigned WavefrontSizeLog2 = ST.getWavefrontSizeLog2();
1181
1182 const RegisterBank *SizeBank = getRegBank(AllocSize, MRI, *TRI);
1183
1184 if (SizeBank != &AMDGPU::SGPRRegBank) {
1185 auto WaveReduction =
1186 B.buildIntrinsic(Intrinsic::amdgcn_wave_reduce_umax, {LLT::scalar(32)})
1187 .addUse(AllocSize)
1188 .addImm(0);
1189 AllocSize = WaveReduction.getReg(0);
1190 }
1191
1192 LLT PtrTy = MRI.getType(Dst);
1194
1196 Register SPReg = Info->getStackPtrOffsetReg();
1197 ApplyRegBankMapping ApplyBank(B, *this, MRI, &AMDGPU::SGPRRegBank);
1198
1199 Register ScaledSize = AllocSize;
1200 if (!HasFlatScratch) {
1201 auto WaveSize = B.buildConstant(LLT::scalar(32), WavefrontSizeLog2);
1202 ScaledSize = B.buildShl(IntPtrTy, AllocSize, WaveSize).getReg(0);
1203 }
1204
1205 auto OldSP = B.buildCopy(PtrTy, SPReg);
1206 if (Alignment > TFI.getStackAlign()) {
1207 const uint64_t ScaledAlignment =
1208 HasFlatScratch ? Alignment.value()
1209 : (Alignment.value() << WavefrontSizeLog2);
1210 const uint64_t StackAlignMask = ScaledAlignment - 1;
1211 auto Tmp1 = B.buildPtrAdd(PtrTy, OldSP,
1212 B.buildConstant(LLT::scalar(32), StackAlignMask));
1213 B.buildMaskLowPtrBits(Dst, Tmp1,
1214 (HasFlatScratch
1215 ? Log2(Alignment)
1216 : Log2(Alignment) + WavefrontSizeLog2));
1217 } else {
1218 B.buildCopy(Dst, OldSP);
1219 }
1220 auto PtrAdd = B.buildPtrAdd(PtrTy, Dst, ScaledSize);
1221 B.buildCopy(SPReg, PtrAdd);
1222 MI.eraseFromParent();
1223 return true;
1224}
1225
1229 int RsrcIdx) const {
1230 const int NumDefs = MI.getNumExplicitDefs();
1231
1232 // The reported argument index is relative to the IR intrinsic call arguments,
1233 // so we need to shift by the number of defs and the intrinsic ID.
1234 RsrcIdx += NumDefs + 1;
1235
1236 // Insert copies to VGPR arguments.
1237 applyDefaultMapping(OpdMapper);
1238
1239 // Fixup any SGPR arguments.
1240 SmallVector<unsigned, 4> SGPRIndexes;
1241 for (int I = NumDefs, NumOps = MI.getNumOperands(); I != NumOps; ++I) {
1242 if (!MI.getOperand(I).isReg())
1243 continue;
1244
1245 // If this intrinsic has a sampler, it immediately follows rsrc.
1246 if (I == RsrcIdx || I == RsrcIdx + 1)
1247 SGPRIndexes.push_back(I);
1248 }
1249
1250 executeInWaterfallLoop(B, MI, SGPRIndexes);
1251 return true;
1252}
1253
1254// Analyze a combined offset from an llvm.amdgcn.s.buffer intrinsic and store
1255// the three offsets (voffset, soffset and instoffset)
1257 MachineIRBuilder &B, Register CombinedOffset, Register &VOffsetReg,
1258 Register &SOffsetReg, int64_t &InstOffsetVal, Align Alignment) const {
1259 const LLT S32 = LLT::scalar(32);
1260 MachineRegisterInfo *MRI = B.getMRI();
1261
1262 if (std::optional<int64_t> Imm =
1263 getIConstantVRegSExtVal(CombinedOffset, *MRI)) {
1264 uint32_t SOffset, ImmOffset;
1265 if (TII->splitMUBUFOffset(*Imm, SOffset, ImmOffset, Alignment)) {
1266 VOffsetReg = B.buildConstant(S32, 0).getReg(0);
1267 SOffsetReg = B.buildConstant(S32, SOffset).getReg(0);
1268 InstOffsetVal = ImmOffset;
1269
1270 B.getMRI()->setRegBank(VOffsetReg, AMDGPU::VGPRRegBank);
1271 B.getMRI()->setRegBank(SOffsetReg, AMDGPU::SGPRRegBank);
1272 return SOffset + ImmOffset;
1273 }
1274 }
1275
1276 const bool CheckNUW = Subtarget.hasGFX1250Insts();
1277 Register Base;
1278 unsigned Offset;
1279
1280 std::tie(Base, Offset) =
1281 AMDGPU::getBaseWithConstantOffset(*MRI, CombinedOffset,
1282 /*KnownBits=*/nullptr,
1283 /*CheckNUW=*/CheckNUW);
1284
1285 uint32_t SOffset, ImmOffset;
1286 if (static_cast<int32_t>(Offset) > 0 &&
1287 TII->splitMUBUFOffset(Offset, SOffset, ImmOffset, Alignment)) {
1288 if (getRegBank(Base, *MRI, *TRI) == &AMDGPU::VGPRRegBank) {
1289 VOffsetReg = Base;
1290 SOffsetReg = B.buildConstant(S32, SOffset).getReg(0);
1291 B.getMRI()->setRegBank(SOffsetReg, AMDGPU::SGPRRegBank);
1292 InstOffsetVal = ImmOffset;
1293 return 0; // XXX - Why is this 0?
1294 }
1295
1296 // If we have SGPR base, we can use it for soffset.
1297 if (SOffset == 0) {
1298 VOffsetReg = B.buildConstant(S32, 0).getReg(0);
1299 B.getMRI()->setRegBank(VOffsetReg, AMDGPU::VGPRRegBank);
1300 SOffsetReg = Base;
1301 InstOffsetVal = ImmOffset;
1302 return 0; // XXX - Why is this 0?
1303 }
1304 }
1305
1306 // Handle the variable sgpr + vgpr case.
1307 MachineInstr *Add = getOpcodeDef(AMDGPU::G_ADD, CombinedOffset, *MRI);
1308 if (Add && static_cast<int32_t>(Offset) >= 0 &&
1309 (!CheckNUW || Add->getFlag(MachineInstr::NoUWrap))) {
1310 Register Src0 = getSrcRegIgnoringCopies(Add->getOperand(1).getReg(), *MRI);
1311 Register Src1 = getSrcRegIgnoringCopies(Add->getOperand(2).getReg(), *MRI);
1312
1313 const RegisterBank *Src0Bank = getRegBank(Src0, *MRI, *TRI);
1314 const RegisterBank *Src1Bank = getRegBank(Src1, *MRI, *TRI);
1315
1316 if (Src0Bank == &AMDGPU::VGPRRegBank && Src1Bank == &AMDGPU::SGPRRegBank) {
1317 VOffsetReg = Src0;
1318 SOffsetReg = Src1;
1319 return 0;
1320 }
1321
1322 if (Src0Bank == &AMDGPU::SGPRRegBank && Src1Bank == &AMDGPU::VGPRRegBank) {
1323 VOffsetReg = Src1;
1324 SOffsetReg = Src0;
1325 return 0;
1326 }
1327 }
1328
1329 // Ensure we have a VGPR for the combined offset. This could be an issue if we
1330 // have an SGPR offset and a VGPR resource.
1331 if (getRegBank(CombinedOffset, *MRI, *TRI) == &AMDGPU::VGPRRegBank) {
1332 VOffsetReg = CombinedOffset;
1333 } else {
1334 VOffsetReg = B.buildCopy(S32, CombinedOffset).getReg(0);
1335 B.getMRI()->setRegBank(VOffsetReg, AMDGPU::VGPRRegBank);
1336 }
1337
1338 SOffsetReg = B.buildConstant(S32, 0).getReg(0);
1339 B.getMRI()->setRegBank(SOffsetReg, AMDGPU::SGPRRegBank);
1340 return 0;
1341}
1342
1344 switch (Opc) {
1345 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD:
1346 return AMDGPU::G_AMDGPU_BUFFER_LOAD;
1347 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_UBYTE:
1348 return AMDGPU::G_AMDGPU_BUFFER_LOAD_UBYTE;
1349 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_SBYTE:
1350 return AMDGPU::G_AMDGPU_BUFFER_LOAD_SBYTE;
1351 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_USHORT:
1352 return AMDGPU::G_AMDGPU_BUFFER_LOAD_USHORT;
1353 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_SSHORT:
1354 return AMDGPU::G_AMDGPU_BUFFER_LOAD_SSHORT;
1355 default:
1356 break;
1357 }
1358 llvm_unreachable("Unexpected s_buffer_load opcode");
1359}
1360
1362 MachineIRBuilder &B, const OperandsMapper &OpdMapper) const {
1363 MachineInstr &MI = OpdMapper.getMI();
1364 MachineRegisterInfo &MRI = OpdMapper.getMRI();
1365
1366 const LLT S32 = LLT::scalar(32);
1367 Register Dst = MI.getOperand(0).getReg();
1368 LLT Ty = MRI.getType(Dst);
1369
1370 const RegisterBank *RSrcBank =
1371 OpdMapper.getInstrMapping().getOperandMapping(1).BreakDown[0].RegBank;
1372 const RegisterBank *OffsetBank =
1373 OpdMapper.getInstrMapping().getOperandMapping(2).BreakDown[0].RegBank;
1374 if (RSrcBank == &AMDGPU::SGPRRegBank &&
1375 OffsetBank == &AMDGPU::SGPRRegBank)
1376 return true; // Legal mapping
1377
1378 // FIXME: 96-bit case was widened during legalize. We need to narrow it back
1379 // here but don't have an MMO.
1380
1381 unsigned LoadSize = Ty.getSizeInBits();
1382 int NumLoads = 1;
1383 if (LoadSize == 256 || LoadSize == 512) {
1384 NumLoads = LoadSize / 128;
1385 Ty = Ty.divide(NumLoads);
1386 }
1387
1388 // Use the alignment to ensure that the required offsets will fit into the
1389 // immediate offsets.
1390 const Align Alignment = NumLoads > 1 ? Align(16 * NumLoads) : Align(1);
1391
1392 MachineFunction &MF = B.getMF();
1393
1394 Register SOffset;
1395 Register VOffset;
1396 int64_t ImmOffset = 0;
1397
1398 unsigned MMOOffset = setBufferOffsets(B, MI.getOperand(2).getReg(), VOffset,
1399 SOffset, ImmOffset, Alignment);
1400
1401 // TODO: 96-bit loads were widened to 128-bit results. Shrink the result if we
1402 // can, but we need to track an MMO for that.
1403 const unsigned MemSize = (Ty.getSizeInBits() + 7) / 8;
1404 const Align MemAlign(4); // FIXME: ABI type alignment?
1409 MemSize, MemAlign);
1410 if (MMOOffset != 0)
1411 BaseMMO = MF.getMachineMemOperand(BaseMMO, MMOOffset, MemSize);
1412
1413 // If only the offset is divergent, emit a MUBUF buffer load instead. We can
1414 // assume that the buffer is unswizzled.
1415
1416 Register RSrc = MI.getOperand(1).getReg();
1417 Register VIndex = B.buildConstant(S32, 0).getReg(0);
1418 B.getMRI()->setRegBank(VIndex, AMDGPU::VGPRRegBank);
1419 unsigned CachePolicy = MI.getOperand(3).getImm();
1420
1421 SmallVector<Register, 4> LoadParts(NumLoads);
1422
1423 MachineBasicBlock::iterator MII = MI.getIterator();
1424 MachineInstrSpan Span(MII, &B.getMBB());
1425
1426 for (int i = 0; i < NumLoads; ++i) {
1427 if (NumLoads == 1) {
1428 LoadParts[i] = Dst;
1429 } else {
1430 LoadParts[i] = MRI.createGenericVirtualRegister(Ty);
1431 MRI.setRegBank(LoadParts[i], AMDGPU::VGPRRegBank);
1432 }
1433
1434 if (i != 0)
1435 BaseMMO = MF.getMachineMemOperand(BaseMMO, 16, MemSize);
1436
1437 B.buildInstr(getSBufferLoadCorrespondingBufferLoadOpcode(MI.getOpcode()))
1438 .addDef(LoadParts[i]) // vdata
1439 .addUse(RSrc) // rsrc
1440 .addUse(VIndex) // vindex
1441 .addUse(VOffset) // voffset
1442 .addUse(SOffset) // soffset
1443 .addImm(ImmOffset + 16 * i) // offset(imm)
1444 .addImm(CachePolicy) // cachepolicy, swizzled buffer(imm)
1445 .addImm(0) // idxen(imm)
1446 .addMemOperand(BaseMMO);
1447 }
1448
1449 // TODO: If only the resource is a VGPR, it may be better to execute the
1450 // scalar load in the waterfall loop if the resource is expected to frequently
1451 // be dynamically uniform.
1452 if (RSrcBank != &AMDGPU::SGPRRegBank) {
1453 // Remove the original instruction to avoid potentially confusing the
1454 // waterfall loop logic.
1455 B.setInstr(*Span.begin());
1456 MI.eraseFromParent();
1457
1458 SmallSet<Register, 4> OpsToWaterfall;
1459
1460 OpsToWaterfall.insert(RSrc);
1461 executeInWaterfallLoop(B, make_range(Span.begin(), Span.end()),
1462 OpsToWaterfall);
1463 }
1464
1465 if (NumLoads != 1) {
1466 if (Ty.isVector())
1467 B.buildConcatVectors(Dst, LoadParts);
1468 else
1469 B.buildMergeLikeInstr(Dst, LoadParts);
1470 }
1471
1472 // We removed the instruction earlier with a waterfall loop.
1473 if (RSrcBank == &AMDGPU::SGPRRegBank)
1474 MI.eraseFromParent();
1475
1476 return true;
1477}
1478
1480 const OperandsMapper &OpdMapper,
1481 bool Signed) const {
1482 MachineInstr &MI = OpdMapper.getMI();
1483 MachineRegisterInfo &MRI = OpdMapper.getMRI();
1484
1485 // Insert basic copies
1486 applyDefaultMapping(OpdMapper);
1487
1488 Register DstReg = MI.getOperand(0).getReg();
1489 LLT Ty = MRI.getType(DstReg);
1490
1491 const LLT S32 = LLT::scalar(32);
1492
1493 unsigned FirstOpnd = isa<GIntrinsic>(MI) ? 2 : 1;
1494 Register SrcReg = MI.getOperand(FirstOpnd).getReg();
1495 Register OffsetReg = MI.getOperand(FirstOpnd + 1).getReg();
1496 Register WidthReg = MI.getOperand(FirstOpnd + 2).getReg();
1497
1498 const RegisterBank *DstBank =
1499 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
1500 if (DstBank == &AMDGPU::VGPRRegBank) {
1501 if (Ty == S32)
1502 return true;
1503
1504 // There is no 64-bit vgpr bitfield extract instructions so the operation
1505 // is expanded to a sequence of instructions that implement the operation.
1506 ApplyRegBankMapping ApplyBank(B, *this, MRI, &AMDGPU::VGPRRegBank);
1507
1508 const LLT S64 = LLT::scalar(64);
1509 // Shift the source operand so that extracted bits start at bit 0.
1510 auto ShiftOffset = Signed ? B.buildAShr(S64, SrcReg, OffsetReg)
1511 : B.buildLShr(S64, SrcReg, OffsetReg);
1512 auto UnmergeSOffset = B.buildUnmerge({S32, S32}, ShiftOffset);
1513
1514 // A 64-bit bitfield extract uses the 32-bit bitfield extract instructions
1515 // if the width is a constant.
1516 if (auto ConstWidth = getIConstantVRegValWithLookThrough(WidthReg, MRI)) {
1517 // Use the 32-bit bitfield extract instruction if the width is a constant.
1518 // Depending on the width size, use either the low or high 32-bits.
1519 auto Zero = B.buildConstant(S32, 0);
1520 auto WidthImm = ConstWidth->Value.getZExtValue();
1521 if (WidthImm <= 32) {
1522 // Use bitfield extract on the lower 32-bit source, and then sign-extend
1523 // or clear the upper 32-bits.
1524 auto Extract =
1525 Signed ? B.buildSbfx(S32, UnmergeSOffset.getReg(0), Zero, WidthReg)
1526 : B.buildUbfx(S32, UnmergeSOffset.getReg(0), Zero, WidthReg);
1527 auto Extend =
1528 Signed ? B.buildAShr(S32, Extract, B.buildConstant(S32, 31)) : Zero;
1529 B.buildMergeLikeInstr(DstReg, {Extract, Extend});
1530 } else {
1531 // Use bitfield extract on upper 32-bit source, and combine with lower
1532 // 32-bit source.
1533 auto UpperWidth = B.buildConstant(S32, WidthImm - 32);
1534 auto Extract =
1535 Signed
1536 ? B.buildSbfx(S32, UnmergeSOffset.getReg(1), Zero, UpperWidth)
1537 : B.buildUbfx(S32, UnmergeSOffset.getReg(1), Zero, UpperWidth);
1538 B.buildMergeLikeInstr(DstReg, {UnmergeSOffset.getReg(0), Extract});
1539 }
1540 MI.eraseFromParent();
1541 return true;
1542 }
1543
1544 // Expand to Src >> Offset << (64 - Width) >> (64 - Width) using 64-bit
1545 // operations.
1546 auto ExtShift = B.buildSub(S32, B.buildConstant(S32, 64), WidthReg);
1547 auto SignBit = B.buildShl(S64, ShiftOffset, ExtShift);
1548 if (Signed)
1549 B.buildAShr(S64, SignBit, ExtShift);
1550 else
1551 B.buildLShr(S64, SignBit, ExtShift);
1552 MI.eraseFromParent();
1553 return true;
1554 }
1555
1556 // The scalar form packs the offset and width in a single operand.
1557
1558 ApplyRegBankMapping ApplyBank(B, *this, MRI, &AMDGPU::SGPRRegBank);
1559
1560 // Ensure the high bits are clear to insert the offset.
1561 auto OffsetMask = B.buildConstant(S32, maskTrailingOnes<unsigned>(6));
1562 auto ClampOffset = B.buildAnd(S32, OffsetReg, OffsetMask);
1563
1564 // Zeros out the low bits, so don't bother clamping the input value.
1565 auto ShiftWidth = B.buildShl(S32, WidthReg, B.buildConstant(S32, 16));
1566
1567 // Transformation function, pack the offset and width of a BFE into
1568 // the format expected by the S_BFE_I32 / S_BFE_U32. In the second
1569 // source, bits [5:0] contain the offset and bits [22:16] the width.
1570 auto MergedInputs = B.buildOr(S32, ClampOffset, ShiftWidth);
1571
1572 // TODO: It might be worth using a pseudo here to avoid scc clobber and
1573 // register class constraints.
1574 unsigned Opc = Ty == S32 ? (Signed ? AMDGPU::S_BFE_I32 : AMDGPU::S_BFE_U32) :
1575 (Signed ? AMDGPU::S_BFE_I64 : AMDGPU::S_BFE_U64);
1576
1577 auto MIB = B.buildInstr(Opc, {DstReg}, {SrcReg, MergedInputs});
1578 constrainSelectedInstRegOperands(*MIB, *TII, *TRI, *this);
1579
1580 MI.eraseFromParent();
1581 return true;
1582}
1583
1585 MachineIRBuilder &B, const OperandsMapper &OpdMapper) const {
1586 MachineInstr &MI = OpdMapper.getMI();
1587 MachineRegisterInfo &MRI = OpdMapper.getMRI();
1588
1589 // Insert basic copies.
1590 applyDefaultMapping(OpdMapper);
1591
1592 Register Dst0 = MI.getOperand(0).getReg();
1593 Register Dst1 = MI.getOperand(1).getReg();
1594 Register Src0 = MI.getOperand(2).getReg();
1595 Register Src1 = MI.getOperand(3).getReg();
1596 Register Src2 = MI.getOperand(4).getReg();
1597
1598 if (MRI.getRegBankOrNull(Src0) == &AMDGPU::VGPRRegBank)
1599 return true;
1600
1601 bool IsUnsigned = MI.getOpcode() == AMDGPU::G_AMDGPU_MAD_U64_U32;
1602 LLT S1 = LLT::scalar(1);
1603 LLT S32 = LLT::scalar(32);
1604
1605 bool DstOnValu = MRI.getRegBankOrNull(Src2) == &AMDGPU::VGPRRegBank;
1606 bool Accumulate = true;
1607
1608 if (!DstOnValu) {
1609 if (mi_match(Src2, MRI, m_ZeroInt()))
1610 Accumulate = false;
1611 }
1612
1613 // Keep the multiplication on the SALU.
1614 Register DstHi;
1615 Register DstLo = B.buildMul(S32, Src0, Src1).getReg(0);
1616 bool MulHiInVgpr = false;
1617
1618 MRI.setRegBank(DstLo, AMDGPU::SGPRRegBank);
1619
1620 if (Subtarget.hasSMulHi()) {
1621 DstHi = IsUnsigned ? B.buildUMulH(S32, Src0, Src1).getReg(0)
1622 : B.buildSMulH(S32, Src0, Src1).getReg(0);
1623 MRI.setRegBank(DstHi, AMDGPU::SGPRRegBank);
1624 } else {
1625 Register VSrc0 = B.buildCopy(S32, Src0).getReg(0);
1626 Register VSrc1 = B.buildCopy(S32, Src1).getReg(0);
1627
1628 MRI.setRegBank(VSrc0, AMDGPU::VGPRRegBank);
1629 MRI.setRegBank(VSrc1, AMDGPU::VGPRRegBank);
1630
1631 DstHi = IsUnsigned ? B.buildUMulH(S32, VSrc0, VSrc1).getReg(0)
1632 : B.buildSMulH(S32, VSrc0, VSrc1).getReg(0);
1633 MRI.setRegBank(DstHi, AMDGPU::VGPRRegBank);
1634
1635 if (!DstOnValu) {
1636 DstHi = buildReadFirstLane(B, MRI, DstHi);
1637 } else {
1638 MulHiInVgpr = true;
1639 }
1640 }
1641
1642 // Accumulate and produce the "carry-out" bit.
1643 //
1644 // The "carry-out" is defined as bit 64 of the result when computed as a
1645 // big integer. For unsigned multiply-add, this matches the usual definition
1646 // of carry-out. For signed multiply-add, bit 64 is the sign bit of the
1647 // result, which is determined as:
1648 // sign(Src0 * Src1) + sign(Src2) + carry-out from unsigned 64-bit add
1649 LLT CarryType = DstOnValu ? S1 : S32;
1650 const RegisterBank &CarryBank =
1651 DstOnValu ? AMDGPU::VCCRegBank : AMDGPU::SGPRRegBank;
1652 const RegisterBank &DstBank =
1653 DstOnValu ? AMDGPU::VGPRRegBank : AMDGPU::SGPRRegBank;
1654 Register Carry;
1655 Register Zero;
1656
1657 if (!IsUnsigned) {
1658 Zero = B.buildConstant(S32, 0).getReg(0);
1659 MRI.setRegBank(Zero,
1660 MulHiInVgpr ? AMDGPU::VGPRRegBank : AMDGPU::SGPRRegBank);
1661
1662 Carry = B.buildICmp(CmpInst::ICMP_SLT, MulHiInVgpr ? S1 : S32, DstHi, Zero)
1663 .getReg(0);
1664 MRI.setRegBank(Carry, MulHiInVgpr ? AMDGPU::VCCRegBank
1665 : AMDGPU::SGPRRegBank);
1666
1667 if (DstOnValu && !MulHiInVgpr) {
1668 Carry = B.buildTrunc(S1, Carry).getReg(0);
1669 MRI.setRegBank(Carry, AMDGPU::VCCRegBank);
1670 }
1671 }
1672
1673 if (Accumulate) {
1674 if (DstOnValu) {
1675 DstLo = B.buildCopy(S32, DstLo).getReg(0);
1676 DstHi = B.buildCopy(S32, DstHi).getReg(0);
1677 MRI.setRegBank(DstLo, AMDGPU::VGPRRegBank);
1678 MRI.setRegBank(DstHi, AMDGPU::VGPRRegBank);
1679 }
1680
1681 auto Unmerge = B.buildUnmerge(S32, Src2);
1682 Register Src2Lo = Unmerge.getReg(0);
1683 Register Src2Hi = Unmerge.getReg(1);
1684 MRI.setRegBank(Src2Lo, DstBank);
1685 MRI.setRegBank(Src2Hi, DstBank);
1686
1687 if (!IsUnsigned) {
1688 auto Src2Sign = B.buildICmp(CmpInst::ICMP_SLT, CarryType, Src2Hi, Zero);
1689 MRI.setRegBank(Src2Sign.getReg(0), CarryBank);
1690
1691 Carry = B.buildXor(CarryType, Carry, Src2Sign).getReg(0);
1692 MRI.setRegBank(Carry, CarryBank);
1693 }
1694
1695 auto AddLo = B.buildUAddo(S32, CarryType, DstLo, Src2Lo);
1696 DstLo = AddLo.getReg(0);
1697 Register CarryLo = AddLo.getReg(1);
1698 MRI.setRegBank(DstLo, DstBank);
1699 MRI.setRegBank(CarryLo, CarryBank);
1700
1701 auto AddHi = B.buildUAdde(S32, CarryType, DstHi, Src2Hi, CarryLo);
1702 DstHi = AddHi.getReg(0);
1703 MRI.setRegBank(DstHi, DstBank);
1704
1705 Register CarryHi = AddHi.getReg(1);
1706 MRI.setRegBank(CarryHi, CarryBank);
1707
1708 if (IsUnsigned) {
1709 Carry = CarryHi;
1710 } else {
1711 Carry = B.buildXor(CarryType, Carry, CarryHi).getReg(0);
1712 MRI.setRegBank(Carry, CarryBank);
1713 }
1714 } else {
1715 if (IsUnsigned) {
1716 Carry = B.buildConstant(CarryType, 0).getReg(0);
1717 MRI.setRegBank(Carry, CarryBank);
1718 }
1719 }
1720
1721 B.buildMergeLikeInstr(Dst0, {DstLo, DstHi});
1722
1723 if (DstOnValu) {
1724 B.buildCopy(Dst1, Carry);
1725 } else {
1726 B.buildTrunc(Dst1, Carry);
1727 }
1728
1729 MI.eraseFromParent();
1730 return true;
1731}
1732
1733// Return a suitable opcode for extending the operands of Opc when widening.
1734static unsigned getExtendOp(unsigned Opc) {
1735 switch (Opc) {
1736 case TargetOpcode::G_ASHR:
1737 case TargetOpcode::G_SMIN:
1738 case TargetOpcode::G_SMAX:
1739 return TargetOpcode::G_SEXT;
1740 case TargetOpcode::G_LSHR:
1741 case TargetOpcode::G_UMIN:
1742 case TargetOpcode::G_UMAX:
1743 return TargetOpcode::G_ZEXT;
1744 default:
1745 return TargetOpcode::G_ANYEXT;
1746 }
1747}
1748
1749// Emit a legalized extension from <2 x s16> to 2 32-bit components, avoiding
1750// any illegal vector extend or unmerge operations.
1751static std::pair<Register, Register>
1752unpackV2S16ToS32(MachineIRBuilder &B, Register Src, unsigned ExtOpcode) {
1753 const LLT S32 = LLT::scalar(32);
1754 auto Bitcast = B.buildBitcast(S32, Src);
1755
1756 if (ExtOpcode == TargetOpcode::G_SEXT) {
1757 auto ExtLo = B.buildSExtInReg(S32, Bitcast, 16);
1758 auto ShiftHi = B.buildAShr(S32, Bitcast, B.buildConstant(S32, 16));
1759 return std::pair(ExtLo.getReg(0), ShiftHi.getReg(0));
1760 }
1761
1762 auto ShiftHi = B.buildLShr(S32, Bitcast, B.buildConstant(S32, 16));
1763 if (ExtOpcode == TargetOpcode::G_ZEXT) {
1764 auto ExtLo = B.buildAnd(S32, Bitcast, B.buildConstant(S32, 0xffff));
1765 return std::pair(ExtLo.getReg(0), ShiftHi.getReg(0));
1766 }
1767
1768 assert(ExtOpcode == TargetOpcode::G_ANYEXT);
1769 return std::pair(Bitcast.getReg(0), ShiftHi.getReg(0));
1770}
1771
1772// For cases where only a single copy is inserted for matching register banks.
1773// Replace the register in the instruction operand
1775 const AMDGPURegisterBankInfo::OperandsMapper &OpdMapper, unsigned OpIdx) {
1776 SmallVector<unsigned, 1> SrcReg(OpdMapper.getVRegs(OpIdx));
1777 if (!SrcReg.empty()) {
1778 assert(SrcReg.size() == 1);
1779 OpdMapper.getMI().getOperand(OpIdx).setReg(SrcReg[0]);
1780 return true;
1781 }
1782
1783 return false;
1784}
1785
1786/// Handle register layout difference for f16 images for some subtargets.
1789 Register Reg) const {
1790 if (!Subtarget.hasUnpackedD16VMem())
1791 return Reg;
1792
1793 const LLT S16 = LLT::scalar(16);
1794 LLT StoreVT = MRI.getType(Reg);
1795 if (!StoreVT.isVector() || StoreVT.getElementType() != S16)
1796 return Reg;
1797
1798 auto Unmerge = B.buildUnmerge(S16, Reg);
1799
1800
1801 SmallVector<Register, 4> WideRegs;
1802 for (int I = 0, E = Unmerge->getNumOperands() - 1; I != E; ++I)
1803 WideRegs.push_back(Unmerge.getReg(I));
1804
1805 const LLT S32 = LLT::scalar(32);
1806 int NumElts = StoreVT.getNumElements();
1807
1808 return B.buildMergeLikeInstr(LLT::fixed_vector(NumElts, S32), WideRegs)
1809 .getReg(0);
1810}
1811
1812static std::pair<Register, unsigned>
1814 int64_t Const;
1815 if (mi_match(Reg, MRI, m_ICst(Const)))
1816 return std::pair(Register(), Const);
1817
1818 Register Base;
1819 if (mi_match(Reg, MRI, m_GAdd(m_Reg(Base), m_ICst(Const))))
1820 return std::pair(Base, Const);
1821
1822 // TODO: Handle G_OR used for add case
1823 return std::pair(Reg, 0);
1824}
1825
1826std::pair<Register, unsigned>
1828 Register OrigOffset) const {
1829 const unsigned MaxImm = SIInstrInfo::getMaxMUBUFImmOffset(Subtarget);
1830 Register BaseReg;
1831 unsigned ImmOffset;
1832 const LLT S32 = LLT::scalar(32);
1833
1834 // TODO: Use AMDGPU::getBaseWithConstantOffset() instead.
1835 std::tie(BaseReg, ImmOffset) = getBaseWithConstantOffset(*B.getMRI(),
1836 OrigOffset);
1837
1838 unsigned C1 = 0;
1839 if (ImmOffset != 0) {
1840 // If the immediate value is too big for the immoffset field, put only bits
1841 // that would normally fit in the immoffset field. The remaining value that
1842 // is copied/added for the voffset field is a large power of 2, and it
1843 // stands more chance of being CSEd with the copy/add for another similar
1844 // load/store.
1845 // However, do not do that rounding down if that is a negative
1846 // number, as it appears to be illegal to have a negative offset in the
1847 // vgpr, even if adding the immediate offset makes it positive.
1848 unsigned Overflow = ImmOffset & ~MaxImm;
1849 ImmOffset -= Overflow;
1850 if (static_cast<int32_t>(Overflow) < 0) {
1851 Overflow += ImmOffset;
1852 ImmOffset = 0;
1853 }
1854
1855 C1 = ImmOffset;
1856 if (Overflow != 0) {
1857 if (!BaseReg)
1858 BaseReg = B.buildConstant(S32, Overflow).getReg(0);
1859 else {
1860 auto OverflowVal = B.buildConstant(S32, Overflow);
1861 BaseReg = B.buildAdd(S32, BaseReg, OverflowVal).getReg(0);
1862 }
1863 }
1864 }
1865
1866 if (!BaseReg)
1867 BaseReg = B.buildConstant(S32, 0).getReg(0);
1868
1869 return {BaseReg, C1};
1870}
1871
1873 Register SrcReg) const {
1874 MachineRegisterInfo &MRI = *B.getMRI();
1875 LLT SrcTy = MRI.getType(SrcReg);
1876 if (SrcTy.getSizeInBits() == 32) {
1877 // Use a v_mov_b32 here to make the exec dependency explicit.
1878 B.buildInstr(AMDGPU::V_MOV_B32_e32)
1879 .addDef(DstReg)
1880 .addUse(SrcReg);
1881 return constrainGenericRegister(DstReg, AMDGPU::VGPR_32RegClass, MRI) &&
1882 constrainGenericRegister(SrcReg, AMDGPU::SReg_32RegClass, MRI);
1883 }
1884
1885 Register TmpReg0 = MRI.createVirtualRegister(&AMDGPU::VGPR_32RegClass);
1886 Register TmpReg1 = MRI.createVirtualRegister(&AMDGPU::VGPR_32RegClass);
1887
1888 B.buildInstr(AMDGPU::V_MOV_B32_e32)
1889 .addDef(TmpReg0)
1890 .addUse(SrcReg, {}, AMDGPU::sub0);
1891 B.buildInstr(AMDGPU::V_MOV_B32_e32)
1892 .addDef(TmpReg1)
1893 .addUse(SrcReg, {}, AMDGPU::sub1);
1894 B.buildInstr(AMDGPU::REG_SEQUENCE)
1895 .addDef(DstReg)
1896 .addUse(TmpReg0)
1897 .addImm(AMDGPU::sub0)
1898 .addUse(TmpReg1)
1899 .addImm(AMDGPU::sub1);
1900
1901 return constrainGenericRegister(SrcReg, AMDGPU::SReg_64RegClass, MRI) &&
1902 constrainGenericRegister(DstReg, AMDGPU::VReg_64RegClass, MRI);
1903}
1904
1905/// Utility function for pushing dynamic vector indexes with a constant offset
1906/// into waterfall loops.
1908 MachineInstr &IdxUseInstr,
1909 unsigned OpIdx,
1910 unsigned ConstOffset) {
1911 MachineRegisterInfo &MRI = *B.getMRI();
1912 const LLT S32 = LLT::scalar(32);
1913 Register WaterfallIdx = IdxUseInstr.getOperand(OpIdx).getReg();
1914 B.setInsertPt(*IdxUseInstr.getParent(), IdxUseInstr.getIterator());
1915
1916 auto MaterializedOffset = B.buildConstant(S32, ConstOffset);
1917
1918 auto Add = B.buildAdd(S32, WaterfallIdx, MaterializedOffset);
1919 MRI.setRegBank(MaterializedOffset.getReg(0), AMDGPU::SGPRRegBank);
1920 MRI.setRegBank(Add.getReg(0), AMDGPU::SGPRRegBank);
1921 IdxUseInstr.getOperand(OpIdx).setReg(Add.getReg(0));
1922}
1923
1924/// Implement extending a 32-bit value to a 64-bit value. \p Lo32Reg is the
1925/// original 32-bit source value (to be inserted in the low part of the combined
1926/// 64-bit result), and \p Hi32Reg is the high half of the combined 64-bit
1927/// value.
1929 Register Hi32Reg, Register Lo32Reg,
1930 unsigned ExtOpc,
1931 const RegisterBank &RegBank,
1932 bool IsBooleanSrc = false) {
1933 if (ExtOpc == AMDGPU::G_ZEXT) {
1934 B.buildConstant(Hi32Reg, 0);
1935 } else if (ExtOpc == AMDGPU::G_SEXT) {
1936 if (IsBooleanSrc) {
1937 // If we know the original source was an s1, the high half is the same as
1938 // the low.
1939 B.buildCopy(Hi32Reg, Lo32Reg);
1940 } else {
1941 // Replicate sign bit from 32-bit extended part.
1942 auto ShiftAmt = B.buildConstant(LLT::scalar(32), 31);
1943 B.getMRI()->setRegBank(ShiftAmt.getReg(0), RegBank);
1944 B.buildAShr(Hi32Reg, Lo32Reg, ShiftAmt);
1945 }
1946 } else {
1947 assert(ExtOpc == AMDGPU::G_ANYEXT && "not an integer extension");
1948 B.buildUndef(Hi32Reg);
1949 }
1950}
1951
1952bool AMDGPURegisterBankInfo::foldExtractEltToCmpSelect(
1954 const OperandsMapper &OpdMapper) const {
1955 MachineRegisterInfo &MRI = *B.getMRI();
1956
1957 Register VecReg = MI.getOperand(1).getReg();
1958 Register Idx = MI.getOperand(2).getReg();
1959
1960 const RegisterBank &IdxBank =
1961 *OpdMapper.getInstrMapping().getOperandMapping(2).BreakDown[0].RegBank;
1962
1963 bool IsDivergentIdx = IdxBank != AMDGPU::SGPRRegBank;
1964
1965 LLT VecTy = MRI.getType(VecReg);
1966 unsigned EltSize = VecTy.getScalarSizeInBits();
1967 unsigned NumElem = VecTy.getNumElements();
1968
1969 if (!SITargetLowering::shouldExpandVectorDynExt(EltSize, NumElem,
1970 IsDivergentIdx, &Subtarget))
1971 return false;
1972
1973 LLT S32 = LLT::scalar(32);
1974
1975 const RegisterBank &DstBank =
1976 *OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
1977 const RegisterBank &SrcBank =
1978 *OpdMapper.getInstrMapping().getOperandMapping(1).BreakDown[0].RegBank;
1979
1980 const RegisterBank &CCBank =
1981 (DstBank == AMDGPU::SGPRRegBank &&
1982 SrcBank == AMDGPU::SGPRRegBank &&
1983 IdxBank == AMDGPU::SGPRRegBank) ? AMDGPU::SGPRRegBank
1984 : AMDGPU::VCCRegBank;
1985 LLT CCTy = (CCBank == AMDGPU::SGPRRegBank) ? S32 : LLT::scalar(1);
1986
1987 if (CCBank == AMDGPU::VCCRegBank && IdxBank == AMDGPU::SGPRRegBank) {
1988 Idx = B.buildCopy(S32, Idx)->getOperand(0).getReg();
1989 MRI.setRegBank(Idx, AMDGPU::VGPRRegBank);
1990 }
1991
1992 LLT EltTy = VecTy.getScalarType();
1993 SmallVector<Register, 2> DstRegs(OpdMapper.getVRegs(0));
1994 unsigned NumLanes = DstRegs.size();
1995 if (!NumLanes)
1996 NumLanes = 1;
1997 else
1998 EltTy = MRI.getType(DstRegs[0]);
1999
2000 auto UnmergeToEltTy = B.buildUnmerge(EltTy, VecReg);
2001 SmallVector<Register, 2> Res(NumLanes);
2002 for (unsigned L = 0; L < NumLanes; ++L)
2003 Res[L] = UnmergeToEltTy.getReg(L);
2004
2005 for (unsigned I = 1; I < NumElem; ++I) {
2006 auto IC = B.buildConstant(S32, I);
2007 MRI.setRegBank(IC->getOperand(0).getReg(), AMDGPU::SGPRRegBank);
2008 auto Cmp = B.buildICmp(CmpInst::ICMP_EQ, CCTy, Idx, IC);
2009 MRI.setRegBank(Cmp->getOperand(0).getReg(), CCBank);
2010
2011 for (unsigned L = 0; L < NumLanes; ++L) {
2012 auto S = B.buildSelect(EltTy, Cmp,
2013 UnmergeToEltTy.getReg(I * NumLanes + L), Res[L]);
2014
2015 for (unsigned N : { 0, 2, 3 })
2016 MRI.setRegBank(S->getOperand(N).getReg(), DstBank);
2017
2018 Res[L] = S->getOperand(0).getReg();
2019 }
2020 }
2021
2022 for (unsigned L = 0; L < NumLanes; ++L) {
2023 Register DstReg = (NumLanes == 1) ? MI.getOperand(0).getReg() : DstRegs[L];
2024 B.buildCopy(DstReg, Res[L]);
2025 MRI.setRegBank(DstReg, DstBank);
2026 }
2027
2028 MRI.setRegBank(MI.getOperand(0).getReg(), DstBank);
2029 MI.eraseFromParent();
2030
2031 return true;
2032}
2033
2034// Insert a cross regbank copy for a register if it already has a bank that
2035// differs from the one we want to set.
2038 const RegisterBank &Bank) {
2039 const RegisterBank *CurrBank = MRI.getRegBankOrNull(Reg);
2040 if (CurrBank && *CurrBank != Bank) {
2041 Register Copy = B.buildCopy(MRI.getType(Reg), Reg).getReg(0);
2042 MRI.setRegBank(Copy, Bank);
2043 return Copy;
2044 }
2045
2046 MRI.setRegBank(Reg, Bank);
2047 return Reg;
2048}
2049
2050bool AMDGPURegisterBankInfo::foldInsertEltToCmpSelect(
2052 const OperandsMapper &OpdMapper) const {
2053
2054 MachineRegisterInfo &MRI = *B.getMRI();
2055 Register VecReg = MI.getOperand(1).getReg();
2056 Register Idx = MI.getOperand(3).getReg();
2057
2058 const RegisterBank &IdxBank =
2059 *OpdMapper.getInstrMapping().getOperandMapping(3).BreakDown[0].RegBank;
2060
2061 bool IsDivergentIdx = IdxBank != AMDGPU::SGPRRegBank;
2062
2063 LLT VecTy = MRI.getType(VecReg);
2064 unsigned EltSize = VecTy.getScalarSizeInBits();
2065 unsigned NumElem = VecTy.getNumElements();
2066
2067 if (!SITargetLowering::shouldExpandVectorDynExt(EltSize, NumElem,
2068 IsDivergentIdx, &Subtarget))
2069 return false;
2070
2071 LLT S32 = LLT::scalar(32);
2072
2073 const RegisterBank &DstBank =
2074 *OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2075 const RegisterBank &SrcBank =
2076 *OpdMapper.getInstrMapping().getOperandMapping(1).BreakDown[0].RegBank;
2077 const RegisterBank &InsBank =
2078 *OpdMapper.getInstrMapping().getOperandMapping(2).BreakDown[0].RegBank;
2079
2080 const RegisterBank &CCBank =
2081 (DstBank == AMDGPU::SGPRRegBank &&
2082 SrcBank == AMDGPU::SGPRRegBank &&
2083 InsBank == AMDGPU::SGPRRegBank &&
2084 IdxBank == AMDGPU::SGPRRegBank) ? AMDGPU::SGPRRegBank
2085 : AMDGPU::VCCRegBank;
2086 LLT CCTy = (CCBank == AMDGPU::SGPRRegBank) ? S32 : LLT::scalar(1);
2087
2088 if (CCBank == AMDGPU::VCCRegBank && IdxBank == AMDGPU::SGPRRegBank) {
2089 Idx = B.buildCopy(S32, Idx)->getOperand(0).getReg();
2090 MRI.setRegBank(Idx, AMDGPU::VGPRRegBank);
2091 }
2092
2093 LLT EltTy = VecTy.getScalarType();
2094 SmallVector<Register, 2> InsRegs(OpdMapper.getVRegs(2));
2095 unsigned NumLanes = InsRegs.size();
2096 if (!NumLanes) {
2097 NumLanes = 1;
2098 InsRegs.push_back(MI.getOperand(2).getReg());
2099 } else {
2100 EltTy = MRI.getType(InsRegs[0]);
2101 }
2102
2103 auto UnmergeToEltTy = B.buildUnmerge(EltTy, VecReg);
2104 SmallVector<Register, 16> Ops(NumElem * NumLanes);
2105
2106 for (unsigned I = 0; I < NumElem; ++I) {
2107 auto IC = B.buildConstant(S32, I);
2108 MRI.setRegBank(IC->getOperand(0).getReg(), AMDGPU::SGPRRegBank);
2109 auto Cmp = B.buildICmp(CmpInst::ICMP_EQ, CCTy, Idx, IC);
2110 MRI.setRegBank(Cmp->getOperand(0).getReg(), CCBank);
2111
2112 for (unsigned L = 0; L < NumLanes; ++L) {
2113 Register Op0 = constrainRegToBank(MRI, B, InsRegs[L], DstBank);
2114 Register Op1 = UnmergeToEltTy.getReg(I * NumLanes + L);
2115 Op1 = constrainRegToBank(MRI, B, Op1, DstBank);
2116
2117 Register Select = B.buildSelect(EltTy, Cmp, Op0, Op1).getReg(0);
2118 MRI.setRegBank(Select, DstBank);
2119
2120 Ops[I * NumLanes + L] = Select;
2121 }
2122 }
2123
2124 LLT MergeTy = LLT::fixed_vector(Ops.size(), EltTy);
2125 if (MergeTy == MRI.getType(MI.getOperand(0).getReg())) {
2126 B.buildBuildVector(MI.getOperand(0), Ops);
2127 } else {
2128 auto Vec = B.buildBuildVector(MergeTy, Ops);
2129 MRI.setRegBank(Vec->getOperand(0).getReg(), DstBank);
2130 B.buildBitcast(MI.getOperand(0).getReg(), Vec);
2131 }
2132
2133 MRI.setRegBank(MI.getOperand(0).getReg(), DstBank);
2134 MI.eraseFromParent();
2135
2136 return true;
2137}
2138
2139// Break s_mul_u64 into 32-bit vector operations.
2141 MachineIRBuilder &B, const OperandsMapper &OpdMapper) const {
2142 SmallVector<Register, 2> DefRegs(OpdMapper.getVRegs(0));
2143 SmallVector<Register, 2> Src0Regs(OpdMapper.getVRegs(1));
2144 SmallVector<Register, 2> Src1Regs(OpdMapper.getVRegs(2));
2145
2146 // All inputs are SGPRs, nothing special to do.
2147 if (DefRegs.empty()) {
2148 assert(Src0Regs.empty() && Src1Regs.empty());
2149 applyDefaultMapping(OpdMapper);
2150 return;
2151 }
2152
2153 assert(DefRegs.size() == 2);
2154 assert(Src0Regs.size() == Src1Regs.size() &&
2155 (Src0Regs.empty() || Src0Regs.size() == 2));
2156
2157 MachineRegisterInfo &MRI = OpdMapper.getMRI();
2158 MachineInstr &MI = OpdMapper.getMI();
2159 Register DstReg = MI.getOperand(0).getReg();
2160 LLT HalfTy = LLT::scalar(32);
2161
2162 // Depending on where the source registers came from, the generic code may
2163 // have decided to split the inputs already or not. If not, we still need to
2164 // extract the values.
2165
2166 if (Src0Regs.empty())
2167 split64BitValueForMapping(B, Src0Regs, HalfTy, MI.getOperand(1).getReg());
2168 else
2169 setRegsToType(MRI, Src0Regs, HalfTy);
2170
2171 if (Src1Regs.empty())
2172 split64BitValueForMapping(B, Src1Regs, HalfTy, MI.getOperand(2).getReg());
2173 else
2174 setRegsToType(MRI, Src1Regs, HalfTy);
2175
2176 setRegsToType(MRI, DefRegs, HalfTy);
2177
2178 // The multiplication is done as follows:
2179 //
2180 // Op1H Op1L
2181 // * Op0H Op0L
2182 // --------------------
2183 // Op1H*Op0L Op1L*Op0L
2184 // + Op1H*Op0H Op1L*Op0H
2185 // -----------------------------------------
2186 // (Op1H*Op0L + Op1L*Op0H + carry) Op1L*Op0L
2187 //
2188 // We drop Op1H*Op0H because the result of the multiplication is a 64-bit
2189 // value and that would overflow.
2190 // The low 32-bit value is Op1L*Op0L.
2191 // The high 32-bit value is Op1H*Op0L + Op1L*Op0H + carry (from
2192 // Op1L*Op0L).
2193
2194 ApplyRegBankMapping ApplyBank(B, *this, MRI, &AMDGPU::VGPRRegBank);
2195
2196 Register Hi = B.buildUMulH(HalfTy, Src0Regs[0], Src1Regs[0]).getReg(0);
2197 Register MulLoHi = B.buildMul(HalfTy, Src0Regs[0], Src1Regs[1]).getReg(0);
2198 Register Add = B.buildAdd(HalfTy, Hi, MulLoHi).getReg(0);
2199 Register MulHiLo = B.buildMul(HalfTy, Src0Regs[1], Src1Regs[0]).getReg(0);
2200 B.buildAdd(DefRegs[1], Add, MulHiLo);
2201 B.buildMul(DefRegs[0], Src0Regs[0], Src1Regs[0]);
2202
2203 MRI.setRegBank(DstReg, AMDGPU::VGPRRegBank);
2204 MI.eraseFromParent();
2205}
2206
2208 MachineIRBuilder &B, const OperandsMapper &OpdMapper) const {
2209 MachineInstr &MI = OpdMapper.getMI();
2210 B.setInstrAndDebugLoc(MI);
2211 unsigned Opc = MI.getOpcode();
2212 MachineRegisterInfo &MRI = OpdMapper.getMRI();
2213 switch (Opc) {
2214 case AMDGPU::G_CONSTANT:
2215 case AMDGPU::G_IMPLICIT_DEF: {
2216 Register DstReg = MI.getOperand(0).getReg();
2217 LLT DstTy = MRI.getType(DstReg);
2218 if (DstTy != LLT::scalar(1))
2219 break;
2220
2221 const RegisterBank *DstBank =
2222 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2223 if (DstBank == &AMDGPU::VCCRegBank)
2224 break;
2225 SmallVector<Register, 1> DefRegs(OpdMapper.getVRegs(0));
2226 if (DefRegs.empty())
2227 DefRegs.push_back(DstReg);
2228
2229 B.setInsertPt(*MI.getParent(), ++MI.getIterator());
2230
2232 LLVMContext &Ctx = B.getMF().getFunction().getContext();
2233
2234 MI.getOperand(0).setReg(NewDstReg);
2235 if (Opc != AMDGPU::G_IMPLICIT_DEF) {
2236 uint64_t ConstVal = MI.getOperand(1).getCImm()->getZExtValue();
2237 MI.getOperand(1).setCImm(
2238 ConstantInt::get(IntegerType::getInt32Ty(Ctx), ConstVal));
2239 }
2240
2241 MRI.setRegBank(NewDstReg, *DstBank);
2242 B.buildTrunc(DefRegs[0], NewDstReg);
2243 return;
2244 }
2245 case AMDGPU::G_PHI: {
2246 Register DstReg = MI.getOperand(0).getReg();
2247 LLT DstTy = MRI.getType(DstReg);
2248 if (DstTy != LLT::scalar(1))
2249 break;
2250
2251 const LLT S32 = LLT::scalar(32);
2252 const RegisterBank *DstBank =
2253 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2254 if (DstBank == &AMDGPU::VCCRegBank) {
2255 applyDefaultMapping(OpdMapper);
2256 // The standard handling only considers the result register bank for
2257 // phis. For VCC, blindly inserting a copy when the phi is lowered will
2258 // produce an invalid copy. We can only copy with some kind of compare to
2259 // get a vector boolean result. Insert a register bank copy that will be
2260 // correctly lowered to a compare.
2261 for (unsigned I = 1, E = MI.getNumOperands(); I != E; I += 2) {
2262 Register SrcReg = MI.getOperand(I).getReg();
2263 const RegisterBank *SrcBank = getRegBank(SrcReg, MRI, *TRI);
2264
2265 if (SrcBank != &AMDGPU::VCCRegBank) {
2266 MachineBasicBlock *SrcMBB = MI.getOperand(I + 1).getMBB();
2267 B.setInsertPt(*SrcMBB, SrcMBB->getFirstTerminator());
2268
2269 auto Copy = B.buildCopy(LLT::scalar(1), SrcReg);
2270 MRI.setRegBank(Copy.getReg(0), AMDGPU::VCCRegBank);
2271 MI.getOperand(I).setReg(Copy.getReg(0));
2272 }
2273 }
2274
2275 return;
2276 }
2277
2278 // Phi handling is strange and only considers the bank of the destination.
2279 substituteSimpleCopyRegs(OpdMapper, 0);
2280
2281 // Promote SGPR/VGPR booleans to s32
2282 ApplyRegBankMapping ApplyBank(B, *this, MRI, DstBank);
2283 B.setInsertPt(B.getMBB(), MI);
2284 LegalizerHelper Helper(B.getMF(), ApplyBank, B);
2285
2286 if (Helper.widenScalar(MI, 0, S32) != LegalizerHelper::Legalized)
2287 llvm_unreachable("widen scalar should have succeeded");
2288
2289 return;
2290 }
2291 case AMDGPU::G_FCMP:
2292 if (!Subtarget.hasSALUFloatInsts())
2293 break;
2294 [[fallthrough]];
2295 case AMDGPU::G_ICMP:
2296 case AMDGPU::G_UADDO:
2297 case AMDGPU::G_USUBO:
2298 case AMDGPU::G_UADDE:
2299 case AMDGPU::G_SADDE:
2300 case AMDGPU::G_USUBE:
2301 case AMDGPU::G_SSUBE: {
2302 unsigned BoolDstOp =
2303 (Opc == AMDGPU::G_ICMP || Opc == AMDGPU::G_FCMP) ? 0 : 1;
2304 Register DstReg = MI.getOperand(BoolDstOp).getReg();
2305
2306 const RegisterBank *DstBank =
2307 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2308 if (DstBank != &AMDGPU::SGPRRegBank)
2309 break;
2310
2311 const bool HasCarryIn = MI.getNumOperands() == 5;
2312
2313 // If this is a scalar compare, promote the result to s32, as the selection
2314 // will end up using a copy to a 32-bit vreg.
2315 const LLT S32 = LLT::scalar(32);
2316 Register NewDstReg = MRI.createGenericVirtualRegister(S32);
2317 MRI.setRegBank(NewDstReg, AMDGPU::SGPRRegBank);
2318 MI.getOperand(BoolDstOp).setReg(NewDstReg);
2319
2320 if (HasCarryIn) {
2321 Register NewSrcReg = MRI.createGenericVirtualRegister(S32);
2322 MRI.setRegBank(NewSrcReg, AMDGPU::SGPRRegBank);
2323 B.buildZExt(NewSrcReg, MI.getOperand(4).getReg());
2324 MI.getOperand(4).setReg(NewSrcReg);
2325 }
2326
2327 MachineBasicBlock *MBB = MI.getParent();
2328 B.setInsertPt(*MBB, std::next(MI.getIterator()));
2329
2330 // If we had a constrained VCC result register, a copy was inserted to VCC
2331 // from SGPR.
2332 SmallVector<Register, 1> DefRegs(OpdMapper.getVRegs(0));
2333 if (DefRegs.empty())
2334 DefRegs.push_back(DstReg);
2335 B.buildTrunc(DefRegs[0], NewDstReg);
2336 return;
2337 }
2338 case AMDGPU::G_SELECT: {
2339 Register DstReg = MI.getOperand(0).getReg();
2340 LLT DstTy = MRI.getType(DstReg);
2341
2342 SmallVector<Register, 1> CondRegs(OpdMapper.getVRegs(1));
2343 if (CondRegs.empty())
2344 CondRegs.push_back(MI.getOperand(1).getReg());
2345 else {
2346 assert(CondRegs.size() == 1);
2347 }
2348
2349 const RegisterBank *CondBank = getRegBank(CondRegs[0], MRI, *TRI);
2350 if (CondBank == &AMDGPU::SGPRRegBank) {
2351 const LLT S32 = LLT::scalar(32);
2352 Register NewCondReg = MRI.createGenericVirtualRegister(S32);
2353 MRI.setRegBank(NewCondReg, AMDGPU::SGPRRegBank);
2354
2355 MI.getOperand(1).setReg(NewCondReg);
2356 B.buildZExt(NewCondReg, CondRegs[0]);
2357 }
2358
2359 if (DstTy.getSizeInBits() != 64)
2360 break;
2361
2362 LLT HalfTy = getHalfSizedType(DstTy);
2363
2364 SmallVector<Register, 2> DefRegs(OpdMapper.getVRegs(0));
2365 SmallVector<Register, 2> Src1Regs(OpdMapper.getVRegs(2));
2366 SmallVector<Register, 2> Src2Regs(OpdMapper.getVRegs(3));
2367
2368 // All inputs are SGPRs, nothing special to do.
2369 if (DefRegs.empty()) {
2370 assert(Src1Regs.empty() && Src2Regs.empty());
2371 break;
2372 }
2373
2374 if (Src1Regs.empty())
2375 split64BitValueForMapping(B, Src1Regs, HalfTy, MI.getOperand(2).getReg());
2376 else {
2377 setRegsToType(MRI, Src1Regs, HalfTy);
2378 }
2379
2380 if (Src2Regs.empty())
2381 split64BitValueForMapping(B, Src2Regs, HalfTy, MI.getOperand(3).getReg());
2382 else
2383 setRegsToType(MRI, Src2Regs, HalfTy);
2384
2385 setRegsToType(MRI, DefRegs, HalfTy);
2386
2387 auto Flags = MI.getFlags();
2388 B.buildSelect(DefRegs[0], CondRegs[0], Src1Regs[0], Src2Regs[0], Flags);
2389 B.buildSelect(DefRegs[1], CondRegs[0], Src1Regs[1], Src2Regs[1], Flags);
2390
2391 MRI.setRegBank(DstReg, AMDGPU::VGPRRegBank);
2392 MI.eraseFromParent();
2393 return;
2394 }
2395 case AMDGPU::G_BRCOND: {
2396 Register CondReg = MI.getOperand(0).getReg();
2397 // FIXME: Should use legalizer helper, but should change bool ext type.
2398 const RegisterBank *CondBank =
2399 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2400
2401 if (CondBank == &AMDGPU::SGPRRegBank) {
2402 const LLT S32 = LLT::scalar(32);
2403 Register NewCondReg = MRI.createGenericVirtualRegister(S32);
2404 MRI.setRegBank(NewCondReg, AMDGPU::SGPRRegBank);
2405
2406 MI.getOperand(0).setReg(NewCondReg);
2407 B.buildZExt(NewCondReg, CondReg);
2408 return;
2409 }
2410
2411 break;
2412 }
2413 case AMDGPU::G_AND:
2414 case AMDGPU::G_OR:
2415 case AMDGPU::G_XOR: {
2416 // 64-bit and is only available on the SALU, so split into 2 32-bit ops if
2417 // there is a VGPR input.
2418 Register DstReg = MI.getOperand(0).getReg();
2419 LLT DstTy = MRI.getType(DstReg);
2420
2421 const RegisterBank *DstBank =
2422 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2423
2424 if (DstTy.getSizeInBits() == 1) {
2425 if (DstBank == &AMDGPU::VCCRegBank)
2426 break;
2427
2428 MachineFunction *MF = MI.getMF();
2429 ApplyRegBankMapping ApplyBank(B, *this, MRI, DstBank);
2430 LegalizerHelper Helper(*MF, ApplyBank, B);
2431
2432 if (Helper.widenScalar(MI, 0, LLT::scalar(32)) !=
2434 llvm_unreachable("widen scalar should have succeeded");
2435 return;
2436 }
2437
2438 if (DstTy.getSizeInBits() == 16 && DstBank == &AMDGPU::SGPRRegBank) {
2439 const LLT S32 = LLT::scalar(32);
2440 MachineBasicBlock *MBB = MI.getParent();
2441 MachineFunction *MF = MBB->getParent();
2442 ApplyRegBankMapping ApplySALU(B, *this, MRI, &AMDGPU::SGPRRegBank);
2443 LegalizerHelper Helper(*MF, ApplySALU, B);
2444 // Widen to S32, but handle `G_XOR x, -1` differently. Legalizer widening
2445 // will use a G_ANYEXT to extend the -1 which prevents matching G_XOR -1
2446 // as "not".
2447 if (MI.getOpcode() == AMDGPU::G_XOR &&
2448 mi_match(MI.getOperand(2).getReg(), MRI, m_SpecificICstOrSplat(-1))) {
2449 Helper.widenScalarSrc(MI, S32, 1, AMDGPU::G_ANYEXT);
2450 Helper.widenScalarSrc(MI, S32, 2, AMDGPU::G_SEXT);
2451 Helper.widenScalarDst(MI, S32);
2452 } else {
2453 if (Helper.widenScalar(MI, 0, S32) != LegalizerHelper::Legalized)
2454 llvm_unreachable("widen scalar should have succeeded");
2455 }
2456 return;
2457 }
2458
2459 if (DstTy.getSizeInBits() != 64)
2460 break;
2461
2462 LLT HalfTy = getHalfSizedType(DstTy);
2463 SmallVector<Register, 2> DefRegs(OpdMapper.getVRegs(0));
2464 SmallVector<Register, 2> Src0Regs(OpdMapper.getVRegs(1));
2465 SmallVector<Register, 2> Src1Regs(OpdMapper.getVRegs(2));
2466
2467 // All inputs are SGPRs, nothing special to do.
2468 if (DefRegs.empty()) {
2469 assert(Src0Regs.empty() && Src1Regs.empty());
2470 break;
2471 }
2472
2473 assert(DefRegs.size() == 2);
2474 assert(Src0Regs.size() == Src1Regs.size() &&
2475 (Src0Regs.empty() || Src0Regs.size() == 2));
2476
2477 // Depending on where the source registers came from, the generic code may
2478 // have decided to split the inputs already or not. If not, we still need to
2479 // extract the values.
2480
2481 if (Src0Regs.empty())
2482 split64BitValueForMapping(B, Src0Regs, HalfTy, MI.getOperand(1).getReg());
2483 else
2484 setRegsToType(MRI, Src0Regs, HalfTy);
2485
2486 if (Src1Regs.empty())
2487 split64BitValueForMapping(B, Src1Regs, HalfTy, MI.getOperand(2).getReg());
2488 else
2489 setRegsToType(MRI, Src1Regs, HalfTy);
2490
2491 setRegsToType(MRI, DefRegs, HalfTy);
2492
2493 auto Flags = MI.getFlags();
2494 B.buildInstr(Opc, {DefRegs[0]}, {Src0Regs[0], Src1Regs[0]}, Flags);
2495 B.buildInstr(Opc, {DefRegs[1]}, {Src0Regs[1], Src1Regs[1]}, Flags);
2496
2497 MRI.setRegBank(DstReg, AMDGPU::VGPRRegBank);
2498 MI.eraseFromParent();
2499 return;
2500 }
2501 case AMDGPU::G_ABS: {
2502 Register SrcReg = MI.getOperand(1).getReg();
2503 const RegisterBank *SrcBank = MRI.getRegBankOrNull(SrcReg);
2504
2505 // There is no VALU abs instruction so we need to replace it with a sub and
2506 // max combination.
2507 if (SrcBank && SrcBank == &AMDGPU::VGPRRegBank) {
2508 MachineFunction *MF = MI.getMF();
2509 ApplyRegBankMapping Apply(B, *this, MRI, &AMDGPU::VGPRRegBank);
2510 LegalizerHelper Helper(*MF, Apply, B);
2511
2513 llvm_unreachable("lowerAbsToMaxNeg should have succeeded");
2514 return;
2515 }
2516 [[fallthrough]];
2517 }
2518 case AMDGPU::G_ADD:
2519 case AMDGPU::G_SUB:
2520 case AMDGPU::G_MUL:
2521 case AMDGPU::G_SHL:
2522 case AMDGPU::G_LSHR:
2523 case AMDGPU::G_ASHR:
2524 case AMDGPU::G_SMIN:
2525 case AMDGPU::G_SMAX:
2526 case AMDGPU::G_UMIN:
2527 case AMDGPU::G_UMAX: {
2528 Register DstReg = MI.getOperand(0).getReg();
2529 LLT DstTy = MRI.getType(DstReg);
2530
2531 // Special case for s_mul_u64. There is not a vector equivalent of
2532 // s_mul_u64. Hence, we have to break down s_mul_u64 into 32-bit vector
2533 // multiplications.
2534 if (!Subtarget.hasVMulU64Inst() && Opc == AMDGPU::G_MUL &&
2535 DstTy.getSizeInBits() == 64) {
2536 applyMappingSMULU64(B, OpdMapper);
2537 return;
2538 }
2539
2540 // 16-bit operations are VALU only, but can be promoted to 32-bit SALU.
2541 // Packed 16-bit operations need to be scalarized and promoted.
2542 if (DstTy != LLT::scalar(16) && DstTy != LLT::fixed_vector(2, 16))
2543 break;
2544
2545 const RegisterBank *DstBank =
2546 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2547 if (DstBank == &AMDGPU::VGPRRegBank)
2548 break;
2549
2550 const LLT S32 = LLT::scalar(32);
2551 MachineBasicBlock *MBB = MI.getParent();
2552 MachineFunction *MF = MBB->getParent();
2553 ApplyRegBankMapping ApplySALU(B, *this, MRI, &AMDGPU::SGPRRegBank);
2554
2555 if (DstTy.isVector() && Opc == AMDGPU::G_ABS) {
2556 Register WideSrcLo, WideSrcHi;
2557
2558 std::tie(WideSrcLo, WideSrcHi) =
2559 unpackV2S16ToS32(B, MI.getOperand(1).getReg(), TargetOpcode::G_SEXT);
2560 auto Lo = B.buildInstr(AMDGPU::G_ABS, {S32}, {WideSrcLo});
2561 auto Hi = B.buildInstr(AMDGPU::G_ABS, {S32}, {WideSrcHi});
2562 B.buildBuildVectorTrunc(DstReg, {Lo.getReg(0), Hi.getReg(0)});
2563 MI.eraseFromParent();
2564 return;
2565 }
2566
2567 if (DstTy.isVector()) {
2568 Register WideSrc0Lo, WideSrc0Hi;
2569 Register WideSrc1Lo, WideSrc1Hi;
2570
2571 unsigned ExtendOp = getExtendOp(MI.getOpcode());
2572 std::tie(WideSrc0Lo, WideSrc0Hi)
2573 = unpackV2S16ToS32(B, MI.getOperand(1).getReg(), ExtendOp);
2574 std::tie(WideSrc1Lo, WideSrc1Hi)
2575 = unpackV2S16ToS32(B, MI.getOperand(2).getReg(), ExtendOp);
2576 auto Lo = B.buildInstr(MI.getOpcode(), {S32}, {WideSrc0Lo, WideSrc1Lo});
2577 auto Hi = B.buildInstr(MI.getOpcode(), {S32}, {WideSrc0Hi, WideSrc1Hi});
2578 B.buildBuildVectorTrunc(DstReg, {Lo.getReg(0), Hi.getReg(0)});
2579 MI.eraseFromParent();
2580 } else {
2581 LegalizerHelper Helper(*MF, ApplySALU, B);
2582
2583 if (Helper.widenScalar(MI, 0, S32) != LegalizerHelper::Legalized)
2584 llvm_unreachable("widen scalar should have succeeded");
2585
2586 // FIXME: s16 shift amounts should be legal.
2587 if (Opc == AMDGPU::G_SHL || Opc == AMDGPU::G_LSHR ||
2588 Opc == AMDGPU::G_ASHR) {
2589 B.setInsertPt(*MBB, MI.getIterator());
2590 if (Helper.widenScalar(MI, 1, S32) != LegalizerHelper::Legalized)
2591 llvm_unreachable("widen scalar should have succeeded");
2592 }
2593 }
2594
2595 return;
2596 }
2597 case AMDGPU::G_AMDGPU_S_MUL_I64_I32:
2598 case AMDGPU::G_AMDGPU_S_MUL_U64_U32: {
2599 // This is a special case for s_mul_u64. We use
2600 // G_AMDGPU_S_MUL_I64_I32 opcode to represent an s_mul_u64 operation
2601 // where the 33 higher bits are sign-extended and
2602 // G_AMDGPU_S_MUL_U64_U32 opcode to represent an s_mul_u64 operation
2603 // where the 32 higher bits are zero-extended. In case scalar registers are
2604 // selected, both opcodes are lowered as s_mul_u64. If the vector registers
2605 // are selected, then G_AMDGPU_S_MUL_I64_I32 and
2606 // G_AMDGPU_S_MUL_U64_U32 are lowered with a vector mad instruction.
2607
2608 // Insert basic copies.
2609 applyDefaultMapping(OpdMapper);
2610
2611 Register DstReg = MI.getOperand(0).getReg();
2612 Register SrcReg0 = MI.getOperand(1).getReg();
2613 Register SrcReg1 = MI.getOperand(2).getReg();
2614 const LLT S32 = LLT::scalar(32);
2615 const LLT S64 = LLT::scalar(64);
2616 assert(MRI.getType(DstReg) == S64 && "This is a special case for s_mul_u64 "
2617 "that handles only 64-bit operands.");
2618 const RegisterBank *DstBank =
2619 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2620
2621 // Replace G_AMDGPU_S_MUL_I64_I32 and G_AMDGPU_S_MUL_U64_U32
2622 // with s_mul_u64 operation.
2623 if (DstBank == &AMDGPU::SGPRRegBank) {
2624 MI.setDesc(TII->get(AMDGPU::S_MUL_U64));
2625 MRI.setRegClass(DstReg, &AMDGPU::SGPR_64RegClass);
2626 MRI.setRegClass(SrcReg0, &AMDGPU::SGPR_64RegClass);
2627 MRI.setRegClass(SrcReg1, &AMDGPU::SGPR_64RegClass);
2628 return;
2629 }
2630
2631 // Replace G_AMDGPU_S_MUL_I64_I32 and G_AMDGPU_S_MUL_U64_U32
2632 // with a vector mad.
2633 assert(MRI.getRegBankOrNull(DstReg) == &AMDGPU::VGPRRegBank &&
2634 "The destination operand should be in vector registers.");
2635
2636 // Extract the lower subregister from the first operand.
2637 Register Op0L = MRI.createVirtualRegister(&AMDGPU::VGPR_32RegClass);
2638 MRI.setRegClass(Op0L, &AMDGPU::VGPR_32RegClass);
2639 MRI.setType(Op0L, S32);
2640 B.buildTrunc(Op0L, SrcReg0);
2641
2642 // Extract the lower subregister from the second operand.
2643 Register Op1L = MRI.createVirtualRegister(&AMDGPU::VGPR_32RegClass);
2644 MRI.setRegClass(Op1L, &AMDGPU::VGPR_32RegClass);
2645 MRI.setType(Op1L, S32);
2646 B.buildTrunc(Op1L, SrcReg1);
2647
2648 unsigned NewOpc = Opc == AMDGPU::G_AMDGPU_S_MUL_U64_U32
2649 ? AMDGPU::G_AMDGPU_MAD_U64_U32
2650 : AMDGPU::G_AMDGPU_MAD_I64_I32;
2651
2653 Register Zero64 = B.buildConstant(S64, 0).getReg(0);
2654 MRI.setRegClass(Zero64, &AMDGPU::VReg_64RegClass);
2655 Register CarryOut = MRI.createVirtualRegister(&AMDGPU::VReg_64RegClass);
2656 MRI.setRegClass(CarryOut, &AMDGPU::VReg_64RegClass);
2657 B.buildInstr(NewOpc, {DstReg, CarryOut}, {Op0L, Op1L, Zero64});
2658 MI.eraseFromParent();
2659 return;
2660 }
2661 case AMDGPU::G_SEXT_INREG: {
2662 SmallVector<Register, 2> SrcRegs(OpdMapper.getVRegs(1));
2663 if (SrcRegs.empty())
2664 break; // Nothing to repair
2665
2666 const LLT S32 = LLT::scalar(32);
2667 ApplyRegBankMapping O(B, *this, MRI, &AMDGPU::VGPRRegBank);
2668
2669 // Don't use LegalizerHelper's narrowScalar. It produces unwanted G_SEXTs
2670 // we would need to further expand, and doesn't let us directly set the
2671 // result registers.
2672 SmallVector<Register, 2> DstRegs(OpdMapper.getVRegs(0));
2673
2674 int Amt = MI.getOperand(2).getImm();
2675 if (Amt <= 32) {
2676 // Downstream users have expectations for the high bit behavior, so freeze
2677 // incoming undefined bits.
2678 if (Amt == 32) {
2679 // The low bits are unchanged.
2680 B.buildFreeze(DstRegs[0], SrcRegs[0]);
2681 } else {
2682 auto Freeze = B.buildFreeze(S32, SrcRegs[0]);
2683 // Extend in the low bits and propagate the sign bit to the high half.
2684 B.buildSExtInReg(DstRegs[0], Freeze, Amt);
2685 }
2686
2687 B.buildAShr(DstRegs[1], DstRegs[0], B.buildConstant(S32, 31));
2688 } else {
2689 // The low bits are unchanged, and extend in the high bits.
2690 // No freeze required
2691 B.buildCopy(DstRegs[0], SrcRegs[0]);
2692 B.buildSExtInReg(DstRegs[1], DstRegs[0], Amt - 32);
2693 }
2694
2695 Register DstReg = MI.getOperand(0).getReg();
2696 MRI.setRegBank(DstReg, AMDGPU::VGPRRegBank);
2697 MI.eraseFromParent();
2698 return;
2699 }
2700 case AMDGPU::G_CTPOP:
2701 case AMDGPU::G_BITREVERSE: {
2702 const RegisterBank *DstBank =
2703 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2704 if (DstBank == &AMDGPU::SGPRRegBank)
2705 break;
2706
2707 Register SrcReg = MI.getOperand(1).getReg();
2708 const LLT S32 = LLT::scalar(32);
2709 LLT Ty = MRI.getType(SrcReg);
2710 if (Ty == S32)
2711 break;
2712
2713 ApplyRegBankMapping ApplyVALU(B, *this, MRI, &AMDGPU::VGPRRegBank);
2714
2715 MachineFunction &MF = B.getMF();
2716 LegalizerHelper Helper(MF, ApplyVALU, B);
2717
2718 if (Helper.narrowScalar(MI, 1, S32) != LegalizerHelper::Legalized)
2719 llvm_unreachable("narrowScalar should have succeeded");
2720 return;
2721 }
2722 case AMDGPU::G_AMDGPU_FFBH_U32:
2723 case AMDGPU::G_AMDGPU_FFBL_B32:
2724 case AMDGPU::G_CTLZ_ZERO_POISON:
2725 case AMDGPU::G_CTTZ_ZERO_POISON: {
2726 const RegisterBank *DstBank =
2727 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
2728 if (DstBank == &AMDGPU::SGPRRegBank)
2729 break;
2730
2731 Register SrcReg = MI.getOperand(1).getReg();
2732 const LLT S32 = LLT::scalar(32);
2733 LLT Ty = MRI.getType(SrcReg);
2734 if (Ty == S32)
2735 break;
2736
2737 // We can narrow this more efficiently than Helper can by using ffbh/ffbl
2738 // which return -1 when the input is zero:
2739 // (ctlz_zero_poison hi:lo) -> (umin (ffbh hi), (add (ffbh lo), 32))
2740 // (cttz_zero_poison hi:lo) -> (umin (add (ffbl hi), 32), (ffbl lo))
2741 // (ffbh hi:lo) -> (umin (ffbh hi), (uaddsat (ffbh lo), 32))
2742 // (ffbl hi:lo) -> (umin (uaddsat (ffbh hi), 32), (ffbh lo))
2743 ApplyRegBankMapping ApplyVALU(B, *this, MRI, &AMDGPU::VGPRRegBank);
2744 SmallVector<Register, 2> SrcRegs(OpdMapper.getVRegs(1));
2745 unsigned NewOpc = Opc == AMDGPU::G_CTLZ_ZERO_POISON
2746 ? (unsigned)AMDGPU::G_AMDGPU_FFBH_U32
2747 : Opc == AMDGPU::G_CTTZ_ZERO_POISON
2748 ? (unsigned)AMDGPU::G_AMDGPU_FFBL_B32
2749 : Opc;
2750 unsigned Idx = NewOpc == AMDGPU::G_AMDGPU_FFBH_U32;
2751 auto X = B.buildInstr(NewOpc, {S32}, {SrcRegs[Idx]});
2752 auto Y = B.buildInstr(NewOpc, {S32}, {SrcRegs[Idx ^ 1]});
2753 unsigned AddOpc =
2754 Opc == AMDGPU::G_CTLZ_ZERO_POISON || Opc == AMDGPU::G_CTTZ_ZERO_POISON
2755 ? AMDGPU::G_ADD
2756 : AMDGPU::G_UADDSAT;
2757 Y = B.buildInstr(AddOpc, {S32}, {Y, B.buildConstant(S32, 32)});
2758 Register DstReg = MI.getOperand(0).getReg();
2759 B.buildUMin(DstReg, X, Y);
2760 MI.eraseFromParent();
2761 return;
2762 }
2763 case AMDGPU::G_SEXT:
2764 case AMDGPU::G_ZEXT:
2765 case AMDGPU::G_ANYEXT: {
2766 Register SrcReg = MI.getOperand(1).getReg();
2767 LLT SrcTy = MRI.getType(SrcReg);
2768 const bool Signed = Opc == AMDGPU::G_SEXT;
2769
2770 assert(OpdMapper.getVRegs(1).empty());
2771
2772 const RegisterBank *SrcBank =
2773 OpdMapper.getInstrMapping().getOperandMapping(1).BreakDown[0].RegBank;
2774
2775 Register DstReg = MI.getOperand(0).getReg();
2776 LLT DstTy = MRI.getType(DstReg);
2777 if (DstTy.isScalar() &&
2778 SrcBank != &AMDGPU::SGPRRegBank &&
2779 SrcBank != &AMDGPU::VCCRegBank &&
2780 // FIXME: Should handle any type that round to s64 when irregular
2781 // breakdowns supported.
2782 DstTy.getSizeInBits() == 64 &&
2783 SrcTy.getSizeInBits() <= 32) {
2784 SmallVector<Register, 2> DefRegs(OpdMapper.getVRegs(0));
2785
2786 // Extend to 32-bit, and then extend the low half.
2787 if (Signed) {
2788 // TODO: Should really be buildSExtOrCopy
2789 B.buildSExtOrTrunc(DefRegs[0], SrcReg);
2790 } else if (Opc == AMDGPU::G_ZEXT) {
2791 B.buildZExtOrTrunc(DefRegs[0], SrcReg);
2792 } else {
2793 B.buildAnyExtOrTrunc(DefRegs[0], SrcReg);
2794 }
2795
2796 extendLow32IntoHigh32(B, DefRegs[1], DefRegs[0], Opc, *SrcBank);
2797 MRI.setRegBank(DstReg, *SrcBank);
2798 MI.eraseFromParent();
2799 return;
2800 }
2801
2802 if (SrcTy != LLT::scalar(1))
2803 return;
2804
2805 // It is not legal to have a legalization artifact with a VCC source. Rather
2806 // than introducing a copy, insert the select we would have to select the
2807 // copy to.
2808 if (SrcBank == &AMDGPU::VCCRegBank) {
2809 SmallVector<Register, 2> DefRegs(OpdMapper.getVRegs(0));
2810
2811 const RegisterBank *DstBank = &AMDGPU::VGPRRegBank;
2812
2813 unsigned DstSize = DstTy.getSizeInBits();
2814 // 64-bit select is SGPR only
2815 const bool UseSel64 = DstSize > 32 &&
2816 SrcBank->getID() == AMDGPU::SGPRRegBankID;
2817
2818 // TODO: Should s16 select be legal?
2819 LLT SelType = UseSel64 ? LLT::scalar(64) : LLT::scalar(32);
2820 auto True = B.buildConstant(SelType, Signed ? -1 : 1);
2821 auto False = B.buildConstant(SelType, 0);
2822
2823 MRI.setRegBank(True.getReg(0), *DstBank);
2824 MRI.setRegBank(False.getReg(0), *DstBank);
2825 MRI.setRegBank(DstReg, *DstBank);
2826
2827 if (DstSize > 32) {
2828 B.buildSelect(DefRegs[0], SrcReg, True, False);
2829 extendLow32IntoHigh32(B, DefRegs[1], DefRegs[0], Opc, *SrcBank, true);
2830 } else if (DstSize < 32) {
2831 auto Sel = B.buildSelect(SelType, SrcReg, True, False);
2832 MRI.setRegBank(Sel.getReg(0), *DstBank);
2833 B.buildTrunc(DstReg, Sel);
2834 } else {
2835 B.buildSelect(DstReg, SrcReg, True, False);
2836 }
2837
2838 MI.eraseFromParent();
2839 return;
2840 }
2841
2842 break;
2843 }
2844 case AMDGPU::G_EXTRACT_VECTOR_ELT: {
2845 SmallVector<Register, 2> DstRegs(OpdMapper.getVRegs(0));
2846
2847 assert(OpdMapper.getVRegs(1).empty() && OpdMapper.getVRegs(2).empty());
2848
2849 Register DstReg = MI.getOperand(0).getReg();
2850 Register SrcReg = MI.getOperand(1).getReg();
2851
2852 const LLT S32 = LLT::scalar(32);
2853 LLT DstTy = MRI.getType(DstReg);
2854 LLT SrcTy = MRI.getType(SrcReg);
2855
2856 if (foldExtractEltToCmpSelect(B, MI, OpdMapper))
2857 return;
2858
2859 const ValueMapping &DstMapping
2860 = OpdMapper.getInstrMapping().getOperandMapping(0);
2861 const RegisterBank *DstBank = DstMapping.BreakDown[0].RegBank;
2862 const RegisterBank *SrcBank =
2863 OpdMapper.getInstrMapping().getOperandMapping(1).BreakDown[0].RegBank;
2864 const RegisterBank *IdxBank =
2865 OpdMapper.getInstrMapping().getOperandMapping(2).BreakDown[0].RegBank;
2866
2867 Register BaseIdxReg;
2868 unsigned ConstOffset;
2869 std::tie(BaseIdxReg, ConstOffset) =
2870 AMDGPU::getBaseWithConstantOffset(MRI, MI.getOperand(2).getReg());
2871
2872 // See if the index is an add of a constant which will be foldable by moving
2873 // the base register of the index later if this is going to be executed in a
2874 // waterfall loop. This is essentially to reassociate the add of a constant
2875 // with the readfirstlane.
2876 bool ShouldMoveIndexIntoLoop = IdxBank != &AMDGPU::SGPRRegBank &&
2877 ConstOffset > 0 &&
2878 ConstOffset < SrcTy.getNumElements();
2879
2880 // Move the base register. We'll re-insert the add later.
2881 if (ShouldMoveIndexIntoLoop)
2882 MI.getOperand(2).setReg(BaseIdxReg);
2883
2884 // If this is a VGPR result only because the index was a VGPR result, the
2885 // actual indexing will be done on the SGPR source vector, which will
2886 // produce a scalar result. We need to copy to the VGPR result inside the
2887 // waterfall loop.
2888 const bool NeedCopyToVGPR = DstBank == &AMDGPU::VGPRRegBank &&
2889 SrcBank == &AMDGPU::SGPRRegBank;
2890 if (DstRegs.empty()) {
2891 applyDefaultMapping(OpdMapper);
2892
2894
2895 if (NeedCopyToVGPR) {
2896 // We don't want a phi for this temporary reg.
2897 Register TmpReg = MRI.createGenericVirtualRegister(DstTy);
2898 MRI.setRegBank(TmpReg, AMDGPU::SGPRRegBank);
2899 MI.getOperand(0).setReg(TmpReg);
2900 B.setInsertPt(*MI.getParent(), ++MI.getIterator());
2901
2902 // Use a v_mov_b32 here to make the exec dependency explicit.
2903 buildVCopy(B, DstReg, TmpReg);
2904 }
2905
2906 // Re-insert the constant offset add inside the waterfall loop.
2907 if (ShouldMoveIndexIntoLoop)
2908 reinsertVectorIndexAdd(B, MI, 2, ConstOffset);
2909
2910 return;
2911 }
2912
2913 assert(DstTy.getSizeInBits() == 64);
2914
2915 LLT Vec32 = LLT::fixed_vector(2 * SrcTy.getNumElements(), 32);
2916
2917 auto CastSrc = B.buildBitcast(Vec32, SrcReg);
2918 auto One = B.buildConstant(S32, 1);
2919
2920 MachineBasicBlock::iterator MII = MI.getIterator();
2921
2922 // Split the vector index into 32-bit pieces. Prepare to move all of the
2923 // new instructions into a waterfall loop if necessary.
2924 //
2925 // Don't put the bitcast or constant in the loop.
2926 MachineInstrSpan Span(MII, &B.getMBB());
2927
2928 // Compute 32-bit element indices, (2 * OrigIdx, 2 * OrigIdx + 1).
2929 auto IdxLo = B.buildShl(S32, BaseIdxReg, One);
2930 auto IdxHi = B.buildAdd(S32, IdxLo, One);
2931
2932 auto Extract0 = B.buildExtractVectorElement(DstRegs[0], CastSrc, IdxLo);
2933 auto Extract1 = B.buildExtractVectorElement(DstRegs[1], CastSrc, IdxHi);
2934
2935 MRI.setRegBank(DstReg, *DstBank);
2936 MRI.setRegBank(CastSrc.getReg(0), *SrcBank);
2937 MRI.setRegBank(One.getReg(0), AMDGPU::SGPRRegBank);
2938 MRI.setRegBank(IdxLo.getReg(0), AMDGPU::SGPRRegBank);
2939 MRI.setRegBank(IdxHi.getReg(0), AMDGPU::SGPRRegBank);
2940
2941 SmallSet<Register, 4> OpsToWaterfall;
2942 if (!collectWaterfallOperands(OpsToWaterfall, MI, MRI, { 2 })) {
2943 MI.eraseFromParent();
2944 return;
2945 }
2946
2947 // Remove the original instruction to avoid potentially confusing the
2948 // waterfall loop logic.
2949 B.setInstr(*Span.begin());
2950 MI.eraseFromParent();
2951 executeInWaterfallLoop(B, make_range(Span.begin(), Span.end()),
2952 OpsToWaterfall);
2953
2954 if (NeedCopyToVGPR) {
2955 MachineBasicBlock *LoopBB = Extract1->getParent();
2958 MRI.setRegBank(TmpReg0, AMDGPU::SGPRRegBank);
2959 MRI.setRegBank(TmpReg1, AMDGPU::SGPRRegBank);
2960
2961 Extract0->getOperand(0).setReg(TmpReg0);
2962 Extract1->getOperand(0).setReg(TmpReg1);
2963
2964 B.setInsertPt(*LoopBB, ++Extract1->getIterator());
2965
2966 buildVCopy(B, DstRegs[0], TmpReg0);
2967 buildVCopy(B, DstRegs[1], TmpReg1);
2968 }
2969
2970 if (ShouldMoveIndexIntoLoop)
2971 reinsertVectorIndexAdd(B, *IdxLo, 1, ConstOffset);
2972
2973 return;
2974 }
2975 case AMDGPU::G_INSERT_VECTOR_ELT: {
2976 SmallVector<Register, 2> InsRegs(OpdMapper.getVRegs(2));
2977
2978 Register DstReg = MI.getOperand(0).getReg();
2979 LLT VecTy = MRI.getType(DstReg);
2980
2981 assert(OpdMapper.getVRegs(0).empty());
2982 assert(OpdMapper.getVRegs(3).empty());
2983
2984 if (substituteSimpleCopyRegs(OpdMapper, 1))
2985 MRI.setType(MI.getOperand(1).getReg(), VecTy);
2986
2987 if (foldInsertEltToCmpSelect(B, MI, OpdMapper))
2988 return;
2989
2990 const RegisterBank *IdxBank =
2991 OpdMapper.getInstrMapping().getOperandMapping(3).BreakDown[0].RegBank;
2992
2993 Register SrcReg = MI.getOperand(1).getReg();
2994 Register InsReg = MI.getOperand(2).getReg();
2995 LLT InsTy = MRI.getType(InsReg);
2996 (void)InsTy;
2997
2998 Register BaseIdxReg;
2999 unsigned ConstOffset;
3000 std::tie(BaseIdxReg, ConstOffset) =
3001 AMDGPU::getBaseWithConstantOffset(MRI, MI.getOperand(3).getReg());
3002
3003 // See if the index is an add of a constant which will be foldable by moving
3004 // the base register of the index later if this is going to be executed in a
3005 // waterfall loop. This is essentially to reassociate the add of a constant
3006 // with the readfirstlane.
3007 bool ShouldMoveIndexIntoLoop = IdxBank != &AMDGPU::SGPRRegBank &&
3008 ConstOffset > 0 &&
3009 ConstOffset < VecTy.getNumElements();
3010
3011 // Move the base register. We'll re-insert the add later.
3012 if (ShouldMoveIndexIntoLoop)
3013 MI.getOperand(3).setReg(BaseIdxReg);
3014
3015
3016 if (InsRegs.empty()) {
3018
3019 // Re-insert the constant offset add inside the waterfall loop.
3020 if (ShouldMoveIndexIntoLoop) {
3021 reinsertVectorIndexAdd(B, MI, 3, ConstOffset);
3022 }
3023
3024 return;
3025 }
3026
3027 assert(InsTy.getSizeInBits() == 64);
3028
3029 const LLT S32 = LLT::scalar(32);
3030 LLT Vec32 = LLT::fixed_vector(2 * VecTy.getNumElements(), 32);
3031
3032 auto CastSrc = B.buildBitcast(Vec32, SrcReg);
3033 auto One = B.buildConstant(S32, 1);
3034
3035 // Split the vector index into 32-bit pieces. Prepare to move all of the
3036 // new instructions into a waterfall loop if necessary.
3037 //
3038 // Don't put the bitcast or constant in the loop.
3040
3041 // Compute 32-bit element indices, (2 * OrigIdx, 2 * OrigIdx + 1).
3042 auto IdxLo = B.buildShl(S32, BaseIdxReg, One);
3043 auto IdxHi = B.buildAdd(S32, IdxLo, One);
3044
3045 auto InsLo = B.buildInsertVectorElement(Vec32, CastSrc, InsRegs[0], IdxLo);
3046 auto InsHi = B.buildInsertVectorElement(Vec32, InsLo, InsRegs[1], IdxHi);
3047
3048 const RegisterBank *DstBank =
3049 OpdMapper.getInstrMapping().getOperandMapping(0).BreakDown[0].RegBank;
3050 const RegisterBank *SrcBank =
3051 OpdMapper.getInstrMapping().getOperandMapping(1).BreakDown[0].RegBank;
3052 const RegisterBank *InsSrcBank =
3053 OpdMapper.getInstrMapping().getOperandMapping(2).BreakDown[0].RegBank;
3054
3055 MRI.setRegBank(InsReg, *InsSrcBank);
3056 MRI.setRegBank(CastSrc.getReg(0), *SrcBank);
3057 MRI.setRegBank(InsLo.getReg(0), *DstBank);
3058 MRI.setRegBank(InsHi.getReg(0), *DstBank);
3059 MRI.setRegBank(One.getReg(0), AMDGPU::SGPRRegBank);
3060 MRI.setRegBank(IdxLo.getReg(0), AMDGPU::SGPRRegBank);
3061 MRI.setRegBank(IdxHi.getReg(0), AMDGPU::SGPRRegBank);
3062
3063
3064 SmallSet<Register, 4> OpsToWaterfall;
3065 if (!collectWaterfallOperands(OpsToWaterfall, MI, MRI, { 3 })) {
3066 B.setInsertPt(B.getMBB(), MI);
3067 B.buildBitcast(DstReg, InsHi);
3068 MI.eraseFromParent();
3069 return;
3070 }
3071
3072 B.setInstr(*Span.begin());
3073 MI.eraseFromParent();
3074
3075 // Figure out the point after the waterfall loop before mangling the control
3076 // flow.
3077 executeInWaterfallLoop(B, make_range(Span.begin(), Span.end()),
3078 OpsToWaterfall);
3079
3080 // The insertion point is now right after the original instruction.
3081 //
3082 // Keep the bitcast to the original vector type out of the loop. Doing this
3083 // saved an extra phi we don't need inside the loop.
3084 B.buildBitcast(DstReg, InsHi);
3085
3086 // Re-insert the constant offset add inside the waterfall loop.
3087 if (ShouldMoveIndexIntoLoop)
3088 reinsertVectorIndexAdd(B, *IdxLo, 1, ConstOffset);
3089
3090 return;
3091 }
3092 case AMDGPU::G_AMDGPU_BUFFER_LOAD:
3093 case AMDGPU::G_AMDGPU_BUFFER_LOAD_USHORT:
3094 case AMDGPU::G_AMDGPU_BUFFER_LOAD_SSHORT:
3095 case AMDGPU::G_AMDGPU_BUFFER_LOAD_UBYTE:
3096 case AMDGPU::G_AMDGPU_BUFFER_LOAD_SBYTE:
3097 case AMDGPU::G_AMDGPU_BUFFER_LOAD_TFE:
3098 case AMDGPU::G_AMDGPU_BUFFER_LOAD_USHORT_TFE:
3099 case AMDGPU::G_AMDGPU_BUFFER_LOAD_SSHORT_TFE:
3100 case AMDGPU::G_AMDGPU_BUFFER_LOAD_UBYTE_TFE:
3101 case AMDGPU::G_AMDGPU_BUFFER_LOAD_SBYTE_TFE:
3102 case AMDGPU::G_AMDGPU_BUFFER_LOAD_FORMAT:
3103 case AMDGPU::G_AMDGPU_BUFFER_LOAD_FORMAT_TFE:
3104 case AMDGPU::G_AMDGPU_BUFFER_LOAD_FORMAT_D16:
3105 case AMDGPU::G_AMDGPU_TBUFFER_LOAD_FORMAT:
3106 case AMDGPU::G_AMDGPU_TBUFFER_LOAD_FORMAT_D16:
3107 case AMDGPU::G_AMDGPU_BUFFER_STORE:
3108 case AMDGPU::G_AMDGPU_BUFFER_STORE_BYTE:
3109 case AMDGPU::G_AMDGPU_BUFFER_STORE_SHORT:
3110 case AMDGPU::G_AMDGPU_BUFFER_STORE_FORMAT:
3111 case AMDGPU::G_AMDGPU_BUFFER_STORE_FORMAT_D16:
3112 case AMDGPU::G_AMDGPU_TBUFFER_STORE_FORMAT:
3113 case AMDGPU::G_AMDGPU_TBUFFER_STORE_FORMAT_D16: {
3114 applyDefaultMapping(OpdMapper);
3115 executeInWaterfallLoop(B, MI, {1, 4});
3116 return;
3117 }
3118 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SWAP:
3119 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_ADD:
3120 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SUB:
3121 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SMIN:
3122 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_UMIN:
3123 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SMAX:
3124 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_UMAX:
3125 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_AND:
3126 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_OR:
3127 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_XOR:
3128 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_INC:
3129 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_DEC:
3130 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SUB_CLAMP_U32:
3131 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_COND_SUB_U32:
3132 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_FADD:
3133 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_FMIN:
3134 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_FMAX: {
3135 applyDefaultMapping(OpdMapper);
3136 executeInWaterfallLoop(B, MI, {2, 5});
3137 return;
3138 }
3139 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_CMPSWAP: {
3140 applyDefaultMapping(OpdMapper);
3141 executeInWaterfallLoop(B, MI, {3, 6});
3142 return;
3143 }
3144 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD:
3145 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_UBYTE:
3146 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_SBYTE:
3147 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_USHORT:
3148 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_SSHORT: {
3149 applyMappingSBufferLoad(B, OpdMapper);
3150 return;
3151 }
3152 case AMDGPU::G_AMDGPU_S_BUFFER_PREFETCH:
3155 return;
3156 case AMDGPU::G_INTRINSIC:
3157 case AMDGPU::G_INTRINSIC_CONVERGENT: {
3158 switch (cast<GIntrinsic>(MI).getIntrinsicID()) {
3159 case Intrinsic::amdgcn_readlane: {
3160 substituteSimpleCopyRegs(OpdMapper, 2);
3161
3162 assert(OpdMapper.getVRegs(0).empty());
3163 assert(OpdMapper.getVRegs(3).empty());
3164
3165 // Make sure the index is an SGPR. It doesn't make sense to run this in a
3166 // waterfall loop, so assume it's a uniform value.
3167 constrainOpWithReadfirstlane(B, MI, 3); // Index
3168 return;
3169 }
3170 case Intrinsic::amdgcn_writelane: {
3171 assert(OpdMapper.getVRegs(0).empty());
3172 assert(OpdMapper.getVRegs(2).empty());
3173 assert(OpdMapper.getVRegs(3).empty());
3174
3175 substituteSimpleCopyRegs(OpdMapper, 4); // VGPR input val
3176 constrainOpWithReadfirstlane(B, MI, 2); // Source value
3177 constrainOpWithReadfirstlane(B, MI, 3); // Index
3178 return;
3179 }
3180 case Intrinsic::amdgcn_interp_p1:
3181 case Intrinsic::amdgcn_interp_p2:
3182 case Intrinsic::amdgcn_interp_mov:
3183 case Intrinsic::amdgcn_interp_p1_f16:
3184 case Intrinsic::amdgcn_interp_p2_f16:
3185 case Intrinsic::amdgcn_lds_param_load: {
3186 applyDefaultMapping(OpdMapper);
3187
3188 // Readlane for m0 value, which is always the last operand.
3189 // FIXME: Should this be a waterfall loop instead?
3190 constrainOpWithReadfirstlane(B, MI, MI.getNumOperands() - 1); // Index
3191 return;
3192 }
3193 case Intrinsic::amdgcn_interp_inreg_p10:
3194 case Intrinsic::amdgcn_interp_inreg_p2:
3195 case Intrinsic::amdgcn_interp_inreg_p10_f16:
3196 case Intrinsic::amdgcn_interp_inreg_p2_f16:
3197 case Intrinsic::amdgcn_interp_p10_rtz_f16:
3198 case Intrinsic::amdgcn_interp_p2_rtz_f16:
3199 case Intrinsic::amdgcn_permlane16_swap:
3200 case Intrinsic::amdgcn_permlane32_swap:
3201 applyDefaultMapping(OpdMapper);
3202 return;
3203 case Intrinsic::amdgcn_permlane16:
3204 case Intrinsic::amdgcn_permlanex16: {
3205 // Doing a waterfall loop over these wouldn't make any sense.
3206 substituteSimpleCopyRegs(OpdMapper, 2);
3207 substituteSimpleCopyRegs(OpdMapper, 3);
3210 return;
3211 }
3212 case Intrinsic::amdgcn_permlane_bcast:
3213 case Intrinsic::amdgcn_permlane_up:
3214 case Intrinsic::amdgcn_permlane_down:
3215 case Intrinsic::amdgcn_permlane_xor:
3216 // Doing a waterfall loop over these wouldn't make any sense.
3219 return;
3220 case Intrinsic::amdgcn_permlane_idx_gen: {
3222 return;
3223 }
3224 case Intrinsic::amdgcn_sbfe:
3225 applyMappingBFE(B, OpdMapper, true);
3226 return;
3227 case Intrinsic::amdgcn_ubfe:
3228 applyMappingBFE(B, OpdMapper, false);
3229 return;
3230 case Intrinsic::amdgcn_inverse_ballot:
3231 case Intrinsic::amdgcn_s_bitreplicate:
3232 case Intrinsic::amdgcn_s_quadmask:
3233 case Intrinsic::amdgcn_s_wqm:
3234 applyDefaultMapping(OpdMapper);
3235 constrainOpWithReadfirstlane(B, MI, 2); // Mask
3236 return;
3237 case Intrinsic::amdgcn_ballot:
3238 // Use default handling and insert copy to vcc source.
3239 break;
3240 }
3241 break;
3242 }
3243 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_LOAD:
3244 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_LOAD_D16:
3245 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_LOAD_NORET:
3246 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_STORE:
3247 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_STORE_D16: {
3248 const AMDGPU::RsrcIntrinsic *RSrcIntrin =
3250 assert(RSrcIntrin && RSrcIntrin->IsImage);
3251 // Non-images can have complications from operands that allow both SGPR
3252 // and VGPR. For now it's too complicated to figure out the final opcode
3253 // to derive the register bank from the MCInstrDesc.
3254 applyMappingImage(B, MI, OpdMapper, RSrcIntrin->RsrcArg);
3255 return;
3256 }
3257 case AMDGPU::G_AMDGPU_BVH_INTERSECT_RAY:
3258 case AMDGPU::G_AMDGPU_BVH8_INTERSECT_RAY:
3259 case AMDGPU::G_AMDGPU_BVH_DUAL_INTERSECT_RAY: {
3260 bool IsDualOrBVH8 =
3261 MI.getOpcode() == AMDGPU::G_AMDGPU_BVH_DUAL_INTERSECT_RAY ||
3262 MI.getOpcode() == AMDGPU::G_AMDGPU_BVH8_INTERSECT_RAY;
3263 unsigned NumMods = IsDualOrBVH8 ? 0 : 1; // Has A16 modifier
3264 unsigned LastRegOpIdx = MI.getNumExplicitOperands() - 1 - NumMods;
3265 applyDefaultMapping(OpdMapper);
3266 executeInWaterfallLoop(B, MI, {LastRegOpIdx});
3267 return;
3268 }
3269 case AMDGPU::G_INTRINSIC_W_SIDE_EFFECTS:
3270 case AMDGPU::G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS: {
3271 auto IntrID = cast<GIntrinsic>(MI).getIntrinsicID();
3272 switch (IntrID) {
3273 case Intrinsic::amdgcn_ds_ordered_add:
3274 case Intrinsic::amdgcn_ds_ordered_swap: {
3275 // This is only allowed to execute with 1 lane, so readfirstlane is safe.
3276 assert(OpdMapper.getVRegs(0).empty());
3277 substituteSimpleCopyRegs(OpdMapper, 3);
3279 return;
3280 }
3281 case Intrinsic::amdgcn_ds_gws_init:
3282 case Intrinsic::amdgcn_ds_gws_barrier:
3283 case Intrinsic::amdgcn_ds_gws_sema_br: {
3284 // Only the first lane is executes, so readfirstlane is safe.
3285 substituteSimpleCopyRegs(OpdMapper, 1);
3287 return;
3288 }
3289 case Intrinsic::amdgcn_ds_gws_sema_v:
3290 case Intrinsic::amdgcn_ds_gws_sema_p:
3291 case Intrinsic::amdgcn_ds_gws_sema_release_all: {
3292 // Only the first lane is executes, so readfirstlane is safe.
3294 return;
3295 }
3296 case Intrinsic::amdgcn_ds_append:
3297 case Intrinsic::amdgcn_ds_consume: {
3299 return;
3300 }
3301 case Intrinsic::amdgcn_s_alloc_vgpr:
3303 return;
3304 case Intrinsic::amdgcn_s_sendmsg:
3305 case Intrinsic::amdgcn_s_sendmsghalt: {
3306 // FIXME: Should this use a waterfall loop?
3308 return;
3309 }
3310 case Intrinsic::amdgcn_s_setreg: {
3312 return;
3313 }
3314 case Intrinsic::amdgcn_s_ttracedata:
3316 return;
3317 case Intrinsic::amdgcn_raw_buffer_load_lds:
3318 case Intrinsic::amdgcn_raw_buffer_load_async_lds:
3319 case Intrinsic::amdgcn_raw_ptr_buffer_load_lds:
3320 case Intrinsic::amdgcn_raw_ptr_buffer_load_async_lds: {
3321 applyDefaultMapping(OpdMapper);
3322 constrainOpWithReadfirstlane(B, MI, 1); // rsrc
3324 constrainOpWithReadfirstlane(B, MI, 5); // soffset
3325 return;
3326 }
3327 case Intrinsic::amdgcn_struct_buffer_load_lds:
3328 case Intrinsic::amdgcn_struct_buffer_load_async_lds:
3329 case Intrinsic::amdgcn_struct_ptr_buffer_load_lds:
3330 case Intrinsic::amdgcn_struct_ptr_buffer_load_async_lds: {
3331 applyDefaultMapping(OpdMapper);
3332 constrainOpWithReadfirstlane(B, MI, 1); // rsrc
3334 constrainOpWithReadfirstlane(B, MI, 6); // soffset
3335 return;
3336 }
3337 case Intrinsic::amdgcn_cluster_load_async_to_lds_b8:
3338 case Intrinsic::amdgcn_cluster_load_async_to_lds_b32:
3339 case Intrinsic::amdgcn_cluster_load_async_to_lds_b64:
3340 case Intrinsic::amdgcn_cluster_load_async_to_lds_b128: {
3341 applyDefaultMapping(OpdMapper);
3343 return;
3344 }
3345 case Intrinsic::amdgcn_load_to_lds:
3346 case Intrinsic::amdgcn_load_async_to_lds:
3347 case Intrinsic::amdgcn_global_load_lds:
3348 case Intrinsic::amdgcn_global_load_async_lds: {
3349 applyDefaultMapping(OpdMapper);
3351 return;
3352 }
3353 case Intrinsic::amdgcn_lds_direct_load: {
3354 applyDefaultMapping(OpdMapper);
3355 // Readlane for m0 value, which is always the last operand.
3356 constrainOpWithReadfirstlane(B, MI, MI.getNumOperands() - 1); // Index
3357 return;
3358 }
3359 case Intrinsic::amdgcn_exp_row:
3360 applyDefaultMapping(OpdMapper);
3362 return;
3363 case Intrinsic::amdgcn_cluster_load_b32:
3364 case Intrinsic::amdgcn_cluster_load_b64:
3365 case Intrinsic::amdgcn_cluster_load_b128: {
3366 applyDefaultMapping(OpdMapper);
3368 return;
3369 }
3370 case Intrinsic::amdgcn_s_sleep_var:
3371 assert(OpdMapper.getVRegs(1).empty());
3373 return;
3374 case Intrinsic::amdgcn_s_barrier_join:
3375 case Intrinsic::amdgcn_s_wakeup_barrier:
3377 return;
3378 case Intrinsic::amdgcn_s_barrier_init:
3379 case Intrinsic::amdgcn_s_barrier_signal_var:
3382 return;
3383 case Intrinsic::amdgcn_s_get_barrier_state:
3384 case Intrinsic::amdgcn_s_get_named_barrier_state: {
3386 return;
3387 }
3388 case Intrinsic::amdgcn_s_prefetch_data:
3389 case Intrinsic::amdgcn_s_prefetch_inst: {
3390 Register PtrReg = MI.getOperand(1).getReg();
3391 unsigned AS = MRI.getType(PtrReg).getAddressSpace();
3395 } else
3396 MI.eraseFromParent();
3397 return;
3398 }
3399 case Intrinsic::amdgcn_tensor_load_to_lds:
3400 case Intrinsic::amdgcn_tensor_store_from_lds: {
3406 return;
3407 }
3408 default: {
3409 if (const AMDGPU::RsrcIntrinsic *RSrcIntrin =
3411 // Non-images can have complications from operands that allow both SGPR
3412 // and VGPR. For now it's too complicated to figure out the final opcode
3413 // to derive the register bank from the MCInstrDesc.
3414 if (RSrcIntrin->IsImage) {
3415 applyMappingImage(B, MI, OpdMapper, RSrcIntrin->RsrcArg);
3416 return;
3417 }
3418 }
3419
3420 break;
3421 }
3422 }
3423 break;
3424 }
3425 case AMDGPU::G_SI_CALL: {
3426 // Use a set to avoid extra readfirstlanes in the case where multiple
3427 // operands are the same register.
3428 SmallSet<Register, 4> SGPROperandRegs;
3429
3430 if (!collectWaterfallOperands(SGPROperandRegs, MI, MRI, {1}))
3431 break;
3432
3433 // Move all copies to physical SGPRs that are used by the call instruction
3434 // into the loop block. Start searching for these copies until the
3435 // ADJCALLSTACKUP.
3436 unsigned FrameSetupOpcode = AMDGPU::ADJCALLSTACKUP;
3437 unsigned FrameDestroyOpcode = AMDGPU::ADJCALLSTACKDOWN;
3438
3439 // Move all non-copies before the copies, so that a complete range can be
3440 // moved into the waterfall loop.
3441 SmallVector<MachineInstr *, 4> NonCopyInstrs;
3442 // Count of NonCopyInstrs found until the current LastCopy.
3443 unsigned NonCopyInstrsLen = 0;
3445 MachineBasicBlock::iterator LastCopy = Start;
3446 MachineBasicBlock *MBB = MI.getParent();
3447 const SIMachineFunctionInfo *Info =
3448 MBB->getParent()->getInfo<SIMachineFunctionInfo>();
3449 while (Start->getOpcode() != FrameSetupOpcode) {
3450 --Start;
3451 bool IsCopy = false;
3452 if (Start->getOpcode() == AMDGPU::COPY) {
3453 auto &Dst = Start->getOperand(0);
3454 if (Dst.isReg()) {
3455 Register Reg = Dst.getReg();
3456 if (Reg.isPhysical() && MI.readsRegister(Reg, TRI)) {
3457 IsCopy = true;
3458 } else {
3459 // Also move the copy from the scratch rsrc descriptor into the loop
3460 // to allow it to be optimized away.
3461 auto &Src = Start->getOperand(1);
3462 if (Src.isReg()) {
3463 Reg = Src.getReg();
3464 IsCopy = Info->getScratchRSrcReg() == Reg;
3465 }
3466 }
3467 }
3468 }
3469
3470 if (IsCopy) {
3471 LastCopy = Start;
3472 NonCopyInstrsLen = NonCopyInstrs.size();
3473 } else {
3474 NonCopyInstrs.push_back(&*Start);
3475 }
3476 }
3477 NonCopyInstrs.resize(NonCopyInstrsLen);
3478
3479 for (auto *NonCopy : reverse(NonCopyInstrs)) {
3480 MBB->splice(LastCopy, MBB, NonCopy->getIterator());
3481 }
3482 Start = LastCopy;
3483
3484 // Do the same for copies after the loop
3485 NonCopyInstrs.clear();
3486 NonCopyInstrsLen = 0;
3488 LastCopy = End;
3489 while (End->getOpcode() != FrameDestroyOpcode) {
3490 ++End;
3491 bool IsCopy = false;
3492 if (End->getOpcode() == AMDGPU::COPY) {
3493 auto &Src = End->getOperand(1);
3494 if (Src.isReg()) {
3495 Register Reg = Src.getReg();
3496 IsCopy = Reg.isPhysical() && MI.modifiesRegister(Reg, TRI);
3497 }
3498 }
3499
3500 if (IsCopy) {
3501 LastCopy = End;
3502 NonCopyInstrsLen = NonCopyInstrs.size();
3503 } else {
3504 NonCopyInstrs.push_back(&*End);
3505 }
3506 }
3507 NonCopyInstrs.resize(NonCopyInstrsLen);
3508
3509 End = LastCopy;
3510 ++LastCopy;
3511 for (auto *NonCopy : reverse(NonCopyInstrs)) {
3512 MBB->splice(LastCopy, MBB, NonCopy->getIterator());
3513 }
3514
3515 ++End;
3516 B.setInsertPt(B.getMBB(), Start);
3517 executeInWaterfallLoop(B, make_range(Start, End), SGPROperandRegs);
3518 break;
3519 }
3520 case AMDGPU::G_AMDGPU_FLAT_LOAD_MONITOR:
3521 case AMDGPU::G_AMDGPU_GLOBAL_LOAD_MONITOR:
3522 case AMDGPU::G_LOAD:
3523 case AMDGPU::G_ZEXTLOAD:
3524 case AMDGPU::G_SEXTLOAD: {
3525 if (applyMappingLoad(B, OpdMapper, MI))
3526 return;
3527 break;
3528 }
3529 case AMDGPU::G_DYN_STACKALLOC:
3530 applyMappingDynStackAlloc(B, OpdMapper, MI);
3531 return;
3532 case AMDGPU::G_STACKRESTORE: {
3533 applyDefaultMapping(OpdMapper);
3535 return;
3536 }
3537 case AMDGPU::G_SBFX:
3538 applyMappingBFE(B, OpdMapper, /*Signed*/ true);
3539 return;
3540 case AMDGPU::G_UBFX:
3541 applyMappingBFE(B, OpdMapper, /*Signed*/ false);
3542 return;
3543 case AMDGPU::G_AMDGPU_MAD_U64_U32:
3544 case AMDGPU::G_AMDGPU_MAD_I64_I32:
3545 applyMappingMAD_64_32(B, OpdMapper);
3546 return;
3547 case AMDGPU::G_PREFETCH: {
3548 if (!Subtarget.hasSafeSmemPrefetch() && !Subtarget.hasVmemPrefInsts()) {
3549 MI.eraseFromParent();
3550 return;
3551 }
3552 Register PtrReg = MI.getOperand(0).getReg();
3553 unsigned PtrBank = getRegBankID(PtrReg, MRI, AMDGPU::SGPRRegBankID);
3554 if (PtrBank == AMDGPU::VGPRRegBankID &&
3555 (!Subtarget.hasVmemPrefInsts() || !MI.getOperand(3).getImm())) {
3556 // Cannot do I$ prefetch with divergent pointer.
3557 MI.eraseFromParent();
3558 return;
3559 }
3560 unsigned AS = MRI.getType(PtrReg).getAddressSpace();
3563 (!Subtarget.hasSafeSmemPrefetch() &&
3565 !MI.getOperand(3).getImm() /* I$ prefetch */))) {
3566 MI.eraseFromParent();
3567 return;
3568 }
3569 applyDefaultMapping(OpdMapper);
3570 return;
3571 }
3572 default:
3573 break;
3574 }
3575
3576 return applyDefaultMapping(OpdMapper);
3577}
3578
3579// vgpr, sgpr -> vgpr
3580// vgpr, agpr -> vgpr
3581// agpr, agpr -> agpr
3582// agpr, sgpr -> vgpr
3583static unsigned regBankUnion(unsigned RB0, unsigned RB1) {
3584 if (RB0 == AMDGPU::InvalidRegBankID)
3585 return RB1;
3586 if (RB1 == AMDGPU::InvalidRegBankID)
3587 return RB0;
3588
3589 if (RB0 == AMDGPU::SGPRRegBankID && RB1 == AMDGPU::SGPRRegBankID)
3590 return AMDGPU::SGPRRegBankID;
3591
3592 if (RB0 == AMDGPU::AGPRRegBankID && RB1 == AMDGPU::AGPRRegBankID)
3593 return AMDGPU::AGPRRegBankID;
3594
3595 return AMDGPU::VGPRRegBankID;
3596}
3597
3598static unsigned regBankBoolUnion(unsigned RB0, unsigned RB1) {
3599 if (RB0 == AMDGPU::InvalidRegBankID)
3600 return RB1;
3601 if (RB1 == AMDGPU::InvalidRegBankID)
3602 return RB0;
3603
3604 // vcc, vcc -> vcc
3605 // vcc, sgpr -> vcc
3606 // vcc, vgpr -> vcc
3607 if (RB0 == AMDGPU::VCCRegBankID || RB1 == AMDGPU::VCCRegBankID)
3608 return AMDGPU::VCCRegBankID;
3609
3610 // vcc, vgpr -> vgpr
3611 return regBankUnion(RB0, RB1);
3612}
3613
3615 const MachineInstr &MI) const {
3616 unsigned RegBank = AMDGPU::InvalidRegBankID;
3617
3618 for (const MachineOperand &MO : MI.operands()) {
3619 if (!MO.isReg())
3620 continue;
3621 Register Reg = MO.getReg();
3622 if (const RegisterBank *Bank = getRegBank(Reg, MRI, *TRI)) {
3623 RegBank = regBankUnion(RegBank, Bank->getID());
3624 if (RegBank == AMDGPU::VGPRRegBankID)
3625 break;
3626 }
3627 }
3628
3629 return RegBank;
3630}
3631
3633 const MachineFunction &MF = *MI.getMF();
3634 const MachineRegisterInfo &MRI = MF.getRegInfo();
3635 for (const MachineOperand &MO : MI.operands()) {
3636 if (!MO.isReg())
3637 continue;
3638 Register Reg = MO.getReg();
3639 if (const RegisterBank *Bank = getRegBank(Reg, MRI, *TRI)) {
3640 if (Bank->getID() != AMDGPU::SGPRRegBankID)
3641 return false;
3642 }
3643 }
3644 return true;
3645}
3646
3649 const MachineFunction &MF = *MI.getMF();
3650 const MachineRegisterInfo &MRI = MF.getRegInfo();
3651 SmallVector<const ValueMapping*, 8> OpdsMapping(MI.getNumOperands());
3652
3653 for (unsigned i = 0, e = MI.getNumOperands(); i != e; ++i) {
3654 const MachineOperand &SrcOp = MI.getOperand(i);
3655 if (!SrcOp.isReg())
3656 continue;
3657
3658 unsigned Size = getSizeInBits(SrcOp.getReg(), MRI, *TRI);
3659 OpdsMapping[i] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
3660 }
3661 return getInstructionMapping(1, 1, getOperandsMapping(OpdsMapping),
3662 MI.getNumOperands());
3663}
3664
3667 const MachineFunction &MF = *MI.getMF();
3668 const MachineRegisterInfo &MRI = MF.getRegInfo();
3669 SmallVector<const ValueMapping*, 8> OpdsMapping(MI.getNumOperands());
3670
3671 // Even though we technically could use SGPRs, this would require knowledge of
3672 // the constant bus restriction. Force all sources to VGPR (except for VCC).
3673 //
3674 // TODO: Unary ops are trivially OK, so accept SGPRs?
3675 for (unsigned i = 0, e = MI.getNumOperands(); i != e; ++i) {
3676 const MachineOperand &Src = MI.getOperand(i);
3677 if (!Src.isReg())
3678 continue;
3679
3680 unsigned Size = getSizeInBits(Src.getReg(), MRI, *TRI);
3681 unsigned BankID = Size == 1 ? AMDGPU::VCCRegBankID : AMDGPU::VGPRRegBankID;
3682 OpdsMapping[i] = AMDGPU::getValueMapping(BankID, Size);
3683 }
3684
3685 return getInstructionMapping(1, 1, getOperandsMapping(OpdsMapping),
3686 MI.getNumOperands());
3687}
3688
3691 const MachineFunction &MF = *MI.getMF();
3692 const MachineRegisterInfo &MRI = MF.getRegInfo();
3693 SmallVector<const ValueMapping*, 8> OpdsMapping(MI.getNumOperands());
3694
3695 for (unsigned I = 0, E = MI.getNumOperands(); I != E; ++I) {
3696 const MachineOperand &Op = MI.getOperand(I);
3697 if (!Op.isReg())
3698 continue;
3699
3700 unsigned Size = getSizeInBits(Op.getReg(), MRI, *TRI);
3701 OpdsMapping[I] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
3702 }
3703
3704 return getInstructionMapping(1, 1, getOperandsMapping(OpdsMapping),
3705 MI.getNumOperands());
3706}
3707
3710 const MachineInstr &MI,
3711 int RsrcIdx) const {
3712 // The reported argument index is relative to the IR intrinsic call arguments,
3713 // so we need to shift by the number of defs and the intrinsic ID.
3714 RsrcIdx += MI.getNumExplicitDefs() + 1;
3715
3716 const int NumOps = MI.getNumOperands();
3718
3719 // TODO: Should packed/unpacked D16 difference be reported here as part of
3720 // the value mapping?
3721 for (int I = 0; I != NumOps; ++I) {
3722 if (!MI.getOperand(I).isReg())
3723 continue;
3724
3725 Register OpReg = MI.getOperand(I).getReg();
3726 // We replace some dead address operands with $noreg
3727 if (!OpReg)
3728 continue;
3729
3730 unsigned Size = getSizeInBits(OpReg, MRI, *TRI);
3731
3732 // FIXME: Probably need a new intrinsic register bank searchable table to
3733 // handle arbitrary intrinsics easily.
3734 //
3735 // If this has a sampler, it immediately follows rsrc.
3736 const bool MustBeSGPR = I == RsrcIdx || I == RsrcIdx + 1;
3737
3738 if (MustBeSGPR) {
3739 // If this must be an SGPR, so we must report whatever it is as legal.
3740 unsigned NewBank = getRegBankID(OpReg, MRI, AMDGPU::SGPRRegBankID);
3741 OpdsMapping[I] = AMDGPU::getValueMapping(NewBank, Size);
3742 } else {
3743 // Some operands must be VGPR, and these are easy to copy to.
3744 OpdsMapping[I] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
3745 }
3746 }
3747
3748 return getInstructionMapping(1, 1, getOperandsMapping(OpdsMapping), NumOps);
3749}
3750
3751/// Return the mapping for a pointer argument.
3754 Register PtrReg) const {
3755 LLT PtrTy = MRI.getType(PtrReg);
3756 unsigned Size = PtrTy.getSizeInBits();
3757 if (Subtarget.useFlatForGlobal() ||
3759 return AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
3760
3761 // If we're using MUBUF instructions for global memory, an SGPR base register
3762 // is possible. Otherwise this needs to be a VGPR.
3763 const RegisterBank *PtrBank = getRegBank(PtrReg, MRI, *TRI);
3764 return AMDGPU::getValueMapping(PtrBank->getID(), Size);
3765}
3766
3769
3770 const MachineFunction &MF = *MI.getMF();
3771 const MachineRegisterInfo &MRI = MF.getRegInfo();
3773 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
3774 Register PtrReg = MI.getOperand(1).getReg();
3775 LLT PtrTy = MRI.getType(PtrReg);
3776 unsigned AS = PtrTy.getAddressSpace();
3777 unsigned PtrSize = PtrTy.getSizeInBits();
3778
3779 const ValueMapping *ValMapping;
3780 const ValueMapping *PtrMapping;
3781
3782 const RegisterBank *PtrBank = getRegBank(PtrReg, MRI, *TRI);
3783
3784 if (PtrBank == &AMDGPU::SGPRRegBank && AMDGPU::isFlatGlobalAddrSpace(AS)) {
3785 if (isScalarLoadLegal(MI)) {
3786 // We have a uniform instruction so we want to use an SMRD load
3787 ValMapping = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
3788 PtrMapping = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, PtrSize);
3789 } else {
3790 ValMapping = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
3791
3792 // If we're using MUBUF instructions for global memory, an SGPR base
3793 // register is possible. Otherwise this needs to be a VGPR.
3794 unsigned PtrBankID = Subtarget.useFlatForGlobal() ?
3795 AMDGPU::VGPRRegBankID : AMDGPU::SGPRRegBankID;
3796
3797 PtrMapping = AMDGPU::getValueMapping(PtrBankID, PtrSize);
3798 }
3799 } else {
3800 ValMapping = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
3801 PtrMapping = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, PtrSize);
3802 }
3803
3804 OpdsMapping[0] = ValMapping;
3805 OpdsMapping[1] = PtrMapping;
3807 1, 1, getOperandsMapping(OpdsMapping), MI.getNumOperands());
3808 return Mapping;
3809
3810 // FIXME: Do we want to add a mapping for FLAT load, or should we just
3811 // handle that during instruction selection?
3812}
3813
3814unsigned
3816 const MachineRegisterInfo &MRI,
3817 unsigned Default) const {
3818 const RegisterBank *Bank = getRegBank(Reg, MRI, *TRI);
3819 return Bank ? Bank->getID() : Default;
3820}
3821
3824 const MachineRegisterInfo &MRI,
3825 const TargetRegisterInfo &TRI) const {
3826 // Lie and claim anything is legal, even though this needs to be an SGPR
3827 // applyMapping will have to deal with it as a waterfall loop.
3828 unsigned Bank = getRegBankID(Reg, MRI, AMDGPU::SGPRRegBankID);
3829 unsigned Size = getSizeInBits(Reg, MRI, TRI);
3830 return AMDGPU::getValueMapping(Bank, Size);
3831}
3832
3835 const MachineRegisterInfo &MRI,
3836 const TargetRegisterInfo &TRI) const {
3837 unsigned Size = getSizeInBits(Reg, MRI, TRI);
3838 return AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
3839}
3840
3843 const MachineRegisterInfo &MRI,
3844 const TargetRegisterInfo &TRI) const {
3845 unsigned Size = getSizeInBits(Reg, MRI, TRI);
3846 return AMDGPU::getValueMapping(AMDGPU::AGPRRegBankID, Size);
3847}
3848
3849///
3850/// This function must return a legal mapping, because
3851/// AMDGPURegisterBankInfo::getInstrAlternativeMappings() is not called
3852/// in RegBankSelect::Mode::Fast. Any mapping that would cause a
3853/// VGPR to SGPR generated is illegal.
3854///
3855// Operands that must be SGPRs must accept potentially divergent VGPRs as
3856// legal. These will be dealt with in applyMappingImpl.
3857//
3860 const MachineFunction &MF = *MI.getMF();
3861 const MachineRegisterInfo &MRI = MF.getRegInfo();
3862
3863 if (MI.isCopy() || MI.getOpcode() == AMDGPU::G_FREEZE) {
3864 Register DstReg = MI.getOperand(0).getReg();
3865 Register SrcReg = MI.getOperand(1).getReg();
3866
3867 // The default logic bothers to analyze impossible alternative mappings. We
3868 // want the most straightforward mapping, so just directly handle this.
3869 const RegisterBank *DstBank = getRegBank(DstReg, MRI, *TRI);
3870 const RegisterBank *SrcBank = getRegBank(SrcReg, MRI, *TRI);
3871
3872 // For COPY between a physical reg and an s1, there is no type associated so
3873 // we need to take the virtual register's type as a hint on how to interpret
3874 // s1 values.
3875 unsigned Size;
3876 if (!SrcReg.isVirtual() && !DstBank &&
3877 MRI.getType(DstReg) == LLT::scalar(1)) {
3878 DstBank = &AMDGPU::VCCRegBank;
3879 Size = 1;
3880 } else if (!DstReg.isVirtual() && MRI.getType(SrcReg) == LLT::scalar(1)) {
3881 DstBank = &AMDGPU::VCCRegBank;
3882 Size = 1;
3883 } else {
3884 Size = getSizeInBits(DstReg, MRI, *TRI);
3885 }
3886
3887 if (!DstBank)
3888 DstBank = SrcBank;
3889 else if (!SrcBank)
3890 SrcBank = DstBank;
3891
3892 if (MI.getOpcode() != AMDGPU::G_FREEZE &&
3893 cannotCopy(*DstBank, *SrcBank, TypeSize::getFixed(Size)))
3895
3896 const ValueMapping &ValMap = getValueMapping(0, Size, *DstBank);
3897 unsigned OpdsMappingSize = MI.isCopy() ? 1 : 2;
3898 SmallVector<const ValueMapping *, 1> OpdsMapping(OpdsMappingSize);
3899 OpdsMapping[0] = &ValMap;
3900 if (MI.getOpcode() == AMDGPU::G_FREEZE)
3901 OpdsMapping[1] = &ValMap;
3902
3903 return getInstructionMapping(
3904 1, /*Cost*/ 1,
3905 /*OperandsMapping*/ getOperandsMapping(OpdsMapping), OpdsMappingSize);
3906 }
3907
3908 if (MI.isRegSequence()) {
3909 // If any input is a VGPR, the result must be a VGPR. The default handling
3910 // assumes any copy between banks is legal.
3911 unsigned BankID = AMDGPU::SGPRRegBankID;
3912
3913 for (unsigned I = 1, E = MI.getNumOperands(); I != E; I += 2) {
3914 auto OpBank = getRegBankID(MI.getOperand(I).getReg(), MRI);
3915 // It doesn't make sense to use vcc or scc banks here, so just ignore
3916 // them.
3917 if (OpBank != AMDGPU::SGPRRegBankID) {
3918 BankID = AMDGPU::VGPRRegBankID;
3919 break;
3920 }
3921 }
3922 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
3923
3924 const ValueMapping &ValMap = getValueMapping(0, Size, getRegBank(BankID));
3925 return getInstructionMapping(
3926 1, /*Cost*/ 1,
3927 /*OperandsMapping*/ getOperandsMapping({&ValMap}), 1);
3928 }
3929
3930 // The default handling is broken and doesn't handle illegal SGPR->VGPR copies
3931 // properly.
3932 //
3933 // TODO: There are additional exec masking dependencies to analyze.
3934 if (auto *PHI = dyn_cast<GPhi>(&MI)) {
3935 unsigned ResultBank = AMDGPU::InvalidRegBankID;
3936 Register DstReg = PHI->getReg(0);
3937
3938 // Sometimes the result may have already been assigned a bank.
3939 if (const RegisterBank *DstBank = getRegBank(DstReg, MRI, *TRI))
3940 ResultBank = DstBank->getID();
3941
3942 for (unsigned I = 0; I < PHI->getNumIncomingValues(); ++I) {
3943 Register Reg = PHI->getIncomingValue(I);
3944 const RegisterBank *Bank = getRegBank(Reg, MRI, *TRI);
3945
3946 // FIXME: Assuming VGPR for any undetermined inputs.
3947 if (!Bank || Bank->getID() == AMDGPU::VGPRRegBankID) {
3948 ResultBank = AMDGPU::VGPRRegBankID;
3949 break;
3950 }
3951
3952 // FIXME: Need to promote SGPR case to s32
3953 unsigned OpBank = Bank->getID();
3954 ResultBank = regBankBoolUnion(ResultBank, OpBank);
3955 }
3956
3957 assert(ResultBank != AMDGPU::InvalidRegBankID);
3958
3959 unsigned Size = MRI.getType(DstReg).getSizeInBits();
3960
3961 const ValueMapping &ValMap =
3962 getValueMapping(0, Size, getRegBank(ResultBank));
3963 return getInstructionMapping(
3964 1, /*Cost*/ 1,
3965 /*OperandsMapping*/ getOperandsMapping({&ValMap}), 1);
3966 }
3967
3969 if (Mapping.isValid())
3970 return Mapping;
3971
3972 SmallVector<const ValueMapping*, 8> OpdsMapping(MI.getNumOperands());
3973
3974 switch (MI.getOpcode()) {
3975 default:
3977
3978 case AMDGPU::G_AND:
3979 case AMDGPU::G_OR:
3980 case AMDGPU::G_XOR:
3981 case AMDGPU::G_MUL: {
3982 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
3983 if (Size == 1) {
3984 const RegisterBank *DstBank
3985 = getRegBank(MI.getOperand(0).getReg(), MRI, *TRI);
3986
3987 unsigned TargetBankID = AMDGPU::InvalidRegBankID;
3988 unsigned BankLHS = AMDGPU::InvalidRegBankID;
3989 unsigned BankRHS = AMDGPU::InvalidRegBankID;
3990 if (DstBank) {
3991 TargetBankID = DstBank->getID();
3992 if (DstBank == &AMDGPU::VCCRegBank) {
3993 TargetBankID = AMDGPU::VCCRegBankID;
3994 BankLHS = AMDGPU::VCCRegBankID;
3995 BankRHS = AMDGPU::VCCRegBankID;
3996 } else {
3997 BankLHS = getRegBankID(MI.getOperand(1).getReg(), MRI,
3998 AMDGPU::SGPRRegBankID);
3999 BankRHS = getRegBankID(MI.getOperand(2).getReg(), MRI,
4000 AMDGPU::SGPRRegBankID);
4001 }
4002 } else {
4003 BankLHS = getRegBankID(MI.getOperand(1).getReg(), MRI,
4004 AMDGPU::VCCRegBankID);
4005 BankRHS = getRegBankID(MI.getOperand(2).getReg(), MRI,
4006 AMDGPU::VCCRegBankID);
4007
4008 // Both inputs should be true booleans to produce a boolean result.
4009 if (BankLHS == AMDGPU::VGPRRegBankID || BankRHS == AMDGPU::VGPRRegBankID) {
4010 TargetBankID = AMDGPU::VGPRRegBankID;
4011 } else if (BankLHS == AMDGPU::VCCRegBankID || BankRHS == AMDGPU::VCCRegBankID) {
4012 TargetBankID = AMDGPU::VCCRegBankID;
4013 BankLHS = AMDGPU::VCCRegBankID;
4014 BankRHS = AMDGPU::VCCRegBankID;
4015 } else if (BankLHS == AMDGPU::SGPRRegBankID && BankRHS == AMDGPU::SGPRRegBankID) {
4016 TargetBankID = AMDGPU::SGPRRegBankID;
4017 }
4018 }
4019
4020 OpdsMapping[0] = AMDGPU::getValueMapping(TargetBankID, Size);
4021 OpdsMapping[1] = AMDGPU::getValueMapping(BankLHS, Size);
4022 OpdsMapping[2] = AMDGPU::getValueMapping(BankRHS, Size);
4023 break;
4024 }
4025
4026 if (Size == 64) {
4027
4028 if (isSALUMapping(MI)) {
4029 OpdsMapping[0] = getValueMappingSGPR64Only(AMDGPU::SGPRRegBankID, Size);
4030 OpdsMapping[1] = OpdsMapping[2] = OpdsMapping[0];
4031 } else {
4032 if (MI.getOpcode() == AMDGPU::G_MUL && Subtarget.hasVMulU64Inst())
4033 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
4034 else
4035 OpdsMapping[0] =
4036 getValueMappingSGPR64Only(AMDGPU::VGPRRegBankID, Size);
4037 unsigned Bank1 = getRegBankID(MI.getOperand(1).getReg(), MRI /*, DefaultBankID*/);
4038 OpdsMapping[1] = AMDGPU::getValueMapping(Bank1, Size);
4039
4040 unsigned Bank2 = getRegBankID(MI.getOperand(2).getReg(), MRI /*, DefaultBankID*/);
4041 OpdsMapping[2] = AMDGPU::getValueMapping(Bank2, Size);
4042 }
4043
4044 break;
4045 }
4046
4047 [[fallthrough]];
4048 }
4049 case AMDGPU::G_PTR_ADD:
4050 case AMDGPU::G_PTRMASK:
4051 case AMDGPU::G_ADD:
4052 case AMDGPU::G_SUB:
4053 case AMDGPU::G_SHL:
4054 case AMDGPU::G_LSHR:
4055 case AMDGPU::G_ASHR:
4056 case AMDGPU::G_UADDO:
4057 case AMDGPU::G_USUBO:
4058 case AMDGPU::G_UADDE:
4059 case AMDGPU::G_SADDE:
4060 case AMDGPU::G_USUBE:
4061 case AMDGPU::G_SSUBE:
4062 case AMDGPU::G_ABS:
4063 case AMDGPU::G_SHUFFLE_VECTOR:
4064 case AMDGPU::G_SBFX:
4065 case AMDGPU::G_UBFX:
4066 case AMDGPU::G_AMDGPU_S_MUL_I64_I32:
4067 case AMDGPU::G_AMDGPU_S_MUL_U64_U32:
4068 if (isSALUMapping(MI)) {
4069 LLT Ty = MRI.getType(MI.getOperand(0).getReg());
4070 unsigned Size = Ty.getSizeInBits();
4071 // Packed add and sub are VALU only.
4072 if (Subtarget.hasAnyPackedU64Ops() && Ty.isVector() && Size == 128)
4073 return getDefaultMappingVOP(MI);
4074 return getDefaultMappingSOP(MI);
4075 }
4076 return getDefaultMappingVOP(MI);
4077 case AMDGPU::G_SMIN:
4078 case AMDGPU::G_SMAX:
4079 case AMDGPU::G_UMIN:
4080 case AMDGPU::G_UMAX:
4081 if (isSALUMapping(MI)) {
4082 // There are no scalar 64-bit min and max, use vector instruction instead.
4083 if (MRI.getType(MI.getOperand(0).getReg()).getSizeInBits() == 64 &&
4084 Subtarget.hasMinMaxI64Insts())
4085 return getDefaultMappingVOP(MI);
4086 return getDefaultMappingSOP(MI);
4087 }
4088 return getDefaultMappingVOP(MI);
4089 case AMDGPU::G_FADD:
4090 case AMDGPU::G_FSUB:
4091 case AMDGPU::G_FMUL:
4092 case AMDGPU::G_FMA:
4093 case AMDGPU::G_FFLOOR:
4094 case AMDGPU::G_FCEIL:
4095 case AMDGPU::G_INTRINSIC_ROUNDEVEN:
4096 case AMDGPU::G_FMINNUM:
4097 case AMDGPU::G_FMAXNUM:
4098 case AMDGPU::G_FMINIMUMNUM:
4099 case AMDGPU::G_FMAXIMUMNUM:
4100 case AMDGPU::G_INTRINSIC_TRUNC:
4101 case AMDGPU::G_STRICT_FADD:
4102 case AMDGPU::G_STRICT_FSUB:
4103 case AMDGPU::G_STRICT_FMUL:
4104 case AMDGPU::G_STRICT_FMA: {
4105 LLT Ty = MRI.getType(MI.getOperand(0).getReg());
4106 unsigned Size = Ty.getSizeInBits();
4107 if (Subtarget.hasSALUFloatInsts() && Ty.isScalar() &&
4108 (Size == 32 || Size == 16) && isSALUMapping(MI))
4109 return getDefaultMappingSOP(MI);
4110 return getDefaultMappingVOP(MI);
4111 }
4112 case AMDGPU::G_FMINIMUM:
4113 case AMDGPU::G_FMAXIMUM: {
4114 LLT Ty = MRI.getType(MI.getOperand(0).getReg());
4115 unsigned Size = Ty.getSizeInBits();
4116 if (Subtarget.hasSALUMinimumMaximumInsts() && Ty.isScalar() &&
4117 (Size == 32 || Size == 16) && isSALUMapping(MI))
4118 return getDefaultMappingSOP(MI);
4119 return getDefaultMappingVOP(MI);
4120 }
4121 case AMDGPU::G_FPTOSI:
4122 case AMDGPU::G_FPTOUI:
4123 case AMDGPU::G_FPTOSI_SAT:
4124 case AMDGPU::G_FPTOUI_SAT:
4125 case AMDGPU::G_SITOFP:
4126 case AMDGPU::G_UITOFP: {
4127 unsigned SizeDst = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4128 unsigned SizeSrc = MRI.getType(MI.getOperand(1).getReg()).getSizeInBits();
4129 if (Subtarget.hasSALUFloatInsts() && SizeDst == 32 && SizeSrc == 32 &&
4131 return getDefaultMappingSOP(MI);
4132 return getDefaultMappingVOP(MI);
4133 }
4134 case AMDGPU::G_FPTRUNC:
4135 case AMDGPU::G_FPEXT: {
4136 unsigned SizeDst = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4137 unsigned SizeSrc = MRI.getType(MI.getOperand(1).getReg()).getSizeInBits();
4138 if (Subtarget.hasSALUFloatInsts() && SizeDst != 64 && SizeSrc != 64 &&
4140 return getDefaultMappingSOP(MI);
4141 return getDefaultMappingVOP(MI);
4142 }
4143 case AMDGPU::G_FSQRT:
4144 case AMDGPU::G_FEXP2:
4145 case AMDGPU::G_FLOG2: {
4146 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4147 if (Subtarget.hasPseudoScalarTrans() && (Size == 16 || Size == 32) &&
4149 return getDefaultMappingSOP(MI);
4150 return getDefaultMappingVOP(MI);
4151 }
4152 case AMDGPU::G_SADDSAT: // FIXME: Could lower sat ops for SALU
4153 case AMDGPU::G_SSUBSAT:
4154 case AMDGPU::G_UADDSAT:
4155 case AMDGPU::G_USUBSAT:
4156 case AMDGPU::G_FMAD:
4157 case AMDGPU::G_FLDEXP:
4158 case AMDGPU::G_FMINNUM_IEEE:
4159 case AMDGPU::G_FMAXNUM_IEEE:
4160 case AMDGPU::G_FCANONICALIZE:
4161 case AMDGPU::G_STRICT_FLDEXP:
4162 case AMDGPU::G_BSWAP: // TODO: Somehow expand for scalar?
4163 case AMDGPU::G_FSHR: // TODO: Expand for scalar
4164 case AMDGPU::G_AMDGPU_FMIN_LEGACY:
4165 case AMDGPU::G_AMDGPU_FMAX_LEGACY:
4166 case AMDGPU::G_AMDGPU_RCP_IFLAG:
4167 case AMDGPU::G_AMDGPU_CVT_F32_UBYTE0:
4168 case AMDGPU::G_AMDGPU_CVT_F32_UBYTE1:
4169 case AMDGPU::G_AMDGPU_CVT_F32_UBYTE2:
4170 case AMDGPU::G_AMDGPU_CVT_F32_UBYTE3:
4171 case AMDGPU::G_AMDGPU_CVT_PK_I16_I32:
4172 case AMDGPU::G_AMDGPU_SMED3:
4173 case AMDGPU::G_AMDGPU_FMED3:
4174 return getDefaultMappingVOP(MI);
4175 case AMDGPU::G_UMULH:
4176 case AMDGPU::G_SMULH: {
4177 if (Subtarget.hasScalarMulHiInsts() && isSALUMapping(MI))
4178 return getDefaultMappingSOP(MI);
4179 return getDefaultMappingVOP(MI);
4180 }
4181 case AMDGPU::G_AMDGPU_MAD_U64_U32:
4182 case AMDGPU::G_AMDGPU_MAD_I64_I32: {
4183 // Three possible mappings:
4184 //
4185 // - Default SOP
4186 // - Default VOP
4187 // - Scalar multiply: src0 and src1 are SGPRs, the rest is VOP.
4188 //
4189 // This allows instruction selection to keep the multiplication part of the
4190 // instruction on the SALU.
4191 bool AllSalu = true;
4192 bool MulSalu = true;
4193 for (unsigned i = 0; i < 5; ++i) {
4194 Register Reg = MI.getOperand(i).getReg();
4195 if (const RegisterBank *Bank = getRegBank(Reg, MRI, *TRI)) {
4196 if (Bank->getID() != AMDGPU::SGPRRegBankID) {
4197 AllSalu = false;
4198 if (i == 2 || i == 3) {
4199 MulSalu = false;
4200 break;
4201 }
4202 }
4203 }
4204 }
4205
4206 if (AllSalu)
4207 return getDefaultMappingSOP(MI);
4208
4209 // If the multiply-add is full-rate in VALU, use that even if the
4210 // multiplication part is scalar. Accumulating separately on the VALU would
4211 // take two instructions.
4212 if (!MulSalu || Subtarget.hasFullRate64Ops())
4213 return getDefaultMappingVOP(MI);
4214
4215 // Keep the multiplication on the SALU, then accumulate on the VALU.
4216 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 64);
4217 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1);
4218 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 32);
4219 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 32);
4220 OpdsMapping[4] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 64);
4221 break;
4222 }
4223 case AMDGPU::G_IMPLICIT_DEF: {
4224 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4225 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
4226 break;
4227 }
4228 case AMDGPU::G_FCONSTANT:
4229 case AMDGPU::G_CONSTANT:
4230 case AMDGPU::G_GLOBAL_VALUE:
4231 case AMDGPU::G_FRAME_INDEX:
4232 case AMDGPU::G_BLOCK_ADDR:
4233 case AMDGPU::G_READSTEADYCOUNTER:
4234 case AMDGPU::G_READCYCLECOUNTER: {
4235 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4236 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
4237 break;
4238 }
4239 case AMDGPU::G_DYN_STACKALLOC: {
4240 // Result is always uniform, and a wave reduction is needed for the source.
4241 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 32);
4242 unsigned SrcBankID = getRegBankID(MI.getOperand(1).getReg(), MRI);
4243 OpdsMapping[1] = AMDGPU::getValueMapping(SrcBankID, 32);
4244 break;
4245 }
4246 case AMDGPU::G_AMDGPU_WAVE_ADDRESS: {
4247 // This case is weird because we expect a physical register in the source,
4248 // but need to set a bank anyway.
4249 //
4250 // TODO: We could select the result to SGPR or VGPR
4251 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 32);
4252 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 32);
4253 break;
4254 }
4255 case AMDGPU::G_INSERT: {
4256 unsigned BankID = getMappingType(MRI, MI);
4257 unsigned DstSize = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
4258 unsigned SrcSize = getSizeInBits(MI.getOperand(1).getReg(), MRI, *TRI);
4259 unsigned EltSize = getSizeInBits(MI.getOperand(2).getReg(), MRI, *TRI);
4260 OpdsMapping[0] = AMDGPU::getValueMapping(BankID, DstSize);
4261 OpdsMapping[1] = AMDGPU::getValueMapping(BankID, SrcSize);
4262 OpdsMapping[2] = AMDGPU::getValueMapping(BankID, EltSize);
4263 OpdsMapping[3] = nullptr;
4264 break;
4265 }
4266 case AMDGPU::G_EXTRACT: {
4267 unsigned BankID = getRegBankID(MI.getOperand(1).getReg(), MRI);
4268 unsigned DstSize = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
4269 unsigned SrcSize = getSizeInBits(MI.getOperand(1).getReg(), MRI, *TRI);
4270 OpdsMapping[0] = AMDGPU::getValueMapping(BankID, DstSize);
4271 OpdsMapping[1] = AMDGPU::getValueMapping(BankID, SrcSize);
4272 OpdsMapping[2] = nullptr;
4273 break;
4274 }
4275 case AMDGPU::G_BUILD_VECTOR:
4276 case AMDGPU::G_BUILD_VECTOR_TRUNC: {
4277 LLT DstTy = MRI.getType(MI.getOperand(0).getReg());
4278 if (DstTy == LLT::fixed_vector(2, 16)) {
4279 unsigned DstSize = DstTy.getSizeInBits();
4280 unsigned SrcSize = MRI.getType(MI.getOperand(1).getReg()).getSizeInBits();
4281 unsigned Src0BankID = getRegBankID(MI.getOperand(1).getReg(), MRI);
4282 unsigned Src1BankID = getRegBankID(MI.getOperand(2).getReg(), MRI);
4283 unsigned DstBankID = regBankUnion(Src0BankID, Src1BankID);
4284
4285 OpdsMapping[0] = AMDGPU::getValueMapping(DstBankID, DstSize);
4286 OpdsMapping[1] = AMDGPU::getValueMapping(Src0BankID, SrcSize);
4287 OpdsMapping[2] = AMDGPU::getValueMapping(Src1BankID, SrcSize);
4288 break;
4289 }
4290
4291 [[fallthrough]];
4292 }
4293 case AMDGPU::G_MERGE_VALUES:
4294 case AMDGPU::G_CONCAT_VECTORS: {
4295 unsigned Bank = getMappingType(MRI, MI);
4296 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4297 unsigned SrcSize = MRI.getType(MI.getOperand(1).getReg()).getSizeInBits();
4298
4299 OpdsMapping[0] = AMDGPU::getValueMapping(Bank, DstSize);
4300 // Op1 and Dst should use the same register bank.
4301 for (unsigned i = 1, e = MI.getNumOperands(); i != e; ++i)
4302 OpdsMapping[i] = AMDGPU::getValueMapping(Bank, SrcSize);
4303 break;
4304 }
4305 case AMDGPU::G_BITREVERSE:
4306 case AMDGPU::G_BITCAST:
4307 case AMDGPU::G_INTTOPTR:
4308 case AMDGPU::G_PTRTOINT:
4309 case AMDGPU::G_FABS:
4310 case AMDGPU::G_FNEG: {
4311 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4312 unsigned BankID = getRegBankID(MI.getOperand(1).getReg(), MRI);
4313 OpdsMapping[0] = OpdsMapping[1] = AMDGPU::getValueMapping(BankID, Size);
4314 break;
4315 }
4316 case AMDGPU::G_AMDGPU_FFBH_U32:
4317 case AMDGPU::G_AMDGPU_FFBL_B32:
4318 case AMDGPU::G_CTLZ_ZERO_POISON:
4319 case AMDGPU::G_CTTZ_ZERO_POISON: {
4320 unsigned Size = MRI.getType(MI.getOperand(1).getReg()).getSizeInBits();
4321 unsigned BankID = getRegBankID(MI.getOperand(1).getReg(), MRI);
4322 OpdsMapping[0] = AMDGPU::getValueMapping(BankID, 32);
4323 OpdsMapping[1] = AMDGPU::getValueMappingSGPR64Only(BankID, Size);
4324 break;
4325 }
4326 case AMDGPU::G_CTPOP: {
4327 unsigned Size = MRI.getType(MI.getOperand(1).getReg()).getSizeInBits();
4328 unsigned BankID = getRegBankID(MI.getOperand(1).getReg(), MRI);
4329 OpdsMapping[0] = AMDGPU::getValueMapping(BankID, 32);
4330
4331 // This should really be getValueMappingSGPR64Only, but allowing the generic
4332 // code to handle the register split just makes using LegalizerHelper more
4333 // difficult.
4334 OpdsMapping[1] = AMDGPU::getValueMapping(BankID, Size);
4335 break;
4336 }
4337 case AMDGPU::G_TRUNC: {
4338 Register Dst = MI.getOperand(0).getReg();
4339 Register Src = MI.getOperand(1).getReg();
4340 unsigned Bank = getRegBankID(Src, MRI);
4341 unsigned DstSize = getSizeInBits(Dst, MRI, *TRI);
4342 unsigned SrcSize = getSizeInBits(Src, MRI, *TRI);
4343 OpdsMapping[0] = AMDGPU::getValueMapping(Bank, DstSize);
4344 OpdsMapping[1] = AMDGPU::getValueMapping(Bank, SrcSize);
4345 break;
4346 }
4347 case AMDGPU::G_ZEXT:
4348 case AMDGPU::G_SEXT:
4349 case AMDGPU::G_ANYEXT:
4350 case AMDGPU::G_SEXT_INREG: {
4351 Register Dst = MI.getOperand(0).getReg();
4352 Register Src = MI.getOperand(1).getReg();
4353 unsigned DstSize = getSizeInBits(Dst, MRI, *TRI);
4354 unsigned SrcSize = getSizeInBits(Src, MRI, *TRI);
4355
4356 unsigned DstBank;
4357 const RegisterBank *SrcBank = getRegBank(Src, MRI, *TRI);
4358 assert(SrcBank);
4359 switch (SrcBank->getID()) {
4360 case AMDGPU::SGPRRegBankID:
4361 DstBank = AMDGPU::SGPRRegBankID;
4362 break;
4363 default:
4364 DstBank = AMDGPU::VGPRRegBankID;
4365 break;
4366 }
4367
4368 // Scalar extend can use 64-bit BFE, but VGPRs require extending to
4369 // 32-bits, and then to 64.
4370 OpdsMapping[0] = AMDGPU::getValueMappingSGPR64Only(DstBank, DstSize);
4371 OpdsMapping[1] = AMDGPU::getValueMappingSGPR64Only(SrcBank->getID(),
4372 SrcSize);
4373 break;
4374 }
4375 case AMDGPU::G_IS_FPCLASS: {
4376 Register SrcReg = MI.getOperand(1).getReg();
4377 unsigned SrcSize = MRI.getType(SrcReg).getSizeInBits();
4378 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4379 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, DstSize);
4380 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SrcSize);
4381 break;
4382 }
4383 case AMDGPU::G_STORE: {
4384 assert(MI.getOperand(0).isReg());
4385 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4386
4387 // FIXME: We need to specify a different reg bank once scalar stores are
4388 // supported.
4389 const ValueMapping *ValMapping =
4390 AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
4391 OpdsMapping[0] = ValMapping;
4392 OpdsMapping[1] = getValueMappingForPtr(MRI, MI.getOperand(1).getReg());
4393 break;
4394 }
4395 case AMDGPU::G_ICMP:
4396 case AMDGPU::G_FCMP: {
4397 unsigned Size = MRI.getType(MI.getOperand(2).getReg()).getSizeInBits();
4398
4399 // See if the result register has already been constrained to vcc, which may
4400 // happen due to control flow intrinsic lowering.
4401 unsigned DstBank = getRegBankID(MI.getOperand(0).getReg(), MRI,
4402 AMDGPU::SGPRRegBankID);
4403 unsigned Op2Bank = getRegBankID(MI.getOperand(2).getReg(), MRI);
4404 unsigned Op3Bank = getRegBankID(MI.getOperand(3).getReg(), MRI);
4405
4406 auto canUseSCCICMP = [&]() {
4407 auto Pred =
4408 static_cast<CmpInst::Predicate>(MI.getOperand(1).getPredicate());
4409 return Size == 32 ||
4410 (Size == 64 &&
4411 (Pred == CmpInst::ICMP_EQ || Pred == CmpInst::ICMP_NE) &&
4412 Subtarget.hasScalarCompareEq64());
4413 };
4414 auto canUseSCCFCMP = [&]() {
4415 return Subtarget.hasSALUFloatInsts() && (Size == 32 || Size == 16);
4416 };
4417
4418 bool isICMP = MI.getOpcode() == AMDGPU::G_ICMP;
4419 bool CanUseSCC = DstBank == AMDGPU::SGPRRegBankID &&
4420 Op2Bank == AMDGPU::SGPRRegBankID &&
4421 Op3Bank == AMDGPU::SGPRRegBankID &&
4422 (isICMP ? canUseSCCICMP() : canUseSCCFCMP());
4423
4424 DstBank = CanUseSCC ? AMDGPU::SGPRRegBankID : AMDGPU::VCCRegBankID;
4425 unsigned SrcBank = CanUseSCC ? AMDGPU::SGPRRegBankID : AMDGPU::VGPRRegBankID;
4426
4427 // TODO: Use 32-bit for scalar output size.
4428 // SCC results will need to be copied to a 32-bit SGPR virtual register.
4429 const unsigned ResultSize = 1;
4430
4431 OpdsMapping[0] = AMDGPU::getValueMapping(DstBank, ResultSize);
4432 OpdsMapping[1] = nullptr; // Predicate Operand.
4433 OpdsMapping[2] = AMDGPU::getValueMapping(SrcBank, Size);
4434 OpdsMapping[3] = AMDGPU::getValueMapping(SrcBank, Size);
4435 break;
4436 }
4437 case AMDGPU::G_EXTRACT_VECTOR_ELT: {
4438 // VGPR index can be used for waterfall when indexing a SGPR vector.
4439 unsigned SrcBankID = getRegBankID(MI.getOperand(1).getReg(), MRI);
4440 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4441 unsigned SrcSize = MRI.getType(MI.getOperand(1).getReg()).getSizeInBits();
4442 unsigned IdxSize = MRI.getType(MI.getOperand(2).getReg()).getSizeInBits();
4443 unsigned IdxBank = getRegBankID(MI.getOperand(2).getReg(), MRI);
4444 unsigned OutputBankID = regBankUnion(SrcBankID, IdxBank);
4445
4446 OpdsMapping[0] = AMDGPU::getValueMappingSGPR64Only(OutputBankID, DstSize);
4447 OpdsMapping[1] = AMDGPU::getValueMapping(SrcBankID, SrcSize);
4448
4449 // The index can be either if the source vector is VGPR.
4450 OpdsMapping[2] = AMDGPU::getValueMapping(IdxBank, IdxSize);
4451 break;
4452 }
4453 case AMDGPU::G_INSERT_VECTOR_ELT: {
4454 unsigned OutputBankID = isSALUMapping(MI) ?
4455 AMDGPU::SGPRRegBankID : AMDGPU::VGPRRegBankID;
4456
4457 unsigned VecSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4458 unsigned InsertSize = MRI.getType(MI.getOperand(2).getReg()).getSizeInBits();
4459 unsigned IdxSize = MRI.getType(MI.getOperand(3).getReg()).getSizeInBits();
4460 unsigned InsertEltBankID = getRegBankID(MI.getOperand(2).getReg(), MRI);
4461 unsigned IdxBankID = getRegBankID(MI.getOperand(3).getReg(), MRI);
4462
4463 OpdsMapping[0] = AMDGPU::getValueMapping(OutputBankID, VecSize);
4464 OpdsMapping[1] = AMDGPU::getValueMapping(OutputBankID, VecSize);
4465
4466 // This is a weird case, because we need to break down the mapping based on
4467 // the register bank of a different operand.
4468 if (InsertSize == 64 && OutputBankID == AMDGPU::VGPRRegBankID) {
4469 OpdsMapping[2] = AMDGPU::getValueMappingSplit64(InsertEltBankID,
4470 InsertSize);
4471 } else {
4472 assert(InsertSize == 32 || InsertSize == 64);
4473 OpdsMapping[2] = AMDGPU::getValueMapping(InsertEltBankID, InsertSize);
4474 }
4475
4476 // The index can be either if the source vector is VGPR.
4477 OpdsMapping[3] = AMDGPU::getValueMapping(IdxBankID, IdxSize);
4478 break;
4479 }
4480 case AMDGPU::G_UNMERGE_VALUES: {
4481 unsigned Bank = getMappingType(MRI, MI);
4482
4483 // Op1 and Dst should use the same register bank.
4484 // FIXME: Shouldn't this be the default? Why do we need to handle this?
4485 for (unsigned i = 0, e = MI.getNumOperands(); i != e; ++i) {
4486 unsigned Size = getSizeInBits(MI.getOperand(i).getReg(), MRI, *TRI);
4487 OpdsMapping[i] = AMDGPU::getValueMapping(Bank, Size);
4488 }
4489 break;
4490 }
4491 case AMDGPU::G_AMDGPU_BUFFER_LOAD:
4492 case AMDGPU::G_AMDGPU_BUFFER_LOAD_UBYTE:
4493 case AMDGPU::G_AMDGPU_BUFFER_LOAD_SBYTE:
4494 case AMDGPU::G_AMDGPU_BUFFER_LOAD_USHORT:
4495 case AMDGPU::G_AMDGPU_BUFFER_LOAD_SSHORT:
4496 case AMDGPU::G_AMDGPU_BUFFER_LOAD_TFE:
4497 case AMDGPU::G_AMDGPU_BUFFER_LOAD_UBYTE_TFE:
4498 case AMDGPU::G_AMDGPU_BUFFER_LOAD_SBYTE_TFE:
4499 case AMDGPU::G_AMDGPU_BUFFER_LOAD_USHORT_TFE:
4500 case AMDGPU::G_AMDGPU_BUFFER_LOAD_SSHORT_TFE:
4501 case AMDGPU::G_AMDGPU_BUFFER_LOAD_FORMAT:
4502 case AMDGPU::G_AMDGPU_BUFFER_LOAD_FORMAT_TFE:
4503 case AMDGPU::G_AMDGPU_BUFFER_LOAD_FORMAT_D16:
4504 case AMDGPU::G_AMDGPU_TBUFFER_LOAD_FORMAT:
4505 case AMDGPU::G_AMDGPU_TBUFFER_LOAD_FORMAT_D16:
4506 case AMDGPU::G_AMDGPU_TBUFFER_STORE_FORMAT:
4507 case AMDGPU::G_AMDGPU_TBUFFER_STORE_FORMAT_D16:
4508 case AMDGPU::G_AMDGPU_BUFFER_STORE:
4509 case AMDGPU::G_AMDGPU_BUFFER_STORE_BYTE:
4510 case AMDGPU::G_AMDGPU_BUFFER_STORE_SHORT:
4511 case AMDGPU::G_AMDGPU_BUFFER_STORE_FORMAT:
4512 case AMDGPU::G_AMDGPU_BUFFER_STORE_FORMAT_D16: {
4513 OpdsMapping[0] = getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
4514
4515 // rsrc
4516 OpdsMapping[1] = getSGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
4517
4518 // vindex
4519 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
4520
4521 // voffset
4522 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
4523
4524 // soffset
4525 OpdsMapping[4] = getSGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
4526
4527 // Any remaining operands are immediates and were correctly null
4528 // initialized.
4529 break;
4530 }
4531 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SWAP:
4532 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_ADD:
4533 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SUB:
4534 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SMIN:
4535 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_UMIN:
4536 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SMAX:
4537 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_UMAX:
4538 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_AND:
4539 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_OR:
4540 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_XOR:
4541 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_INC:
4542 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_DEC:
4543 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_SUB_CLAMP_U32:
4544 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_COND_SUB_U32:
4545 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_FADD:
4546 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_FMIN:
4547 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_FMAX: {
4548 // vdata_out
4549 OpdsMapping[0] = getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
4550
4551 // vdata_in
4552 OpdsMapping[1] = getVGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
4553
4554 // rsrc
4555 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
4556
4557 // vindex
4558 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
4559
4560 // voffset
4561 OpdsMapping[4] = getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
4562
4563 // soffset
4564 OpdsMapping[5] = getSGPROpMapping(MI.getOperand(5).getReg(), MRI, *TRI);
4565
4566 // Any remaining operands are immediates and were correctly null
4567 // initialized.
4568 break;
4569 }
4570 case AMDGPU::G_AMDGPU_BUFFER_ATOMIC_CMPSWAP: {
4571 // vdata_out
4572 OpdsMapping[0] = getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
4573
4574 // vdata_in
4575 OpdsMapping[1] = getVGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
4576
4577 // cmp
4578 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
4579
4580 // rsrc
4581 OpdsMapping[3] = getSGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
4582
4583 // vindex
4584 OpdsMapping[4] = getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
4585
4586 // voffset
4587 OpdsMapping[5] = getVGPROpMapping(MI.getOperand(5).getReg(), MRI, *TRI);
4588
4589 // soffset
4590 OpdsMapping[6] = getSGPROpMapping(MI.getOperand(6).getReg(), MRI, *TRI);
4591
4592 // Any remaining operands are immediates and were correctly null
4593 // initialized.
4594 break;
4595 }
4596 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD:
4597 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_UBYTE:
4598 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_SBYTE:
4599 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_USHORT:
4600 case AMDGPU::G_AMDGPU_S_BUFFER_LOAD_SSHORT: {
4601 // Lie and claim everything is legal, even though some need to be
4602 // SGPRs. applyMapping will have to deal with it as a waterfall loop.
4603 OpdsMapping[1] = getSGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
4604 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
4605
4606 // We need to convert this to a MUBUF if either the resource of offset is
4607 // VGPR.
4608 unsigned RSrcBank = OpdsMapping[1]->BreakDown[0].RegBank->getID();
4609 unsigned OffsetBank = OpdsMapping[2]->BreakDown[0].RegBank->getID();
4610 unsigned ResultBank = regBankUnion(RSrcBank, OffsetBank);
4611
4612 unsigned Size0 = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4613 OpdsMapping[0] = AMDGPU::getValueMapping(ResultBank, Size0);
4614 break;
4615 }
4616 case AMDGPU::G_AMDGPU_S_BUFFER_PREFETCH:
4617 OpdsMapping[0] = getSGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
4618 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
4619 break;
4620 case AMDGPU::G_AMDGPU_SPONENTRY: {
4621 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4622 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
4623 break;
4624 }
4625 case AMDGPU::G_INTRINSIC:
4626 case AMDGPU::G_INTRINSIC_CONVERGENT: {
4627 switch (cast<GIntrinsic>(MI).getIntrinsicID()) {
4628 default:
4630 case Intrinsic::amdgcn_div_fmas:
4631 case Intrinsic::amdgcn_div_fixup:
4632 case Intrinsic::amdgcn_trig_preop:
4633 case Intrinsic::amdgcn_sin:
4634 case Intrinsic::amdgcn_cos:
4635 case Intrinsic::amdgcn_log_clamp:
4636 case Intrinsic::amdgcn_rcp_legacy:
4637 case Intrinsic::amdgcn_rsq_legacy:
4638 case Intrinsic::amdgcn_rsq_clamp:
4639 case Intrinsic::amdgcn_tanh:
4640 case Intrinsic::amdgcn_fmul_legacy:
4641 case Intrinsic::amdgcn_fma_legacy:
4642 case Intrinsic::amdgcn_frexp_mant:
4643 case Intrinsic::amdgcn_frexp_exp:
4644 case Intrinsic::amdgcn_fract:
4645 case Intrinsic::amdgcn_cvt_pknorm_i16:
4646 case Intrinsic::amdgcn_cvt_pknorm_u16:
4647 case Intrinsic::amdgcn_cvt_pk_i16:
4648 case Intrinsic::amdgcn_cvt_pk_u16:
4649 case Intrinsic::amdgcn_cvt_sr_pk_f16_f32:
4650 case Intrinsic::amdgcn_cvt_sr_pk_bf16_f32:
4651 case Intrinsic::amdgcn_cvt_pk_f16_fp8:
4652 case Intrinsic::amdgcn_cvt_pk_f16_bf8:
4653 case Intrinsic::amdgcn_cvt_pk_fp8_f16:
4654 case Intrinsic::amdgcn_cvt_pk_bf8_f16:
4655 case Intrinsic::amdgcn_cvt_sr_fp8_f16:
4656 case Intrinsic::amdgcn_cvt_sr_bf8_f16:
4657 case Intrinsic::amdgcn_cvt_scale_pk8_f16_fp8:
4658 case Intrinsic::amdgcn_cvt_scale_pk8_bf16_fp8:
4659 case Intrinsic::amdgcn_cvt_scale_pk8_f16_bf8:
4660 case Intrinsic::amdgcn_cvt_scale_pk8_bf16_bf8:
4661 case Intrinsic::amdgcn_cvt_scale_pk8_f16_fp4:
4662 case Intrinsic::amdgcn_cvt_scale_pk8_bf16_fp4:
4663 case Intrinsic::amdgcn_cvt_scale_pk8_f32_fp8:
4664 case Intrinsic::amdgcn_cvt_scale_pk8_f32_bf8:
4665 case Intrinsic::amdgcn_cvt_scale_pk8_f32_fp4:
4666 case Intrinsic::amdgcn_cvt_scale_pk16_f16_fp6:
4667 case Intrinsic::amdgcn_cvt_scale_pk16_bf16_fp6:
4668 case Intrinsic::amdgcn_cvt_scale_pk16_f16_bf6:
4669 case Intrinsic::amdgcn_cvt_scale_pk16_bf16_bf6:
4670 case Intrinsic::amdgcn_cvt_scale_pk16_f32_fp6:
4671 case Intrinsic::amdgcn_cvt_scale_pk16_f32_bf6:
4672 case Intrinsic::amdgcn_cvt_scalef32_pk8_fp8_bf16:
4673 case Intrinsic::amdgcn_cvt_scalef32_pk8_bf8_bf16:
4674 case Intrinsic::amdgcn_cvt_scalef32_pk8_fp8_f16:
4675 case Intrinsic::amdgcn_cvt_scalef32_pk8_bf8_f16:
4676 case Intrinsic::amdgcn_cvt_scalef32_pk8_fp8_f32:
4677 case Intrinsic::amdgcn_cvt_scalef32_pk8_bf8_f32:
4678 case Intrinsic::amdgcn_cvt_scalef32_pk8_fp4_f32:
4679 case Intrinsic::amdgcn_cvt_scalef32_pk8_fp4_f16:
4680 case Intrinsic::amdgcn_cvt_scalef32_pk8_fp4_bf16:
4681 case Intrinsic::amdgcn_cvt_scalef32_pk16_fp6_f32:
4682 case Intrinsic::amdgcn_cvt_scalef32_pk16_bf6_f32:
4683 case Intrinsic::amdgcn_cvt_scalef32_pk16_fp6_f16:
4684 case Intrinsic::amdgcn_cvt_scalef32_pk16_bf6_f16:
4685 case Intrinsic::amdgcn_cvt_scalef32_pk16_fp6_bf16:
4686 case Intrinsic::amdgcn_cvt_scalef32_pk16_bf6_bf16:
4687 case Intrinsic::amdgcn_cvt_scalef32_sr_pk8_fp8_bf16:
4688 case Intrinsic::amdgcn_cvt_scalef32_sr_pk8_bf8_bf16:
4689 case Intrinsic::amdgcn_cvt_scalef32_sr_pk8_fp8_f16:
4690 case Intrinsic::amdgcn_cvt_scalef32_sr_pk8_bf8_f16:
4691 case Intrinsic::amdgcn_cvt_scalef32_sr_pk8_fp8_f32:
4692 case Intrinsic::amdgcn_cvt_scalef32_sr_pk8_bf8_f32:
4693 case Intrinsic::amdgcn_cvt_scalef32_sr_pk8_fp4_f32:
4694 case Intrinsic::amdgcn_cvt_scalef32_sr_pk8_fp4_f16:
4695 case Intrinsic::amdgcn_cvt_scalef32_sr_pk8_fp4_bf16:
4696 case Intrinsic::amdgcn_cvt_scalef32_sr_pk16_fp6_f32:
4697 case Intrinsic::amdgcn_cvt_scalef32_sr_pk16_bf6_f32:
4698 case Intrinsic::amdgcn_cvt_scalef32_sr_pk16_fp6_f16:
4699 case Intrinsic::amdgcn_cvt_scalef32_sr_pk16_bf6_f16:
4700 case Intrinsic::amdgcn_cvt_scalef32_sr_pk16_fp6_bf16:
4701 case Intrinsic::amdgcn_cvt_scalef32_sr_pk16_bf6_bf16:
4702 case Intrinsic::amdgcn_sat_pk4_i4_i8:
4703 case Intrinsic::amdgcn_sat_pk4_u4_u8:
4704 case Intrinsic::amdgcn_fmed3:
4705 case Intrinsic::amdgcn_cubeid:
4706 case Intrinsic::amdgcn_cubema:
4707 case Intrinsic::amdgcn_cubesc:
4708 case Intrinsic::amdgcn_cubetc:
4709 case Intrinsic::amdgcn_sffbh:
4710 case Intrinsic::amdgcn_fmad_ftz:
4711 case Intrinsic::amdgcn_mbcnt_lo:
4712 case Intrinsic::amdgcn_mbcnt_hi:
4713 case Intrinsic::amdgcn_mul_u24:
4714 case Intrinsic::amdgcn_mul_i24:
4715 case Intrinsic::amdgcn_mulhi_u24:
4716 case Intrinsic::amdgcn_mulhi_i24:
4717 case Intrinsic::amdgcn_lerp:
4718 case Intrinsic::amdgcn_sad_u8:
4719 case Intrinsic::amdgcn_msad_u8:
4720 case Intrinsic::amdgcn_sad_hi_u8:
4721 case Intrinsic::amdgcn_sad_u16:
4722 case Intrinsic::amdgcn_qsad_pk_u16_u8:
4723 case Intrinsic::amdgcn_mqsad_pk_u16_u8:
4724 case Intrinsic::amdgcn_mqsad_u32_u8:
4725 case Intrinsic::amdgcn_cvt_pk_u8_f32:
4726 case Intrinsic::amdgcn_alignbyte:
4727 case Intrinsic::amdgcn_perm:
4728 case Intrinsic::amdgcn_prng_b32:
4729 case Intrinsic::amdgcn_fdot2:
4730 case Intrinsic::amdgcn_sdot2:
4731 case Intrinsic::amdgcn_udot2:
4732 case Intrinsic::amdgcn_sdot4:
4733 case Intrinsic::amdgcn_udot4:
4734 case Intrinsic::amdgcn_sdot8:
4735 case Intrinsic::amdgcn_udot8:
4736 case Intrinsic::amdgcn_fdot2_bf16_bf16:
4737 case Intrinsic::amdgcn_fdot2_f16_f16:
4738 case Intrinsic::amdgcn_fdot2_f32_bf16:
4739 case Intrinsic::amdgcn_fdot2c_f32_bf16:
4740 case Intrinsic::amdgcn_sudot4:
4741 case Intrinsic::amdgcn_sudot8:
4742 case Intrinsic::amdgcn_dot4_f32_fp8_bf8:
4743 case Intrinsic::amdgcn_dot4_f32_bf8_fp8:
4744 case Intrinsic::amdgcn_dot4_f32_fp8_fp8:
4745 case Intrinsic::amdgcn_dot4_f32_bf8_bf8:
4746 case Intrinsic::amdgcn_cvt_f32_fp8:
4747 case Intrinsic::amdgcn_cvt_f32_fp8_e5m3:
4748 case Intrinsic::amdgcn_cvt_f32_bf8:
4749 case Intrinsic::amdgcn_cvt_off_f32_i4:
4750 case Intrinsic::amdgcn_cvt_pk_f32_fp8:
4751 case Intrinsic::amdgcn_cvt_pk_f32_bf8:
4752 case Intrinsic::amdgcn_cvt_pk_fp8_f32:
4753 case Intrinsic::amdgcn_cvt_pk_fp8_f32_e5m3:
4754 case Intrinsic::amdgcn_cvt_pk_bf8_f32:
4755 case Intrinsic::amdgcn_cvt_sr_fp8_f32:
4756 case Intrinsic::amdgcn_cvt_sr_fp8_f32_e5m3:
4757 case Intrinsic::amdgcn_cvt_sr_bf8_f32:
4758 case Intrinsic::amdgcn_cvt_sr_bf16_f32:
4759 case Intrinsic::amdgcn_cvt_sr_f16_f32:
4760 case Intrinsic::amdgcn_cvt_f16_fp8:
4761 case Intrinsic::amdgcn_cvt_f16_bf8:
4762 case Intrinsic::amdgcn_cvt_scalef32_pk32_fp6_f16:
4763 case Intrinsic::amdgcn_cvt_scalef32_pk32_bf6_f16:
4764 case Intrinsic::amdgcn_cvt_scalef32_pk32_fp6_bf16:
4765 case Intrinsic::amdgcn_cvt_scalef32_pk32_bf6_bf16:
4766 case Intrinsic::amdgcn_cvt_scalef32_pk32_fp6_f32:
4767 case Intrinsic::amdgcn_cvt_scalef32_pk32_bf6_f32:
4768 case Intrinsic::amdgcn_cvt_scalef32_f16_fp8:
4769 case Intrinsic::amdgcn_cvt_scalef32_f16_bf8:
4770 case Intrinsic::amdgcn_cvt_scalef32_f32_fp8:
4771 case Intrinsic::amdgcn_cvt_scalef32_f32_bf8:
4772 case Intrinsic::amdgcn_cvt_scalef32_pk_fp8_f32:
4773 case Intrinsic::amdgcn_cvt_scalef32_pk_bf8_f32:
4774 case Intrinsic::amdgcn_cvt_scalef32_pk_f32_fp8:
4775 case Intrinsic::amdgcn_cvt_scalef32_pk_f32_bf8:
4776 case Intrinsic::amdgcn_cvt_scalef32_pk_fp8_f16:
4777 case Intrinsic::amdgcn_cvt_scalef32_pk_fp8_bf16:
4778 case Intrinsic::amdgcn_cvt_scalef32_pk_bf8_f16:
4779 case Intrinsic::amdgcn_cvt_scalef32_pk_bf8_bf16:
4780 case Intrinsic::amdgcn_cvt_scalef32_pk_f32_fp4:
4781 case Intrinsic::amdgcn_cvt_scalef32_pk_fp4_f32:
4782 case Intrinsic::amdgcn_cvt_scalef32_pk_f16_fp4:
4783 case Intrinsic::amdgcn_cvt_scalef32_pk_bf16_fp4:
4784 case Intrinsic::amdgcn_cvt_scalef32_pk32_f32_fp6:
4785 case Intrinsic::amdgcn_cvt_scalef32_pk32_f32_bf6:
4786 case Intrinsic::amdgcn_cvt_scalef32_pk32_f16_bf6:
4787 case Intrinsic::amdgcn_cvt_scalef32_pk32_bf16_bf6:
4788 case Intrinsic::amdgcn_cvt_scalef32_pk32_f16_fp6:
4789 case Intrinsic::amdgcn_cvt_scalef32_pk32_bf16_fp6:
4790 case Intrinsic::amdgcn_cvt_scalef32_pk_f16_bf8:
4791 case Intrinsic::amdgcn_cvt_scalef32_pk_bf16_bf8:
4792 case Intrinsic::amdgcn_cvt_scalef32_pk_f16_fp8:
4793 case Intrinsic::amdgcn_cvt_scalef32_pk_bf16_fp8:
4794 case Intrinsic::amdgcn_cvt_scalef32_pk_fp4_f16:
4795 case Intrinsic::amdgcn_cvt_scalef32_pk_fp4_bf16:
4796 case Intrinsic::amdgcn_cvt_scalef32_sr_pk_fp4_f16:
4797 case Intrinsic::amdgcn_cvt_scalef32_sr_pk_fp4_bf16:
4798 case Intrinsic::amdgcn_cvt_scalef32_sr_pk_fp4_f32:
4799 case Intrinsic::amdgcn_cvt_scalef32_sr_pk32_bf6_bf16:
4800 case Intrinsic::amdgcn_cvt_scalef32_sr_pk32_bf6_f16:
4801 case Intrinsic::amdgcn_cvt_scalef32_sr_pk32_bf6_f32:
4802 case Intrinsic::amdgcn_cvt_scalef32_sr_pk32_fp6_bf16:
4803 case Intrinsic::amdgcn_cvt_scalef32_sr_pk32_fp6_f16:
4804 case Intrinsic::amdgcn_cvt_scalef32_sr_pk32_fp6_f32:
4805 case Intrinsic::amdgcn_cvt_scalef32_sr_bf8_bf16:
4806 case Intrinsic::amdgcn_cvt_scalef32_sr_bf8_f16:
4807 case Intrinsic::amdgcn_cvt_scalef32_sr_bf8_f32:
4808 case Intrinsic::amdgcn_cvt_scalef32_sr_fp8_bf16:
4809 case Intrinsic::amdgcn_cvt_scalef32_sr_fp8_f16:
4810 case Intrinsic::amdgcn_cvt_scalef32_sr_fp8_f32:
4811 case Intrinsic::amdgcn_ashr_pk_i8_i32:
4812 case Intrinsic::amdgcn_ashr_pk_u8_i32:
4813 case Intrinsic::amdgcn_cvt_scalef32_2xpk16_fp6_f32:
4814 case Intrinsic::amdgcn_cvt_scalef32_2xpk16_bf6_f32:
4815 case Intrinsic::amdgcn_wmma_bf16_16x16x16_bf16:
4816 case Intrinsic::amdgcn_wmma_f16_16x16x16_f16:
4817 case Intrinsic::amdgcn_wmma_bf16_16x16x16_bf16_tied:
4818 case Intrinsic::amdgcn_wmma_f16_16x16x16_f16_tied:
4819 case Intrinsic::amdgcn_wmma_f32_16x16x16_bf16:
4820 case Intrinsic::amdgcn_wmma_f32_16x16x16_f16:
4821 case Intrinsic::amdgcn_wmma_i32_16x16x16_iu4:
4822 case Intrinsic::amdgcn_wmma_i32_16x16x16_iu8:
4823 case Intrinsic::amdgcn_wmma_f32_16x16x16_fp8_fp8:
4824 case Intrinsic::amdgcn_wmma_f32_16x16x16_fp8_bf8:
4825 case Intrinsic::amdgcn_wmma_f32_16x16x16_bf8_fp8:
4826 case Intrinsic::amdgcn_wmma_f32_16x16x16_bf8_bf8:
4827 case Intrinsic::amdgcn_wmma_i32_16x16x32_iu4:
4828 case Intrinsic::amdgcn_swmmac_f32_16x16x32_f16:
4829 case Intrinsic::amdgcn_swmmac_f32_16x16x32_bf16:
4830 case Intrinsic::amdgcn_swmmac_f16_16x16x32_f16:
4831 case Intrinsic::amdgcn_swmmac_bf16_16x16x32_bf16:
4832 case Intrinsic::amdgcn_swmmac_i32_16x16x32_iu8:
4833 case Intrinsic::amdgcn_swmmac_i32_16x16x32_iu4:
4834 case Intrinsic::amdgcn_swmmac_i32_16x16x64_iu4:
4835 case Intrinsic::amdgcn_swmmac_f32_16x16x32_fp8_fp8:
4836 case Intrinsic::amdgcn_swmmac_f32_16x16x32_fp8_bf8:
4837 case Intrinsic::amdgcn_swmmac_f32_16x16x32_bf8_fp8:
4838 case Intrinsic::amdgcn_swmmac_f32_16x16x32_bf8_bf8:
4839 case Intrinsic::amdgcn_wmma_f64_16x16x4_f64:
4840 case Intrinsic::amdgcn_wmma_f32_16x16x4_f32:
4841 case Intrinsic::amdgcn_wmma_f32_16x16x32_bf16:
4842 case Intrinsic::amdgcn_wmma_f32_16x16x32_f16:
4843 case Intrinsic::amdgcn_wmma_f16_16x16x32_f16:
4844 case Intrinsic::amdgcn_wmma_bf16_16x16x32_bf16:
4845 case Intrinsic::amdgcn_wmma_bf16f32_16x16x32_bf16:
4846 case Intrinsic::amdgcn_wmma_f32_16x16x64_fp8_fp8:
4847 case Intrinsic::amdgcn_wmma_f32_16x16x64_fp8_bf8:
4848 case Intrinsic::amdgcn_wmma_f32_16x16x64_bf8_fp8:
4849 case Intrinsic::amdgcn_wmma_f32_16x16x64_bf8_bf8:
4850 case Intrinsic::amdgcn_wmma_f16_16x16x64_fp8_fp8:
4851 case Intrinsic::amdgcn_wmma_f16_16x16x64_fp8_bf8:
4852 case Intrinsic::amdgcn_wmma_f16_16x16x64_bf8_fp8:
4853 case Intrinsic::amdgcn_wmma_f16_16x16x64_bf8_bf8:
4854 case Intrinsic::amdgcn_wmma_f16_16x16x128_fp8_fp8:
4855 case Intrinsic::amdgcn_wmma_f16_16x16x128_fp8_bf8:
4856 case Intrinsic::amdgcn_wmma_f16_16x16x128_bf8_fp8:
4857 case Intrinsic::amdgcn_wmma_f16_16x16x128_bf8_bf8:
4858 case Intrinsic::amdgcn_wmma_f32_16x16x128_fp8_fp8:
4859 case Intrinsic::amdgcn_wmma_f32_16x16x128_fp8_bf8:
4860 case Intrinsic::amdgcn_wmma_f32_16x16x128_bf8_fp8:
4861 case Intrinsic::amdgcn_wmma_f32_16x16x128_bf8_bf8:
4862 case Intrinsic::amdgcn_wmma_i32_16x16x64_iu8:
4863 case Intrinsic::amdgcn_wmma_f32_16x16x128_f8f6f4:
4864 case Intrinsic::amdgcn_wmma_scale_f32_16x16x128_f8f6f4:
4865 case Intrinsic::amdgcn_wmma_scale16_f32_16x16x128_f8f6f4:
4866 case Intrinsic::amdgcn_wmma_f32_32x16x128_f4:
4867 case Intrinsic::amdgcn_wmma_scale_f32_32x16x128_f4:
4868 case Intrinsic::amdgcn_wmma_scale16_f32_32x16x128_f4:
4869 case Intrinsic::amdgcn_swmmac_f16_16x16x64_f16:
4870 case Intrinsic::amdgcn_swmmac_bf16_16x16x64_bf16:
4871 case Intrinsic::amdgcn_swmmac_f32_16x16x64_bf16:
4872 case Intrinsic::amdgcn_swmmac_bf16f32_16x16x64_bf16:
4873 case Intrinsic::amdgcn_swmmac_f32_16x16x64_f16:
4874 case Intrinsic::amdgcn_swmmac_f32_16x16x128_fp8_fp8:
4875 case Intrinsic::amdgcn_swmmac_f32_16x16x128_fp8_bf8:
4876 case Intrinsic::amdgcn_swmmac_f32_16x16x128_bf8_fp8:
4877 case Intrinsic::amdgcn_swmmac_f32_16x16x128_bf8_bf8:
4878 case Intrinsic::amdgcn_swmmac_f16_16x16x128_fp8_fp8:
4879 case Intrinsic::amdgcn_swmmac_f16_16x16x128_fp8_bf8:
4880 case Intrinsic::amdgcn_swmmac_f16_16x16x128_bf8_fp8:
4881 case Intrinsic::amdgcn_swmmac_f16_16x16x128_bf8_bf8:
4882 case Intrinsic::amdgcn_swmmac_i32_16x16x128_iu8:
4883 case Intrinsic::amdgcn_perm_pk16_b4_u4:
4884 case Intrinsic::amdgcn_perm_pk16_b6_u4:
4885 case Intrinsic::amdgcn_perm_pk16_b8_u4:
4886 case Intrinsic::amdgcn_add_max_i32:
4887 case Intrinsic::amdgcn_add_max_u32:
4888 case Intrinsic::amdgcn_add_min_i32:
4889 case Intrinsic::amdgcn_add_min_u32:
4890 case Intrinsic::amdgcn_pk_add_max_i16:
4891 case Intrinsic::amdgcn_pk_add_max_u16:
4892 case Intrinsic::amdgcn_pk_add_min_i16:
4893 case Intrinsic::amdgcn_pk_add_min_u16:
4894 return getDefaultMappingVOP(MI);
4895 case Intrinsic::amdgcn_log:
4896 case Intrinsic::amdgcn_exp2:
4897 case Intrinsic::amdgcn_rcp:
4898 case Intrinsic::amdgcn_rsq:
4899 case Intrinsic::amdgcn_sqrt: {
4900 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4901 if (Subtarget.hasPseudoScalarTrans() && (Size == 16 || Size == 32) &&
4903 return getDefaultMappingSOP(MI);
4904 return getDefaultMappingVOP(MI);
4905 }
4906 case Intrinsic::amdgcn_sbfe:
4907 case Intrinsic::amdgcn_ubfe:
4908 if (isSALUMapping(MI))
4909 return getDefaultMappingSOP(MI);
4910 return getDefaultMappingVOP(MI);
4911 case Intrinsic::amdgcn_ds_swizzle:
4912 case Intrinsic::amdgcn_ds_permute:
4913 case Intrinsic::amdgcn_ds_bpermute:
4914 case Intrinsic::amdgcn_update_dpp:
4915 case Intrinsic::amdgcn_mov_dpp8:
4916 case Intrinsic::amdgcn_mov_dpp:
4917 case Intrinsic::amdgcn_strict_wwm:
4918 case Intrinsic::amdgcn_wwm:
4919 case Intrinsic::amdgcn_strict_wqm:
4920 case Intrinsic::amdgcn_wqm:
4921 case Intrinsic::amdgcn_softwqm:
4922 case Intrinsic::amdgcn_set_inactive:
4923 case Intrinsic::amdgcn_set_inactive_chain_arg:
4924 case Intrinsic::amdgcn_permlane64:
4925 case Intrinsic::amdgcn_ds_bpermute_fi_b32:
4927 case Intrinsic::amdgcn_cvt_pkrtz:
4928 if (Subtarget.hasSALUFloatInsts() && isSALUMapping(MI))
4929 return getDefaultMappingSOP(MI);
4930 return getDefaultMappingVOP(MI);
4931 case Intrinsic::amdgcn_kernarg_segment_ptr:
4932 case Intrinsic::amdgcn_s_getpc:
4933 case Intrinsic::amdgcn_groupstaticsize:
4934 case Intrinsic::amdgcn_reloc_constant:
4935 case Intrinsic::returnaddress: {
4936 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4937 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
4938 break;
4939 }
4940 case Intrinsic::amdgcn_wqm_vote: {
4941 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4942 OpdsMapping[0] = OpdsMapping[2]
4943 = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, Size);
4944 break;
4945 }
4946 case Intrinsic::amdgcn_ps_live: {
4947 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1);
4948 break;
4949 }
4950 case Intrinsic::amdgcn_div_scale: {
4951 unsigned Dst0Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4952 unsigned Dst1Size = MRI.getType(MI.getOperand(1).getReg()).getSizeInBits();
4953 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Dst0Size);
4954 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, Dst1Size);
4955
4956 unsigned SrcSize = MRI.getType(MI.getOperand(3).getReg()).getSizeInBits();
4957 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SrcSize);
4958 OpdsMapping[4] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SrcSize);
4959 break;
4960 }
4961 case Intrinsic::amdgcn_class: {
4962 Register Src0Reg = MI.getOperand(2).getReg();
4963 Register Src1Reg = MI.getOperand(3).getReg();
4964 unsigned Src0Size = MRI.getType(Src0Reg).getSizeInBits();
4965 unsigned Src1Size = MRI.getType(Src1Reg).getSizeInBits();
4966 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4967 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, DstSize);
4968 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Src0Size);
4969 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Src1Size);
4970 break;
4971 }
4972 case Intrinsic::amdgcn_icmp:
4973 case Intrinsic::amdgcn_fcmp: {
4974 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4975 // This is not VCCRegBank because this is not used in boolean contexts.
4976 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, DstSize);
4977 unsigned OpSize = MRI.getType(MI.getOperand(2).getReg()).getSizeInBits();
4978 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, OpSize);
4979 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, OpSize);
4980 break;
4981 }
4982 case Intrinsic::amdgcn_readlane: {
4983 // This must be an SGPR, but accept a VGPR.
4984 Register IdxReg = MI.getOperand(3).getReg();
4985 unsigned IdxSize = MRI.getType(IdxReg).getSizeInBits();
4986 unsigned IdxBank = getRegBankID(IdxReg, MRI, AMDGPU::SGPRRegBankID);
4987 OpdsMapping[3] = AMDGPU::getValueMapping(IdxBank, IdxSize);
4988 [[fallthrough]];
4989 }
4990 case Intrinsic::amdgcn_readfirstlane: {
4991 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4992 unsigned SrcSize = MRI.getType(MI.getOperand(2).getReg()).getSizeInBits();
4993 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, DstSize);
4994 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SrcSize);
4995 break;
4996 }
4997 case Intrinsic::amdgcn_writelane: {
4998 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
4999 Register SrcReg = MI.getOperand(2).getReg();
5000 unsigned SrcSize = MRI.getType(SrcReg).getSizeInBits();
5001 unsigned SrcBank = getRegBankID(SrcReg, MRI, AMDGPU::SGPRRegBankID);
5002 Register IdxReg = MI.getOperand(3).getReg();
5003 unsigned IdxSize = MRI.getType(IdxReg).getSizeInBits();
5004 unsigned IdxBank = getRegBankID(IdxReg, MRI, AMDGPU::SGPRRegBankID);
5005 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, DstSize);
5006
5007 // These 2 must be SGPRs, but accept VGPRs. Readfirstlane will be inserted
5008 // to legalize.
5009 OpdsMapping[2] = AMDGPU::getValueMapping(SrcBank, SrcSize);
5010 OpdsMapping[3] = AMDGPU::getValueMapping(IdxBank, IdxSize);
5011 OpdsMapping[4] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, SrcSize);
5012 break;
5013 }
5014 case Intrinsic::amdgcn_if_break: {
5015 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
5016 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
5017 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1);
5018 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
5019 break;
5020 }
5021 case Intrinsic::amdgcn_permlane16:
5022 case Intrinsic::amdgcn_permlanex16: {
5023 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
5024 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5025 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5026 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5027 OpdsMapping[4] = getSGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5028 OpdsMapping[5] = getSGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5029 break;
5030 }
5031 case Intrinsic::amdgcn_permlane_bcast:
5032 case Intrinsic::amdgcn_permlane_up:
5033 case Intrinsic::amdgcn_permlane_down:
5034 case Intrinsic::amdgcn_permlane_xor: {
5035 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
5036 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5037 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5038 OpdsMapping[3] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5039 OpdsMapping[4] = getSGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5040 break;
5041 }
5042 case Intrinsic::amdgcn_permlane_idx_gen: {
5043 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
5044 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5045 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5046 OpdsMapping[3] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5047 break;
5048 }
5049 case Intrinsic::amdgcn_permlane16_var:
5050 case Intrinsic::amdgcn_permlanex16_var: {
5051 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
5052 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5053 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5054 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5055 OpdsMapping[4] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5056 break;
5057 }
5058 case Intrinsic::amdgcn_mfma_f32_4x4x1f32:
5059 case Intrinsic::amdgcn_mfma_f32_4x4x4f16:
5060 case Intrinsic::amdgcn_mfma_i32_4x4x4i8:
5061 case Intrinsic::amdgcn_mfma_f32_4x4x2bf16:
5062 case Intrinsic::amdgcn_mfma_f32_16x16x1f32:
5063 case Intrinsic::amdgcn_mfma_f32_16x16x4f32:
5064 case Intrinsic::amdgcn_mfma_f32_16x16x4f16:
5065 case Intrinsic::amdgcn_mfma_f32_16x16x16f16:
5066 case Intrinsic::amdgcn_mfma_i32_16x16x4i8:
5067 case Intrinsic::amdgcn_mfma_i32_16x16x16i8:
5068 case Intrinsic::amdgcn_mfma_f32_16x16x2bf16:
5069 case Intrinsic::amdgcn_mfma_f32_16x16x8bf16:
5070 case Intrinsic::amdgcn_mfma_f32_32x32x1f32:
5071 case Intrinsic::amdgcn_mfma_f32_32x32x2f32:
5072 case Intrinsic::amdgcn_mfma_f32_32x32x4f16:
5073 case Intrinsic::amdgcn_mfma_f32_32x32x8f16:
5074 case Intrinsic::amdgcn_mfma_i32_32x32x4i8:
5075 case Intrinsic::amdgcn_mfma_i32_32x32x8i8:
5076 case Intrinsic::amdgcn_mfma_f32_32x32x2bf16:
5077 case Intrinsic::amdgcn_mfma_f32_32x32x4bf16:
5078 case Intrinsic::amdgcn_mfma_f32_32x32x4bf16_1k:
5079 case Intrinsic::amdgcn_mfma_f32_16x16x4bf16_1k:
5080 case Intrinsic::amdgcn_mfma_f32_4x4x4bf16_1k:
5081 case Intrinsic::amdgcn_mfma_f32_32x32x8bf16_1k:
5082 case Intrinsic::amdgcn_mfma_f32_16x16x16bf16_1k:
5083 case Intrinsic::amdgcn_mfma_f64_16x16x4f64:
5084 case Intrinsic::amdgcn_mfma_f64_4x4x4f64:
5085 case Intrinsic::amdgcn_mfma_i32_16x16x32_i8:
5086 case Intrinsic::amdgcn_mfma_i32_32x32x16_i8:
5087 case Intrinsic::amdgcn_mfma_f32_16x16x8_xf32:
5088 case Intrinsic::amdgcn_mfma_f32_32x32x4_xf32:
5089 case Intrinsic::amdgcn_mfma_f32_16x16x32_bf8_bf8:
5090 case Intrinsic::amdgcn_mfma_f32_16x16x32_bf8_fp8:
5091 case Intrinsic::amdgcn_mfma_f32_16x16x32_fp8_bf8:
5092 case Intrinsic::amdgcn_mfma_f32_16x16x32_fp8_fp8:
5093 case Intrinsic::amdgcn_mfma_f32_32x32x16_bf8_bf8:
5094 case Intrinsic::amdgcn_mfma_f32_32x32x16_bf8_fp8:
5095 case Intrinsic::amdgcn_mfma_f32_32x32x16_fp8_bf8:
5096 case Intrinsic::amdgcn_mfma_f32_32x32x16_fp8_fp8:
5097 case Intrinsic::amdgcn_mfma_f32_16x16x32_f16:
5098 case Intrinsic::amdgcn_mfma_f32_32x32x16_f16:
5099 case Intrinsic::amdgcn_mfma_i32_16x16x64_i8:
5100 case Intrinsic::amdgcn_mfma_i32_32x32x32_i8:
5101 case Intrinsic::amdgcn_mfma_f32_16x16x32_bf16: {
5102 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5103 unsigned MinNumRegsRequired = DstSize / 32;
5104
5105 // Default for MAI intrinsics.
5106 // srcC can also be an immediate which can be folded later.
5107 // FIXME: Should we eventually add an alternative mapping with AGPR src
5108 // for srcA/srcB?
5109 //
5110 // vdst, srcA, srcB, srcC
5112
5113 bool UseAGPRForm = !Subtarget.hasGFX90AInsts() ||
5114 Info->selectAGPRFormMFMA(MinNumRegsRequired);
5115
5116 OpdsMapping[0] =
5117 UseAGPRForm ? getAGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI)
5118 : getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5119 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5120 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5121 OpdsMapping[4] =
5122 UseAGPRForm ? getAGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI)
5123 : getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5124 break;
5125 }
5126 case Intrinsic::amdgcn_mfma_scale_f32_16x16x128_f8f6f4:
5127 case Intrinsic::amdgcn_mfma_scale_f32_32x32x64_f8f6f4: {
5128 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5129 unsigned MinNumRegsRequired = DstSize / 32;
5130
5132 bool UseAGPRForm = Info->selectAGPRFormMFMA(MinNumRegsRequired);
5133
5134 OpdsMapping[0] =
5135 UseAGPRForm ? getAGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI)
5136 : getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5137
5138 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5139 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5140 OpdsMapping[4] =
5141 UseAGPRForm ? getAGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI)
5142 : getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5143
5144 OpdsMapping[8] = getVGPROpMapping(MI.getOperand(8).getReg(), MRI, *TRI);
5145 OpdsMapping[10] = getVGPROpMapping(MI.getOperand(10).getReg(), MRI, *TRI);
5146 break;
5147 }
5148 case Intrinsic::amdgcn_smfmac_f32_16x16x32_f16:
5149 case Intrinsic::amdgcn_smfmac_f32_32x32x16_f16:
5150 case Intrinsic::amdgcn_smfmac_f32_16x16x32_bf16:
5151 case Intrinsic::amdgcn_smfmac_f32_32x32x16_bf16:
5152 case Intrinsic::amdgcn_smfmac_i32_16x16x64_i8:
5153 case Intrinsic::amdgcn_smfmac_i32_32x32x32_i8:
5154 case Intrinsic::amdgcn_smfmac_f32_16x16x64_bf8_bf8:
5155 case Intrinsic::amdgcn_smfmac_f32_16x16x64_bf8_fp8:
5156 case Intrinsic::amdgcn_smfmac_f32_16x16x64_fp8_bf8:
5157 case Intrinsic::amdgcn_smfmac_f32_16x16x64_fp8_fp8:
5158 case Intrinsic::amdgcn_smfmac_f32_32x32x32_bf8_bf8:
5159 case Intrinsic::amdgcn_smfmac_f32_32x32x32_bf8_fp8:
5160 case Intrinsic::amdgcn_smfmac_f32_32x32x32_fp8_bf8:
5161 case Intrinsic::amdgcn_smfmac_f32_32x32x32_fp8_fp8:
5162 case Intrinsic::amdgcn_smfmac_f32_16x16x64_f16:
5163 case Intrinsic::amdgcn_smfmac_f32_32x32x32_f16:
5164 case Intrinsic::amdgcn_smfmac_f32_16x16x64_bf16:
5165 case Intrinsic::amdgcn_smfmac_f32_32x32x32_bf16:
5166 case Intrinsic::amdgcn_smfmac_i32_16x16x128_i8:
5167 case Intrinsic::amdgcn_smfmac_i32_32x32x64_i8:
5168 case Intrinsic::amdgcn_smfmac_f32_16x16x128_bf8_bf8:
5169 case Intrinsic::amdgcn_smfmac_f32_16x16x128_bf8_fp8:
5170 case Intrinsic::amdgcn_smfmac_f32_16x16x128_fp8_bf8:
5171 case Intrinsic::amdgcn_smfmac_f32_16x16x128_fp8_fp8:
5172 case Intrinsic::amdgcn_smfmac_f32_32x32x64_bf8_bf8:
5173 case Intrinsic::amdgcn_smfmac_f32_32x32x64_bf8_fp8:
5174 case Intrinsic::amdgcn_smfmac_f32_32x32x64_fp8_bf8:
5175 case Intrinsic::amdgcn_smfmac_f32_32x32x64_fp8_fp8: {
5176 Register DstReg = MI.getOperand(0).getReg();
5177 unsigned DstSize = MRI.getType(DstReg).getSizeInBits();
5178 unsigned MinNumRegsRequired = DstSize / 32;
5180 bool UseAGPRForm = Info->selectAGPRFormMFMA(MinNumRegsRequired);
5181
5182 // vdst, srcA, srcB, srcC, idx
5183 OpdsMapping[0] = UseAGPRForm ? getAGPROpMapping(DstReg, MRI, *TRI)
5184 : getVGPROpMapping(DstReg, MRI, *TRI);
5185
5186 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5187 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5188 OpdsMapping[4] =
5189 UseAGPRForm ? getAGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI)
5190 : getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5191 OpdsMapping[5] = getVGPROpMapping(MI.getOperand(5).getReg(), MRI, *TRI);
5192 break;
5193 }
5194 case Intrinsic::amdgcn_interp_p1:
5195 case Intrinsic::amdgcn_interp_p2:
5196 case Intrinsic::amdgcn_interp_mov:
5197 case Intrinsic::amdgcn_interp_p1_f16:
5198 case Intrinsic::amdgcn_interp_p2_f16:
5199 case Intrinsic::amdgcn_lds_param_load: {
5200 const int M0Idx = MI.getNumOperands() - 1;
5201 Register M0Reg = MI.getOperand(M0Idx).getReg();
5202 unsigned M0Bank = getRegBankID(M0Reg, MRI, AMDGPU::SGPRRegBankID);
5203 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5204
5205 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, DstSize);
5206 for (int I = 2; I != M0Idx && MI.getOperand(I).isReg(); ++I)
5207 OpdsMapping[I] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5208
5209 // Must be SGPR, but we must take whatever the original bank is and fix it
5210 // later.
5211 OpdsMapping[M0Idx] = AMDGPU::getValueMapping(M0Bank, 32);
5212 break;
5213 }
5214 case Intrinsic::amdgcn_interp_inreg_p10:
5215 case Intrinsic::amdgcn_interp_inreg_p2:
5216 case Intrinsic::amdgcn_interp_inreg_p10_f16:
5217 case Intrinsic::amdgcn_interp_inreg_p2_f16:
5218 case Intrinsic::amdgcn_interp_p10_rtz_f16:
5219 case Intrinsic::amdgcn_interp_p2_rtz_f16: {
5220 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5221 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, DstSize);
5222 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5223 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5224 OpdsMapping[4] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5225 break;
5226 }
5227 case Intrinsic::amdgcn_permlane16_swap:
5228 case Intrinsic::amdgcn_permlane32_swap: {
5229 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5230 OpdsMapping[0] = OpdsMapping[1] = OpdsMapping[3] = OpdsMapping[4] =
5231 AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, DstSize);
5232 break;
5233 }
5234 case Intrinsic::amdgcn_ballot: {
5235 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5236 unsigned SrcSize = MRI.getType(MI.getOperand(2).getReg()).getSizeInBits();
5237 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, DstSize);
5238 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, SrcSize);
5239 break;
5240 }
5241 case Intrinsic::amdgcn_inverse_ballot: {
5242 // This must be an SGPR, but accept a VGPR.
5243 Register MaskReg = MI.getOperand(2).getReg();
5244 unsigned MaskSize = MRI.getType(MaskReg).getSizeInBits();
5245 unsigned MaskBank = getRegBankID(MaskReg, MRI, AMDGPU::SGPRRegBankID);
5246 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1);
5247 OpdsMapping[2] = AMDGPU::getValueMapping(MaskBank, MaskSize);
5248 break;
5249 }
5250 case Intrinsic::amdgcn_bitop3: {
5251 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
5252 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5253 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5254 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5255 OpdsMapping[4] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5256 break;
5257 }
5258 case Intrinsic::amdgcn_s_quadmask:
5259 case Intrinsic::amdgcn_s_wqm: {
5260 Register MaskReg = MI.getOperand(2).getReg();
5261 unsigned MaskSize = MRI.getType(MaskReg).getSizeInBits();
5262 unsigned MaskBank = getRegBankID(MaskReg, MRI, AMDGPU::SGPRRegBankID);
5263 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, MaskSize);
5264 OpdsMapping[2] = AMDGPU::getValueMapping(MaskBank, MaskSize);
5265 break;
5266 }
5267 case Intrinsic::amdgcn_wave_reduce_add:
5268 case Intrinsic::amdgcn_wave_reduce_fadd:
5269 case Intrinsic::amdgcn_wave_reduce_sub:
5270 case Intrinsic::amdgcn_wave_reduce_fsub:
5271 case Intrinsic::amdgcn_wave_reduce_min:
5272 case Intrinsic::amdgcn_wave_reduce_umin:
5273 case Intrinsic::amdgcn_wave_reduce_fmin:
5274 case Intrinsic::amdgcn_wave_reduce_max:
5275 case Intrinsic::amdgcn_wave_reduce_umax:
5276 case Intrinsic::amdgcn_wave_reduce_fmax:
5277 case Intrinsic::amdgcn_wave_reduce_and:
5278 case Intrinsic::amdgcn_wave_reduce_or:
5279 case Intrinsic::amdgcn_wave_reduce_xor: {
5280 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5281 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, DstSize);
5282 unsigned OpSize = MRI.getType(MI.getOperand(2).getReg()).getSizeInBits();
5283 auto regBankID =
5284 isSALUMapping(MI) ? AMDGPU::SGPRRegBankID : AMDGPU::VGPRRegBankID;
5285 OpdsMapping[2] = AMDGPU::getValueMapping(regBankID, OpSize);
5286 break;
5287 }
5288 case Intrinsic::amdgcn_s_bitreplicate: {
5289 Register MaskReg = MI.getOperand(2).getReg();
5290 unsigned MaskBank = getRegBankID(MaskReg, MRI, AMDGPU::SGPRRegBankID);
5291 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 64);
5292 OpdsMapping[2] = AMDGPU::getValueMapping(MaskBank, 32);
5293 break;
5294 }
5295 case Intrinsic::amdgcn_wave_shuffle: {
5296 unsigned OpSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5297 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, OpSize);
5298 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, OpSize);
5299 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, OpSize);
5300 break;
5301 }
5302 }
5303 break;
5304 }
5305 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_LOAD:
5306 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_LOAD_D16:
5307 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_LOAD_NORET:
5308 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_STORE:
5309 case AMDGPU::G_AMDGPU_INTRIN_IMAGE_STORE_D16: {
5310 auto IntrID = AMDGPU::getIntrinsicID(MI);
5311 const AMDGPU::RsrcIntrinsic *RSrcIntrin = AMDGPU::lookupRsrcIntrinsic(IntrID);
5312 assert(RSrcIntrin && "missing RsrcIntrinsic for image intrinsic");
5313 // Non-images can have complications from operands that allow both SGPR
5314 // and VGPR. For now it's too complicated to figure out the final opcode
5315 // to derive the register bank from the MCInstrDesc.
5316 assert(RSrcIntrin->IsImage);
5317 return getImageMapping(MRI, MI, RSrcIntrin->RsrcArg);
5318 }
5319 case AMDGPU::G_AMDGPU_BVH_INTERSECT_RAY:
5320 case AMDGPU::G_AMDGPU_BVH8_INTERSECT_RAY:
5321 case AMDGPU::G_AMDGPU_BVH_DUAL_INTERSECT_RAY: {
5322 bool IsDualOrBVH8 =
5323 MI.getOpcode() == AMDGPU::G_AMDGPU_BVH_DUAL_INTERSECT_RAY ||
5324 MI.getOpcode() == AMDGPU::G_AMDGPU_BVH8_INTERSECT_RAY;
5325 unsigned NumMods = IsDualOrBVH8 ? 0 : 1; // Has A16 modifier
5326 unsigned LastRegOpIdx = MI.getNumExplicitOperands() - 1 - NumMods;
5327 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5328 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, DstSize);
5329 if (IsDualOrBVH8) {
5330 OpdsMapping[1] = AMDGPU::getValueMapping(
5331 AMDGPU::VGPRRegBankID,
5332 MRI.getType(MI.getOperand(1).getReg()).getSizeInBits());
5333 OpdsMapping[2] = AMDGPU::getValueMapping(
5334 AMDGPU::VGPRRegBankID,
5335 MRI.getType(MI.getOperand(2).getReg()).getSizeInBits());
5336 }
5337 OpdsMapping[LastRegOpIdx] =
5338 getSGPROpMapping(MI.getOperand(LastRegOpIdx).getReg(), MRI, *TRI);
5339 if (LastRegOpIdx == 3) {
5340 // Sequential form: all operands combined into VGPR256/VGPR512
5341 unsigned Size = MRI.getType(MI.getOperand(2).getReg()).getSizeInBits();
5342 if (Size > 256)
5343 Size = 512;
5344 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5345 } else {
5346 // NSA form
5347 unsigned FirstSrcOpIdx = IsDualOrBVH8 ? 4 : 2;
5348 for (unsigned I = FirstSrcOpIdx; I < LastRegOpIdx; ++I) {
5349 unsigned Size = MRI.getType(MI.getOperand(I).getReg()).getSizeInBits();
5350 OpdsMapping[I] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5351 }
5352 }
5353 break;
5354 }
5355 case AMDGPU::G_INTRINSIC_W_SIDE_EFFECTS:
5356 case AMDGPU::G_INTRINSIC_CONVERGENT_W_SIDE_EFFECTS: {
5357 auto IntrID = cast<GIntrinsic>(MI).getIntrinsicID();
5358 switch (IntrID) {
5359 case Intrinsic::amdgcn_s_getreg:
5360 case Intrinsic::amdgcn_s_memtime:
5361 case Intrinsic::amdgcn_s_memrealtime:
5362 case Intrinsic::amdgcn_s_get_waveid_in_workgroup:
5363 case Intrinsic::amdgcn_s_sendmsg_rtn: {
5364 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5365 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
5366 break;
5367 }
5368 case Intrinsic::amdgcn_global_atomic_fmin_num:
5369 case Intrinsic::amdgcn_global_atomic_fmax_num:
5370 case Intrinsic::amdgcn_flat_atomic_fmin_num:
5371 case Intrinsic::amdgcn_flat_atomic_fmax_num:
5372 case Intrinsic::amdgcn_global_atomic_ordered_add_b64:
5373 case Intrinsic::amdgcn_global_load_tr_b64:
5374 case Intrinsic::amdgcn_global_load_tr_b128:
5375 case Intrinsic::amdgcn_global_load_tr4_b64:
5376 case Intrinsic::amdgcn_global_load_tr6_b96:
5377 case Intrinsic::amdgcn_ds_load_tr8_b64:
5378 case Intrinsic::amdgcn_ds_load_tr16_b128:
5379 case Intrinsic::amdgcn_ds_load_tr4_b64:
5380 case Intrinsic::amdgcn_ds_load_tr6_b96:
5381 case Intrinsic::amdgcn_ds_read_tr4_b64:
5382 case Intrinsic::amdgcn_ds_read_tr6_b96:
5383 case Intrinsic::amdgcn_ds_read_tr8_b64:
5384 case Intrinsic::amdgcn_ds_read_tr16_b64:
5385 case Intrinsic::amdgcn_ds_atomic_async_barrier_arrive_b64:
5386 case Intrinsic::amdgcn_ds_atomic_barrier_arrive_rtn_b64:
5388 case Intrinsic::amdgcn_ds_ordered_add:
5389 case Intrinsic::amdgcn_ds_ordered_swap: {
5390 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5391 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, DstSize);
5392 unsigned M0Bank = getRegBankID(MI.getOperand(2).getReg(), MRI,
5393 AMDGPU::SGPRRegBankID);
5394 OpdsMapping[2] = AMDGPU::getValueMapping(M0Bank, 32);
5395 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5396 break;
5397 }
5398 case Intrinsic::amdgcn_ds_append:
5399 case Intrinsic::amdgcn_ds_consume: {
5400 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5401 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, DstSize);
5402 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5403 break;
5404 }
5405 case Intrinsic::amdgcn_exp_compr:
5406 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5407 OpdsMapping[4] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5408 break;
5409 case Intrinsic::amdgcn_exp:
5410 // FIXME: Could we support packed types here?
5411 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5412 OpdsMapping[4] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5413 OpdsMapping[5] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5414 OpdsMapping[6] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5415 break;
5416 case Intrinsic::amdgcn_exp_row:
5417 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5418 OpdsMapping[4] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5419 OpdsMapping[5] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5420 OpdsMapping[6] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5421 OpdsMapping[8] = getSGPROpMapping(MI.getOperand(8).getReg(), MRI, *TRI);
5422 break;
5423 case Intrinsic::amdgcn_s_alloc_vgpr:
5424 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 1);
5425 OpdsMapping[2] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 32);
5426 break;
5427 case Intrinsic::amdgcn_s_sendmsg:
5428 case Intrinsic::amdgcn_s_sendmsghalt: {
5429 // This must be an SGPR, but accept a VGPR.
5430 unsigned Bank = getRegBankID(MI.getOperand(2).getReg(), MRI,
5431 AMDGPU::SGPRRegBankID);
5432 OpdsMapping[2] = AMDGPU::getValueMapping(Bank, 32);
5433 break;
5434 }
5435 case Intrinsic::amdgcn_s_setreg: {
5436 // This must be an SGPR, but accept a VGPR.
5437 unsigned Bank = getRegBankID(MI.getOperand(2).getReg(), MRI,
5438 AMDGPU::SGPRRegBankID);
5439 OpdsMapping[2] = AMDGPU::getValueMapping(Bank, 32);
5440 break;
5441 }
5442 case Intrinsic::amdgcn_s_ttracedata: {
5443 // This must be an SGPR, but accept a VGPR.
5444 unsigned Bank =
5445 getRegBankID(MI.getOperand(1).getReg(), MRI, AMDGPU::SGPRRegBankID);
5446 OpdsMapping[1] = AMDGPU::getValueMapping(Bank, 32);
5447 break;
5448 }
5449 case Intrinsic::amdgcn_end_cf: {
5450 unsigned Size = getSizeInBits(MI.getOperand(1).getReg(), MRI, *TRI);
5451 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
5452 break;
5453 }
5454 case Intrinsic::amdgcn_else: {
5455 unsigned WaveSize = getSizeInBits(MI.getOperand(1).getReg(), MRI, *TRI);
5456 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1);
5457 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, WaveSize);
5458 OpdsMapping[3] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, WaveSize);
5459 break;
5460 }
5461 case Intrinsic::amdgcn_init_whole_wave:
5462 case Intrinsic::amdgcn_live_mask: {
5463 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1);
5464 break;
5465 }
5466 case Intrinsic::amdgcn_wqm_demote:
5467 case Intrinsic::amdgcn_kill: {
5468 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1);
5469 break;
5470 }
5471 case Intrinsic::amdgcn_ptr_s_buffer_load: {
5472 OpdsMapping[0] = getSGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5473 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5474 OpdsMapping[3] = getSGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5475 break;
5476 }
5477 case Intrinsic::amdgcn_raw_buffer_load:
5478 case Intrinsic::amdgcn_raw_ptr_buffer_load:
5479 case Intrinsic::amdgcn_raw_atomic_buffer_load:
5480 case Intrinsic::amdgcn_raw_ptr_atomic_buffer_load:
5481 case Intrinsic::amdgcn_raw_tbuffer_load:
5482 case Intrinsic::amdgcn_raw_ptr_tbuffer_load: {
5483 // FIXME: Should make intrinsic ID the last operand of the instruction,
5484 // then this would be the same as store
5485 OpdsMapping[0] = getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5486 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5487 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5488 OpdsMapping[4] = getSGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5489 break;
5490 }
5491 case Intrinsic::amdgcn_raw_buffer_load_lds:
5492 case Intrinsic::amdgcn_raw_buffer_load_async_lds:
5493 case Intrinsic::amdgcn_raw_ptr_buffer_load_lds:
5494 case Intrinsic::amdgcn_raw_ptr_buffer_load_async_lds: {
5495 OpdsMapping[1] = getSGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5496 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5497 OpdsMapping[4] = getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5498 OpdsMapping[5] = getSGPROpMapping(MI.getOperand(5).getReg(), MRI, *TRI);
5499 break;
5500 }
5501 case Intrinsic::amdgcn_raw_buffer_store:
5502 case Intrinsic::amdgcn_raw_ptr_buffer_store:
5503 case Intrinsic::amdgcn_raw_buffer_store_format:
5504 case Intrinsic::amdgcn_raw_ptr_buffer_store_format:
5505 case Intrinsic::amdgcn_raw_tbuffer_store:
5506 case Intrinsic::amdgcn_raw_ptr_tbuffer_store: {
5507 OpdsMapping[1] = getVGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5508 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5509 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5510 OpdsMapping[4] = getSGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5511 break;
5512 }
5513 case Intrinsic::amdgcn_struct_buffer_load:
5514 case Intrinsic::amdgcn_struct_ptr_buffer_load:
5515 case Intrinsic::amdgcn_struct_tbuffer_load:
5516 case Intrinsic::amdgcn_struct_ptr_tbuffer_load:
5517 case Intrinsic::amdgcn_struct_atomic_buffer_load:
5518 case Intrinsic::amdgcn_struct_ptr_atomic_buffer_load: {
5519 OpdsMapping[0] = getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5520 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5521 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5522 OpdsMapping[4] = getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5523 OpdsMapping[5] = getSGPROpMapping(MI.getOperand(5).getReg(), MRI, *TRI);
5524 break;
5525 }
5526 case Intrinsic::amdgcn_struct_buffer_load_lds:
5527 case Intrinsic::amdgcn_struct_buffer_load_async_lds:
5528 case Intrinsic::amdgcn_struct_ptr_buffer_load_lds:
5529 case Intrinsic::amdgcn_struct_ptr_buffer_load_async_lds: {
5530 OpdsMapping[1] = getSGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5531 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5532 OpdsMapping[4] = getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5533 OpdsMapping[5] = getVGPROpMapping(MI.getOperand(5).getReg(), MRI, *TRI);
5534 OpdsMapping[6] = getSGPROpMapping(MI.getOperand(6).getReg(), MRI, *TRI);
5535 break;
5536 }
5537 case Intrinsic::amdgcn_struct_buffer_store:
5538 case Intrinsic::amdgcn_struct_ptr_buffer_store:
5539 case Intrinsic::amdgcn_struct_tbuffer_store:
5540 case Intrinsic::amdgcn_struct_ptr_tbuffer_store: {
5541 OpdsMapping[1] = getVGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5542 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5543 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5544 OpdsMapping[4] = getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI);
5545 OpdsMapping[5] = getSGPROpMapping(MI.getOperand(5).getReg(), MRI, *TRI);
5546 break;
5547 }
5548 case Intrinsic::amdgcn_init_exec_from_input: {
5549 unsigned Size = getSizeInBits(MI.getOperand(1).getReg(), MRI, *TRI);
5550 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, Size);
5551 break;
5552 }
5553 case Intrinsic::amdgcn_ds_gws_init:
5554 case Intrinsic::amdgcn_ds_gws_barrier:
5555 case Intrinsic::amdgcn_ds_gws_sema_br: {
5556 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5557
5558 // This must be an SGPR, but accept a VGPR.
5559 unsigned Bank = getRegBankID(MI.getOperand(2).getReg(), MRI,
5560 AMDGPU::SGPRRegBankID);
5561 OpdsMapping[2] = AMDGPU::getValueMapping(Bank, 32);
5562 break;
5563 }
5564 case Intrinsic::amdgcn_ds_gws_sema_v:
5565 case Intrinsic::amdgcn_ds_gws_sema_p:
5566 case Intrinsic::amdgcn_ds_gws_sema_release_all: {
5567 // This must be an SGPR, but accept a VGPR.
5568 unsigned Bank = getRegBankID(MI.getOperand(1).getReg(), MRI,
5569 AMDGPU::SGPRRegBankID);
5570 OpdsMapping[1] = AMDGPU::getValueMapping(Bank, 32);
5571 break;
5572 }
5573 case Intrinsic::amdgcn_cluster_load_b32:
5574 case Intrinsic::amdgcn_cluster_load_b64:
5575 case Intrinsic::amdgcn_cluster_load_b128: {
5576 OpdsMapping[0] = getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5577 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5578 unsigned M0Bank =
5579 getRegBankID(MI.getOperand(4).getReg(), MRI, AMDGPU::SGPRRegBankID);
5580 OpdsMapping[4] = AMDGPU::getValueMapping(M0Bank, 32);
5581 break;
5582 }
5583 case Intrinsic::amdgcn_cluster_load_async_to_lds_b8:
5584 case Intrinsic::amdgcn_cluster_load_async_to_lds_b32:
5585 case Intrinsic::amdgcn_cluster_load_async_to_lds_b64:
5586 case Intrinsic::amdgcn_cluster_load_async_to_lds_b128: {
5587 OpdsMapping[1] = getVGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5588 // LDS address goes into $vdst (VGPR).
5589 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5590 unsigned M0Bank =
5591 getRegBankID(MI.getOperand(5).getReg(), MRI, AMDGPU::SGPRRegBankID);
5592 OpdsMapping[5] = AMDGPU::getValueMapping(M0Bank, 32);
5593 break;
5594 }
5595 case Intrinsic::amdgcn_global_store_async_from_lds_b8:
5596 case Intrinsic::amdgcn_global_store_async_from_lds_b32:
5597 case Intrinsic::amdgcn_global_store_async_from_lds_b64:
5598 case Intrinsic::amdgcn_global_store_async_from_lds_b128:
5599 case Intrinsic::amdgcn_global_load_async_to_lds_b8:
5600 case Intrinsic::amdgcn_global_load_async_to_lds_b32:
5601 case Intrinsic::amdgcn_global_load_async_to_lds_b64:
5602 case Intrinsic::amdgcn_global_load_async_to_lds_b128: {
5603 OpdsMapping[1] = getVGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5604 // LDS address goes into $vdst/$vdata (VGPR).
5605 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5606 break;
5607 }
5608 case Intrinsic::amdgcn_load_to_lds:
5609 case Intrinsic::amdgcn_load_async_to_lds:
5610 case Intrinsic::amdgcn_global_load_lds:
5611 case Intrinsic::amdgcn_global_load_async_lds: {
5612 OpdsMapping[1] = getVGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5613 // LDS address goes into M0 (SGPR).
5614 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5615 break;
5616 }
5617 case Intrinsic::amdgcn_lds_direct_load: {
5618 const int M0Idx = MI.getNumOperands() - 1;
5619 Register M0Reg = MI.getOperand(M0Idx).getReg();
5620 unsigned M0Bank = getRegBankID(M0Reg, MRI, AMDGPU::SGPRRegBankID);
5621 unsigned DstSize = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5622
5623 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, DstSize);
5624 for (int I = 2; I != M0Idx && MI.getOperand(I).isReg(); ++I)
5625 OpdsMapping[I] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, 32);
5626
5627 // Must be SGPR, but we must take whatever the original bank is and fix it
5628 // later.
5629 OpdsMapping[M0Idx] = AMDGPU::getValueMapping(M0Bank, 32);
5630 break;
5631 }
5632 case Intrinsic::amdgcn_ds_add_gs_reg_rtn:
5633 case Intrinsic::amdgcn_ds_sub_gs_reg_rtn:
5634 OpdsMapping[0] = getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5635 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5636 break;
5637 case Intrinsic::amdgcn_ds_bvh_stack_rtn:
5638 case Intrinsic::amdgcn_ds_bvh_stack_push4_pop1_rtn:
5639 case Intrinsic::amdgcn_ds_bvh_stack_push8_pop1_rtn:
5640 case Intrinsic::amdgcn_ds_bvh_stack_push8_pop2_rtn: {
5641 OpdsMapping[0] =
5642 getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI); // %vdst
5643 OpdsMapping[1] =
5644 getVGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI); // %addr
5645 OpdsMapping[3] =
5646 getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI); // %addr
5647 OpdsMapping[4] =
5648 getVGPROpMapping(MI.getOperand(4).getReg(), MRI, *TRI); // %data0
5649 OpdsMapping[5] =
5650 getVGPROpMapping(MI.getOperand(5).getReg(), MRI, *TRI); // %data1
5651 break;
5652 }
5653 case Intrinsic::amdgcn_s_sleep_var:
5654 OpdsMapping[1] = getSGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5655 break;
5656 case Intrinsic::amdgcn_s_barrier_join:
5657 case Intrinsic::amdgcn_s_wakeup_barrier:
5658 OpdsMapping[1] = getSGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5659 break;
5660 case Intrinsic::amdgcn_s_barrier_init:
5661 case Intrinsic::amdgcn_s_barrier_signal_var:
5662 OpdsMapping[1] = getSGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5663 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5664 break;
5665 case Intrinsic::amdgcn_s_barrier_signal_isfirst: {
5666 const unsigned ResultSize = 1;
5667 OpdsMapping[0] =
5668 AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, ResultSize);
5669 break;
5670 }
5671 case Intrinsic::amdgcn_s_get_barrier_state:
5672 case Intrinsic::amdgcn_s_get_named_barrier_state: {
5673 OpdsMapping[0] = getSGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5674 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5675 break;
5676 }
5677 case Intrinsic::amdgcn_pops_exiting_wave_id:
5678 return getDefaultMappingSOP(MI);
5679 case Intrinsic::amdgcn_tensor_load_to_lds:
5680 case Intrinsic::amdgcn_tensor_store_from_lds: {
5681 // Lie and claim everything is legal, even all operands need to be
5682 // SGPRs. applyMapping will have to deal with it with readfirstlane.
5683 for (unsigned I = 1; I < MI.getNumOperands(); ++I) {
5684 if (MI.getOperand(I).isReg()) {
5685 Register Reg = MI.getOperand(I).getReg();
5686 auto OpBank = getRegBankID(Reg, MRI);
5687 unsigned Size = getSizeInBits(Reg, MRI, *TRI);
5688 OpdsMapping[I] = AMDGPU::getValueMapping(OpBank, Size);
5689 }
5690 }
5691 break;
5692 }
5693 case Intrinsic::amdgcn_s_prefetch_data:
5694 case Intrinsic::amdgcn_s_prefetch_inst: {
5695 OpdsMapping[1] = getSGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5696 OpdsMapping[2] = getSGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5697 break;
5698 }
5699 case Intrinsic::amdgcn_flat_prefetch:
5700 case Intrinsic::amdgcn_global_prefetch:
5701 return getDefaultMappingVOP(MI);
5702 default:
5704 }
5705 break;
5706 }
5707 case AMDGPU::G_SELECT: {
5708 unsigned Size = MRI.getType(MI.getOperand(0).getReg()).getSizeInBits();
5709 unsigned Op2Bank = getRegBankID(MI.getOperand(2).getReg(), MRI,
5710 AMDGPU::SGPRRegBankID);
5711 unsigned Op3Bank = getRegBankID(MI.getOperand(3).getReg(), MRI,
5712 AMDGPU::SGPRRegBankID);
5713 bool SGPRSrcs = Op2Bank == AMDGPU::SGPRRegBankID &&
5714 Op3Bank == AMDGPU::SGPRRegBankID;
5715
5716 unsigned CondBankDefault = SGPRSrcs ?
5717 AMDGPU::SGPRRegBankID : AMDGPU::VCCRegBankID;
5718 unsigned CondBank = getRegBankID(MI.getOperand(1).getReg(), MRI,
5719 CondBankDefault);
5720 if (CondBank == AMDGPU::SGPRRegBankID)
5721 CondBank = SGPRSrcs ? AMDGPU::SGPRRegBankID : AMDGPU::VCCRegBankID;
5722 else if (CondBank == AMDGPU::VGPRRegBankID)
5723 CondBank = AMDGPU::VCCRegBankID;
5724
5725 unsigned Bank = SGPRSrcs && CondBank == AMDGPU::SGPRRegBankID ?
5726 AMDGPU::SGPRRegBankID : AMDGPU::VGPRRegBankID;
5727
5728 assert(CondBank == AMDGPU::VCCRegBankID || CondBank == AMDGPU::SGPRRegBankID);
5729
5730 // TODO: Should report 32-bit for scalar condition type.
5731 if (Size == 64) {
5732 OpdsMapping[0] = AMDGPU::getValueMappingSGPR64Only(Bank, Size);
5733 OpdsMapping[1] = AMDGPU::getValueMapping(CondBank, 1);
5734 OpdsMapping[2] = AMDGPU::getValueMappingSGPR64Only(Bank, Size);
5735 OpdsMapping[3] = AMDGPU::getValueMappingSGPR64Only(Bank, Size);
5736 } else {
5737 OpdsMapping[0] = AMDGPU::getValueMapping(Bank, Size);
5738 OpdsMapping[1] = AMDGPU::getValueMapping(CondBank, 1);
5739 OpdsMapping[2] = AMDGPU::getValueMapping(Bank, Size);
5740 OpdsMapping[3] = AMDGPU::getValueMapping(Bank, Size);
5741 }
5742
5743 break;
5744 }
5745
5746 case AMDGPU::G_SI_CALL: {
5747 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::SGPRRegBankID, 64);
5748 // Lie and claim everything is legal, even though some need to be
5749 // SGPRs. applyMapping will have to deal with it as a waterfall loop.
5750 OpdsMapping[1] = getSGPROpMapping(MI.getOperand(1).getReg(), MRI, *TRI);
5751
5752 // Allow anything for implicit arguments
5753 for (unsigned I = 4; I < MI.getNumOperands(); ++I) {
5754 if (MI.getOperand(I).isReg()) {
5755 Register Reg = MI.getOperand(I).getReg();
5756 auto OpBank = getRegBankID(Reg, MRI);
5757 unsigned Size = getSizeInBits(Reg, MRI, *TRI);
5758 OpdsMapping[I] = AMDGPU::getValueMapping(OpBank, Size);
5759 }
5760 }
5761 break;
5762 }
5763 case AMDGPU::G_LOAD:
5764 case AMDGPU::G_ZEXTLOAD:
5765 case AMDGPU::G_SEXTLOAD:
5766 return getInstrMappingForLoad(MI);
5767
5768 case AMDGPU::G_ATOMICRMW_XCHG:
5769 case AMDGPU::G_ATOMICRMW_ADD:
5770 case AMDGPU::G_ATOMICRMW_SUB:
5771 case AMDGPU::G_ATOMICRMW_AND:
5772 case AMDGPU::G_ATOMICRMW_OR:
5773 case AMDGPU::G_ATOMICRMW_XOR:
5774 case AMDGPU::G_ATOMICRMW_MAX:
5775 case AMDGPU::G_ATOMICRMW_MIN:
5776 case AMDGPU::G_ATOMICRMW_UMAX:
5777 case AMDGPU::G_ATOMICRMW_UMIN:
5778 case AMDGPU::G_ATOMICRMW_FADD:
5779 case AMDGPU::G_ATOMICRMW_FMIN:
5780 case AMDGPU::G_ATOMICRMW_FMAX:
5781 case AMDGPU::G_ATOMICRMW_UINC_WRAP:
5782 case AMDGPU::G_ATOMICRMW_UDEC_WRAP:
5783 case AMDGPU::G_ATOMICRMW_USUB_COND:
5784 case AMDGPU::G_ATOMICRMW_USUB_SAT:
5785 case AMDGPU::G_AMDGPU_ATOMIC_CMPXCHG: {
5786 OpdsMapping[0] = getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5787 OpdsMapping[1] = getValueMappingForPtr(MRI, MI.getOperand(1).getReg());
5788 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5789 break;
5790 }
5791 case AMDGPU::G_ATOMIC_CMPXCHG: {
5792 OpdsMapping[0] = getVGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5793 OpdsMapping[1] = getValueMappingForPtr(MRI, MI.getOperand(1).getReg());
5794 OpdsMapping[2] = getVGPROpMapping(MI.getOperand(2).getReg(), MRI, *TRI);
5795 OpdsMapping[3] = getVGPROpMapping(MI.getOperand(3).getReg(), MRI, *TRI);
5796 break;
5797 }
5798 case AMDGPU::G_BRCOND: {
5799 unsigned Bank = getRegBankID(MI.getOperand(0).getReg(), MRI,
5800 AMDGPU::SGPRRegBankID);
5801 assert(MRI.getType(MI.getOperand(0).getReg()).getSizeInBits() == 1);
5802 if (Bank != AMDGPU::SGPRRegBankID)
5803 Bank = AMDGPU::VCCRegBankID;
5804
5805 OpdsMapping[0] = AMDGPU::getValueMapping(Bank, 1);
5806 break;
5807 }
5808 case AMDGPU::G_INTRINSIC_FPTRUNC_ROUND:
5809 return getDefaultMappingVOP(MI);
5810 case AMDGPU::G_PREFETCH:
5811 OpdsMapping[0] = getSGPROpMapping(MI.getOperand(0).getReg(), MRI, *TRI);
5812 break;
5813 case AMDGPU::G_AMDGPU_WHOLE_WAVE_FUNC_SETUP:
5814 case AMDGPU::G_AMDGPU_WHOLE_WAVE_FUNC_RETURN:
5815 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VCCRegBankID, 1);
5816 break;
5817 case AMDGPU::G_AMDGPU_FLAT_LOAD_MONITOR:
5818 case AMDGPU::G_AMDGPU_GLOBAL_LOAD_MONITOR: {
5819 unsigned Size = getSizeInBits(MI.getOperand(0).getReg(), MRI, *TRI);
5820 unsigned PtrSize = getSizeInBits(MI.getOperand(1).getReg(), MRI, *TRI);
5821 OpdsMapping[0] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, Size);
5822 OpdsMapping[1] = AMDGPU::getValueMapping(AMDGPU::VGPRRegBankID, PtrSize);
5823 break;
5824 }
5825 }
5826
5827 return getInstructionMapping(/*ID*/1, /*Cost*/1,
5828 getOperandsMapping(OpdsMapping),
5829 MI.getNumOperands());
5830}
static unsigned getIntrinsicID(const SDNode *N)
assert(UImm &&(UImm !=~static_cast< T >(0)) &&"Invalid immediate!")
Contains the definition of a TargetInstrInfo class that is common to all AMD GPUs.
constexpr LLT S16
constexpr LLT S1
constexpr LLT S32
constexpr LLT S64
AMDGPU Register Bank Select
static bool substituteSimpleCopyRegs(const AMDGPURegisterBankInfo::OperandsMapper &OpdMapper, unsigned OpIdx)
static unsigned regBankBoolUnion(unsigned RB0, unsigned RB1)
static std::pair< Register, unsigned > getBaseWithConstantOffset(MachineRegisterInfo &MRI, Register Reg)
static Register constrainRegToBank(MachineRegisterInfo &MRI, MachineIRBuilder &B, Register &Reg, const RegisterBank &Bank)
static std::pair< Register, Register > unpackV2S16ToS32(MachineIRBuilder &B, Register Src, unsigned ExtOpcode)
static void extendLow32IntoHigh32(MachineIRBuilder &B, Register Hi32Reg, Register Lo32Reg, unsigned ExtOpc, const RegisterBank &RegBank, bool IsBooleanSrc=false)
Implement extending a 32-bit value to a 64-bit value.
static unsigned getExtendOp(unsigned Opc)
static bool isVectorRegisterBank(const RegisterBank &Bank)
static unsigned regBankUnion(unsigned RB0, unsigned RB1)
static std::pair< LLT, LLT > splitUnequalType(LLT Ty, unsigned FirstSize)
Split Ty into 2 pieces.
static void setRegsToType(MachineRegisterInfo &MRI, ArrayRef< Register > Regs, LLT NewTy)
Replace the current type each register in Regs has with NewTy.
static void reinsertVectorIndexAdd(MachineIRBuilder &B, MachineInstr &IdxUseInstr, unsigned OpIdx, unsigned ConstOffset)
Utility function for pushing dynamic vector indexes with a constant offset into waterfall loops.
static LLT widen96To128(LLT Ty)
static LLT getHalfSizedType(LLT Ty)
static unsigned getSBufferLoadCorrespondingBufferLoadOpcode(unsigned Opc)
This file declares the targeting of the RegisterBankInfo class for AMDGPU.
Rewrite undef for PHI
MachineBasicBlock & MBB
MachineBasicBlock MachineBasicBlock::iterator DebugLoc DL
MachineBasicBlock MachineBasicBlock::iterator MBBI
#define X(NUM, ENUM, NAME)
Definition ELF.h:856
static GCRegistry::Add< OcamlGC > B("ocaml", "ocaml 3.10-compatible GC")
AMD GCN specific subclass of TargetSubtarget.
Declares convenience wrapper classes for interpreting MachineInstr instances as specific generic oper...
IRTranslator LLVM IR MI
const size_t AbstractManglingParser< Derived, Alloc >::NumOps
const AbstractManglingParser< Derived, Alloc >::OperatorInfo AbstractManglingParser< Derived, Alloc >::Ops[]
#define I(x, y, z)
Definition MD5.cpp:57
Contains matchers for matching SSA Machine Instructions.
This file declares the MachineIRBuilder class.
Register Reg
Promote Memory to Register
Definition Mem2Reg.cpp:110
static bool isReg(const MCInst &MI, unsigned OpNo)
MachineInstr unsigned OpIdx
ConstantRange Range(APInt(BitWidth, Low), APInt(BitWidth, High))
static constexpr MCPhysReg SPReg
Interface definition for SIRegisterInfo.
static TableGen::Emitter::Opt Y("gen-skeleton-entry", EmitSkeleton, "Generate example skeleton entry")
bool applyMappingDynStackAlloc(MachineIRBuilder &B, const OperandsMapper &OpdMapper, MachineInstr &MI) const
std::pair< Register, unsigned > splitBufferOffsets(MachineIRBuilder &B, Register Offset) const
bool collectWaterfallOperands(SmallSet< Register, 4 > &SGPROperandRegs, MachineInstr &MI, MachineRegisterInfo &MRI, ArrayRef< unsigned > OpIndices) const
const InstructionMapping & getImageMapping(const MachineRegisterInfo &MRI, const MachineInstr &MI, int RsrcIdx) const
InstructionMappings addMappingFromTable(const MachineInstr &MI, const MachineRegisterInfo &MRI, const std::array< unsigned, NumOps > RegSrcOpIdx, ArrayRef< OpRegBankEntry< NumOps > > Table) const
unsigned copyCost(const RegisterBank &A, const RegisterBank &B, TypeSize Size) const override
Get the cost of a copy from B to A, or put differently, get the cost of A = COPY B.
RegisterBankInfo::InstructionMappings getInstrAlternativeMappingsIntrinsicWSideEffects(const MachineInstr &MI, const MachineRegisterInfo &MRI) const
bool buildVCopy(MachineIRBuilder &B, Register DstReg, Register SrcReg) const
bool executeInWaterfallLoop(MachineIRBuilder &B, iterator_range< MachineBasicBlock::iterator > Range, SmallSet< Register, 4 > &SGPROperandRegs) const
Legalize instruction MI where operands in OpIndices must be SGPRs.
const RegisterBank & getRegBankFromRegClass(const TargetRegisterClass &RC, LLT) const override
Get a register bank that covers RC.
AMDGPURegisterBankInfo(const GCNSubtarget &STI)
bool applyMappingMAD_64_32(MachineIRBuilder &B, const OperandsMapper &OpdMapper) const
unsigned getRegBankID(Register Reg, const MachineRegisterInfo &MRI, unsigned Default=AMDGPU::VGPRRegBankID) const
Register handleD16VData(MachineIRBuilder &B, MachineRegisterInfo &MRI, Register Reg) const
Handle register layout difference for f16 images for some subtargets.
const RegisterBankInfo::InstructionMapping & getInstrMappingForLoad(const MachineInstr &MI) const
void applyMappingImpl(MachineIRBuilder &Builder, const OperandsMapper &OpdMapper) const override
See RegisterBankInfo::applyMapping.
bool applyMappingBFE(MachineIRBuilder &B, const OperandsMapper &OpdMapper, bool Signed) const
bool applyMappingImage(MachineIRBuilder &B, MachineInstr &MI, const OperandsMapper &OpdMapper, int RSrcIdx) const
const ValueMapping * getVGPROpMapping(Register Reg, const MachineRegisterInfo &MRI, const TargetRegisterInfo &TRI) const
bool isScalarLoadLegal(const MachineInstr &MI) const
unsigned setBufferOffsets(MachineIRBuilder &B, Register CombinedOffset, Register &VOffsetReg, Register &SOffsetReg, int64_t &InstOffsetVal, Align Alignment) const
const ValueMapping * getSGPROpMapping(Register Reg, const MachineRegisterInfo &MRI, const TargetRegisterInfo &TRI) const
bool applyMappingLoad(MachineIRBuilder &B, const OperandsMapper &OpdMapper, MachineInstr &MI) const
void split64BitValueForMapping(MachineIRBuilder &B, SmallVector< Register, 2 > &Regs, LLT HalfTy, Register Reg) const
Split 64-bit value Reg into two 32-bit halves and populate them into Regs.
const ValueMapping * getValueMappingForPtr(const MachineRegisterInfo &MRI, Register Ptr) const
Return the mapping for a pointer argument.
unsigned getMappingType(const MachineRegisterInfo &MRI, const MachineInstr &MI) const
RegisterBankInfo::InstructionMappings getInstrAlternativeMappingsIntrinsic(const MachineInstr &MI, const MachineRegisterInfo &MRI) const
bool isDivergentRegBank(const RegisterBank *RB) const override
Returns true if the register bank is considered divergent.
void constrainOpWithReadfirstlane(MachineIRBuilder &B, MachineInstr &MI, unsigned OpIdx) const
InstructionMappings getInstrAlternativeMappings(const MachineInstr &MI) const override
Get the alternative mappings for MI.
const InstructionMapping & getDefaultMappingSOP(const MachineInstr &MI) const
const InstructionMapping & getDefaultMappingAllVGPR(const MachineInstr &MI) const
const InstructionMapping & getInstrMapping(const MachineInstr &MI) const override
This function must return a legal mapping, because AMDGPURegisterBankInfo::getInstrAlternativeMapping...
unsigned getBreakDownCost(const ValueMapping &ValMapping, const RegisterBank *CurBank=nullptr) const override
Get the cost of using ValMapping to decompose a register.
const ValueMapping * getAGPROpMapping(Register Reg, const MachineRegisterInfo &MRI, const TargetRegisterInfo &TRI) const
const InstructionMapping & getDefaultMappingVOP(const MachineInstr &MI) const
bool isSALUMapping(const MachineInstr &MI) const
Register buildReadFirstLane(MachineIRBuilder &B, MachineRegisterInfo &MRI, Register Src) const
bool applyMappingSBufferLoad(MachineIRBuilder &B, const OperandsMapper &OpdMapper) const
void applyMappingSMULU64(MachineIRBuilder &B, const OperandsMapper &OpdMapper) const
static const LaneMaskConstants & get(const GCNSubtarget &ST)
Represent a constant reference to an array (0 or more elements consecutively in memory),...
Definition ArrayRef.h:40
Predicate
This enumeration lists the possible predicates for CmpInst subclasses.
Definition InstrTypes.h:740
@ ICMP_SLT
signed less than
Definition InstrTypes.h:769
@ ICMP_NE
not equal
Definition InstrTypes.h:762
A debug info location.
Definition DebugLoc.h:126
iterator find(const_arg_type_t< KeyT > Val)
Definition DenseMap.h:223
iterator end()
Definition DenseMap.h:141
std::pair< iterator, bool > insert(const std::pair< KeyT, ValueT > &KV)
Definition DenseMap.h:284
static constexpr ElementCount getFixed(ScalarTy MinVal)
Definition TypeSize.h:309
Abstract class that contains various methods for clients to notify about changes.
constexpr unsigned getScalarSizeInBits() const
constexpr bool isScalar() const
LLT getScalarType() const
static constexpr LLT scalar(unsigned SizeInBits)
Get a low-level scalar or aggregate "bag of bits".
constexpr uint16_t getNumElements() const
Returns the number of elements in a vector LLT.
constexpr bool isVector() const
constexpr TypeSize getSizeInBits() const
Returns the total size of the type. Must only be called on sized types.
LLT divide(int Factor) const
Return a type that is Factor times smaller.
constexpr unsigned getAddressSpace() const
static constexpr LLT fixed_vector(unsigned NumElements, unsigned ScalarSizeInBits)
Get a low-level fixed-width vector of some number of elements and element width.
LLT getElementType() const
Returns the vector's element type. Only valid for vector types.
static constexpr LLT scalarOrVector(ElementCount EC, LLT ScalarTy)
This is an important class for using LLVM in a threaded context.
Definition LLVMContext.h:68
LLVM_ABI void widenScalarSrc(MachineInstr &MI, LLT WideTy, unsigned OpIdx, unsigned ExtOpcode)
Legalize a single operand OpIdx of the machine instruction MI as a Use by extending the operand's typ...
LLVM_ABI LegalizeResult lowerAbsToMaxNeg(MachineInstr &MI)
LLVM_ABI LegalizeResult narrowScalar(MachineInstr &MI, unsigned TypeIdx, LLT NarrowTy)
Legalize an instruction by reducing the width of the underlying scalar type.
LLVM_ABI LegalizeResult reduceLoadStoreWidth(GLoadStore &MI, unsigned TypeIdx, LLT NarrowTy)
@ Legalized
Instruction has been legalized and the MachineFunction changed.
LLVM_ABI LegalizeResult fewerElementsVector(MachineInstr &MI, unsigned TypeIdx, LLT NarrowTy)
Legalize a vector instruction by splitting into multiple components, each acting on the same scalar t...
LLVM_ABI LegalizeResult widenScalar(MachineInstr &MI, unsigned TypeIdx, LLT WideTy)
Legalize an instruction by performing the operation on a wider scalar type (for example a 16-bit addi...
LLVM_ABI void widenScalarDst(MachineInstr &MI, LLT WideTy, unsigned OpIdx=0, unsigned TruncOpcode=TargetOpcode::G_TRUNC)
Legalize a single operand OpIdx of the machine instruction MI as a Def by extending the operand's typ...
TypeSize getValue() const
LLVM_ABI void transferSuccessorsAndUpdatePHIs(MachineBasicBlock *FromMBB)
Transfers all the successors, as in transferSuccessors, and update PHI operands in the successor bloc...
LLVM_ABI iterator getFirstTerminator()
Returns an iterator to the first terminator instruction of this basic block.
LLVM_ABI void addSuccessor(MachineBasicBlock *Succ, BranchProbability Prob=BranchProbability::getUnknown())
Add Succ as a successor of this MachineBasicBlock.
const MachineFunction * getParent() const
Return the MachineFunction containing this basic block.
void splice(iterator Where, MachineBasicBlock *Other, iterator From)
Take an instruction from MBB 'Other' at the position From, and insert it into this MBB right before '...
MachineInstrBundleIterator< MachineInstr > iterator
const TargetSubtargetInfo & getSubtarget() const
getSubtarget - Return the subtarget for which this machine code is being compiled.
MachineMemOperand * getMachineMemOperand(MachinePointerInfo PtrInfo, MachineMemOperand::Flags f, LLT MemTy, Align base_alignment, const AAMDNodes &AAInfo=AAMDNodes(), const MDNode *Ranges=nullptr, SyncScope::ID SSID=SyncScope::System, AtomicOrdering Ordering=AtomicOrdering::NotAtomic, AtomicOrdering FailureOrdering=AtomicOrdering::NotAtomic)
getMachineMemOperand - Allocate a new MachineMemOperand.
MachineRegisterInfo & getRegInfo()
getRegInfo - Return information about the registers currently in use.
BasicBlockListType::iterator iterator
Ty * getInfo()
getInfo - Keep track of various per-function pieces of information for backends that would like to do...
MachineBasicBlock * CreateMachineBasicBlock(const BasicBlock *BB=nullptr, std::optional< UniqueBBID > BBID=std::nullopt)
CreateMachineInstr - Allocate a new MachineInstr.
void insert(iterator MBBI, MachineBasicBlock *MBB)
Helper class to build MachineInstr.
const MachineInstrBuilder & addReg(Register RegNo, RegState Flags={}, unsigned SubReg=0) const
Add a new virtual register operand.
MachineInstrSpan provides an interface to get an iteration range containing the instruction it was in...
MachineBasicBlock::iterator begin()
MachineBasicBlock::iterator end()
Representation of each machine instruction.
const MachineBasicBlock * getParent() const
const MachineOperand & getOperand(unsigned i) const
A description of a memory reference used in the backend.
LocationSize getSize() const
Return the size in bytes of the memory reference.
unsigned getAddrSpace() const
bool isAtomic() const
Returns true if this operation has an atomic ordering requirement of unordered or higher,...
@ MODereferenceable
The memory access is dereferenceable (i.e., doesn't trap).
@ MOLoad
The memory access reads data.
@ MOInvariant
The memory access always returns the same value (or traps).
Flags getFlags() const
Return the raw flags of the source value,.
LLVM_ABI Align getAlign() const
Return the minimum known alignment in bytes of the actual memory reference.
MachineOperand class - Representation of each machine instruction operand.
LLVM_ABI void setReg(Register Reg)
Change the register this operand corresponds to.
Register getReg() const
getReg - Returns the register number.
MachineRegisterInfo - Keep track of information for virtual and physical registers,...
const RegClassOrRegBank & getRegClassOrRegBank(Register Reg) const
Return the register bank or register class of Reg.
LLVM_ABI Register createVirtualRegister(const TargetRegisterClass *RegClass, StringRef Name="")
createVirtualRegister - Create and return a new virtual register in the function with the specified r...
LLT getType(Register Reg) const
Get the low-level type of Reg or LLT{} if Reg is not a generic (target independent) virtual register.
const RegisterBank * getRegBankOrNull(Register Reg) const
Return the register bank of Reg, or null if Reg has not been assigned a register bank or has been ass...
LLVM_ABI void setRegBank(Register Reg, const RegisterBank &RegBank)
Set the register bank to RegBank for Reg.
LLVM_ABI void setType(Register VReg, LLT Ty)
Set the low-level type of VReg to Ty.
LLVM_ABI void setRegClass(Register Reg, const TargetRegisterClass *RC)
setRegClass - Set the register class of the specified virtual register.
LLVM_ABI Register createGenericVirtualRegister(LLT Ty, StringRef Name="")
Create and return a new generic virtual register with low-level type Ty.
void setSimpleHint(Register VReg, Register PrefReg)
Specify the preferred (target independent) register allocation hint for the specified virtual registe...
Helper class that represents how the value of an instruction may be mapped and what is the related co...
bool isValid() const
Check whether this object is valid.
Helper class used to get/create the virtual registers that will be used to replace the MachineOperand...
const InstructionMapping & getInstrMapping() const
The final mapping of the instruction.
MachineRegisterInfo & getMRI() const
The MachineRegisterInfo we used to realize the mapping.
LLVM_ABI iterator_range< SmallVectorImpl< Register >::const_iterator > getVRegs(unsigned OpIdx, bool ForDebug=false) const
Get all the virtual registers required to map the OpIdx-th operand of the instruction.
virtual InstructionMappings getInstrAlternativeMappings(const MachineInstr &MI) const
Get the alternative mappings for MI.
static const TargetRegisterClass * constrainGenericRegister(Register Reg, const TargetRegisterClass &RC, MachineRegisterInfo &MRI)
Constrain the (possibly generic) virtual register Reg to RC.
const InstructionMapping & getInstructionMapping(unsigned ID, unsigned Cost, const ValueMapping *OperandsMapping, unsigned NumOperands) const
Method to get a uniquely generated InstructionMapping.
static void applyDefaultMapping(const OperandsMapper &OpdMapper)
Helper method to apply something that is like the default mapping.
const ValueMapping & getValueMapping(unsigned StartIdx, unsigned Length, const RegisterBank &RegBank) const
The most common ValueMapping consists of a single PartialMapping.
const InstructionMapping & getInvalidInstructionMapping() const
Method to get a uniquely generated invalid InstructionMapping.
const RegisterBank & getRegBank(unsigned ID)
Get the register bank identified by ID.
const unsigned * Sizes
Hold the sizes of the register banks for all HwModes.
bool cannotCopy(const RegisterBank &Dst, const RegisterBank &Src, TypeSize Size) const
TypeSize getSizeInBits(Register Reg, const MachineRegisterInfo &MRI, const TargetRegisterInfo &TRI) const
Get the size in bits of Reg.
const ValueMapping * getOperandsMapping(Iterator Begin, Iterator End) const
Get the uniquely generated array of ValueMapping for the elements of between Begin and End.
SmallVector< const InstructionMapping *, 4 > InstructionMappings
Convenient type to represent the alternatives for mapping an instruction.
virtual unsigned copyCost(const RegisterBank &A, const RegisterBank &B, TypeSize Size) const
Get the cost of a copy from B to A, or put differently, get the cost of A = COPY B.
const InstructionMapping & getInstrMappingImpl(const MachineInstr &MI) const
Try to get the mapping of MI.
This class implements the register bank concept.
unsigned getID() const
Get the identifier of this register bank.
Wrapper class representing virtual and physical registers.
Definition Register.h:20
constexpr bool isVirtual() const
Return true if the specified register number is in the virtual register namespace.
Definition Register.h:79
static unsigned getMaxMUBUFImmOffset(const GCNSubtarget &ST)
This class keeps track of the SPI_SP_INPUT_ADDR config register, which tells the hardware which inter...
bool selectAGPRFormMFMA(unsigned NumRegs) const
Return true if an MFMA that requires at least NumRegs should select to the AGPR form,...
static bool shouldExpandVectorDynExt(unsigned EltSize, unsigned NumElem, bool IsDivergentIdx, const GCNSubtarget *Subtarget)
Check if EXTRACT_VECTOR_ELT/INSERT_VECTOR_ELT (<n x e>, var-idx) should be expanded into a set of cmp...
SmallSet - This maintains a set of unique values, optimizing for the case when the set is small (less...
Definition SmallSet.h:134
size_type count(const T &V) const
count - Return 1 if the element is in the set, 0 otherwise.
Definition SmallSet.h:176
bool empty() const
Definition SmallSet.h:169
std::pair< const_iterator, bool > insert(const T &V)
insert - Insert an element into the set if it isn't already there.
Definition SmallSet.h:184
void resize(size_type N)
void push_back(const T &Elt)
This is a 'vector' (really, a variable-sized array), optimized for the case when the array is small.
Register getReg() const
TargetRegisterInfo base class - We assume that the target defines a static array of TargetRegisterDes...
static constexpr TypeSize getFixed(ScalarTy ExactSize)
Definition TypeSize.h:343
static LLVM_ABI IntegerType * getInt32Ty(LLVMContext &C)
Definition Type.cpp:309
self_iterator getIterator()
Definition ilist_node.h:123
A range adaptor for a pair of iterators.
#define llvm_unreachable(msg)
Marks that the current location is not supposed to be reachable.
@ CONSTANT_ADDRESS_32BIT
Address space for 32-bit constant memory.
@ REGION_ADDRESS
Address space for region memory. (GDS)
@ LOCAL_ADDRESS
Address space for local memory.
@ CONSTANT_ADDRESS
Address space for constant memory (VTX2).
@ PRIVATE_ADDRESS
Address space for private memory.
@ BUFFER_RESOURCE
Address space for 128-bit buffer resources.
bool isFlatGlobalAddrSpace(unsigned AS)
bool isUniformMMO(const MachineMemOperand *MMO)
bool isExtendedGlobalAddrSpace(unsigned AS)
Intrinsic::ID getIntrinsicID(const MachineInstr &I)
Return the intrinsic ID for opcodes with the G_AMDGPU_INTRIN_ prefix.
std::pair< Register, unsigned > getBaseWithConstantOffset(MachineRegisterInfo &MRI, Register Reg, GISelValueTracking *ValueTracking=nullptr, bool CheckNUW=false)
Returns base register and constant offset.
const RsrcIntrinsic * lookupRsrcIntrinsic(unsigned Intr)
operand_type_match m_Reg()
SpecificConstantMatch m_ZeroInt()
Convenience matchers for specific integer values.
ConstantMatch< APInt > m_ICst(APInt &Cst)
BinaryOp_match< LHS, RHS, TargetOpcode::G_ADD, true > m_GAdd(const LHS &L, const RHS &R)
bool mi_match(Reg R, const MachineRegisterInfo &MRI, Pattern &&P)
SpecificConstantOrSplatMatch m_SpecificICstOrSplat(const APInt &RequestedValue)
Matches a RequestedValue constant or a constant splat of RequestedValue.
This is an optimization pass for GlobalISel generic memory operations.
@ Offset
Definition DWP.cpp:578
LLVM_ABI MachineInstr * getOpcodeDef(unsigned Opcode, Register Reg, const MachineRegisterInfo &MRI)
See if Reg is defined by an single def instruction that is Opcode.
Definition Utils.cpp:656
MachineInstrBuilder BuildMI(MachineFunction &MF, const MIMetadata &MIMD, const MCInstrDesc &MCID)
Builder interface. Specify how to create the initial instruction itself.
@ Kill
The last use of a register.
decltype(auto) dyn_cast(const From &Val)
dyn_cast<X> - Return the argument parameter cast to the specified type.
Definition Casting.h:643
LLVM_ABI void constrainSelectedInstRegOperands(MachineInstr &I, const TargetInstrInfo &TII, const TargetRegisterInfo &TRI, const RegisterBankInfo &RBI)
Mutate the newly-selected instruction I to constrain its (possibly generic) virtual register operands...
Definition Utils.cpp:159
iterator_range< T > make_range(T x, T y)
Convenience function for iterating over sub-ranges.
LLVM_ABI std::optional< int64_t > getIConstantVRegSExtVal(Register VReg, const MachineRegisterInfo &MRI)
If VReg is defined by a G_CONSTANT fits in int64_t returns it.
Definition Utils.cpp:317
static const MachineMemOperand::Flags MONoClobber
Mark the MMO of a uniform load if there are no potentially clobbering stores on any path from the sta...
Definition SIInstrInfo.h:46
auto reverse(ContainerTy &&C)
Definition STLExtras.h:407
class LLVM_GSL_OWNER SmallVector
Forward declaration of SmallVector so that calculateSmallVectorDefaultInlinedElements can reference s...
bool isa(const From &Val)
isa<X> - Return true if the parameter to the template is an instance of one of the template type argu...
Definition Casting.h:547
@ Add
Sum of integers.
IntPtrTy
Definition InstrProf.h:82
DWARFExpression::Operation Op
void call_once(once_flag &flag, Function &&F, Args &&... ArgList)
Execute the function specified as a parameter once.
Definition Threading.h:86
decltype(auto) cast(const From &Val)
cast<X> - Return the argument parameter cast to the specified type.
Definition Casting.h:559
LLVM_ABI std::optional< ValueAndVReg > getIConstantVRegValWithLookThrough(Register VReg, const MachineRegisterInfo &MRI, bool LookThroughInstrs=true)
If VReg is defined by a statically evaluable chain of instructions rooted on a G_CONSTANT returns its...
Definition Utils.cpp:436
Align assumeAligned(uint64_t Value)
Treats the value 0 as a 1, so Align is always at least 1.
Definition Alignment.h:100
unsigned Log2(Align A)
Returns the log2 of the alignment.
Definition Alignment.h:197
LLVM_ABI Register getSrcRegIgnoringCopies(Register Reg, const MachineRegisterInfo &MRI)
Find the source register for Reg, folding away any trivial copies.
Definition Utils.cpp:504
constexpr T maskTrailingOnes(unsigned N)
Create a bitmask with the N right-most bits set to 1, and all other bits set to 0.
Definition MathExtras.h:78
@ Default
The result value is uniform if and only if all operands are uniform.
Definition Uniformity.h:20
MCRegisterClass TargetRegisterClass
Definition FastISel.h:58
#define N
This struct is a compact representation of a valid (non-zero power of two) alignment.
Definition Alignment.h:39
constexpr uint64_t value() const
This is a hole in the type system and should not be abused.
Definition Alignment.h:77
This class contains a discriminated union of information about pointers in memory operands,...
unsigned StartIdx
Number of bits at which this partial mapping starts in the original value.
const RegisterBank * RegBank
Register bank where the partial value lives.
unsigned Length
Length of this mapping in bits.
Helper struct that represents how a value is mapped through different register banks.
unsigned NumBreakDowns
Number of partial mapping to break down this value.
const PartialMapping * BreakDown
How the value is broken down between the different register banks.
The llvm::once_flag structure.
Definition Threading.h:67