LLVM 24.0.0git
AArch64FrameLowering.cpp
Go to the documentation of this file.
1//===- AArch64FrameLowering.cpp - AArch64 Frame Lowering -------*- C++ -*-====//
2//
3// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
4// See https://llvm.org/LICENSE.txt for license information.
5// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
6//
7//===----------------------------------------------------------------------===//
8//
9// This file contains the AArch64 implementation of TargetFrameLowering class.
10//
11// On AArch64, stack frames are structured as follows:
12//
13// The stack grows downward.
14//
15// All of the individual frame areas on the frame below are optional, i.e. it's
16// possible to create a function so that the particular area isn't present
17// in the frame.
18//
19// At function entry, the "frame" looks as follows:
20//
21// | | Higher address
22// |-----------------------------------|
23// | |
24// | arguments passed on the stack |
25// | |
26// |-----------------------------------| <- sp
27// | | Lower address
28//
29//
30// After the prologue has run, the frame has the following general structure.
31// Note that this doesn't depict the case where a red-zone is used. Also,
32// technically the last frame area (VLAs) doesn't get created until in the
33// main function body, after the prologue is run. However, it's depicted here
34// for completeness.
35//
36// | | Higher address
37// |-----------------------------------|
38// | |
39// | arguments passed on the stack |
40// | |
41// |-----------------------------------|
42// | |
43// | (Win64 only) varargs from reg |
44// | |
45// |-----------------------------------|
46// | |
47// | (Win64 only) callee-saved SVE reg |
48// | |
49// |-----------------------------------|
50// | |
51// | callee-saved gpr registers | <--.
52// | | | On Darwin platforms these
53// |- - - - - - - - - - - - - - - - - -| | callee saves are swapped,
54// | prev_lr | | (frame record first)
55// | prev_fp | <--'
56// | async context if needed |
57// | (a.k.a. "frame record") |
58// |-----------------------------------| <- fp(=x29)
59// Default SVE stack layout Split SVE objects
60// (aarch64-split-sve-objects=false) (aarch64-split-sve-objects=true)
61// |-----------------------------------| |-----------------------------------|
62// | <hazard padding> | | callee-saved PPR registers |
63// |-----------------------------------| |-----------------------------------|
64// | | | PPR stack objects |
65// | callee-saved fp/simd/SVE regs | |-----------------------------------|
66// | | | <hazard padding> |
67// |-----------------------------------| |-----------------------------------|
68// | | | callee-saved ZPR/FPR registers |
69// | SVE stack objects | |-----------------------------------|
70// | | | ZPR stack objects |
71// |-----------------------------------| |-----------------------------------|
72// ^ NB: FPR CSRs are promoted to ZPRs
73// |-----------------------------------|
74// |.empty.space.to.make.part.below....|
75// |.aligned.in.case.it.needs.more.than| (size of this area is unknown at
76// |.the.standard.16-byte.alignment....| compile time; if present)
77// |-----------------------------------|
78// | local variables of fixed size |
79// | including spill slots |
80// | <FPR> |
81// | <hazard padding> |
82// | <GPR> |
83// |-----------------------------------| <- bp(not defined by ABI,
84// |.variable-sized.local.variables....| LLVM chooses X19)
85// |.(VLAs)............................| (size of this area is unknown at
86// |...................................| compile time)
87// |-----------------------------------| <- sp
88// | | Lower address
89//
90//
91// To access the data in a frame, at-compile time, a constant offset must be
92// computable from one of the pointers (fp, bp, sp) to access it. The size
93// of the areas with a dotted background cannot be computed at compile-time
94// if they are present, making it required to have all three of fp, bp and
95// sp to be set up to be able to access all contents in the frame areas,
96// assuming all of the frame areas are non-empty.
97//
98// For most functions, some of the frame areas are empty. For those functions,
99// it may not be necessary to set up fp or bp:
100// * A base pointer is definitely needed when there are both VLAs and local
101// variables with more-than-default alignment requirements.
102// * A frame pointer is definitely needed when there are local variables with
103// more-than-default alignment requirements.
104//
105// For Darwin platforms the frame-record (fp, lr) is stored at the top of the
106// callee-saved area, since the unwind encoding does not allow for encoding
107// this dynamically and existing tools depend on this layout. For other
108// platforms, the frame-record is stored at the bottom of the (gpr) callee-saved
109// area to allow SVE stack objects (allocated directly below the callee-saves,
110// if available) to be accessed directly from the framepointer.
111// The SVE spill/fill instructions have VL-scaled addressing modes such
112// as:
113// ldr z8, [fp, #-7 mul vl]
114// For SVE the size of the vector length (VL) is not known at compile-time, so
115// '#-7 mul vl' is an offset that can only be evaluated at runtime. With this
116// layout, we don't need to add an unscaled offset to the framepointer before
117// accessing the SVE object in the frame.
118//
119// In some cases when a base pointer is not strictly needed, it is generated
120// anyway when offsets from the frame pointer to access local variables become
121// so large that the offset can't be encoded in the immediate fields of loads
122// or stores.
123//
124// Outgoing function arguments must be at the bottom of the stack frame when
125// calling another function. If we do not have variable-sized stack objects, we
126// can allocate a "reserved call frame" area at the bottom of the local
127// variable area, large enough for all outgoing calls. If we do have VLAs, then
128// the stack pointer must be decremented and incremented around each call to
129// make space for the arguments below the VLAs.
130//
131// FIXME: also explain the redzone concept.
132//
133// About stack hazards: Under some SME contexts, a coprocessor with its own
134// separate cache can used for FP operations. This can create hazards if the CPU
135// and the SME unit try to access the same area of memory, including if the
136// access is to an area of the stack. To try to alleviate this we attempt to
137// introduce extra padding into the stack frame between FP and GPR accesses,
138// controlled by the aarch64-stack-hazard-size option. Without changing the
139// layout of the stack frame in the diagram above, a stack object of size
140// aarch64-stack-hazard-size is added between GPR and FPR CSRs. Another is added
141// to the stack objects section, and stack objects are sorted so that FPR >
142// Hazard padding slot > GPRs (where possible). Unfortunately some things are
143// not handled well (VLA area, arguments on the stack, objects with both GPR and
144// FPR accesses), but if those are controlled by the user then the entire stack
145// frame becomes GPR at the start/end with FPR in the middle, surrounded by
146// Hazard padding.
147//
148// An example of the prologue:
149//
150// .globl __foo
151// .align 2
152// __foo:
153// Ltmp0:
154// .cfi_startproc
155// .cfi_personality 155, ___gxx_personality_v0
156// Leh_func_begin:
157// .cfi_lsda 16, Lexception33
158//
159// stp xa,bx, [sp, -#offset]!
160// ...
161// stp x28, x27, [sp, #offset-32]
162// stp fp, lr, [sp, #offset-16]
163// add fp, sp, #offset - 16
164// sub sp, sp, #1360
165//
166// The Stack:
167// +-------------------------------------------+
168// 10000 | ........ | ........ | ........ | ........ |
169// 10004 | ........ | ........ | ........ | ........ |
170// +-------------------------------------------+
171// 10008 | ........ | ........ | ........ | ........ |
172// 1000c | ........ | ........ | ........ | ........ |
173// +===========================================+
174// 10010 | X28 Register |
175// 10014 | X28 Register |
176// +-------------------------------------------+
177// 10018 | X27 Register |
178// 1001c | X27 Register |
179// +===========================================+
180// 10020 | Frame Pointer |
181// 10024 | Frame Pointer |
182// +-------------------------------------------+
183// 10028 | Link Register |
184// 1002c | Link Register |
185// +===========================================+
186// 10030 | ........ | ........ | ........ | ........ |
187// 10034 | ........ | ........ | ........ | ........ |
188// +-------------------------------------------+
189// 10038 | ........ | ........ | ........ | ........ |
190// 1003c | ........ | ........ | ........ | ........ |
191// +-------------------------------------------+
192//
193// [sp] = 10030 :: >>initial value<<
194// sp = 10020 :: stp fp, lr, [sp, #-16]!
195// fp = sp == 10020 :: mov fp, sp
196// [sp] == 10020 :: stp x28, x27, [sp, #-16]!
197// sp == 10010 :: >>final value<<
198//
199// The frame pointer (w29) points to address 10020. If we use an offset of
200// '16' from 'w29', we get the CFI offsets of -8 for w30, -16 for w29, -24
201// for w27, and -32 for w28:
202//
203// Ltmp1:
204// .cfi_def_cfa w29, 16
205// Ltmp2:
206// .cfi_offset w30, -8
207// Ltmp3:
208// .cfi_offset w29, -16
209// Ltmp4:
210// .cfi_offset w27, -24
211// Ltmp5:
212// .cfi_offset w28, -32
213//
214//===----------------------------------------------------------------------===//
215
216#include "AArch64FrameLowering.h"
217#include "AArch64InstrInfo.h"
220#include "AArch64RegisterInfo.h"
221#include "AArch64SMEAttributes.h"
222#include "AArch64Subtarget.h"
225#include "llvm/ADT/ScopeExit.h"
226#include "llvm/ADT/SmallVector.h"
244#include "llvm/IR/Attributes.h"
245#include "llvm/IR/CallingConv.h"
246#include "llvm/IR/DataLayout.h"
247#include "llvm/IR/DebugLoc.h"
248#include "llvm/IR/Function.h"
249#include "llvm/MC/MCAsmInfo.h"
250#include "llvm/MC/MCDwarf.h"
252#include "llvm/Support/Debug.h"
258#include <cassert>
259#include <cstdint>
260#include <iterator>
261#include <optional>
262#include <vector>
263
264using namespace llvm;
265
266#define DEBUG_TYPE "frame-info"
267
268static cl::opt<bool> EnableRedZone("aarch64-redzone",
269 cl::desc("enable use of redzone on AArch64"),
270 cl::init(false), cl::Hidden);
271
273 "stack-tagging-merge-settag",
274 cl::desc("merge settag instruction in function epilog"), cl::init(true),
275 cl::Hidden);
276
277static cl::opt<bool> OrderFrameObjects("aarch64-order-frame-objects",
278 cl::desc("sort stack allocations"),
279 cl::init(true), cl::Hidden);
280
281static cl::opt<bool>
282 SplitSVEObjects("aarch64-split-sve-objects",
283 cl::desc("Split allocation of ZPR & PPR objects"),
284 cl::init(true), cl::Hidden);
285
287 "homogeneous-prolog-epilog", cl::Hidden,
288 cl::desc("Emit homogeneous prologue and epilogue for the size "
289 "optimization (default = off)"));
290
291// Stack hazard size for analysis remarks. StackHazardSize takes precedence.
293 StackHazardRemarkSize("aarch64-stack-hazard-remark-size", cl::init(0),
294 cl::Hidden);
295// Whether to insert padding into non-streaming functions (for testing).
296static cl::opt<bool>
297 StackHazardInNonStreaming("aarch64-stack-hazard-in-non-streaming",
298 cl::init(false), cl::Hidden);
299
301 "aarch64-disable-multivector-spill-fill",
302 cl::desc("Disable use of LD/ST pairs for SME2 or SVE2p1"), cl::init(false),
303 cl::Hidden);
304
305int64_t
307 MachineBasicBlock &MBB) const {
308 MachineBasicBlock::iterator MBBI = MBB.getLastNonDebugInstr();
310 bool IsTailCallReturn = (MBB.end() != MBBI)
312 : false;
313
314 int64_t ArgumentPopSize = 0;
315 if (IsTailCallReturn) {
316 MachineOperand &StackAdjust = MBBI->getOperand(1);
317
318 // For a tail-call in a callee-pops-arguments environment, some or all of
319 // the stack may actually be in use for the call's arguments, this is
320 // calculated during LowerCall and consumed here...
321 ArgumentPopSize = StackAdjust.getImm();
322 } else {
323 // ... otherwise the amount to pop is *all* of the argument space,
324 // conveniently stored in the MachineFunctionInfo by
325 // LowerFormalArguments. This will, of course, be zero for the C calling
326 // convention.
327 ArgumentPopSize = AFI->getArgumentStackToRestore();
328 }
329
330 return ArgumentPopSize;
331}
332
334 MachineFunction &MF);
335
336enum class AssignObjectOffsets { No, Yes };
337/// Process all the SVE stack objects and the SVE stack size and offsets for
338/// each object. If AssignOffsets is "Yes", the offsets get assigned (and SVE
339/// stack sizes set). Returns the size of the SVE stack.
341 AssignObjectOffsets AssignOffsets);
342
343static unsigned getStackHazardSize(const MachineFunction &MF) {
344 return MF.getSubtarget<AArch64Subtarget>().getStreamingHazardSize();
345}
346
352
355 // With split SVE objects, the hazard padding is added to the PPR region,
356 // which places it between the [GPR, PPR] area and the [ZPR, FPR] area. This
357 // avoids hazards between both GPRs and FPRs and ZPRs and PPRs.
360 : 0,
361 AFI->getStackSizePPR());
362}
363
364// Conservatively, returns true if the function is likely to have SVE vectors
365// on the stack. This function is safe to be called before callee-saves or
366// object offsets have been determined.
368 const MachineFunction &MF) {
369 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
370 if (AFI->isSVECC())
371 return true;
372
373 if (AFI->hasCalculatedStackSizeSVE())
374 return bool(AFL.getSVEStackSize(MF));
375
376 const MachineFrameInfo &MFI = MF.getFrameInfo();
377 for (int FI = MFI.getObjectIndexBegin(); FI < MFI.getObjectIndexEnd(); FI++) {
378 if (MFI.hasScalableStackID(FI))
379 return true;
380 }
381
382 return false;
383}
384
385static bool isTargetWindows(const MachineFunction &MF) {
386 // TODO: Should this include targets like UEFI (which use Windows CFI)?
387 // Note: Currently, there is not AArch64 support for UEFI. The value returned
388 // here must align with the predicate used for returning the list of callee
389 // saved regs in AArch64RegisterInfo::getCalleeSavedRegs(), so that we use
390 // invalidateWindowsRegisterPairing() where appropriate.
392}
393
395 const MachineFunction &MF) const {
396 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
397 return isTargetWindows(MF) && AFI->getSVECalleeSavedStackSize();
398}
399
400/// Returns true if a homogeneous prolog or epilog code can be emitted
401/// for the size optimization. If possible, a frame helper call is injected.
402/// When Exit block is given, this check is for epilog.
403bool AArch64FrameLowering::homogeneousPrologEpilog(
404 MachineFunction &MF, MachineBasicBlock *Exit) const {
405 if (!MF.getFunction().hasMinSize())
406 return false;
408 return false;
409 if (EnableRedZone)
410 return false;
411
412 // TODO: Window is supported yet.
413 if (isTargetWindows(MF))
414 return false;
415
416 // TODO: SVE is not supported yet.
417 if (isLikelyToHaveSVEStack(*this, MF))
418 return false;
419
420 // Bail on stack adjustment needed on return for simplicity.
421 const MachineFrameInfo &MFI = MF.getFrameInfo();
422 const TargetRegisterInfo *RegInfo = MF.getSubtarget().getRegisterInfo();
423 if (MFI.hasVarSizedObjects() || RegInfo->hasStackRealignment(MF))
424 return false;
425 if (Exit && getArgumentStackToRestore(MF, *Exit))
426 return false;
427
428 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
429 if (AFI->hasSwiftAsyncContext() || AFI->hasStreamingModeChanges())
430 return false;
431
432 // If there are an odd number of GPRs before LR and FP in the CSRs list,
433 // they will not be paired into one RegPairInfo, which is incompatible with
434 // the assumption made by the homogeneous prolog epilog pass.
435 const MCPhysReg *CSRegs = MF.getRegInfo().getCalleeSavedRegs();
436 unsigned NumGPRs = 0;
437 for (unsigned I = 0; CSRegs[I]; ++I) {
438 Register Reg = CSRegs[I];
439 if (Reg == AArch64::LR) {
440 assert(CSRegs[I + 1] == AArch64::FP);
441 if (NumGPRs % 2 != 0)
442 return false;
443 break;
444 }
445 if (AArch64::GPR64RegClass.contains(Reg))
446 ++NumGPRs;
447 }
448
449 return true;
450}
451
452/// Returns true if CSRs should be paired.
453bool AArch64FrameLowering::producePairRegisters(MachineFunction &MF) const {
454 return produceCompactUnwindFrame(*this, MF) || homogeneousPrologEpilog(MF);
455}
456
457/// This is the biggest offset to the stack pointer we can encode in aarch64
458/// instructions (without using a separate calculation and a temp register).
459/// Note that the exception here are vector stores/loads which cannot encode any
460/// displacements (see estimateRSStackSizeLimit(), isAArch64FrameOffsetLegal()).
461static const unsigned DefaultSafeSPDisplacement = 255;
462
463/// Look at each instruction that references stack frames and return the stack
464/// size limit beyond which some of these instructions will require a scratch
465/// register during their expansion later.
467 // FIXME: For now, just conservatively guesstimate based on unscaled indexing
468 // range. We'll end up allocating an unnecessary spill slot a lot, but
469 // realistically that's not a big deal at this stage of the game.
470 for (MachineBasicBlock &MBB : MF) {
471 for (MachineInstr &MI : MBB) {
472 if (MI.isDebugInstr() || MI.isPseudo() ||
473 MI.getOpcode() == AArch64::ADDXri ||
474 MI.getOpcode() == AArch64::ADDSXri)
475 continue;
476
477 for (const MachineOperand &MO : MI.operands()) {
478 if (!MO.isFI())
479 continue;
480
482 if (isAArch64FrameOffsetLegal(MI, Offset, nullptr, nullptr, nullptr) ==
484 return 0;
485 }
486 }
487 }
489}
490
495
496unsigned
497AArch64FrameLowering::getFixedObjectSize(const MachineFunction &MF,
498 const AArch64FunctionInfo *AFI,
499 bool IsWin64, bool IsFunclet) const {
500 assert(AFI->getTailCallReservedStack() % 16 == 0 &&
501 "Tail call reserved stack must be aligned to 16 bytes");
502 if (!IsWin64 || IsFunclet) {
503 return AFI->getTailCallReservedStack();
504 } else {
505 if (AFI->getTailCallReservedStack() != 0 &&
506 !MF.getFunction().getAttributes().hasAttrSomewhere(
507 Attribute::SwiftAsync))
508 report_fatal_error("cannot generate ABI-changing tail call for Win64");
509 unsigned FixedObjectSize = AFI->getTailCallReservedStack();
510
511 // Var args are stored here in the primary function.
512 FixedObjectSize += AFI->getVarArgsGPRSize();
513
514 if (MF.hasEHFunclets()) {
515 // Catch objects are stored here in the primary function.
516 const MachineFrameInfo &MFI = MF.getFrameInfo();
517 const WinEHFuncInfo &EHInfo = *MF.getWinEHFuncInfo();
518 SmallSetVector<int, 8> CatchObjFrameIndices;
519 for (const WinEHTryBlockMapEntry &TBME : EHInfo.TryBlockMap) {
520 for (const WinEHHandlerType &H : TBME.HandlerArray) {
521 int FrameIndex = H.CatchObj.FrameIndex;
522 if ((FrameIndex != INT_MAX) &&
523 CatchObjFrameIndices.insert(FrameIndex)) {
524 FixedObjectSize = alignTo(FixedObjectSize,
525 MFI.getObjectAlign(FrameIndex).value()) +
526 MFI.getObjectSize(FrameIndex);
527 }
528 }
529 }
530 // To support EH funclets we allocate an UnwindHelp object
531 FixedObjectSize += 8;
532 }
533 return alignTo(FixedObjectSize, 16);
534 }
535}
536
538 if (!EnableRedZone)
539 return false;
540
541 // Don't use the red zone if the function explicitly asks us not to.
542 // This is typically used for kernel code.
543 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
544 const unsigned RedZoneSize =
546 if (!RedZoneSize)
547 return false;
548
549 const MachineFrameInfo &MFI = MF.getFrameInfo();
551 uint64_t NumBytes = AFI->getLocalStackSize();
552
553 // If neither NEON or SVE are available, a COPY from one Q-reg to
554 // another requires a spill -> reload sequence. We can do that
555 // using a pre-decrementing store/post-decrementing load, but
556 // if we do so, we can't use the Red Zone.
557 bool LowerQRegCopyThroughMem = Subtarget.hasFPARMv8() &&
558 !Subtarget.isNeonAvailable() &&
559 !Subtarget.hasSVE();
560
561 return !(MFI.hasCalls() || hasFP(MF) || NumBytes > RedZoneSize ||
562 AFI->hasSVEStackSize() || LowerQRegCopyThroughMem);
563}
564
565/// hasFPImpl - Return true if the specified function should have a dedicated
566/// frame pointer register.
568 const MachineFrameInfo &MFI = MF.getFrameInfo();
569 const TargetRegisterInfo *RegInfo = MF.getSubtarget().getRegisterInfo();
571
572 // Win64 EH requires a frame pointer if funclets are present, as the locals
573 // are accessed off the frame pointer in both the parent function and the
574 // funclets.
575 if (MF.hasEHFunclets())
576 return true;
577
578 // When the stack guard is mixed with the frame pointer, a dedicated FP is
579 // required so the guard value remains stable in the presence of dynamic
580 // stack allocations (e.g. _alloca on MSVCRT).
581 if (MFI.hasStackProtectorIndex()) {
582 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
583 if (Subtarget.getTargetLowering()->useStackGuardMixFP())
584 return true;
585 }
586
587 // Retain behavior of always omitting the FP for leaf functions when possible.
589 return true;
590 if (MFI.hasVarSizedObjects() || MFI.isFrameAddressTaken() ||
591 MFI.hasStackMap() || MFI.hasPatchPoint() ||
592 RegInfo->hasStackRealignment(MF))
593 return true;
594
595 // If we:
596 //
597 // 1. Have streaming mode changes
598 // OR:
599 // 2. Have a streaming body with SVE stack objects
600 //
601 // Then the value of VG restored when unwinding to this function may not match
602 // the value of VG used to set up the stack.
603 //
604 // This is a problem as the CFA can be described with an expression of the
605 // form: CFA = SP + NumBytes + VG * NumScalableBytes.
606 //
607 // If the value of VG used in that expression does not match the value used to
608 // set up the stack, an incorrect address for the CFA will be computed, and
609 // unwinding will fail.
610 //
611 // We work around this issue by ensuring the frame-pointer can describe the
612 // CFA in either of these cases.
613 if (AFI.needsDwarfUnwindInfo(MF) &&
616 return true;
617 // With large callframes around we may need to use FP to access the scavenging
618 // emergency spillslot.
619 //
620 // Unfortunately some calls to hasFP() like machine verifier ->
621 // getReservedReg() -> hasFP in the middle of global isel are too early
622 // to know the max call frame size. Hopefully conservatively returning "true"
623 // in those cases is fine.
624 // DefaultSafeSPDisplacement is fine as we only emergency spill GP regs.
625 if (!MFI.isMaxCallFrameSizeComputed() ||
627 return true;
628
629 return false;
630}
631
632/// Should the Frame Pointer be reserved for the current function?
634 const TargetMachine &TM = MF.getTarget();
635 const Triple &TT = TM.getTargetTriple();
636
637 // These OSes require the frame chain is valid, even if the current frame does
638 // not use a frame pointer.
639 if (TT.isOSDarwin() || TT.isOSWindows())
640 return true;
641
642 // If the function has a frame pointer, it is reserved.
643 if (hasFP(MF))
644 return true;
645
646 // Frontend has requested to preserve the frame pointer.
647 if (MF.framePointerIsReserved())
648 return true;
649
650 return false;
651}
652
653/// hasReservedCallFrame - Under normal circumstances, when a frame pointer is
654/// not required, we reserve argument space for call sites in the function
655/// immediately on entry to the current function. This eliminates the need for
656/// add/sub sp brackets around call sites. Returns true if the call frame is
657/// included as part of the stack frame.
659 const MachineFunction &MF) const {
660 // The stack probing code for the dynamically allocated outgoing arguments
661 // area assumes that the stack is probed at the top - either by the prologue
662 // code, which issues a probe if `hasVarSizedObjects` return true, or by the
663 // most recent variable-sized object allocation. Changing the condition here
664 // may need to be followed up by changes to the probe issuing logic.
665 return !MF.getFrameInfo().hasVarSizedObjects();
666}
667
671
672 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
673 const AArch64InstrInfo *TII = Subtarget.getInstrInfo();
674 const AArch64TargetLowering *TLI = Subtarget.getTargetLowering();
675 [[maybe_unused]] MachineFrameInfo &MFI = MF.getFrameInfo();
676 DebugLoc DL = I->getDebugLoc();
677 unsigned Opc = I->getOpcode();
678 bool IsDestroy = Opc == TII->getCallFrameDestroyOpcode();
679 uint64_t CalleePopAmount = IsDestroy ? I->getOperand(1).getImm() : 0;
680
681 if (!hasReservedCallFrame(MF)) {
682 int64_t Amount = I->getOperand(0).getImm();
683 Amount = alignTo(Amount, getStackAlign());
684 if (!IsDestroy)
685 Amount = -Amount;
686
687 // N.b. if CalleePopAmount is valid but zero (i.e. callee would pop, but it
688 // doesn't have to pop anything), then the first operand will be zero too so
689 // this adjustment is a no-op.
690 if (CalleePopAmount == 0) {
691 // FIXME: in-function stack adjustment for calls is limited to 24-bits
692 // because there's no guaranteed temporary register available.
693 //
694 // ADD/SUB (immediate) has only LSL #0 and LSL #12 available.
695 // 1) For offset <= 12-bit, we use LSL #0
696 // 2) For 12-bit <= offset <= 24-bit, we use two instructions. One uses
697 // LSL #0, and the other uses LSL #12.
698 //
699 // Most call frames will be allocated at the start of a function so
700 // this is OK, but it is a limitation that needs dealing with.
701 assert(Amount > -0xffffff && Amount < 0xffffff && "call frame too large");
702
703 if (TLI->hasInlineStackProbe(MF) &&
705 // When stack probing is enabled, the decrement of SP may need to be
706 // probed. We only need to do this if the call site needs 1024 bytes of
707 // space or more, because a region smaller than that is allowed to be
708 // unprobed at an ABI boundary. We rely on the fact that SP has been
709 // probed exactly at this point, either by the prologue or most recent
710 // dynamic allocation.
712 "non-reserved call frame without var sized objects?");
713 Register ScratchReg =
714 MF.getRegInfo().createVirtualRegister(&AArch64::GPR64RegClass);
715 inlineStackProbeFixed(I, ScratchReg, -Amount, StackOffset::get(0, 0));
716 } else {
717 emitFrameOffset(MBB, I, DL, AArch64::SP, AArch64::SP,
718 StackOffset::getFixed(Amount), TII);
719 }
720 }
721 } else if (CalleePopAmount != 0) {
722 // If the calling convention demands that the callee pops arguments from the
723 // stack, we want to add it back if we have a reserved call frame.
724 assert(CalleePopAmount < 0xffffff && "call frame too large");
725 emitFrameOffset(MBB, I, DL, AArch64::SP, AArch64::SP,
726 StackOffset::getFixed(-(int64_t)CalleePopAmount), TII);
727 }
728 return MBB.erase(I);
729}
730
732 MachineBasicBlock &MBB) const {
733
734 MachineFunction &MF = *MBB.getParent();
735 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
736 const auto &TRI = *Subtarget.getRegisterInfo();
737 const auto &MFI = *MF.getInfo<AArch64FunctionInfo>();
738
739 CFIInstBuilder CFIBuilder(MBB, MBB.begin(), MachineInstr::NoFlags);
740
741 // Reset the CFA to `SP + 0`.
742 CFIBuilder.buildDefCFA(AArch64::SP, 0);
743
744 // Flip the RA sign state.
745 if (MFI.shouldSignReturnAddress(MF)) {
746 if (MFI.branchProtectionPAuthLR()) {
747 CFIBuilder.buildNegateRAStateWithPC();
748 } else if (!MF.getTarget().getTargetTriple().isOSBinFormatMachO()) {
749 CFIBuilder.buildNegateRAState();
750 }
751 }
752
753 // Shadow call stack uses X18, reset it.
754 if (MFI.needsShadowCallStackPrologueEpilogue(MF))
755 CFIBuilder.buildSameValue(AArch64::X18);
756
757 // Emit .cfi_same_value for callee-saved registers.
758 const std::vector<CalleeSavedInfo> &CSI =
760 for (const auto &Info : CSI) {
761 MCRegister Reg = Info.getReg();
762 if (!TRI.regNeedsCFI(Reg, Reg))
763 continue;
764 CFIBuilder.buildSameValue(Reg);
765 }
766}
767
769 switch (Reg.id()) {
770 default:
771 // The called routine is expected to preserve r19-r28
772 // r29 and r30 are used as frame pointer and link register resp.
773 return 0;
774
775 // GPRs
776#define CASE(n) \
777 case AArch64::W##n: \
778 case AArch64::X##n: \
779 return AArch64::X##n
780 CASE(0);
781 CASE(1);
782 CASE(2);
783 CASE(3);
784 CASE(4);
785 CASE(5);
786 CASE(6);
787 CASE(7);
788 CASE(8);
789 CASE(9);
790 CASE(10);
791 CASE(11);
792 CASE(12);
793 CASE(13);
794 CASE(14);
795 CASE(15);
796 CASE(16);
797 CASE(17);
798 CASE(18);
799#undef CASE
800
801 // FPRs
802#define CASE(n) \
803 case AArch64::B##n: \
804 case AArch64::H##n: \
805 case AArch64::S##n: \
806 case AArch64::D##n: \
807 case AArch64::Q##n: \
808 return HasSVE ? AArch64::Z##n : AArch64::Q##n
809 CASE(0);
810 CASE(1);
811 CASE(2);
812 CASE(3);
813 CASE(4);
814 CASE(5);
815 CASE(6);
816 CASE(7);
817 CASE(8);
818 CASE(9);
819 CASE(10);
820 CASE(11);
821 CASE(12);
822 CASE(13);
823 CASE(14);
824 CASE(15);
825 CASE(16);
826 CASE(17);
827 CASE(18);
828 CASE(19);
829 CASE(20);
830 CASE(21);
831 CASE(22);
832 CASE(23);
833 CASE(24);
834 CASE(25);
835 CASE(26);
836 CASE(27);
837 CASE(28);
838 CASE(29);
839 CASE(30);
840 CASE(31);
841#undef CASE
842 }
843}
844
845void AArch64FrameLowering::emitZeroCallUsedRegs(BitVector RegsToZero,
847 RegScavenger *) const {
848 // Insertion point.
850
851 // Fake a debug loc.
852 DebugLoc DL;
853 if (MBBI != MBB.end())
854 DL = MBBI->getDebugLoc();
855
856 const MachineFunction &MF = *MBB.getParent();
857 const AArch64Subtarget &STI = MF.getSubtarget<AArch64Subtarget>();
858 const AArch64RegisterInfo &TRI = *STI.getRegisterInfo();
859
860 BitVector GPRsToZero(TRI.getNumRegs());
861 BitVector FPRsToZero(TRI.getNumRegs());
862 bool HasSVE = STI.isSVEorStreamingSVEAvailable();
863 // Without an FP unit (e.g. -mgeneral-regs-only) the FP/vector registers can't
864 // hold a value and there is no instruction to clear them, so leave them out.
865 bool HasFPR = STI.hasFPARMv8();
866 for (MCRegister Reg : RegsToZero.set_bits()) {
867 if (TRI.isGeneralPurposeRegister(MF, Reg)) {
868 // For GPRs, we only care to clear out the 64-bit register.
869 if (MCRegister XReg = getRegisterOrZero(Reg, HasSVE))
870 GPRsToZero.set(XReg);
871 } else if (HasFPR && AArch64InstrInfo::isFpOrNEON(Reg)) {
872 // For FPRs,
873 if (MCRegister XReg = getRegisterOrZero(Reg, HasSVE))
874 FPRsToZero.set(XReg);
875 }
876 }
877
878 const AArch64InstrInfo &TII = *STI.getInstrInfo();
879
880 // Zero out GPRs.
881 for (MCRegister Reg : GPRsToZero.set_bits())
882 TII.buildClearRegister(Reg, MBB, MBBI, DL);
883
884 // Zero out FP/vector registers.
885 for (MCRegister Reg : FPRsToZero.set_bits())
886 TII.buildClearRegister(Reg, MBB, MBBI, DL);
887
888 if (HasSVE) {
889 for (MCRegister PReg :
890 {AArch64::P0, AArch64::P1, AArch64::P2, AArch64::P3, AArch64::P4,
891 AArch64::P5, AArch64::P6, AArch64::P7, AArch64::P8, AArch64::P9,
892 AArch64::P10, AArch64::P11, AArch64::P12, AArch64::P13, AArch64::P14,
893 AArch64::P15}) {
894 if (RegsToZero[PReg])
895 BuildMI(MBB, MBBI, DL, TII.get(AArch64::PFALSE), PReg);
896 }
897 }
898}
899
900bool AArch64FrameLowering::windowsRequiresStackProbe(
901 const MachineFunction &MF, uint64_t StackSizeInBytes) const {
902 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
903 const AArch64FunctionInfo &MFI = *MF.getInfo<AArch64FunctionInfo>();
904 // TODO: When implementing stack protectors, take that into account
905 // for the probe threshold.
906 return Subtarget.isTargetWindows() && MFI.hasStackProbing() &&
907 StackSizeInBytes >= uint64_t(MFI.getStackProbeSize());
908}
909
911 const MachineBasicBlock &MBB) {
912 const MachineFunction *MF = MBB.getParent();
913 LiveRegs.addLiveIns(MBB);
914 // Mark callee saved registers as used so we will not choose them.
915 const MCPhysReg *CSRegs = MF->getRegInfo().getCalleeSavedRegs();
916 for (unsigned i = 0; CSRegs[i]; ++i)
917 LiveRegs.addReg(CSRegs[i]);
918}
919
921AArch64FrameLowering::findScratchNonCalleeSaveRegister(MachineBasicBlock *MBB,
922 bool HasCall) const {
924
925 // If MBB is an entry block, use X9 as the scratch register
926 // preserve_none functions may be using X9 to pass arguments,
927 // so prefer to pick an available register below.
928 if (&MF->front() == MBB &&
930 return AArch64::X9;
931
932 const AArch64Subtarget &Subtarget = MF->getSubtarget<AArch64Subtarget>();
933 const AArch64RegisterInfo &TRI = *Subtarget.getRegisterInfo();
934 LivePhysRegs LiveRegs(TRI);
935 getLiveRegsForEntryMBB(LiveRegs, *MBB);
936 if (HasCall) {
937 LiveRegs.addReg(AArch64::X16);
938 LiveRegs.addReg(AArch64::X17);
939 LiveRegs.addReg(AArch64::X18);
940 }
941
942 // Prefer X9 since it was historically used for the prologue scratch reg.
943 const MachineRegisterInfo &MRI = MF->getRegInfo();
944 if (LiveRegs.available(MRI, AArch64::X9))
945 return AArch64::X9;
946
947 for (unsigned Reg : AArch64::GPR64RegClass) {
948 if (LiveRegs.available(MRI, Reg))
949 return Reg;
950 }
951 return AArch64::NoRegister;
952}
953
955 const MachineBasicBlock &MBB) const {
956 const MachineFunction *MF = MBB.getParent();
957 MachineBasicBlock *TmpMBB = const_cast<MachineBasicBlock *>(&MBB);
958 const AArch64Subtarget &Subtarget = MF->getSubtarget<AArch64Subtarget>();
959 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
960 const AArch64TargetLowering *TLI = Subtarget.getTargetLowering();
962
963 if (AFI->hasSwiftAsyncContext()) {
964 const AArch64RegisterInfo &TRI = *Subtarget.getRegisterInfo();
965 const MachineRegisterInfo &MRI = MF->getRegInfo();
968 // The StoreSwiftAsyncContext clobbers X16 and X17. Make sure they are
969 // available.
970 if (!LiveRegs.available(MRI, AArch64::X16) ||
971 !LiveRegs.available(MRI, AArch64::X17))
972 return false;
973 }
974
975 // Certain stack probing sequences might clobber flags, then we can't use
976 // the block as a prologue if the flags register is a live-in.
978 MBB.isLiveIn(AArch64::NZCV))
979 return false;
980
981 if (RegInfo->hasStackRealignment(*MF) || TLI->hasInlineStackProbe(*MF))
982 if (findScratchNonCalleeSaveRegister(TmpMBB) == AArch64::NoRegister)
983 return false;
984
985 // May need a scratch register (for return value) if require making a special
986 // call
987 if (requiresSaveVG(*MF) ||
988 windowsRequiresStackProbe(*MF, std::numeric_limits<uint64_t>::max()))
989 if (findScratchNonCalleeSaveRegister(TmpMBB, true) == AArch64::NoRegister)
990 return false;
991
992 return true;
993}
994
996 const Function &F = MF.getFunction();
997 return MF.getTarget().getMCAsmInfo().usesWindowsCFI() &&
998 F.needsUnwindTableEntry();
999}
1000
1001bool AArch64FrameLowering::shouldSignReturnAddressEverywhere(
1002 const MachineFunction &MF) const {
1003 // FIXME: With WinCFI, extra care should be taken to place SEH_PACSignLR
1004 // and SEH_EpilogEnd instructions in the correct order.
1006 return false;
1009}
1010
1011// Given a load or a store instruction, generate an appropriate unwinding SEH
1012// code on Windows.
1014AArch64FrameLowering::insertSEH(MachineBasicBlock::iterator MBBI,
1015 const AArch64InstrInfo &TII,
1016 MachineInstr::MIFlag Flag) const {
1017 unsigned Opc = MBBI->getOpcode();
1018 MachineBasicBlock *MBB = MBBI->getParent();
1019 MachineFunction &MF = *MBB->getParent();
1020 DebugLoc DL = MBBI->getDebugLoc();
1021 unsigned ImmIdx = MBBI->getNumOperands() - 1;
1022 int Imm = MBBI->getOperand(ImmIdx).getImm();
1024 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1025 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
1026
1027 switch (Opc) {
1028 default:
1029 report_fatal_error("No SEH Opcode for this instruction");
1030 case AArch64::STR_ZXI:
1031 case AArch64::LDR_ZXI: {
1032 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1033 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveZReg))
1034 .addImm(Reg0)
1035 .addImm(Imm)
1036 .setMIFlag(Flag);
1037 break;
1038 }
1039 case AArch64::STR_PXI:
1040 case AArch64::LDR_PXI: {
1041 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1042 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SavePReg))
1043 .addImm(Reg0)
1044 .addImm(Imm)
1045 .setMIFlag(Flag);
1046 break;
1047 }
1048 case AArch64::LDPDpost:
1049 Imm = -Imm;
1050 [[fallthrough]];
1051 case AArch64::STPDpre: {
1052 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1053 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(2).getReg());
1054 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFRegP_X))
1055 .addImm(Reg0)
1056 .addImm(Reg1)
1057 .addImm(Imm * 8)
1058 .setMIFlag(Flag);
1059 break;
1060 }
1061 case AArch64::LDPXpost:
1062 Imm = -Imm;
1063 [[fallthrough]];
1064 case AArch64::STPXpre: {
1065 Register Reg0 = MBBI->getOperand(1).getReg();
1066 Register Reg1 = MBBI->getOperand(2).getReg();
1067 if (Reg0 == AArch64::FP && Reg1 == AArch64::LR)
1068 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFPLR_X))
1069 .addImm(Imm * 8)
1070 .setMIFlag(Flag);
1071 else
1072 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveRegP_X))
1073 .addImm(RegInfo->getSEHRegNum(Reg0))
1074 .addImm(RegInfo->getSEHRegNum(Reg1))
1075 .addImm(Imm * 8)
1076 .setMIFlag(Flag);
1077 break;
1078 }
1079 case AArch64::LDRDpost:
1080 Imm = -Imm;
1081 [[fallthrough]];
1082 case AArch64::STRDpre: {
1083 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1084 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFReg_X))
1085 .addImm(Reg)
1086 .addImm(Imm)
1087 .setMIFlag(Flag);
1088 break;
1089 }
1090 case AArch64::LDRXpost:
1091 Imm = -Imm;
1092 [[fallthrough]];
1093 case AArch64::STRXpre: {
1094 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1095 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveReg_X))
1096 .addImm(Reg)
1097 .addImm(Imm)
1098 .setMIFlag(Flag);
1099 break;
1100 }
1101 case AArch64::STPDi:
1102 case AArch64::LDPDi: {
1103 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1104 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1105 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFRegP))
1106 .addImm(Reg0)
1107 .addImm(Reg1)
1108 .addImm(Imm * 8)
1109 .setMIFlag(Flag);
1110 break;
1111 }
1112 case AArch64::STPXi:
1113 case AArch64::LDPXi: {
1114 Register Reg0 = MBBI->getOperand(0).getReg();
1115 Register Reg1 = MBBI->getOperand(1).getReg();
1116
1117 int SEHReg0 = RegInfo->getSEHRegNum(Reg0);
1118 int SEHReg1 = RegInfo->getSEHRegNum(Reg1);
1119
1120 if (Reg0 == AArch64::FP && Reg1 == AArch64::LR)
1121 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFPLR))
1122 .addImm(Imm * 8)
1123 .setMIFlag(Flag);
1124 else if (SEHReg0 >= 19 && SEHReg1 >= 19)
1125 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveRegP))
1126 .addImm(SEHReg0)
1127 .addImm(SEHReg1)
1128 .addImm(Imm * 8)
1129 .setMIFlag(Flag);
1130 else
1131 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegIP))
1132 .addImm(SEHReg0)
1133 .addImm(SEHReg1)
1134 .addImm(Imm * 8)
1135 .setMIFlag(Flag);
1136 break;
1137 }
1138 case AArch64::STRXui:
1139 case AArch64::LDRXui: {
1140 int Reg = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1141 if (Reg >= 19)
1142 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveReg))
1143 .addImm(Reg)
1144 .addImm(Imm * 8)
1145 .setMIFlag(Flag);
1146 else
1147 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegI))
1148 .addImm(Reg)
1149 .addImm(Imm * 8)
1150 .setMIFlag(Flag);
1151 break;
1152 }
1153 case AArch64::STRDui:
1154 case AArch64::LDRDui: {
1155 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1156 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFReg))
1157 .addImm(Reg)
1158 .addImm(Imm * 8)
1159 .setMIFlag(Flag);
1160 break;
1161 }
1162 case AArch64::STPQi:
1163 case AArch64::LDPQi: {
1164 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1165 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1166 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegQP))
1167 .addImm(Reg0)
1168 .addImm(Reg1)
1169 .addImm(Imm * 16)
1170 .setMIFlag(Flag);
1171 break;
1172 }
1173 case AArch64::LDPQpost:
1174 Imm = -Imm;
1175 [[fallthrough]];
1176 case AArch64::STPQpre: {
1177 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1178 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(2).getReg());
1179 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegQPX))
1180 .addImm(Reg0)
1181 .addImm(Reg1)
1182 .addImm(Imm * 16)
1183 .setMIFlag(Flag);
1184 break;
1185 }
1186 }
1187 auto I = MBB->insertAfter(MBBI, MIB);
1188 return I;
1189}
1190
1193 if (!AFI->needsDwarfUnwindInfo(MF) || !AFI->hasStreamingModeChanges())
1194 return false;
1195 // For Darwin platforms we don't save VG for non-SVE functions, even if SME
1196 // is enabled with streaming mode changes.
1197 auto &ST = MF.getSubtarget<AArch64Subtarget>();
1198 if (ST.isTargetDarwin())
1199 return ST.hasSVE();
1200 return true;
1201}
1202
1204 MachineFunction &MF) const {
1205 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1206 const AArch64InstrInfo *TII = Subtarget.getInstrInfo();
1207
1208 auto EmitSignRA = [&](MachineBasicBlock &MBB) {
1209 DebugLoc DL; // Set debug location to unknown.
1211
1212 BuildMI(MBB, MBBI, DL, TII->get(AArch64::PAUTH_PROLOGUE))
1214 };
1215
1216 auto EmitAuthRA = [&](MachineBasicBlock &MBB) {
1217 DebugLoc DL;
1218 MachineBasicBlock::iterator MBBI = MBB.getFirstTerminator();
1219 if (MBBI != MBB.end())
1220 DL = MBBI->getDebugLoc();
1221
1222 TII->createPauthEpilogueInstr(MBB, DL);
1223 };
1224
1225 // This should be in sync with PEIImpl::calculateSaveRestoreBlocks.
1226 EmitSignRA(MF.front());
1227 for (MachineBasicBlock &MBB : MF) {
1228 if (MBB.isEHFuncletEntry())
1229 EmitSignRA(MBB);
1230 if (MBB.isReturnBlock())
1231 EmitAuthRA(MBB);
1232 }
1233}
1234
1236 MachineBasicBlock &MBB) const {
1237 AArch64PrologueEmitter PrologueEmitter(MF, MBB, *this);
1238 PrologueEmitter.emitPrologue();
1239}
1240
1242 MachineBasicBlock &MBB) const {
1243 AArch64EpilogueEmitter EpilogueEmitter(MF, MBB, *this);
1244 EpilogueEmitter.emitEpilogue();
1245}
1246
1249 MF.getInfo<AArch64FunctionInfo>()->needsDwarfUnwindInfo(MF);
1250}
1251
1253 return enableCFIFixup(MF) &&
1254 MF.getInfo<AArch64FunctionInfo>()->needsAsyncDwarfUnwindInfo(MF);
1255}
1256
1257/// getFrameIndexReference - Provide a base+offset reference to an FI slot for
1258/// debug info. It's the same as what we use for resolving the code-gen
1259/// references for now. FIXME: This can go wrong when references are
1260/// SP-relative and simple call frames aren't used.
1263 Register &FrameReg) const {
1265 MF, FI, FrameReg,
1266 /*PreferFP=*/
1267 MF.getFunction().hasFnAttribute(Attribute::SanitizeHWAddress) ||
1268 MF.getFunction().hasFnAttribute(Attribute::SanitizeMemTag),
1269 /*ForSimm=*/false);
1270}
1271
1274 int FI) const {
1275 // This function serves to provide a comparable offset from a single reference
1276 // point (the value of SP at function entry) that can be used for analysis,
1277 // e.g. the stack-frame-layout analysis pass. It is not guaranteed to be
1278 // correct for all objects in the presence of VLA-area objects or dynamic
1279 // stack re-alignment.
1280
1281 const auto &MFI = MF.getFrameInfo();
1282
1283 int64_t ObjectOffset = MFI.getObjectOffset(FI);
1284 StackOffset ZPRStackSize = getZPRStackSize(MF);
1285 StackOffset PPRStackSize = getPPRStackSize(MF);
1286 StackOffset SVEStackSize = ZPRStackSize + PPRStackSize;
1287
1288 // For VLA-area objects, just emit an offset at the end of the stack frame.
1289 // Whilst not quite correct, these objects do live at the end of the frame and
1290 // so it is more useful for analysis for the offset to reflect this.
1291 if (MFI.isVariableSizedObjectIndex(FI)) {
1292 return StackOffset::getFixed(-((int64_t)MFI.getStackSize())) - SVEStackSize;
1293 }
1294
1295 // This is correct in the absence of any SVE stack objects.
1296 if (!SVEStackSize)
1297 return StackOffset::getFixed(ObjectOffset - getOffsetOfLocalArea());
1298
1299 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1300 bool FPAfterSVECalleeSaves = hasSVECalleeSavesAboveFrameRecord(MF);
1301 if (MFI.hasScalableStackID(FI)) {
1302 if (FPAfterSVECalleeSaves &&
1303 -ObjectOffset <= (int64_t)AFI->getSVECalleeSavedStackSize()) {
1304 assert(!AFI->hasSplitSVEObjects() &&
1305 "split-sve-objects not supported with FPAfterSVECalleeSaves");
1306 return StackOffset::getScalable(ObjectOffset);
1307 }
1308 StackOffset AccessOffset{};
1309 // The scalable vectors are below (lower address) the scalable predicates
1310 // with split SVE objects, so we must subtract the size of the predicates.
1311 if (AFI->hasSplitSVEObjects() &&
1312 MFI.getStackID(FI) == TargetStackID::ScalableVector)
1313 AccessOffset = -PPRStackSize;
1314 return AccessOffset +
1315 StackOffset::get(-((int64_t)AFI->getCalleeSavedStackSize()),
1316 ObjectOffset);
1317 }
1318
1319 bool IsFixed = MFI.isFixedObjectIndex(FI);
1320 bool IsCSR =
1321 !IsFixed && ObjectOffset >= -((int)AFI->getCalleeSavedStackSize(MFI));
1322
1323 StackOffset ScalableOffset = {};
1324 if (!IsFixed && !IsCSR) {
1325 ScalableOffset = -SVEStackSize;
1326 } else if (FPAfterSVECalleeSaves && IsCSR) {
1327 ScalableOffset =
1329 }
1330
1331 return StackOffset::getFixed(ObjectOffset) + ScalableOffset;
1332}
1333
1339
1340StackOffset AArch64FrameLowering::getFPOffset(const MachineFunction &MF,
1341 int64_t ObjectOffset) const {
1342 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1343 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1344 const Function &F = MF.getFunction();
1345 bool IsWin64 = Subtarget.isCallingConvWin64(F.getCallingConv(), F.isVarArg());
1346 unsigned FixedObject =
1347 getFixedObjectSize(MF, AFI, IsWin64, /*IsFunclet=*/false);
1348 int64_t CalleeSaveSize = AFI->getCalleeSavedStackSize(MF.getFrameInfo());
1349 int64_t FPAdjust =
1350 CalleeSaveSize - AFI->getCalleeSaveBaseToFrameRecordOffset();
1351 return StackOffset::getFixed(ObjectOffset + FixedObject + FPAdjust);
1352}
1353
1354StackOffset AArch64FrameLowering::getStackOffset(const MachineFunction &MF,
1355 int64_t ObjectOffset) const {
1356 const auto &MFI = MF.getFrameInfo();
1357 return StackOffset::getFixed(ObjectOffset + (int64_t)MFI.getStackSize());
1358}
1359
1360// TODO: This function currently does not work for scalable vectors.
1362 int FI) const {
1363 const AArch64RegisterInfo *RegInfo =
1364 MF.getSubtarget<AArch64Subtarget>().getRegisterInfo();
1365 int ObjectOffset = MF.getFrameInfo().getObjectOffset(FI);
1366 return RegInfo->getLocalAddressRegister(MF) == AArch64::FP
1367 ? getFPOffset(MF, ObjectOffset).getFixed()
1368 : getStackOffset(MF, ObjectOffset).getFixed();
1369}
1370
1372 const MachineFunction &MF, int FI, Register &FrameReg, bool PreferFP,
1373 bool ForSimm) const {
1374 const auto &MFI = MF.getFrameInfo();
1375 int64_t ObjectOffset = MFI.getObjectOffset(FI);
1376 bool isFixed = MFI.isFixedObjectIndex(FI);
1377 auto StackID = static_cast<TargetStackID::Value>(MFI.getStackID(FI));
1378 return resolveFrameOffsetReference(MF, ObjectOffset, isFixed, StackID,
1379 FrameReg, PreferFP, ForSimm);
1380}
1381
1383 const MachineFunction &MF, int64_t ObjectOffset, bool isFixed,
1384 TargetStackID::Value StackID, Register &FrameReg, bool PreferFP,
1385 bool ForSimm) const {
1386 const auto &MFI = MF.getFrameInfo();
1387 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1388 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
1389 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1390
1391 int64_t FPOffset = getFPOffset(MF, ObjectOffset).getFixed();
1392 int64_t Offset = getStackOffset(MF, ObjectOffset).getFixed();
1393 bool isCSR =
1394 !isFixed && ObjectOffset >= -((int)AFI->getCalleeSavedStackSize(MFI));
1395 bool isSVE = MFI.isScalableStackID(StackID);
1396
1397 StackOffset ZPRStackSize = getZPRStackSize(MF);
1398 StackOffset PPRStackSize = getPPRStackSize(MF);
1399 StackOffset SVEStackSize = ZPRStackSize + PPRStackSize;
1400
1401 // Use frame pointer to reference fixed objects. Use it for locals if
1402 // there are VLAs or a dynamically realigned SP (and thus the SP isn't
1403 // reliable as a base). Make sure useFPForScavengingIndex() does the
1404 // right thing for the emergency spill slot.
1405 bool UseFP = false;
1406 if (AFI->hasStackFrame() && !isSVE) {
1407 // We shouldn't prefer using the FP to access fixed-sized stack objects when
1408 // there are scalable (SVE) objects in between the FP and the fixed-sized
1409 // objects.
1410 PreferFP &= !SVEStackSize;
1411
1412 // Note: Keeping the following as multiple 'if' statements rather than
1413 // merging to a single expression for readability.
1414 //
1415 // Argument access should always use the FP.
1416 if (isFixed) {
1417 UseFP = hasFP(MF);
1418 } else if (isCSR && RegInfo->hasStackRealignment(MF)) {
1419 // References to the CSR area must use FP if we're re-aligning the stack
1420 // since the dynamically-sized alignment padding is between the SP/BP and
1421 // the CSR area.
1422 assert(hasFP(MF) && "Re-aligned stack must have frame pointer");
1423 UseFP = true;
1424 } else if (hasFP(MF) && !RegInfo->hasStackRealignment(MF)) {
1425 // If the FPOffset is negative and we're producing a signed immediate, we
1426 // have to keep in mind that the available offset range for negative
1427 // offsets is smaller than for positive ones. If an offset is available
1428 // via the FP and the SP, use whichever is closest.
1429 bool FPOffsetFits = !ForSimm || FPOffset >= -256;
1430 PreferFP |= Offset > -FPOffset && !SVEStackSize;
1431
1432 if (FPOffset >= 0) {
1433 // If the FPOffset is positive, that'll always be best, as the SP/BP
1434 // will be even further away.
1435 UseFP = true;
1436 } else if (MFI.hasVarSizedObjects()) {
1437 // If we have variable sized objects, we can use either FP or BP, as the
1438 // SP offset is unknown. We can use the base pointer if we have one and
1439 // FP is not preferred. If not, we're stuck with using FP.
1440 bool CanUseBP = RegInfo->hasBasePointer(MF);
1441 if (FPOffsetFits && CanUseBP) // Both are ok. Pick the best.
1442 UseFP = PreferFP;
1443 else if (!CanUseBP) // Can't use BP. Forced to use FP.
1444 UseFP = true;
1445 // else we can use BP and FP, but the offset from FP won't fit.
1446 // That will make us scavenge registers which we can probably avoid by
1447 // using BP. If it won't fit for BP either, we'll scavenge anyway.
1448 } else if (MF.hasEHFunclets() && !RegInfo->hasBasePointer(MF)) {
1449 // Funclets access the locals contained in the parent's stack frame
1450 // via the frame pointer, so we have to use the FP in the parent
1451 // function.
1452 (void) Subtarget;
1453 assert(Subtarget.isCallingConvWin64(MF.getFunction().getCallingConv(),
1454 MF.getFunction().isVarArg()) &&
1455 "Funclets should only be present on Win64");
1456 UseFP = true;
1457 } else {
1458 // We have the choice between FP and (SP or BP).
1459 if (FPOffsetFits && PreferFP) // If FP is the best fit, use it.
1460 UseFP = true;
1461 }
1462 }
1463 }
1464
1465 assert(
1466 ((isFixed || isCSR) || !RegInfo->hasStackRealignment(MF) || !UseFP) &&
1467 "In the presence of dynamic stack pointer realignment, "
1468 "non-argument/CSR objects cannot be accessed through the frame pointer");
1469
1470 bool FPAfterSVECalleeSaves = hasSVECalleeSavesAboveFrameRecord(MF);
1471
1472 if (isSVE) {
1473 StackOffset FPOffset = StackOffset::get(
1474 -AFI->getCalleeSaveBaseToFrameRecordOffset(), ObjectOffset);
1475 StackOffset SPOffset =
1476 SVEStackSize +
1477 StackOffset::get(MFI.getStackSize() - AFI->getCalleeSavedStackSize(),
1478 ObjectOffset);
1479
1480 // With split SVE objects the ObjectOffset is relative to the split area
1481 // (i.e. the PPR area or ZPR area respectively).
1482 if (AFI->hasSplitSVEObjects() && StackID == TargetStackID::ScalableVector) {
1483 // If we're accessing an SVE vector with split SVE objects...
1484 // - From the FP we need to move down past the PPR area:
1485 FPOffset -= PPRStackSize;
1486 // - From the SP we only need to move up to the ZPR area:
1487 SPOffset -= PPRStackSize;
1488 // Note: `SPOffset = SVEStackSize + ...`, so `-= PPRStackSize` results in
1489 // `SPOffset = ZPRStackSize + ...`.
1490 }
1491
1492 if (FPAfterSVECalleeSaves) {
1494 if (-ObjectOffset <= (int64_t)AFI->getSVECalleeSavedStackSize()) {
1497 }
1498 }
1499
1500 // Always use the FP for SVE spills if available and beneficial.
1501 if (hasFP(MF) && (SPOffset.getFixed() ||
1502 FPOffset.getScalable() < SPOffset.getScalable() ||
1503 RegInfo->hasStackRealignment(MF))) {
1504 FrameReg = RegInfo->getFrameRegister(MF);
1505 return FPOffset;
1506 }
1507 FrameReg = RegInfo->hasBasePointer(MF) ? RegInfo->getBaseRegister()
1508 : MCRegister(AArch64::SP);
1509
1510 return SPOffset;
1511 }
1512
1513 StackOffset SVEAreaOffset = {};
1514 if (FPAfterSVECalleeSaves) {
1515 // In this stack layout, the FP is in between the callee saves and other
1516 // SVE allocations.
1517 StackOffset SVECalleeSavedStack =
1519 if (UseFP) {
1520 if (isFixed)
1521 SVEAreaOffset = SVECalleeSavedStack;
1522 else if (!isCSR)
1523 SVEAreaOffset = SVECalleeSavedStack - SVEStackSize;
1524 } else {
1525 if (isFixed)
1526 SVEAreaOffset = SVEStackSize;
1527 else if (isCSR)
1528 SVEAreaOffset = SVEStackSize - SVECalleeSavedStack;
1529 }
1530 } else {
1531 if (UseFP && !(isFixed || isCSR))
1532 SVEAreaOffset = -SVEStackSize;
1533 if (!UseFP && (isFixed || isCSR))
1534 SVEAreaOffset = SVEStackSize;
1535 }
1536
1537 if (UseFP) {
1538 FrameReg = RegInfo->getFrameRegister(MF);
1539 return StackOffset::getFixed(FPOffset) + SVEAreaOffset;
1540 }
1541
1542 // Use the base pointer if we have one.
1543 if (RegInfo->hasBasePointer(MF))
1544 FrameReg = RegInfo->getBaseRegister();
1545 else {
1546 assert(!MFI.hasVarSizedObjects() &&
1547 "Can't use SP when we have var sized objects.");
1548 FrameReg = AArch64::SP;
1549 // If we're using the red zone for this function, the SP won't actually
1550 // be adjusted, so the offsets will be negative. They're also all
1551 // within range of the signed 9-bit immediate instructions.
1552 if (canUseRedZone(MF))
1553 Offset -= AFI->getLocalStackSize();
1554 }
1555
1556 return StackOffset::getFixed(Offset) + SVEAreaOffset;
1557}
1558
1560 // Do not set a kill flag on values that are also marked as live-in. This
1561 // happens with the @llvm-returnaddress intrinsic and with arguments passed in
1562 // callee saved registers.
1563 // Omitting the kill flags is conservatively correct even if the live-in
1564 // is not used after all.
1565 bool IsLiveIn = MF.getRegInfo().isLiveIn(Reg);
1566 return getKillRegState(!IsLiveIn);
1567}
1568
1570 MachineFunction &MF) {
1571 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1574 return Subtarget.isTargetMachO() &&
1575 !(Subtarget.getTargetLowering()->supportSwiftError() &&
1576 Attrs.hasAttrSomewhere(Attribute::SwiftError)) &&
1578 !AFL.requiresSaveVG(MF) && !AFI->isSVECC();
1579}
1580
1581static bool invalidateWindowsRegisterPairing(bool SpillExtendedVolatile,
1582 unsigned SpillCount, unsigned Reg1,
1583 unsigned Reg2, bool NeedsWinCFI,
1584 const TargetRegisterInfo *TRI) {
1585 // If we are generating register pairs for a Windows function that requires
1586 // EH support, then pair consecutive registers only. There are no unwind
1587 // opcodes for saves/restores of non-consecutive register pairs.
1588 // The unwind opcodes are save_regp, save_regp_x, save_fregp, save_frepg_x,
1589 // save_lrpair.
1590 // https://docs.microsoft.com/en-us/cpp/build/arm64-exception-handling
1591
1592 if (Reg2 == AArch64::FP)
1593 return true;
1594 if (!NeedsWinCFI)
1595 return false;
1596
1597 // ARM64EC introduced `save_any_regp`, which expects 16-byte alignment.
1598 // This is handled by only allowing paired spills for registers spilled at
1599 // even positions (which should be 16-byte aligned, as other GPRs/FPRs are
1600 // 8-bytes). We carve out an exception for {FP,LR}, which does not require
1601 // 16-byte alignment in the uop representation.
1602 if (TRI->getEncodingValue(Reg2) == TRI->getEncodingValue(Reg1) + 1)
1603 return SpillExtendedVolatile
1604 ? !((Reg1 == AArch64::FP && Reg2 == AArch64::LR) ||
1605 (SpillCount % 2) == 0)
1606 : false;
1607
1608 // If pairing a GPR with LR, the pair can be described by the save_lrpair
1609 // opcode. The save_lrpair opcode requires the first register to be odd.
1610 if (Reg1 >= AArch64::X19 && Reg1 <= AArch64::X27 &&
1611 (Reg1 - AArch64::X19) % 2 == 0 && Reg2 == AArch64::LR)
1612 return false;
1613 return true;
1614}
1615
1616/// Returns true if Reg1 and Reg2 cannot be paired using a ldp/stp instruction.
1617/// WindowsCFI requires that only consecutive registers can be paired.
1618/// LR and FP need to be allocated together when the frame needs to save
1619/// the frame-record. This means any other register pairing with LR is invalid.
1620static bool invalidateRegisterPairing(bool SpillExtendedVolatile,
1621 unsigned SpillCount, unsigned Reg1,
1622 unsigned Reg2, bool UsesWinAAPCS,
1623 bool NeedsWinCFI, bool NeedsFrameRecord,
1624 const TargetRegisterInfo *TRI) {
1625 if (UsesWinAAPCS)
1626 return invalidateWindowsRegisterPairing(SpillExtendedVolatile, SpillCount,
1627 Reg1, Reg2, NeedsWinCFI, TRI);
1628
1629 // If we need to store the frame record, don't pair any register
1630 // with LR other than FP.
1631 if (NeedsFrameRecord)
1632 return Reg2 == AArch64::LR;
1633
1634 return false;
1635}
1636
1637namespace {
1638
1639struct RegPairInfo {
1640 Register Reg1;
1641 Register Reg2;
1642 int FrameIdx;
1643 int Offset;
1644 enum RegType { GPR, FPR64, FPR128, PPR, ZPR, VG } Type;
1645 const TargetRegisterClass *RC;
1646
1647 RegPairInfo() = default;
1648
1649 bool isPaired() const { return Reg2.isValid(); }
1650
1651 bool isScalable() const { return Type == PPR || Type == ZPR; }
1652};
1653
1654} // end anonymous namespace
1655
1657 for (unsigned PReg = AArch64::P8; PReg <= AArch64::P15; ++PReg) {
1658 if (SavedRegs.test(PReg)) {
1659 unsigned PNReg = PReg - AArch64::P0 + AArch64::PN0;
1660 return MCRegister(PNReg);
1661 }
1662 }
1663 return MCRegister();
1664}
1665
1666// The multivector LD/ST are available only for SME or SVE2p1 targets
1668 MachineFunction &MF) {
1670 return false;
1671
1672 SMEAttrs FuncAttrs = MF.getInfo<AArch64FunctionInfo>()->getSMEFnAttrs();
1673 bool IsLocallyStreaming =
1674 FuncAttrs.hasStreamingBody() && !FuncAttrs.hasStreamingInterface();
1675
1676 // Only when in streaming mode SME2 instructions can be safely used.
1677 // It is not safe to use SME2 instructions when in streaming compatible or
1678 // locally streaming mode.
1679 return Subtarget.hasSVE2p1() ||
1680 (Subtarget.hasSME2() &&
1681 (!IsLocallyStreaming && Subtarget.isStreaming()));
1682}
1683
1685 MachineFunction &MF,
1687 const TargetRegisterInfo *TRI,
1689 bool NeedsFrameRecord) {
1690
1691 if (CSI.empty())
1692 return;
1693
1694 bool IsWindows = isTargetWindows(MF);
1696 unsigned StackHazardSize = getStackHazardSize(MF);
1697 MachineFrameInfo &MFI = MF.getFrameInfo();
1699 unsigned Count = CSI.size();
1700 (void)CC;
1701 // MachO's compact unwind format relies on all registers being stored in
1702 // pairs.
1703 assert((!produceCompactUnwindFrame(AFL, MF) ||
1706 (Count & 1) == 0) &&
1707 "Odd number of callee-saved regs to spill!");
1708 int ByteOffset = AFI->getCalleeSavedStackSize();
1709 int StackFillDir = -1;
1710 int RegInc = 1;
1711 unsigned FirstReg = 0;
1712 if (IsWindows) {
1713 // For WinCFI, fill the stack from the bottom up.
1714 ByteOffset = 0;
1715 StackFillDir = 1;
1716 // As the CSI array is reversed to match PrologEpilogInserter, iterate
1717 // backwards, to pair up registers starting from lower numbered registers.
1718 RegInc = -1;
1719 FirstReg = Count - 1;
1720 }
1721
1722 bool FPAfterSVECalleeSaves = AFL.hasSVECalleeSavesAboveFrameRecord(MF);
1723 // Windows AAPCS has x9-x15 as volatile registers, x16-x17 as intra-procedural
1724 // scratch, x18 as platform reserved. However, clang has extended calling
1725 // convensions such as preserve_most and preserve_all which treat these as
1726 // CSR. As such, the ARM64 unwind uOPs bias registers by 19. We use ARM64EC
1727 // uOPs which have separate restrictions. We need to check for that.
1728 //
1729 // NOTE: we currently do not account for the D registers as LLVM does not
1730 // support non-ABI compliant D register spills.
1731 bool SpillExtendedVolatile =
1732 IsWindows && llvm::any_of(CSI, [](const CalleeSavedInfo &CSI) {
1733 const auto &Reg = CSI.getReg();
1734 return Reg >= AArch64::X0 && Reg <= AArch64::X18;
1735 });
1736
1737 int ZPRByteOffset = 0;
1738 int PPRByteOffset = 0;
1739 bool SplitPPRs = AFI->hasSplitSVEObjects();
1740 if (SplitPPRs) {
1741 ZPRByteOffset = AFI->getZPRCalleeSavedStackSize();
1742 PPRByteOffset = AFI->getPPRCalleeSavedStackSize();
1743 } else if (!FPAfterSVECalleeSaves) {
1744 ZPRByteOffset =
1746 // Unused: Everything goes in ZPR space.
1747 PPRByteOffset = 0;
1748 }
1749
1750 bool NeedGapToAlignStack = AFI->hasCalleeSaveStackFreeSpace();
1751 Register LastReg = 0;
1752 bool HasCSHazardPadding = AFI->hasStackHazardSlotIndex() && !SplitPPRs;
1753
1754 auto AlignOffset = [StackFillDir](int Offset, int Align) {
1755 if (StackFillDir < 0)
1756 return alignDown(Offset, Align);
1757 return alignTo(Offset, Align);
1758 };
1759
1760 // When iterating backwards, the loop condition relies on unsigned wraparound.
1761 for (unsigned i = FirstReg; i < Count; i += RegInc) {
1762 RegPairInfo RPI;
1763 RPI.Reg1 = CSI[i].getReg();
1764
1765 if (AArch64::GPR64RegClass.contains(RPI.Reg1)) {
1766 RPI.Type = RegPairInfo::GPR;
1767 RPI.RC = &AArch64::GPR64RegClass;
1768 } else if (AArch64::FPR64RegClass.contains(RPI.Reg1)) {
1769 RPI.Type = RegPairInfo::FPR64;
1770 RPI.RC = &AArch64::FPR64RegClass;
1771 } else if (AArch64::FPR128RegClass.contains(RPI.Reg1)) {
1772 RPI.Type = RegPairInfo::FPR128;
1773 RPI.RC = &AArch64::FPR128RegClass;
1774 } else if (AArch64::ZPRRegClass.contains(RPI.Reg1)) {
1775 RPI.Type = RegPairInfo::ZPR;
1776 RPI.RC = &AArch64::ZPRRegClass;
1777 } else if (AArch64::PPRRegClass.contains(RPI.Reg1)) {
1778 RPI.Type = RegPairInfo::PPR;
1779 RPI.RC = &AArch64::PPRRegClass;
1780 } else if (RPI.Reg1 == AArch64::VG) {
1781 RPI.Type = RegPairInfo::VG;
1782 RPI.RC = &AArch64::FIXED_REGSRegClass;
1783 } else {
1784 llvm_unreachable("Unsupported register class.");
1785 }
1786
1787 int &ScalableByteOffset = RPI.Type == RegPairInfo::PPR && SplitPPRs
1788 ? PPRByteOffset
1789 : ZPRByteOffset;
1790
1791 // Add the stack hazard size as we transition from GPR->FPR CSRs.
1792 if (HasCSHazardPadding &&
1793 (!LastReg || !AArch64InstrInfo::isFpOrNEON(LastReg)) &&
1795 ByteOffset += StackFillDir * StackHazardSize;
1796 LastReg = RPI.Reg1;
1797
1798 bool NeedsWinCFI = AFL.needsWinCFI(MF);
1799 int Scale = TRI->getSpillSize(*RPI.RC);
1800 // Add the next reg to the pair if it is in the same register class.
1801 if (unsigned(i + RegInc) < Count && !HasCSHazardPadding) {
1802 MCRegister NextReg = CSI[i + RegInc].getReg();
1803 unsigned SpillCount = NeedsWinCFI ? FirstReg - i : i;
1804 int Aligned = AlignOffset(ByteOffset, Scale);
1805 int PairOffset = IsWindows ? Aligned : Aligned + StackFillDir * 2 * Scale;
1806 bool PairFitsImmRange =
1807 PairOffset / Scale >= -64 && PairOffset / Scale <= 63;
1808 switch (RPI.Type) {
1809 case RegPairInfo::GPR:
1810 if (AArch64::GPR64RegClass.contains(NextReg) && PairFitsImmRange &&
1811 !invalidateRegisterPairing(SpillExtendedVolatile, SpillCount,
1812 RPI.Reg1, NextReg, IsWindows,
1813 NeedsWinCFI, NeedsFrameRecord, TRI))
1814 RPI.Reg2 = NextReg;
1815 break;
1816 case RegPairInfo::FPR64:
1817 if (AArch64::FPR64RegClass.contains(NextReg) && PairFitsImmRange &&
1818 !invalidateRegisterPairing(SpillExtendedVolatile, SpillCount,
1819 RPI.Reg1, NextReg, IsWindows,
1820 NeedsWinCFI, NeedsFrameRecord, TRI))
1821 RPI.Reg2 = NextReg;
1822 break;
1823 case RegPairInfo::FPR128:
1824 if (AArch64::FPR128RegClass.contains(NextReg) && PairFitsImmRange)
1825 RPI.Reg2 = NextReg;
1826 break;
1827 case RegPairInfo::PPR:
1828 break;
1829 case RegPairInfo::ZPR:
1830 if (!NeedsWinCFI && AFI->getPredicateRegForFillSpill() != 0 &&
1831 ((RPI.Reg1 - AArch64::Z0) & 1) == 0 && (NextReg == RPI.Reg1 + 1)) {
1832 // Calculate offset of register pair to see if pair instruction can be
1833 // used.
1834 int Offset = (ScalableByteOffset + StackFillDir * 2 * Scale) / Scale;
1835 if ((-16 <= Offset && Offset <= 14) && (Offset % 2 == 0))
1836 RPI.Reg2 = NextReg;
1837 }
1838 break;
1839 case RegPairInfo::VG:
1840 break;
1841 }
1842 }
1843
1844 // GPRs and FPRs are saved in pairs of 64-bit regs. We expect the CSI
1845 // list to come in sorted by frame index so that we can issue the store
1846 // pair instructions directly. Assert if we see anything otherwise.
1847 //
1848 // The order of the registers in the list is controlled by
1849 // getCalleeSavedRegs(), so they will always be in-order, as well.
1850 assert((!RPI.isPaired() ||
1851 (CSI[i].getFrameIdx() + RegInc == CSI[i + RegInc].getFrameIdx())) &&
1852 "Out of order callee saved regs!");
1853
1854 assert((!RPI.isPaired() || !NeedsFrameRecord || RPI.Reg2 != AArch64::FP ||
1855 RPI.Reg1 == AArch64::LR) &&
1856 "FrameRecord must be allocated together with LR");
1857
1858 // Windows AAPCS has FP and LR reversed.
1859 assert((!RPI.isPaired() || !NeedsFrameRecord || RPI.Reg1 != AArch64::FP ||
1860 RPI.Reg2 == AArch64::LR) &&
1861 "FrameRecord must be allocated together with LR");
1862
1863 // MachO's compact unwind format relies on all registers being stored in
1864 // adjacent register pairs.
1865 assert((!produceCompactUnwindFrame(AFL, MF) ||
1868 (RPI.isPaired() &&
1869 ((RPI.Reg1 == AArch64::LR && RPI.Reg2 == AArch64::FP) ||
1870 RPI.Reg1 + 1 == RPI.Reg2))) &&
1871 "Callee-save registers not saved as adjacent register pair!");
1872
1873 RPI.FrameIdx = CSI[i].getFrameIdx();
1874 if (IsWindows &&
1875 RPI.isPaired()) // RPI.FrameIdx must be the lower index of the pair
1876 RPI.FrameIdx = CSI[i + RegInc].getFrameIdx();
1877
1878 // Realign the scalable offset if necessary. This is relevant when spilling
1879 // predicates on Windows.
1880 if (RPI.isScalable() && ScalableByteOffset % Scale != 0)
1881 ScalableByteOffset = AlignOffset(ScalableByteOffset, Scale);
1882
1883 // Realign the fixed offset if necessary. This is relevant when spilling Q
1884 // registers after spilling an odd amount of X registers.
1885 if (!RPI.isScalable() && ByteOffset % Scale != 0)
1886 ByteOffset = AlignOffset(ByteOffset, Scale);
1887
1888 int OffsetPre = RPI.isScalable() ? ScalableByteOffset : ByteOffset;
1889 assert(OffsetPre % Scale == 0);
1890
1891 if (RPI.isScalable())
1892 ScalableByteOffset += StackFillDir * (RPI.isPaired() ? 2 * Scale : Scale);
1893 else
1894 ByteOffset += StackFillDir * (RPI.isPaired() ? 2 * Scale : Scale);
1895
1896 // Swift's async context is directly before FP, so allocate an extra
1897 // 8 bytes for it.
1898 if (NeedsFrameRecord && AFI->hasSwiftAsyncContext() &&
1899 ((!IsWindows && RPI.Reg2 == AArch64::FP) ||
1900 (IsWindows && RPI.Reg2 == AArch64::LR)))
1901 ByteOffset += StackFillDir * 8;
1902
1903 // Round up size of non-pair to pair size if we need to pad the
1904 // callee-save area to ensure 16-byte alignment.
1905 if (NeedGapToAlignStack && !IsWindows && !RPI.isScalable() &&
1906 RPI.Type != RegPairInfo::FPR128 && !RPI.isPaired() &&
1907 ByteOffset % 16 != 0) {
1908 ByteOffset += 8 * StackFillDir;
1909 assert(MFI.getObjectAlign(RPI.FrameIdx) <= Align(16));
1910 // A stack frame with a gap looks like this, bottom up:
1911 // d9, d8. x21, gap, x20, x19.
1912 // Set extra alignment on the x21 object to create the gap above it.
1913 MFI.setObjectAlignment(RPI.FrameIdx, Align(16));
1914 NeedGapToAlignStack = false;
1915 }
1916
1917 int OffsetPost = RPI.isScalable() ? ScalableByteOffset : ByteOffset;
1918 assert(OffsetPost % Scale == 0);
1919 // If filling top down (default), we want the offset after incrementing it.
1920 // If filling bottom up (WinCFI) we need the original offset.
1921 int Offset = IsWindows ? OffsetPre : OffsetPost;
1922
1923 // The FP, LR pair goes 8 bytes into our expanded 24-byte slot so that the
1924 // Swift context can directly precede FP.
1925 if (NeedsFrameRecord && AFI->hasSwiftAsyncContext() &&
1926 ((!IsWindows && RPI.Reg2 == AArch64::FP) ||
1927 (IsWindows && RPI.Reg2 == AArch64::LR)))
1928 Offset += 8;
1929 RPI.Offset = Offset / Scale;
1930
1931 assert((!RPI.isPaired() ||
1932 (!RPI.isScalable() && RPI.Offset >= -64 && RPI.Offset <= 63) ||
1933 (RPI.isScalable() && RPI.Offset >= -256 && RPI.Offset <= 255)) &&
1934 "Offset out of bounds for LDP/STP immediate");
1935
1936 auto isFrameRecord = [&] {
1937 if (RPI.isPaired())
1938 return IsWindows ? RPI.Reg1 == AArch64::FP && RPI.Reg2 == AArch64::LR
1939 : RPI.Reg1 == AArch64::LR && RPI.Reg2 == AArch64::FP;
1940 // Otherwise, look for the frame record as two unpaired registers. This is
1941 // needed for -aarch64-stack-hazard-size=<val>, which disables register
1942 // pairing (as the padding may be too large for the LDP/STP offset). Note:
1943 // On Windows, this check works out as current reg == FP, next reg == LR,
1944 // and on other platforms current reg == FP, previous reg == LR. This
1945 // works out as the correct pre-increment or post-increment offsets
1946 // respectively.
1947 return i > 0 && RPI.Reg1 == AArch64::FP &&
1948 CSI[i - 1].getReg() == AArch64::LR;
1949 };
1950
1951 // Save the offset to frame record so that the FP register can point to the
1952 // innermost frame record (spilled FP and LR registers).
1953 if (NeedsFrameRecord && isFrameRecord())
1955
1956 RegPairs.push_back(RPI);
1957 if (RPI.isPaired())
1958 i += RegInc;
1959 }
1960 if (IsWindows) {
1961 // If we need an alignment gap in the stack, align the topmost stack
1962 // object. A stack frame with a gap looks like this, bottom up:
1963 // x19, d8. d9, gap.
1964 // Set extra alignment on the topmost stack object (the first element in
1965 // CSI, which goes top down), to create the gap above it.
1966 if (AFI->hasCalleeSaveStackFreeSpace())
1967 MFI.setObjectAlignment(CSI[0].getFrameIdx(), Align(16));
1968 // We iterated bottom up over the registers; flip RegPairs back to top
1969 // down order.
1970 std::reverse(RegPairs.begin(), RegPairs.end());
1971 }
1972}
1973
1977 MachineFunction &MF = *MBB.getParent();
1978 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1979 auto &TLI = *Subtarget.getTargetLowering();
1980 const AArch64InstrInfo &TII = *Subtarget.getInstrInfo();
1981 bool NeedsWinCFI = needsWinCFI(MF);
1982 DebugLoc DL;
1984
1985 computeCalleeSaveRegisterPairs(*this, MF, CSI, TRI, RegPairs, hasFP(MF));
1986
1987 MachineRegisterInfo &MRI = MF.getRegInfo();
1988 // Refresh the reserved regs in case there are any potential changes since the
1989 // last freeze.
1990 MRI.freezeReservedRegs();
1991
1992 if (homogeneousPrologEpilog(MF)) {
1993 auto MIB = BuildMI(MBB, MI, DL, TII.get(AArch64::HOM_Prolog))
1995
1996 for (auto &RPI : RegPairs) {
1997 MIB.addReg(RPI.Reg1);
1998 MIB.addReg(RPI.Reg2);
1999
2000 // Update register live in.
2001 if (!MRI.isReserved(RPI.Reg1))
2002 MBB.addLiveIn(RPI.Reg1);
2003 if (RPI.isPaired() && !MRI.isReserved(RPI.Reg2))
2004 MBB.addLiveIn(RPI.Reg2);
2005 }
2006 return true;
2007 }
2008 bool PTrueCreated = false;
2009 for (const RegPairInfo &RPI : llvm::reverse(RegPairs)) {
2010 Register Reg1 = RPI.Reg1;
2011 Register Reg2 = RPI.Reg2;
2012 unsigned StrOpc;
2013
2014 // Issue sequence of spills for cs regs. The first spill may be converted
2015 // to a pre-decrement store later by emitPrologue if the callee-save stack
2016 // area allocation can't be combined with the local stack area allocation.
2017 // For example:
2018 // stp x22, x21, [sp, #0] // addImm(+0)
2019 // stp x20, x19, [sp, #16] // addImm(+2)
2020 // stp fp, lr, [sp, #32] // addImm(+4)
2021 // Rationale: This sequence saves uop updates compared to a sequence of
2022 // pre-increment spills like stp xi,xj,[sp,#-16]!
2023 // Note: Similar rationale and sequence for restores in epilog.
2024 unsigned Size = TRI->getSpillSize(*RPI.RC);
2025 Align Alignment = TRI->getSpillAlign(*RPI.RC);
2026 switch (RPI.Type) {
2027 case RegPairInfo::GPR:
2028 StrOpc = RPI.isPaired() ? AArch64::STPXi : AArch64::STRXui;
2029 break;
2030 case RegPairInfo::FPR64:
2031 StrOpc = RPI.isPaired() ? AArch64::STPDi : AArch64::STRDui;
2032 break;
2033 case RegPairInfo::FPR128:
2034 StrOpc = RPI.isPaired() ? AArch64::STPQi : AArch64::STRQui;
2035 break;
2036 case RegPairInfo::ZPR:
2037 StrOpc = RPI.isPaired() ? AArch64::ST1B_2Z_IMM : AArch64::STR_ZXI;
2038 break;
2039 case RegPairInfo::PPR:
2040 StrOpc = AArch64::STR_PXI;
2041 break;
2042 case RegPairInfo::VG:
2043 StrOpc = AArch64::STRXui;
2044 break;
2045 }
2046
2047 Register X0Scratch;
2048 llvm::scope_exit RestoreX0([&] {
2049 if (X0Scratch != AArch64::NoRegister)
2050 BuildMI(MBB, MI, DL, TII.get(TargetOpcode::COPY), AArch64::X0)
2051 .addReg(X0Scratch)
2053 });
2054
2055 if (Reg1 == AArch64::VG) {
2056 // Find an available register to store value of VG to.
2057 Reg1 = findScratchNonCalleeSaveRegister(&MBB, true);
2058 assert(Reg1 != AArch64::NoRegister);
2059 if (MF.getSubtarget<AArch64Subtarget>().hasSVE()) {
2060 BuildMI(MBB, MI, DL, TII.get(AArch64::CNTD_XPiI), Reg1)
2061 .addImm(31)
2062 .addImm(1)
2064 } else {
2066 if (any_of(MBB.liveins(),
2067 [&STI](const MachineBasicBlock::RegisterMaskPair &LiveIn) {
2068 return STI.getRegisterInfo()->isSuperOrSubRegisterEq(
2069 AArch64::X0, LiveIn.PhysReg);
2070 })) {
2071 X0Scratch = Reg1;
2072 BuildMI(MBB, MI, DL, TII.get(TargetOpcode::COPY), X0Scratch)
2073 .addReg(AArch64::X0)
2075 }
2076
2077 RTLIB::Libcall LC = RTLIB::SMEABI_GET_CURRENT_VG;
2078 const uint32_t *RegMask =
2079 TRI->getCallPreservedMask(MF, TLI.getLibcallCallingConv(LC));
2080 BuildMI(MBB, MI, DL, TII.get(AArch64::BL))
2081 .addExternalSymbol(TLI.getLibcallName(LC))
2082 .addRegMask(RegMask)
2083 .addReg(AArch64::X0, RegState::ImplicitDefine)
2085 Reg1 = AArch64::X0;
2086 }
2087 }
2088
2089 LLVM_DEBUG({
2090 dbgs() << "CSR spill: (" << printReg(Reg1, TRI);
2091 if (RPI.isPaired())
2092 dbgs() << ", " << printReg(Reg2, TRI);
2093 dbgs() << ") -> fi#(" << RPI.FrameIdx;
2094 if (RPI.isPaired())
2095 dbgs() << ", " << RPI.FrameIdx + 1;
2096 dbgs() << ")\n";
2097 });
2098
2099 assert((!isTargetWindows(MF) ||
2100 !(Reg1 == AArch64::LR && Reg2 == AArch64::FP)) &&
2101 "Windows unwdinding requires a consecutive (FP,LR) pair");
2102 // Windows unwind codes require consecutive registers if registers are
2103 // paired. Make the switch here, so that the code below will save (x,x+1)
2104 // and not (x+1,x).
2105 unsigned FrameIdxReg1 = RPI.FrameIdx;
2106 unsigned FrameIdxReg2 = RPI.FrameIdx + 1;
2107 if (isTargetWindows(MF) && RPI.isPaired()) {
2108 std::swap(Reg1, Reg2);
2109 std::swap(FrameIdxReg1, FrameIdxReg2);
2110 }
2111
2112 if (RPI.isPaired() && RPI.isScalable()) {
2113 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2116 unsigned PnReg = AFI->getPredicateRegForFillSpill();
2117 assert((PnReg != 0 && enableMultiVectorSpillFill(Subtarget, MF)) &&
2118 "Expects SVE2.1 or SME2 target and a predicate register");
2119#ifdef EXPENSIVE_CHECKS
2120 auto IsPPR = [](const RegPairInfo &c) {
2121 return c.Reg1 == RegPairInfo::PPR;
2122 };
2123 auto PPRBegin = std::find_if(RegPairs.begin(), RegPairs.end(), IsPPR);
2124 auto IsZPR = [](const RegPairInfo &c) {
2125 return c.Type == RegPairInfo::ZPR;
2126 };
2127 auto ZPRBegin = std::find_if(RegPairs.begin(), RegPairs.end(), IsZPR);
2128 assert(!(PPRBegin < ZPRBegin) &&
2129 "Expected callee save predicate to be handled first");
2130#endif
2131 if (!PTrueCreated) {
2132 PTrueCreated = true;
2133 BuildMI(MBB, MI, DL, TII.get(AArch64::PTRUE_C_B), PnReg)
2135 }
2136 MachineInstrBuilder MIB = BuildMI(MBB, MI, DL, TII.get(StrOpc));
2137 if (!MRI.isReserved(Reg1))
2138 MBB.addLiveIn(Reg1);
2139 if (!MRI.isReserved(Reg2))
2140 MBB.addLiveIn(Reg2);
2141 MIB.addReg(/*PairRegs*/ AArch64::Z0_Z1 + (RPI.Reg1 - AArch64::Z0));
2143 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2144 MachineMemOperand::MOStore, Size, Alignment));
2145 MIB.addReg(PnReg);
2146 MIB.addReg(AArch64::SP)
2147 .addImm(RPI.Offset / 2) // [sp, #imm*2*vscale],
2148 // where 2*vscale is implicit
2151 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2152 MachineMemOperand::MOStore, Size, Alignment));
2153 if (NeedsWinCFI)
2154 insertSEH(MIB, TII, MachineInstr::FrameSetup);
2155 } else { // The code when the pair of ZReg is not present
2156 MachineInstrBuilder MIB = BuildMI(MBB, MI, DL, TII.get(StrOpc));
2157 if (!MRI.isReserved(Reg1))
2158 MBB.addLiveIn(Reg1);
2159 if (RPI.isPaired()) {
2160 if (!MRI.isReserved(Reg2))
2161 MBB.addLiveIn(Reg2);
2162 MIB.addReg(Reg2, getPrologueDeath(MF, Reg2));
2164 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2165 MachineMemOperand::MOStore, Size, Alignment));
2166 }
2167 MIB.addReg(Reg1, getPrologueDeath(MF, Reg1))
2168 .addReg(AArch64::SP)
2169 .addImm(RPI.Offset) // [sp, #offset*vscale],
2170 // where factor*vscale is implicit
2173 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2174 MachineMemOperand::MOStore, Size, Alignment));
2175 if (NeedsWinCFI)
2176 insertSEH(MIB, TII, MachineInstr::FrameSetup);
2177 }
2178 // Update the StackIDs of the SVE stack slots.
2179 MachineFrameInfo &MFI = MF.getFrameInfo();
2180 if (RPI.Type == RegPairInfo::ZPR) {
2181 MFI.setStackID(FrameIdxReg1, TargetStackID::ScalableVector);
2182 if (RPI.isPaired())
2183 MFI.setStackID(FrameIdxReg2, TargetStackID::ScalableVector);
2184 } else if (RPI.Type == RegPairInfo::PPR) {
2186 if (RPI.isPaired())
2188 }
2189 }
2190 return true;
2191}
2192
2196 MachineFunction &MF = *MBB.getParent();
2197 const AArch64InstrInfo &TII =
2198 *MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
2199 DebugLoc DL;
2201 bool NeedsWinCFI = needsWinCFI(MF);
2202
2203 if (MBBI != MBB.end())
2204 DL = MBBI->getDebugLoc();
2205
2206 computeCalleeSaveRegisterPairs(*this, MF, CSI, TRI, RegPairs, hasFP(MF));
2207 if (homogeneousPrologEpilog(MF, &MBB)) {
2208 auto MIB = BuildMI(MBB, MBBI, DL, TII.get(AArch64::HOM_Epilog))
2210 for (auto &RPI : RegPairs) {
2211 MIB.addReg(RPI.Reg1, RegState::Define);
2212 MIB.addReg(RPI.Reg2, RegState::Define);
2213 }
2214 return true;
2215 }
2216
2217 // For performance reasons restore SVE register in increasing order
2218 auto IsPPR = [](const RegPairInfo &c) { return c.Type == RegPairInfo::PPR; };
2219 auto PPRBegin = llvm::find_if(RegPairs, IsPPR);
2220 auto PPREnd = std::find_if_not(PPRBegin, RegPairs.end(), IsPPR);
2221 std::reverse(PPRBegin, PPREnd);
2222 auto IsZPR = [](const RegPairInfo &c) { return c.Type == RegPairInfo::ZPR; };
2223 auto ZPRBegin = llvm::find_if(RegPairs, IsZPR);
2224 auto ZPREnd = std::find_if_not(ZPRBegin, RegPairs.end(), IsZPR);
2225 std::reverse(ZPRBegin, ZPREnd);
2226
2227 bool PTrueCreated = false;
2228 for (const RegPairInfo &RPI : RegPairs) {
2229 Register Reg1 = RPI.Reg1;
2230 Register Reg2 = RPI.Reg2;
2231
2232 // Issue sequence of restores for cs regs. The last restore may be converted
2233 // to a post-increment load later by emitEpilogue if the callee-save stack
2234 // area allocation can't be combined with the local stack area allocation.
2235 // For example:
2236 // ldp fp, lr, [sp, #32] // addImm(+4)
2237 // ldp x20, x19, [sp, #16] // addImm(+2)
2238 // ldp x22, x21, [sp, #0] // addImm(+0)
2239 // Note: see comment in spillCalleeSavedRegisters()
2240 unsigned LdrOpc;
2241 unsigned Size = TRI->getSpillSize(*RPI.RC);
2242 Align Alignment = TRI->getSpillAlign(*RPI.RC);
2243 switch (RPI.Type) {
2244 case RegPairInfo::GPR:
2245 LdrOpc = RPI.isPaired() ? AArch64::LDPXi : AArch64::LDRXui;
2246 break;
2247 case RegPairInfo::FPR64:
2248 LdrOpc = RPI.isPaired() ? AArch64::LDPDi : AArch64::LDRDui;
2249 break;
2250 case RegPairInfo::FPR128:
2251 LdrOpc = RPI.isPaired() ? AArch64::LDPQi : AArch64::LDRQui;
2252 break;
2253 case RegPairInfo::ZPR:
2254 LdrOpc = RPI.isPaired() ? AArch64::LD1B_2Z_IMM : AArch64::LDR_ZXI;
2255 break;
2256 case RegPairInfo::PPR:
2257 LdrOpc = AArch64::LDR_PXI;
2258 break;
2259 case RegPairInfo::VG:
2260 continue;
2261 }
2262 LLVM_DEBUG({
2263 dbgs() << "CSR restore: (" << printReg(Reg1, TRI);
2264 if (RPI.isPaired())
2265 dbgs() << ", " << printReg(Reg2, TRI);
2266 dbgs() << ") -> fi#(" << RPI.FrameIdx;
2267 if (RPI.isPaired())
2268 dbgs() << ", " << RPI.FrameIdx + 1;
2269 dbgs() << ")\n";
2270 });
2271
2272 // Windows unwind codes require consecutive registers if registers are
2273 // paired. Make the switch here, so that the code below will save (x,x+1)
2274 // and not (x+1,x).
2275 unsigned FrameIdxReg1 = RPI.FrameIdx;
2276 unsigned FrameIdxReg2 = RPI.FrameIdx + 1;
2277 if (isTargetWindows(MF) && RPI.isPaired()) {
2278 std::swap(Reg1, Reg2);
2279 std::swap(FrameIdxReg1, FrameIdxReg2);
2280 }
2281
2283 if (RPI.isPaired() && RPI.isScalable()) {
2284 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2286 unsigned PnReg = AFI->getPredicateRegForFillSpill();
2287 assert((PnReg != 0 && enableMultiVectorSpillFill(Subtarget, MF)) &&
2288 "Expects SVE2.1 or SME2 target and a predicate register");
2289#ifdef EXPENSIVE_CHECKS
2290 assert(!(PPRBegin < ZPRBegin) &&
2291 "Expected callee save predicate to be handled first");
2292#endif
2293 if (!PTrueCreated) {
2294 PTrueCreated = true;
2295 BuildMI(MBB, MBBI, DL, TII.get(AArch64::PTRUE_C_B), PnReg)
2297 }
2298 MachineInstrBuilder MIB = BuildMI(MBB, MBBI, DL, TII.get(LdrOpc));
2299 MIB.addReg(/*PairRegs*/ AArch64::Z0_Z1 + (RPI.Reg1 - AArch64::Z0),
2300 getDefRegState(true));
2302 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2303 MachineMemOperand::MOLoad, Size, Alignment));
2304 MIB.addReg(PnReg);
2305 MIB.addReg(AArch64::SP)
2306 .addImm(RPI.Offset / 2) // [sp, #imm*2*vscale]
2307 // where 2*vscale is implicit
2310 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2311 MachineMemOperand::MOLoad, Size, Alignment));
2312 if (NeedsWinCFI)
2313 insertSEH(MIB, TII, MachineInstr::FrameDestroy);
2314 } else {
2315 MachineInstrBuilder MIB = BuildMI(MBB, MBBI, DL, TII.get(LdrOpc));
2316 if (RPI.isPaired()) {
2317 MIB.addReg(Reg2, getDefRegState(true));
2319 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2320 MachineMemOperand::MOLoad, Size, Alignment));
2321 }
2322 MIB.addReg(Reg1, getDefRegState(true));
2323 MIB.addReg(AArch64::SP)
2324 .addImm(RPI.Offset) // [sp, #offset*vscale]
2325 // where factor*vscale is implicit
2328 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2329 MachineMemOperand::MOLoad, Size, Alignment));
2330 if (NeedsWinCFI)
2331 insertSEH(MIB, TII, MachineInstr::FrameDestroy);
2332 }
2333 }
2334 return true;
2335}
2336
2337// Return the FrameID for a MMO.
2338static std::optional<int> getMMOFrameID(MachineMemOperand *MMO,
2339 const MachineFrameInfo &MFI) {
2340 auto *PSV =
2342 if (PSV)
2343 return std::optional<int>(PSV->getFrameIndex());
2344
2345 if (MMO->getValue()) {
2346 if (auto *Al = dyn_cast<AllocaInst>(getUnderlyingObject(MMO->getValue()))) {
2347 for (int FI = MFI.getObjectIndexBegin(); FI < MFI.getObjectIndexEnd();
2348 FI++)
2349 if (MFI.getObjectAllocation(FI) == Al)
2350 return FI;
2351 }
2352 }
2353
2354 return std::nullopt;
2355}
2356
2357// Return the FrameID for a Load/Store instruction by looking at the first MMO.
2358static std::optional<int> getLdStFrameID(const MachineInstr &MI,
2359 const MachineFrameInfo &MFI) {
2360 if (!MI.mayLoadOrStore() || MI.getNumMemOperands() < 1)
2361 return std::nullopt;
2362
2363 return getMMOFrameID(*MI.memoperands_begin(), MFI);
2364}
2365
2366// Returns true if the LDST MachineInstr \p MI is a PPR access.
2367static bool isPPRAccess(const MachineInstr &MI) {
2368 return AArch64::PPRRegClass.contains(MI.getOperand(0).getReg());
2369}
2370
2371// Check if a Hazard slot is needed for the current function, and if so create
2372// one for it. The index is stored in AArch64FunctionInfo->StackHazardSlotIndex,
2373// which can be used to determine if any hazard padding is needed.
2374void AArch64FrameLowering::determineStackHazardSlot(
2375 MachineFunction &MF, BitVector &SavedRegs) const {
2376 unsigned StackHazardSize = getStackHazardSize(MF);
2377 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2378 if (StackHazardSize == 0 || StackHazardSize % 16 != 0 ||
2380 return;
2381
2382 // Stack hazards are only needed in streaming functions.
2383 SMEAttrs Attrs = AFI->getSMEFnAttrs();
2384 if (!StackHazardInNonStreaming && Attrs.hasNonStreamingInterfaceAndBody())
2385 return;
2386
2387 MachineFrameInfo &MFI = MF.getFrameInfo();
2388
2389 // Add a hazard slot if there are any CSR FPR registers, or are any fp-only
2390 // stack objects.
2391 bool HasFPRCSRs = any_of(SavedRegs.set_bits(), [](unsigned Reg) {
2392 return AArch64::FPR64RegClass.contains(Reg) ||
2393 AArch64::FPR128RegClass.contains(Reg) ||
2394 AArch64::ZPRRegClass.contains(Reg);
2395 });
2396 bool HasPPRCSRs = any_of(SavedRegs.set_bits(), [](unsigned Reg) {
2397 return AArch64::PPRRegClass.contains(Reg);
2398 });
2399 bool HasFPRStackObjects = false;
2400 bool HasPPRStackObjects = false;
2401 if (!HasFPRCSRs || SplitSVEObjects) {
2402 enum SlotType : uint8_t {
2403 Unknown = 0,
2404 ZPRorFPR = 1 << 0,
2405 PPR = 1 << 1,
2406 GPR = 1 << 2,
2408 };
2409
2410 // Find stack slots solely used for one kind of register (ZPR, PPR, etc.),
2411 // based on the kinds of accesses used in the function.
2412 SmallVector<SlotType> SlotTypes(MFI.getObjectIndexEnd(), SlotType::Unknown);
2413 for (auto &MBB : MF) {
2414 for (auto &MI : MBB) {
2415 std::optional<int> FI = getLdStFrameID(MI, MFI);
2416 if (!FI || FI < 0 || FI > int(SlotTypes.size()))
2417 continue;
2418 if (MFI.hasScalableStackID(*FI)) {
2419 SlotTypes[*FI] |=
2420 isPPRAccess(MI) ? SlotType::PPR : SlotType::ZPRorFPR;
2421 } else {
2422 SlotTypes[*FI] |= AArch64InstrInfo::isFpOrNEON(MI)
2423 ? SlotType::ZPRorFPR
2424 : SlotType::GPR;
2425 }
2426 }
2427 }
2428
2429 for (int FI = 0; FI < int(SlotTypes.size()); ++FI) {
2430 HasFPRStackObjects |= SlotTypes[FI] == SlotType::ZPRorFPR;
2431 // For SplitSVEObjects remember that this stack slot is a predicate, this
2432 // will be needed later when determining the frame layout.
2433 if (SlotTypes[FI] == SlotType::PPR) {
2435 HasPPRStackObjects = true;
2436 }
2437 }
2438 }
2439
2440 if (HasFPRCSRs || HasFPRStackObjects) {
2441 int ID = MFI.CreateStackObject(StackHazardSize, Align(16), false);
2442 LLVM_DEBUG(dbgs() << "Created Hazard slot at " << ID << " size "
2443 << StackHazardSize << "\n");
2444 AFI->setStackHazardSlotIndex(ID);
2445 }
2446
2447 if (!AFI->hasStackHazardSlotIndex())
2448 return;
2449
2450 if (SplitSVEObjects) {
2451 CallingConv::ID CC = MF.getFunction().getCallingConv();
2452 if (AFI->isSVECC() || CC == CallingConv::AArch64_SVE_VectorCall) {
2453 AFI->setSplitSVEObjects(true);
2454 LLVM_DEBUG(dbgs() << "Using SplitSVEObjects for SVE CC function\n");
2455 return;
2456 }
2457
2458 // We only use SplitSVEObjects in non-SVE CC functions if there's a
2459 // possibility of a stack hazard between PPRs and ZPRs/FPRs.
2460 LLVM_DEBUG(dbgs() << "Determining if SplitSVEObjects should be used in "
2461 "non-SVE CC function...\n");
2462
2463 // If another calling convention is explicitly set FPRs can't be promoted to
2464 // ZPR callee-saves.
2466 LLVM_DEBUG(
2467 dbgs()
2468 << "Calling convention is not supported with SplitSVEObjects\n");
2469 return;
2470 }
2471
2472 if (!HasPPRCSRs && !HasPPRStackObjects) {
2473 LLVM_DEBUG(
2474 dbgs() << "Not using SplitSVEObjects as no PPRs are on the stack\n");
2475 return;
2476 }
2477
2478 if (!HasFPRCSRs && !HasFPRStackObjects) {
2479 LLVM_DEBUG(
2480 dbgs()
2481 << "Not using SplitSVEObjects as no FPRs or ZPRs are on the stack\n");
2482 return;
2483 }
2484
2485 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2486 MF.getSubtarget<AArch64Subtarget>();
2488 "Expected SVE to be available for PPRs");
2489
2490 const TargetRegisterInfo *TRI = MF.getSubtarget().getRegisterInfo();
2491 // With SplitSVEObjects the CS hazard padding is placed between the
2492 // PPRs and ZPRs. If there are any FPR CS there would be a hazard between
2493 // them and the CS GRPs. Avoid this by promoting all FPR CS to ZPRs.
2494 BitVector FPRZRegs(SavedRegs.size());
2495 for (size_t Reg = 0, E = SavedRegs.size(); HasFPRCSRs && Reg < E; ++Reg) {
2496 BitVector::reference RegBit = SavedRegs[Reg];
2497 if (!RegBit)
2498 continue;
2499 unsigned SubRegIdx = 0;
2500 if (AArch64::FPR64RegClass.contains(Reg))
2501 SubRegIdx = AArch64::dsub;
2502 else if (AArch64::FPR128RegClass.contains(Reg))
2503 SubRegIdx = AArch64::zsub;
2504 else
2505 continue;
2506 // Clear the bit for the FPR save.
2507 RegBit = false;
2508 // Mark that we should save the corresponding ZPR.
2509 Register ZReg =
2510 TRI->getMatchingSuperReg(Reg, SubRegIdx, &AArch64::ZPRRegClass);
2511 FPRZRegs.set(ZReg);
2512 }
2513 SavedRegs |= FPRZRegs;
2514
2515 AFI->setSplitSVEObjects(true);
2516 LLVM_DEBUG(dbgs() << "SplitSVEObjects enabled!\n");
2517 }
2518}
2519
2521 BitVector &SavedRegs,
2522 RegScavenger *RS) const {
2523 // All calls are tail calls in GHC calling conv, and functions have no
2524 // prologue/epilogue.
2526 return;
2527
2528 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
2529
2531 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
2533 unsigned UnspilledCSGPR = AArch64::NoRegister;
2534 unsigned UnspilledCSGPRPaired = AArch64::NoRegister;
2535
2536 MachineFrameInfo &MFI = MF.getFrameInfo();
2537 const MCPhysReg *CSRegs = MF.getRegInfo().getCalleeSavedRegs();
2538
2539 MCRegister BasePointerReg =
2540 RegInfo->hasBasePointer(MF) ? RegInfo->getBaseRegister() : MCRegister();
2541
2542 unsigned ExtraCSSpill = 0;
2543 bool HasUnpairedGPR64 = false;
2544 bool HasPairZReg = false;
2545 BitVector UserReservedRegs = RegInfo->getUserReservedRegs(MF);
2546 BitVector ReservedRegs = RegInfo->getReservedRegs(MF);
2547
2548 // Figure out which callee-saved registers to save/restore.
2549 for (unsigned i = 0; CSRegs[i]; ++i) {
2550 const MCRegister Reg = CSRegs[i];
2551
2552 // Add the base pointer register to SavedRegs if it is callee-save.
2553 if (Reg == BasePointerReg)
2554 SavedRegs.set(Reg);
2555
2556 // Don't save manually reserved registers set through +reserve-x#i,
2557 // even for callee-saved registers, as per GCC's behavior.
2558 if (UserReservedRegs[Reg]) {
2559 SavedRegs.reset(Reg);
2560 continue;
2561 }
2562
2563 bool RegUsed = SavedRegs.test(Reg);
2564 MCRegister PairedReg;
2565 const bool RegIsGPR64 = AArch64::GPR64RegClass.contains(Reg);
2566 if (RegIsGPR64 || AArch64::FPR64RegClass.contains(Reg) ||
2567 AArch64::FPR128RegClass.contains(Reg)) {
2568 // Compensate for odd numbers of GP CSRs.
2569 // For now, all the known cases of odd number of CSRs are of GPRs.
2570 if (HasUnpairedGPR64)
2571 PairedReg = CSRegs[i % 2 == 0 ? i - 1 : i + 1];
2572 else
2573 PairedReg = CSRegs[i ^ 1];
2574 }
2575
2576 // If the function requires all the GP registers to save (SavedRegs),
2577 // and there are an odd number of GP CSRs at the same time (CSRegs),
2578 // PairedReg could be in a different register class from Reg, which would
2579 // lead to a FPR (usually D8) accidentally being marked saved.
2580 if (RegIsGPR64 && !AArch64::GPR64RegClass.contains(PairedReg)) {
2581 PairedReg = AArch64::NoRegister;
2582 HasUnpairedGPR64 = true;
2583 }
2584 assert(PairedReg == AArch64::NoRegister ||
2585 AArch64::GPR64RegClass.contains(Reg, PairedReg) ||
2586 AArch64::FPR64RegClass.contains(Reg, PairedReg) ||
2587 AArch64::FPR128RegClass.contains(Reg, PairedReg));
2588
2589 if (!RegUsed) {
2590 if (AArch64::GPR64RegClass.contains(Reg) && !ReservedRegs[Reg]) {
2591 UnspilledCSGPR = Reg;
2592 UnspilledCSGPRPaired = PairedReg;
2593 }
2594 continue;
2595 }
2596
2597 // MachO's compact unwind format relies on all registers being stored in
2598 // pairs.
2599 // FIXME: the usual format is actually better if unwinding isn't needed.
2600 if (producePairRegisters(MF) && PairedReg != AArch64::NoRegister &&
2601 !SavedRegs.test(PairedReg)) {
2602 SavedRegs.set(PairedReg);
2603 if (AArch64::GPR64RegClass.contains(PairedReg) &&
2604 !ReservedRegs[PairedReg])
2605 ExtraCSSpill = PairedReg;
2606 }
2607 // Check if there is a pair of ZRegs, so it can select PReg for spill/fill
2608 HasPairZReg |= (AArch64::ZPRRegClass.contains(Reg, CSRegs[i ^ 1]) &&
2609 SavedRegs.test(CSRegs[i ^ 1]));
2610 }
2611
2612 if (HasPairZReg && enableMultiVectorSpillFill(Subtarget, MF)) {
2614 // Find a suitable predicate register for the multi-vector spill/fill
2615 // instructions.
2616 MCRegister PnReg = findFreePredicateReg(SavedRegs);
2617 if (PnReg.isValid())
2618 AFI->setPredicateRegForFillSpill(PnReg);
2619 // If no free callee-save has been found assign one.
2620 if (!AFI->getPredicateRegForFillSpill() &&
2621 MF.getFunction().getCallingConv() ==
2623 SavedRegs.set(AArch64::P8);
2624 AFI->setPredicateRegForFillSpill(AArch64::PN8);
2625 }
2626
2627 assert(!ReservedRegs[AFI->getPredicateRegForFillSpill()] &&
2628 "Predicate cannot be a reserved register");
2629 }
2630
2632 !Subtarget.isTargetWindows()) {
2633 // For Windows calling convention on a non-windows OS, where X18 is treated
2634 // as reserved, back up X18 when entering non-windows code (marked with the
2635 // Windows calling convention) and restore when returning regardless of
2636 // whether the individual function uses it - it might call other functions
2637 // that clobber it.
2638 SavedRegs.set(AArch64::X18);
2639 }
2640
2641 // Determine if a Hazard slot should be used and where it should go.
2642 // If SplitSVEObjects is used, the hazard padding is placed between the PPRs
2643 // and ZPRs. Otherwise, it goes in the callee save area.
2644 determineStackHazardSlot(MF, SavedRegs);
2645
2646 // Calculates the callee saved stack size.
2647 unsigned CSStackSize = 0;
2648 unsigned ZPRCSStackSize = 0;
2649 unsigned PPRCSStackSize = 0;
2651 for (unsigned Reg : SavedRegs.set_bits()) {
2652 auto *RC = TRI->getMinimalPhysRegClass(MCRegister(Reg));
2653 assert(RC && "expected register class!");
2654 auto SpillSize = TRI->getSpillSize(*RC);
2655 bool IsZPR = AArch64::ZPRRegClass.contains(Reg);
2656 bool IsPPR = !IsZPR && AArch64::PPRRegClass.contains(Reg);
2657 if (IsZPR)
2658 ZPRCSStackSize += SpillSize;
2659 else if (IsPPR)
2660 PPRCSStackSize += SpillSize;
2661 else {
2662 // A register and its super-register can both appear in SavedRegs.
2663 // Only the widest register is actually spilled, so skip such
2664 // sub-registers here to avoid double-counting the overlap.
2665 bool SavedSuper = any_of(TRI->superregs(Reg), [&](MCPhysReg SuperReg) {
2666 return SavedRegs.test(SuperReg);
2667 });
2668 if (!SavedSuper)
2669 CSStackSize += SpillSize;
2670 }
2671 }
2672
2673 // Save number of saved regs, so we can easily update CSStackSize later to
2674 // account for any additional 64-bit GPR saves. Note: After this point
2675 // only 64-bit GPRs can be added to SavedRegs.
2676 unsigned NumSavedRegs = SavedRegs.count();
2677
2678 // If we have hazard padding in the CS area add that to the size.
2680 CSStackSize += getStackHazardSize(MF);
2681
2682 // Increase the callee-saved stack size if the function has streaming mode
2683 // changes, as we will need to spill the value of the VG register.
2684 if (requiresSaveVG(MF))
2685 CSStackSize += 8;
2686
2687 // If we must call __arm_get_current_vg in the prologue preserve the LR.
2688 if (requiresSaveVG(MF) && !Subtarget.hasSVE())
2689 SavedRegs.set(AArch64::LR);
2690
2691 // The frame record needs to be created by saving the appropriate registers
2692 uint64_t EstimatedStackSize = MFI.estimateStackSize(MF);
2693 if (hasFP(MF) ||
2694 windowsRequiresStackProbe(MF, EstimatedStackSize + CSStackSize + 16)) {
2695 SavedRegs.set(AArch64::FP);
2696 SavedRegs.set(AArch64::LR);
2697 }
2698
2699 LLVM_DEBUG({
2700 dbgs() << "*** determineCalleeSaves\nSaved CSRs:";
2701 for (unsigned Reg : SavedRegs.set_bits())
2702 dbgs() << ' ' << printReg(MCRegister(Reg), RegInfo);
2703 dbgs() << "\n";
2704 });
2705
2706 // If any callee-saved registers are used, the frame cannot be eliminated.
2707 auto [ZPRLocalStackSize, PPRLocalStackSize] =
2709 uint64_t SVELocals = ZPRLocalStackSize + PPRLocalStackSize;
2710 uint64_t SVEStackSize =
2711 alignTo(ZPRCSStackSize + PPRCSStackSize + SVELocals, 16);
2712 bool CanEliminateFrame = (SavedRegs.count() == 0) && !SVEStackSize;
2713
2714 // The CSR spill slots have not been allocated yet, so estimateStackSize
2715 // won't include them.
2716 unsigned EstimatedStackSizeLimit = estimateRSStackSizeLimit(MF);
2717
2718 // We may address some of the stack above the canonical frame address, either
2719 // for our own arguments or during a call. Include that in calculating whether
2720 // we have complicated addressing concerns.
2721 int64_t CalleeStackUsed = 0;
2722 for (int I = MFI.getObjectIndexBegin(); I != 0; ++I) {
2723 int64_t FixedOff = MFI.getObjectOffset(I);
2724 if (FixedOff > CalleeStackUsed)
2725 CalleeStackUsed = FixedOff;
2726 }
2727
2728 // Conservatively always assume BigStack when there are SVE spills.
2729 bool BigStack = SVEStackSize || (EstimatedStackSize + CSStackSize +
2730 CalleeStackUsed) > EstimatedStackSizeLimit;
2731 if (BigStack || !CanEliminateFrame || RegInfo->cannotEliminateFrame(MF))
2732 AFI->setHasStackFrame(true);
2733
2734 // Estimate if we might need to scavenge a register at some point in order
2735 // to materialize a stack offset. If so, either spill one additional
2736 // callee-saved register or reserve a special spill slot to facilitate
2737 // register scavenging. If we already spilled an extra callee-saved register
2738 // above to keep the number of spills even, we don't need to do anything else
2739 // here.
2740 if (BigStack) {
2741 if (!ExtraCSSpill && UnspilledCSGPR != AArch64::NoRegister) {
2742 LLVM_DEBUG(dbgs() << "Spilling " << printReg(UnspilledCSGPR, RegInfo)
2743 << " to get a scratch register.\n");
2744 SavedRegs.set(UnspilledCSGPR);
2745 ExtraCSSpill = UnspilledCSGPR;
2746
2747 // MachO's compact unwind format relies on all registers being stored in
2748 // pairs, so if we need to spill one extra for BigStack, then we need to
2749 // store the pair.
2750 if (producePairRegisters(MF)) {
2751 if (UnspilledCSGPRPaired == AArch64::NoRegister) {
2752 // Failed to make a pair for compact unwind format, revert spilling.
2753 if (produceCompactUnwindFrame(*this, MF)) {
2754 SavedRegs.reset(UnspilledCSGPR);
2755 ExtraCSSpill = AArch64::NoRegister;
2756 }
2757 } else
2758 SavedRegs.set(UnspilledCSGPRPaired);
2759 }
2760 }
2761
2762 // If we didn't find an extra callee-saved register to spill, create
2763 // an emergency spill slot.
2764 if (!ExtraCSSpill || MF.getRegInfo().isPhysRegUsed(ExtraCSSpill)) {
2766 const TargetRegisterClass &RC = AArch64::GPR64RegClass;
2767 unsigned Size = TRI->getSpillSize(RC);
2768 Align Alignment = TRI->getSpillAlign(RC);
2769 int FI = MFI.CreateSpillStackObject(Size, Alignment);
2770 RS->addScavengingFrameIndex(FI);
2771 LLVM_DEBUG(dbgs() << "No available CS registers, allocated fi#" << FI
2772 << " as the emergency spill slot.\n");
2773 }
2774 }
2775
2776 // Adding the size of additional 64bit GPR saves.
2777 CSStackSize += 8 * (SavedRegs.count() - NumSavedRegs);
2778
2779 // A Swift asynchronous context extends the frame record with a pointer
2780 // directly before FP.
2781 if (hasFP(MF) && AFI->hasSwiftAsyncContext())
2782 CSStackSize += 8;
2783
2784 uint64_t AlignedCSStackSize = alignTo(CSStackSize, 16);
2785 LLVM_DEBUG(dbgs() << "Estimated stack frame size: "
2786 << EstimatedStackSize + AlignedCSStackSize << " bytes.\n");
2787
2789 AFI->getCalleeSavedStackSize() == AlignedCSStackSize) &&
2790 "Should not invalidate callee saved info");
2791
2792 // Round up to register pair alignment to avoid additional SP adjustment
2793 // instructions.
2794 AFI->setCalleeSavedStackSize(AlignedCSStackSize);
2795 AFI->setCalleeSaveStackHasFreeSpace(AlignedCSStackSize != CSStackSize);
2796 AFI->setSVECalleeSavedStackSize(ZPRCSStackSize, alignTo(PPRCSStackSize, 16));
2797}
2798
2801 std::vector<CalleeSavedInfo> &CSI) {
2802 // Reorder callee-saved ZPRs to maximize pairing which requires
2803 // consecutive even/odd registers at even scaled stack offsets.
2804 // Additional requirements are checked when the register pairs are formed.
2805 assert(!isTargetWindows(MF) &&
2806 "ZPR callee-save reordering not supported on Windows");
2807
2808 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2809 if (!AFI->getPredicateRegForFillSpill())
2810 return;
2811
2813 SmallVector<size_t> ZPRPositions;
2814
2815 for (auto [Index, CS] : llvm::enumerate(CSI)) {
2816 if (AArch64::ZPRRegClass.contains(CS.getReg())) {
2817 ZPRSaves.push_back(CS);
2818 ZPRPositions.push_back(Index);
2819 }
2820 }
2821
2822 if (ZPRSaves.size() < 2)
2823 return;
2824
2825 llvm::sort(ZPRSaves, [](const auto &A, const auto &B) {
2826 return A.getReg() < B.getReg();
2827 });
2828
2831 for (size_t i = 0; i < ZPRSaves.size();) {
2832 if (i + 1 < ZPRSaves.size() &&
2833 (ZPRSaves[i].getReg() + 1 == ZPRSaves[i + 1].getReg()) &&
2834 (ZPRSaves[i].getReg() - AArch64::Z0) % 2 == 0) {
2835 Pairs.emplace_back(ZPRSaves[i], ZPRSaves[i + 1]);
2836 i += 2;
2837 } else {
2838 Singles.push_back(ZPRSaves[i++]);
2839 }
2840 }
2841
2842 // If the lowest offset is odd, select one register to spill here
2843 // so subsequent pairs begin at an even offset.
2844 int ZPRByteOffset = AFI->getZPRCalleeSavedStackSize();
2845 if (!AFI->hasSplitSVEObjects())
2846 ZPRByteOffset += AFI->getPPRCalleeSavedStackSize();
2847
2848 const int Scale = RegInfo->getSpillSize(AArch64::ZPRRegClass);
2849 const int LowestOffset =
2850 (ZPRByteOffset / Scale) - static_cast<int>(ZPRSaves.size());
2851
2852 // Prefer the highest single otherwise split the highest pair.
2853 std::optional<CalleeSavedInfo> AlignmentSingle;
2854 if (LowestOffset % 2 != 0) {
2855 if (!Singles.empty()) {
2856 AlignmentSingle = Singles.pop_back_val();
2857 } else {
2858 assert(!Pairs.empty() && "Expected a ZPR pair to split");
2859 auto [Even, Odd] = Pairs.pop_back_val();
2860 AlignmentSingle = Odd;
2861 Singles.push_back(Even);
2862 }
2863 }
2864
2865 if (Pairs.empty())
2866 return;
2867
2868 // Build ZPRs so reverse spill emission processes the leading single first,
2869 // followed by the candidate even/odd pairs and remaining singles.
2870 SmallVector<CalleeSavedInfo> ZPRSavesInCSIOrder;
2871 llvm::append_range(ZPRSavesInCSIOrder, Singles);
2872
2873 for (const auto &[Even, Odd] : Pairs) {
2874 ZPRSavesInCSIOrder.push_back(Even);
2875 ZPRSavesInCSIOrder.push_back(Odd);
2876 }
2877
2878 if (AlignmentSingle)
2879 ZPRSavesInCSIOrder.push_back(*AlignmentSingle);
2880
2881 assert(ZPRSavesInCSIOrder.size() == ZPRPositions.size() &&
2882 "Reordering should not change the number of ZPR spills");
2883 for (auto [Position, CS] : llvm::zip(ZPRPositions, ZPRSavesInCSIOrder))
2884 CSI[Position] = CS;
2885}
2886
2888 MachineFunction &MF, const TargetRegisterInfo *RegInfo,
2889 std::vector<CalleeSavedInfo> &CSI) const {
2890 bool IsWindows = isTargetWindows(MF);
2891 unsigned StackHazardSize = getStackHazardSize(MF);
2892 // To match the canonical windows frame layout, reverse the list of
2893 // callee saved registers to get them laid out by PrologEpilogInserter
2894 // in the right order. (PrologEpilogInserter allocates stack objects top
2895 // down. Windows canonical prologs store higher numbered registers at
2896 // the top, thus have the CSI array start from the highest registers.)
2897 if (IsWindows)
2898 std::reverse(CSI.begin(), CSI.end());
2899
2900 if (CSI.empty())
2901 return true; // Early exit if no callee saved registers are modified!
2902
2903 // Now that we know which registers need to be saved and restored, allocate
2904 // stack slots for them.
2905 MachineFrameInfo &MFI = MF.getFrameInfo();
2906 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2907
2908 // Insert VG into the list of CSRs, immediately before LR if saved.
2909 if (requiresSaveVG(MF)) {
2910 CalleeSavedInfo VGInfo(AArch64::VG);
2911 auto It =
2912 find_if(CSI, [](auto &Info) { return Info.getReg() == AArch64::LR; });
2913 if (It != CSI.end())
2914 CSI.insert(It, VGInfo);
2915 else
2916 CSI.push_back(VGInfo);
2917 }
2918
2919 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
2920 if (!IsWindows && enableMultiVectorSpillFill(Subtarget, MF))
2921 // The Windows stack layout is not supported by this reordering function
2922 // yet.
2923 orderZPRCalleeSavesForPairs(MF, RegInfo, CSI);
2924
2925 Register LastReg = 0;
2926 int HazardSlotIndex = std::numeric_limits<int>::max();
2927 for (auto &CS : CSI) {
2928 MCRegister Reg = CS.getReg();
2929 const TargetRegisterClass *RC = RegInfo->getMinimalPhysRegClass(Reg);
2930
2931 // Create a hazard slot as we switch between GPR and FPR CSRs.
2933 (!LastReg || !AArch64InstrInfo::isFpOrNEON(LastReg)) &&
2935 assert(HazardSlotIndex == std::numeric_limits<int>::max() &&
2936 "Unexpected register order for hazard slot");
2937 HazardSlotIndex = MFI.CreateStackObject(StackHazardSize, Align(8), true);
2938 LLVM_DEBUG(dbgs() << "Created CSR Hazard at slot " << HazardSlotIndex
2939 << "\n");
2940 AFI->setStackHazardCSRSlotIndex(HazardSlotIndex);
2941 MFI.setIsCalleeSavedObjectIndex(HazardSlotIndex, true);
2942 }
2943
2944 unsigned Size = RegInfo->getSpillSize(*RC);
2945 Align Alignment(RegInfo->getSpillAlign(*RC));
2946 int FrameIdx = MFI.CreateStackObject(Size, Alignment, true);
2947 CS.setFrameIdx(FrameIdx);
2948 MFI.setIsCalleeSavedObjectIndex(FrameIdx, true);
2949
2950 // Grab 8 bytes below FP for the extended asynchronous frame info.
2951 if (hasFP(MF) && AFI->hasSwiftAsyncContext() && Reg == AArch64::FP) {
2952 FrameIdx = MFI.CreateStackObject(8, Alignment, true);
2953 AFI->setSwiftAsyncContextFrameIdx(FrameIdx);
2954 MFI.setIsCalleeSavedObjectIndex(FrameIdx, true);
2955 }
2956 LastReg = Reg;
2957 }
2958
2959 // Add hazard slot in the case where no FPR CSRs are present.
2961 HazardSlotIndex == std::numeric_limits<int>::max()) {
2962 HazardSlotIndex = MFI.CreateStackObject(StackHazardSize, Align(8), true);
2963 LLVM_DEBUG(dbgs() << "Created CSR Hazard at slot " << HazardSlotIndex
2964 << "\n");
2965 AFI->setStackHazardCSRSlotIndex(HazardSlotIndex);
2966 MFI.setIsCalleeSavedObjectIndex(HazardSlotIndex, true);
2967 }
2968
2969 return true;
2970}
2971
2973 const MachineFunction &MF) const {
2975 // If the function has streaming-mode changes, don't scavenge a
2976 // spillslot in the callee-save area, as that might require an
2977 // 'addvl' in the streaming-mode-changing call-sequence when the
2978 // function doesn't use a FP.
2979 if (AFI->hasStreamingModeChanges() && !hasFP(MF))
2980 return false;
2981 // Don't allow register salvaging with hazard slots, in case it moves objects
2982 // into the wrong place.
2983 if (AFI->hasStackHazardSlotIndex())
2984 return false;
2985 return AFI->hasCalleeSaveStackFreeSpace();
2986}
2987
2988/// returns true if there are any SVE callee saves.
2990 int &Min, int &Max) {
2991 Min = std::numeric_limits<int>::max();
2992 Max = std::numeric_limits<int>::min();
2993
2994 if (!MFI.isCalleeSavedInfoValid())
2995 return false;
2996
2997 const std::vector<CalleeSavedInfo> &CSI = MFI.getCalleeSavedInfo();
2998 for (auto &CS : CSI) {
2999 if (AArch64::ZPRRegClass.contains(CS.getReg()) ||
3000 AArch64::PPRRegClass.contains(CS.getReg())) {
3001 assert((Max == std::numeric_limits<int>::min() ||
3002 Max + 1 == CS.getFrameIdx()) &&
3003 "SVE CalleeSaves are not consecutive");
3004 Min = std::min(Min, CS.getFrameIdx());
3005 Max = std::max(Max, CS.getFrameIdx());
3006 }
3007 }
3008 return Min != std::numeric_limits<int>::max();
3009}
3010
3012 AssignObjectOffsets AssignOffsets) {
3013 MachineFrameInfo &MFI = MF.getFrameInfo();
3014 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
3015
3016 SVEStackSizes SVEStack{};
3017
3018 // With SplitSVEObjects we maintain separate stack offsets for predicates
3019 // (PPRs) and SVE vectors (ZPRs). When SplitSVEObjects is disabled predicates
3020 // are included in the SVE vector area.
3021 uint64_t &ZPRStackTop = SVEStack.ZPRStackSize;
3022 uint64_t &PPRStackTop =
3023 AFI->hasSplitSVEObjects() ? SVEStack.PPRStackSize : SVEStack.ZPRStackSize;
3024
3025#ifndef NDEBUG
3026 // First process all fixed stack objects.
3027 for (int I = MFI.getObjectIndexBegin(); I != 0; ++I)
3028 assert(!MFI.hasScalableStackID(I) &&
3029 "SVE vectors should never be passed on the stack by value, only by "
3030 "reference.");
3031#endif
3032
3033 auto AllocateObject = [&](int FI) {
3035 ? ZPRStackTop
3036 : PPRStackTop;
3037
3038 // FIXME: Given that the length of SVE vectors is not necessarily a power of
3039 // two, we'd need to align every object dynamically at runtime if the
3040 // alignment is larger than 16. This is not yet supported.
3041 Align Alignment = MFI.getObjectAlign(FI);
3042 if (Alignment > Align(16))
3044 "Alignment of scalable vectors > 16 bytes is not yet supported");
3045
3046 StackTop += MFI.getObjectSize(FI);
3047 StackTop = alignTo(StackTop, Alignment);
3048
3049 assert(StackTop < (uint64_t)std::numeric_limits<int64_t>::max() &&
3050 "SVE StackTop far too large?!");
3051
3052 int64_t Offset = -int64_t(StackTop);
3053 if (AssignOffsets == AssignObjectOffsets::Yes)
3054 MFI.setObjectOffset(FI, Offset);
3055
3056 LLVM_DEBUG(dbgs() << "alloc FI(" << FI << ") at SP[" << Offset << "]\n");
3057 };
3058
3059 // Then process all callee saved slots.
3060 int MinCSFrameIndex, MaxCSFrameIndex;
3061 if (getSVECalleeSaveSlotRange(MFI, MinCSFrameIndex, MaxCSFrameIndex)) {
3062 for (int FI = MinCSFrameIndex; FI <= MaxCSFrameIndex; ++FI)
3063 AllocateObject(FI);
3064 }
3065
3066 // Ensure the CS area is 16-byte aligned.
3067 PPRStackTop = alignTo(PPRStackTop, Align(16U));
3068 ZPRStackTop = alignTo(ZPRStackTop, Align(16U));
3069
3070 // Create a buffer of SVE objects to allocate and sort it.
3071 SmallVector<int, 8> ObjectsToAllocate;
3072 // If we have a stack protector, and we've previously decided that we have SVE
3073 // objects on the stack and thus need it to go in the SVE stack area, then it
3074 // needs to go first.
3075 int StackProtectorFI = -1;
3076 if (MFI.hasStackProtectorIndex()) {
3077 StackProtectorFI = MFI.getStackProtectorIndex();
3078 if (MFI.getStackID(StackProtectorFI) == TargetStackID::ScalableVector)
3079 ObjectsToAllocate.push_back(StackProtectorFI);
3080 }
3081
3082 for (int FI = 0, E = MFI.getObjectIndexEnd(); FI != E; ++FI) {
3083 if (FI == StackProtectorFI || MFI.isDeadObjectIndex(FI) ||
3085 continue;
3086
3089 continue;
3090
3091 ObjectsToAllocate.push_back(FI);
3092 }
3093
3094 // Allocate all SVE locals and spills
3095 for (unsigned FI : ObjectsToAllocate)
3096 AllocateObject(FI);
3097
3098 PPRStackTop = alignTo(PPRStackTop, Align(16U));
3099 ZPRStackTop = alignTo(ZPRStackTop, Align(16U));
3100
3101 if (AssignOffsets == AssignObjectOffsets::Yes)
3102 AFI->setStackSizeSVE(SVEStack.ZPRStackSize, SVEStack.PPRStackSize);
3103
3104 return SVEStack;
3105}
3106
3108 MachineFunction &MF, RegScavenger *RS) const {
3110 "Upwards growing stack unsupported");
3111
3113
3114 // If this function isn't doing Win64-style C++ EH, we don't need to do
3115 // anything.
3116 if (!MF.hasEHFunclets())
3117 return;
3118
3119 MachineFrameInfo &MFI = MF.getFrameInfo();
3120 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
3121
3122 // Win64 C++ EH needs to allocate space for the catch objects in the fixed
3123 // object area right next to the UnwindHelp object.
3124 WinEHFuncInfo &EHInfo = *MF.getWinEHFuncInfo();
3125 int64_t CurrentOffset =
3127 for (WinEHTryBlockMapEntry &TBME : EHInfo.TryBlockMap) {
3128 for (WinEHHandlerType &H : TBME.HandlerArray) {
3129 int FrameIndex = H.CatchObj.FrameIndex;
3130 if ((FrameIndex != INT_MAX) && MFI.getObjectOffset(FrameIndex) == 0) {
3131 CurrentOffset =
3132 alignTo(CurrentOffset, MFI.getObjectAlign(FrameIndex).value());
3133 CurrentOffset += MFI.getObjectSize(FrameIndex);
3134 MFI.setObjectOffset(FrameIndex, -CurrentOffset);
3135 }
3136 }
3137 }
3138
3139 // Create an UnwindHelp object.
3140 // The UnwindHelp object is allocated at the start of the fixed object area
3141 int64_t UnwindHelpOffset = alignTo(CurrentOffset + 8, Align(16));
3142 assert(UnwindHelpOffset == getFixedObjectSize(MF, AFI, /*IsWin64*/ true,
3143 /*IsFunclet*/ false) &&
3144 "UnwindHelpOffset must be at the start of the fixed object area");
3145 int UnwindHelpFI = MFI.CreateFixedObject(/*Size*/ 8, -UnwindHelpOffset,
3146 /*IsImmutable=*/false);
3147 EHInfo.UnwindHelpFrameIdx = UnwindHelpFI;
3148
3149 MachineBasicBlock &MBB = MF.front();
3150 auto MBBI = MBB.begin();
3151 while (MBBI != MBB.end() && MBBI->getFlag(MachineInstr::FrameSetup))
3152 ++MBBI;
3153
3154 // We need to store -2 into the UnwindHelp object at the start of the
3155 // function.
3156 DebugLoc DL;
3157 RS->enterBasicBlockEnd(MBB);
3158 RS->backward(MBBI);
3159 Register DstReg = RS->FindUnusedReg(&AArch64::GPR64commonRegClass);
3160 assert(DstReg && "There must be a free register after frame setup");
3161 const AArch64InstrInfo &TII =
3162 *MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3163 BuildMI(MBB, MBBI, DL, TII.get(AArch64::MOVi64imm), DstReg).addImm(-2);
3164 BuildMI(MBB, MBBI, DL, TII.get(AArch64::STURXi))
3165 .addReg(DstReg, getKillRegState(true))
3166 .addFrameIndex(UnwindHelpFI)
3167 .addImm(0);
3168}
3169
3170namespace {
3171struct TagStoreInstr {
3173 int64_t Offset, Size;
3174 explicit TagStoreInstr(MachineInstr *MI, int64_t Offset, int64_t Size)
3175 : MI(MI), Offset(Offset), Size(Size) {}
3176};
3177
3178class TagStoreEdit {
3179 MachineFunction *MF;
3180 MachineBasicBlock *MBB;
3181 MachineRegisterInfo *MRI;
3182 // Tag store instructions that are being replaced.
3184 // Combined memref arguments of the above instructions.
3186
3187 // Replace allocation tags in [FrameReg + FrameRegOffset, FrameReg +
3188 // FrameRegOffset + Size) with the address tag of SP.
3189 Register FrameReg;
3190 StackOffset FrameRegOffset;
3191 int64_t Size;
3192 // If not std::nullopt, move FrameReg to (FrameReg + FrameRegUpdate) at the
3193 // end.
3194 std::optional<int64_t> FrameRegUpdate;
3195 // MIFlags for any FrameReg updating instructions.
3196 unsigned FrameRegUpdateFlags;
3197
3198 // Use zeroing instruction variants.
3199 bool ZeroData;
3200 DebugLoc DL;
3201
3202 void emitUnrolled(MachineBasicBlock::iterator InsertI);
3203 void emitLoop(MachineBasicBlock::iterator InsertI);
3204
3205public:
3206 TagStoreEdit(MachineBasicBlock *MBB, bool ZeroData)
3207 : MBB(MBB), ZeroData(ZeroData) {
3208 MF = MBB->getParent();
3209 MRI = &MF->getRegInfo();
3210 }
3211 // Add an instruction to be replaced. Instructions must be added in the
3212 // ascending order of Offset, and have to be adjacent.
3213 void addInstruction(TagStoreInstr I) {
3214 assert((TagStores.empty() ||
3215 TagStores.back().Offset + TagStores.back().Size == I.Offset) &&
3216 "Non-adjacent tag store instructions.");
3217 TagStores.push_back(I);
3218 }
3219 void clear() { TagStores.clear(); }
3220 // Emit equivalent code at the given location, and erase the current set of
3221 // instructions. May skip if the replacement is not profitable. May invalidate
3222 // the input iterator and replace it with a valid one.
3223 void emitCode(MachineBasicBlock::iterator &InsertI,
3224 const AArch64FrameLowering *TFI, bool TryMergeSPUpdate);
3225};
3226
3227void TagStoreEdit::emitUnrolled(MachineBasicBlock::iterator InsertI) {
3228 const AArch64InstrInfo *TII =
3229 MF->getSubtarget<AArch64Subtarget>().getInstrInfo();
3230
3231 const int64_t kMinOffset = -256 * 16;
3232 const int64_t kMaxOffset = 255 * 16;
3233
3234 Register BaseReg = FrameReg;
3235 int64_t BaseRegOffsetBytes = FrameRegOffset.getFixed();
3236 if (BaseRegOffsetBytes < kMinOffset ||
3237 BaseRegOffsetBytes + (Size - Size % 32) > kMaxOffset ||
3238 // BaseReg can be FP, which is not necessarily aligned to 16-bytes. In
3239 // that case, BaseRegOffsetBytes will not be aligned to 16 bytes, which
3240 // is required for the offset of ST2G.
3241 BaseRegOffsetBytes % 16 != 0) {
3242 Register ScratchReg = MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3243 emitFrameOffset(*MBB, InsertI, DL, ScratchReg, BaseReg,
3244 StackOffset::getFixed(BaseRegOffsetBytes), TII);
3245 BaseReg = ScratchReg;
3246 BaseRegOffsetBytes = 0;
3247 }
3248
3249 MachineInstr *LastI = nullptr;
3250 while (Size) {
3251 int64_t InstrSize = (Size > 16) ? 32 : 16;
3252 unsigned Opcode =
3253 InstrSize == 16
3254 ? (ZeroData ? AArch64::STZGi : AArch64::STGi)
3255 : (ZeroData ? AArch64::STZ2Gi : AArch64::ST2Gi);
3256 assert(BaseRegOffsetBytes % 16 == 0);
3257 MachineInstr *I = BuildMI(*MBB, InsertI, DL, TII->get(Opcode))
3258 .addReg(AArch64::SP)
3259 .addReg(BaseReg)
3260 .addImm(BaseRegOffsetBytes / 16)
3261 .setMemRefs(CombinedMemRefs);
3262 // A store to [BaseReg, #0] should go last for an opportunity to fold the
3263 // final SP adjustment in the epilogue.
3264 if (BaseRegOffsetBytes == 0)
3265 LastI = I;
3266 BaseRegOffsetBytes += InstrSize;
3267 Size -= InstrSize;
3268 }
3269
3270 if (LastI)
3271 MBB->splice(InsertI, MBB, LastI);
3272}
3273
3274void TagStoreEdit::emitLoop(MachineBasicBlock::iterator InsertI) {
3275 const AArch64InstrInfo *TII =
3276 MF->getSubtarget<AArch64Subtarget>().getInstrInfo();
3277
3278 Register BaseReg = FrameRegUpdate
3279 ? FrameReg
3280 : MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3281 Register SizeReg = MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3282
3283 emitFrameOffset(*MBB, InsertI, DL, BaseReg, FrameReg, FrameRegOffset, TII);
3284
3285 int64_t LoopSize = Size;
3286 // If the loop size is not a multiple of 32, split off one 16-byte store at
3287 // the end to fold BaseReg update into.
3288 if (FrameRegUpdate && *FrameRegUpdate)
3289 LoopSize -= LoopSize % 32;
3290 MachineInstr *LoopI = BuildMI(*MBB, InsertI, DL,
3291 TII->get(ZeroData ? AArch64::STZGloop_wback
3292 : AArch64::STGloop_wback))
3293 .addDef(SizeReg)
3294 .addDef(BaseReg)
3295 .addImm(LoopSize)
3296 .addReg(BaseReg)
3297 .setMemRefs(CombinedMemRefs);
3298 if (FrameRegUpdate)
3299 LoopI->setFlags(FrameRegUpdateFlags);
3300
3301 int64_t ExtraBaseRegUpdate =
3302 FrameRegUpdate ? (*FrameRegUpdate - FrameRegOffset.getFixed() - Size) : 0;
3303 LLVM_DEBUG(dbgs() << "TagStoreEdit::emitLoop: LoopSize=" << LoopSize
3304 << ", Size=" << Size
3305 << ", ExtraBaseRegUpdate=" << ExtraBaseRegUpdate
3306 << ", FrameRegUpdate=" << FrameRegUpdate
3307 << ", FrameRegOffset.getFixed()="
3308 << FrameRegOffset.getFixed() << "\n");
3309 if (LoopSize < Size) {
3310 assert(FrameRegUpdate);
3311 assert(Size - LoopSize == 16);
3312 // Tag 16 more bytes at BaseReg and update BaseReg.
3313 int64_t STGOffset = ExtraBaseRegUpdate + 16;
3314 assert(STGOffset % 16 == 0 && STGOffset >= -4096 && STGOffset <= 4080 &&
3315 "STG immediate out of range");
3316 BuildMI(*MBB, InsertI, DL,
3317 TII->get(ZeroData ? AArch64::STZGPostIndex : AArch64::STGPostIndex))
3318 .addDef(BaseReg)
3319 .addReg(BaseReg)
3320 .addReg(BaseReg)
3321 .addImm(STGOffset / 16)
3322 .setMemRefs(CombinedMemRefs)
3323 .setMIFlags(FrameRegUpdateFlags);
3324 } else if (ExtraBaseRegUpdate) {
3325 // Update BaseReg.
3326 int64_t AddSubOffset = std::abs(ExtraBaseRegUpdate);
3327 assert(AddSubOffset <= 4095 && "ADD/SUB immediate out of range");
3328 BuildMI(
3329 *MBB, InsertI, DL,
3330 TII->get(ExtraBaseRegUpdate > 0 ? AArch64::ADDXri : AArch64::SUBXri))
3331 .addDef(BaseReg)
3332 .addReg(BaseReg)
3333 .addImm(AddSubOffset)
3334 .addImm(0)
3335 .setMIFlags(FrameRegUpdateFlags);
3336 }
3337}
3338
3339// Check if *II is a register update that can be merged into STGloop that ends
3340// at (Reg + Size). RemainingOffset is the required adjustment to Reg after the
3341// end of the loop.
3342bool canMergeRegUpdate(MachineBasicBlock::iterator II, unsigned Reg,
3343 int64_t Size, int64_t *TotalOffset) {
3344 MachineInstr &MI = *II;
3345 if ((MI.getOpcode() == AArch64::ADDXri ||
3346 MI.getOpcode() == AArch64::SUBXri) &&
3347 MI.getOperand(0).getReg() == Reg && MI.getOperand(1).getReg() == Reg) {
3348 unsigned Shift = AArch64_AM::getShiftValue(MI.getOperand(3).getImm());
3349 int64_t Offset = MI.getOperand(2).getImm() << Shift;
3350 if (MI.getOpcode() == AArch64::SUBXri)
3351 Offset = -Offset;
3352 int64_t PostOffset = Offset - Size;
3353 // TagStoreEdit::emitLoop might emit either an ADD/SUB after the loop, or
3354 // an STGPostIndex which does the last 16 bytes of tag write. Which one is
3355 // chosen depends on the alignment of the loop size, but the difference
3356 // between the valid ranges for the two instructions is small, so we
3357 // conservatively assume that it could be either case here.
3358 //
3359 // Max offset of STGPostIndex, minus the 16 byte tag write folded into that
3360 // instruction.
3361 const int64_t kMaxOffset = 4080 - 16;
3362 // Max offset of SUBXri.
3363 const int64_t kMinOffset = -4095;
3364 if (PostOffset <= kMaxOffset && PostOffset >= kMinOffset &&
3365 PostOffset % 16 == 0) {
3366 *TotalOffset = Offset;
3367 return true;
3368 }
3369 }
3370 return false;
3371}
3372
3373void mergeMemRefs(const SmallVectorImpl<TagStoreInstr> &TSE,
3375 MemRefs.clear();
3376 for (auto &TS : TSE) {
3377 MachineInstr *MI = TS.MI;
3378 // An instruction without memory operands may access anything. Be
3379 // conservative and return an empty list.
3380 if (MI->memoperands_empty()) {
3381 MemRefs.clear();
3382 return;
3383 }
3384 MemRefs.append(MI->memoperands_begin(), MI->memoperands_end());
3385 }
3386}
3387
3388void TagStoreEdit::emitCode(MachineBasicBlock::iterator &InsertI,
3389 const AArch64FrameLowering *TFI,
3390 bool TryMergeSPUpdate) {
3391 if (TagStores.empty())
3392 return;
3393 TagStoreInstr &FirstTagStore = TagStores[0];
3394 TagStoreInstr &LastTagStore = TagStores[TagStores.size() - 1];
3395 Size = LastTagStore.Offset - FirstTagStore.Offset + LastTagStore.Size;
3396 DL = TagStores[0].MI->getDebugLoc();
3397
3398 Register Reg;
3399 FrameRegOffset = TFI->resolveFrameOffsetReference(
3400 *MF, FirstTagStore.Offset, false /*isFixed*/,
3401 TargetStackID::Default /*StackID*/, Reg,
3402 /*PreferFP=*/false, /*ForSimm=*/true);
3403 FrameReg = Reg;
3404 FrameRegUpdate = std::nullopt;
3405
3406 mergeMemRefs(TagStores, CombinedMemRefs);
3407
3408 LLVM_DEBUG({
3409 dbgs() << "Replacing adjacent STG instructions:\n";
3410 for (const auto &Instr : TagStores) {
3411 dbgs() << " " << *Instr.MI;
3412 }
3413 });
3414
3415 // Size threshold where a loop becomes shorter than a linear sequence of
3416 // tagging instructions.
3417 const int kSetTagLoopThreshold = 176;
3418 if (Size < kSetTagLoopThreshold) {
3419 if (TagStores.size() < 2)
3420 return;
3421 emitUnrolled(InsertI);
3422 } else {
3423 MachineInstr *UpdateInstr = nullptr;
3424 int64_t TotalOffset = 0;
3425 if (TryMergeSPUpdate) {
3426 // See if we can merge base register update into the STGloop.
3427 // This is done in AArch64LoadStoreOptimizer for "normal" stores,
3428 // but STGloop is way too unusual for that, and also it only
3429 // realistically happens in function epilogue. Also, STGloop is expanded
3430 // before that pass.
3431 if (InsertI != MBB->end() &&
3432 canMergeRegUpdate(InsertI, FrameReg, FrameRegOffset.getFixed() + Size,
3433 &TotalOffset)) {
3434 UpdateInstr = &*InsertI++;
3435 LLVM_DEBUG(dbgs() << "Folding SP update into loop:\n "
3436 << *UpdateInstr);
3437 }
3438 }
3439
3440 if (!UpdateInstr && TagStores.size() < 2)
3441 return;
3442
3443 if (UpdateInstr) {
3444 FrameRegUpdate = TotalOffset;
3445 FrameRegUpdateFlags = UpdateInstr->getFlags();
3446 }
3447 emitLoop(InsertI);
3448 if (UpdateInstr)
3449 UpdateInstr->eraseFromParent();
3450 }
3451
3452 for (auto &TS : TagStores)
3453 TS.MI->eraseFromParent();
3454}
3455
3456bool isMergeableStackTaggingInstruction(MachineInstr &MI, int64_t &Offset,
3457 int64_t &Size, bool &ZeroData) {
3458 MachineFunction &MF = *MI.getParent()->getParent();
3459 const MachineFrameInfo &MFI = MF.getFrameInfo();
3460
3461 unsigned Opcode = MI.getOpcode();
3462 ZeroData = (Opcode == AArch64::STZGloop || Opcode == AArch64::STZGi ||
3463 Opcode == AArch64::STZ2Gi);
3464
3465 if (Opcode == AArch64::STGloop || Opcode == AArch64::STZGloop) {
3466 if (!MI.getOperand(0).isDead() || !MI.getOperand(1).isDead())
3467 return false;
3468 if (!MI.getOperand(2).isImm() || !MI.getOperand(3).isFI())
3469 return false;
3470 Offset = MFI.getObjectOffset(MI.getOperand(3).getIndex());
3471 Size = MI.getOperand(2).getImm();
3472 return true;
3473 }
3474
3475 if (Opcode == AArch64::STGi || Opcode == AArch64::STZGi)
3476 Size = 16;
3477 else if (Opcode == AArch64::ST2Gi || Opcode == AArch64::STZ2Gi)
3478 Size = 32;
3479 else
3480 return false;
3481
3482 if (MI.getOperand(0).getReg() != AArch64::SP || !MI.getOperand(1).isFI())
3483 return false;
3484
3485 Offset = MFI.getObjectOffset(MI.getOperand(1).getIndex()) +
3486 16 * MI.getOperand(2).getImm();
3487 return true;
3488}
3489
3490static size_t countAvailableScavengerSlots(LivePhysRegs &LiveRegs,
3492 RegScavenger *RS) {
3493 auto FreeGPRs =
3494 llvm::count_if(AArch64::GPR64RegClass, [&LiveRegs, &MRI](auto Reg) {
3495 return LiveRegs.available(MRI, Reg);
3496 });
3497
3498 size_t NumEmergencySlots = 0;
3499 if (RS)
3500 NumEmergencySlots = RS->getNumScavengingFrameIndices();
3501
3502 return FreeGPRs + NumEmergencySlots;
3503}
3504
3505// Detect a run of memory tagging instructions for adjacent stack frame slots,
3506// and replace them with a shorter instruction sequence:
3507// * replace STG + STG with ST2G
3508// * replace STGloop + STGloop with STGloop
3509// This code needs to run when stack slot offsets are already known, but before
3510// FrameIndex operands in STG instructions are eliminated.
3512 const AArch64FrameLowering *TFI,
3513 RegScavenger *RS) {
3514 bool FirstZeroData;
3515 int64_t Size, Offset;
3516 MachineInstr &MI = *II;
3519 if (&MI == &MBB->instr_back())
3520 return II;
3521 if (!isMergeableStackTaggingInstruction(MI, Offset, Size, FirstZeroData))
3522 return II;
3523
3525 Instrs.emplace_back(&MI, Offset, Size);
3526
3527 constexpr int kScanLimit = 10;
3528 int Count = 0;
3530 NextI != E && Count < kScanLimit; ++NextI) {
3531 MachineInstr &MI = *NextI;
3532 bool ZeroData;
3533 int64_t Size, Offset;
3534 // Collect instructions that update memory tags with a FrameIndex operand
3535 // and (when applicable) constant size, and whose output registers are dead
3536 // (the latter is almost always the case in practice). Since these
3537 // instructions effectively have no inputs or outputs, we are free to skip
3538 // any non-aliasing instructions in between without tracking used registers.
3539 if (isMergeableStackTaggingInstruction(MI, Offset, Size, ZeroData)) {
3540 if (ZeroData != FirstZeroData)
3541 break;
3542 Instrs.emplace_back(&MI, Offset, Size);
3543 continue;
3544 }
3545
3546 // Only count non-transient, non-tagging instructions toward the scan
3547 // limit.
3548 if (!MI.isTransient())
3549 ++Count;
3550
3551 // Just in case, stop before the epilogue code starts.
3552 if (MI.getFlag(MachineInstr::FrameSetup) ||
3554 break;
3555
3556 // Reject anything that may alias the collected instructions.
3557 if (MI.mayLoadOrStore() || MI.hasUnmodeledSideEffects() || MI.isCall())
3558 break;
3559 }
3560
3561 // New code will be inserted after the last tagging instruction we've found.
3562 MachineBasicBlock::iterator InsertI = Instrs.back().MI;
3563
3564 // All the gathered stack tag instructions are merged and placed after
3565 // last tag store in the list. The check should be made if the nzcv
3566 // flag is live at the point where we are trying to insert. Otherwise
3567 // the nzcv flag might get clobbered if any stg loops are present.
3568
3569 // FIXME : This approach of bailing out from merge is conservative in
3570 // some ways like even if stg loops are not present after merge the
3571 // insert list, this liveness check is done (which is not needed).
3573 LiveRegs.addLiveOuts(*MBB);
3574 for (auto I = MBB->rbegin();; ++I) {
3575 MachineInstr &MI = *I;
3576 if (MI == InsertI)
3577 break;
3578 LiveRegs.stepBackward(*I);
3579 }
3580 InsertI++;
3581 if (LiveRegs.contains(AArch64::NZCV))
3582 return InsertI;
3583
3584 // Emitting an MTE loop requires two physical registers (BaseReg and
3585 // SizeReg). If the function is under register pressure, the register
3586 // scavenger will crash trying to allocate them. If we don't have at least
3587 // two free slots (free registers + emergency slots), bail out and fall back
3588 // to the unrolled sequence.
3589 if (countAvailableScavengerSlots(LiveRegs, MBB->getParent()->getRegInfo(),
3590 RS) < 2) {
3591 LLVM_DEBUG(
3592 dbgs() << "Failed to merge MTE stack tagging instructions into loop "
3593 << "due to high register pressure.\n");
3594 return InsertI;
3595 }
3596
3597 llvm::stable_sort(Instrs,
3598 [](const TagStoreInstr &Left, const TagStoreInstr &Right) {
3599 return Left.Offset < Right.Offset;
3600 });
3601
3602 // Make sure that we don't have any overlapping stores.
3603 int64_t CurOffset = Instrs[0].Offset;
3604 for (auto &Instr : Instrs) {
3605 if (CurOffset > Instr.Offset)
3606 return NextI;
3607 CurOffset = Instr.Offset + Instr.Size;
3608 }
3609
3610 // Find contiguous runs of tagged memory and emit shorter instruction
3611 // sequences for them when possible.
3612 TagStoreEdit TSE(MBB, FirstZeroData);
3613 std::optional<int64_t> EndOffset;
3614 for (auto &Instr : Instrs) {
3615 if (EndOffset && *EndOffset != Instr.Offset) {
3616 // Found a gap.
3617 TSE.emitCode(InsertI, TFI, /*TryMergeSPUpdate = */ false);
3618 TSE.clear();
3619 }
3620
3621 TSE.addInstruction(Instr);
3622 EndOffset = Instr.Offset + Instr.Size;
3623 }
3624
3625 const MachineFunction *MF = MBB->getParent();
3626 // Multiple FP/SP updates in a loop cannot be described by CFI instructions.
3627 TSE.emitCode(
3628 InsertI, TFI, /*TryMergeSPUpdate = */
3630
3631 return InsertI;
3632}
3633} // namespace
3634
3636 MachineFunction &MF, RegScavenger *RS = nullptr) const {
3637 for (auto &BB : MF)
3638 for (MachineBasicBlock::iterator II = BB.begin(); II != BB.end();) {
3640 II = tryMergeAdjacentSTG(II, this, RS);
3641 }
3642
3643 // By the time this method is called, most of the prologue/epilogue code is
3644 // already emitted, whether its location was affected by the shrink-wrapping
3645 // optimization or not.
3646 if (!MF.getFunction().hasFnAttribute(Attribute::Naked) &&
3647 shouldSignReturnAddressEverywhere(MF))
3649}
3650
3651/// For Win64 AArch64 EH, the offset to the Unwind object is from the SP
3652/// before the update. This is easily retrieved as it is exactly the offset
3653/// that is set in processFunctionBeforeFrameFinalized.
3655 const MachineFunction &MF, int FI, Register &FrameReg,
3656 bool IgnoreSPUpdates) const {
3657 const MachineFrameInfo &MFI = MF.getFrameInfo();
3658 if (IgnoreSPUpdates) {
3659 LLVM_DEBUG(dbgs() << "Offset from the SP for " << FI << " is "
3660 << MFI.getObjectOffset(FI) << "\n");
3661 FrameReg = AArch64::SP;
3662 return StackOffset::getFixed(MFI.getObjectOffset(FI));
3663 }
3664
3665 // Go to common code if we cannot provide sp + offset.
3666 if (MFI.hasVarSizedObjects() ||
3669 return getFrameIndexReference(MF, FI, FrameReg);
3670
3671 FrameReg = AArch64::SP;
3672 return getStackOffset(MF, MFI.getObjectOffset(FI));
3673}
3674
3675/// The parent frame offset (aka dispFrame) is only used on X86_64 to retrieve
3676/// the parent's frame pointer
3678 const MachineFunction &MF) const {
3679 return 0;
3680}
3681
3682/// Funclets only need to account for space for the callee saved registers,
3683/// as the locals are accounted for in the parent's stack frame.
3685 const MachineFunction &MF) const {
3686 // This is the size of the pushed CSRs.
3687 unsigned CSSize =
3688 MF.getInfo<AArch64FunctionInfo>()->getCalleeSavedStackSize();
3689 // This is the amount of stack a funclet needs to allocate.
3690 return alignTo(CSSize + MF.getFrameInfo().getMaxCallFrameSize(),
3691 getStackAlign());
3692}
3693
3694namespace {
3695struct FrameObject {
3696 bool IsValid = false;
3697 // Index of the object in MFI.
3698 int ObjectIndex = 0;
3699 // Group ID this object belongs to.
3700 int GroupIndex = -1;
3701 // This object should be placed first (closest to SP).
3702 bool ObjectFirst = false;
3703 // This object's group (which always contains the object with
3704 // ObjectFirst==true) should be placed first.
3705 bool GroupFirst = false;
3706
3707 // Used to distinguish between FP and GPR accesses. The values are decided so
3708 // that they sort FPR < Hazard < GPR and they can be or'd together.
3709 unsigned Accesses = 0;
3710 enum { AccessFPR = 1, AccessHazard = 2, AccessGPR = 4 };
3711};
3712
3713class GroupBuilder {
3714 SmallVector<int, 8> CurrentMembers;
3715 int NextGroupIndex = 0;
3716 std::vector<FrameObject> &Objects;
3717
3718public:
3719 GroupBuilder(std::vector<FrameObject> &Objects) : Objects(Objects) {}
3720 void AddMember(int Index) { CurrentMembers.push_back(Index); }
3721 void EndCurrentGroup() {
3722 if (CurrentMembers.size() > 1) {
3723 // Create a new group with the current member list. This might remove them
3724 // from their pre-existing groups. That's OK, dealing with overlapping
3725 // groups is too hard and unlikely to make a difference.
3726 LLVM_DEBUG(dbgs() << "group:");
3727 for (int Index : CurrentMembers) {
3728 Objects[Index].GroupIndex = NextGroupIndex;
3729 LLVM_DEBUG(dbgs() << " " << Index);
3730 }
3731 LLVM_DEBUG(dbgs() << "\n");
3732 NextGroupIndex++;
3733 }
3734 CurrentMembers.clear();
3735 }
3736};
3737
3738bool FrameObjectCompare(const FrameObject &A, const FrameObject &B) {
3739 // Objects at a lower index are closer to FP; objects at a higher index are
3740 // closer to SP.
3741 //
3742 // For consistency in our comparison, all invalid objects are placed
3743 // at the end. This also allows us to stop walking when we hit the
3744 // first invalid item after it's all sorted.
3745 //
3746 // If we want to include a stack hazard region, order FPR accesses < the
3747 // hazard object < GPRs accesses in order to create a separation between the
3748 // two. For the Accesses field 1 = FPR, 2 = Hazard Object, 4 = GPR.
3749 //
3750 // Otherwise the "first" object goes first (closest to SP), followed by the
3751 // members of the "first" group.
3752 //
3753 // The rest are sorted by the group index to keep the groups together.
3754 // Higher numbered groups are more likely to be around longer (i.e. untagged
3755 // in the function epilogue and not at some earlier point). Place them closer
3756 // to SP.
3757 //
3758 // If all else equal, sort by the object index to keep the objects in the
3759 // original order.
3760 return std::make_tuple(!A.IsValid, A.Accesses, A.ObjectFirst, A.GroupFirst,
3761 A.GroupIndex, A.ObjectIndex) <
3762 std::make_tuple(!B.IsValid, B.Accesses, B.ObjectFirst, B.GroupFirst,
3763 B.GroupIndex, B.ObjectIndex);
3764}
3765} // namespace
3766
3768 const MachineFunction &MF, SmallVectorImpl<int> &ObjectsToAllocate) const {
3770
3771 if ((!OrderFrameObjects && !AFI.hasSplitSVEObjects()) ||
3772 ObjectsToAllocate.empty())
3773 return;
3774
3775 const MachineFrameInfo &MFI = MF.getFrameInfo();
3776 std::vector<FrameObject> FrameObjects(MFI.getObjectIndexEnd());
3777 for (auto &Obj : ObjectsToAllocate) {
3778 FrameObjects[Obj].IsValid = true;
3779 FrameObjects[Obj].ObjectIndex = Obj;
3780 }
3781
3782 // Identify FPR vs GPR slots for hazards, and stack slots that are tagged at
3783 // the same time.
3784 GroupBuilder GB(FrameObjects);
3785 for (auto &MBB : MF) {
3786 for (auto &MI : MBB) {
3787 if (MI.isDebugInstr())
3788 continue;
3789
3790 if (AFI.hasStackHazardSlotIndex()) {
3791 std::optional<int> FI = getLdStFrameID(MI, MFI);
3792 if (FI && *FI >= 0 && *FI < (int)FrameObjects.size()) {
3793 if (MFI.getStackID(*FI) == TargetStackID::ScalableVector ||
3795 FrameObjects[*FI].Accesses |= FrameObject::AccessFPR;
3796 else
3797 FrameObjects[*FI].Accesses |= FrameObject::AccessGPR;
3798 }
3799 }
3800
3801 int OpIndex;
3802 switch (MI.getOpcode()) {
3803 case AArch64::STGloop:
3804 case AArch64::STZGloop:
3805 OpIndex = 3;
3806 break;
3807 case AArch64::STGi:
3808 case AArch64::STZGi:
3809 case AArch64::ST2Gi:
3810 case AArch64::STZ2Gi:
3811 OpIndex = 1;
3812 break;
3813 default:
3814 OpIndex = -1;
3815 }
3816
3817 int TaggedFI = -1;
3818 if (OpIndex >= 0) {
3819 const MachineOperand &MO = MI.getOperand(OpIndex);
3820 if (MO.isFI()) {
3821 int FI = MO.getIndex();
3822 if (FI >= 0 && FI < MFI.getObjectIndexEnd() &&
3823 FrameObjects[FI].IsValid)
3824 TaggedFI = FI;
3825 }
3826 }
3827
3828 // If this is a stack tagging instruction for a slot that is not part of a
3829 // group yet, either start a new group or add it to the current one.
3830 if (TaggedFI >= 0)
3831 GB.AddMember(TaggedFI);
3832 else
3833 GB.EndCurrentGroup();
3834 }
3835 // Groups should never span multiple basic blocks.
3836 GB.EndCurrentGroup();
3837 }
3838
3839 if (AFI.hasStackHazardSlotIndex()) {
3840 FrameObjects[AFI.getStackHazardSlotIndex()].Accesses =
3841 FrameObject::AccessHazard;
3842 // If a stack object is unknown or both GPR and FPR, sort it into GPR.
3843 for (auto &Obj : FrameObjects)
3844 if (!Obj.Accesses ||
3845 Obj.Accesses == (FrameObject::AccessGPR | FrameObject::AccessFPR))
3846 Obj.Accesses = FrameObject::AccessGPR;
3847 }
3848
3849 // If the function's tagged base pointer is pinned to a stack slot, we want to
3850 // put that slot first when possible. This will likely place it at SP + 0,
3851 // and save one instruction when generating the base pointer because IRG does
3852 // not allow an immediate offset.
3853 std::optional<int> TBPI = AFI.getTaggedBasePointerIndex();
3854 if (TBPI) {
3855 FrameObjects[*TBPI].ObjectFirst = true;
3856 FrameObjects[*TBPI].GroupFirst = true;
3857 int FirstGroupIndex = FrameObjects[*TBPI].GroupIndex;
3858 if (FirstGroupIndex >= 0)
3859 for (FrameObject &Object : FrameObjects)
3860 if (Object.GroupIndex == FirstGroupIndex)
3861 Object.GroupFirst = true;
3862 }
3863
3864 llvm::stable_sort(FrameObjects, FrameObjectCompare);
3865
3866 int i = 0;
3867 for (auto &Obj : FrameObjects) {
3868 // All invalid items are sorted at the end, so it's safe to stop.
3869 if (!Obj.IsValid)
3870 break;
3871 ObjectsToAllocate[i++] = Obj.ObjectIndex;
3872 }
3873
3874 LLVM_DEBUG({
3875 dbgs() << "Final frame order:\n";
3876 for (auto &Obj : FrameObjects) {
3877 if (!Obj.IsValid)
3878 break;
3879 dbgs() << " " << Obj.ObjectIndex << ": group " << Obj.GroupIndex;
3880 if (Obj.ObjectFirst)
3881 dbgs() << ", first";
3882 if (Obj.GroupFirst)
3883 dbgs() << ", group-first";
3884 dbgs() << "\n";
3885 }
3886 });
3887}
3888
3889/// Emit a loop to decrement SP until it is equal to TargetReg, with probes at
3890/// least every ProbeSize bytes. Returns an iterator of the first instruction
3891/// after the loop. The difference between SP and TargetReg must be an exact
3892/// multiple of ProbeSize.
3894AArch64FrameLowering::inlineStackProbeLoopExactMultiple(
3895 MachineBasicBlock::iterator MBBI, int64_t ProbeSize,
3896 Register TargetReg) const {
3897 MachineBasicBlock &MBB = *MBBI->getParent();
3898 MachineFunction &MF = *MBB.getParent();
3899 const AArch64InstrInfo *TII =
3900 MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3901 DebugLoc DL = MBB.findDebugLoc(MBBI);
3902
3903 MachineFunction::iterator MBBInsertPoint = std::next(MBB.getIterator());
3904 MachineBasicBlock *LoopMBB = MF.CreateMachineBasicBlock(MBB.getBasicBlock());
3905 MF.insert(MBBInsertPoint, LoopMBB);
3906 MachineBasicBlock *ExitMBB = MF.CreateMachineBasicBlock(MBB.getBasicBlock());
3907 MF.insert(MBBInsertPoint, ExitMBB);
3908
3909 // SUB SP, SP, #ProbeSize (or equivalent if ProbeSize is not encodable
3910 // in SUB).
3911 emitFrameOffset(*LoopMBB, LoopMBB->end(), DL, AArch64::SP, AArch64::SP,
3912 StackOffset::getFixed(-ProbeSize), TII,
3914 // LDR XZR, [SP]
3915 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::LDRXui))
3916 .addDef(AArch64::XZR)
3917 .addReg(AArch64::SP)
3918 .addImm(0)
3922 Align(8)))
3924 // CMP SP, TargetReg
3925 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::SUBSXrx64),
3926 AArch64::XZR)
3927 .addReg(AArch64::SP)
3928 .addReg(TargetReg)
3931 // B.CC Loop
3932 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::Bcc))
3934 .addMBB(LoopMBB)
3936
3937 LoopMBB->addSuccessor(ExitMBB);
3938 LoopMBB->addSuccessor(LoopMBB);
3939 // Synthesize the exit MBB.
3940 ExitMBB->splice(ExitMBB->end(), &MBB, MBBI, MBB.end());
3942 MBB.addSuccessor(LoopMBB);
3943 // Update liveins.
3944 fullyRecomputeLiveIns({ExitMBB, LoopMBB});
3945
3946 return ExitMBB->begin();
3947}
3948
3949void AArch64FrameLowering::inlineStackProbeFixed(
3950 MachineBasicBlock::iterator MBBI, Register ScratchReg, int64_t FrameSize,
3951 StackOffset CFAOffset) const {
3952 MachineBasicBlock *MBB = MBBI->getParent();
3953 MachineFunction &MF = *MBB->getParent();
3954 const AArch64InstrInfo *TII =
3955 MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3956 AArch64FunctionInfo *AFI = MF.getInfo<AArch64FunctionInfo>();
3957 bool EmitAsyncCFI = AFI->needsAsyncDwarfUnwindInfo(MF);
3958 bool HasFP = hasFP(MF);
3959
3960 DebugLoc DL;
3961 int64_t ProbeSize = MF.getInfo<AArch64FunctionInfo>()->getStackProbeSize();
3962 int64_t NumBlocks = FrameSize / ProbeSize;
3963 int64_t ResidualSize = FrameSize % ProbeSize;
3964
3965 LLVM_DEBUG(dbgs() << "Stack probing: total " << FrameSize << " bytes, "
3966 << NumBlocks << " blocks of " << ProbeSize
3967 << " bytes, plus " << ResidualSize << " bytes\n");
3968
3969 // Decrement SP by NumBlock * ProbeSize bytes, with either unrolled or
3970 // ordinary loop.
3971 if (NumBlocks <= AArch64::StackProbeMaxLoopUnroll) {
3972 for (int i = 0; i < NumBlocks; ++i) {
3973 // SUB SP, SP, #ProbeSize (or equivalent if ProbeSize is not
3974 // encodable in a SUB).
3975 emitFrameOffset(*MBB, MBBI, DL, AArch64::SP, AArch64::SP,
3976 StackOffset::getFixed(-ProbeSize), TII,
3977 MachineInstr::FrameSetup, false, false, nullptr,
3978 EmitAsyncCFI && !HasFP, CFAOffset);
3979 CFAOffset += StackOffset::getFixed(ProbeSize);
3980 // LDR XZR, [SP]
3981 BuildMI(*MBB, MBBI, DL, TII->get(AArch64::LDRXui))
3982 .addDef(AArch64::XZR)
3983 .addReg(AArch64::SP)
3984 .addImm(0)
3988 Align(8)))
3990 }
3991 } else if (NumBlocks != 0) {
3992 // SUB ScratchReg, SP, #FrameSize (or equivalent if FrameSize is not
3993 // encodable in ADD). ScrathReg may temporarily become the CFA register.
3994 emitFrameOffset(*MBB, MBBI, DL, ScratchReg, AArch64::SP,
3995 StackOffset::getFixed(-ProbeSize * NumBlocks), TII,
3996 MachineInstr::FrameSetup, false, false, nullptr,
3997 EmitAsyncCFI && !HasFP, CFAOffset);
3998 CFAOffset += StackOffset::getFixed(ProbeSize * NumBlocks);
3999 MBBI = inlineStackProbeLoopExactMultiple(MBBI, ProbeSize, ScratchReg);
4000 MBB = MBBI->getParent();
4001 if (EmitAsyncCFI && !HasFP) {
4002 // Set the CFA register back to SP.
4003 CFIInstBuilder(*MBB, MBBI, MachineInstr::FrameSetup)
4004 .buildDefCFARegister(AArch64::SP);
4005 }
4006 }
4007
4008 if (ResidualSize != 0) {
4009 // SUB SP, SP, #ResidualSize (or equivalent if ResidualSize is not encodable
4010 // in SUB).
4011 emitFrameOffset(*MBB, MBBI, DL, AArch64::SP, AArch64::SP,
4012 StackOffset::getFixed(-ResidualSize), TII,
4013 MachineInstr::FrameSetup, false, false, nullptr,
4014 EmitAsyncCFI && !HasFP, CFAOffset);
4015 if (ResidualSize > AArch64::StackProbeMaxUnprobedStack) {
4016 // LDR XZR, [SP]
4017 BuildMI(*MBB, MBBI, DL, TII->get(AArch64::LDRXui))
4018 .addDef(AArch64::XZR)
4019 .addReg(AArch64::SP)
4020 .addImm(0)
4024 Align(8)))
4026 }
4027 }
4028}
4029
4030void AArch64FrameLowering::inlineStackProbe(MachineFunction &MF,
4031 MachineBasicBlock &MBB) const {
4032 // Get the instructions that need to be replaced. We emit at most two of
4033 // these. Remember them in order to avoid complications coming from the need
4034 // to traverse the block while potentially creating more blocks.
4035 SmallVector<MachineInstr *, 4> ToReplace;
4036 for (MachineInstr &MI : MBB)
4037 if (MI.getOpcode() == AArch64::PROBED_STACKALLOC ||
4038 MI.getOpcode() == AArch64::PROBED_STACKALLOC_VAR)
4039 ToReplace.push_back(&MI);
4040
4041 for (MachineInstr *MI : ToReplace) {
4042 if (MI->getOpcode() == AArch64::PROBED_STACKALLOC) {
4043 Register ScratchReg = MI->getOperand(0).getReg();
4044 int64_t FrameSize = MI->getOperand(1).getImm();
4045 StackOffset CFAOffset = StackOffset::get(MI->getOperand(2).getImm(),
4046 MI->getOperand(3).getImm());
4047 inlineStackProbeFixed(MI->getIterator(), ScratchReg, FrameSize,
4048 CFAOffset);
4049 } else {
4050 assert(MI->getOpcode() == AArch64::PROBED_STACKALLOC_VAR &&
4051 "Stack probe pseudo-instruction expected");
4052 const AArch64InstrInfo *TII =
4053 MI->getMF()->getSubtarget<AArch64Subtarget>().getInstrInfo();
4054 Register TargetReg = MI->getOperand(0).getReg();
4055 (void)TII->probedStackAlloc(MI->getIterator(), TargetReg, true);
4056 }
4057 MI->eraseFromParent();
4058 }
4059}
4060
4063 NotAccessed = 0, // Stack object not accessed by load/store instructions.
4064 GPR = 1 << 0, // A general purpose register.
4065 PPR = 1 << 1, // A predicate register.
4066 FPR = 1 << 2, // A floating point/Neon/SVE register.
4067 };
4068
4069 int Idx;
4071 int64_t Size;
4072 unsigned AccessTypes;
4073
4075
4076 bool operator<(const StackAccess &Rhs) const {
4077 return std::make_tuple(start(), Idx) <
4078 std::make_tuple(Rhs.start(), Rhs.Idx);
4079 }
4080
4081 bool isCPU() const {
4082 // Predicate register load and store instructions execute on the CPU.
4084 }
4085 bool isSME() const { return AccessTypes & AccessType::FPR; }
4086 bool isMixed() const { return isCPU() && isSME(); }
4087
4088 int64_t start() const { return Offset.getFixed() + Offset.getScalable(); }
4089 int64_t end() const { return start() + Size; }
4090
4091 std::string getTypeString() const {
4092 switch (AccessTypes) {
4093 case AccessType::FPR:
4094 return "FPR";
4095 case AccessType::PPR:
4096 return "PPR";
4097 case AccessType::GPR:
4098 return "GPR";
4100 return "NA";
4101 default:
4102 return "Mixed";
4103 }
4104 }
4105
4106 void print(raw_ostream &OS) const {
4107 OS << getTypeString() << " stack object at [SP"
4108 << (Offset.getFixed() < 0 ? "" : "+") << Offset.getFixed();
4109 if (Offset.getScalable())
4110 OS << (Offset.getScalable() < 0 ? "" : "+") << Offset.getScalable()
4111 << " * vscale";
4112 OS << "]";
4113 }
4114};
4115
4116static inline raw_ostream &operator<<(raw_ostream &OS, const StackAccess &SA) {
4117 SA.print(OS);
4118 return OS;
4119}
4120
4121void AArch64FrameLowering::emitRemarks(
4122 const MachineFunction &MF, MachineOptimizationRemarkEmitter *ORE) const {
4123
4124 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
4126 return;
4127
4128 unsigned StackHazardSize = getStackHazardSize(MF);
4129 const uint64_t HazardSize =
4130 (StackHazardSize) ? StackHazardSize : StackHazardRemarkSize;
4131
4132 if (HazardSize == 0)
4133 return;
4134
4135 const MachineFrameInfo &MFI = MF.getFrameInfo();
4136 // Bail if function has no stack objects.
4137 if (!MFI.hasStackObjects())
4138 return;
4139
4140 std::vector<StackAccess> StackAccesses(MFI.getNumObjects());
4141
4142 size_t NumFPLdSt = 0;
4143 size_t NumNonFPLdSt = 0;
4144
4145 // Collect stack accesses via Load/Store instructions.
4146 for (const MachineBasicBlock &MBB : MF) {
4147 for (const MachineInstr &MI : MBB) {
4148 if (!MI.mayLoadOrStore() || MI.getNumMemOperands() < 1)
4149 continue;
4150 for (MachineMemOperand *MMO : MI.memoperands()) {
4151 std::optional<int> FI = getMMOFrameID(MMO, MFI);
4152 if (FI && !MFI.isDeadObjectIndex(*FI)) {
4153 int FrameIdx = *FI;
4154
4155 size_t ArrIdx = FrameIdx + MFI.getNumFixedObjects();
4156 if (StackAccesses[ArrIdx].AccessTypes == StackAccess::NotAccessed) {
4157 StackAccesses[ArrIdx].Idx = FrameIdx;
4158 StackAccesses[ArrIdx].Offset =
4159 getFrameIndexReferenceFromSP(MF, FrameIdx);
4160 StackAccesses[ArrIdx].Size = MFI.getObjectSize(FrameIdx);
4161 }
4162
4163 unsigned RegTy = StackAccess::AccessType::GPR;
4164 if (MFI.hasScalableStackID(FrameIdx))
4167 RegTy = StackAccess::FPR;
4168
4169 StackAccesses[ArrIdx].AccessTypes |= RegTy;
4170
4171 if (RegTy == StackAccess::FPR)
4172 ++NumFPLdSt;
4173 else
4174 ++NumNonFPLdSt;
4175 }
4176 }
4177 }
4178 }
4179
4180 if (NumFPLdSt == 0 || NumNonFPLdSt == 0)
4181 return;
4182
4183 llvm::sort(StackAccesses);
4184 llvm::erase_if(StackAccesses, [](const StackAccess &S) {
4186 });
4187
4190
4191 if (StackAccesses.front().isMixed())
4192 MixedObjects.push_back(&StackAccesses.front());
4193
4194 for (auto It = StackAccesses.begin(), End = std::prev(StackAccesses.end());
4195 It != End; ++It) {
4196 const auto &First = *It;
4197 const auto &Second = *(It + 1);
4198
4199 if (Second.isMixed())
4200 MixedObjects.push_back(&Second);
4201
4202 if ((First.isSME() && Second.isCPU()) ||
4203 (First.isCPU() && Second.isSME())) {
4204 uint64_t Distance = static_cast<uint64_t>(Second.start() - First.end());
4205 if (Distance < HazardSize)
4206 HazardPairs.emplace_back(&First, &Second);
4207 }
4208 }
4209
4210 auto EmitRemark = [&](llvm::StringRef Str) {
4211 ORE->emit([&]() {
4212 auto R = MachineOptimizationRemarkAnalysis(
4213 "sme", "StackHazard", MF.getFunction().getSubprogram(), &MF.front());
4214 return R << formatv("stack hazard in '{0}': ", MF.getName()).str() << Str;
4215 });
4216 };
4217
4218 for (const auto &P : HazardPairs)
4219 EmitRemark(formatv("{0} is too close to {1}", *P.first, *P.second).str());
4220
4221 for (const auto *Obj : MixedObjects)
4222 EmitRemark(
4223 formatv("{0} accessed by both GP and FP instructions", *Obj).str());
4224}
static void getLiveRegsForEntryMBB(LivePhysRegs &LiveRegs, const MachineBasicBlock &MBB)
static const unsigned DefaultSafeSPDisplacement
This is the biggest offset to the stack pointer we can encode in aarch64 instructions (without using ...
static void orderZPRCalleeSavesForPairs(MachineFunction &MF, const TargetRegisterInfo *RegInfo, std::vector< CalleeSavedInfo > &CSI)
static RegState getPrologueDeath(MachineFunction &MF, unsigned Reg)
static bool produceCompactUnwindFrame(const AArch64FrameLowering &, MachineFunction &MF)
static cl::opt< bool > StackTaggingMergeSetTag("stack-tagging-merge-settag", cl::desc("merge settag instruction in function epilog"), cl::init(true), cl::Hidden)
bool enableMultiVectorSpillFill(const AArch64Subtarget &Subtarget, MachineFunction &MF)
static std::optional< int > getLdStFrameID(const MachineInstr &MI, const MachineFrameInfo &MFI)
static cl::opt< bool > SplitSVEObjects("aarch64-split-sve-objects", cl::desc("Split allocation of ZPR & PPR objects"), cl::init(true), cl::Hidden)
static cl::opt< bool > StackHazardInNonStreaming("aarch64-stack-hazard-in-non-streaming", cl::init(false), cl::Hidden)
void computeCalleeSaveRegisterPairs(const AArch64FrameLowering &AFL, MachineFunction &MF, ArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI, SmallVectorImpl< RegPairInfo > &RegPairs, bool NeedsFrameRecord)
static cl::opt< bool > OrderFrameObjects("aarch64-order-frame-objects", cl::desc("sort stack allocations"), cl::init(true), cl::Hidden)
static cl::opt< bool > DisableMultiVectorSpillFill("aarch64-disable-multivector-spill-fill", cl::desc("Disable use of LD/ST pairs for SME2 or SVE2p1"), cl::init(false), cl::Hidden)
static cl::opt< bool > EnableRedZone("aarch64-redzone", cl::desc("enable use of redzone on AArch64"), cl::init(false), cl::Hidden)
static bool invalidateRegisterPairing(bool SpillExtendedVolatile, unsigned SpillCount, unsigned Reg1, unsigned Reg2, bool UsesWinAAPCS, bool NeedsWinCFI, bool NeedsFrameRecord, const TargetRegisterInfo *TRI)
Returns true if Reg1 and Reg2 cannot be paired using a ldp/stp instruction.
cl::opt< bool > EnableHomogeneousPrologEpilog("homogeneous-prolog-epilog", cl::Hidden, cl::desc("Emit homogeneous prologue and epilogue for the size " "optimization (default = off)"))
static bool isLikelyToHaveSVEStack(const AArch64FrameLowering &AFL, const MachineFunction &MF)
static bool invalidateWindowsRegisterPairing(bool SpillExtendedVolatile, unsigned SpillCount, unsigned Reg1, unsigned Reg2, bool NeedsWinCFI, const TargetRegisterInfo *TRI)
static SVEStackSizes determineSVEStackSizes(MachineFunction &MF, AssignObjectOffsets AssignOffsets)
Process all the SVE stack objects and the SVE stack size and offsets for each object.
static bool isTargetWindows(const MachineFunction &MF)
static unsigned estimateRSStackSizeLimit(MachineFunction &MF)
Look at each instruction that references stack frames and return the stack size limit beyond which so...
static bool getSVECalleeSaveSlotRange(const MachineFrameInfo &MFI, int &Min, int &Max)
returns true if there are any SVE callee saves.
static cl::opt< unsigned > StackHazardRemarkSize("aarch64-stack-hazard-remark-size", cl::init(0), cl::Hidden)
static MCRegister getRegisterOrZero(MCRegister Reg, bool HasSVE)
static unsigned getStackHazardSize(const MachineFunction &MF)
MCRegister findFreePredicateReg(BitVector &SavedRegs)
static bool isPPRAccess(const MachineInstr &MI)
static std::optional< int > getMMOFrameID(MachineMemOperand *MMO, const MachineFrameInfo &MFI)
assert(UImm &&(UImm !=~static_cast< T >(0)) &&"Invalid immediate!")
This file contains the declaration of the AArch64PrologueEmitter and AArch64EpilogueEmitter classes,...
static const int kSetTagLoopThreshold
unsigned Imm
unsigned uint64_t
MachineBasicBlock & MBB
MachineBasicBlock MachineBasicBlock::iterator DebugLoc DL
MachineBasicBlock MachineBasicBlock::iterator MBBI
This file contains the simple types necessary to represent the attributes associated with functions a...
#define CASE(ATTRNAME, AANAME,...)
static GCRegistry::Add< ErlangGC > A("erlang", "erlang-compatible garbage collector")
static GCRegistry::Add< CoreCLRGC > E("coreclr", "CoreCLR-compatible GC")
static GCRegistry::Add< OcamlGC > B("ocaml", "ocaml 3.10-compatible GC")
DXIL Forward Handle Accesses
const HexagonInstrInfo * TII
IRTranslator LLVM IR MI
static std::string getTypeString(Type *T)
Definition LLParser.cpp:68
This file implements the LivePhysRegs utility for tracking liveness of physical registers.
#define F(x, y, z)
Definition MD5.cpp:54
#define I(x, y, z)
Definition MD5.cpp:57
#define H(x, y, z)
Definition MD5.cpp:56
Register Reg
Register const TargetRegisterInfo * TRI
Promote Memory to Register
Definition Mem2Reg.cpp:110
uint64_t IntrinsicInst * II
#define P(N)
This file declares the machine register scavenger class.
static bool contains(SmallPtrSetImpl< ConstantExpr * > &Cache, ConstantExpr *Expr, Constant *C)
Definition Value.cpp:484
This file defines the scope_exit class, which executes user-defined cleanup logic at scope exit.
This file defines the SmallVector class.
#define LLVM_DEBUG(...)
Definition Debug.h:119
StackOffset getSVEStackSize(const MachineFunction &MF) const
Returns the size of the entire SVE stackframe (PPRs + ZPRs).
StackOffset getZPRStackSize(const MachineFunction &MF) const
Returns the size of the entire ZPR stackframe (calleesaves + spills).
void processFunctionBeforeFrameIndicesReplaced(MachineFunction &MF, RegScavenger *RS) const override
processFunctionBeforeFrameIndicesReplaced - This method is called immediately before MO_FrameIndex op...
MachineBasicBlock::iterator eliminateCallFramePseudoInstr(MachineFunction &MF, MachineBasicBlock &MBB, MachineBasicBlock::iterator I) const override
This method is called during prolog/epilog code insertion to eliminate call frame setup and destroy p...
bool canUseAsPrologue(const MachineBasicBlock &MBB) const override
Check whether or not the given MBB can be used as a prologue for the target.
bool enableStackSlotScavenging(const MachineFunction &MF) const override
Returns true if the stack slot holes in the fixed and callee-save stack area should be used when allo...
bool assignCalleeSavedSpillSlots(MachineFunction &MF, const TargetRegisterInfo *TRI, std::vector< CalleeSavedInfo > &CSI) const override
assignCalleeSavedSpillSlots - Allows target to override spill slot assignment logic.
bool spillCalleeSavedRegisters(MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, ArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI) const override
spillCalleeSavedRegisters - Issues instruction(s) to spill all callee saved registers and returns tru...
bool restoreCalleeSavedRegisters(MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, MutableArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI) const override
restoreCalleeSavedRegisters - Issues instruction(s) to restore all callee saved registers and returns...
bool enableFullCFIFixup(const MachineFunction &MF) const override
enableFullCFIFixup - Returns true if we may need to fix the unwind information such that it is accura...
StackOffset getFrameIndexReferenceFromSP(const MachineFunction &MF, int FI) const override
getFrameIndexReferenceFromSP - This method returns the offset from the stack pointer to the slot of t...
bool enableCFIFixup(const MachineFunction &MF) const override
Returns true if we may need to fix the unwind information for the function.
StackOffset getNonLocalFrameIndexReference(const MachineFunction &MF, int FI) const override
getNonLocalFrameIndexReference - This method returns the offset used to reference a frame index locat...
TargetStackID::Value getStackIDForScalableVectors() const override
Returns the StackID that scalable vectors should be associated with.
bool hasFPImpl(const MachineFunction &MF) const override
hasFPImpl - Return true if the specified function should have a dedicated frame pointer register.
void emitPrologue(MachineFunction &MF, MachineBasicBlock &MBB) const override
emitProlog/emitEpilog - These methods insert prolog and epilog code into the function.
void resetCFIToInitialState(MachineBasicBlock &MBB) const override
Emit CFI instructions that recreate the state of the unwind information upon function entry.
bool hasReservedCallFrame(const MachineFunction &MF) const override
hasReservedCallFrame - Under normal circumstances, when a frame pointer is not required,...
bool hasSVECalleeSavesAboveFrameRecord(const MachineFunction &MF) const
StackOffset resolveFrameOffsetReference(const MachineFunction &MF, int64_t ObjectOffset, bool isFixed, TargetStackID::Value StackID, Register &FrameReg, bool PreferFP, bool ForSimm) const
bool canUseRedZone(const MachineFunction &MF) const
Can this function use the red zone for local allocations.
bool needsWinCFI(const MachineFunction &MF) const
bool isFPReserved(const MachineFunction &MF) const
Should the Frame Pointer be reserved for the current function?
void processFunctionBeforeFrameFinalized(MachineFunction &MF, RegScavenger *RS) const override
processFunctionBeforeFrameFinalized - This method is called immediately before the specified function...
int getSEHFrameIndexOffset(const MachineFunction &MF, int FI) const
unsigned getWinEHFuncletFrameSize(const MachineFunction &MF) const
Funclets only need to account for space for the callee saved registers, as the locals are accounted f...
void orderFrameObjects(const MachineFunction &MF, SmallVectorImpl< int > &ObjectsToAllocate) const override
Order the symbols in the local stack frame.
void emitEpilogue(MachineFunction &MF, MachineBasicBlock &MBB) const override
StackOffset getPPRStackSize(const MachineFunction &MF) const
Returns the size of the entire PPR stackframe (calleesaves + spills + hazard padding).
int64_t getArgumentStackToRestore(MachineFunction &MF, MachineBasicBlock &MBB) const
Returns how much of the incoming argument stack area (in bytes) we should clean up in an epilogue.
void determineCalleeSaves(MachineFunction &MF, BitVector &SavedRegs, RegScavenger *RS) const override
This method determines which of the registers reported by TargetRegisterInfo::getCalleeSavedRegs() sh...
StackOffset getFrameIndexReference(const MachineFunction &MF, int FI, Register &FrameReg) const override
getFrameIndexReference - Provide a base+offset reference to an FI slot for debug info.
StackOffset getFrameIndexReferencePreferSP(const MachineFunction &MF, int FI, Register &FrameReg, bool IgnoreSPUpdates) const override
For Win64 AArch64 EH, the offset to the Unwind object is from the SP before the update.
StackOffset resolveFrameIndexReference(const MachineFunction &MF, int FI, Register &FrameReg, bool PreferFP, bool ForSimm) const
unsigned getWinEHParentFrameOffset(const MachineFunction &MF) const override
The parent frame offset (aka dispFrame) is only used on X86_64 to retrieve the parent's frame pointer...
bool requiresSaveVG(const MachineFunction &MF) const
void emitPacRetPlusLeafHardening(MachineFunction &MF) const
Harden the entire function with pac-ret.
AArch64FunctionInfo - This class is derived from MachineFunctionInfo and contains private AArch64-spe...
unsigned getCalleeSavedStackSize(const MachineFrameInfo &MFI) const
void setCalleeSaveBaseToFrameRecordOffset(int Offset)
SignReturnAddress getSignReturnAddressCondition() const
void setStackSizeSVE(uint64_t ZPR, uint64_t PPR)
std::optional< int > getTaggedBasePointerIndex() const
bool needsDwarfUnwindInfo(const MachineFunction &MF) const
void setSVECalleeSavedStackSize(unsigned ZPR, unsigned PPR)
bool needsAsyncDwarfUnwindInfo(const MachineFunction &MF) const
static bool isTailCallReturnInst(const MachineInstr &MI)
Returns true if MI is one of the TCRETURN* instructions.
static bool isFpOrNEON(Register Reg)
Returns whether the physical register is FP or NEON.
const AArch64RegisterInfo * getRegisterInfo() const override
bool isNeonAvailable() const
Returns true if the target has NEON and the function at runtime is known to have NEON enabled (e....
const AArch64InstrInfo * getInstrInfo() const override
const AArch64TargetLowering * getTargetLowering() const override
bool isSVEorStreamingSVEAvailable() const
Returns true if the target has access to either the full range of SVE instructions,...
bool isStreaming() const
Returns true if the function has a streaming body.
bool hasInlineStackProbe(const MachineFunction &MF) const override
True if stack clash protection is enabled for this functions.
unsigned getRedZoneSize(const Function &F) const
Represent a constant reference to an array (0 or more elements consecutively in memory),...
Definition ArrayRef.h:40
size_t size() const
Get the array size.
Definition ArrayRef.h:141
bool empty() const
Check if the array is empty.
Definition ArrayRef.h:136
bool test(unsigned Idx) const
Returns true if bit Idx is set.
Definition BitVector.h:482
BitVector & reset()
Reset all bits in the bitvector.
Definition BitVector.h:409
size_type count() const
Returns the number of bits which are set.
Definition BitVector.h:181
BitVector & set()
Set all bits in the bitvector.
Definition BitVector.h:366
iterator_range< const_set_bits_iterator > set_bits() const
Definition BitVector.h:159
size_type size() const
Returns the number of bits in this bitvector.
Definition BitVector.h:178
Helper class for creating CFI instructions and inserting them into MIR.
The CalleeSavedInfo class tracks the information need to locate where a callee saved register is in t...
A debug info location.
Definition DebugLoc.h:126
bool hasMinSize() const
Optimize this function for minimum size (-Oz).
Definition Function.h:696
CallingConv::ID getCallingConv() const
getCallingConv()/setCallingConv(CC) - These method get and set the calling convention of this functio...
Definition Function.h:273
AttributeList getAttributes() const
Return the attribute list for this Function.
Definition Function.h:329
bool isVarArg() const
isVarArg - Return true if this function takes a variable number of arguments.
Definition Function.h:230
bool hasFnAttribute(Attribute::AttrKind Kind) const
Return true if the function has the attribute.
Definition Function.cpp:730
A set of physical registers with utility functions to track liveness when walking backward/forward th...
bool usesWindowsCFI() const
Definition MCAsmInfo.h:675
Wrapper class representing physical registers. Should be passed by value.
Definition MCRegister.h:41
LLVM_ABI void transferSuccessorsAndUpdatePHIs(MachineBasicBlock *FromMBB)
Transfers all the successors, as in transferSuccessors, and update PHI operands in the successor bloc...
LLVM_ABI iterator getFirstTerminator()
Returns an iterator to the first terminator instruction of this basic block.
LLVM_ABI void addSuccessor(MachineBasicBlock *Succ, BranchProbability Prob=BranchProbability::getUnknown())
Add Succ as a successor of this MachineBasicBlock.
const MachineFunction * getParent() const
Return the MachineFunction containing this basic block.
reverse_iterator rbegin()
iterator insertAfter(iterator I, MachineInstr *MI)
Insert MI into the instruction list after I.
void splice(iterator Where, MachineBasicBlock *Other, iterator From)
Take an instruction from MBB 'Other' at the position From, and insert it into this MBB right before '...
MachineInstrBundleIterator< MachineInstr > iterator
The MachineFrameInfo class represents an abstract stack frame until prolog/epilog code is inserted.
LLVM_ABI int CreateFixedObject(uint64_t Size, int64_t SPOffset, bool IsImmutable, bool isAliased=false)
Create a new object at a fixed location on the stack.
bool hasVarSizedObjects() const
This method may be called any time after instruction selection is complete to determine if the stack ...
const AllocaInst * getObjectAllocation(int ObjectIdx) const
Return the underlying Alloca of the specified stack object if it exists.
LLVM_ABI int CreateStackObject(uint64_t Size, Align Alignment, bool isSpillSlot, const AllocaInst *Alloca=nullptr, uint8_t ID=0)
Create a new statically sized stack object, returning a nonnegative identifier to represent it.
bool hasCalls() const
Return true if the current function has any function calls.
bool isFrameAddressTaken() const
This method may be called any time after instruction selection is complete to determine if there is a...
void setObjectOffset(int ObjectIdx, int64_t SPOffset)
Set the stack frame offset of the specified object.
bool isCalleeSavedObjectIndex(int ObjectIdx) const
uint64_t getMaxCallFrameSize() const
Return the maximum size of a call frame that must be allocated for an outgoing function call.
bool hasPatchPoint() const
This method may be called any time after instruction selection is complete to determine if there is a...
bool hasScalableStackID(int ObjectIdx) const
int getStackProtectorIndex() const
Return the index for the stack protector object.
LLVM_ABI uint64_t estimateStackSize(const MachineFunction &MF) const
Estimate and return the size of the stack frame.
void setStackID(int ObjectIdx, uint8_t ID)
bool isCalleeSavedInfoValid() const
Has the callee saved info been calculated yet?
Align getObjectAlign(int ObjectIdx) const
Return the alignment of the specified stack object.
int64_t getObjectSize(int ObjectIdx) const
Return the size of the specified object.
bool isMaxCallFrameSizeComputed() const
bool hasStackMap() const
This method may be called any time after instruction selection is complete to determine if there is a...
LLVM_ABI int CreateSpillStackObject(uint64_t Size, Align Alignment, TargetStackID::Value StackID=TargetStackID::Default)
Create a new statically sized stack object that represents a spill slot, returning a nonnegative iden...
const std::vector< CalleeSavedInfo > & getCalleeSavedInfo() const
Returns a reference to call saved info vector for the current function.
unsigned getNumObjects() const
Return the number of objects.
int getObjectIndexEnd() const
Return one past the maximum frame object index.
bool hasStackProtectorIndex() const
bool hasStackObjects() const
Return true if there are any stack objects in this function.
uint8_t getStackID(int ObjectIdx) const
unsigned getNumFixedObjects() const
Return the number of fixed objects.
void setIsCalleeSavedObjectIndex(int ObjectIdx, bool IsCalleeSaved)
int64_t getObjectOffset(int ObjectIdx) const
Return the assigned stack offset of the specified object from the incoming stack pointer.
int getObjectIndexBegin() const
Return the minimum frame object index.
void setObjectAlignment(int ObjectIdx, Align Alignment)
setObjectAlignment - Change the alignment of the specified stack object.
bool isDeadObjectIndex(int ObjectIdx) const
Returns true if the specified index corresponds to a dead object.
const WinEHFuncInfo * getWinEHFuncInfo() const
getWinEHFuncInfo - Return information about how the current function uses Windows exception handling.
const TargetSubtargetInfo & getSubtarget() const
getSubtarget - Return the subtarget for which this machine code is being compiled.
LLVM_ABI bool framePointerIsReserved() const
Returns true if the frame pointer must always either point to a new frame record or be un-modified in...
MachineFrameInfo & getFrameInfo()
getFrameInfo - Return the frame info object for the current function.
MachineRegisterInfo & getRegInfo()
getRegInfo - Return information about the registers currently in use.
Function & getFunction()
Return the LLVM function that this machine code represents.
BasicBlockListType::iterator iterator
LLVM_ABI bool disableFramePointerElim() const
Returns true if frame pointer elimination should be disabled for this function.
Ty * getInfo()
getInfo - Keep track of various per-function pieces of information for backends that would like to do...
const MachineBasicBlock & front() const
MachineMemOperand * getMachineMemOperand(MachinePointerInfo PtrInfo, MachineMemOperand::Flags F, LLT MemTy, Align BaseAlignment, const MMOMetadata &Metadata=MMOMetadata(), SyncScope::ID SSID=SyncScope::System, AtomicOrdering Ordering=AtomicOrdering::NotAtomic, AtomicOrdering FailureOrdering=AtomicOrdering::NotAtomic)
getMachineMemOperand - Allocate a new MachineMemOperand.
MachineBasicBlock * CreateMachineBasicBlock(const BasicBlock *BB=nullptr, std::optional< UniqueBBID > BBID=std::nullopt)
CreateMachineInstr - Allocate a new MachineInstr.
void insert(iterator MBBI, MachineBasicBlock *MBB)
const TargetMachine & getTarget() const
getTarget - Return the target machine this machine code is compiled with
const MachineInstrBuilder & setMemRefs(ArrayRef< MachineMemOperand * > MMOs) const
const MachineInstrBuilder & addExternalSymbol(const char *FnName, unsigned TargetFlags=0) const
const MachineInstrBuilder & addReg(Register RegNo, RegState Flags={}, unsigned SubReg=0) const
Add a new virtual register operand.
const MachineInstrBuilder & setMIFlag(MachineInstr::MIFlag Flag) const
const MachineInstrBuilder & addImm(int64_t Val) const
Add a new immediate operand.
const MachineInstrBuilder & addFrameIndex(int Idx) const
const MachineInstrBuilder & addRegMask(const uint32_t *Mask) const
const MachineInstrBuilder & addMBB(MachineBasicBlock *MBB, unsigned TargetFlags=0) const
const MachineInstrBuilder & addDef(Register RegNo, RegState Flags={}, unsigned SubReg=0) const
Add a virtual register definition operand.
const MachineInstrBuilder & setMIFlags(unsigned Flags) const
const MachineInstrBuilder & addMemOperand(MachineMemOperand *MMO) const
Representation of each machine instruction.
void setFlags(unsigned flags)
uint32_t getFlags() const
Return the MI flags bitvector.
LLVM_ABI MachineInstrBundleIterator< MachineInstr > eraseFromParent()
Unlink 'this' from the containing basic block and delete it.
A description of a memory reference used in the backend.
const PseudoSourceValue * getPseudoValue() const
@ MOVolatile
The memory access is volatile.
@ MOLoad
The memory access reads data.
@ MOStore
The memory access writes data.
const Value * getValue() const
Return the base address of the memory access.
MachineOperand class - Representation of each machine instruction operand.
int64_t getImm() const
bool isFI() const
isFI - Tests if this is a MO_FrameIndex operand.
LLVM_ABI void emit(DiagnosticInfoOptimizationBase &OptDiag)
Emit an optimization remark.
MachineRegisterInfo - Keep track of information for virtual and physical registers,...
LLVM_ABI void freezeReservedRegs()
freezeReservedRegs - Called by the register allocator to freeze the set of reserved registers before ...
bool isReserved(MCRegister PhysReg) const
isReserved - Returns true when PhysReg is a reserved register.
LLVM_ABI Register createVirtualRegister(const TargetRegisterClass *RegClass, StringRef Name="")
createVirtualRegister - Create and return a new virtual register in the function with the specified r...
LLVM_ABI bool isLiveIn(Register Reg) const
LLVM_ABI const MCPhysReg * getCalleeSavedRegs() const
Returns list of callee saved registers.
LLVM_ABI bool isPhysRegUsed(MCRegister PhysReg, bool SkipRegMaskTest=false) const
Return true if the specified register is modified or read in this function.
Represent a mutable reference to an array (0 or more elements consecutively in memory),...
Definition ArrayRef.h:294
Wrapper class representing virtual and physical registers.
Definition Register.h:20
constexpr bool isValid() const
Definition Register.h:112
SMEAttrs is a utility class to parse the SME ACLE attributes on functions.
bool hasStreamingInterface() const
bool hasNonStreamingInterfaceAndBody() const
bool hasStreamingBody() const
bool insert(const value_type &X)
Insert a new element into the SetVector.
Definition SetVector.h:157
A SetVector that performs no allocations if smaller than a certain size.
Definition SetVector.h:345
This class consists of common code factored out of the SmallVector class to reduce code duplication b...
reference emplace_back(ArgTypes &&... Args)
void append(ItTy in_start, ItTy in_end)
Add the specified range to the end of the SmallVector.
void push_back(const T &Elt)
This is a 'vector' (really, a variable-sized array), optimized for the case when the array is small.
StackOffset holds a fixed and a scalable offset in bytes.
Definition TypeSize.h:30
int64_t getFixed() const
Returns the fixed component of the stack.
Definition TypeSize.h:46
int64_t getScalable() const
Returns the scalable component of the stack.
Definition TypeSize.h:49
static StackOffset get(int64_t Fixed, int64_t Scalable)
Definition TypeSize.h:41
static StackOffset getScalable(int64_t Scalable)
Definition TypeSize.h:40
static StackOffset getFixed(int64_t Fixed)
Definition TypeSize.h:39
bool hasFP(const MachineFunction &MF) const
hasFP - Return true if the specified function should have a dedicated frame pointer register.
virtual void determineCalleeSaves(MachineFunction &MF, BitVector &SavedRegs, RegScavenger *RS=nullptr) const
This method determines which of the registers reported by TargetRegisterInfo::getCalleeSavedRegs() sh...
int getOffsetOfLocalArea() const
getOffsetOfLocalArea - This method returns the offset of the local area from the stack pointer on ent...
Align getStackAlign() const
getStackAlignment - This method returns the number of bytes to which the stack pointer must be aligne...
StackDirection getStackGrowthDirection() const
getStackGrowthDirection - Return the direction the stack grows
virtual bool enableCFIFixup(const MachineFunction &MF) const
Returns true if we may need to fix the unwind information for the function.
Primary interface to the complete machine description for the target machine.
const Triple & getTargetTriple() const
const MCAsmInfo & getMCAsmInfo() const
Return target specific asm information.
TargetRegisterInfo base class - We assume that the target defines a static array of TargetRegisterDes...
bool hasStackRealignment(const MachineFunction &MF) const
True if stack realignment is required and still possible.
virtual const TargetRegisterInfo * getRegisterInfo() const =0
Return the target's register information.
Triple - Helper class for working with autoconf configuration names.
Definition Triple.h:48
bool isOSBinFormatMachO() const
Tests whether the environment is MachO.
Definition Triple.h:875
This class implements an extremely fast bulk output stream that can only output to a stream.
Definition raw_ostream.h:53
#define llvm_unreachable(msg)
Marks that the current location is not supposed to be reachable.
static unsigned getShiftValue(unsigned Imm)
getShiftValue - Extract the shift value.
static unsigned getArithExtendImm(AArch64_AM::ShiftExtendType ET, unsigned Imm)
getArithExtendImm - Encode the extend type and shift amount for an arithmetic instruction: imm: 3-bit...
const unsigned StackProbeMaxLoopUnroll
Maximum number of iterations to unroll for a constant size probing loop.
const unsigned StackProbeMaxUnprobedStack
Maximum allowed number of unprobed bytes above SP at an ABI boundary.
constexpr char Align[]
Key for Kernel::Arg::Metadata::mAlign.
constexpr char Attrs[]
Key for Kernel::Metadata::mAttrs.
unsigned ID
LLVM IR allows to use arbitrary numbers as calling convention identifiers.
Definition CallingConv.h:24
@ AArch64_SVE_VectorCall
Used between AArch64 SVE functions.
@ PreserveMost
Used for runtime calls that preserves most registers.
Definition CallingConv.h:63
@ CXX_FAST_TLS
Used for access functions.
Definition CallingConv.h:72
@ GHC
Used by the Glasgow Haskell Compiler (GHC).
Definition CallingConv.h:50
@ PreserveAll
Used for runtime calls that preserves (almost) all registers.
Definition CallingConv.h:66
@ Fast
Attempts to make calls as fast as possible (e.g.
Definition CallingConv.h:41
@ PreserveNone
Used for runtime calls that preserves none general registers.
Definition CallingConv.h:90
@ Win64
The C convention as implemented on Windows/x86-64 and AArch64.
@ SwiftTail
This follows the Swift calling convention in how arguments are passed but guarantees tail calls will ...
Definition CallingConv.h:87
@ C
The default llvm calling convention, compatible with C.
Definition CallingConv.h:34
initializer< Ty > init(const Ty &Val)
NodeAddr< InstrNode * > Instr
Definition RDFGraph.h:389
BaseReg
Stack frame base register. Bit 0 of FREInfo.Info.
Definition SFrame.h:77
This is an optimization pass for GlobalISel generic memory operations.
@ Offset
Definition DWP.cpp:577
detail::zippy< detail::zip_shortest, T, U, Args... > zip(T &&t, U &&u, Args &&...args)
zip iterator for two or more iteratable types.
Definition STLExtras.h:830
void stable_sort(R &&Range)
Definition STLExtras.h:2116
MachineInstrBuilder BuildMI(MachineFunction &MF, const MIMetadata &MIMD, const MCInstrDesc &MCID)
Builder interface. Specify how to create the initial instruction itself.
int isAArch64FrameOffsetLegal(const MachineInstr &MI, StackOffset &Offset, bool *OutUseUnscaledOp=nullptr, unsigned *OutUnscaledOp=nullptr, int64_t *EmittableOffset=nullptr)
Check if the Offset is a valid frame offset for MI.
@ Unknown
Not known to have no common set bits.
RegState
Flags to represent properties of register accesses.
@ Define
Register definition.
auto enumerate(FirstRange &&First, RestRanges &&...Rest)
Given two or more input ranges, returns a new range whose values are tuples (A, B,...
Definition STLExtras.h:2554
constexpr RegState getKillRegState(bool B)
decltype(auto) dyn_cast(const From &Val)
dyn_cast<X> - Return the argument parameter cast to the specified type.
Definition Casting.h:643
void append_range(Container &C, Range &&R)
Wrapper function to append range R to container C.
Definition STLExtras.h:2208
@ AArch64FrameOffsetCannotUpdate
Offset cannot apply.
constexpr T alignDown(U Value, V Align, W Skew=0)
Returns the largest unsigned integer less than or equal to Value and is Skew mod Align.
Definition MathExtras.h:541
auto dyn_cast_or_null(const Y &Val)
Definition Casting.h:753
bool any_of(R &&range, UnaryPredicate P)
Provide wrappers to std::any_of which take ranges instead of having to pass begin/end explicitly.
Definition STLExtras.h:1746
auto formatv(bool Validate, const char *Fmt, Ts &&...Vals)
auto reverse(ContainerTy &&C)
Definition STLExtras.h:407
void sort(IteratorTy Start, IteratorTy End)
Definition STLExtras.h:1636
LLVM_ABI raw_ostream & dbgs()
dbgs() - This returns a reference to a raw_ostream for debugging messages.
Definition Debug.cpp:209
void emitFrameOffset(MachineBasicBlock &MBB, MachineBasicBlock::iterator MBBI, const DebugLoc &DL, unsigned DestReg, unsigned SrcReg, StackOffset Offset, const TargetInstrInfo *TII, MachineInstr::MIFlag=MachineInstr::NoFlags, bool SetNZCV=false, bool NeedsWinCFI=false, bool *HasWinCFI=nullptr, bool EmitCFAOffset=false, StackOffset InitialOffset={}, unsigned FrameReg=AArch64::SP)
emitFrameOffset - Emit instructions as needed to set DestReg to SrcReg plus Offset.
LLVM_ABI void report_fatal_error(Error Err, bool gen_crash_diag=true)
Definition Error.cpp:163
constexpr uint64_t alignTo(uint64_t Size, Align A)
Returns a multiple of A needed to store Size bytes.
Definition Alignment.h:144
constexpr RegState getDefRegState(bool B)
class LLVM_GSL_OWNER SmallVector
Forward declaration of SmallVector so that calculateSmallVectorDefaultInlinedElements can reference s...
@ First
Helpers to iterate all locations in the MemoryEffectsBase class.
Definition ModRef.h:74
uint16_t MCPhysReg
An unsigned integer type large enough to represent all physical registers, but not necessarily virtua...
Definition MCRegister.h:21
RelativeUniformCounterPtr ValuesPtrExpr VTableAddr Count
Definition InstrProf.h:145
raw_ostream & operator<<(raw_ostream &OS, const APFixedPoint &FX)
auto count_if(R &&Range, UnaryPredicate P)
Wrapper function around std::count_if to count the number of times an element satisfying a given pred...
Definition STLExtras.h:2019
auto find_if(R &&Range, UnaryPredicate P)
Provide wrappers to std::find_if which take ranges instead of having to pass begin/end explicitly.
Definition STLExtras.h:1772
void erase_if(Container &C, UnaryPredicate P)
Provide a container algorithm similar to C++ Library Fundamentals v2's erase_if which is equivalent t...
Definition STLExtras.h:2192
bool is_contained(R &&Range, const E &Element)
Returns true if Element is found in Range.
Definition STLExtras.h:1947
LLVM_ABI const Value * getUnderlyingObject(const Value *V, unsigned MaxLookup=MaxLookupSearchDepth)
This method strips off any GEP address adjustments, pointer casts or llvm.threadlocal....
void fullyRecomputeLiveIns(ArrayRef< MachineBasicBlock * > MBBs)
Convenience function for recomputing live-in's for a set of MBBs until the computation converges.
LLVM_ABI Printable printReg(Register Reg, const TargetRegisterInfo *TRI=nullptr, unsigned SubIdx=0, const MachineRegisterInfo *MRI=nullptr)
Prints virtual and physical registers with or without a TRI instance.
MCRegisterClass TargetRegisterClass
Definition FastISel.h:58
void swap(llvm::BitVector &LHS, llvm::BitVector &RHS)
Implement std::swap in terms of BitVector swap.
Definition BitVector.h:880
bool operator<(const StackAccess &Rhs) const
void print(raw_ostream &OS) const
int64_t start() const
std::string getTypeString() const
int64_t end() const
This struct is a compact representation of a valid (non-zero power of two) alignment.
Definition Alignment.h:39
constexpr uint64_t value() const
This is a hole in the type system and should not be abused.
Definition Alignment.h:77
Pair of physical register and lane mask.
static LLVM_ABI MachinePointerInfo getUnknownStack(MachineFunction &MF)
Stack memory without other information.
static LLVM_ABI MachinePointerInfo getFixedStack(MachineFunction &MF, int FI, int64_t Offset=0)
Return a MachinePointerInfo record that refers to the specified FrameIndex.
SmallVector< WinEHTryBlockMapEntry, 4 > TryBlockMap
SmallVector< WinEHHandlerType, 1 > HandlerArray