LLVM 24.0.0git
AArch64FrameLowering.cpp
Go to the documentation of this file.
1//===- AArch64FrameLowering.cpp - AArch64 Frame Lowering -------*- C++ -*-====//
2//
3// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
4// See https://llvm.org/LICENSE.txt for license information.
5// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
6//
7//===----------------------------------------------------------------------===//
8//
9// This file contains the AArch64 implementation of TargetFrameLowering class.
10//
11// On AArch64, stack frames are structured as follows:
12//
13// The stack grows downward.
14//
15// All of the individual frame areas on the frame below are optional, i.e. it's
16// possible to create a function so that the particular area isn't present
17// in the frame.
18//
19// At function entry, the "frame" looks as follows:
20//
21// | | Higher address
22// |-----------------------------------|
23// | |
24// | arguments passed on the stack |
25// | |
26// |-----------------------------------| <- sp
27// | | Lower address
28//
29//
30// After the prologue has run, the frame has the following general structure.
31// Note that this doesn't depict the case where a red-zone is used. Also,
32// technically the last frame area (VLAs) doesn't get created until in the
33// main function body, after the prologue is run. However, it's depicted here
34// for completeness.
35//
36// | | Higher address
37// |-----------------------------------|
38// | |
39// | arguments passed on the stack |
40// | |
41// |-----------------------------------|
42// | |
43// | (Win64 only) varargs from reg |
44// | |
45// |-----------------------------------|
46// | |
47// | (Win64 only) callee-saved SVE reg |
48// | |
49// |-----------------------------------|
50// | |
51// | callee-saved gpr registers | <--.
52// | | | On Darwin platforms these
53// |- - - - - - - - - - - - - - - - - -| | callee saves are swapped,
54// | prev_lr | | (frame record first)
55// | prev_fp | <--'
56// | async context if needed |
57// | (a.k.a. "frame record") |
58// |-----------------------------------| <- fp(=x29)
59// Default SVE stack layout Split SVE objects
60// (aarch64-split-sve-objects=false) (aarch64-split-sve-objects=true)
61// |-----------------------------------| |-----------------------------------|
62// | <hazard padding> | | callee-saved PPR registers |
63// |-----------------------------------| |-----------------------------------|
64// | | | PPR stack objects |
65// | callee-saved fp/simd/SVE regs | |-----------------------------------|
66// | | | <hazard padding> |
67// |-----------------------------------| |-----------------------------------|
68// | | | callee-saved ZPR/FPR registers |
69// | SVE stack objects | |-----------------------------------|
70// | | | ZPR stack objects |
71// |-----------------------------------| |-----------------------------------|
72// ^ NB: FPR CSRs are promoted to ZPRs
73// |-----------------------------------|
74// |.empty.space.to.make.part.below....|
75// |.aligned.in.case.it.needs.more.than| (size of this area is unknown at
76// |.the.standard.16-byte.alignment....| compile time; if present)
77// |-----------------------------------|
78// | local variables of fixed size |
79// | including spill slots |
80// | <FPR> |
81// | <hazard padding> |
82// | <GPR> |
83// |-----------------------------------| <- bp(not defined by ABI,
84// |.variable-sized.local.variables....| LLVM chooses X19)
85// |.(VLAs)............................| (size of this area is unknown at
86// |...................................| compile time)
87// |-----------------------------------| <- sp
88// | | Lower address
89//
90//
91// To access the data in a frame, at-compile time, a constant offset must be
92// computable from one of the pointers (fp, bp, sp) to access it. The size
93// of the areas with a dotted background cannot be computed at compile-time
94// if they are present, making it required to have all three of fp, bp and
95// sp to be set up to be able to access all contents in the frame areas,
96// assuming all of the frame areas are non-empty.
97//
98// For most functions, some of the frame areas are empty. For those functions,
99// it may not be necessary to set up fp or bp:
100// * A base pointer is definitely needed when there are both VLAs and local
101// variables with more-than-default alignment requirements.
102// * A frame pointer is definitely needed when there are local variables with
103// more-than-default alignment requirements.
104//
105// For Darwin platforms the frame-record (fp, lr) is stored at the top of the
106// callee-saved area, since the unwind encoding does not allow for encoding
107// this dynamically and existing tools depend on this layout. For other
108// platforms, the frame-record is stored at the bottom of the (gpr) callee-saved
109// area to allow SVE stack objects (allocated directly below the callee-saves,
110// if available) to be accessed directly from the framepointer.
111// The SVE spill/fill instructions have VL-scaled addressing modes such
112// as:
113// ldr z8, [fp, #-7 mul vl]
114// For SVE the size of the vector length (VL) is not known at compile-time, so
115// '#-7 mul vl' is an offset that can only be evaluated at runtime. With this
116// layout, we don't need to add an unscaled offset to the framepointer before
117// accessing the SVE object in the frame.
118//
119// In some cases when a base pointer is not strictly needed, it is generated
120// anyway when offsets from the frame pointer to access local variables become
121// so large that the offset can't be encoded in the immediate fields of loads
122// or stores.
123//
124// Outgoing function arguments must be at the bottom of the stack frame when
125// calling another function. If we do not have variable-sized stack objects, we
126// can allocate a "reserved call frame" area at the bottom of the local
127// variable area, large enough for all outgoing calls. If we do have VLAs, then
128// the stack pointer must be decremented and incremented around each call to
129// make space for the arguments below the VLAs.
130//
131// FIXME: also explain the redzone concept.
132//
133// About stack hazards: Under some SME contexts, a coprocessor with its own
134// separate cache can used for FP operations. This can create hazards if the CPU
135// and the SME unit try to access the same area of memory, including if the
136// access is to an area of the stack. To try to alleviate this we attempt to
137// introduce extra padding into the stack frame between FP and GPR accesses,
138// controlled by the aarch64-stack-hazard-size option. Without changing the
139// layout of the stack frame in the diagram above, a stack object of size
140// aarch64-stack-hazard-size is added between GPR and FPR CSRs. Another is added
141// to the stack objects section, and stack objects are sorted so that FPR >
142// Hazard padding slot > GPRs (where possible). Unfortunately some things are
143// not handled well (VLA area, arguments on the stack, objects with both GPR and
144// FPR accesses), but if those are controlled by the user then the entire stack
145// frame becomes GPR at the start/end with FPR in the middle, surrounded by
146// Hazard padding.
147//
148// An example of the prologue:
149//
150// .globl __foo
151// .align 2
152// __foo:
153// Ltmp0:
154// .cfi_startproc
155// .cfi_personality 155, ___gxx_personality_v0
156// Leh_func_begin:
157// .cfi_lsda 16, Lexception33
158//
159// stp xa,bx, [sp, -#offset]!
160// ...
161// stp x28, x27, [sp, #offset-32]
162// stp fp, lr, [sp, #offset-16]
163// add fp, sp, #offset - 16
164// sub sp, sp, #1360
165//
166// The Stack:
167// +-------------------------------------------+
168// 10000 | ........ | ........ | ........ | ........ |
169// 10004 | ........ | ........ | ........ | ........ |
170// +-------------------------------------------+
171// 10008 | ........ | ........ | ........ | ........ |
172// 1000c | ........ | ........ | ........ | ........ |
173// +===========================================+
174// 10010 | X28 Register |
175// 10014 | X28 Register |
176// +-------------------------------------------+
177// 10018 | X27 Register |
178// 1001c | X27 Register |
179// +===========================================+
180// 10020 | Frame Pointer |
181// 10024 | Frame Pointer |
182// +-------------------------------------------+
183// 10028 | Link Register |
184// 1002c | Link Register |
185// +===========================================+
186// 10030 | ........ | ........ | ........ | ........ |
187// 10034 | ........ | ........ | ........ | ........ |
188// +-------------------------------------------+
189// 10038 | ........ | ........ | ........ | ........ |
190// 1003c | ........ | ........ | ........ | ........ |
191// +-------------------------------------------+
192//
193// [sp] = 10030 :: >>initial value<<
194// sp = 10020 :: stp fp, lr, [sp, #-16]!
195// fp = sp == 10020 :: mov fp, sp
196// [sp] == 10020 :: stp x28, x27, [sp, #-16]!
197// sp == 10010 :: >>final value<<
198//
199// The frame pointer (w29) points to address 10020. If we use an offset of
200// '16' from 'w29', we get the CFI offsets of -8 for w30, -16 for w29, -24
201// for w27, and -32 for w28:
202//
203// Ltmp1:
204// .cfi_def_cfa w29, 16
205// Ltmp2:
206// .cfi_offset w30, -8
207// Ltmp3:
208// .cfi_offset w29, -16
209// Ltmp4:
210// .cfi_offset w27, -24
211// Ltmp5:
212// .cfi_offset w28, -32
213//
214//===----------------------------------------------------------------------===//
215
216#include "AArch64FrameLowering.h"
217#include "AArch64InstrInfo.h"
220#include "AArch64RegisterInfo.h"
221#include "AArch64SMEAttributes.h"
222#include "AArch64Subtarget.h"
225#include "llvm/ADT/ScopeExit.h"
226#include "llvm/ADT/SmallVector.h"
244#include "llvm/IR/Attributes.h"
245#include "llvm/IR/CallingConv.h"
246#include "llvm/IR/DataLayout.h"
247#include "llvm/IR/DebugLoc.h"
248#include "llvm/IR/Function.h"
249#include "llvm/MC/MCAsmInfo.h"
250#include "llvm/MC/MCDwarf.h"
252#include "llvm/Support/Debug.h"
259#include <cassert>
260#include <cstdint>
261#include <iterator>
262#include <optional>
263#include <vector>
264
265using namespace llvm;
266
267#define DEBUG_TYPE "frame-info"
268
269static cl::opt<bool> EnableRedZone("aarch64-redzone",
270 cl::desc("enable use of redzone on AArch64"),
271 cl::init(false), cl::Hidden);
272
274 "stack-tagging-merge-settag",
275 cl::desc("merge settag instruction in function epilog"), cl::init(true),
276 cl::Hidden);
277
278static cl::opt<bool> OrderFrameObjects("aarch64-order-frame-objects",
279 cl::desc("sort stack allocations"),
280 cl::init(true), cl::Hidden);
281
282static cl::opt<bool>
283 SplitSVEObjects("aarch64-split-sve-objects",
284 cl::desc("Split allocation of ZPR & PPR objects"),
285 cl::init(true), cl::Hidden);
286
288 "homogeneous-prolog-epilog", cl::Hidden,
289 cl::desc("Emit homogeneous prologue and epilogue for the size "
290 "optimization (default = off)"));
291
292// Stack hazard size for analysis remarks. StackHazardSize takes precedence.
294 StackHazardRemarkSize("aarch64-stack-hazard-remark-size", cl::init(0),
295 cl::Hidden);
296// Whether to insert padding into non-streaming functions (for testing).
297static cl::opt<bool>
298 StackHazardInNonStreaming("aarch64-stack-hazard-in-non-streaming",
299 cl::init(false), cl::Hidden);
300
302 "aarch64-disable-multivector-spill-fill",
303 cl::desc("Disable use of LD/ST pairs for SME2 or SVE2p1"), cl::init(false),
304 cl::Hidden);
305
306int64_t
308 MachineBasicBlock &MBB) const {
309 MachineBasicBlock::iterator MBBI = MBB.getLastNonDebugInstr();
311 bool IsTailCallReturn = (MBB.end() != MBBI)
313 : false;
314
315 int64_t ArgumentPopSize = 0;
316 if (IsTailCallReturn) {
317 MachineOperand &StackAdjust = MBBI->getOperand(1);
318
319 // For a tail-call in a callee-pops-arguments environment, some or all of
320 // the stack may actually be in use for the call's arguments, this is
321 // calculated during LowerCall and consumed here...
322 ArgumentPopSize = StackAdjust.getImm();
323 } else {
324 // ... otherwise the amount to pop is *all* of the argument space,
325 // conveniently stored in the MachineFunctionInfo by
326 // LowerFormalArguments. This will, of course, be zero for the C calling
327 // convention.
328 ArgumentPopSize = AFI->getArgumentStackToRestore();
329 }
330
331 return ArgumentPopSize;
332}
333
335 MachineFunction &MF);
336
337enum class AssignObjectOffsets { No, Yes };
338/// Process all the SVE stack objects and the SVE stack size and offsets for
339/// each object. If AssignOffsets is "Yes", the offsets get assigned (and SVE
340/// stack sizes set). Returns the size of the SVE stack.
342 AssignObjectOffsets AssignOffsets);
343
344static unsigned getStackHazardSize(const MachineFunction &MF) {
345 return MF.getSubtarget<AArch64Subtarget>().getStreamingHazardSize();
346}
347
353
356 // With split SVE objects, the hazard padding is added to the PPR region,
357 // which places it between the [GPR, PPR] area and the [ZPR, FPR] area. This
358 // avoids hazards between both GPRs and FPRs and ZPRs and PPRs.
361 : 0,
362 AFI->getStackSizePPR());
363}
364
365// Conservatively, returns true if the function is likely to have SVE vectors
366// on the stack. This function is safe to be called before callee-saves or
367// object offsets have been determined.
369 const MachineFunction &MF) {
370 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
371 if (AFI->isSVECC())
372 return true;
373
374 if (AFI->hasCalculatedStackSizeSVE())
375 return bool(AFL.getSVEStackSize(MF));
376
377 const MachineFrameInfo &MFI = MF.getFrameInfo();
378 for (int FI = MFI.getObjectIndexBegin(); FI < MFI.getObjectIndexEnd(); FI++) {
379 if (MFI.hasScalableStackID(FI))
380 return true;
381 }
382
383 return false;
384}
385
386static bool isTargetWindows(const MachineFunction &MF) {
387 // TODO: Should this include targets like UEFI (which use Windows CFI)?
388 // Note: Currently, there is not AArch64 support for UEFI. The value returned
389 // here must align with the predicate used for returning the list of callee
390 // saved regs in AArch64RegisterInfo::getCalleeSavedRegs(), so that we use
391 // invalidateWindowsRegisterPairing() where appropriate.
393}
394
396 const MachineFunction &MF) const {
397 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
398 return isTargetWindows(MF) && AFI->getSVECalleeSavedStackSize();
399}
400
401/// Returns true if a homogeneous prolog or epilog code can be emitted
402/// for the size optimization. If possible, a frame helper call is injected.
403/// When Exit block is given, this check is for epilog.
404bool AArch64FrameLowering::homogeneousPrologEpilog(
405 MachineFunction &MF, MachineBasicBlock *Exit) const {
406 if (!MF.getFunction().hasMinSize())
407 return false;
409 return false;
410 if (EnableRedZone)
411 return false;
412
413 // TODO: Window is supported yet.
414 if (isTargetWindows(MF))
415 return false;
416
417 // TODO: SVE is not supported yet.
418 if (isLikelyToHaveSVEStack(*this, MF))
419 return false;
420
421 // Bail on stack adjustment needed on return for simplicity.
422 const MachineFrameInfo &MFI = MF.getFrameInfo();
423 const TargetRegisterInfo *RegInfo = MF.getSubtarget().getRegisterInfo();
424 if (MFI.hasVarSizedObjects() || RegInfo->hasStackRealignment(MF))
425 return false;
426 if (Exit && getArgumentStackToRestore(MF, *Exit))
427 return false;
428
429 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
430 if (AFI->hasSwiftAsyncContext() || AFI->hasStreamingModeChanges())
431 return false;
432
433 // If there are an odd number of GPRs before LR and FP in the CSRs list,
434 // they will not be paired into one RegPairInfo, which is incompatible with
435 // the assumption made by the homogeneous prolog epilog pass.
436 const MCPhysReg *CSRegs = MF.getRegInfo().getCalleeSavedRegs();
437 unsigned NumGPRs = 0;
438 for (unsigned I = 0; CSRegs[I]; ++I) {
439 Register Reg = CSRegs[I];
440 if (Reg == AArch64::LR) {
441 assert(CSRegs[I + 1] == AArch64::FP);
442 if (NumGPRs % 2 != 0)
443 return false;
444 break;
445 }
446 if (AArch64::GPR64RegClass.contains(Reg))
447 ++NumGPRs;
448 }
449
450 return true;
451}
452
453/// Returns true if CSRs should be paired.
454bool AArch64FrameLowering::producePairRegisters(MachineFunction &MF) const {
455 return produceCompactUnwindFrame(*this, MF) || homogeneousPrologEpilog(MF);
456}
457
458/// This is the biggest offset to the stack pointer we can encode in aarch64
459/// instructions (without using a separate calculation and a temp register).
460/// Note that the exception here are vector stores/loads which cannot encode any
461/// displacements (see estimateRSStackSizeLimit(), isAArch64FrameOffsetLegal()).
462static const unsigned DefaultSafeSPDisplacement = 255;
463
464/// Look at each instruction that references stack frames and return the stack
465/// size limit beyond which some of these instructions will require a scratch
466/// register during their expansion later.
468 // FIXME: For now, just conservatively guesstimate based on unscaled indexing
469 // range. We'll end up allocating an unnecessary spill slot a lot, but
470 // realistically that's not a big deal at this stage of the game.
471 for (MachineBasicBlock &MBB : MF) {
472 for (MachineInstr &MI : MBB) {
473 if (MI.isDebugInstr() || MI.isPseudo() ||
474 MI.getOpcode() == AArch64::ADDXri ||
475 MI.getOpcode() == AArch64::ADDSXri)
476 continue;
477
478 for (const MachineOperand &MO : MI.operands()) {
479 if (!MO.isFI())
480 continue;
481
483 if (isAArch64FrameOffsetLegal(MI, Offset, nullptr, nullptr, nullptr) ==
485 return 0;
486 }
487 }
488 }
490}
491
496
497unsigned
498AArch64FrameLowering::getFixedObjectSize(const MachineFunction &MF,
499 const AArch64FunctionInfo *AFI,
500 bool IsWin64, bool IsFunclet) const {
501 assert(AFI->getTailCallReservedStack() % 16 == 0 &&
502 "Tail call reserved stack must be aligned to 16 bytes");
503 if (!IsWin64 || IsFunclet) {
504 return AFI->getTailCallReservedStack();
505 } else {
506 if (AFI->getTailCallReservedStack() != 0 &&
507 !MF.getFunction().getAttributes().hasAttrSomewhere(
508 Attribute::SwiftAsync))
509 report_fatal_error("cannot generate ABI-changing tail call for Win64");
510 unsigned FixedObjectSize = AFI->getTailCallReservedStack();
511
512 // Var args are stored here in the primary function.
513 FixedObjectSize += AFI->getVarArgsGPRSize();
514
515 if (MF.hasEHFunclets()) {
516 // Catch objects are stored here in the primary function.
517 const MachineFrameInfo &MFI = MF.getFrameInfo();
518 const WinEHFuncInfo &EHInfo = *MF.getWinEHFuncInfo();
519 SmallSetVector<int, 8> CatchObjFrameIndices;
520 for (const WinEHTryBlockMapEntry &TBME : EHInfo.TryBlockMap) {
521 for (const WinEHHandlerType &H : TBME.HandlerArray) {
522 int FrameIndex = H.CatchObj.FrameIndex;
523 if ((FrameIndex != INT_MAX) &&
524 CatchObjFrameIndices.insert(FrameIndex)) {
525 FixedObjectSize = alignTo(FixedObjectSize,
526 MFI.getObjectAlign(FrameIndex).value()) +
527 MFI.getObjectSize(FrameIndex);
528 }
529 }
530 }
531 // To support EH funclets we allocate an UnwindHelp object
532 FixedObjectSize += 8;
533 }
534 return alignTo(FixedObjectSize, 16);
535 }
536}
537
539 if (!EnableRedZone)
540 return false;
541
542 // Don't use the red zone if the function explicitly asks us not to.
543 // This is typically used for kernel code.
544 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
545 const unsigned RedZoneSize =
547 if (!RedZoneSize)
548 return false;
549
550 const MachineFrameInfo &MFI = MF.getFrameInfo();
552 uint64_t NumBytes = AFI->getLocalStackSize();
553
554 // If neither NEON or SVE are available, a COPY from one Q-reg to
555 // another requires a spill -> reload sequence. We can do that
556 // using a pre-decrementing store/post-decrementing load, but
557 // if we do so, we can't use the Red Zone.
558 bool LowerQRegCopyThroughMem = Subtarget.hasFPARMv8() &&
559 !Subtarget.isNeonAvailable() &&
560 !Subtarget.hasSVE();
561
562 return !(MFI.hasCalls() || hasFP(MF) || NumBytes > RedZoneSize ||
563 AFI->hasSVEStackSize() || LowerQRegCopyThroughMem);
564}
565
566/// hasFPImpl - Return true if the specified function should have a dedicated
567/// frame pointer register.
569 const MachineFrameInfo &MFI = MF.getFrameInfo();
570 const TargetRegisterInfo *RegInfo = MF.getSubtarget().getRegisterInfo();
572
573 // Win64 EH requires a frame pointer if funclets are present, as the locals
574 // are accessed off the frame pointer in both the parent function and the
575 // funclets.
576 if (MF.hasEHFunclets())
577 return true;
578
579 // When the stack guard is mixed with the frame pointer, a dedicated FP is
580 // required so the guard value remains stable in the presence of dynamic
581 // stack allocations (e.g. _alloca on MSVCRT).
582 if (MFI.hasStackProtectorIndex()) {
583 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
584 if (Subtarget.getTargetLowering()->useStackGuardMixFP())
585 return true;
586 }
587
588 // Retain behavior of always omitting the FP for leaf functions when possible.
590 return true;
591 if (MFI.hasVarSizedObjects() || MFI.isFrameAddressTaken() ||
592 MFI.hasStackMap() || MFI.hasPatchPoint() ||
593 RegInfo->hasStackRealignment(MF))
594 return true;
595
596 // If we:
597 //
598 // 1. Have streaming mode changes
599 // OR:
600 // 2. Have a streaming body with SVE stack objects
601 //
602 // Then the value of VG restored when unwinding to this function may not match
603 // the value of VG used to set up the stack.
604 //
605 // This is a problem as the CFA can be described with an expression of the
606 // form: CFA = SP + NumBytes + VG * NumScalableBytes.
607 //
608 // If the value of VG used in that expression does not match the value used to
609 // set up the stack, an incorrect address for the CFA will be computed, and
610 // unwinding will fail.
611 //
612 // We work around this issue by ensuring the frame-pointer can describe the
613 // CFA in either of these cases.
614 if (AFI.needsDwarfUnwindInfo(MF) &&
617 return true;
618 // With large callframes around we may need to use FP to access the scavenging
619 // emergency spillslot.
620 //
621 // Unfortunately some calls to hasFP() like machine verifier ->
622 // getReservedReg() -> hasFP in the middle of global isel are too early
623 // to know the max call frame size. Hopefully conservatively returning "true"
624 // in those cases is fine.
625 // DefaultSafeSPDisplacement is fine as we only emergency spill GP regs.
626 if (!MFI.isMaxCallFrameSizeComputed() ||
628 return true;
629
630 return false;
631}
632
633/// Should the Frame Pointer be reserved for the current function?
635 const TargetMachine &TM = MF.getTarget();
636 const Triple &TT = TM.getTargetTriple();
637
638 // These OSes require the frame chain is valid, even if the current frame does
639 // not use a frame pointer.
640 if (TT.isOSDarwin() || TT.isOSWindows())
641 return true;
642
643 // If the function has a frame pointer, it is reserved.
644 if (hasFP(MF))
645 return true;
646
647 // Frontend has requested to preserve the frame pointer.
649 return true;
650
651 return false;
652}
653
654/// hasReservedCallFrame - Under normal circumstances, when a frame pointer is
655/// not required, we reserve argument space for call sites in the function
656/// immediately on entry to the current function. This eliminates the need for
657/// add/sub sp brackets around call sites. Returns true if the call frame is
658/// included as part of the stack frame.
660 const MachineFunction &MF) const {
661 // The stack probing code for the dynamically allocated outgoing arguments
662 // area assumes that the stack is probed at the top - either by the prologue
663 // code, which issues a probe if `hasVarSizedObjects` return true, or by the
664 // most recent variable-sized object allocation. Changing the condition here
665 // may need to be followed up by changes to the probe issuing logic.
666 return !MF.getFrameInfo().hasVarSizedObjects();
667}
668
672
673 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
674 const AArch64InstrInfo *TII = Subtarget.getInstrInfo();
675 const AArch64TargetLowering *TLI = Subtarget.getTargetLowering();
676 [[maybe_unused]] MachineFrameInfo &MFI = MF.getFrameInfo();
677 DebugLoc DL = I->getDebugLoc();
678 unsigned Opc = I->getOpcode();
679 bool IsDestroy = Opc == TII->getCallFrameDestroyOpcode();
680 uint64_t CalleePopAmount = IsDestroy ? I->getOperand(1).getImm() : 0;
681
682 if (!hasReservedCallFrame(MF)) {
683 int64_t Amount = I->getOperand(0).getImm();
684 Amount = alignTo(Amount, getStackAlign());
685 if (!IsDestroy)
686 Amount = -Amount;
687
688 // N.b. if CalleePopAmount is valid but zero (i.e. callee would pop, but it
689 // doesn't have to pop anything), then the first operand will be zero too so
690 // this adjustment is a no-op.
691 if (CalleePopAmount == 0) {
692 // FIXME: in-function stack adjustment for calls is limited to 24-bits
693 // because there's no guaranteed temporary register available.
694 //
695 // ADD/SUB (immediate) has only LSL #0 and LSL #12 available.
696 // 1) For offset <= 12-bit, we use LSL #0
697 // 2) For 12-bit <= offset <= 24-bit, we use two instructions. One uses
698 // LSL #0, and the other uses LSL #12.
699 //
700 // Most call frames will be allocated at the start of a function so
701 // this is OK, but it is a limitation that needs dealing with.
702 assert(Amount > -0xffffff && Amount < 0xffffff && "call frame too large");
703
704 if (TLI->hasInlineStackProbe(MF) &&
706 // When stack probing is enabled, the decrement of SP may need to be
707 // probed. We only need to do this if the call site needs 1024 bytes of
708 // space or more, because a region smaller than that is allowed to be
709 // unprobed at an ABI boundary. We rely on the fact that SP has been
710 // probed exactly at this point, either by the prologue or most recent
711 // dynamic allocation.
713 "non-reserved call frame without var sized objects?");
714 Register ScratchReg =
715 MF.getRegInfo().createVirtualRegister(&AArch64::GPR64RegClass);
716 inlineStackProbeFixed(I, ScratchReg, -Amount, StackOffset::get(0, 0));
717 } else {
718 emitFrameOffset(MBB, I, DL, AArch64::SP, AArch64::SP,
719 StackOffset::getFixed(Amount), TII);
720 }
721 }
722 } else if (CalleePopAmount != 0) {
723 // If the calling convention demands that the callee pops arguments from the
724 // stack, we want to add it back if we have a reserved call frame.
725 assert(CalleePopAmount < 0xffffff && "call frame too large");
726 emitFrameOffset(MBB, I, DL, AArch64::SP, AArch64::SP,
727 StackOffset::getFixed(-(int64_t)CalleePopAmount), TII);
728 }
729 return MBB.erase(I);
730}
731
733 MachineBasicBlock &MBB) const {
734
735 MachineFunction &MF = *MBB.getParent();
736 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
737 const auto &TRI = *Subtarget.getRegisterInfo();
738 const auto &MFI = *MF.getInfo<AArch64FunctionInfo>();
739
740 CFIInstBuilder CFIBuilder(MBB, MBB.begin(), MachineInstr::NoFlags);
741
742 // Reset the CFA to `SP + 0`.
743 CFIBuilder.buildDefCFA(AArch64::SP, 0);
744
745 // Flip the RA sign state.
746 if (MFI.shouldSignReturnAddress(MF)) {
747 if (MFI.branchProtectionPAuthLR()) {
748 CFIBuilder.buildNegateRAStateWithPC();
749 } else if (!MF.getTarget().getTargetTriple().isOSBinFormatMachO()) {
750 CFIBuilder.buildNegateRAState();
751 }
752 }
753
754 // Shadow call stack uses X18, reset it.
755 if (MFI.needsShadowCallStackPrologueEpilogue(MF))
756 CFIBuilder.buildSameValue(AArch64::X18);
757
758 // Emit .cfi_same_value for callee-saved registers.
759 const std::vector<CalleeSavedInfo> &CSI =
761 for (const auto &Info : CSI) {
762 MCRegister Reg = Info.getReg();
763 if (!TRI.regNeedsCFI(Reg, Reg))
764 continue;
765 CFIBuilder.buildSameValue(Reg);
766 }
767}
768
770 switch (Reg.id()) {
771 default:
772 // The called routine is expected to preserve r19-r28
773 // r29 and r30 are used as frame pointer and link register resp.
774 return 0;
775
776 // GPRs
777#define CASE(n) \
778 case AArch64::W##n: \
779 case AArch64::X##n: \
780 return AArch64::X##n
781 CASE(0);
782 CASE(1);
783 CASE(2);
784 CASE(3);
785 CASE(4);
786 CASE(5);
787 CASE(6);
788 CASE(7);
789 CASE(8);
790 CASE(9);
791 CASE(10);
792 CASE(11);
793 CASE(12);
794 CASE(13);
795 CASE(14);
796 CASE(15);
797 CASE(16);
798 CASE(17);
799 CASE(18);
800#undef CASE
801
802 // FPRs
803#define CASE(n) \
804 case AArch64::B##n: \
805 case AArch64::H##n: \
806 case AArch64::S##n: \
807 case AArch64::D##n: \
808 case AArch64::Q##n: \
809 return HasSVE ? AArch64::Z##n : AArch64::Q##n
810 CASE(0);
811 CASE(1);
812 CASE(2);
813 CASE(3);
814 CASE(4);
815 CASE(5);
816 CASE(6);
817 CASE(7);
818 CASE(8);
819 CASE(9);
820 CASE(10);
821 CASE(11);
822 CASE(12);
823 CASE(13);
824 CASE(14);
825 CASE(15);
826 CASE(16);
827 CASE(17);
828 CASE(18);
829 CASE(19);
830 CASE(20);
831 CASE(21);
832 CASE(22);
833 CASE(23);
834 CASE(24);
835 CASE(25);
836 CASE(26);
837 CASE(27);
838 CASE(28);
839 CASE(29);
840 CASE(30);
841 CASE(31);
842#undef CASE
843 }
844}
845
846void AArch64FrameLowering::emitZeroCallUsedRegs(BitVector RegsToZero,
848 RegScavenger *) const {
849 // Insertion point.
851
852 // Fake a debug loc.
853 DebugLoc DL;
854 if (MBBI != MBB.end())
855 DL = MBBI->getDebugLoc();
856
857 const MachineFunction &MF = *MBB.getParent();
858 const AArch64Subtarget &STI = MF.getSubtarget<AArch64Subtarget>();
859 const AArch64RegisterInfo &TRI = *STI.getRegisterInfo();
860
861 BitVector GPRsToZero(TRI.getNumRegs());
862 BitVector FPRsToZero(TRI.getNumRegs());
863 bool HasSVE = STI.isSVEorStreamingSVEAvailable();
864 // Without an FP unit (e.g. -mgeneral-regs-only) the FP/vector registers can't
865 // hold a value and there is no instruction to clear them, so leave them out.
866 bool HasFPR = STI.hasFPARMv8();
867 for (MCRegister Reg : RegsToZero.set_bits()) {
868 if (TRI.isGeneralPurposeRegister(MF, Reg)) {
869 // For GPRs, we only care to clear out the 64-bit register.
870 if (MCRegister XReg = getRegisterOrZero(Reg, HasSVE))
871 GPRsToZero.set(XReg);
872 } else if (HasFPR && AArch64InstrInfo::isFpOrNEON(Reg)) {
873 // For FPRs,
874 if (MCRegister XReg = getRegisterOrZero(Reg, HasSVE))
875 FPRsToZero.set(XReg);
876 }
877 }
878
879 const AArch64InstrInfo &TII = *STI.getInstrInfo();
880
881 // Zero out GPRs.
882 for (MCRegister Reg : GPRsToZero.set_bits())
883 TII.buildClearRegister(Reg, MBB, MBBI, DL);
884
885 // Zero out FP/vector registers.
886 for (MCRegister Reg : FPRsToZero.set_bits())
887 TII.buildClearRegister(Reg, MBB, MBBI, DL);
888
889 if (HasSVE) {
890 for (MCRegister PReg :
891 {AArch64::P0, AArch64::P1, AArch64::P2, AArch64::P3, AArch64::P4,
892 AArch64::P5, AArch64::P6, AArch64::P7, AArch64::P8, AArch64::P9,
893 AArch64::P10, AArch64::P11, AArch64::P12, AArch64::P13, AArch64::P14,
894 AArch64::P15}) {
895 if (RegsToZero[PReg])
896 BuildMI(MBB, MBBI, DL, TII.get(AArch64::PFALSE), PReg);
897 }
898 }
899}
900
901bool AArch64FrameLowering::windowsRequiresStackProbe(
902 const MachineFunction &MF, uint64_t StackSizeInBytes) const {
903 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
904 const AArch64FunctionInfo &MFI = *MF.getInfo<AArch64FunctionInfo>();
905 // TODO: When implementing stack protectors, take that into account
906 // for the probe threshold.
907 return Subtarget.isTargetWindows() && MFI.hasStackProbing() &&
908 StackSizeInBytes >= uint64_t(MFI.getStackProbeSize());
909}
910
912 const MachineBasicBlock &MBB) {
913 const MachineFunction *MF = MBB.getParent();
914 LiveRegs.addLiveIns(MBB);
915 // Mark callee saved registers as used so we will not choose them.
916 const MCPhysReg *CSRegs = MF->getRegInfo().getCalleeSavedRegs();
917 for (unsigned i = 0; CSRegs[i]; ++i)
918 LiveRegs.addReg(CSRegs[i]);
919}
920
922AArch64FrameLowering::findScratchNonCalleeSaveRegister(MachineBasicBlock *MBB,
923 bool HasCall) const {
924 MachineFunction *MF = MBB->getParent();
925
926 // If MBB is an entry block, use X9 as the scratch register
927 // preserve_none functions may be using X9 to pass arguments,
928 // so prefer to pick an available register below.
929 if (&MF->front() == MBB &&
931 return AArch64::X9;
932
933 const AArch64Subtarget &Subtarget = MF->getSubtarget<AArch64Subtarget>();
934 const AArch64RegisterInfo &TRI = *Subtarget.getRegisterInfo();
935 LivePhysRegs LiveRegs(TRI);
936 getLiveRegsForEntryMBB(LiveRegs, *MBB);
937 if (HasCall) {
938 LiveRegs.addReg(AArch64::X16);
939 LiveRegs.addReg(AArch64::X17);
940 LiveRegs.addReg(AArch64::X18);
941 }
942
943 // Prefer X9 since it was historically used for the prologue scratch reg.
944 const MachineRegisterInfo &MRI = MF->getRegInfo();
945 if (LiveRegs.available(MRI, AArch64::X9))
946 return AArch64::X9;
947
948 for (unsigned Reg : AArch64::GPR64RegClass) {
949 if (LiveRegs.available(MRI, Reg))
950 return Reg;
951 }
952 return AArch64::NoRegister;
953}
954
956 const MachineBasicBlock &MBB) const {
957 const MachineFunction *MF = MBB.getParent();
958 MachineBasicBlock *TmpMBB = const_cast<MachineBasicBlock *>(&MBB);
959 const AArch64Subtarget &Subtarget = MF->getSubtarget<AArch64Subtarget>();
960 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
961 const AArch64TargetLowering *TLI = Subtarget.getTargetLowering();
963
964 if (AFI->hasSwiftAsyncContext()) {
965 const AArch64RegisterInfo &TRI = *Subtarget.getRegisterInfo();
966 const MachineRegisterInfo &MRI = MF->getRegInfo();
969 // The StoreSwiftAsyncContext clobbers X16 and X17. Make sure they are
970 // available.
971 if (!LiveRegs.available(MRI, AArch64::X16) ||
972 !LiveRegs.available(MRI, AArch64::X17))
973 return false;
974 }
975
976 // Certain stack probing sequences might clobber flags, then we can't use
977 // the block as a prologue if the flags register is a live-in.
979 MBB.isLiveIn(AArch64::NZCV))
980 return false;
981
982 if (RegInfo->hasStackRealignment(*MF) || TLI->hasInlineStackProbe(*MF))
983 if (findScratchNonCalleeSaveRegister(TmpMBB) == AArch64::NoRegister)
984 return false;
985
986 // May need a scratch register (for return value) if require making a special
987 // call
988 if (requiresSaveVG(*MF) ||
989 windowsRequiresStackProbe(*MF, std::numeric_limits<uint64_t>::max()))
990 if (findScratchNonCalleeSaveRegister(TmpMBB, true) == AArch64::NoRegister)
991 return false;
992
993 return true;
994}
995
997 const Function &F = MF.getFunction();
998 return MF.getTarget().getMCAsmInfo().usesWindowsCFI() &&
999 F.needsUnwindTableEntry();
1000}
1001
1002bool AArch64FrameLowering::shouldSignReturnAddressEverywhere(
1003 const MachineFunction &MF) const {
1004 // FIXME: With WinCFI, extra care should be taken to place SEH_PACSignLR
1005 // and SEH_EpilogEnd instructions in the correct order.
1007 return false;
1010}
1011
1012// Given a load or a store instruction, generate an appropriate unwinding SEH
1013// code on Windows.
1015AArch64FrameLowering::insertSEH(MachineBasicBlock::iterator MBBI,
1016 const AArch64InstrInfo &TII,
1017 MachineInstr::MIFlag Flag) const {
1018 unsigned Opc = MBBI->getOpcode();
1019 MachineBasicBlock *MBB = MBBI->getParent();
1020 MachineFunction &MF = *MBB->getParent();
1021 DebugLoc DL = MBBI->getDebugLoc();
1022 unsigned ImmIdx = MBBI->getNumOperands() - 1;
1023 int Imm = MBBI->getOperand(ImmIdx).getImm();
1025 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1026 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
1027
1028 switch (Opc) {
1029 default:
1030 report_fatal_error("No SEH Opcode for this instruction");
1031 case AArch64::STR_ZXI:
1032 case AArch64::LDR_ZXI: {
1033 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1034 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveZReg))
1035 .addImm(Reg0)
1036 .addImm(Imm)
1037 .setMIFlag(Flag);
1038 break;
1039 }
1040 case AArch64::STR_PXI:
1041 case AArch64::LDR_PXI: {
1042 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1043 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SavePReg))
1044 .addImm(Reg0)
1045 .addImm(Imm)
1046 .setMIFlag(Flag);
1047 break;
1048 }
1049 case AArch64::LDPDpost:
1050 Imm = -Imm;
1051 [[fallthrough]];
1052 case AArch64::STPDpre: {
1053 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1054 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(2).getReg());
1055 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFRegP_X))
1056 .addImm(Reg0)
1057 .addImm(Reg1)
1058 .addImm(Imm * 8)
1059 .setMIFlag(Flag);
1060 break;
1061 }
1062 case AArch64::LDPXpost:
1063 Imm = -Imm;
1064 [[fallthrough]];
1065 case AArch64::STPXpre: {
1066 Register Reg0 = MBBI->getOperand(1).getReg();
1067 Register Reg1 = MBBI->getOperand(2).getReg();
1068 if (Reg0 == AArch64::FP && Reg1 == AArch64::LR)
1069 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFPLR_X))
1070 .addImm(Imm * 8)
1071 .setMIFlag(Flag);
1072 else
1073 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveRegP_X))
1074 .addImm(RegInfo->getSEHRegNum(Reg0))
1075 .addImm(RegInfo->getSEHRegNum(Reg1))
1076 .addImm(Imm * 8)
1077 .setMIFlag(Flag);
1078 break;
1079 }
1080 case AArch64::LDRDpost:
1081 Imm = -Imm;
1082 [[fallthrough]];
1083 case AArch64::STRDpre: {
1084 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1085 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFReg_X))
1086 .addImm(Reg)
1087 .addImm(Imm)
1088 .setMIFlag(Flag);
1089 break;
1090 }
1091 case AArch64::LDRXpost:
1092 Imm = -Imm;
1093 [[fallthrough]];
1094 case AArch64::STRXpre: {
1095 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1096 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveReg_X))
1097 .addImm(Reg)
1098 .addImm(Imm)
1099 .setMIFlag(Flag);
1100 break;
1101 }
1102 case AArch64::STPDi:
1103 case AArch64::LDPDi: {
1104 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1105 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1106 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFRegP))
1107 .addImm(Reg0)
1108 .addImm(Reg1)
1109 .addImm(Imm * 8)
1110 .setMIFlag(Flag);
1111 break;
1112 }
1113 case AArch64::STPXi:
1114 case AArch64::LDPXi: {
1115 Register Reg0 = MBBI->getOperand(0).getReg();
1116 Register Reg1 = MBBI->getOperand(1).getReg();
1117
1118 int SEHReg0 = RegInfo->getSEHRegNum(Reg0);
1119 int SEHReg1 = RegInfo->getSEHRegNum(Reg1);
1120
1121 if (Reg0 == AArch64::FP && Reg1 == AArch64::LR)
1122 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFPLR))
1123 .addImm(Imm * 8)
1124 .setMIFlag(Flag);
1125 else if (SEHReg0 >= 19 && SEHReg1 >= 19)
1126 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveRegP))
1127 .addImm(SEHReg0)
1128 .addImm(SEHReg1)
1129 .addImm(Imm * 8)
1130 .setMIFlag(Flag);
1131 else
1132 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegIP))
1133 .addImm(SEHReg0)
1134 .addImm(SEHReg1)
1135 .addImm(Imm * 8)
1136 .setMIFlag(Flag);
1137 break;
1138 }
1139 case AArch64::STRXui:
1140 case AArch64::LDRXui: {
1141 int Reg = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1142 if (Reg >= 19)
1143 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveReg))
1144 .addImm(Reg)
1145 .addImm(Imm * 8)
1146 .setMIFlag(Flag);
1147 else
1148 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegI))
1149 .addImm(Reg)
1150 .addImm(Imm * 8)
1151 .setMIFlag(Flag);
1152 break;
1153 }
1154 case AArch64::STRDui:
1155 case AArch64::LDRDui: {
1156 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1157 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFReg))
1158 .addImm(Reg)
1159 .addImm(Imm * 8)
1160 .setMIFlag(Flag);
1161 break;
1162 }
1163 case AArch64::STPQi:
1164 case AArch64::LDPQi: {
1165 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1166 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1167 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegQP))
1168 .addImm(Reg0)
1169 .addImm(Reg1)
1170 .addImm(Imm * 16)
1171 .setMIFlag(Flag);
1172 break;
1173 }
1174 case AArch64::LDPQpost:
1175 Imm = -Imm;
1176 [[fallthrough]];
1177 case AArch64::STPQpre: {
1178 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1179 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(2).getReg());
1180 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegQPX))
1181 .addImm(Reg0)
1182 .addImm(Reg1)
1183 .addImm(Imm * 16)
1184 .setMIFlag(Flag);
1185 break;
1186 }
1187 }
1188 auto I = MBB->insertAfter(MBBI, MIB);
1189 return I;
1190}
1191
1194 if (!AFI->needsDwarfUnwindInfo(MF) || !AFI->hasStreamingModeChanges())
1195 return false;
1196 // For Darwin platforms we don't save VG for non-SVE functions, even if SME
1197 // is enabled with streaming mode changes.
1198 auto &ST = MF.getSubtarget<AArch64Subtarget>();
1199 if (ST.isTargetDarwin())
1200 return ST.hasSVE();
1201 return true;
1202}
1203
1205 MachineFunction &MF) const {
1206 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1207 const AArch64InstrInfo *TII = Subtarget.getInstrInfo();
1208
1209 auto EmitSignRA = [&](MachineBasicBlock &MBB) {
1210 DebugLoc DL; // Set debug location to unknown.
1212
1213 BuildMI(MBB, MBBI, DL, TII->get(AArch64::PAUTH_PROLOGUE))
1215 };
1216
1217 auto EmitAuthRA = [&](MachineBasicBlock &MBB) {
1218 DebugLoc DL;
1219 MachineBasicBlock::iterator MBBI = MBB.getFirstTerminator();
1220 if (MBBI != MBB.end())
1221 DL = MBBI->getDebugLoc();
1222
1223 TII->createPauthEpilogueInstr(MBB, DL);
1224 };
1225
1226 // This should be in sync with PEIImpl::calculateSaveRestoreBlocks.
1227 EmitSignRA(MF.front());
1228 for (MachineBasicBlock &MBB : MF) {
1229 if (MBB.isEHFuncletEntry())
1230 EmitSignRA(MBB);
1231 if (MBB.isReturnBlock())
1232 EmitAuthRA(MBB);
1233 }
1234}
1235
1237 MachineBasicBlock &MBB) const {
1238 AArch64PrologueEmitter PrologueEmitter(MF, MBB, *this);
1239 PrologueEmitter.emitPrologue();
1240}
1241
1243 MachineBasicBlock &MBB) const {
1244 AArch64EpilogueEmitter EpilogueEmitter(MF, MBB, *this);
1245 EpilogueEmitter.emitEpilogue();
1246}
1247
1250 MF.getInfo<AArch64FunctionInfo>()->needsDwarfUnwindInfo(MF);
1251}
1252
1254 return enableCFIFixup(MF) &&
1255 MF.getInfo<AArch64FunctionInfo>()->needsAsyncDwarfUnwindInfo(MF);
1256}
1257
1258/// getFrameIndexReference - Provide a base+offset reference to an FI slot for
1259/// debug info. It's the same as what we use for resolving the code-gen
1260/// references for now. FIXME: This can go wrong when references are
1261/// SP-relative and simple call frames aren't used.
1264 Register &FrameReg) const {
1266 MF, FI, FrameReg,
1267 /*PreferFP=*/
1268 MF.getFunction().hasFnAttribute(Attribute::SanitizeHWAddress) ||
1269 MF.getFunction().hasFnAttribute(Attribute::SanitizeMemTag),
1270 /*ForSimm=*/false);
1271}
1272
1275 int FI) const {
1276 // This function serves to provide a comparable offset from a single reference
1277 // point (the value of SP at function entry) that can be used for analysis,
1278 // e.g. the stack-frame-layout analysis pass. It is not guaranteed to be
1279 // correct for all objects in the presence of VLA-area objects or dynamic
1280 // stack re-alignment.
1281
1282 const auto &MFI = MF.getFrameInfo();
1283
1284 int64_t ObjectOffset = MFI.getObjectOffset(FI);
1285 StackOffset ZPRStackSize = getZPRStackSize(MF);
1286 StackOffset PPRStackSize = getPPRStackSize(MF);
1287 StackOffset SVEStackSize = ZPRStackSize + PPRStackSize;
1288
1289 // For VLA-area objects, just emit an offset at the end of the stack frame.
1290 // Whilst not quite correct, these objects do live at the end of the frame and
1291 // so it is more useful for analysis for the offset to reflect this.
1292 if (MFI.isVariableSizedObjectIndex(FI)) {
1293 return StackOffset::getFixed(-((int64_t)MFI.getStackSize())) - SVEStackSize;
1294 }
1295
1296 // This is correct in the absence of any SVE stack objects.
1297 if (!SVEStackSize)
1298 return StackOffset::getFixed(ObjectOffset - getOffsetOfLocalArea());
1299
1300 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1301 bool FPAfterSVECalleeSaves = hasSVECalleeSavesAboveFrameRecord(MF);
1302 if (MFI.hasScalableStackID(FI)) {
1303 if (FPAfterSVECalleeSaves &&
1304 -ObjectOffset <= (int64_t)AFI->getSVECalleeSavedStackSize()) {
1305 assert(!AFI->hasSplitSVEObjects() &&
1306 "split-sve-objects not supported with FPAfterSVECalleeSaves");
1307 return StackOffset::getScalable(ObjectOffset);
1308 }
1309 StackOffset AccessOffset{};
1310 // The scalable vectors are below (lower address) the scalable predicates
1311 // with split SVE objects, so we must subtract the size of the predicates.
1312 if (AFI->hasSplitSVEObjects() &&
1313 MFI.getStackID(FI) == TargetStackID::ScalableVector)
1314 AccessOffset = -PPRStackSize;
1315 return AccessOffset +
1316 StackOffset::get(-((int64_t)AFI->getCalleeSavedStackSize()),
1317 ObjectOffset);
1318 }
1319
1320 bool IsFixed = MFI.isFixedObjectIndex(FI);
1321 bool IsCSR =
1322 !IsFixed && ObjectOffset >= -((int)AFI->getCalleeSavedStackSize(MFI));
1323
1324 StackOffset ScalableOffset = {};
1325 if (!IsFixed && !IsCSR) {
1326 ScalableOffset = -SVEStackSize;
1327 } else if (FPAfterSVECalleeSaves && IsCSR) {
1328 ScalableOffset =
1330 }
1331
1332 return StackOffset::getFixed(ObjectOffset) + ScalableOffset;
1333}
1334
1340
1341StackOffset AArch64FrameLowering::getFPOffset(const MachineFunction &MF,
1342 int64_t ObjectOffset) const {
1343 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1344 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1345 const Function &F = MF.getFunction();
1346 bool IsWin64 = Subtarget.isCallingConvWin64(F.getCallingConv(), F.isVarArg());
1347 unsigned FixedObject =
1348 getFixedObjectSize(MF, AFI, IsWin64, /*IsFunclet=*/false);
1349 int64_t CalleeSaveSize = AFI->getCalleeSavedStackSize(MF.getFrameInfo());
1350 int64_t FPAdjust =
1351 CalleeSaveSize - AFI->getCalleeSaveBaseToFrameRecordOffset();
1352 return StackOffset::getFixed(ObjectOffset + FixedObject + FPAdjust);
1353}
1354
1355StackOffset AArch64FrameLowering::getStackOffset(const MachineFunction &MF,
1356 int64_t ObjectOffset) const {
1357 const auto &MFI = MF.getFrameInfo();
1358 return StackOffset::getFixed(ObjectOffset + (int64_t)MFI.getStackSize());
1359}
1360
1361// TODO: This function currently does not work for scalable vectors.
1363 int FI) const {
1364 const AArch64RegisterInfo *RegInfo =
1365 MF.getSubtarget<AArch64Subtarget>().getRegisterInfo();
1366 int ObjectOffset = MF.getFrameInfo().getObjectOffset(FI);
1367 return RegInfo->getLocalAddressRegister(MF) == AArch64::FP
1368 ? getFPOffset(MF, ObjectOffset).getFixed()
1369 : getStackOffset(MF, ObjectOffset).getFixed();
1370}
1371
1373 const MachineFunction &MF, int FI, Register &FrameReg, bool PreferFP,
1374 bool ForSimm) const {
1375 const auto &MFI = MF.getFrameInfo();
1376 int64_t ObjectOffset = MFI.getObjectOffset(FI);
1377 bool isFixed = MFI.isFixedObjectIndex(FI);
1378 auto StackID = static_cast<TargetStackID::Value>(MFI.getStackID(FI));
1379 return resolveFrameOffsetReference(MF, ObjectOffset, isFixed, StackID,
1380 FrameReg, PreferFP, ForSimm);
1381}
1382
1384 const MachineFunction &MF, int64_t ObjectOffset, bool isFixed,
1385 TargetStackID::Value StackID, Register &FrameReg, bool PreferFP,
1386 bool ForSimm) const {
1387 const auto &MFI = MF.getFrameInfo();
1388 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1389 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
1390 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1391
1392 int64_t FPOffset = getFPOffset(MF, ObjectOffset).getFixed();
1393 int64_t Offset = getStackOffset(MF, ObjectOffset).getFixed();
1394 bool isCSR =
1395 !isFixed && ObjectOffset >= -((int)AFI->getCalleeSavedStackSize(MFI));
1396 bool isSVE = MFI.isScalableStackID(StackID);
1397
1398 StackOffset ZPRStackSize = getZPRStackSize(MF);
1399 StackOffset PPRStackSize = getPPRStackSize(MF);
1400 StackOffset SVEStackSize = ZPRStackSize + PPRStackSize;
1401
1402 // Use frame pointer to reference fixed objects. Use it for locals if
1403 // there are VLAs or a dynamically realigned SP (and thus the SP isn't
1404 // reliable as a base). Make sure useFPForScavengingIndex() does the
1405 // right thing for the emergency spill slot.
1406 bool UseFP = false;
1407 if (AFI->hasStackFrame() && !isSVE) {
1408 // We shouldn't prefer using the FP to access fixed-sized stack objects when
1409 // there are scalable (SVE) objects in between the FP and the fixed-sized
1410 // objects.
1411 PreferFP &= !SVEStackSize;
1412
1413 // Note: Keeping the following as multiple 'if' statements rather than
1414 // merging to a single expression for readability.
1415 //
1416 // Argument access should always use the FP.
1417 if (isFixed) {
1418 UseFP = hasFP(MF);
1419 } else if (isCSR && RegInfo->hasStackRealignment(MF)) {
1420 // References to the CSR area must use FP if we're re-aligning the stack
1421 // since the dynamically-sized alignment padding is between the SP/BP and
1422 // the CSR area.
1423 assert(hasFP(MF) && "Re-aligned stack must have frame pointer");
1424 UseFP = true;
1425 } else if (hasFP(MF) && !RegInfo->hasStackRealignment(MF)) {
1426 // If the FPOffset is negative and we're producing a signed immediate, we
1427 // have to keep in mind that the available offset range for negative
1428 // offsets is smaller than for positive ones. If an offset is available
1429 // via the FP and the SP, use whichever is closest.
1430 bool FPOffsetFits = !ForSimm || FPOffset >= -256;
1431 PreferFP |= Offset > -FPOffset && !SVEStackSize;
1432
1433 if (FPOffset >= 0) {
1434 // If the FPOffset is positive, that'll always be best, as the SP/BP
1435 // will be even further away.
1436 UseFP = true;
1437 } else if (MFI.hasVarSizedObjects()) {
1438 // If we have variable sized objects, we can use either FP or BP, as the
1439 // SP offset is unknown. We can use the base pointer if we have one and
1440 // FP is not preferred. If not, we're stuck with using FP.
1441 bool CanUseBP = RegInfo->hasBasePointer(MF);
1442 if (FPOffsetFits && CanUseBP) // Both are ok. Pick the best.
1443 UseFP = PreferFP;
1444 else if (!CanUseBP) // Can't use BP. Forced to use FP.
1445 UseFP = true;
1446 // else we can use BP and FP, but the offset from FP won't fit.
1447 // That will make us scavenge registers which we can probably avoid by
1448 // using BP. If it won't fit for BP either, we'll scavenge anyway.
1449 } else if (MF.hasEHFunclets() && !RegInfo->hasBasePointer(MF)) {
1450 // Funclets access the locals contained in the parent's stack frame
1451 // via the frame pointer, so we have to use the FP in the parent
1452 // function.
1453 (void) Subtarget;
1454 assert(Subtarget.isCallingConvWin64(MF.getFunction().getCallingConv(),
1455 MF.getFunction().isVarArg()) &&
1456 "Funclets should only be present on Win64");
1457 UseFP = true;
1458 } else {
1459 // We have the choice between FP and (SP or BP).
1460 if (FPOffsetFits && PreferFP) // If FP is the best fit, use it.
1461 UseFP = true;
1462 }
1463 }
1464 }
1465
1466 assert(
1467 ((isFixed || isCSR) || !RegInfo->hasStackRealignment(MF) || !UseFP) &&
1468 "In the presence of dynamic stack pointer realignment, "
1469 "non-argument/CSR objects cannot be accessed through the frame pointer");
1470
1471 bool FPAfterSVECalleeSaves = hasSVECalleeSavesAboveFrameRecord(MF);
1472
1473 if (isSVE) {
1474 StackOffset FPOffset = StackOffset::get(
1475 -AFI->getCalleeSaveBaseToFrameRecordOffset(), ObjectOffset);
1476 StackOffset SPOffset =
1477 SVEStackSize +
1478 StackOffset::get(MFI.getStackSize() - AFI->getCalleeSavedStackSize(),
1479 ObjectOffset);
1480
1481 // With split SVE objects the ObjectOffset is relative to the split area
1482 // (i.e. the PPR area or ZPR area respectively).
1483 if (AFI->hasSplitSVEObjects() && StackID == TargetStackID::ScalableVector) {
1484 // If we're accessing an SVE vector with split SVE objects...
1485 // - From the FP we need to move down past the PPR area:
1486 FPOffset -= PPRStackSize;
1487 // - From the SP we only need to move up to the ZPR area:
1488 SPOffset -= PPRStackSize;
1489 // Note: `SPOffset = SVEStackSize + ...`, so `-= PPRStackSize` results in
1490 // `SPOffset = ZPRStackSize + ...`.
1491 }
1492
1493 if (FPAfterSVECalleeSaves) {
1495 if (-ObjectOffset <= (int64_t)AFI->getSVECalleeSavedStackSize()) {
1498 }
1499 }
1500
1501 // Always use the FP for SVE spills if available and beneficial.
1502 if (hasFP(MF) && (SPOffset.getFixed() ||
1503 FPOffset.getScalable() < SPOffset.getScalable() ||
1504 RegInfo->hasStackRealignment(MF))) {
1505 FrameReg = RegInfo->getFrameRegister(MF);
1506 return FPOffset;
1507 }
1508 FrameReg = RegInfo->hasBasePointer(MF) ? RegInfo->getBaseRegister()
1509 : MCRegister(AArch64::SP);
1510
1511 return SPOffset;
1512 }
1513
1514 StackOffset SVEAreaOffset = {};
1515 if (FPAfterSVECalleeSaves) {
1516 // In this stack layout, the FP is in between the callee saves and other
1517 // SVE allocations.
1518 StackOffset SVECalleeSavedStack =
1520 if (UseFP) {
1521 if (isFixed)
1522 SVEAreaOffset = SVECalleeSavedStack;
1523 else if (!isCSR)
1524 SVEAreaOffset = SVECalleeSavedStack - SVEStackSize;
1525 } else {
1526 if (isFixed)
1527 SVEAreaOffset = SVEStackSize;
1528 else if (isCSR)
1529 SVEAreaOffset = SVEStackSize - SVECalleeSavedStack;
1530 }
1531 } else {
1532 if (UseFP && !(isFixed || isCSR))
1533 SVEAreaOffset = -SVEStackSize;
1534 if (!UseFP && (isFixed || isCSR))
1535 SVEAreaOffset = SVEStackSize;
1536 }
1537
1538 if (UseFP) {
1539 FrameReg = RegInfo->getFrameRegister(MF);
1540 return StackOffset::getFixed(FPOffset) + SVEAreaOffset;
1541 }
1542
1543 // Use the base pointer if we have one.
1544 if (RegInfo->hasBasePointer(MF))
1545 FrameReg = RegInfo->getBaseRegister();
1546 else {
1547 assert(!MFI.hasVarSizedObjects() &&
1548 "Can't use SP when we have var sized objects.");
1549 FrameReg = AArch64::SP;
1550 // If we're using the red zone for this function, the SP won't actually
1551 // be adjusted, so the offsets will be negative. They're also all
1552 // within range of the signed 9-bit immediate instructions.
1553 if (canUseRedZone(MF))
1554 Offset -= AFI->getLocalStackSize();
1555 }
1556
1557 return StackOffset::getFixed(Offset) + SVEAreaOffset;
1558}
1559
1561 // Do not set a kill flag on values that are also marked as live-in. This
1562 // happens with the @llvm-returnaddress intrinsic and with arguments passed in
1563 // callee saved registers.
1564 // Omitting the kill flags is conservatively correct even if the live-in
1565 // is not used after all.
1566 bool IsLiveIn = MF.getRegInfo().isLiveIn(Reg);
1567 return getKillRegState(!IsLiveIn);
1568}
1569
1571 MachineFunction &MF) {
1572 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1573 AttributeList Attrs = MF.getFunction().getAttributes();
1575 return Subtarget.isTargetMachO() &&
1576 !(Subtarget.getTargetLowering()->supportSwiftError() &&
1577 Attrs.hasAttrSomewhere(Attribute::SwiftError)) &&
1579 !AFL.requiresSaveVG(MF) && !AFI->isSVECC();
1580}
1581
1582static bool invalidateWindowsRegisterPairing(bool SpillExtendedVolatile,
1583 unsigned SpillCount, unsigned Reg1,
1584 unsigned Reg2, bool NeedsWinCFI,
1585 const TargetRegisterInfo *TRI) {
1586 // If we are generating register pairs for a Windows function that requires
1587 // EH support, then pair consecutive registers only. There are no unwind
1588 // opcodes for saves/restores of non-consecutive register pairs.
1589 // The unwind opcodes are save_regp, save_regp_x, save_fregp, save_frepg_x,
1590 // save_lrpair.
1591 // https://docs.microsoft.com/en-us/cpp/build/arm64-exception-handling
1592
1593 if (Reg2 == AArch64::FP)
1594 return true;
1595 if (!NeedsWinCFI)
1596 return false;
1597
1598 // ARM64EC introduced `save_any_regp`, which expects 16-byte alignment.
1599 // This is handled by only allowing paired spills for registers spilled at
1600 // even positions (which should be 16-byte aligned, as other GPRs/FPRs are
1601 // 8-bytes). We carve out an exception for {FP,LR}, which does not require
1602 // 16-byte alignment in the uop representation.
1603 if (TRI->getEncodingValue(Reg2) == TRI->getEncodingValue(Reg1) + 1)
1604 return SpillExtendedVolatile
1605 ? !((Reg1 == AArch64::FP && Reg2 == AArch64::LR) ||
1606 (SpillCount % 2) == 0)
1607 : false;
1608
1609 // If pairing a GPR with LR, the pair can be described by the save_lrpair
1610 // opcode. The save_lrpair opcode requires the first register to be odd.
1611 if (Reg1 >= AArch64::X19 && Reg1 <= AArch64::X27 &&
1612 (Reg1 - AArch64::X19) % 2 == 0 && Reg2 == AArch64::LR)
1613 return false;
1614 return true;
1615}
1616
1617/// Returns true if Reg1 and Reg2 cannot be paired using a ldp/stp instruction.
1618/// WindowsCFI requires that only consecutive registers can be paired.
1619/// LR and FP need to be allocated together when the frame needs to save
1620/// the frame-record. This means any other register pairing with LR is invalid.
1621static bool invalidateRegisterPairing(bool SpillExtendedVolatile,
1622 unsigned SpillCount, unsigned Reg1,
1623 unsigned Reg2, bool UsesWinAAPCS,
1624 bool NeedsWinCFI, bool NeedsFrameRecord,
1625 const TargetRegisterInfo *TRI) {
1626 if (UsesWinAAPCS)
1627 return invalidateWindowsRegisterPairing(SpillExtendedVolatile, SpillCount,
1628 Reg1, Reg2, NeedsWinCFI, TRI);
1629
1630 // If we need to store the frame record, don't pair any register
1631 // with LR other than FP.
1632 if (NeedsFrameRecord)
1633 return Reg2 == AArch64::LR;
1634
1635 return false;
1636}
1637
1638namespace {
1639
1640struct RegPairInfo {
1641 Register Reg1;
1642 Register Reg2;
1643 int FrameIdx;
1644 int Offset;
1645 enum RegType { GPR, FPR64, FPR128, PPR, ZPR, VG } Type;
1646 const TargetRegisterClass *RC;
1647
1648 RegPairInfo() = default;
1649
1650 bool isPaired() const { return Reg2.isValid(); }
1651
1652 bool isScalable() const { return Type == PPR || Type == ZPR; }
1653};
1654
1655} // end anonymous namespace
1656
1658 for (unsigned PReg = AArch64::P8; PReg <= AArch64::P15; ++PReg) {
1659 if (SavedRegs.test(PReg)) {
1660 unsigned PNReg = PReg - AArch64::P0 + AArch64::PN0;
1661 return MCRegister(PNReg);
1662 }
1663 }
1664 return MCRegister();
1665}
1666
1667// The multivector LD/ST are available only for SME or SVE2p1 targets
1669 MachineFunction &MF) {
1671 return false;
1672
1673 SMEAttrs FuncAttrs = MF.getInfo<AArch64FunctionInfo>()->getSMEFnAttrs();
1674 bool IsLocallyStreaming =
1675 FuncAttrs.hasStreamingBody() && !FuncAttrs.hasStreamingInterface();
1676
1677 // Only when in streaming mode SME2 instructions can be safely used.
1678 // It is not safe to use SME2 instructions when in streaming compatible or
1679 // locally streaming mode.
1680 return Subtarget.hasSVE2p1() ||
1681 (Subtarget.hasSME2() &&
1682 (!IsLocallyStreaming && Subtarget.isStreaming()));
1683}
1684
1686 MachineFunction &MF,
1688 const TargetRegisterInfo *TRI,
1690 bool NeedsFrameRecord) {
1691
1692 if (CSI.empty())
1693 return;
1694
1695 bool IsWindows = isTargetWindows(MF);
1697 unsigned StackHazardSize = getStackHazardSize(MF);
1698 MachineFrameInfo &MFI = MF.getFrameInfo();
1700 unsigned Count = CSI.size();
1701 (void)CC;
1702 // MachO's compact unwind format relies on all registers being stored in
1703 // pairs.
1704 assert((!produceCompactUnwindFrame(AFL, MF) ||
1707 (Count & 1) == 0) &&
1708 "Odd number of callee-saved regs to spill!");
1709 int ByteOffset = AFI->getCalleeSavedStackSize();
1710 int StackFillDir = -1;
1711 int RegInc = 1;
1712 unsigned FirstReg = 0;
1713 if (IsWindows) {
1714 // For WinCFI, fill the stack from the bottom up.
1715 ByteOffset = 0;
1716 StackFillDir = 1;
1717 // As the CSI array is reversed to match PrologEpilogInserter, iterate
1718 // backwards, to pair up registers starting from lower numbered registers.
1719 RegInc = -1;
1720 FirstReg = Count - 1;
1721 }
1722
1723 bool FPAfterSVECalleeSaves = AFL.hasSVECalleeSavesAboveFrameRecord(MF);
1724 // Windows AAPCS has x9-x15 as volatile registers, x16-x17 as intra-procedural
1725 // scratch, x18 as platform reserved. However, clang has extended calling
1726 // convensions such as preserve_most and preserve_all which treat these as
1727 // CSR. As such, the ARM64 unwind uOPs bias registers by 19. We use ARM64EC
1728 // uOPs which have separate restrictions. We need to check for that.
1729 //
1730 // NOTE: we currently do not account for the D registers as LLVM does not
1731 // support non-ABI compliant D register spills.
1732 bool SpillExtendedVolatile =
1733 IsWindows && llvm::any_of(CSI, [](const CalleeSavedInfo &CSI) {
1734 const auto &Reg = CSI.getReg();
1735 return Reg >= AArch64::X0 && Reg <= AArch64::X18;
1736 });
1737
1738 int ZPRByteOffset = 0;
1739 int PPRByteOffset = 0;
1740 bool SplitPPRs = AFI->hasSplitSVEObjects();
1741 if (SplitPPRs) {
1742 ZPRByteOffset = AFI->getZPRCalleeSavedStackSize();
1743 PPRByteOffset = AFI->getPPRCalleeSavedStackSize();
1744 } else if (!FPAfterSVECalleeSaves) {
1745 ZPRByteOffset =
1747 // Unused: Everything goes in ZPR space.
1748 PPRByteOffset = 0;
1749 }
1750
1751 bool NeedGapToAlignStack = AFI->hasCalleeSaveStackFreeSpace();
1752 Register LastReg = 0;
1753 bool HasCSHazardPadding = AFI->hasStackHazardSlotIndex() && !SplitPPRs;
1754
1755 auto AlignOffset = [StackFillDir](int Offset, int Align) {
1756 if (StackFillDir < 0)
1757 return alignDown(Offset, Align);
1758 return alignTo(Offset, Align);
1759 };
1760
1761 // When iterating backwards, the loop condition relies on unsigned wraparound.
1762 for (unsigned i = FirstReg; i < Count; i += RegInc) {
1763 RegPairInfo RPI;
1764 RPI.Reg1 = CSI[i].getReg();
1765
1766 if (AArch64::GPR64RegClass.contains(RPI.Reg1)) {
1767 RPI.Type = RegPairInfo::GPR;
1768 RPI.RC = &AArch64::GPR64RegClass;
1769 } else if (AArch64::FPR64RegClass.contains(RPI.Reg1)) {
1770 RPI.Type = RegPairInfo::FPR64;
1771 RPI.RC = &AArch64::FPR64RegClass;
1772 } else if (AArch64::FPR128RegClass.contains(RPI.Reg1)) {
1773 RPI.Type = RegPairInfo::FPR128;
1774 RPI.RC = &AArch64::FPR128RegClass;
1775 } else if (AArch64::ZPRRegClass.contains(RPI.Reg1)) {
1776 RPI.Type = RegPairInfo::ZPR;
1777 RPI.RC = &AArch64::ZPRRegClass;
1778 } else if (AArch64::PPRRegClass.contains(RPI.Reg1)) {
1779 RPI.Type = RegPairInfo::PPR;
1780 RPI.RC = &AArch64::PPRRegClass;
1781 } else if (RPI.Reg1 == AArch64::VG) {
1782 RPI.Type = RegPairInfo::VG;
1783 RPI.RC = &AArch64::FIXED_REGSRegClass;
1784 } else {
1785 llvm_unreachable("Unsupported register class.");
1786 }
1787
1788 int &ScalableByteOffset = RPI.Type == RegPairInfo::PPR && SplitPPRs
1789 ? PPRByteOffset
1790 : ZPRByteOffset;
1791
1792 // Add the stack hazard size as we transition from GPR->FPR CSRs.
1793 if (HasCSHazardPadding &&
1794 (!LastReg || !AArch64InstrInfo::isFpOrNEON(LastReg)) &&
1796 ByteOffset += StackFillDir * StackHazardSize;
1797 LastReg = RPI.Reg1;
1798
1799 bool NeedsWinCFI = AFL.needsWinCFI(MF);
1800 int Scale = TRI->getSpillSize(*RPI.RC);
1801 // Add the next reg to the pair if it is in the same register class.
1802 if (unsigned(i + RegInc) < Count && !HasCSHazardPadding) {
1803 MCRegister NextReg = CSI[i + RegInc].getReg();
1804 unsigned SpillCount = NeedsWinCFI ? FirstReg - i : i;
1805 int Aligned = AlignOffset(ByteOffset, Scale);
1806 int PairOffset = IsWindows ? Aligned : Aligned + StackFillDir * 2 * Scale;
1807 bool PairFitsImmRange =
1808 PairOffset / Scale >= -64 && PairOffset / Scale <= 63;
1809 switch (RPI.Type) {
1810 case RegPairInfo::GPR:
1811 if (AArch64::GPR64RegClass.contains(NextReg) && PairFitsImmRange &&
1812 !invalidateRegisterPairing(SpillExtendedVolatile, SpillCount,
1813 RPI.Reg1, NextReg, IsWindows,
1814 NeedsWinCFI, NeedsFrameRecord, TRI))
1815 RPI.Reg2 = NextReg;
1816 break;
1817 case RegPairInfo::FPR64:
1818 if (AArch64::FPR64RegClass.contains(NextReg) && PairFitsImmRange &&
1819 !invalidateRegisterPairing(SpillExtendedVolatile, SpillCount,
1820 RPI.Reg1, NextReg, IsWindows,
1821 NeedsWinCFI, NeedsFrameRecord, TRI))
1822 RPI.Reg2 = NextReg;
1823 break;
1824 case RegPairInfo::FPR128:
1825 if (AArch64::FPR128RegClass.contains(NextReg) && PairFitsImmRange)
1826 RPI.Reg2 = NextReg;
1827 break;
1828 case RegPairInfo::PPR:
1829 break;
1830 case RegPairInfo::ZPR:
1831 if (AFI->getPredicateRegForFillSpill() != 0 &&
1832 ((RPI.Reg1 - AArch64::Z0) & 1) == 0 && (NextReg == RPI.Reg1 + 1)) {
1833 // Calculate offset of register pair to see if pair instruction can be
1834 // used.
1835 int Offset = (ScalableByteOffset + StackFillDir * 2 * Scale) / Scale;
1836 if ((-16 <= Offset && Offset <= 14) && (Offset % 2 == 0))
1837 RPI.Reg2 = NextReg;
1838 }
1839 break;
1840 case RegPairInfo::VG:
1841 break;
1842 }
1843 }
1844
1845 // GPRs and FPRs are saved in pairs of 64-bit regs. We expect the CSI
1846 // list to come in sorted by frame index so that we can issue the store
1847 // pair instructions directly. Assert if we see anything otherwise.
1848 //
1849 // The order of the registers in the list is controlled by
1850 // getCalleeSavedRegs(), so they will always be in-order, as well.
1851 assert((!RPI.isPaired() ||
1852 (CSI[i].getFrameIdx() + RegInc == CSI[i + RegInc].getFrameIdx())) &&
1853 "Out of order callee saved regs!");
1854
1855 assert((!RPI.isPaired() || !NeedsFrameRecord || RPI.Reg2 != AArch64::FP ||
1856 RPI.Reg1 == AArch64::LR) &&
1857 "FrameRecord must be allocated together with LR");
1858
1859 // Windows AAPCS has FP and LR reversed.
1860 assert((!RPI.isPaired() || !NeedsFrameRecord || RPI.Reg1 != AArch64::FP ||
1861 RPI.Reg2 == AArch64::LR) &&
1862 "FrameRecord must be allocated together with LR");
1863
1864 // MachO's compact unwind format relies on all registers being stored in
1865 // adjacent register pairs.
1866 assert((!produceCompactUnwindFrame(AFL, MF) ||
1869 (RPI.isPaired() &&
1870 ((RPI.Reg1 == AArch64::LR && RPI.Reg2 == AArch64::FP) ||
1871 RPI.Reg1 + 1 == RPI.Reg2))) &&
1872 "Callee-save registers not saved as adjacent register pair!");
1873
1874 RPI.FrameIdx = CSI[i].getFrameIdx();
1875 if (IsWindows &&
1876 RPI.isPaired()) // RPI.FrameIdx must be the lower index of the pair
1877 RPI.FrameIdx = CSI[i + RegInc].getFrameIdx();
1878
1879 // Realign the scalable offset if necessary. This is relevant when spilling
1880 // predicates on Windows.
1881 if (RPI.isScalable() && ScalableByteOffset % Scale != 0)
1882 ScalableByteOffset = AlignOffset(ScalableByteOffset, Scale);
1883
1884 // Realign the fixed offset if necessary. This is relevant when spilling Q
1885 // registers after spilling an odd amount of X registers.
1886 if (!RPI.isScalable() && ByteOffset % Scale != 0)
1887 ByteOffset = AlignOffset(ByteOffset, Scale);
1888
1889 int OffsetPre = RPI.isScalable() ? ScalableByteOffset : ByteOffset;
1890 assert(OffsetPre % Scale == 0);
1891
1892 if (RPI.isScalable())
1893 ScalableByteOffset += StackFillDir * (RPI.isPaired() ? 2 * Scale : Scale);
1894 else
1895 ByteOffset += StackFillDir * (RPI.isPaired() ? 2 * Scale : Scale);
1896
1897 // Swift's async context is directly before FP, so allocate an extra
1898 // 8 bytes for it.
1899 if (NeedsFrameRecord && AFI->hasSwiftAsyncContext() &&
1900 ((!IsWindows && RPI.Reg2 == AArch64::FP) ||
1901 (IsWindows && RPI.Reg2 == AArch64::LR)))
1902 ByteOffset += StackFillDir * 8;
1903
1904 // Round up size of non-pair to pair size if we need to pad the
1905 // callee-save area to ensure 16-byte alignment.
1906 if (NeedGapToAlignStack && !IsWindows && !RPI.isScalable() &&
1907 RPI.Type != RegPairInfo::FPR128 && !RPI.isPaired() &&
1908 ByteOffset % 16 != 0) {
1909 ByteOffset += 8 * StackFillDir;
1910 assert(MFI.getObjectAlign(RPI.FrameIdx) <= Align(16));
1911 // A stack frame with a gap looks like this, bottom up:
1912 // d9, d8. x21, gap, x20, x19.
1913 // Set extra alignment on the x21 object to create the gap above it.
1914 MFI.setObjectAlignment(RPI.FrameIdx, Align(16));
1915 NeedGapToAlignStack = false;
1916 }
1917
1918 int OffsetPost = RPI.isScalable() ? ScalableByteOffset : ByteOffset;
1919 assert(OffsetPost % Scale == 0);
1920 // If filling top down (default), we want the offset after incrementing it.
1921 // If filling bottom up (WinCFI) we need the original offset.
1922 int Offset = IsWindows ? OffsetPre : OffsetPost;
1923
1924 // The FP, LR pair goes 8 bytes into our expanded 24-byte slot so that the
1925 // Swift context can directly precede FP.
1926 if (NeedsFrameRecord && AFI->hasSwiftAsyncContext() &&
1927 ((!IsWindows && RPI.Reg2 == AArch64::FP) ||
1928 (IsWindows && RPI.Reg2 == AArch64::LR)))
1929 Offset += 8;
1930 RPI.Offset = Offset / Scale;
1931
1932 assert((!RPI.isPaired() ||
1933 (!RPI.isScalable() && RPI.Offset >= -64 && RPI.Offset <= 63) ||
1934 (RPI.isScalable() && RPI.Offset >= -256 && RPI.Offset <= 255)) &&
1935 "Offset out of bounds for LDP/STP immediate");
1936
1937 auto isFrameRecord = [&] {
1938 if (RPI.isPaired())
1939 return IsWindows ? RPI.Reg1 == AArch64::FP && RPI.Reg2 == AArch64::LR
1940 : RPI.Reg1 == AArch64::LR && RPI.Reg2 == AArch64::FP;
1941 // Otherwise, look for the frame record as two unpaired registers. This is
1942 // needed for -aarch64-stack-hazard-size=<val>, which disables register
1943 // pairing (as the padding may be too large for the LDP/STP offset). Note:
1944 // On Windows, this check works out as current reg == FP, next reg == LR,
1945 // and on other platforms current reg == FP, previous reg == LR. This
1946 // works out as the correct pre-increment or post-increment offsets
1947 // respectively.
1948 return i > 0 && RPI.Reg1 == AArch64::FP &&
1949 CSI[i - 1].getReg() == AArch64::LR;
1950 };
1951
1952 // Save the offset to frame record so that the FP register can point to the
1953 // innermost frame record (spilled FP and LR registers).
1954 if (NeedsFrameRecord && isFrameRecord())
1956
1957 RegPairs.push_back(RPI);
1958 if (RPI.isPaired())
1959 i += RegInc;
1960 }
1961 if (IsWindows) {
1962 // If we need an alignment gap in the stack, align the topmost stack
1963 // object. A stack frame with a gap looks like this, bottom up:
1964 // x19, d8. d9, gap.
1965 // Set extra alignment on the topmost stack object (the first element in
1966 // CSI, which goes top down), to create the gap above it.
1967 if (AFI->hasCalleeSaveStackFreeSpace())
1968 MFI.setObjectAlignment(CSI[0].getFrameIdx(), Align(16));
1969 // We iterated bottom up over the registers; flip RegPairs back to top
1970 // down order.
1971 std::reverse(RegPairs.begin(), RegPairs.end());
1972 }
1973}
1974
1978 MachineFunction &MF = *MBB.getParent();
1979 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1980 auto &TLI = *Subtarget.getTargetLowering();
1981 const AArch64InstrInfo &TII = *Subtarget.getInstrInfo();
1982 bool NeedsWinCFI = needsWinCFI(MF);
1983 DebugLoc DL;
1985
1986 computeCalleeSaveRegisterPairs(*this, MF, CSI, TRI, RegPairs, hasFP(MF));
1987
1988 MachineRegisterInfo &MRI = MF.getRegInfo();
1989 // Refresh the reserved regs in case there are any potential changes since the
1990 // last freeze.
1991 MRI.freezeReservedRegs();
1992
1993 if (homogeneousPrologEpilog(MF)) {
1994 auto MIB = BuildMI(MBB, MI, DL, TII.get(AArch64::HOM_Prolog))
1996
1997 for (auto &RPI : RegPairs) {
1998 MIB.addReg(RPI.Reg1);
1999 MIB.addReg(RPI.Reg2);
2000
2001 // Update register live in.
2002 if (!MRI.isReserved(RPI.Reg1))
2003 MBB.addLiveIn(RPI.Reg1);
2004 if (RPI.isPaired() && !MRI.isReserved(RPI.Reg2))
2005 MBB.addLiveIn(RPI.Reg2);
2006 }
2007 return true;
2008 }
2009 bool PTrueCreated = false;
2010 for (const RegPairInfo &RPI : llvm::reverse(RegPairs)) {
2011 Register Reg1 = RPI.Reg1;
2012 Register Reg2 = RPI.Reg2;
2013 unsigned StrOpc;
2014
2015 // Issue sequence of spills for cs regs. The first spill may be converted
2016 // to a pre-decrement store later by emitPrologue if the callee-save stack
2017 // area allocation can't be combined with the local stack area allocation.
2018 // For example:
2019 // stp x22, x21, [sp, #0] // addImm(+0)
2020 // stp x20, x19, [sp, #16] // addImm(+2)
2021 // stp fp, lr, [sp, #32] // addImm(+4)
2022 // Rationale: This sequence saves uop updates compared to a sequence of
2023 // pre-increment spills like stp xi,xj,[sp,#-16]!
2024 // Note: Similar rationale and sequence for restores in epilog.
2025 unsigned Size = TRI->getSpillSize(*RPI.RC);
2026 Align Alignment = TRI->getSpillAlign(*RPI.RC);
2027 switch (RPI.Type) {
2028 case RegPairInfo::GPR:
2029 StrOpc = RPI.isPaired() ? AArch64::STPXi : AArch64::STRXui;
2030 break;
2031 case RegPairInfo::FPR64:
2032 StrOpc = RPI.isPaired() ? AArch64::STPDi : AArch64::STRDui;
2033 break;
2034 case RegPairInfo::FPR128:
2035 StrOpc = RPI.isPaired() ? AArch64::STPQi : AArch64::STRQui;
2036 break;
2037 case RegPairInfo::ZPR:
2038 StrOpc = RPI.isPaired() ? AArch64::ST1B_2Z_IMM : AArch64::STR_ZXI;
2039 break;
2040 case RegPairInfo::PPR:
2041 StrOpc = AArch64::STR_PXI;
2042 break;
2043 case RegPairInfo::VG:
2044 StrOpc = AArch64::STRXui;
2045 break;
2046 }
2047
2048 Register X0Scratch;
2049 llvm::scope_exit RestoreX0([&] {
2050 if (X0Scratch != AArch64::NoRegister)
2051 BuildMI(MBB, MI, DL, TII.get(TargetOpcode::COPY), AArch64::X0)
2052 .addReg(X0Scratch)
2054 });
2055
2056 if (Reg1 == AArch64::VG) {
2057 // Find an available register to store value of VG to.
2058 Reg1 = findScratchNonCalleeSaveRegister(&MBB, true);
2059 assert(Reg1 != AArch64::NoRegister);
2060 if (MF.getSubtarget<AArch64Subtarget>().hasSVE()) {
2061 BuildMI(MBB, MI, DL, TII.get(AArch64::CNTD_XPiI), Reg1)
2062 .addImm(31)
2063 .addImm(1)
2065 } else {
2067 if (any_of(MBB.liveins(),
2068 [&STI](const MachineBasicBlock::RegisterMaskPair &LiveIn) {
2069 return STI.getRegisterInfo()->isSuperOrSubRegisterEq(
2070 AArch64::X0, LiveIn.PhysReg);
2071 })) {
2072 X0Scratch = Reg1;
2073 BuildMI(MBB, MI, DL, TII.get(TargetOpcode::COPY), X0Scratch)
2074 .addReg(AArch64::X0)
2076 }
2077
2078 RTLIB::Libcall LC = RTLIB::SMEABI_GET_CURRENT_VG;
2079 const uint32_t *RegMask =
2080 TRI->getCallPreservedMask(MF, TLI.getLibcallCallingConv(LC));
2081 BuildMI(MBB, MI, DL, TII.get(AArch64::BL))
2082 .addExternalSymbol(TLI.getLibcallName(LC))
2083 .addRegMask(RegMask)
2084 .addReg(AArch64::X0, RegState::ImplicitDefine)
2086 Reg1 = AArch64::X0;
2087 }
2088 }
2089
2090 LLVM_DEBUG({
2091 dbgs() << "CSR spill: (" << printReg(Reg1, TRI);
2092 if (RPI.isPaired())
2093 dbgs() << ", " << printReg(Reg2, TRI);
2094 dbgs() << ") -> fi#(" << RPI.FrameIdx;
2095 if (RPI.isPaired())
2096 dbgs() << ", " << RPI.FrameIdx + 1;
2097 dbgs() << ")\n";
2098 });
2099
2100 assert((!isTargetWindows(MF) ||
2101 !(Reg1 == AArch64::LR && Reg2 == AArch64::FP)) &&
2102 "Windows unwdinding requires a consecutive (FP,LR) pair");
2103 // Windows unwind codes require consecutive registers if registers are
2104 // paired. Make the switch here, so that the code below will save (x,x+1)
2105 // and not (x+1,x).
2106 unsigned FrameIdxReg1 = RPI.FrameIdx;
2107 unsigned FrameIdxReg2 = RPI.FrameIdx + 1;
2108 if (isTargetWindows(MF) && RPI.isPaired()) {
2109 std::swap(Reg1, Reg2);
2110 std::swap(FrameIdxReg1, FrameIdxReg2);
2111 }
2112
2113 if (RPI.isPaired() && RPI.isScalable()) {
2114 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2117 unsigned PnReg = AFI->getPredicateRegForFillSpill();
2118 assert((PnReg != 0 && enableMultiVectorSpillFill(Subtarget, MF)) &&
2119 "Expects SVE2.1 or SME2 target and a predicate register");
2120#ifdef EXPENSIVE_CHECKS
2121 auto IsPPR = [](const RegPairInfo &c) {
2122 return c.Reg1 == RegPairInfo::PPR;
2123 };
2124 auto PPRBegin = std::find_if(RegPairs.begin(), RegPairs.end(), IsPPR);
2125 auto IsZPR = [](const RegPairInfo &c) {
2126 return c.Type == RegPairInfo::ZPR;
2127 };
2128 auto ZPRBegin = std::find_if(RegPairs.begin(), RegPairs.end(), IsZPR);
2129 assert(!(PPRBegin < ZPRBegin) &&
2130 "Expected callee save predicate to be handled first");
2131#endif
2132 if (!PTrueCreated) {
2133 PTrueCreated = true;
2134 BuildMI(MBB, MI, DL, TII.get(AArch64::PTRUE_C_B), PnReg)
2136 }
2137 MachineInstrBuilder MIB = BuildMI(MBB, MI, DL, TII.get(StrOpc));
2138 if (!MRI.isReserved(Reg1))
2139 MBB.addLiveIn(Reg1);
2140 if (!MRI.isReserved(Reg2))
2141 MBB.addLiveIn(Reg2);
2142 MIB.addReg(/*PairRegs*/ AArch64::Z0_Z1 + (RPI.Reg1 - AArch64::Z0));
2144 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2145 MachineMemOperand::MOStore, Size, Alignment));
2146 MIB.addReg(PnReg);
2147 MIB.addReg(AArch64::SP)
2148 .addImm(RPI.Offset / 2) // [sp, #imm*2*vscale],
2149 // where 2*vscale is implicit
2152 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2153 MachineMemOperand::MOStore, Size, Alignment));
2154 if (NeedsWinCFI)
2155 insertSEH(MIB, TII, MachineInstr::FrameSetup);
2156 } else { // The code when the pair of ZReg is not present
2157 MachineInstrBuilder MIB = BuildMI(MBB, MI, DL, TII.get(StrOpc));
2158 if (!MRI.isReserved(Reg1))
2159 MBB.addLiveIn(Reg1);
2160 if (RPI.isPaired()) {
2161 if (!MRI.isReserved(Reg2))
2162 MBB.addLiveIn(Reg2);
2163 MIB.addReg(Reg2, getPrologueDeath(MF, Reg2));
2165 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2166 MachineMemOperand::MOStore, Size, Alignment));
2167 }
2168 MIB.addReg(Reg1, getPrologueDeath(MF, Reg1))
2169 .addReg(AArch64::SP)
2170 .addImm(RPI.Offset) // [sp, #offset*vscale],
2171 // where factor*vscale is implicit
2174 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2175 MachineMemOperand::MOStore, Size, Alignment));
2176 if (NeedsWinCFI)
2177 insertSEH(MIB, TII, MachineInstr::FrameSetup);
2178 }
2179 // Update the StackIDs of the SVE stack slots.
2180 MachineFrameInfo &MFI = MF.getFrameInfo();
2181 if (RPI.Type == RegPairInfo::ZPR) {
2182 MFI.setStackID(FrameIdxReg1, TargetStackID::ScalableVector);
2183 if (RPI.isPaired())
2184 MFI.setStackID(FrameIdxReg2, TargetStackID::ScalableVector);
2185 } else if (RPI.Type == RegPairInfo::PPR) {
2187 if (RPI.isPaired())
2189 }
2190 }
2191 return true;
2192}
2193
2197 MachineFunction &MF = *MBB.getParent();
2198 const AArch64InstrInfo &TII =
2199 *MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
2200 DebugLoc DL;
2202 bool NeedsWinCFI = needsWinCFI(MF);
2203
2204 if (MBBI != MBB.end())
2205 DL = MBBI->getDebugLoc();
2206
2207 computeCalleeSaveRegisterPairs(*this, MF, CSI, TRI, RegPairs, hasFP(MF));
2208 if (homogeneousPrologEpilog(MF, &MBB)) {
2209 auto MIB = BuildMI(MBB, MBBI, DL, TII.get(AArch64::HOM_Epilog))
2211 for (auto &RPI : RegPairs) {
2212 MIB.addReg(RPI.Reg1, RegState::Define);
2213 MIB.addReg(RPI.Reg2, RegState::Define);
2214 }
2215 return true;
2216 }
2217
2218 // For performance reasons restore SVE register in increasing order
2219 auto IsPPR = [](const RegPairInfo &c) { return c.Type == RegPairInfo::PPR; };
2220 auto PPRBegin = llvm::find_if(RegPairs, IsPPR);
2221 auto PPREnd = std::find_if_not(PPRBegin, RegPairs.end(), IsPPR);
2222 std::reverse(PPRBegin, PPREnd);
2223 auto IsZPR = [](const RegPairInfo &c) { return c.Type == RegPairInfo::ZPR; };
2224 auto ZPRBegin = llvm::find_if(RegPairs, IsZPR);
2225 auto ZPREnd = std::find_if_not(ZPRBegin, RegPairs.end(), IsZPR);
2226 std::reverse(ZPRBegin, ZPREnd);
2227
2228 bool PTrueCreated = false;
2229 for (const RegPairInfo &RPI : RegPairs) {
2230 Register Reg1 = RPI.Reg1;
2231 Register Reg2 = RPI.Reg2;
2232
2233 // Issue sequence of restores for cs regs. The last restore may be converted
2234 // to a post-increment load later by emitEpilogue if the callee-save stack
2235 // area allocation can't be combined with the local stack area allocation.
2236 // For example:
2237 // ldp fp, lr, [sp, #32] // addImm(+4)
2238 // ldp x20, x19, [sp, #16] // addImm(+2)
2239 // ldp x22, x21, [sp, #0] // addImm(+0)
2240 // Note: see comment in spillCalleeSavedRegisters()
2241 unsigned LdrOpc;
2242 unsigned Size = TRI->getSpillSize(*RPI.RC);
2243 Align Alignment = TRI->getSpillAlign(*RPI.RC);
2244 switch (RPI.Type) {
2245 case RegPairInfo::GPR:
2246 LdrOpc = RPI.isPaired() ? AArch64::LDPXi : AArch64::LDRXui;
2247 break;
2248 case RegPairInfo::FPR64:
2249 LdrOpc = RPI.isPaired() ? AArch64::LDPDi : AArch64::LDRDui;
2250 break;
2251 case RegPairInfo::FPR128:
2252 LdrOpc = RPI.isPaired() ? AArch64::LDPQi : AArch64::LDRQui;
2253 break;
2254 case RegPairInfo::ZPR:
2255 LdrOpc = RPI.isPaired() ? AArch64::LD1B_2Z_IMM : AArch64::LDR_ZXI;
2256 break;
2257 case RegPairInfo::PPR:
2258 LdrOpc = AArch64::LDR_PXI;
2259 break;
2260 case RegPairInfo::VG:
2261 continue;
2262 }
2263 LLVM_DEBUG({
2264 dbgs() << "CSR restore: (" << printReg(Reg1, TRI);
2265 if (RPI.isPaired())
2266 dbgs() << ", " << printReg(Reg2, TRI);
2267 dbgs() << ") -> fi#(" << RPI.FrameIdx;
2268 if (RPI.isPaired())
2269 dbgs() << ", " << RPI.FrameIdx + 1;
2270 dbgs() << ")\n";
2271 });
2272
2273 // Windows unwind codes require consecutive registers if registers are
2274 // paired. Make the switch here, so that the code below will save (x,x+1)
2275 // and not (x+1,x).
2276 unsigned FrameIdxReg1 = RPI.FrameIdx;
2277 unsigned FrameIdxReg2 = RPI.FrameIdx + 1;
2278 if (isTargetWindows(MF) && RPI.isPaired()) {
2279 std::swap(Reg1, Reg2);
2280 std::swap(FrameIdxReg1, FrameIdxReg2);
2281 }
2282
2284 if (RPI.isPaired() && RPI.isScalable()) {
2285 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2287 unsigned PnReg = AFI->getPredicateRegForFillSpill();
2288 assert((PnReg != 0 && enableMultiVectorSpillFill(Subtarget, MF)) &&
2289 "Expects SVE2.1 or SME2 target and a predicate register");
2290#ifdef EXPENSIVE_CHECKS
2291 assert(!(PPRBegin < ZPRBegin) &&
2292 "Expected callee save predicate to be handled first");
2293#endif
2294 if (!PTrueCreated) {
2295 PTrueCreated = true;
2296 BuildMI(MBB, MBBI, DL, TII.get(AArch64::PTRUE_C_B), PnReg)
2298 }
2299 MachineInstrBuilder MIB = BuildMI(MBB, MBBI, DL, TII.get(LdrOpc));
2300 MIB.addReg(/*PairRegs*/ AArch64::Z0_Z1 + (RPI.Reg1 - AArch64::Z0),
2301 getDefRegState(true));
2303 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2304 MachineMemOperand::MOLoad, Size, Alignment));
2305 MIB.addReg(PnReg);
2306 MIB.addReg(AArch64::SP)
2307 .addImm(RPI.Offset / 2) // [sp, #imm*2*vscale]
2308 // where 2*vscale is implicit
2311 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2312 MachineMemOperand::MOLoad, Size, Alignment));
2313 if (NeedsWinCFI)
2314 insertSEH(MIB, TII, MachineInstr::FrameDestroy);
2315 } else {
2316 MachineInstrBuilder MIB = BuildMI(MBB, MBBI, DL, TII.get(LdrOpc));
2317 if (RPI.isPaired()) {
2318 MIB.addReg(Reg2, getDefRegState(true));
2320 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2321 MachineMemOperand::MOLoad, Size, Alignment));
2322 }
2323 MIB.addReg(Reg1, getDefRegState(true));
2324 MIB.addReg(AArch64::SP)
2325 .addImm(RPI.Offset) // [sp, #offset*vscale]
2326 // where factor*vscale is implicit
2329 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2330 MachineMemOperand::MOLoad, Size, Alignment));
2331 if (NeedsWinCFI)
2332 insertSEH(MIB, TII, MachineInstr::FrameDestroy);
2333 }
2334 }
2335 return true;
2336}
2337
2338// Return the FrameID for a MMO.
2339static std::optional<int> getMMOFrameID(MachineMemOperand *MMO,
2340 const MachineFrameInfo &MFI) {
2341 auto *PSV =
2343 if (PSV)
2344 return std::optional<int>(PSV->getFrameIndex());
2345
2346 if (MMO->getValue()) {
2347 if (auto *Al = dyn_cast<AllocaInst>(getUnderlyingObject(MMO->getValue()))) {
2348 for (int FI = MFI.getObjectIndexBegin(); FI < MFI.getObjectIndexEnd();
2349 FI++)
2350 if (MFI.getObjectAllocation(FI) == Al)
2351 return FI;
2352 }
2353 }
2354
2355 return std::nullopt;
2356}
2357
2358// Return the FrameID for a Load/Store instruction by looking at the first MMO.
2359static std::optional<int> getLdStFrameID(const MachineInstr &MI,
2360 const MachineFrameInfo &MFI) {
2361 if (!MI.mayLoadOrStore() || MI.getNumMemOperands() < 1)
2362 return std::nullopt;
2363
2364 return getMMOFrameID(*MI.memoperands_begin(), MFI);
2365}
2366
2367// Returns true if the LDST MachineInstr \p MI is a PPR access.
2368static bool isPPRAccess(const MachineInstr &MI) {
2369 return AArch64::PPRRegClass.contains(MI.getOperand(0).getReg());
2370}
2371
2372// Check if a Hazard slot is needed for the current function, and if so create
2373// one for it. The index is stored in AArch64FunctionInfo->StackHazardSlotIndex,
2374// which can be used to determine if any hazard padding is needed.
2375void AArch64FrameLowering::determineStackHazardSlot(
2376 MachineFunction &MF, BitVector &SavedRegs) const {
2377 unsigned StackHazardSize = getStackHazardSize(MF);
2378 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2379 if (StackHazardSize == 0 || StackHazardSize % 16 != 0 ||
2381 return;
2382
2383 // Stack hazards are only needed in streaming functions.
2384 SMEAttrs Attrs = AFI->getSMEFnAttrs();
2385 if (!StackHazardInNonStreaming && Attrs.hasNonStreamingInterfaceAndBody())
2386 return;
2387
2388 MachineFrameInfo &MFI = MF.getFrameInfo();
2389
2390 // Add a hazard slot if there are any CSR FPR registers, or are any fp-only
2391 // stack objects.
2392 bool HasFPRCSRs = any_of(SavedRegs.set_bits(), [](unsigned Reg) {
2393 return AArch64::FPR64RegClass.contains(Reg) ||
2394 AArch64::FPR128RegClass.contains(Reg) ||
2395 AArch64::ZPRRegClass.contains(Reg);
2396 });
2397 bool HasPPRCSRs = any_of(SavedRegs.set_bits(), [](unsigned Reg) {
2398 return AArch64::PPRRegClass.contains(Reg);
2399 });
2400 bool HasFPRStackObjects = false;
2401 bool HasPPRStackObjects = false;
2402 if (!HasFPRCSRs || SplitSVEObjects) {
2403 enum SlotType : uint8_t {
2404 Unknown = 0,
2405 ZPRorFPR = 1 << 0,
2406 PPR = 1 << 1,
2407 GPR = 1 << 2,
2409 };
2410
2411 // Find stack slots solely used for one kind of register (ZPR, PPR, etc.),
2412 // based on the kinds of accesses used in the function.
2413 SmallVector<SlotType> SlotTypes(MFI.getObjectIndexEnd(), SlotType::Unknown);
2414 for (auto &MBB : MF) {
2415 for (auto &MI : MBB) {
2416 std::optional<int> FI = getLdStFrameID(MI, MFI);
2417 if (!FI || FI < 0 || FI > int(SlotTypes.size()))
2418 continue;
2419 if (MFI.hasScalableStackID(*FI)) {
2420 SlotTypes[*FI] |=
2421 isPPRAccess(MI) ? SlotType::PPR : SlotType::ZPRorFPR;
2422 } else {
2423 SlotTypes[*FI] |= AArch64InstrInfo::isFpOrNEON(MI)
2424 ? SlotType::ZPRorFPR
2425 : SlotType::GPR;
2426 }
2427 }
2428 }
2429
2430 for (int FI = 0; FI < int(SlotTypes.size()); ++FI) {
2431 HasFPRStackObjects |= SlotTypes[FI] == SlotType::ZPRorFPR;
2432 // For SplitSVEObjects remember that this stack slot is a predicate, this
2433 // will be needed later when determining the frame layout.
2434 if (SlotTypes[FI] == SlotType::PPR) {
2436 HasPPRStackObjects = true;
2437 }
2438 }
2439 }
2440
2441 if (HasFPRCSRs || HasFPRStackObjects) {
2442 int ID = MFI.CreateStackObject(StackHazardSize, Align(16), false);
2443 LLVM_DEBUG(dbgs() << "Created Hazard slot at " << ID << " size "
2444 << StackHazardSize << "\n");
2445 AFI->setStackHazardSlotIndex(ID);
2446 }
2447
2448 if (!AFI->hasStackHazardSlotIndex())
2449 return;
2450
2451 if (SplitSVEObjects) {
2452 CallingConv::ID CC = MF.getFunction().getCallingConv();
2453 if (AFI->isSVECC() || CC == CallingConv::AArch64_SVE_VectorCall) {
2454 AFI->setSplitSVEObjects(true);
2455 LLVM_DEBUG(dbgs() << "Using SplitSVEObjects for SVE CC function\n");
2456 return;
2457 }
2458
2459 // We only use SplitSVEObjects in non-SVE CC functions if there's a
2460 // possibility of a stack hazard between PPRs and ZPRs/FPRs.
2461 LLVM_DEBUG(dbgs() << "Determining if SplitSVEObjects should be used in "
2462 "non-SVE CC function...\n");
2463
2464 // If another calling convention is explicitly set FPRs can't be promoted to
2465 // ZPR callee-saves.
2467 LLVM_DEBUG(
2468 dbgs()
2469 << "Calling convention is not supported with SplitSVEObjects\n");
2470 return;
2471 }
2472
2473 if (!HasPPRCSRs && !HasPPRStackObjects) {
2474 LLVM_DEBUG(
2475 dbgs() << "Not using SplitSVEObjects as no PPRs are on the stack\n");
2476 return;
2477 }
2478
2479 if (!HasFPRCSRs && !HasFPRStackObjects) {
2480 LLVM_DEBUG(
2481 dbgs()
2482 << "Not using SplitSVEObjects as no FPRs or ZPRs are on the stack\n");
2483 return;
2484 }
2485
2486 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2487 MF.getSubtarget<AArch64Subtarget>();
2489 "Expected SVE to be available for PPRs");
2490
2491 const TargetRegisterInfo *TRI = MF.getSubtarget().getRegisterInfo();
2492 // With SplitSVEObjects the CS hazard padding is placed between the
2493 // PPRs and ZPRs. If there are any FPR CS there would be a hazard between
2494 // them and the CS GRPs. Avoid this by promoting all FPR CS to ZPRs.
2495 BitVector FPRZRegs(SavedRegs.size());
2496 for (size_t Reg = 0, E = SavedRegs.size(); HasFPRCSRs && Reg < E; ++Reg) {
2497 BitVector::reference RegBit = SavedRegs[Reg];
2498 if (!RegBit)
2499 continue;
2500 unsigned SubRegIdx = 0;
2501 if (AArch64::FPR64RegClass.contains(Reg))
2502 SubRegIdx = AArch64::dsub;
2503 else if (AArch64::FPR128RegClass.contains(Reg))
2504 SubRegIdx = AArch64::zsub;
2505 else
2506 continue;
2507 // Clear the bit for the FPR save.
2508 RegBit = false;
2509 // Mark that we should save the corresponding ZPR.
2510 Register ZReg =
2511 TRI->getMatchingSuperReg(Reg, SubRegIdx, &AArch64::ZPRRegClass);
2512 FPRZRegs.set(ZReg);
2513 }
2514 SavedRegs |= FPRZRegs;
2515
2516 AFI->setSplitSVEObjects(true);
2517 LLVM_DEBUG(dbgs() << "SplitSVEObjects enabled!\n");
2518 }
2519}
2520
2522 BitVector &SavedRegs,
2523 RegScavenger *RS) const {
2524 // All calls are tail calls in GHC calling conv, and functions have no
2525 // prologue/epilogue.
2527 return;
2528
2529 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
2530
2532 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
2534 unsigned UnspilledCSGPR = AArch64::NoRegister;
2535 unsigned UnspilledCSGPRPaired = AArch64::NoRegister;
2536
2537 MachineFrameInfo &MFI = MF.getFrameInfo();
2538 const MCPhysReg *CSRegs = MF.getRegInfo().getCalleeSavedRegs();
2539
2540 MCRegister BasePointerReg =
2541 RegInfo->hasBasePointer(MF) ? RegInfo->getBaseRegister() : MCRegister();
2542
2543 unsigned ExtraCSSpill = 0;
2544 bool HasUnpairedGPR64 = false;
2545 bool HasPairZReg = false;
2546 BitVector UserReservedRegs = RegInfo->getUserReservedRegs(MF);
2547 BitVector ReservedRegs = RegInfo->getReservedRegs(MF);
2548
2549 // Figure out which callee-saved registers to save/restore.
2550 for (unsigned i = 0; CSRegs[i]; ++i) {
2551 const MCRegister Reg = CSRegs[i];
2552
2553 // Add the base pointer register to SavedRegs if it is callee-save.
2554 if (Reg == BasePointerReg)
2555 SavedRegs.set(Reg);
2556
2557 // Don't save manually reserved registers set through +reserve-x#i,
2558 // even for callee-saved registers, as per GCC's behavior.
2559 if (UserReservedRegs[Reg]) {
2560 SavedRegs.reset(Reg);
2561 continue;
2562 }
2563
2564 bool RegUsed = SavedRegs.test(Reg);
2565 MCRegister PairedReg;
2566 const bool RegIsGPR64 = AArch64::GPR64RegClass.contains(Reg);
2567 if (RegIsGPR64 || AArch64::FPR64RegClass.contains(Reg) ||
2568 AArch64::FPR128RegClass.contains(Reg)) {
2569 // Compensate for odd numbers of GP CSRs.
2570 // For now, all the known cases of odd number of CSRs are of GPRs.
2571 if (HasUnpairedGPR64)
2572 PairedReg = CSRegs[i % 2 == 0 ? i - 1 : i + 1];
2573 else
2574 PairedReg = CSRegs[i ^ 1];
2575 }
2576
2577 // If the function requires all the GP registers to save (SavedRegs),
2578 // and there are an odd number of GP CSRs at the same time (CSRegs),
2579 // PairedReg could be in a different register class from Reg, which would
2580 // lead to a FPR (usually D8) accidentally being marked saved.
2581 if (RegIsGPR64 && !AArch64::GPR64RegClass.contains(PairedReg)) {
2582 PairedReg = AArch64::NoRegister;
2583 HasUnpairedGPR64 = true;
2584 }
2585 assert(PairedReg == AArch64::NoRegister ||
2586 AArch64::GPR64RegClass.contains(Reg, PairedReg) ||
2587 AArch64::FPR64RegClass.contains(Reg, PairedReg) ||
2588 AArch64::FPR128RegClass.contains(Reg, PairedReg));
2589
2590 if (!RegUsed) {
2591 if (AArch64::GPR64RegClass.contains(Reg) && !ReservedRegs[Reg]) {
2592 UnspilledCSGPR = Reg;
2593 UnspilledCSGPRPaired = PairedReg;
2594 }
2595 continue;
2596 }
2597
2598 // MachO's compact unwind format relies on all registers being stored in
2599 // pairs.
2600 // FIXME: the usual format is actually better if unwinding isn't needed.
2601 if (producePairRegisters(MF) && PairedReg != AArch64::NoRegister &&
2602 !SavedRegs.test(PairedReg)) {
2603 SavedRegs.set(PairedReg);
2604 if (AArch64::GPR64RegClass.contains(PairedReg) &&
2605 !ReservedRegs[PairedReg])
2606 ExtraCSSpill = PairedReg;
2607 }
2608 // Check if there is a pair of ZRegs, so it can select PReg for spill/fill
2609 HasPairZReg |= (AArch64::ZPRRegClass.contains(Reg, CSRegs[i ^ 1]) &&
2610 SavedRegs.test(CSRegs[i ^ 1]));
2611 }
2612
2613 if (HasPairZReg && enableMultiVectorSpillFill(Subtarget, MF)) {
2615 // Find a suitable predicate register for the multi-vector spill/fill
2616 // instructions.
2617 MCRegister PnReg = findFreePredicateReg(SavedRegs);
2618 if (PnReg.isValid())
2619 AFI->setPredicateRegForFillSpill(PnReg);
2620 // If no free callee-save has been found assign one.
2621 if (!AFI->getPredicateRegForFillSpill() &&
2622 MF.getFunction().getCallingConv() ==
2624 SavedRegs.set(AArch64::P8);
2625 AFI->setPredicateRegForFillSpill(AArch64::PN8);
2626 }
2627
2628 assert(!ReservedRegs[AFI->getPredicateRegForFillSpill()] &&
2629 "Predicate cannot be a reserved register");
2630 }
2631
2633 !Subtarget.isTargetWindows()) {
2634 // For Windows calling convention on a non-windows OS, where X18 is treated
2635 // as reserved, back up X18 when entering non-windows code (marked with the
2636 // Windows calling convention) and restore when returning regardless of
2637 // whether the individual function uses it - it might call other functions
2638 // that clobber it.
2639 SavedRegs.set(AArch64::X18);
2640 }
2641
2642 // Determine if a Hazard slot should be used and where it should go.
2643 // If SplitSVEObjects is used, the hazard padding is placed between the PPRs
2644 // and ZPRs. Otherwise, it goes in the callee save area.
2645 determineStackHazardSlot(MF, SavedRegs);
2646
2647 // Calculates the callee saved stack size.
2648 unsigned CSStackSize = 0;
2649 unsigned ZPRCSStackSize = 0;
2650 unsigned PPRCSStackSize = 0;
2652 for (unsigned Reg : SavedRegs.set_bits()) {
2653 auto *RC = TRI->getMinimalPhysRegClass(MCRegister(Reg));
2654 assert(RC && "expected register class!");
2655 auto SpillSize = TRI->getSpillSize(*RC);
2656 bool IsZPR = AArch64::ZPRRegClass.contains(Reg);
2657 bool IsPPR = !IsZPR && AArch64::PPRRegClass.contains(Reg);
2658 if (IsZPR)
2659 ZPRCSStackSize += SpillSize;
2660 else if (IsPPR)
2661 PPRCSStackSize += SpillSize;
2662 else {
2663 // A register and its super-register can both appear in SavedRegs.
2664 // Only the widest register is actually spilled, so skip such
2665 // sub-registers here to avoid double-counting the overlap.
2666 bool SavedSuper = any_of(TRI->superregs(Reg), [&](MCPhysReg SuperReg) {
2667 return SavedRegs.test(SuperReg);
2668 });
2669 if (!SavedSuper)
2670 CSStackSize += SpillSize;
2671 }
2672 }
2673
2674 // Save number of saved regs, so we can easily update CSStackSize later to
2675 // account for any additional 64-bit GPR saves. Note: After this point
2676 // only 64-bit GPRs can be added to SavedRegs.
2677 unsigned NumSavedRegs = SavedRegs.count();
2678
2679 // If we have hazard padding in the CS area add that to the size.
2681 CSStackSize += getStackHazardSize(MF);
2682
2683 // Increase the callee-saved stack size if the function has streaming mode
2684 // changes, as we will need to spill the value of the VG register.
2685 if (requiresSaveVG(MF))
2686 CSStackSize += 8;
2687
2688 // If we must call __arm_get_current_vg in the prologue preserve the LR.
2689 if (requiresSaveVG(MF) && !Subtarget.hasSVE())
2690 SavedRegs.set(AArch64::LR);
2691
2692 // The frame record needs to be created by saving the appropriate registers
2693 uint64_t EstimatedStackSize = MFI.estimateStackSize(MF);
2694 if (hasFP(MF) ||
2695 windowsRequiresStackProbe(MF, EstimatedStackSize + CSStackSize + 16)) {
2696 SavedRegs.set(AArch64::FP);
2697 SavedRegs.set(AArch64::LR);
2698 }
2699
2700 LLVM_DEBUG({
2701 dbgs() << "*** determineCalleeSaves\nSaved CSRs:";
2702 for (unsigned Reg : SavedRegs.set_bits())
2703 dbgs() << ' ' << printReg(MCRegister(Reg), RegInfo);
2704 dbgs() << "\n";
2705 });
2706
2707 // If any callee-saved registers are used, the frame cannot be eliminated.
2708 auto [ZPRLocalStackSize, PPRLocalStackSize] =
2710 uint64_t SVELocals = ZPRLocalStackSize + PPRLocalStackSize;
2711 uint64_t SVEStackSize =
2712 alignTo(ZPRCSStackSize + PPRCSStackSize + SVELocals, 16);
2713 bool CanEliminateFrame = (SavedRegs.count() == 0) && !SVEStackSize;
2714
2715 // The CSR spill slots have not been allocated yet, so estimateStackSize
2716 // won't include them.
2717 unsigned EstimatedStackSizeLimit = estimateRSStackSizeLimit(MF);
2718
2719 // We may address some of the stack above the canonical frame address, either
2720 // for our own arguments or during a call. Include that in calculating whether
2721 // we have complicated addressing concerns.
2722 int64_t CalleeStackUsed = 0;
2723 for (int I = MFI.getObjectIndexBegin(); I != 0; ++I) {
2724 int64_t FixedOff = MFI.getObjectOffset(I);
2725 if (FixedOff > CalleeStackUsed)
2726 CalleeStackUsed = FixedOff;
2727 }
2728
2729 // Conservatively always assume BigStack when there are SVE spills.
2730 bool BigStack = SVEStackSize || (EstimatedStackSize + CSStackSize +
2731 CalleeStackUsed) > EstimatedStackSizeLimit;
2732 if (BigStack || !CanEliminateFrame || RegInfo->cannotEliminateFrame(MF))
2733 AFI->setHasStackFrame(true);
2734
2735 // Estimate if we might need to scavenge a register at some point in order
2736 // to materialize a stack offset. If so, either spill one additional
2737 // callee-saved register or reserve a special spill slot to facilitate
2738 // register scavenging. If we already spilled an extra callee-saved register
2739 // above to keep the number of spills even, we don't need to do anything else
2740 // here.
2741 if (BigStack) {
2742 if (!ExtraCSSpill && UnspilledCSGPR != AArch64::NoRegister) {
2743 LLVM_DEBUG(dbgs() << "Spilling " << printReg(UnspilledCSGPR, RegInfo)
2744 << " to get a scratch register.\n");
2745 SavedRegs.set(UnspilledCSGPR);
2746 ExtraCSSpill = UnspilledCSGPR;
2747
2748 // MachO's compact unwind format relies on all registers being stored in
2749 // pairs, so if we need to spill one extra for BigStack, then we need to
2750 // store the pair.
2751 if (producePairRegisters(MF)) {
2752 if (UnspilledCSGPRPaired == AArch64::NoRegister) {
2753 // Failed to make a pair for compact unwind format, revert spilling.
2754 if (produceCompactUnwindFrame(*this, MF)) {
2755 SavedRegs.reset(UnspilledCSGPR);
2756 ExtraCSSpill = AArch64::NoRegister;
2757 }
2758 } else
2759 SavedRegs.set(UnspilledCSGPRPaired);
2760 }
2761 }
2762
2763 // If we didn't find an extra callee-saved register to spill, create
2764 // an emergency spill slot.
2765 if (!ExtraCSSpill || MF.getRegInfo().isPhysRegUsed(ExtraCSSpill)) {
2767 const TargetRegisterClass &RC = AArch64::GPR64RegClass;
2768 unsigned Size = TRI->getSpillSize(RC);
2769 Align Alignment = TRI->getSpillAlign(RC);
2770 int FI = MFI.CreateSpillStackObject(Size, Alignment);
2771 RS->addScavengingFrameIndex(FI);
2772 LLVM_DEBUG(dbgs() << "No available CS registers, allocated fi#" << FI
2773 << " as the emergency spill slot.\n");
2774 }
2775 }
2776
2777 // Adding the size of additional 64bit GPR saves.
2778 CSStackSize += 8 * (SavedRegs.count() - NumSavedRegs);
2779
2780 // A Swift asynchronous context extends the frame record with a pointer
2781 // directly before FP.
2782 if (hasFP(MF) && AFI->hasSwiftAsyncContext())
2783 CSStackSize += 8;
2784
2785 uint64_t AlignedCSStackSize = alignTo(CSStackSize, 16);
2786 LLVM_DEBUG(dbgs() << "Estimated stack frame size: "
2787 << EstimatedStackSize + AlignedCSStackSize << " bytes.\n");
2788
2790 AFI->getCalleeSavedStackSize() == AlignedCSStackSize) &&
2791 "Should not invalidate callee saved info");
2792
2793 // Round up to register pair alignment to avoid additional SP adjustment
2794 // instructions.
2795 AFI->setCalleeSavedStackSize(AlignedCSStackSize);
2796 AFI->setCalleeSaveStackHasFreeSpace(AlignedCSStackSize != CSStackSize);
2797 AFI->setSVECalleeSavedStackSize(ZPRCSStackSize, alignTo(PPRCSStackSize, 16));
2798}
2799
2801 MachineFunction &MF, const TargetRegisterInfo *RegInfo,
2802 std::vector<CalleeSavedInfo> &CSI) const {
2803 bool IsWindows = isTargetWindows(MF);
2804 unsigned StackHazardSize = getStackHazardSize(MF);
2805 // To match the canonical windows frame layout, reverse the list of
2806 // callee saved registers to get them laid out by PrologEpilogInserter
2807 // in the right order. (PrologEpilogInserter allocates stack objects top
2808 // down. Windows canonical prologs store higher numbered registers at
2809 // the top, thus have the CSI array start from the highest registers.)
2810 if (IsWindows)
2811 std::reverse(CSI.begin(), CSI.end());
2812
2813 if (CSI.empty())
2814 return true; // Early exit if no callee saved registers are modified!
2815
2816 // Now that we know which registers need to be saved and restored, allocate
2817 // stack slots for them.
2818 MachineFrameInfo &MFI = MF.getFrameInfo();
2819 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2820
2821 if (IsWindows && hasFP(MF) && AFI->hasSwiftAsyncContext()) {
2822 int FrameIdx = MFI.CreateStackObject(8, Align(16), true);
2823 AFI->setSwiftAsyncContextFrameIdx(FrameIdx);
2824 MFI.setIsCalleeSavedObjectIndex(FrameIdx, true);
2825 }
2826
2827 // Insert VG into the list of CSRs, immediately before LR if saved.
2828 if (requiresSaveVG(MF)) {
2829 CalleeSavedInfo VGInfo(AArch64::VG);
2830 auto It =
2831 find_if(CSI, [](auto &Info) { return Info.getReg() == AArch64::LR; });
2832 if (It != CSI.end())
2833 CSI.insert(It, VGInfo);
2834 else
2835 CSI.push_back(VGInfo);
2836 }
2837
2838 Register LastReg = 0;
2839 int HazardSlotIndex = std::numeric_limits<int>::max();
2840 for (auto &CS : CSI) {
2841 MCRegister Reg = CS.getReg();
2842 const TargetRegisterClass *RC = RegInfo->getMinimalPhysRegClass(Reg);
2843
2844 // Create a hazard slot as we switch between GPR and FPR CSRs.
2846 (!LastReg || !AArch64InstrInfo::isFpOrNEON(LastReg)) &&
2848 assert(HazardSlotIndex == std::numeric_limits<int>::max() &&
2849 "Unexpected register order for hazard slot");
2850 HazardSlotIndex = MFI.CreateStackObject(StackHazardSize, Align(8), true);
2851 LLVM_DEBUG(dbgs() << "Created CSR Hazard at slot " << HazardSlotIndex
2852 << "\n");
2853 AFI->setStackHazardCSRSlotIndex(HazardSlotIndex);
2854 MFI.setIsCalleeSavedObjectIndex(HazardSlotIndex, true);
2855 }
2856
2857 unsigned Size = RegInfo->getSpillSize(*RC);
2858 Align Alignment(RegInfo->getSpillAlign(*RC));
2859 int FrameIdx = MFI.CreateStackObject(Size, Alignment, true);
2860 CS.setFrameIdx(FrameIdx);
2861 MFI.setIsCalleeSavedObjectIndex(FrameIdx, true);
2862
2863 // Grab 8 bytes below FP for the extended asynchronous frame info.
2864 if (hasFP(MF) && AFI->hasSwiftAsyncContext() && !IsWindows &&
2865 Reg == AArch64::FP) {
2866 FrameIdx = MFI.CreateStackObject(8, Alignment, true);
2867 AFI->setSwiftAsyncContextFrameIdx(FrameIdx);
2868 MFI.setIsCalleeSavedObjectIndex(FrameIdx, true);
2869 }
2870 LastReg = Reg;
2871 }
2872
2873 // Add hazard slot in the case where no FPR CSRs are present.
2875 HazardSlotIndex == std::numeric_limits<int>::max()) {
2876 HazardSlotIndex = MFI.CreateStackObject(StackHazardSize, Align(8), true);
2877 LLVM_DEBUG(dbgs() << "Created CSR Hazard at slot " << HazardSlotIndex
2878 << "\n");
2879 AFI->setStackHazardCSRSlotIndex(HazardSlotIndex);
2880 MFI.setIsCalleeSavedObjectIndex(HazardSlotIndex, true);
2881 }
2882
2883 return true;
2884}
2885
2887 const MachineFunction &MF) const {
2889 // If the function has streaming-mode changes, don't scavenge a
2890 // spillslot in the callee-save area, as that might require an
2891 // 'addvl' in the streaming-mode-changing call-sequence when the
2892 // function doesn't use a FP.
2893 if (AFI->hasStreamingModeChanges() && !hasFP(MF))
2894 return false;
2895 // Don't allow register salvaging with hazard slots, in case it moves objects
2896 // into the wrong place.
2897 if (AFI->hasStackHazardSlotIndex())
2898 return false;
2899 return AFI->hasCalleeSaveStackFreeSpace();
2900}
2901
2902/// returns true if there are any SVE callee saves.
2904 int &Min, int &Max) {
2905 Min = std::numeric_limits<int>::max();
2906 Max = std::numeric_limits<int>::min();
2907
2908 if (!MFI.isCalleeSavedInfoValid())
2909 return false;
2910
2911 const std::vector<CalleeSavedInfo> &CSI = MFI.getCalleeSavedInfo();
2912 for (auto &CS : CSI) {
2913 if (AArch64::ZPRRegClass.contains(CS.getReg()) ||
2914 AArch64::PPRRegClass.contains(CS.getReg())) {
2915 assert((Max == std::numeric_limits<int>::min() ||
2916 Max + 1 == CS.getFrameIdx()) &&
2917 "SVE CalleeSaves are not consecutive");
2918 Min = std::min(Min, CS.getFrameIdx());
2919 Max = std::max(Max, CS.getFrameIdx());
2920 }
2921 }
2922 return Min != std::numeric_limits<int>::max();
2923}
2924
2926 AssignObjectOffsets AssignOffsets) {
2927 MachineFrameInfo &MFI = MF.getFrameInfo();
2928 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2929
2930 SVEStackSizes SVEStack{};
2931
2932 // With SplitSVEObjects we maintain separate stack offsets for predicates
2933 // (PPRs) and SVE vectors (ZPRs). When SplitSVEObjects is disabled predicates
2934 // are included in the SVE vector area.
2935 uint64_t &ZPRStackTop = SVEStack.ZPRStackSize;
2936 uint64_t &PPRStackTop =
2937 AFI->hasSplitSVEObjects() ? SVEStack.PPRStackSize : SVEStack.ZPRStackSize;
2938
2939#ifndef NDEBUG
2940 // First process all fixed stack objects.
2941 for (int I = MFI.getObjectIndexBegin(); I != 0; ++I)
2942 assert(!MFI.hasScalableStackID(I) &&
2943 "SVE vectors should never be passed on the stack by value, only by "
2944 "reference.");
2945#endif
2946
2947 auto AllocateObject = [&](int FI) {
2949 ? ZPRStackTop
2950 : PPRStackTop;
2951
2952 // FIXME: Given that the length of SVE vectors is not necessarily a power of
2953 // two, we'd need to align every object dynamically at runtime if the
2954 // alignment is larger than 16. This is not yet supported.
2955 Align Alignment = MFI.getObjectAlign(FI);
2956 if (Alignment > Align(16))
2958 "Alignment of scalable vectors > 16 bytes is not yet supported");
2959
2960 StackTop += MFI.getObjectSize(FI);
2961 StackTop = alignTo(StackTop, Alignment);
2962
2963 assert(StackTop < (uint64_t)std::numeric_limits<int64_t>::max() &&
2964 "SVE StackTop far too large?!");
2965
2966 int64_t Offset = -int64_t(StackTop);
2967 if (AssignOffsets == AssignObjectOffsets::Yes)
2968 MFI.setObjectOffset(FI, Offset);
2969
2970 LLVM_DEBUG(dbgs() << "alloc FI(" << FI << ") at SP[" << Offset << "]\n");
2971 };
2972
2973 // Then process all callee saved slots.
2974 int MinCSFrameIndex, MaxCSFrameIndex;
2975 if (getSVECalleeSaveSlotRange(MFI, MinCSFrameIndex, MaxCSFrameIndex)) {
2976 for (int FI = MinCSFrameIndex; FI <= MaxCSFrameIndex; ++FI)
2977 AllocateObject(FI);
2978 }
2979
2980 // Ensure the CS area is 16-byte aligned.
2981 PPRStackTop = alignTo(PPRStackTop, Align(16U));
2982 ZPRStackTop = alignTo(ZPRStackTop, Align(16U));
2983
2984 // Create a buffer of SVE objects to allocate and sort it.
2985 SmallVector<int, 8> ObjectsToAllocate;
2986 // If we have a stack protector, and we've previously decided that we have SVE
2987 // objects on the stack and thus need it to go in the SVE stack area, then it
2988 // needs to go first.
2989 int StackProtectorFI = -1;
2990 if (MFI.hasStackProtectorIndex()) {
2991 StackProtectorFI = MFI.getStackProtectorIndex();
2992 if (MFI.getStackID(StackProtectorFI) == TargetStackID::ScalableVector)
2993 ObjectsToAllocate.push_back(StackProtectorFI);
2994 }
2995
2996 for (int FI = 0, E = MFI.getObjectIndexEnd(); FI != E; ++FI) {
2997 if (FI == StackProtectorFI || MFI.isDeadObjectIndex(FI) ||
2999 continue;
3000
3003 continue;
3004
3005 ObjectsToAllocate.push_back(FI);
3006 }
3007
3008 // Allocate all SVE locals and spills
3009 for (unsigned FI : ObjectsToAllocate)
3010 AllocateObject(FI);
3011
3012 PPRStackTop = alignTo(PPRStackTop, Align(16U));
3013 ZPRStackTop = alignTo(ZPRStackTop, Align(16U));
3014
3015 if (AssignOffsets == AssignObjectOffsets::Yes)
3016 AFI->setStackSizeSVE(SVEStack.ZPRStackSize, SVEStack.PPRStackSize);
3017
3018 return SVEStack;
3019}
3020
3022 MachineFunction &MF, RegScavenger *RS) const {
3024 "Upwards growing stack unsupported");
3025
3027
3028 // If this function isn't doing Win64-style C++ EH, we don't need to do
3029 // anything.
3030 if (!MF.hasEHFunclets())
3031 return;
3032
3033 MachineFrameInfo &MFI = MF.getFrameInfo();
3034 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
3035
3036 // Win64 C++ EH needs to allocate space for the catch objects in the fixed
3037 // object area right next to the UnwindHelp object.
3038 WinEHFuncInfo &EHInfo = *MF.getWinEHFuncInfo();
3039 int64_t CurrentOffset =
3041 for (WinEHTryBlockMapEntry &TBME : EHInfo.TryBlockMap) {
3042 for (WinEHHandlerType &H : TBME.HandlerArray) {
3043 int FrameIndex = H.CatchObj.FrameIndex;
3044 if ((FrameIndex != INT_MAX) && MFI.getObjectOffset(FrameIndex) == 0) {
3045 CurrentOffset =
3046 alignTo(CurrentOffset, MFI.getObjectAlign(FrameIndex).value());
3047 CurrentOffset += MFI.getObjectSize(FrameIndex);
3048 MFI.setObjectOffset(FrameIndex, -CurrentOffset);
3049 }
3050 }
3051 }
3052
3053 // Create an UnwindHelp object.
3054 // The UnwindHelp object is allocated at the start of the fixed object area
3055 int64_t UnwindHelpOffset = alignTo(CurrentOffset + 8, Align(16));
3056 assert(UnwindHelpOffset == getFixedObjectSize(MF, AFI, /*IsWin64*/ true,
3057 /*IsFunclet*/ false) &&
3058 "UnwindHelpOffset must be at the start of the fixed object area");
3059 int UnwindHelpFI = MFI.CreateFixedObject(/*Size*/ 8, -UnwindHelpOffset,
3060 /*IsImmutable=*/false);
3061 EHInfo.UnwindHelpFrameIdx = UnwindHelpFI;
3062
3063 MachineBasicBlock &MBB = MF.front();
3064 auto MBBI = MBB.begin();
3065 while (MBBI != MBB.end() && MBBI->getFlag(MachineInstr::FrameSetup))
3066 ++MBBI;
3067
3068 // We need to store -2 into the UnwindHelp object at the start of the
3069 // function.
3070 DebugLoc DL;
3071 RS->enterBasicBlockEnd(MBB);
3072 RS->backward(MBBI);
3073 Register DstReg = RS->FindUnusedReg(&AArch64::GPR64commonRegClass);
3074 assert(DstReg && "There must be a free register after frame setup");
3075 const AArch64InstrInfo &TII =
3076 *MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3077 BuildMI(MBB, MBBI, DL, TII.get(AArch64::MOVi64imm), DstReg).addImm(-2);
3078 BuildMI(MBB, MBBI, DL, TII.get(AArch64::STURXi))
3079 .addReg(DstReg, getKillRegState(true))
3080 .addFrameIndex(UnwindHelpFI)
3081 .addImm(0);
3082}
3083
3084namespace {
3085struct TagStoreInstr {
3087 int64_t Offset, Size;
3088 explicit TagStoreInstr(MachineInstr *MI, int64_t Offset, int64_t Size)
3089 : MI(MI), Offset(Offset), Size(Size) {}
3090};
3091
3092class TagStoreEdit {
3093 MachineFunction *MF;
3094 MachineBasicBlock *MBB;
3095 MachineRegisterInfo *MRI;
3096 // Tag store instructions that are being replaced.
3098 // Combined memref arguments of the above instructions.
3100
3101 // Replace allocation tags in [FrameReg + FrameRegOffset, FrameReg +
3102 // FrameRegOffset + Size) with the address tag of SP.
3103 Register FrameReg;
3104 StackOffset FrameRegOffset;
3105 int64_t Size;
3106 // If not std::nullopt, move FrameReg to (FrameReg + FrameRegUpdate) at the
3107 // end.
3108 std::optional<int64_t> FrameRegUpdate;
3109 // MIFlags for any FrameReg updating instructions.
3110 unsigned FrameRegUpdateFlags;
3111
3112 // Use zeroing instruction variants.
3113 bool ZeroData;
3114 DebugLoc DL;
3115
3116 void emitUnrolled(MachineBasicBlock::iterator InsertI);
3117 void emitLoop(MachineBasicBlock::iterator InsertI);
3118
3119public:
3120 TagStoreEdit(MachineBasicBlock *MBB, bool ZeroData)
3121 : MBB(MBB), ZeroData(ZeroData) {
3122 MF = MBB->getParent();
3123 MRI = &MF->getRegInfo();
3124 }
3125 // Add an instruction to be replaced. Instructions must be added in the
3126 // ascending order of Offset, and have to be adjacent.
3127 void addInstruction(TagStoreInstr I) {
3128 assert((TagStores.empty() ||
3129 TagStores.back().Offset + TagStores.back().Size == I.Offset) &&
3130 "Non-adjacent tag store instructions.");
3131 TagStores.push_back(I);
3132 }
3133 void clear() { TagStores.clear(); }
3134 // Emit equivalent code at the given location, and erase the current set of
3135 // instructions. May skip if the replacement is not profitable. May invalidate
3136 // the input iterator and replace it with a valid one.
3137 void emitCode(MachineBasicBlock::iterator &InsertI,
3138 const AArch64FrameLowering *TFI, bool TryMergeSPUpdate);
3139};
3140
3141void TagStoreEdit::emitUnrolled(MachineBasicBlock::iterator InsertI) {
3142 const AArch64InstrInfo *TII =
3143 MF->getSubtarget<AArch64Subtarget>().getInstrInfo();
3144
3145 const int64_t kMinOffset = -256 * 16;
3146 const int64_t kMaxOffset = 255 * 16;
3147
3148 Register BaseReg = FrameReg;
3149 int64_t BaseRegOffsetBytes = FrameRegOffset.getFixed();
3150 if (BaseRegOffsetBytes < kMinOffset ||
3151 BaseRegOffsetBytes + (Size - Size % 32) > kMaxOffset ||
3152 // BaseReg can be FP, which is not necessarily aligned to 16-bytes. In
3153 // that case, BaseRegOffsetBytes will not be aligned to 16 bytes, which
3154 // is required for the offset of ST2G.
3155 BaseRegOffsetBytes % 16 != 0) {
3156 Register ScratchReg = MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3157 emitFrameOffset(*MBB, InsertI, DL, ScratchReg, BaseReg,
3158 StackOffset::getFixed(BaseRegOffsetBytes), TII);
3159 BaseReg = ScratchReg;
3160 BaseRegOffsetBytes = 0;
3161 }
3162
3163 MachineInstr *LastI = nullptr;
3164 while (Size) {
3165 int64_t InstrSize = (Size > 16) ? 32 : 16;
3166 unsigned Opcode =
3167 InstrSize == 16
3168 ? (ZeroData ? AArch64::STZGi : AArch64::STGi)
3169 : (ZeroData ? AArch64::STZ2Gi : AArch64::ST2Gi);
3170 assert(BaseRegOffsetBytes % 16 == 0);
3171 MachineInstr *I = BuildMI(*MBB, InsertI, DL, TII->get(Opcode))
3172 .addReg(AArch64::SP)
3173 .addReg(BaseReg)
3174 .addImm(BaseRegOffsetBytes / 16)
3175 .setMemRefs(CombinedMemRefs);
3176 // A store to [BaseReg, #0] should go last for an opportunity to fold the
3177 // final SP adjustment in the epilogue.
3178 if (BaseRegOffsetBytes == 0)
3179 LastI = I;
3180 BaseRegOffsetBytes += InstrSize;
3181 Size -= InstrSize;
3182 }
3183
3184 if (LastI)
3185 MBB->splice(InsertI, MBB, LastI);
3186}
3187
3188void TagStoreEdit::emitLoop(MachineBasicBlock::iterator InsertI) {
3189 const AArch64InstrInfo *TII =
3190 MF->getSubtarget<AArch64Subtarget>().getInstrInfo();
3191
3192 Register BaseReg = FrameRegUpdate
3193 ? FrameReg
3194 : MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3195 Register SizeReg = MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3196
3197 emitFrameOffset(*MBB, InsertI, DL, BaseReg, FrameReg, FrameRegOffset, TII);
3198
3199 int64_t LoopSize = Size;
3200 // If the loop size is not a multiple of 32, split off one 16-byte store at
3201 // the end to fold BaseReg update into.
3202 if (FrameRegUpdate && *FrameRegUpdate)
3203 LoopSize -= LoopSize % 32;
3204 MachineInstr *LoopI = BuildMI(*MBB, InsertI, DL,
3205 TII->get(ZeroData ? AArch64::STZGloop_wback
3206 : AArch64::STGloop_wback))
3207 .addDef(SizeReg)
3208 .addDef(BaseReg)
3209 .addImm(LoopSize)
3210 .addReg(BaseReg)
3211 .setMemRefs(CombinedMemRefs);
3212 if (FrameRegUpdate)
3213 LoopI->setFlags(FrameRegUpdateFlags);
3214
3215 int64_t ExtraBaseRegUpdate =
3216 FrameRegUpdate ? (*FrameRegUpdate - FrameRegOffset.getFixed() - Size) : 0;
3217 LLVM_DEBUG(dbgs() << "TagStoreEdit::emitLoop: LoopSize=" << LoopSize
3218 << ", Size=" << Size
3219 << ", ExtraBaseRegUpdate=" << ExtraBaseRegUpdate
3220 << ", FrameRegUpdate=" << FrameRegUpdate
3221 << ", FrameRegOffset.getFixed()="
3222 << FrameRegOffset.getFixed() << "\n");
3223 if (LoopSize < Size) {
3224 assert(FrameRegUpdate);
3225 assert(Size - LoopSize == 16);
3226 // Tag 16 more bytes at BaseReg and update BaseReg.
3227 int64_t STGOffset = ExtraBaseRegUpdate + 16;
3228 assert(STGOffset % 16 == 0 && STGOffset >= -4096 && STGOffset <= 4080 &&
3229 "STG immediate out of range");
3230 BuildMI(*MBB, InsertI, DL,
3231 TII->get(ZeroData ? AArch64::STZGPostIndex : AArch64::STGPostIndex))
3232 .addDef(BaseReg)
3233 .addReg(BaseReg)
3234 .addReg(BaseReg)
3235 .addImm(STGOffset / 16)
3236 .setMemRefs(CombinedMemRefs)
3237 .setMIFlags(FrameRegUpdateFlags);
3238 } else if (ExtraBaseRegUpdate) {
3239 // Update BaseReg.
3240 int64_t AddSubOffset = std::abs(ExtraBaseRegUpdate);
3241 assert(AddSubOffset <= 4095 && "ADD/SUB immediate out of range");
3242 BuildMI(
3243 *MBB, InsertI, DL,
3244 TII->get(ExtraBaseRegUpdate > 0 ? AArch64::ADDXri : AArch64::SUBXri))
3245 .addDef(BaseReg)
3246 .addReg(BaseReg)
3247 .addImm(AddSubOffset)
3248 .addImm(0)
3249 .setMIFlags(FrameRegUpdateFlags);
3250 }
3251}
3252
3253// Check if *II is a register update that can be merged into STGloop that ends
3254// at (Reg + Size). RemainingOffset is the required adjustment to Reg after the
3255// end of the loop.
3256bool canMergeRegUpdate(MachineBasicBlock::iterator II, unsigned Reg,
3257 int64_t Size, int64_t *TotalOffset) {
3258 MachineInstr &MI = *II;
3259 if ((MI.getOpcode() == AArch64::ADDXri ||
3260 MI.getOpcode() == AArch64::SUBXri) &&
3261 MI.getOperand(0).getReg() == Reg && MI.getOperand(1).getReg() == Reg) {
3262 unsigned Shift = AArch64_AM::getShiftValue(MI.getOperand(3).getImm());
3263 int64_t Offset = MI.getOperand(2).getImm() << Shift;
3264 if (MI.getOpcode() == AArch64::SUBXri)
3265 Offset = -Offset;
3266 int64_t PostOffset = Offset - Size;
3267 // TagStoreEdit::emitLoop might emit either an ADD/SUB after the loop, or
3268 // an STGPostIndex which does the last 16 bytes of tag write. Which one is
3269 // chosen depends on the alignment of the loop size, but the difference
3270 // between the valid ranges for the two instructions is small, so we
3271 // conservatively assume that it could be either case here.
3272 //
3273 // Max offset of STGPostIndex, minus the 16 byte tag write folded into that
3274 // instruction.
3275 const int64_t kMaxOffset = 4080 - 16;
3276 // Max offset of SUBXri.
3277 const int64_t kMinOffset = -4095;
3278 if (PostOffset <= kMaxOffset && PostOffset >= kMinOffset &&
3279 PostOffset % 16 == 0) {
3280 *TotalOffset = Offset;
3281 return true;
3282 }
3283 }
3284 return false;
3285}
3286
3287void mergeMemRefs(const SmallVectorImpl<TagStoreInstr> &TSE,
3289 MemRefs.clear();
3290 for (auto &TS : TSE) {
3291 MachineInstr *MI = TS.MI;
3292 // An instruction without memory operands may access anything. Be
3293 // conservative and return an empty list.
3294 if (MI->memoperands_empty()) {
3295 MemRefs.clear();
3296 return;
3297 }
3298 MemRefs.append(MI->memoperands_begin(), MI->memoperands_end());
3299 }
3300}
3301
3302void TagStoreEdit::emitCode(MachineBasicBlock::iterator &InsertI,
3303 const AArch64FrameLowering *TFI,
3304 bool TryMergeSPUpdate) {
3305 if (TagStores.empty())
3306 return;
3307 TagStoreInstr &FirstTagStore = TagStores[0];
3308 TagStoreInstr &LastTagStore = TagStores[TagStores.size() - 1];
3309 Size = LastTagStore.Offset - FirstTagStore.Offset + LastTagStore.Size;
3310 DL = TagStores[0].MI->getDebugLoc();
3311
3312 Register Reg;
3313 FrameRegOffset = TFI->resolveFrameOffsetReference(
3314 *MF, FirstTagStore.Offset, false /*isFixed*/,
3315 TargetStackID::Default /*StackID*/, Reg,
3316 /*PreferFP=*/false, /*ForSimm=*/true);
3317 FrameReg = Reg;
3318 FrameRegUpdate = std::nullopt;
3319
3320 mergeMemRefs(TagStores, CombinedMemRefs);
3321
3322 LLVM_DEBUG({
3323 dbgs() << "Replacing adjacent STG instructions:\n";
3324 for (const auto &Instr : TagStores) {
3325 dbgs() << " " << *Instr.MI;
3326 }
3327 });
3328
3329 // Size threshold where a loop becomes shorter than a linear sequence of
3330 // tagging instructions.
3331 const int kSetTagLoopThreshold = 176;
3332 if (Size < kSetTagLoopThreshold) {
3333 if (TagStores.size() < 2)
3334 return;
3335 emitUnrolled(InsertI);
3336 } else {
3337 MachineInstr *UpdateInstr = nullptr;
3338 int64_t TotalOffset = 0;
3339 if (TryMergeSPUpdate) {
3340 // See if we can merge base register update into the STGloop.
3341 // This is done in AArch64LoadStoreOptimizer for "normal" stores,
3342 // but STGloop is way too unusual for that, and also it only
3343 // realistically happens in function epilogue. Also, STGloop is expanded
3344 // before that pass.
3345 if (InsertI != MBB->end() &&
3346 canMergeRegUpdate(InsertI, FrameReg, FrameRegOffset.getFixed() + Size,
3347 &TotalOffset)) {
3348 UpdateInstr = &*InsertI++;
3349 LLVM_DEBUG(dbgs() << "Folding SP update into loop:\n "
3350 << *UpdateInstr);
3351 }
3352 }
3353
3354 if (!UpdateInstr && TagStores.size() < 2)
3355 return;
3356
3357 if (UpdateInstr) {
3358 FrameRegUpdate = TotalOffset;
3359 FrameRegUpdateFlags = UpdateInstr->getFlags();
3360 }
3361 emitLoop(InsertI);
3362 if (UpdateInstr)
3363 UpdateInstr->eraseFromParent();
3364 }
3365
3366 for (auto &TS : TagStores)
3367 TS.MI->eraseFromParent();
3368}
3369
3370bool isMergeableStackTaggingInstruction(MachineInstr &MI, int64_t &Offset,
3371 int64_t &Size, bool &ZeroData) {
3372 MachineFunction &MF = *MI.getParent()->getParent();
3373 const MachineFrameInfo &MFI = MF.getFrameInfo();
3374
3375 unsigned Opcode = MI.getOpcode();
3376 ZeroData = (Opcode == AArch64::STZGloop || Opcode == AArch64::STZGi ||
3377 Opcode == AArch64::STZ2Gi);
3378
3379 if (Opcode == AArch64::STGloop || Opcode == AArch64::STZGloop) {
3380 if (!MI.getOperand(0).isDead() || !MI.getOperand(1).isDead())
3381 return false;
3382 if (!MI.getOperand(2).isImm() || !MI.getOperand(3).isFI())
3383 return false;
3384 Offset = MFI.getObjectOffset(MI.getOperand(3).getIndex());
3385 Size = MI.getOperand(2).getImm();
3386 return true;
3387 }
3388
3389 if (Opcode == AArch64::STGi || Opcode == AArch64::STZGi)
3390 Size = 16;
3391 else if (Opcode == AArch64::ST2Gi || Opcode == AArch64::STZ2Gi)
3392 Size = 32;
3393 else
3394 return false;
3395
3396 if (MI.getOperand(0).getReg() != AArch64::SP || !MI.getOperand(1).isFI())
3397 return false;
3398
3399 Offset = MFI.getObjectOffset(MI.getOperand(1).getIndex()) +
3400 16 * MI.getOperand(2).getImm();
3401 return true;
3402}
3403
3404static size_t countAvailableScavengerSlots(LivePhysRegs &LiveRegs,
3406 RegScavenger *RS) {
3407 auto FreeGPRs =
3408 llvm::count_if(AArch64::GPR64RegClass, [&LiveRegs, &MRI](auto Reg) {
3409 return LiveRegs.available(MRI, Reg);
3410 });
3411
3412 size_t NumEmergencySlots = 0;
3413 if (RS)
3414 NumEmergencySlots = RS->getNumScavengingFrameIndices();
3415
3416 return FreeGPRs + NumEmergencySlots;
3417}
3418
3419// Detect a run of memory tagging instructions for adjacent stack frame slots,
3420// and replace them with a shorter instruction sequence:
3421// * replace STG + STG with ST2G
3422// * replace STGloop + STGloop with STGloop
3423// This code needs to run when stack slot offsets are already known, but before
3424// FrameIndex operands in STG instructions are eliminated.
3426 const AArch64FrameLowering *TFI,
3427 RegScavenger *RS) {
3428 bool FirstZeroData;
3429 int64_t Size, Offset;
3430 MachineInstr &MI = *II;
3431 MachineBasicBlock *MBB = MI.getParent();
3433 if (&MI == &MBB->instr_back())
3434 return II;
3435 if (!isMergeableStackTaggingInstruction(MI, Offset, Size, FirstZeroData))
3436 return II;
3437
3439 Instrs.emplace_back(&MI, Offset, Size);
3440
3441 constexpr int kScanLimit = 10;
3442 int Count = 0;
3444 NextI != E && Count < kScanLimit; ++NextI) {
3445 MachineInstr &MI = *NextI;
3446 bool ZeroData;
3447 int64_t Size, Offset;
3448 // Collect instructions that update memory tags with a FrameIndex operand
3449 // and (when applicable) constant size, and whose output registers are dead
3450 // (the latter is almost always the case in practice). Since these
3451 // instructions effectively have no inputs or outputs, we are free to skip
3452 // any non-aliasing instructions in between without tracking used registers.
3453 if (isMergeableStackTaggingInstruction(MI, Offset, Size, ZeroData)) {
3454 if (ZeroData != FirstZeroData)
3455 break;
3456 Instrs.emplace_back(&MI, Offset, Size);
3457 continue;
3458 }
3459
3460 // Only count non-transient, non-tagging instructions toward the scan
3461 // limit.
3462 if (!MI.isTransient())
3463 ++Count;
3464
3465 // Just in case, stop before the epilogue code starts.
3466 if (MI.getFlag(MachineInstr::FrameSetup) ||
3468 break;
3469
3470 // Reject anything that may alias the collected instructions.
3471 if (MI.mayLoadOrStore() || MI.hasUnmodeledSideEffects() || MI.isCall())
3472 break;
3473 }
3474
3475 // New code will be inserted after the last tagging instruction we've found.
3476 MachineBasicBlock::iterator InsertI = Instrs.back().MI;
3477
3478 // All the gathered stack tag instructions are merged and placed after
3479 // last tag store in the list. The check should be made if the nzcv
3480 // flag is live at the point where we are trying to insert. Otherwise
3481 // the nzcv flag might get clobbered if any stg loops are present.
3482
3483 // FIXME : This approach of bailing out from merge is conservative in
3484 // some ways like even if stg loops are not present after merge the
3485 // insert list, this liveness check is done (which is not needed).
3487 LiveRegs.addLiveOuts(*MBB);
3488 for (auto I = MBB->rbegin();; ++I) {
3489 MachineInstr &MI = *I;
3490 if (MI == InsertI)
3491 break;
3492 LiveRegs.stepBackward(*I);
3493 }
3494 InsertI++;
3495 if (LiveRegs.contains(AArch64::NZCV))
3496 return InsertI;
3497
3498 // Emitting an MTE loop requires two physical registers (BaseReg and
3499 // SizeReg). If the function is under register pressure, the register
3500 // scavenger will crash trying to allocate them. If we don't have at least
3501 // two free slots (free registers + emergency slots), bail out and fall back
3502 // to the unrolled sequence.
3503 if (countAvailableScavengerSlots(LiveRegs, MBB->getParent()->getRegInfo(),
3504 RS) < 2) {
3505 LLVM_DEBUG(
3506 dbgs() << "Failed to merge MTE stack tagging instructions into loop "
3507 << "due to high register pressure.\n");
3508 return InsertI;
3509 }
3510
3511 llvm::stable_sort(Instrs,
3512 [](const TagStoreInstr &Left, const TagStoreInstr &Right) {
3513 return Left.Offset < Right.Offset;
3514 });
3515
3516 // Make sure that we don't have any overlapping stores.
3517 int64_t CurOffset = Instrs[0].Offset;
3518 for (auto &Instr : Instrs) {
3519 if (CurOffset > Instr.Offset)
3520 return NextI;
3521 CurOffset = Instr.Offset + Instr.Size;
3522 }
3523
3524 // Find contiguous runs of tagged memory and emit shorter instruction
3525 // sequences for them when possible.
3526 TagStoreEdit TSE(MBB, FirstZeroData);
3527 std::optional<int64_t> EndOffset;
3528 for (auto &Instr : Instrs) {
3529 if (EndOffset && *EndOffset != Instr.Offset) {
3530 // Found a gap.
3531 TSE.emitCode(InsertI, TFI, /*TryMergeSPUpdate = */ false);
3532 TSE.clear();
3533 }
3534
3535 TSE.addInstruction(Instr);
3536 EndOffset = Instr.Offset + Instr.Size;
3537 }
3538
3539 const MachineFunction *MF = MBB->getParent();
3540 // Multiple FP/SP updates in a loop cannot be described by CFI instructions.
3541 TSE.emitCode(
3542 InsertI, TFI, /*TryMergeSPUpdate = */
3544
3545 return InsertI;
3546}
3547} // namespace
3548
3550 MachineFunction &MF, RegScavenger *RS = nullptr) const {
3551 for (auto &BB : MF)
3552 for (MachineBasicBlock::iterator II = BB.begin(); II != BB.end();) {
3554 II = tryMergeAdjacentSTG(II, this, RS);
3555 }
3556
3557 // By the time this method is called, most of the prologue/epilogue code is
3558 // already emitted, whether its location was affected by the shrink-wrapping
3559 // optimization or not.
3560 if (!MF.getFunction().hasFnAttribute(Attribute::Naked) &&
3561 shouldSignReturnAddressEverywhere(MF))
3563}
3564
3565/// For Win64 AArch64 EH, the offset to the Unwind object is from the SP
3566/// before the update. This is easily retrieved as it is exactly the offset
3567/// that is set in processFunctionBeforeFrameFinalized.
3569 const MachineFunction &MF, int FI, Register &FrameReg,
3570 bool IgnoreSPUpdates) const {
3571 const MachineFrameInfo &MFI = MF.getFrameInfo();
3572 if (IgnoreSPUpdates) {
3573 LLVM_DEBUG(dbgs() << "Offset from the SP for " << FI << " is "
3574 << MFI.getObjectOffset(FI) << "\n");
3575 FrameReg = AArch64::SP;
3576 return StackOffset::getFixed(MFI.getObjectOffset(FI));
3577 }
3578
3579 // Go to common code if we cannot provide sp + offset.
3580 if (MFI.hasVarSizedObjects() ||
3583 return getFrameIndexReference(MF, FI, FrameReg);
3584
3585 FrameReg = AArch64::SP;
3586 return getStackOffset(MF, MFI.getObjectOffset(FI));
3587}
3588
3589/// The parent frame offset (aka dispFrame) is only used on X86_64 to retrieve
3590/// the parent's frame pointer
3592 const MachineFunction &MF) const {
3593 return 0;
3594}
3595
3596/// Funclets only need to account for space for the callee saved registers,
3597/// as the locals are accounted for in the parent's stack frame.
3599 const MachineFunction &MF) const {
3600 // This is the size of the pushed CSRs.
3601 unsigned CSSize =
3602 MF.getInfo<AArch64FunctionInfo>()->getCalleeSavedStackSize();
3603 // This is the amount of stack a funclet needs to allocate.
3604 return alignTo(CSSize + MF.getFrameInfo().getMaxCallFrameSize(),
3605 getStackAlign());
3606}
3607
3608namespace {
3609struct FrameObject {
3610 bool IsValid = false;
3611 // Index of the object in MFI.
3612 int ObjectIndex = 0;
3613 // Group ID this object belongs to.
3614 int GroupIndex = -1;
3615 // This object should be placed first (closest to SP).
3616 bool ObjectFirst = false;
3617 // This object's group (which always contains the object with
3618 // ObjectFirst==true) should be placed first.
3619 bool GroupFirst = false;
3620
3621 // Used to distinguish between FP and GPR accesses. The values are decided so
3622 // that they sort FPR < Hazard < GPR and they can be or'd together.
3623 unsigned Accesses = 0;
3624 enum { AccessFPR = 1, AccessHazard = 2, AccessGPR = 4 };
3625};
3626
3627class GroupBuilder {
3628 SmallVector<int, 8> CurrentMembers;
3629 int NextGroupIndex = 0;
3630 std::vector<FrameObject> &Objects;
3631
3632public:
3633 GroupBuilder(std::vector<FrameObject> &Objects) : Objects(Objects) {}
3634 void AddMember(int Index) { CurrentMembers.push_back(Index); }
3635 void EndCurrentGroup() {
3636 if (CurrentMembers.size() > 1) {
3637 // Create a new group with the current member list. This might remove them
3638 // from their pre-existing groups. That's OK, dealing with overlapping
3639 // groups is too hard and unlikely to make a difference.
3640 LLVM_DEBUG(dbgs() << "group:");
3641 for (int Index : CurrentMembers) {
3642 Objects[Index].GroupIndex = NextGroupIndex;
3643 LLVM_DEBUG(dbgs() << " " << Index);
3644 }
3645 LLVM_DEBUG(dbgs() << "\n");
3646 NextGroupIndex++;
3647 }
3648 CurrentMembers.clear();
3649 }
3650};
3651
3652bool FrameObjectCompare(const FrameObject &A, const FrameObject &B) {
3653 // Objects at a lower index are closer to FP; objects at a higher index are
3654 // closer to SP.
3655 //
3656 // For consistency in our comparison, all invalid objects are placed
3657 // at the end. This also allows us to stop walking when we hit the
3658 // first invalid item after it's all sorted.
3659 //
3660 // If we want to include a stack hazard region, order FPR accesses < the
3661 // hazard object < GPRs accesses in order to create a separation between the
3662 // two. For the Accesses field 1 = FPR, 2 = Hazard Object, 4 = GPR.
3663 //
3664 // Otherwise the "first" object goes first (closest to SP), followed by the
3665 // members of the "first" group.
3666 //
3667 // The rest are sorted by the group index to keep the groups together.
3668 // Higher numbered groups are more likely to be around longer (i.e. untagged
3669 // in the function epilogue and not at some earlier point). Place them closer
3670 // to SP.
3671 //
3672 // If all else equal, sort by the object index to keep the objects in the
3673 // original order.
3674 return std::make_tuple(!A.IsValid, A.Accesses, A.ObjectFirst, A.GroupFirst,
3675 A.GroupIndex, A.ObjectIndex) <
3676 std::make_tuple(!B.IsValid, B.Accesses, B.ObjectFirst, B.GroupFirst,
3677 B.GroupIndex, B.ObjectIndex);
3678}
3679} // namespace
3680
3682 const MachineFunction &MF, SmallVectorImpl<int> &ObjectsToAllocate) const {
3684
3685 if ((!OrderFrameObjects && !AFI.hasSplitSVEObjects()) ||
3686 ObjectsToAllocate.empty())
3687 return;
3688
3689 const MachineFrameInfo &MFI = MF.getFrameInfo();
3690 std::vector<FrameObject> FrameObjects(MFI.getObjectIndexEnd());
3691 for (auto &Obj : ObjectsToAllocate) {
3692 FrameObjects[Obj].IsValid = true;
3693 FrameObjects[Obj].ObjectIndex = Obj;
3694 }
3695
3696 // Identify FPR vs GPR slots for hazards, and stack slots that are tagged at
3697 // the same time.
3698 GroupBuilder GB(FrameObjects);
3699 for (auto &MBB : MF) {
3700 for (auto &MI : MBB) {
3701 if (MI.isDebugInstr())
3702 continue;
3703
3704 if (AFI.hasStackHazardSlotIndex()) {
3705 std::optional<int> FI = getLdStFrameID(MI, MFI);
3706 if (FI && *FI >= 0 && *FI < (int)FrameObjects.size()) {
3707 if (MFI.getStackID(*FI) == TargetStackID::ScalableVector ||
3709 FrameObjects[*FI].Accesses |= FrameObject::AccessFPR;
3710 else
3711 FrameObjects[*FI].Accesses |= FrameObject::AccessGPR;
3712 }
3713 }
3714
3715 int OpIndex;
3716 switch (MI.getOpcode()) {
3717 case AArch64::STGloop:
3718 case AArch64::STZGloop:
3719 OpIndex = 3;
3720 break;
3721 case AArch64::STGi:
3722 case AArch64::STZGi:
3723 case AArch64::ST2Gi:
3724 case AArch64::STZ2Gi:
3725 OpIndex = 1;
3726 break;
3727 default:
3728 OpIndex = -1;
3729 }
3730
3731 int TaggedFI = -1;
3732 if (OpIndex >= 0) {
3733 const MachineOperand &MO = MI.getOperand(OpIndex);
3734 if (MO.isFI()) {
3735 int FI = MO.getIndex();
3736 if (FI >= 0 && FI < MFI.getObjectIndexEnd() &&
3737 FrameObjects[FI].IsValid)
3738 TaggedFI = FI;
3739 }
3740 }
3741
3742 // If this is a stack tagging instruction for a slot that is not part of a
3743 // group yet, either start a new group or add it to the current one.
3744 if (TaggedFI >= 0)
3745 GB.AddMember(TaggedFI);
3746 else
3747 GB.EndCurrentGroup();
3748 }
3749 // Groups should never span multiple basic blocks.
3750 GB.EndCurrentGroup();
3751 }
3752
3753 if (AFI.hasStackHazardSlotIndex()) {
3754 FrameObjects[AFI.getStackHazardSlotIndex()].Accesses =
3755 FrameObject::AccessHazard;
3756 // If a stack object is unknown or both GPR and FPR, sort it into GPR.
3757 for (auto &Obj : FrameObjects)
3758 if (!Obj.Accesses ||
3759 Obj.Accesses == (FrameObject::AccessGPR | FrameObject::AccessFPR))
3760 Obj.Accesses = FrameObject::AccessGPR;
3761 }
3762
3763 // If the function's tagged base pointer is pinned to a stack slot, we want to
3764 // put that slot first when possible. This will likely place it at SP + 0,
3765 // and save one instruction when generating the base pointer because IRG does
3766 // not allow an immediate offset.
3767 std::optional<int> TBPI = AFI.getTaggedBasePointerIndex();
3768 if (TBPI) {
3769 FrameObjects[*TBPI].ObjectFirst = true;
3770 FrameObjects[*TBPI].GroupFirst = true;
3771 int FirstGroupIndex = FrameObjects[*TBPI].GroupIndex;
3772 if (FirstGroupIndex >= 0)
3773 for (FrameObject &Object : FrameObjects)
3774 if (Object.GroupIndex == FirstGroupIndex)
3775 Object.GroupFirst = true;
3776 }
3777
3778 llvm::stable_sort(FrameObjects, FrameObjectCompare);
3779
3780 int i = 0;
3781 for (auto &Obj : FrameObjects) {
3782 // All invalid items are sorted at the end, so it's safe to stop.
3783 if (!Obj.IsValid)
3784 break;
3785 ObjectsToAllocate[i++] = Obj.ObjectIndex;
3786 }
3787
3788 LLVM_DEBUG({
3789 dbgs() << "Final frame order:\n";
3790 for (auto &Obj : FrameObjects) {
3791 if (!Obj.IsValid)
3792 break;
3793 dbgs() << " " << Obj.ObjectIndex << ": group " << Obj.GroupIndex;
3794 if (Obj.ObjectFirst)
3795 dbgs() << ", first";
3796 if (Obj.GroupFirst)
3797 dbgs() << ", group-first";
3798 dbgs() << "\n";
3799 }
3800 });
3801}
3802
3803/// Emit a loop to decrement SP until it is equal to TargetReg, with probes at
3804/// least every ProbeSize bytes. Returns an iterator of the first instruction
3805/// after the loop. The difference between SP and TargetReg must be an exact
3806/// multiple of ProbeSize.
3808AArch64FrameLowering::inlineStackProbeLoopExactMultiple(
3809 MachineBasicBlock::iterator MBBI, int64_t ProbeSize,
3810 Register TargetReg) const {
3811 MachineBasicBlock &MBB = *MBBI->getParent();
3812 MachineFunction &MF = *MBB.getParent();
3813 const AArch64InstrInfo *TII =
3814 MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3815 DebugLoc DL = MBB.findDebugLoc(MBBI);
3816
3817 MachineFunction::iterator MBBInsertPoint = std::next(MBB.getIterator());
3818 MachineBasicBlock *LoopMBB = MF.CreateMachineBasicBlock(MBB.getBasicBlock());
3819 MF.insert(MBBInsertPoint, LoopMBB);
3820 MachineBasicBlock *ExitMBB = MF.CreateMachineBasicBlock(MBB.getBasicBlock());
3821 MF.insert(MBBInsertPoint, ExitMBB);
3822
3823 // SUB SP, SP, #ProbeSize (or equivalent if ProbeSize is not encodable
3824 // in SUB).
3825 emitFrameOffset(*LoopMBB, LoopMBB->end(), DL, AArch64::SP, AArch64::SP,
3826 StackOffset::getFixed(-ProbeSize), TII,
3828 // LDR XZR, [SP]
3829 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::LDRXui))
3830 .addDef(AArch64::XZR)
3831 .addReg(AArch64::SP)
3832 .addImm(0)
3836 Align(8)))
3838 // CMP SP, TargetReg
3839 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::SUBSXrx64),
3840 AArch64::XZR)
3841 .addReg(AArch64::SP)
3842 .addReg(TargetReg)
3845 // B.CC Loop
3846 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::Bcc))
3848 .addMBB(LoopMBB)
3850
3851 LoopMBB->addSuccessor(ExitMBB);
3852 LoopMBB->addSuccessor(LoopMBB);
3853 // Synthesize the exit MBB.
3854 ExitMBB->splice(ExitMBB->end(), &MBB, MBBI, MBB.end());
3856 MBB.addSuccessor(LoopMBB);
3857 // Update liveins.
3858 fullyRecomputeLiveIns({ExitMBB, LoopMBB});
3859
3860 return ExitMBB->begin();
3861}
3862
3863void AArch64FrameLowering::inlineStackProbeFixed(
3864 MachineBasicBlock::iterator MBBI, Register ScratchReg, int64_t FrameSize,
3865 StackOffset CFAOffset) const {
3866 MachineBasicBlock *MBB = MBBI->getParent();
3867 MachineFunction &MF = *MBB->getParent();
3868 const AArch64InstrInfo *TII =
3869 MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3870 AArch64FunctionInfo *AFI = MF.getInfo<AArch64FunctionInfo>();
3871 bool EmitAsyncCFI = AFI->needsAsyncDwarfUnwindInfo(MF);
3872 bool HasFP = hasFP(MF);
3873
3874 DebugLoc DL;
3875 int64_t ProbeSize = MF.getInfo<AArch64FunctionInfo>()->getStackProbeSize();
3876 int64_t NumBlocks = FrameSize / ProbeSize;
3877 int64_t ResidualSize = FrameSize % ProbeSize;
3878
3879 LLVM_DEBUG(dbgs() << "Stack probing: total " << FrameSize << " bytes, "
3880 << NumBlocks << " blocks of " << ProbeSize
3881 << " bytes, plus " << ResidualSize << " bytes\n");
3882
3883 // Decrement SP by NumBlock * ProbeSize bytes, with either unrolled or
3884 // ordinary loop.
3885 if (NumBlocks <= AArch64::StackProbeMaxLoopUnroll) {
3886 for (int i = 0; i < NumBlocks; ++i) {
3887 // SUB SP, SP, #ProbeSize (or equivalent if ProbeSize is not
3888 // encodable in a SUB).
3889 emitFrameOffset(*MBB, MBBI, DL, AArch64::SP, AArch64::SP,
3890 StackOffset::getFixed(-ProbeSize), TII,
3891 MachineInstr::FrameSetup, false, false, nullptr,
3892 EmitAsyncCFI && !HasFP, CFAOffset);
3893 CFAOffset += StackOffset::getFixed(ProbeSize);
3894 // LDR XZR, [SP]
3895 BuildMI(*MBB, MBBI, DL, TII->get(AArch64::LDRXui))
3896 .addDef(AArch64::XZR)
3897 .addReg(AArch64::SP)
3898 .addImm(0)
3902 Align(8)))
3904 }
3905 } else if (NumBlocks != 0) {
3906 // SUB ScratchReg, SP, #FrameSize (or equivalent if FrameSize is not
3907 // encodable in ADD). ScrathReg may temporarily become the CFA register.
3908 emitFrameOffset(*MBB, MBBI, DL, ScratchReg, AArch64::SP,
3909 StackOffset::getFixed(-ProbeSize * NumBlocks), TII,
3910 MachineInstr::FrameSetup, false, false, nullptr,
3911 EmitAsyncCFI && !HasFP, CFAOffset);
3912 CFAOffset += StackOffset::getFixed(ProbeSize * NumBlocks);
3913 MBBI = inlineStackProbeLoopExactMultiple(MBBI, ProbeSize, ScratchReg);
3914 MBB = MBBI->getParent();
3915 if (EmitAsyncCFI && !HasFP) {
3916 // Set the CFA register back to SP.
3917 CFIInstBuilder(*MBB, MBBI, MachineInstr::FrameSetup)
3918 .buildDefCFARegister(AArch64::SP);
3919 }
3920 }
3921
3922 if (ResidualSize != 0) {
3923 // SUB SP, SP, #ResidualSize (or equivalent if ResidualSize is not encodable
3924 // in SUB).
3925 emitFrameOffset(*MBB, MBBI, DL, AArch64::SP, AArch64::SP,
3926 StackOffset::getFixed(-ResidualSize), TII,
3927 MachineInstr::FrameSetup, false, false, nullptr,
3928 EmitAsyncCFI && !HasFP, CFAOffset);
3929 if (ResidualSize > AArch64::StackProbeMaxUnprobedStack) {
3930 // LDR XZR, [SP]
3931 BuildMI(*MBB, MBBI, DL, TII->get(AArch64::LDRXui))
3932 .addDef(AArch64::XZR)
3933 .addReg(AArch64::SP)
3934 .addImm(0)
3938 Align(8)))
3940 }
3941 }
3942}
3943
3944void AArch64FrameLowering::inlineStackProbe(MachineFunction &MF,
3945 MachineBasicBlock &MBB) const {
3946 // Get the instructions that need to be replaced. We emit at most two of
3947 // these. Remember them in order to avoid complications coming from the need
3948 // to traverse the block while potentially creating more blocks.
3949 SmallVector<MachineInstr *, 4> ToReplace;
3950 for (MachineInstr &MI : MBB)
3951 if (MI.getOpcode() == AArch64::PROBED_STACKALLOC ||
3952 MI.getOpcode() == AArch64::PROBED_STACKALLOC_VAR)
3953 ToReplace.push_back(&MI);
3954
3955 for (MachineInstr *MI : ToReplace) {
3956 if (MI->getOpcode() == AArch64::PROBED_STACKALLOC) {
3957 Register ScratchReg = MI->getOperand(0).getReg();
3958 int64_t FrameSize = MI->getOperand(1).getImm();
3959 StackOffset CFAOffset = StackOffset::get(MI->getOperand(2).getImm(),
3960 MI->getOperand(3).getImm());
3961 inlineStackProbeFixed(MI->getIterator(), ScratchReg, FrameSize,
3962 CFAOffset);
3963 } else {
3964 assert(MI->getOpcode() == AArch64::PROBED_STACKALLOC_VAR &&
3965 "Stack probe pseudo-instruction expected");
3966 const AArch64InstrInfo *TII =
3967 MI->getMF()->getSubtarget<AArch64Subtarget>().getInstrInfo();
3968 Register TargetReg = MI->getOperand(0).getReg();
3969 (void)TII->probedStackAlloc(MI->getIterator(), TargetReg, true);
3970 }
3971 MI->eraseFromParent();
3972 }
3973}
3974
3977 NotAccessed = 0, // Stack object not accessed by load/store instructions.
3978 GPR = 1 << 0, // A general purpose register.
3979 PPR = 1 << 1, // A predicate register.
3980 FPR = 1 << 2, // A floating point/Neon/SVE register.
3981 };
3982
3983 int Idx;
3985 int64_t Size;
3986 unsigned AccessTypes;
3987
3989
3990 bool operator<(const StackAccess &Rhs) const {
3991 return std::make_tuple(start(), Idx) <
3992 std::make_tuple(Rhs.start(), Rhs.Idx);
3993 }
3994
3995 bool isCPU() const {
3996 // Predicate register load and store instructions execute on the CPU.
3998 }
3999 bool isSME() const { return AccessTypes & AccessType::FPR; }
4000 bool isMixed() const { return isCPU() && isSME(); }
4001
4002 int64_t start() const { return Offset.getFixed() + Offset.getScalable(); }
4003 int64_t end() const { return start() + Size; }
4004
4005 std::string getTypeString() const {
4006 switch (AccessTypes) {
4007 case AccessType::FPR:
4008 return "FPR";
4009 case AccessType::PPR:
4010 return "PPR";
4011 case AccessType::GPR:
4012 return "GPR";
4014 return "NA";
4015 default:
4016 return "Mixed";
4017 }
4018 }
4019
4020 void print(raw_ostream &OS) const {
4021 OS << getTypeString() << " stack object at [SP"
4022 << (Offset.getFixed() < 0 ? "" : "+") << Offset.getFixed();
4023 if (Offset.getScalable())
4024 OS << (Offset.getScalable() < 0 ? "" : "+") << Offset.getScalable()
4025 << " * vscale";
4026 OS << "]";
4027 }
4028};
4029
4030static inline raw_ostream &operator<<(raw_ostream &OS, const StackAccess &SA) {
4031 SA.print(OS);
4032 return OS;
4033}
4034
4035void AArch64FrameLowering::emitRemarks(
4036 const MachineFunction &MF, MachineOptimizationRemarkEmitter *ORE) const {
4037
4038 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
4040 return;
4041
4042 unsigned StackHazardSize = getStackHazardSize(MF);
4043 const uint64_t HazardSize =
4044 (StackHazardSize) ? StackHazardSize : StackHazardRemarkSize;
4045
4046 if (HazardSize == 0)
4047 return;
4048
4049 const MachineFrameInfo &MFI = MF.getFrameInfo();
4050 // Bail if function has no stack objects.
4051 if (!MFI.hasStackObjects())
4052 return;
4053
4054 std::vector<StackAccess> StackAccesses(MFI.getNumObjects());
4055
4056 size_t NumFPLdSt = 0;
4057 size_t NumNonFPLdSt = 0;
4058
4059 // Collect stack accesses via Load/Store instructions.
4060 for (const MachineBasicBlock &MBB : MF) {
4061 for (const MachineInstr &MI : MBB) {
4062 if (!MI.mayLoadOrStore() || MI.getNumMemOperands() < 1)
4063 continue;
4064 for (MachineMemOperand *MMO : MI.memoperands()) {
4065 std::optional<int> FI = getMMOFrameID(MMO, MFI);
4066 if (FI && !MFI.isDeadObjectIndex(*FI)) {
4067 int FrameIdx = *FI;
4068
4069 size_t ArrIdx = FrameIdx + MFI.getNumFixedObjects();
4070 if (StackAccesses[ArrIdx].AccessTypes == StackAccess::NotAccessed) {
4071 StackAccesses[ArrIdx].Idx = FrameIdx;
4072 StackAccesses[ArrIdx].Offset =
4073 getFrameIndexReferenceFromSP(MF, FrameIdx);
4074 StackAccesses[ArrIdx].Size = MFI.getObjectSize(FrameIdx);
4075 }
4076
4077 unsigned RegTy = StackAccess::AccessType::GPR;
4078 if (MFI.hasScalableStackID(FrameIdx))
4081 RegTy = StackAccess::FPR;
4082
4083 StackAccesses[ArrIdx].AccessTypes |= RegTy;
4084
4085 if (RegTy == StackAccess::FPR)
4086 ++NumFPLdSt;
4087 else
4088 ++NumNonFPLdSt;
4089 }
4090 }
4091 }
4092 }
4093
4094 if (NumFPLdSt == 0 || NumNonFPLdSt == 0)
4095 return;
4096
4097 llvm::sort(StackAccesses);
4098 llvm::erase_if(StackAccesses, [](const StackAccess &S) {
4100 });
4101
4104
4105 if (StackAccesses.front().isMixed())
4106 MixedObjects.push_back(&StackAccesses.front());
4107
4108 for (auto It = StackAccesses.begin(), End = std::prev(StackAccesses.end());
4109 It != End; ++It) {
4110 const auto &First = *It;
4111 const auto &Second = *(It + 1);
4112
4113 if (Second.isMixed())
4114 MixedObjects.push_back(&Second);
4115
4116 if ((First.isSME() && Second.isCPU()) ||
4117 (First.isCPU() && Second.isSME())) {
4118 uint64_t Distance = static_cast<uint64_t>(Second.start() - First.end());
4119 if (Distance < HazardSize)
4120 HazardPairs.emplace_back(&First, &Second);
4121 }
4122 }
4123
4124 auto EmitRemark = [&](llvm::StringRef Str) {
4125 ORE->emit([&]() {
4126 auto R = MachineOptimizationRemarkAnalysis(
4127 "sme", "StackHazard", MF.getFunction().getSubprogram(), &MF.front());
4128 return R << formatv("stack hazard in '{0}': ", MF.getName()).str() << Str;
4129 });
4130 };
4131
4132 for (const auto &P : HazardPairs)
4133 EmitRemark(formatv("{0} is too close to {1}", *P.first, *P.second).str());
4134
4135 for (const auto *Obj : MixedObjects)
4136 EmitRemark(
4137 formatv("{0} accessed by both GP and FP instructions", *Obj).str());
4138}
static void getLiveRegsForEntryMBB(LivePhysRegs &LiveRegs, const MachineBasicBlock &MBB)
static const unsigned DefaultSafeSPDisplacement
This is the biggest offset to the stack pointer we can encode in aarch64 instructions (without using ...
static RegState getPrologueDeath(MachineFunction &MF, unsigned Reg)
static bool produceCompactUnwindFrame(const AArch64FrameLowering &, MachineFunction &MF)
static cl::opt< bool > StackTaggingMergeSetTag("stack-tagging-merge-settag", cl::desc("merge settag instruction in function epilog"), cl::init(true), cl::Hidden)
bool enableMultiVectorSpillFill(const AArch64Subtarget &Subtarget, MachineFunction &MF)
static std::optional< int > getLdStFrameID(const MachineInstr &MI, const MachineFrameInfo &MFI)
static cl::opt< bool > SplitSVEObjects("aarch64-split-sve-objects", cl::desc("Split allocation of ZPR & PPR objects"), cl::init(true), cl::Hidden)
static cl::opt< bool > StackHazardInNonStreaming("aarch64-stack-hazard-in-non-streaming", cl::init(false), cl::Hidden)
void computeCalleeSaveRegisterPairs(const AArch64FrameLowering &AFL, MachineFunction &MF, ArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI, SmallVectorImpl< RegPairInfo > &RegPairs, bool NeedsFrameRecord)
static cl::opt< bool > OrderFrameObjects("aarch64-order-frame-objects", cl::desc("sort stack allocations"), cl::init(true), cl::Hidden)
static cl::opt< bool > DisableMultiVectorSpillFill("aarch64-disable-multivector-spill-fill", cl::desc("Disable use of LD/ST pairs for SME2 or SVE2p1"), cl::init(false), cl::Hidden)
static cl::opt< bool > EnableRedZone("aarch64-redzone", cl::desc("enable use of redzone on AArch64"), cl::init(false), cl::Hidden)
static bool invalidateRegisterPairing(bool SpillExtendedVolatile, unsigned SpillCount, unsigned Reg1, unsigned Reg2, bool UsesWinAAPCS, bool NeedsWinCFI, bool NeedsFrameRecord, const TargetRegisterInfo *TRI)
Returns true if Reg1 and Reg2 cannot be paired using a ldp/stp instruction.
cl::opt< bool > EnableHomogeneousPrologEpilog("homogeneous-prolog-epilog", cl::Hidden, cl::desc("Emit homogeneous prologue and epilogue for the size " "optimization (default = off)"))
static bool isLikelyToHaveSVEStack(const AArch64FrameLowering &AFL, const MachineFunction &MF)
static bool invalidateWindowsRegisterPairing(bool SpillExtendedVolatile, unsigned SpillCount, unsigned Reg1, unsigned Reg2, bool NeedsWinCFI, const TargetRegisterInfo *TRI)
static SVEStackSizes determineSVEStackSizes(MachineFunction &MF, AssignObjectOffsets AssignOffsets)
Process all the SVE stack objects and the SVE stack size and offsets for each object.
static bool isTargetWindows(const MachineFunction &MF)
static unsigned estimateRSStackSizeLimit(MachineFunction &MF)
Look at each instruction that references stack frames and return the stack size limit beyond which so...
static bool getSVECalleeSaveSlotRange(const MachineFrameInfo &MFI, int &Min, int &Max)
returns true if there are any SVE callee saves.
static cl::opt< unsigned > StackHazardRemarkSize("aarch64-stack-hazard-remark-size", cl::init(0), cl::Hidden)
static MCRegister getRegisterOrZero(MCRegister Reg, bool HasSVE)
static unsigned getStackHazardSize(const MachineFunction &MF)
MCRegister findFreePredicateReg(BitVector &SavedRegs)
static bool isPPRAccess(const MachineInstr &MI)
static std::optional< int > getMMOFrameID(MachineMemOperand *MMO, const MachineFrameInfo &MFI)
assert(UImm &&(UImm !=~static_cast< T >(0)) &&"Invalid immediate!")
This file contains the declaration of the AArch64PrologueEmitter and AArch64EpilogueEmitter classes,...
static const int kSetTagLoopThreshold
MachineBasicBlock & MBB
MachineBasicBlock MachineBasicBlock::iterator DebugLoc DL
MachineBasicBlock MachineBasicBlock::iterator MBBI
This file contains the simple types necessary to represent the attributes associated with functions a...
#define CASE(ATTRNAME, AANAME,...)
static GCRegistry::Add< ErlangGC > A("erlang", "erlang-compatible garbage collector")
static GCRegistry::Add< CoreCLRGC > E("coreclr", "CoreCLR-compatible GC")
static GCRegistry::Add< OcamlGC > B("ocaml", "ocaml 3.10-compatible GC")
DXIL Forward Handle Accesses
const HexagonInstrInfo * TII
IRTranslator LLVM IR MI
static std::string getTypeString(Type *T)
Definition LLParser.cpp:68
This file implements the LivePhysRegs utility for tracking liveness of physical registers.
#define F(x, y, z)
Definition MD5.cpp:54
#define I(x, y, z)
Definition MD5.cpp:57
#define H(x, y, z)
Definition MD5.cpp:56
Register Reg
Register const TargetRegisterInfo * TRI
Promote Memory to Register
Definition Mem2Reg.cpp:110
uint64_t IntrinsicInst * II
#define P(N)
This file declares the machine register scavenger class.
static bool contains(SmallPtrSetImpl< ConstantExpr * > &Cache, ConstantExpr *Expr, Constant *C)
Definition Value.cpp:484
This file defines the scope_exit class, which executes user-defined cleanup logic at scope exit.
This file defines the SmallVector class.
#define LLVM_DEBUG(...)
Definition Debug.h:119
StackOffset getSVEStackSize(const MachineFunction &MF) const
Returns the size of the entire SVE stackframe (PPRs + ZPRs).
StackOffset getZPRStackSize(const MachineFunction &MF) const
Returns the size of the entire ZPR stackframe (calleesaves + spills).
void processFunctionBeforeFrameIndicesReplaced(MachineFunction &MF, RegScavenger *RS) const override
processFunctionBeforeFrameIndicesReplaced - This method is called immediately before MO_FrameIndex op...
MachineBasicBlock::iterator eliminateCallFramePseudoInstr(MachineFunction &MF, MachineBasicBlock &MBB, MachineBasicBlock::iterator I) const override
This method is called during prolog/epilog code insertion to eliminate call frame setup and destroy p...
bool canUseAsPrologue(const MachineBasicBlock &MBB) const override
Check whether or not the given MBB can be used as a prologue for the target.
bool enableStackSlotScavenging(const MachineFunction &MF) const override
Returns true if the stack slot holes in the fixed and callee-save stack area should be used when allo...
bool assignCalleeSavedSpillSlots(MachineFunction &MF, const TargetRegisterInfo *TRI, std::vector< CalleeSavedInfo > &CSI) const override
assignCalleeSavedSpillSlots - Allows target to override spill slot assignment logic.
bool spillCalleeSavedRegisters(MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, ArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI) const override
spillCalleeSavedRegisters - Issues instruction(s) to spill all callee saved registers and returns tru...
bool restoreCalleeSavedRegisters(MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, MutableArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI) const override
restoreCalleeSavedRegisters - Issues instruction(s) to restore all callee saved registers and returns...
bool enableFullCFIFixup(const MachineFunction &MF) const override
enableFullCFIFixup - Returns true if we may need to fix the unwind information such that it is accura...
StackOffset getFrameIndexReferenceFromSP(const MachineFunction &MF, int FI) const override
getFrameIndexReferenceFromSP - This method returns the offset from the stack pointer to the slot of t...
bool enableCFIFixup(const MachineFunction &MF) const override
Returns true if we may need to fix the unwind information for the function.
StackOffset getNonLocalFrameIndexReference(const MachineFunction &MF, int FI) const override
getNonLocalFrameIndexReference - This method returns the offset used to reference a frame index locat...
TargetStackID::Value getStackIDForScalableVectors() const override
Returns the StackID that scalable vectors should be associated with.
bool hasFPImpl(const MachineFunction &MF) const override
hasFPImpl - Return true if the specified function should have a dedicated frame pointer register.
void emitPrologue(MachineFunction &MF, MachineBasicBlock &MBB) const override
emitProlog/emitEpilog - These methods insert prolog and epilog code into the function.
void resetCFIToInitialState(MachineBasicBlock &MBB) const override
Emit CFI instructions that recreate the state of the unwind information upon function entry.
bool hasReservedCallFrame(const MachineFunction &MF) const override
hasReservedCallFrame - Under normal circumstances, when a frame pointer is not required,...
bool hasSVECalleeSavesAboveFrameRecord(const MachineFunction &MF) const
StackOffset resolveFrameOffsetReference(const MachineFunction &MF, int64_t ObjectOffset, bool isFixed, TargetStackID::Value StackID, Register &FrameReg, bool PreferFP, bool ForSimm) const
bool canUseRedZone(const MachineFunction &MF) const
Can this function use the red zone for local allocations.
bool needsWinCFI(const MachineFunction &MF) const
bool isFPReserved(const MachineFunction &MF) const
Should the Frame Pointer be reserved for the current function?
void processFunctionBeforeFrameFinalized(MachineFunction &MF, RegScavenger *RS) const override
processFunctionBeforeFrameFinalized - This method is called immediately before the specified function...
int getSEHFrameIndexOffset(const MachineFunction &MF, int FI) const
unsigned getWinEHFuncletFrameSize(const MachineFunction &MF) const
Funclets only need to account for space for the callee saved registers, as the locals are accounted f...
void orderFrameObjects(const MachineFunction &MF, SmallVectorImpl< int > &ObjectsToAllocate) const override
Order the symbols in the local stack frame.
void emitEpilogue(MachineFunction &MF, MachineBasicBlock &MBB) const override
StackOffset getPPRStackSize(const MachineFunction &MF) const
Returns the size of the entire PPR stackframe (calleesaves + spills + hazard padding).
int64_t getArgumentStackToRestore(MachineFunction &MF, MachineBasicBlock &MBB) const
Returns how much of the incoming argument stack area (in bytes) we should clean up in an epilogue.
void determineCalleeSaves(MachineFunction &MF, BitVector &SavedRegs, RegScavenger *RS) const override
This method determines which of the registers reported by TargetRegisterInfo::getCalleeSavedRegs() sh...
StackOffset getFrameIndexReference(const MachineFunction &MF, int FI, Register &FrameReg) const override
getFrameIndexReference - Provide a base+offset reference to an FI slot for debug info.
StackOffset getFrameIndexReferencePreferSP(const MachineFunction &MF, int FI, Register &FrameReg, bool IgnoreSPUpdates) const override
For Win64 AArch64 EH, the offset to the Unwind object is from the SP before the update.
StackOffset resolveFrameIndexReference(const MachineFunction &MF, int FI, Register &FrameReg, bool PreferFP, bool ForSimm) const
unsigned getWinEHParentFrameOffset(const MachineFunction &MF) const override
The parent frame offset (aka dispFrame) is only used on X86_64 to retrieve the parent's frame pointer...
bool requiresSaveVG(const MachineFunction &MF) const
void emitPacRetPlusLeafHardening(MachineFunction &MF) const
Harden the entire function with pac-ret.
AArch64FunctionInfo - This class is derived from MachineFunctionInfo and contains private AArch64-spe...
unsigned getCalleeSavedStackSize(const MachineFrameInfo &MFI) const
void setCalleeSaveBaseToFrameRecordOffset(int Offset)
SignReturnAddress getSignReturnAddressCondition() const
void setStackSizeSVE(uint64_t ZPR, uint64_t PPR)
std::optional< int > getTaggedBasePointerIndex() const
bool needsDwarfUnwindInfo(const MachineFunction &MF) const
void setSVECalleeSavedStackSize(unsigned ZPR, unsigned PPR)
bool needsAsyncDwarfUnwindInfo(const MachineFunction &MF) const
static bool isTailCallReturnInst(const MachineInstr &MI)
Returns true if MI is one of the TCRETURN* instructions.
static bool isFpOrNEON(Register Reg)
Returns whether the physical register is FP or NEON.
const AArch64RegisterInfo * getRegisterInfo() const override
bool isNeonAvailable() const
Returns true if the target has NEON and the function at runtime is known to have NEON enabled (e....
const AArch64InstrInfo * getInstrInfo() const override
const AArch64TargetLowering * getTargetLowering() const override
bool isSVEorStreamingSVEAvailable() const
Returns true if the target has access to either the full range of SVE instructions,...
bool isStreaming() const
Returns true if the function has a streaming body.
bool hasInlineStackProbe(const MachineFunction &MF) const override
True if stack clash protection is enabled for this functions.
unsigned getRedZoneSize(const Function &F) const
Represent a constant reference to an array (0 or more elements consecutively in memory),...
Definition ArrayRef.h:40
size_t size() const
Get the array size.
Definition ArrayRef.h:141
bool empty() const
Check if the array is empty.
Definition ArrayRef.h:136
bool test(unsigned Idx) const
Returns true if bit Idx is set.
Definition BitVector.h:482
BitVector & reset()
Reset all bits in the bitvector.
Definition BitVector.h:409
size_type count() const
Returns the number of bits which are set.
Definition BitVector.h:181
BitVector & set()
Set all bits in the bitvector.
Definition BitVector.h:366
iterator_range< const_set_bits_iterator > set_bits() const
Definition BitVector.h:159
size_type size() const
Returns the number of bits in this bitvector.
Definition BitVector.h:178
Helper class for creating CFI instructions and inserting them into MIR.
The CalleeSavedInfo class tracks the information need to locate where a callee saved register is in t...
A debug info location.
Definition DebugLoc.h:126
bool hasMinSize() const
Optimize this function for minimum size (-Oz).
Definition Function.h:688
CallingConv::ID getCallingConv() const
getCallingConv()/setCallingConv(CC) - These method get and set the calling convention of this functio...
Definition Function.h:272
AttributeList getAttributes() const
Return the attribute list for this Function.
Definition Function.h:328
bool isVarArg() const
isVarArg - Return true if this function takes a variable number of arguments.
Definition Function.h:229
bool hasFnAttribute(Attribute::AttrKind Kind) const
Return true if the function has the attribute.
Definition Function.cpp:727
A set of physical registers with utility functions to track liveness when walking backward/forward th...
bool usesWindowsCFI() const
Definition MCAsmInfo.h:675
Wrapper class representing physical registers. Should be passed by value.
Definition MCRegister.h:41
LLVM_ABI void transferSuccessorsAndUpdatePHIs(MachineBasicBlock *FromMBB)
Transfers all the successors, as in transferSuccessors, and update PHI operands in the successor bloc...
LLVM_ABI iterator getFirstTerminator()
Returns an iterator to the first terminator instruction of this basic block.
LLVM_ABI void addSuccessor(MachineBasicBlock *Succ, BranchProbability Prob=BranchProbability::getUnknown())
Add Succ as a successor of this MachineBasicBlock.
const MachineFunction * getParent() const
Return the MachineFunction containing this basic block.
reverse_iterator rbegin()
iterator insertAfter(iterator I, MachineInstr *MI)
Insert MI into the instruction list after I.
void splice(iterator Where, MachineBasicBlock *Other, iterator From)
Take an instruction from MBB 'Other' at the position From, and insert it into this MBB right before '...
MachineInstrBundleIterator< MachineInstr > iterator
The MachineFrameInfo class represents an abstract stack frame until prolog/epilog code is inserted.
LLVM_ABI int CreateFixedObject(uint64_t Size, int64_t SPOffset, bool IsImmutable, bool isAliased=false)
Create a new object at a fixed location on the stack.
bool hasVarSizedObjects() const
This method may be called any time after instruction selection is complete to determine if the stack ...
const AllocaInst * getObjectAllocation(int ObjectIdx) const
Return the underlying Alloca of the specified stack object if it exists.
LLVM_ABI int CreateStackObject(uint64_t Size, Align Alignment, bool isSpillSlot, const AllocaInst *Alloca=nullptr, uint8_t ID=0)
Create a new statically sized stack object, returning a nonnegative identifier to represent it.
bool hasCalls() const
Return true if the current function has any function calls.
bool isFrameAddressTaken() const
This method may be called any time after instruction selection is complete to determine if there is a...
void setObjectOffset(int ObjectIdx, int64_t SPOffset)
Set the stack frame offset of the specified object.
bool isCalleeSavedObjectIndex(int ObjectIdx) const
uint64_t getMaxCallFrameSize() const
Return the maximum size of a call frame that must be allocated for an outgoing function call.
bool hasPatchPoint() const
This method may be called any time after instruction selection is complete to determine if there is a...
bool hasScalableStackID(int ObjectIdx) const
int getStackProtectorIndex() const
Return the index for the stack protector object.
LLVM_ABI uint64_t estimateStackSize(const MachineFunction &MF) const
Estimate and return the size of the stack frame.
void setStackID(int ObjectIdx, uint8_t ID)
bool isCalleeSavedInfoValid() const
Has the callee saved info been calculated yet?
Align getObjectAlign(int ObjectIdx) const
Return the alignment of the specified stack object.
int64_t getObjectSize(int ObjectIdx) const
Return the size of the specified object.
bool isMaxCallFrameSizeComputed() const
bool hasStackMap() const
This method may be called any time after instruction selection is complete to determine if there is a...
LLVM_ABI int CreateSpillStackObject(uint64_t Size, Align Alignment, TargetStackID::Value StackID=TargetStackID::Default)
Create a new statically sized stack object that represents a spill slot, returning a nonnegative iden...
const std::vector< CalleeSavedInfo > & getCalleeSavedInfo() const
Returns a reference to call saved info vector for the current function.
unsigned getNumObjects() const
Return the number of objects.
int getObjectIndexEnd() const
Return one past the maximum frame object index.
bool hasStackProtectorIndex() const
bool hasStackObjects() const
Return true if there are any stack objects in this function.
uint8_t getStackID(int ObjectIdx) const
unsigned getNumFixedObjects() const
Return the number of fixed objects.
void setIsCalleeSavedObjectIndex(int ObjectIdx, bool IsCalleeSaved)
int64_t getObjectOffset(int ObjectIdx) const
Return the assigned stack offset of the specified object from the incoming stack pointer.
int getObjectIndexBegin() const
Return the minimum frame object index.
void setObjectAlignment(int ObjectIdx, Align Alignment)
setObjectAlignment - Change the alignment of the specified stack object.
bool isDeadObjectIndex(int ObjectIdx) const
Returns true if the specified index corresponds to a dead object.
const WinEHFuncInfo * getWinEHFuncInfo() const
getWinEHFuncInfo - Return information about how the current function uses Windows exception handling.
const TargetSubtargetInfo & getSubtarget() const
getSubtarget - Return the subtarget for which this machine code is being compiled.
MachineMemOperand * getMachineMemOperand(MachinePointerInfo PtrInfo, MachineMemOperand::Flags f, LLT MemTy, Align base_alignment, const AAMDNodes &AAInfo=AAMDNodes(), const MDNode *Ranges=nullptr, SyncScope::ID SSID=SyncScope::System, AtomicOrdering Ordering=AtomicOrdering::NotAtomic, AtomicOrdering FailureOrdering=AtomicOrdering::NotAtomic)
getMachineMemOperand - Allocate a new MachineMemOperand.
MachineFrameInfo & getFrameInfo()
getFrameInfo - Return the frame info object for the current function.
MachineRegisterInfo & getRegInfo()
getRegInfo - Return information about the registers currently in use.
Function & getFunction()
Return the LLVM function that this machine code represents.
BasicBlockListType::iterator iterator
Ty * getInfo()
getInfo - Keep track of various per-function pieces of information for backends that would like to do...
const MachineBasicBlock & front() const
MachineBasicBlock * CreateMachineBasicBlock(const BasicBlock *BB=nullptr, std::optional< UniqueBBID > BBID=std::nullopt)
CreateMachineInstr - Allocate a new MachineInstr.
void insert(iterator MBBI, MachineBasicBlock *MBB)
const TargetMachine & getTarget() const
getTarget - Return the target machine this machine code is compiled with
const MachineInstrBuilder & setMemRefs(ArrayRef< MachineMemOperand * > MMOs) const
const MachineInstrBuilder & addExternalSymbol(const char *FnName, unsigned TargetFlags=0) const
const MachineInstrBuilder & addReg(Register RegNo, RegState Flags={}, unsigned SubReg=0) const
Add a new virtual register operand.
const MachineInstrBuilder & setMIFlag(MachineInstr::MIFlag Flag) const
const MachineInstrBuilder & addImm(int64_t Val) const
Add a new immediate operand.
const MachineInstrBuilder & addFrameIndex(int Idx) const
const MachineInstrBuilder & addRegMask(const uint32_t *Mask) const
const MachineInstrBuilder & addMBB(MachineBasicBlock *MBB, unsigned TargetFlags=0) const
const MachineInstrBuilder & addDef(Register RegNo, RegState Flags={}, unsigned SubReg=0) const
Add a virtual register definition operand.
const MachineInstrBuilder & setMIFlags(unsigned Flags) const
const MachineInstrBuilder & addMemOperand(MachineMemOperand *MMO) const
Representation of each machine instruction.
void setFlags(unsigned flags)
uint32_t getFlags() const
Return the MI flags bitvector.
LLVM_ABI MachineInstrBundleIterator< MachineInstr > eraseFromParent()
Unlink 'this' from the containing basic block and delete it.
A description of a memory reference used in the backend.
const PseudoSourceValue * getPseudoValue() const
@ MOVolatile
The memory access is volatile.
@ MOLoad
The memory access reads data.
@ MOStore
The memory access writes data.
const Value * getValue() const
Return the base address of the memory access.
MachineOperand class - Representation of each machine instruction operand.
int64_t getImm() const
bool isFI() const
isFI - Tests if this is a MO_FrameIndex operand.
LLVM_ABI void emit(DiagnosticInfoOptimizationBase &OptDiag)
Emit an optimization remark.
MachineRegisterInfo - Keep track of information for virtual and physical registers,...
LLVM_ABI void freezeReservedRegs()
freezeReservedRegs - Called by the register allocator to freeze the set of reserved registers before ...
bool isReserved(MCRegister PhysReg) const
isReserved - Returns true when PhysReg is a reserved register.
LLVM_ABI Register createVirtualRegister(const TargetRegisterClass *RegClass, StringRef Name="")
createVirtualRegister - Create and return a new virtual register in the function with the specified r...
LLVM_ABI bool isLiveIn(Register Reg) const
LLVM_ABI const MCPhysReg * getCalleeSavedRegs() const
Returns list of callee saved registers.
LLVM_ABI bool isPhysRegUsed(MCRegister PhysReg, bool SkipRegMaskTest=false) const
Return true if the specified register is modified or read in this function.
Represent a mutable reference to an array (0 or more elements consecutively in memory),...
Definition ArrayRef.h:294
Wrapper class representing virtual and physical registers.
Definition Register.h:20
constexpr bool isValid() const
Definition Register.h:112
SMEAttrs is a utility class to parse the SME ACLE attributes on functions.
bool hasStreamingInterface() const
bool hasNonStreamingInterfaceAndBody() const
bool hasStreamingBody() const
bool insert(const value_type &X)
Insert a new element into the SetVector.
Definition SetVector.h:157
A SetVector that performs no allocations if smaller than a certain size.
Definition SetVector.h:345
This class consists of common code factored out of the SmallVector class to reduce code duplication b...
reference emplace_back(ArgTypes &&... Args)
void append(ItTy in_start, ItTy in_end)
Add the specified range to the end of the SmallVector.
void push_back(const T &Elt)
This is a 'vector' (really, a variable-sized array), optimized for the case when the array is small.
StackOffset holds a fixed and a scalable offset in bytes.
Definition TypeSize.h:30
int64_t getFixed() const
Returns the fixed component of the stack.
Definition TypeSize.h:46
int64_t getScalable() const
Returns the scalable component of the stack.
Definition TypeSize.h:49
static StackOffset get(int64_t Fixed, int64_t Scalable)
Definition TypeSize.h:41
static StackOffset getScalable(int64_t Scalable)
Definition TypeSize.h:40
static StackOffset getFixed(int64_t Fixed)
Definition TypeSize.h:39
bool hasFP(const MachineFunction &MF) const
hasFP - Return true if the specified function should have a dedicated frame pointer register.
virtual void determineCalleeSaves(MachineFunction &MF, BitVector &SavedRegs, RegScavenger *RS=nullptr) const
This method determines which of the registers reported by TargetRegisterInfo::getCalleeSavedRegs() sh...
int getOffsetOfLocalArea() const
getOffsetOfLocalArea - This method returns the offset of the local area from the stack pointer on ent...
Align getStackAlign() const
getStackAlignment - This method returns the number of bytes to which the stack pointer must be aligne...
StackDirection getStackGrowthDirection() const
getStackGrowthDirection - Return the direction the stack grows
virtual bool enableCFIFixup(const MachineFunction &MF) const
Returns true if we may need to fix the unwind information for the function.
Primary interface to the complete machine description for the target machine.
const Triple & getTargetTriple() const
const MCAsmInfo & getMCAsmInfo() const
Return target specific asm information.
TargetOptions Options
LLVM_ABI bool FramePointerIsReserved(const MachineFunction &MF) const
FramePointerIsReserved - This returns true if the frame pointer must always either point to a new fra...
LLVM_ABI bool DisableFramePointerElim(const MachineFunction &MF) const
DisableFramePointerElim - This returns true if frame pointer elimination optimization should be disab...
TargetRegisterInfo base class - We assume that the target defines a static array of TargetRegisterDes...
bool hasStackRealignment(const MachineFunction &MF) const
True if stack realignment is required and still possible.
virtual const TargetRegisterInfo * getRegisterInfo() const =0
Return the target's register information.
Triple - Helper class for working with autoconf configuration names.
Definition Triple.h:48
bool isOSBinFormatMachO() const
Tests whether the environment is MachO.
Definition Triple.h:873
This class implements an extremely fast bulk output stream that can only output to a stream.
Definition raw_ostream.h:53
#define llvm_unreachable(msg)
Marks that the current location is not supposed to be reachable.
static unsigned getShiftValue(unsigned Imm)
getShiftValue - Extract the shift value.
static unsigned getArithExtendImm(AArch64_AM::ShiftExtendType ET, unsigned Imm)
getArithExtendImm - Encode the extend type and shift amount for an arithmetic instruction: imm: 3-bit...
const unsigned StackProbeMaxLoopUnroll
Maximum number of iterations to unroll for a constant size probing loop.
const unsigned StackProbeMaxUnprobedStack
Maximum allowed number of unprobed bytes above SP at an ABI boundary.
constexpr char Align[]
Key for Kernel::Arg::Metadata::mAlign.
constexpr char Attrs[]
Key for Kernel::Metadata::mAttrs.
unsigned ID
LLVM IR allows to use arbitrary numbers as calling convention identifiers.
Definition CallingConv.h:24
@ AArch64_SVE_VectorCall
Used between AArch64 SVE functions.
@ PreserveMost
Used for runtime calls that preserves most registers.
Definition CallingConv.h:63
@ CXX_FAST_TLS
Used for access functions.
Definition CallingConv.h:72
@ GHC
Used by the Glasgow Haskell Compiler (GHC).
Definition CallingConv.h:50
@ PreserveAll
Used for runtime calls that preserves (almost) all registers.
Definition CallingConv.h:66
@ Fast
Attempts to make calls as fast as possible (e.g.
Definition CallingConv.h:41
@ PreserveNone
Used for runtime calls that preserves none general registers.
Definition CallingConv.h:90
@ Win64
The C convention as implemented on Windows/x86-64 and AArch64.
@ SwiftTail
This follows the Swift calling convention in how arguments are passed but guarantees tail calls will ...
Definition CallingConv.h:87
@ C
The default llvm calling convention, compatible with C.
Definition CallingConv.h:34
initializer< Ty > init(const Ty &Val)
NodeAddr< InstrNode * > Instr
Definition RDFGraph.h:389
BaseReg
Stack frame base register. Bit 0 of FREInfo.Info.
Definition SFrame.h:77
This is an optimization pass for GlobalISel generic memory operations.
@ Offset
Definition DWP.cpp:578
void stable_sort(R &&Range)
Definition STLExtras.h:2116
MachineInstrBuilder BuildMI(MachineFunction &MF, const MIMetadata &MIMD, const MCInstrDesc &MCID)
Builder interface. Specify how to create the initial instruction itself.
int isAArch64FrameOffsetLegal(const MachineInstr &MI, StackOffset &Offset, bool *OutUseUnscaledOp=nullptr, unsigned *OutUnscaledOp=nullptr, int64_t *EmittableOffset=nullptr)
Check if the Offset is a valid frame offset for MI.
@ Unknown
Not known to have no common set bits.
RegState
Flags to represent properties of register accesses.
@ Define
Register definition.
constexpr RegState getKillRegState(bool B)
decltype(auto) dyn_cast(const From &Val)
dyn_cast<X> - Return the argument parameter cast to the specified type.
Definition Casting.h:643
@ AArch64FrameOffsetCannotUpdate
Offset cannot apply.
constexpr T alignDown(U Value, V Align, W Skew=0)
Returns the largest unsigned integer less than or equal to Value and is Skew mod Align.
Definition MathExtras.h:547
auto dyn_cast_or_null(const Y &Val)
Definition Casting.h:753
bool any_of(R &&range, UnaryPredicate P)
Provide wrappers to std::any_of which take ranges instead of having to pass begin/end explicitly.
Definition STLExtras.h:1746
auto formatv(bool Validate, const char *Fmt, Ts &&...Vals)
auto reverse(ContainerTy &&C)
Definition STLExtras.h:407
void sort(IteratorTy Start, IteratorTy End)
Definition STLExtras.h:1636
LLVM_ABI raw_ostream & dbgs()
dbgs() - This returns a reference to a raw_ostream for debugging messages.
Definition Debug.cpp:209
void emitFrameOffset(MachineBasicBlock &MBB, MachineBasicBlock::iterator MBBI, const DebugLoc &DL, unsigned DestReg, unsigned SrcReg, StackOffset Offset, const TargetInstrInfo *TII, MachineInstr::MIFlag=MachineInstr::NoFlags, bool SetNZCV=false, bool NeedsWinCFI=false, bool *HasWinCFI=nullptr, bool EmitCFAOffset=false, StackOffset InitialOffset={}, unsigned FrameReg=AArch64::SP)
emitFrameOffset - Emit instructions as needed to set DestReg to SrcReg plus Offset.
LLVM_ABI void report_fatal_error(Error Err, bool gen_crash_diag=true)
Definition Error.cpp:163
constexpr uint64_t alignTo(uint64_t Size, Align A)
Returns a multiple of A needed to store Size bytes.
Definition Alignment.h:144
constexpr RegState getDefRegState(bool B)
class LLVM_GSL_OWNER SmallVector
Forward declaration of SmallVector so that calculateSmallVectorDefaultInlinedElements can reference s...
@ First
Helpers to iterate all locations in the MemoryEffectsBase class.
Definition ModRef.h:74
uint16_t MCPhysReg
An unsigned integer type large enough to represent all physical registers, but not necessarily virtua...
Definition MCRegister.h:21
RelativeUniformCounterPtr ValuesPtrExpr VTableAddr Count
Definition InstrProf.h:145
raw_ostream & operator<<(raw_ostream &OS, const APFixedPoint &FX)
auto count_if(R &&Range, UnaryPredicate P)
Wrapper function around std::count_if to count the number of times an element satisfying a given pred...
Definition STLExtras.h:2019
auto find_if(R &&Range, UnaryPredicate P)
Provide wrappers to std::find_if which take ranges instead of having to pass begin/end explicitly.
Definition STLExtras.h:1772
void erase_if(Container &C, UnaryPredicate P)
Provide a container algorithm similar to C++ Library Fundamentals v2's erase_if which is equivalent t...
Definition STLExtras.h:2192
bool is_contained(R &&Range, const E &Element)
Returns true if Element is found in Range.
Definition STLExtras.h:1947
LLVM_ABI const Value * getUnderlyingObject(const Value *V, unsigned MaxLookup=MaxLookupSearchDepth)
This method strips off any GEP address adjustments, pointer casts or llvm.threadlocal....
void fullyRecomputeLiveIns(ArrayRef< MachineBasicBlock * > MBBs)
Convenience function for recomputing live-in's for a set of MBBs until the computation converges.
LLVM_ABI Printable printReg(Register Reg, const TargetRegisterInfo *TRI=nullptr, unsigned SubIdx=0, const MachineRegisterInfo *MRI=nullptr)
Prints virtual and physical registers with or without a TRI instance.
MCRegisterClass TargetRegisterClass
Definition FastISel.h:58
void swap(llvm::BitVector &LHS, llvm::BitVector &RHS)
Implement std::swap in terms of BitVector swap.
Definition BitVector.h:880
bool operator<(const StackAccess &Rhs) const
void print(raw_ostream &OS) const
int64_t start() const
std::string getTypeString() const
int64_t end() const
This struct is a compact representation of a valid (non-zero power of two) alignment.
Definition Alignment.h:39
constexpr uint64_t value() const
This is a hole in the type system and should not be abused.
Definition Alignment.h:77
Pair of physical register and lane mask.
static LLVM_ABI MachinePointerInfo getUnknownStack(MachineFunction &MF)
Stack memory without other information.
static LLVM_ABI MachinePointerInfo getFixedStack(MachineFunction &MF, int FI, int64_t Offset=0)
Return a MachinePointerInfo record that refers to the specified FrameIndex.
SmallVector< WinEHTryBlockMapEntry, 4 > TryBlockMap
SmallVector< WinEHHandlerType, 1 > HandlerArray