|
LLVM 24.0.0git
|
CandidateHeuristics contains state and implementations to facilitate making per instruction scheduling decisions; it contains methods used in tryCandidate to decide which instruction to schedule next. More...
#include "Target/AMDGPU/AMDGPUCoExecSchedStrategy.h"
Protected Member Functions | |
| void | collectRegionSummary () |
| Walk over the region and collect characteristics for the various heuristics. | |
| unsigned | getMaxBlockingCycles (const MCSchedClassDesc *SC, const MachineInstr *MI) |
| unsigned | getHWUICyclesForSU (SUnit *SU) |
Compute the blocking cycles for the appropriate HardwareUnit given an SU. | |
| unsigned | getHWUICyclesForMI (MachineInstr *MI) |
Compute the blocking cycles for the appropriate HardwareUnit given an MI. | |
| unsigned | getCarriedLatency (SUnit *SU) |
Estimate the block carried latency from loads for a given SU. | |
Protected Attributes | |
| ScheduleDAGMI * | DAG |
| const SIInstrInfo * | SII |
| const SIRegisterInfo * | SRI |
| const TargetSchedModel * | SchedModel |
| SmallVector< HardwareUnitInfo, 8 > | HWUInfo |
| DenseMap< MachineInstr *, unsigned > | CarriedLatencies |
CandidateHeuristics contains state and implementations to facilitate making per instruction scheduling decisions; it contains methods used in tryCandidate to decide which instruction to schedule next.
Definition at line 199 of file AMDGPUCoExecSchedStrategy.h.
|
default |
References DAG, SchedModel, and TRI.
|
protected |
Walk over the region and collect characteristics for the various heuristics.
Definition at line 708 of file AMDGPUCoExecSchedStrategy.cpp.
References CarriedLatencies, llvm::AMDGPU::classifyFlavor(), DAG, dumpRegionSummary(), getCarriedLatency(), getHWUICyclesForSU(), HWUInfo, LLVM_DEBUG, MI, SchedModel, and SII.
Referenced by initialize().
| void CandidateHeuristics::dumpRegionSummary | ( | ) |
Definition at line 728 of file AMDGPUCoExecSchedStrategy.cpp.
References DAG, llvm::dbgs(), llvm::AMDGPU::getFlavorName(), llvm::MachineBasicBlock::getNumber(), and HWUInfo.
Referenced by collectRegionSummary().
Estimate the block carried latency from loads for a given SU.
This is essentially global scheduling info that our local scheduling infrastructure lacks the necessary infrastructure to accurately measure. Thus, this method just attempts to find a reasonable upper bound for carried load latency to avoid long stalls.
Definition at line 629 of file AMDGPUCoExecSchedStrategy.cpp.
References assert(), BlockCarriedLatency, llvm::AMDGPU::classifyFlavor(), DAG, llvm::AMDGPU::DS, llvm::AMDGPU::Fence, getHWUICyclesForMI(), llvm::SUnit::getInstr(), I, llvm::SlotIndex::isEarlierInstr(), llvm::Latency, MBB, MI, and SII.
Referenced by collectRegionSummary().
|
protected |
Compute the blocking cycles for the appropriate HardwareUnit given an MI.
Definition at line 588 of file AMDGPUCoExecSchedStrategy.cpp.
References assert(), getMaxBlockingCycles(), MI, and SchedModel.
Referenced by getCarriedLatency().
Compute the blocking cycles for the appropriate HardwareUnit given an SU.
Definition at line 580 of file AMDGPUCoExecSchedStrategy.cpp.
References assert(), DAG, llvm::SUnit::getInstr(), getMaxBlockingCycles(), MI, SchedModel, and SII.
Referenced by collectRegionSummary(), and updateForScheduling().
| HardwareUnitInfo * CandidateHeuristics::getHWUIFromFlavor | ( | AMDGPU::InstructionFlavor | Flavor | ) |
Given a Flavor , find the corresponding HardwareUnit.
Definition at line 555 of file AMDGPUCoExecSchedStrategy.cpp.
References HWUInfo.
Referenced by tryEffectiveStall(), and updateForScheduling().
|
protected |
SC. Definition at line 564 of file AMDGPUCoExecSchedStrategy.cpp.
References MI, and SchedModel.
Referenced by getHWUICyclesForMI(), and getHWUICyclesForSU().
| unsigned CandidateHeuristics::getStructuralStallCycles | ( | SchedBoundary & | Zone, |
| SUnit * | SU ) |
Definition at line 766 of file AMDGPUCoExecSchedStrategy.cpp.
References DAG, llvm::SchedBoundary::getCurrCycle(), llvm::SUnit::getInstr(), llvm::SchedBoundary::getNextResourceCycle(), llvm::SUnit::hasReservedResource, llvm::SchedBoundary::HazardRec, llvm::SchedBoundary::isTop(), llvm::make_range(), MI, and SchedModel.
Referenced by tryEffectiveStall().
| void CandidateHeuristics::initialize | ( | ScheduleDAGMI * | DAG, |
| const TargetSchedModel * | SchedModel, | ||
| const TargetRegisterInfo * | TRI ) |
Definition at line 600 of file AMDGPUCoExecSchedStrategy.cpp.
References assert(), collectRegionSummary(), DAG, llvm::AMDGPU::DefaultBufferSizes::DS, llvm::AMDGPU::DS, HWUInfo, I, llvm::AMDGPU::MultiCycleVALU, llvm::AMDGPU::NUM_FLAVORS, SchedModel, SII, SRI, llvm::AMDGPU::TRANS, TRI, and llvm::AMDGPU::WMMA.
| void CandidateHeuristics::sortHWUIResources | ( | ) |
Sort the HardwarUnitInfo vector.
After sorting, the HWUI that are highest priority are first. Priority is determined by maximizing coexecution and keeping the critical HardwareUnit busy.
Definition at line 745 of file AMDGPUCoExecSchedStrategy.cpp.
References A(), B(), HWUInfo, and llvm::sort().
| bool CandidateHeuristics::tryCriticalResource | ( | GenericSchedulerBase::SchedCandidate & | TryCand, |
| GenericSchedulerBase::SchedCandidate & | Cand, | ||
| SchedBoundary * | Zone ) const |
Check for critical resource consumption.
Prefer the candidate that uses the most prioritized HardwareUnit. If both candidates use the same HarwareUnit, prefer the candidate with higher priority on that HardwareUnit.
Definition at line 955 of file AMDGPUCoExecSchedStrategy.cpp.
References llvm::HardwareUnitInfo::contains(), llvm::AMDGPU::DS, llvm::HardwareUnitInfo::getHigherPriority(), llvm::HardwareUnitInfo::getType(), HWUInfo, I, llvm::GenericSchedulerBase::SchedCandidate::Reason, llvm::GenericSchedulerBase::RegCritical, llvm::GenericSchedulerBase::SchedCandidate::SU, and tryCriticalResourceDependency().
| bool CandidateHeuristics::tryCriticalResourceDependency | ( | GenericSchedulerBase::SchedCandidate & | TryCand, |
| GenericSchedulerBase::SchedCandidate & | Cand, | ||
| SchedBoundary * | Zone ) const |
Check for dependencies of instructions that use prioritized HardwareUnits.
Prefer the candidate that is a dependency of an instruction that uses the most prioritized HardwareUnit. If both candidates enable the same HardwareUnit, prefer the candidate that enables the higher priority instruction on that HardwareUnit.
Definition at line 868 of file AMDGPUCoExecSchedStrategy.cpp.
References llvm::AMDGPU::classifyFlavor(), DAG, llvm::AMDGPU::DS, llvm::Enabled, llvm::SUnit::getHeight(), llvm::SUnit::getInstr(), llvm::HardwareUnitInfo::getNextTargetSU(), llvm::HardwareUnitInfo::getType(), HWUInfo, I, llvm::GenericSchedulerBase::SchedCandidate::Reason, llvm::GenericSchedulerBase::RegCritical, SII, llvm::GenericSchedulerBase::SchedCandidate::SU, and llvm::AMDGPU::WMMA.
Referenced by tryCriticalResource().
| bool CandidateHeuristics::tryEffectiveStall | ( | GenericSchedulerBase::SchedCandidate & | Cand, |
| GenericSchedulerBase::SchedCandidate & | TryCand, | ||
| SchedBoundary & | Zone ) |
Definition at line 800 of file AMDGPUCoExecSchedStrategy.cpp.
References assert(), CarriedLatencies, llvm::AMDGPU::classifyFlavor(), DAG, llvm::HardwareUnitInfo::getBufferAvailableCycle(), llvm::HardwareUnitInfo::getBufferSize(), llvm::SchedBoundary::getCurrCycle(), getHWUIFromFlavor(), llvm::SchedBoundary::getLatencyStallCycles(), getStructuralStallCycles(), llvm::SchedBoundary::isTop(), llvm::Latency, LLVM_DEBUG, llvm::GenericSchedulerBase::Stall, llvm::GenericSchedulerBase::SchedCandidate::SU, and llvm::tryLess().
| void CandidateHeuristics::updateForScheduling | ( | SUnit * | SU | ) |
Update the state to reflect that SU is going to be scheduled.
Definition at line 593 of file AMDGPUCoExecSchedStrategy.cpp.
References assert(), llvm::AMDGPU::classifyFlavor(), getHWUICyclesForSU(), getHWUIFromFlavor(), llvm::SUnit::getInstr(), llvm::HardwareUnitInfo::markScheduled(), and SII.
|
protected |
Definition at line 206 of file AMDGPUCoExecSchedStrategy.h.
Referenced by collectRegionSummary(), and tryEffectiveStall().
|
protected |
Definition at line 201 of file AMDGPUCoExecSchedStrategy.h.
Referenced by CandidateHeuristics(), collectRegionSummary(), dumpRegionSummary(), getCarriedLatency(), getHWUICyclesForSU(), getStructuralStallCycles(), initialize(), tryCriticalResourceDependency(), and tryEffectiveStall().
|
protected |
Definition at line 205 of file AMDGPUCoExecSchedStrategy.h.
Referenced by collectRegionSummary(), dumpRegionSummary(), getHWUIFromFlavor(), initialize(), sortHWUIResources(), tryCriticalResource(), and tryCriticalResourceDependency().
|
protected |
Definition at line 204 of file AMDGPUCoExecSchedStrategy.h.
Referenced by CandidateHeuristics(), collectRegionSummary(), getHWUICyclesForMI(), getHWUICyclesForSU(), getMaxBlockingCycles(), getStructuralStallCycles(), and initialize().
|
protected |
Definition at line 202 of file AMDGPUCoExecSchedStrategy.h.
Referenced by collectRegionSummary(), getCarriedLatency(), getHWUICyclesForSU(), initialize(), tryCriticalResourceDependency(), and updateForScheduling().
|
protected |
Definition at line 203 of file AMDGPUCoExecSchedStrategy.h.
Referenced by initialize().