- Jun 11, 2024
-
-
aengelke authored
Copying an MCInst isn't cheap (copies all operands) and the whole instruction is only used for the Intel erratum mitigation, which is off by default. In all other cases, the opcode alone suffices. This slightly pessimizes code that uses moves to segment registers -- but that's uncommon and not performance-sensitive anyway. As a related change, also call canPadInst() only when the result is actually used, which is typically only the case in emitInstrEnd. This gives a minor performance improvement.
-
aengelke authored
This allows tested passes to depend on the AMX model in the function info. Preparatory work for to adopt #94358 for other AMX passes.
-
Simon Pilgrim authored
-
Jay Foad authored
-
c8ef authored
Follow up of #94887. Context: https://github.com/llvm/llvm-project/pull/94887#pullrequestreview-2106213891
-
Simon Pilgrim authored
This was previously set to 4uops which was including the cost of extra register moves in the original test code.
-
Pavel Samolysov authored
This addresses a review comment for PR #94987 Because that PR is a big automatic change, this change was moved in a separate one.
-
Timm Bäder authored
This is still not perfect, but an improvement in general.
-
Jay Foad authored
Use DEFINE: %{res} = ... instead of $(cat ...). Rewrite one use of a subshell to write to a temporary file instead. -
Braden Helmer authored
Fixes #95036 #95033 #94933 #94930
-
Arjun P authored
This matches the other two functions AddOverflow and SubOverflow.
-
Sander de Smalen authored
This was pointed out in PR #93940.
-
Durgadoss R authored
This patch adds APFloat type support for two FP6 data types, E2M3 and E3M2. The definitions for the two formats are detailed in section 5.3.2 of the OCP specification, which can be accessed here: https://www.opencompute.org/documents/ocp-microscaling-formats-mx-v1-0-spec-final-pdf Signed-off-by:
Durgadoss R <durgadossr@nvidia.com>
-
Mikael Holmen authored
Without this gcc (9.3.0) warns with ../lib/Target/AArch64/GISel/AArch64CallLowering.cpp: In function 'unsigned int getCallOpcode(const llvm::MachineFunction&, bool, bool, std::optional<llvm::CallLowering::PtrAuthInfo>&, llvm::MachineRegisterInfo&)': ../lib/Target/AArch64/GISel/AArch64CallLowering.cpp:1025: warning: enumeral and non-enumeral type in conditional expression [-Wextra] 1025 | return IsIndirect ? getBLRCallOpcode(CallerF) : AArch64::BL; | -
Shivam Gupta authored
This commit adds a test for lea_rsp_pattern_p which was previously due as FIXME.
-
Abid Qadeer authored
This PR generates dwarf to extract the information about the arrays from descriptor. The DWARF needs the offset of the fields like `lower_bound` and `extent`. The getComponentOffset has been added to calculate them which pushes the issue of host and target data size into getDescFieldTypeModel. As we use data layout now, some tests needed to be adjusted to have a dummy data layout to avoid failure. With this change in place, GDB is able show the assumed shape arrays correctly. subroutine ff(n, m, arr) integer n, m integer :: arr(:, :) print *, arr do i = 1, n do j = 1, m arr(j, i) = (i * 5) + j + 10 end do end do print *, arr end subroutine ff Breakpoint 1, ff (n=4, m=3, arr=...) at test1.f90:13 13 print *, arr (gdb) p arr $1 = ((6, 7, 8, 9) (11, 12, 13, 14) (16, 17, 18, 19)) (gdb) ptype arr type = integer (4,3) (gdb) c Continuing. 6 7 8 9 11 12 13 14 16 17 18 19 -
jeanPerier authored
-
Vikash Gupta authored
The pass pipeline of some architecture splits register allocation phase based on different register classes. As some analyses need to be computed at the beginning of the register allocation and kept alive till all values are assigned to some physical registers. This poses challenge with objective of introducing StackSlotColoring after partial virtual registers are assigned to physical registers, in order to optimize stack slots usage.As this pass doesn't preserve few analysis yet to be needed by the register allocation of the remaining virtual registers, necessiating them to be kept preserved.
-
martinboehme authored
The patch includes a repro for a case where we were returning a null `FieldDecl` when calling `getReferencedDecls()` on the `InitListExpr` for a union. Also, I noticed while working on this that `RecordInitListHelper` has a bug where it doesn't work correctly for empty unions. This patch also includes a repro and fix for this bug.
-
aengelke authored
This avoids std::map, which is slow, and uses a StringMap. Section name, group name, linked-to name and unique id are encoded into the key for fast lookup. This gives a measurable performance boost for applications that compile many small object files (e.g., functions in JIT compilers). --- Now also the second case works properly. That's what happens when you do that last refactoring without re-running all tests... sorry.
-
martinboehme authored
This is one of the node kinds that should be considered an "original initializer". The patch adds a test that was causing an assertion failure in `assert(Children.size() == 1)` without the fix.
-
Pengcheng Wang authored
According to RVV spec: > In general, the requirement is to support LMUL ≥ SEWMIN/ELEN, > where SEWMIN is the narrowest supported SEW value and ELEN is > the widest supported SEW value. > > For a given supported fractional LMUL setting, implementations > must support SEW settings between SEWMIN and LMUL * ELEN, inclusive. We print a warning if these requirements are not met. Reviewers: kito-cheng, asb, frasercrmck, jrtc27, michaelmaitland, lukel97 Reviewed By: lukel97 Pull Request: https://github.com/llvm/llvm-project/pull/94313
-
Johannes Reifferscheid authored
This is a follow-up to #92506.
-
Nikita Popov authored
Use clang_target_link_libraries() instead of LINK_LIBS when linking clang libraries. This ensures that in CLANG_LINK_CLANG_DYLIB mode we link against libclang-cpp.so (instead of linking against both it and the static libraries). Most places were already doing this correctly, there were just a handful of leftovers.
-
Pengcheng Wang authored
It seems that we have `B` extension again: https://github.com/riscv/riscv-b According to the spec, `B` extension represents the collection of the `Zba`, `Zbb`, `Zbs` extensions. Though it hasn't been ratified, I set its version to `1.0`.
-
Paul Kirth authored
Reverts llvm/llvm-project#86609 This change causes compile-time regressions for stage2 builds (https://llvm-compile-time-tracker.com/compare.php?from=3254f31a66263ea9647c9547f1531c3123444fcd&to=c5978f1eb5eeca8610b9dfce1fcbf1f473911cd8&stat=instructions:u). It also introduced unintended changes to `.text` which should be addressed before relanding.
-
Martin Storsjö authored
GCC usually doesn't warn about unrecognized -Wno-<foo> options, if no diagnostics are printed. However if some diagnostics are printed, it also mentions that there were unrecognized -Wno-<foo> options. Before 4feae05c, we checked for whether -Wnested-anon-types was supported, and added the -Wno-<foo> form if the positive form of the option was supported. As of GCC 14, -Wnested-anon-types isn't supported, thus limit the use of the option to actual Clang (and still only while using the GCC compatible driver). This avoids unnecessary mentions about unrecognized -Wno-<foo> options when building with GCC.
-
Alexander Shaposhnikov authored
Add plumbing for the numerical sanitizer on Clang's side.
-
Mariusz Sikora authored
… instructions
-
Daniil Kovalev authored
Lower global references to ptrauth constants into `@AUTH` `MCExpr`'s. The logic is common for MachO and ELF - test both. --------- Co-authored-by:Ahmed Bougacha <ahmed@bougacha.org>
-
Valentin Clement (バレンタイン クレメン) authored
This is a follow up patch to #94652 and handles the lowering of the reduce intrinsic with DIM argument and non scalar result.
-
Jason Molenda authored
Add comments and a test for delay-init libraries on macOS. I originally added the support in 954d00e8 a month ago, but without these additional clarifications. rdar://126885033
-
PiJoules authored
They were assigned from calls to find_chunk_ptr_for_size which return size_t now.
-
Pavel Samolysov authored
This addresses a clang-tidy suggestion.
-
Fangrui Song authored
-
Fabian Mora authored
This patch updates the lowering of `LaunchFuncOp` in GPU to LLVM to only legalize the operation with the converted operands, effectively removing the lowering used by the old serialization pipeline. It also removes all remaining uses of the old gpu serialization infrastructure in `gpu-to-llvm`. See [Compilation overview | 'gpu' Dialect - MLIR docs](https://mlir.llvm.org/docs/Dialects/GPU/#compilation-overview) for additional information on the target attributes compilation pipeline that replaced the old serialization pipeline.
-
Fangrui Song authored
-
Alexander Shaposhnikov authored
Add sanitize_numerical_stability attribute.
-
Farzon Lotfi authored
Relanding this PR now that https://github.com/llvm/llvm-project/pull/90503 has merged. with `FTAN` landing in [TargetLoweringBase.cpp:L1021](https://github.com/llvm/llvm-project/blob/main/llvm/lib/CodeGen/TargetLoweringBase.cpp#L1020C23-L1021C63 ) There is now a llvm tan intrinsic 32\64\128 Expand case for all llvm backends. In LLVM, the `llvm.experimental.constrained.cos` and `llvm.experimental.constrained.sin` intrinsics are used for performing cosine and sine calculations with additional constraints on floating-point operations. This behavior is expected for all floating-point math intrinsics. This change adds these constraints for the `tan` intrinsic. - `Builtins.td` - replace TanF128 with F16F128MathTemplate - `CGBuiltin.cpp` - map existing tan builtins to `tan` and `constrained_tan` intrinsic - `ConstrainedOps.def` map tan and constrained_tan to an ISDOpcode. resolves #91421 --------- Co-authored-by:
Farzon Lotfi <farzon@farzon.com>
-