- Jun 07, 2024
-
-
Christudasan Devadasan authored
Test CodeGen/AMDGPU/build_vector.ll has the lit patterns partially hand-written and the rest auto-generated. It doesn't look good when changes are required with future patches. Auto-generating the entire pattern. Moved out the R600 test into build_vector-r600.ll.
-
Chuanqi Xu authored
This reverts commit 5c104879. The ArmV7 bot is complaining the change breaks the alignment.
-
Jon Roelofs authored
before: ``` *** Dumping AST Record Layout 0 | class llvm::SUnit 0 | SDNode * Node 8 | MachineInstr * Instr 16 | SUnit * OrigNode 24 | const MCSchedClassDesc * SchedClass 32 | class llvm::SmallVector<class llvm::SDep, 4> Preds 32 | class llvm::SmallVectorImpl<class llvm::SDep> (base) 32 | class llvm::SmallVectorTemplateBase<class llvm::SDep> (base) 32 | class llvm::SmallVectorTemplateCommon<class llvm::SDep> (base) 32 | class llvm::SmallVectorBase<uint32_t> (base) 32 | void * BeginX 40 | unsigned int Size 44 | unsigned int Capacity 48 | struct llvm::SmallVectorStorage<class llvm::SDep, 4> (base) 48 | char[64] InlineElts 112 | class llvm::SmallVector<class llvm::SDep, 4> Succs 112 | class llvm::SmallVectorImpl<class llvm::SDep> (base) 112 | class llvm::SmallVectorTemplateBase<class llvm::SDep> (base) 112 | class llvm::SmallVectorTemplateCommon<class llvm::SDep> (base) 112 | class llvm::SmallVectorBase<uint32_t> (base) 112 | void * BeginX 120 | unsigned int Size 124 | unsigned int Capacity 128 | struct llvm::SmallVectorStorage<class llvm::SDep, 4> (base) 128 | char[64] InlineElts 192 | unsigned int NodeNum 196 | unsigned int NodeQueueId 200 | unsigned int NumPreds 204 | unsigned int NumSuccs 208 | unsigned int NumPredsLeft 212 | unsigned int NumSuccsLeft 216 | unsigned int WeakPredsLeft 220 | unsigned int WeakSuccsLeft 224 | unsigned short NumRegDefsLeft 226 | unsigned short Latency 228:0-0 | _Bool isVRegCycle 228:1-1 | _Bool isCall 228:2-2 | _Bool isCallOp 228:3-3 | _Bool isTwoAddress 228:4-4 | _Bool isCommutable 228:5-5 | _Bool hasPhysRegUses 228:6-6 | _Bool hasPhysRegDefs 228:7-7 | _Bool hasPhysRegClobbers 229:0-0 | _Bool isPending 229:1-1 | _Bool isAvailable 229:2-2 | _Bool isScheduled 229:3-3 | _Bool isScheduleHigh 229:4-4 | _Bool isScheduleLow 229:5-5 | _Bool isCloned 229:6-6 | _Bool isUnbuffered 229:7-7 | _Bool hasReservedResource 232 | Sched::Preference SchedulingPref 236:0-0 | _Bool isDepthCurrent 236:1-1 | _Bool isHeightCurrent 240 | unsigned int Depth 244 | unsigned int Height 248 | unsigned int TopReadyCycle 252 | unsigned int BotReadyCycle 256 | const TargetRegisterClass * CopyDstRC 264 | const TargetRegisterClass * CopySrcRC | [sizeof=272, dsize=272, align=8, | nvsize=272, nvalign=8] ``` after: ``` *** Dumping AST Record Layout 0 | class llvm::SUnit 0 | union llvm::SUnit::(anonymous at /Users/jonathan_roelofs/llvm-upstream/llvm/include/llvm/CodeGen/ScheduleDAG.h:246:5) 0 | SDNode * Node 0 | MachineInstr * Instr 8 | SUnit * OrigNode 16 | const MCSchedClassDesc * SchedClass 24 | const TargetRegisterClass * CopyDstRC 32 | const TargetRegisterClass * CopySrcRC 40 | class llvm::SmallVector<class llvm::SDep, 4> Preds 40 | class llvm::SmallVectorImpl<class llvm::SDep> (base) 40 | class llvm::SmallVectorTemplateBase<class llvm::SDep> (base) 40 | class llvm::SmallVectorTemplateCommon<class llvm::SDep> (base) 40 | class llvm::SmallVectorBase<uint32_t> (base) 40 | void * BeginX 48 | unsigned int Size 52 | unsigned int Capacity 56 | struct llvm::SmallVectorStorage<class llvm::SDep, 4> (base) 56 | char[64] InlineElts 120 | class llvm::SmallVector<class llvm::SDep, 4> Succs 120 | class llvm::SmallVectorImpl<class llvm::SDep> (base) 120 | class llvm::SmallVectorTemplateBase<class llvm::SDep> (base) 120 | class llvm::SmallVectorTemplateCommon<class llvm::SDep> (base) 120 | class llvm::SmallVectorBase<uint32_t> (base) 120 | void * BeginX 128 | unsigned int Size 132 | unsigned int Capacity 136 | struct llvm::SmallVectorStorage<class llvm::SDep, 4> (base) 136 | char[64] InlineElts 200 | unsigned int NodeNum 204 | unsigned int NodeQueueId 208 | unsigned int NumPreds 212 | unsigned int NumSuccs 216 | unsigned int NumPredsLeft 220 | unsigned int NumSuccsLeft 224 | unsigned int WeakPredsLeft 228 | unsigned int WeakSuccsLeft 232 | unsigned int TopReadyCycle 236 | unsigned int BotReadyCycle 240 | unsigned int Depth 244 | unsigned int Height 248:0-0 | _Bool isVRegCycle 248:1-1 | _Bool isCall 248:2-2 | _Bool isCallOp 248:3-3 | _Bool isTwoAddress 248:4-4 | _Bool isCommutable 248:5-5 | _Bool hasPhysRegUses 248:6-6 | _Bool hasPhysRegDefs 248:7-7 | _Bool hasPhysRegClobbers 249:0-0 | _Bool isPending 249:1-1 | _Bool isAvailable 249:2-2 | _Bool isScheduled 249:3-3 | _Bool isScheduleHigh 249:4-4 | _Bool isScheduleLow 249:5-5 | _Bool isCloned 249:6-6 | _Bool isUnbuffered 249:7-7 | _Bool hasReservedResource 250 | unsigned short NumRegDefsLeft 252 | unsigned short Latency 254:0-0 | _Bool isDepthCurrent 254:1-1 | _Bool isHeightCurrent 254:2-2 | _Bool isNode 254:3-3 | _Bool isInst 254:4-7 | Sched::Preference SchedulingPref | [sizeof=256, dsize=255, align=8, | nvsize=255, nvalign=8] ``` -
Chuanqi Xu authored
Following of https://github.com/llvm/llvm-project/pull/86912 The motivation of the patch series is that, for a module interface unit `X`, when the dependent modules of `X` changes, if the changes is not relevant with `X`, we hope the BMI of `X` won't change. For the specific patch, we hope if the changes was about irrelevant declaration changes, we hope the BMI of `X` won't change. **However**, I found the patch itself is not very useful in practice, since the adding or removing declarations, will change the state of identifiers and types in most cases. That said, for the most simple example, ``` // partA.cppm export module m:partA; // partA.v1.cppm export module m:partA; export void a() {} // partB.cppm export module m:partB; export void b() {} // m.cppm export module m; export import :partA; export import :partB; // onlyUseB; export module onlyUseB; import m; export inline void onluUseB() { b(); } ``` the BMI of `onlyUseB` will change after we change the implementation of `partA.cppm` to `partA.v1.cppm`. Since `partA.v1.cppm` introduces new identifiers and types (the function prototype). So in this patch, we have to write the tests as: ``` // partA.cppm export module m:partA; export int getA() { ... } export int getA2(int) { ... } // partA.v1.cppm export module m:partA; export int getA() { ... } export int getA(int) { ... } export int getA2(int) { ... } // partB.cppm export module m:partB; export void b() {} // m.cppm export module m; export import :partA; export import :partB; // onlyUseB; export module onlyUseB; import m; export inline void onluUseB() { b(); } ``` so that the new introduced declaration `int getA(int)` doesn't introduce new identifiers and types, then the BMI of `onlyUseB` can keep unchanged. While it looks not so great, the patch should be the base of the patch to erase the transitive change for identifiers and types since I don't know how can we introduce new types and identifiers without introducing new declarations. Given how tightly the relationship between declarations, types and identifiers, I think we can only reach the ideal state after we made the series for all of the three entties. The design of the patch is similar to https://github.com/llvm/llvm-project/pull/86912, which extends the 32-bit DeclID to 64-bit and use the higher bits to store the module file index and the lower bits to store the Local Decl ID. A slight difference is that we only use 48 bits to store the new DeclID since we try to use the higher 16 bits to store the module ID in the prefix of Decl class. Previously, we use 32 bits to store the module ID and 32 bits to store the DeclID. I don't want to allocate additional space so I tried to make the additional space the same as 64 bits. An potential interesting thing here is about the relationship between the module ID and the module file index. I feel we can get the module file index by the module ID. But I didn't prove it or implement it. Since I want to make the patch itself as small as possible. We can make it in the future if we want. Another change in the patch is the new concept Decl Index, which means the index of the very big array `DeclsLoaded` in ASTReader. Previously, the index of a loaded declaration is simply the Decl ID minus PREDEFINED_DECL_NUMs. So there are some places they got used ambiguously. But this patch tried to split these two concepts. As https://github.com/llvm/llvm-project/pull/86912 did, the change will increase the on-disk PCM file sizes. As the declaration ID may be the most IDs in the PCM file, this can have the biggest impact on the size. In my experiments, this change will bring 6.6% increase of the on-disk PCM size. No compile-time performance regression observed. Given the benefits in the motivation example, I think the cost is worthwhile.
-
Weining Lu authored
-
Gedare Bloom authored
Short-circuit the parsing of tok::colon to label colons found within lines starting with asm as InlineASMColon. Fixes #92616. --------- Co-authored-by:Owen Pan <owenpiano@gmail.com>
-
Nour authored
-
Craig Topper authored
Instead of having multiple places insert into the Features vector independently, check all the conditions in one place. This avoids a subtle ordering requirement that -mstrict-align processing had to be done after the others.
-
Noah Goldstein authored
We don't need the `noundef` check if the new simplification is a constant. This cleans up regressions from folding multiuse: `(icmp eq/ne (sub/xor x, y), 0)` -> `(icmp eq/ne x, y)`. Closes #88298 -
Noah Goldstein authored
-
Fangrui Song authored
-
Fangrui Song authored
Make it easier to add CREL support.
-
Thurston Dang authored
This test case shows a limitation of DFSan's sscanf implementation (introduced in https://reviews.llvm.org/D153775): it simply ignores ordinary characters in the format string, instead of actually comparing them against the input. This may change the semantics of instrumented programs. Importantly, this also means that DFSan's release_shadow_space.c test, which relies on sscanf to scrape the RSS from /proc/maps output, will incorrectly match lines that don't contain RSS information. As a result, it adding together numbers from irrelevant output (e.g., base addresses), resulting in test flakiness (https://github.com/llvm/llvm-project/issues/91287).
-
Owen Pan authored
Fixes #94555.
-
Fangrui Song authored
-
Konstantin Varlamov authored
-
Congcong Cai authored
Fixes: #94634
-
Kazu Hirata authored
Call stacks are a huge portion of the MemProf profile, taking up 70+% of the profile file size. This patch implements a radix tree to compress call stacks, which are known to have long common prefixes. Specifically, CallStackRadixTreeBuilder, introduced in this patch, takes call stacks in the MemProf profile, sorts them in the dictionary order to maximize the common prefix between adjacent call stacks, and then encodes a radix tree into a single array that is ready for serialization. The resulting radix array is essentially a concatenation of call stack arrays, each encoded with its length followed by the payload, except that these arrays contain "instructions" like "skip 7 elements forward" to borrow common prefixes from other call stacks. This patch does not integrate with the MemProf serialization/deserialization infrastructure yet. Once integrated, the radix tree is expected to roughly halve the file size of the MemProf profile.
-
Med Ismail Bennani authored
This patch fixes a build issue following e57308b0 when enabling module build. With that change, we failed to build the LLVM_IR module since GEPNoWrapFlags wasn't defined prior to using it. This patch addressed that issue by including the missing header in `llvm/IR/IRBuilderFolder.h` which uses the `GEPNoWrapFlags` type. This should ensure that we can always build the `LLVM_IR` module. Signed-off-by:
Med Ismail Bennani <ismail@bennani.ma> Signed-off-by:
Med Ismail Bennani <ismail@bennani.ma>
-
Alexander Richardson authored
This function is called during very early startup and which can result in a crash on FreeBSD. The sigaction() function in libc is indirected via a table so that it can be interposed by the threading library rather than calling the syscall directly. In the crash I was observing this table had not yet been relocated, so we ended up jumping to an invalid address. To avoid this problem we can call __sys_sigaction, which calls the syscall directly and in FreeBSD 15 is part of libsys rather than libc, so does not depend on libc being fully initialized.
-
Thurston Dang authored
…ty with high-entropy ASLR With high-entropy ASLR (e.g., 32-bits == 16TB), the allocator base of 0x700000000000 (112TB) may collide with the placement of the libraries (e.g., on Linux, the mmap base could be 128TB - 16TB == 112TB). This results in a segfault in the test case. This patch moves the allocator base below the PIE program segment, inspired by fb77ca05. As per that patch: 1) we are leaving the old behavior for Apple 2) since ASLR cannot be set above 32-bits for x86-64 Linux, we expect this new layout to be durable. Note that this is only changing a test case, not the behavior of sanitizers. Sanitizers have their own settings for initializing the allocator base. Reproducer: 1. ninja check-sanitizer # Just to build the test binary needed below; no need to actually run the tests here 2. sudo sysctl vm.mmap_rnd_bits=32 # Increase ASLR entropy 3. for f in `seq 1 10000`; do echo $f; GTEST_FILTER=*SizeClassAllocator64Dense ./projects/compiler-rt/lib/sanitizer_common/tests/Sanitizer-x86_64-Test > /tmp/x; if [ $? -ne 0 ]; then cat /tmp/x; fi; done
-
Nikita Popov authored
These CHECKs are all checking indices, which must be strictly smaller than the size (otherwise they would go out of bounds).
-
aaryanshukla authored
- addressed https://github.com/llvm/llvm-project/pull/94317#issuecomment-2153103129 - added conditional in cmake file for exit_handler object library Co-authored-by:
Aaryan Shukla <aaryanshukla@google.com>
-
Vitaly Buka authored
-
Alex Langford authored
-
Kazu Hirata authored
This patch replaces llvm::SmallVector<Frame> with std::vector<Frame>. llvm::SmallVector<Frame> sets aside one inline element. Meanwhile, when I sort all call stacks by their lengths, the length at the first percentile is already 2. That is, 99 percent of call stacks do not take advantage of the inline element. Using std::vector<Frame> reduces the cycle and instruction counts by 11% and 22%, respectively, with "llvm-profdata show" modified to deserialize all MemProfRecords.
-
Dave Lee authored
Fixed for more accurate searches of the flag `-Wsystem-headers-in-module=`.
-
Kazu Hirata authored
This patch removes swapToHostOrder in favor of llvm::support::endian::readNext as swapToHostOrder is too thin a wrapper around readNext. Note that there are two variants of readNext: - readNext<type, endian, align>(ptr) - readNext<type, align>(ptr, endian) swapToHostOrder uses the former, but this patch switches to the latter. While we are at it, this patch teaches readNext to default to unaligned just as I did in: commit 568368a4 Author: Kazu Hirata <kazu@google.com> Date: Mon Apr 15 19:05:30 2024 -0700
-
Fangrui Song authored
-
Philip Reames authored
getVNInfoFromReg is expected to return a nullptr if-and-only-if the operand is undef. (This was asserted for.) Reverse the order of the checks to simplify an upcoming set of patches.
-
Joseph Huber authored
Summary: The old COV3 implementation of HSA used to omit the implicit arguments from the kernel argument size. For COV4 and COV5 this is no longer the case so we can simply use the size reported from the symbol information. See https://github.com/ROCm/ROCR-Runtime/issues/117#issuecomment-812758161
-
Alex Langford authored
The summary already includes other size information, e.g. total debug info size in bytes. The only other way I can get this information is by dumping all statistics which can be quite large. Adding it to the summary seems fair.
-
Christopher Bate authored
Relaxes restriction that certain public utility functions only apply to the builtin ModuleOp.
-
Jay Foad authored
Remove `REQUIRES: shell` from some tests that seem fine without it. Tested on Windows and with LIT_USE_INTERNAL_SHELL=1 on Linux.
-
Joseph Huber authored
Summary: We don't have the abs function to link against, just use the builtin.
-
Med Ismail Bennani authored
This PR removes the `target-aarch64` requirement on the crashlog tests to exercice them on Intel bots and make image loading single-threaded temporarily while implementing a fix for a deadlock issue when loading the images in parallel. Signed-off-by:Med Ismail Bennani <ismail@bennani.ma>
-
Haojian Wu authored
-
Fangrui Song authored
https://reviews.llvm.org/D85867 changed the way we assign file offsets (alloc sections first, then non-alloc sections). It also removed a non-alloc special case from `findOrphanPos`. Looking at the memory-nonalloc-no-warn.test change, which would be needed by #93761, it makes sense to restore the previous behavior: when placing non-alloc orphan sections, keep these sections at the end so that the section index order matches the file offset order. This change is cosmetic. In sections-nonalloc.s, GNU ld places the orphan `other3` in the middle and the orphan .symtab/.shstrtab/.strtab at the end. Pull Request: https://github.com/llvm/llvm-project/pull/94519
-
Stanislav Mekhanoshin authored
-
Kazu Hirata authored
Changing the type of Frame::SymbolName from std::optional<std::string> to std::unique<std::string> reduces sizeof(Frame) from 64 to 32. The smaller type reduces the cycle and instruction counts by 23% and 4.4%, respectively, with "llvm-profdata show" modified to deserialize all MemProfRecords in a MemProf V2 profile. The peak memory usage is cut down nearly by half.
-