1. Jun 07, 2024
    • Christudasan Devadasan's avatar
      [AMDGPU] Auto-generating lit test patterns (NFC) (#93837) · 0e1d6e2f
      Christudasan Devadasan authored
      Test CodeGen/AMDGPU/build_vector.ll has the lit patterns partially
      hand-written and the rest auto-generated. It doesn't look good when
      changes are required with future patches. Auto-generating the entire
      pattern. Moved out the R600 test into build_vector-r600.ll.
      0e1d6e2f
    • Chuanqi Xu's avatar
      Revert "[serialization] no transitive decl change (#92083)" · 4f70c5ec
      Chuanqi Xu authored
      This reverts commit 5c104879.
      
      The ArmV7 bot is complaining the change breaks the alignment.
      4f70c5ec
    • Jon Roelofs's avatar
      [llvm][ScheduleDAG] Re-arrange SUnit's members to make it smaller (#94547) · 222e0a04
      Jon Roelofs authored
      before:
      ```
      *** Dumping AST Record Layout
               0 | class llvm::SUnit
               0 |   SDNode * Node
               8 |   MachineInstr * Instr
              16 |   SUnit * OrigNode
              24 |   const MCSchedClassDesc * SchedClass
              32 |   class llvm::SmallVector<class llvm::SDep, 4> Preds
              32 |     class llvm::SmallVectorImpl<class llvm::SDep> (base)
              32 |       class llvm::SmallVectorTemplateBase<class llvm::SDep> (base)
              32 |         class llvm::SmallVectorTemplateCommon<class llvm::SDep> (base)
              32 |           class llvm::SmallVectorBase<uint32_t> (base)
              32 |             void * BeginX
              40 |             unsigned int Size
              44 |             unsigned int Capacity
              48 |     struct llvm::SmallVectorStorage<class llvm::SDep, 4> (base)
              48 |       char[64] InlineElts
             112 |   class llvm::SmallVector<class llvm::SDep, 4> Succs
             112 |     class llvm::SmallVectorImpl<class llvm::SDep> (base)
             112 |       class llvm::SmallVectorTemplateBase<class llvm::SDep> (base)
             112 |         class llvm::SmallVectorTemplateCommon<class llvm::SDep> (base)
             112 |           class llvm::SmallVectorBase<uint32_t> (base)
             112 |             void * BeginX
             120 |             unsigned int Size
             124 |             unsigned int Capacity
             128 |     struct llvm::SmallVectorStorage<class llvm::SDep, 4> (base)
             128 |       char[64] InlineElts
             192 |   unsigned int NodeNum
             196 |   unsigned int NodeQueueId
             200 |   unsigned int NumPreds
             204 |   unsigned int NumSuccs
             208 |   unsigned int NumPredsLeft
             212 |   unsigned int NumSuccsLeft
             216 |   unsigned int WeakPredsLeft
             220 |   unsigned int WeakSuccsLeft
             224 |   unsigned short NumRegDefsLeft
             226 |   unsigned short Latency
         228:0-0 |   _Bool isVRegCycle
         228:1-1 |   _Bool isCall
         228:2-2 |   _Bool isCallOp
         228:3-3 |   _Bool isTwoAddress
         228:4-4 |   _Bool isCommutable
         228:5-5 |   _Bool hasPhysRegUses
         228:6-6 |   _Bool hasPhysRegDefs
         228:7-7 |   _Bool hasPhysRegClobbers
         229:0-0 |   _Bool isPending
         229:1-1 |   _Bool isAvailable
         229:2-2 |   _Bool isScheduled
         229:3-3 |   _Bool isScheduleHigh
         229:4-4 |   _Bool isScheduleLow
         229:5-5 |   _Bool isCloned
         229:6-6 |   _Bool isUnbuffered
         229:7-7 |   _Bool hasReservedResource
             232 |   Sched::Preference SchedulingPref
         236:0-0 |   _Bool isDepthCurrent
         236:1-1 |   _Bool isHeightCurrent
             240 |   unsigned int Depth
             244 |   unsigned int Height
             248 |   unsigned int TopReadyCycle
             252 |   unsigned int BotReadyCycle
             256 |   const TargetRegisterClass * CopyDstRC
             264 |   const TargetRegisterClass * CopySrcRC
                 | [sizeof=272, dsize=272, align=8,
                 |  nvsize=272, nvalign=8]
      ```
      
      after:
      ```
      *** Dumping AST Record Layout
               0 | class llvm::SUnit
               0 |   union llvm::SUnit::(anonymous at /Users/jonathan_roelofs/llvm-upstream/llvm/include/llvm/CodeGen/ScheduleDAG.h:246:5)
               0 |     SDNode * Node
               0 |     MachineInstr * Instr
               8 |   SUnit * OrigNode
              16 |   const MCSchedClassDesc * SchedClass
              24 |   const TargetRegisterClass * CopyDstRC
              32 |   const TargetRegisterClass * CopySrcRC
              40 |   class llvm::SmallVector<class llvm::SDep, 4> Preds
              40 |     class llvm::SmallVectorImpl<class llvm::SDep> (base)
              40 |       class llvm::SmallVectorTemplateBase<class llvm::SDep> (base)
              40 |         class llvm::SmallVectorTemplateCommon<class llvm::SDep> (base)
              40 |           class llvm::SmallVectorBase<uint32_t> (base)
              40 |             void * BeginX
              48 |             unsigned int Size
              52 |             unsigned int Capacity
              56 |     struct llvm::SmallVectorStorage<class llvm::SDep, 4> (base)
              56 |       char[64] InlineElts
             120 |   class llvm::SmallVector<class llvm::SDep, 4> Succs
             120 |     class llvm::SmallVectorImpl<class llvm::SDep> (base)
             120 |       class llvm::SmallVectorTemplateBase<class llvm::SDep> (base)
             120 |         class llvm::SmallVectorTemplateCommon<class llvm::SDep> (base)
             120 |           class llvm::SmallVectorBase<uint32_t> (base)
             120 |             void * BeginX
             128 |             unsigned int Size
             132 |             unsigned int Capacity
             136 |     struct llvm::SmallVectorStorage<class llvm::SDep, 4> (base)
             136 |       char[64] InlineElts
             200 |   unsigned int NodeNum
             204 |   unsigned int NodeQueueId
             208 |   unsigned int NumPreds
             212 |   unsigned int NumSuccs
             216 |   unsigned int NumPredsLeft
             220 |   unsigned int NumSuccsLeft
             224 |   unsigned int WeakPredsLeft
             228 |   unsigned int WeakSuccsLeft
             232 |   unsigned int TopReadyCycle
             236 |   unsigned int BotReadyCycle
             240 |   unsigned int Depth
             244 |   unsigned int Height
         248:0-0 |   _Bool isVRegCycle
         248:1-1 |   _Bool isCall
         248:2-2 |   _Bool isCallOp
         248:3-3 |   _Bool isTwoAddress
         248:4-4 |   _Bool isCommutable
         248:5-5 |   _Bool hasPhysRegUses
         248:6-6 |   _Bool hasPhysRegDefs
         248:7-7 |   _Bool hasPhysRegClobbers
         249:0-0 |   _Bool isPending
         249:1-1 |   _Bool isAvailable
         249:2-2 |   _Bool isScheduled
         249:3-3 |   _Bool isScheduleHigh
         249:4-4 |   _Bool isScheduleLow
         249:5-5 |   _Bool isCloned
         249:6-6 |   _Bool isUnbuffered
         249:7-7 |   _Bool hasReservedResource
             250 |   unsigned short NumRegDefsLeft
             252 |   unsigned short Latency
         254:0-0 |   _Bool isDepthCurrent
         254:1-1 |   _Bool isHeightCurrent
         254:2-2 |   _Bool isNode
         254:3-3 |   _Bool isInst
         254:4-7 |   Sched::Preference SchedulingPref
                 | [sizeof=256, dsize=255, align=8,
                 |  nvsize=255, nvalign=8]
      ```
      222e0a04
    • Chuanqi Xu's avatar
      [serialization] no transitive decl change (#92083) · 5c104879
      Chuanqi Xu authored
      Following of https://github.com/llvm/llvm-project/pull/86912
      
      The motivation of the patch series is that, for a module interface unit
      `X`, when the dependent modules of `X` changes, if the changes is not
      relevant with `X`, we hope the BMI of `X` won't change. For the specific
      patch, we hope if the changes was about irrelevant declaration changes,
      we hope the BMI of `X` won't change. **However**, I found the patch
      itself is not very useful in practice, since the adding or removing
      declarations, will change the state of identifiers and types in most
      cases.
      
      That said, for the most simple example,
      
      ```
      // partA.cppm
      export module m:partA;
      
      // partA.v1.cppm
      export module m:partA;
      export void a() {}
      
      // partB.cppm
      export module m:partB;
      export void b() {}
      
      // m.cppm
      export module m;
      export import :partA;
      export import :partB;
      
      // onlyUseB;
      export module onlyUseB;
      import m;
      export inline void onluUseB() {
          b();
      }
      ```
      
      the BMI of `onlyUseB` will change after we change the implementation of
      `partA.cppm` to `partA.v1.cppm`. Since `partA.v1.cppm` introduces new
      identifiers and types (the function prototype).
      
      So in this patch, we have to write the tests as:
      
      ```
      // partA.cppm
      export module m:partA;
      export int getA() { ... }
      export int getA2(int) { ... }
      
      // partA.v1.cppm
      export module m:partA;
      export int getA() { ... }
      export int getA(int) { ... }
      export int getA2(int) { ... }
      
      // partB.cppm
      export module m:partB;
      export void b() {}
      
      // m.cppm
      export module m;
      export import :partA;
      export import :partB;
      
      // onlyUseB;
      export module onlyUseB;
      import m;
      export inline void onluUseB() {
          b();
      }
      ```
      
      so that the new introduced declaration `int getA(int)` doesn't introduce
      new identifiers and types, then the BMI of `onlyUseB` can keep
      unchanged.
      
      While it looks not so great, the patch should be the base of the patch
      to erase the transitive change for identifiers and types since I don't
      know how can we introduce new types and identifiers without introducing
      new declarations. Given how tightly the relationship between
      declarations, types and identifiers, I think we can only reach the ideal
      state after we made the series for all of the three entties.
      
      The design of the patch is similar to
      https://github.com/llvm/llvm-project/pull/86912, which extends the
      32-bit DeclID to 64-bit and use the higher bits to store the module file
      index and the lower bits to store the Local Decl ID.
      
      A slight difference is that we only use 48 bits to store the new DeclID
      since we try to use the higher 16 bits to store the module ID in the
      prefix of Decl class. Previously, we use 32 bits to store the module ID
      and 32 bits to store the DeclID. I don't want to allocate additional
      space so I tried to make the additional space the same as 64 bits. An
      potential interesting thing here is about the relationship between the
      module ID and the module file index. I feel we can get the module file
      index by the module ID. But I didn't prove it or implement it. Since I
      want to make the patch itself as small as possible. We can make it in
      the future if we want.
      
      Another change in the patch is the new concept Decl Index, which means
      the index of the very big array `DeclsLoaded` in ASTReader. Previously,
      the index of a loaded declaration is simply the Decl ID minus
      PREDEFINED_DECL_NUMs. So there are some places they got used
      ambiguously. But this patch tried to split these two concepts.
      
      As https://github.com/llvm/llvm-project/pull/86912 did, the change will
      increase the on-disk PCM file sizes. As the declaration ID may be the
      most IDs in the PCM file, this can have the biggest impact on the size.
      In my experiments, this change will bring 6.6% increase of the on-disk
      PCM size. No compile-time performance regression observed. Given the
      benefits in the motivation example, I think the cost is worthwhile.
      5c104879
    • Weining Lu's avatar
      f2441b02
    • Gedare Bloom's avatar
      [clang-format]: Annotate colons found in inline assembly (#92617) · a10135f4
      Gedare Bloom authored
      
      
      Short-circuit the parsing of tok::colon to label colons found within
      lines starting with asm as InlineASMColon.
      
      Fixes #92616.
      
      ---------
      
      Co-authored-by: default avatarOwen Pan <owenpiano@gmail.com>
      a10135f4
    • Nour's avatar
    • Craig Topper's avatar
      [RISCV] Unify all the code that adds unaligned-scalar/vector-mem to Features vector. (#94660) · 2d65097b
      Craig Topper authored
      Instead of having multiple places insert into the Features vector
      independently, check all the conditions in one place.
      
      This avoids a subtle ordering requirement that -mstrict-align processing
      had to be done after the others.
      2d65097b
    • Noah Goldstein's avatar
      [InstCombine] Improve coverage of `foldSelectValueEquivalence` for constants · 7e7c29ba
      Noah Goldstein authored
      We don't need the `noundef` check if the new simplification is a
      constant.
      
      This cleans up regressions from folding multiuse:
          `(icmp eq/ne (sub/xor x, y), 0)` -> `(icmp eq/ne x, y)`.
      
      Closes #88298
      7e7c29ba
    • Noah Goldstein's avatar
    • Fangrui Song's avatar
      [SPIRV] Fix -Wunused-but-set-variable. NFC · f42025c2
      Fangrui Song authored
      f42025c2
    • Fangrui Song's avatar
      [ELF] Simplify code. NFC · a1fa43d0
      Fangrui Song authored
      Make it easier to add CREL support.
      a1fa43d0
    • Thurston Dang's avatar
      [dfsan] Add test case for sscanf (#94700) · 79cd6c3d
      Thurston Dang authored
      This test case shows a limitation of DFSan's sscanf implementation
      (introduced in https://reviews.llvm.org/D153775): it simply ignores
      ordinary characters in the format string, instead of actually comparing
      them against the input. This may change the semantics of instrumented
      programs.
      
      Importantly, this also means that DFSan's release_shadow_space.c test,
      which relies on sscanf to scrape the RSS from /proc/maps output, will
      incorrectly match lines that don't contain RSS information. As a result,
      it adding together numbers from irrelevant output (e.g., base
      addresses), resulting in test flakiness
      (https://github.com/llvm/llvm-project/issues/91287).
      79cd6c3d
    • Owen Pan's avatar
      5e0fc93d
    • Fangrui Song's avatar
      [ELF] Improve -r section group tests · aae025a0
      Fangrui Song authored
      aae025a0
    • Konstantin Varlamov's avatar
    • Congcong Cai's avatar
    • Kazu Hirata's avatar
      [memprof] Add CallStackRadixTreeBuilder (#93784) · 5c0df5fe
      Kazu Hirata authored
      Call stacks are a huge portion of the MemProf profile, taking up 70+%
      of the profile file size.
      
      This patch implements a radix tree to compress call stacks, which are
      known to have long common prefixes.  Specifically,
      CallStackRadixTreeBuilder, introduced in this patch, takes call stacks
      in the MemProf profile, sorts them in the dictionary order to maximize
      the common prefix between adjacent call stacks, and then encodes a
      radix tree into a single array that is ready for serialization.
      
      The resulting radix array is essentially a concatenation of call stack
      arrays, each encoded with its length followed by the payload, except
      that these arrays contain "instructions" like "skip 7 elements
      forward" to borrow common prefixes from other call stacks.
      
      This patch does not integrate with the MemProf
      serialization/deserialization infrastructure yet.  Once integrated,
      the radix tree is expected to roughly halve the file size of the
      MemProf profile.
      5c0df5fe
    • Med Ismail Bennani's avatar
      [llvm/IR] Fix module build issue following e57308b0 (NFC) (#94580) · 210692bc
      Med Ismail Bennani authored
      This patch fixes a build issue following e57308b0
      
       when enabling
      module build.
      
      With that change, we failed to build the LLVM_IR module since
      GEPNoWrapFlags wasn't defined prior to using it.
      
      This patch addressed that issue by including the missing header in
      `llvm/IR/IRBuilderFolder.h` which uses the `GEPNoWrapFlags` type.
      This should ensure that we can always build the `LLVM_IR` module.
      
      Signed-off-by: default avatarMed Ismail Bennani <ismail@bennani.ma>
      
      Signed-off-by: default avatarMed Ismail Bennani <ismail@bennani.ma>
      210692bc
    • Alexander Richardson's avatar
      [compiler-rt] Map internal_sigaction to __sys_sigaction on FreeBSD (#84441) · 92a87088
      Alexander Richardson authored
      This function is called during very early startup and which can result
      in a crash on FreeBSD. The sigaction() function in libc is indirected
      via a table so that it can be interposed by the threading library
      rather than calling the syscall directly. In the crash I was observing
      this table had not yet been relocated, so we ended up jumping to an
      invalid address. To avoid this problem we can call __sys_sigaction,
      which calls the syscall directly and in FreeBSD 15 is part of libsys
      rather than libc, so does not depend on libc being fully initialized.
      92a87088
    • Thurston Dang's avatar
      [sanitizer_common] Change allocator base in test case for compatibili… (#93234) · 19bbbcbe
      Thurston Dang authored
      …ty with high-entropy ASLR
      
      With high-entropy ASLR (e.g., 32-bits == 16TB), the allocator base of
      0x700000000000 (112TB) may collide with the placement of the libraries
      (e.g., on Linux, the mmap base could be 128TB - 16TB == 112TB). This
      results in a segfault in the test case.
      
      This patch moves the allocator base below the PIE program segment,
      inspired by fb77ca05. As per that patch:
      1) we are leaving the old behavior for Apple 2) since ASLR cannot be set
      above 32-bits for x86-64 Linux, we expect this new layout to be durable.
      
      Note that this is only changing a test case, not the behavior of
      sanitizers. Sanitizers have their own settings for initializing the
      allocator base.
      
      Reproducer:
      1. ninja check-sanitizer # Just to build the test binary needed below;
      no need to actually run the tests here
      2. sudo sysctl vm.mmap_rnd_bits=32 # Increase ASLR entropy
      3. for f in `seq 1 10000`; do echo $f;
      GTEST_FILTER=*SizeClassAllocator64Dense
      ./projects/compiler-rt/lib/sanitizer_common/tests/Sanitizer-x86_64-Test
      > /tmp/x; if [ $? -ne 0 ]; then cat /tmp/x; fi; done
      19bbbcbe
    • Nikita Popov's avatar
      [sanitizer] Make CHECKs in bitvector more precise (NFC) (#94630) · b1861429
      Nikita Popov authored
      These CHECKs are all checking indices, which must be strictly smaller
      than the size (otherwise they would go out of bounds).
      b1861429
    • aaryanshukla's avatar
      [libc] fixed target issue with exit_handler (#94678) · 855af13e
      aaryanshukla authored
      - addressed
      https://github.com/llvm/llvm-project/pull/94317#issuecomment-2153103129
      
      
      - added conditional in cmake file for exit_handler object library
      
      Co-authored-by: default avatarAaryan Shukla <aaryanshukla@google.com>
      855af13e
    • Vitaly Buka's avatar
      [clang-repl] Disable new test after #89811 · 81e9a3c3
      Vitaly Buka authored
      81e9a3c3
    • Alex Langford's avatar
      ba7f52cc
    • Kazu Hirata's avatar
      [memprof] Use std::vector<Frame> instead of llvm::SmallVector<Frame> (NFC) (#94432) · 4a918f07
      Kazu Hirata authored
      This patch replaces llvm::SmallVector<Frame> with std::vector<Frame>.
      
      llvm::SmallVector<Frame> sets aside one inline element.  Meanwhile,
      when I sort all call stacks by their lengths, the length at the first
      percentile is already 2.  That is, 99 percent of call stacks do not
      take advantage of the inline element.
      
      Using std::vector<Frame> reduces the cycle and instruction counts by
      11% and 22%, respectively, with "llvm-profdata show" modified to
      deserialize all MemProfRecords.
      4a918f07
    • Dave Lee's avatar
      [clang] Fix flag typo in comment · 12ccc245
      Dave Lee authored
      Fixed for more accurate searches of the flag `-Wsystem-headers-in-module=`.
      12ccc245
    • Kazu Hirata's avatar
      [ProfileData] Remove swapToHostOrder (#94665) · 7476c20c
      Kazu Hirata authored
      This patch removes swapToHostOrder in favor of
      llvm::support::endian::readNext as swapToHostOrder is too thin a
      wrapper around readNext.
      
      Note that there are two variants of readNext:
      
      - readNext<type, endian, align>(ptr)
      - readNext<type, align>(ptr, endian)
      
      swapToHostOrder uses the former, but this patch switches to the latter.
      
      While we are at it, this patch teaches readNext to default to
      unaligned just as I did in:
      
        commit 568368a4
        Author: Kazu Hirata <kazu@google.com>
        Date:   Mon Apr 15 19:05:30 2024 -0700
      7476c20c
    • Fangrui Song's avatar
    • Philip Reames's avatar
      [RISCV][InsertVSETVLI] Check for undef register operand directly [nfc] · 06f03b80
      Philip Reames authored
      getVNInfoFromReg is expected to return a nullptr if-and-only-if the
      operand is undef.  (This was asserted for.)  Reverse the order of the
      checks to simplify an upcoming set of patches.
      06f03b80
    • Joseph Huber's avatar
      [Offload] Use the kernel argument size directly in AMDGPU offloading (#94667) · 9e209a4a
      Joseph Huber authored
      Summary:
      The old COV3 implementation of HSA used to omit the implicit arguments
      from the kernel argument size. For COV4 and COV5 this is no longer the
      case so we can simply use the size reported from the symbol information.
      
      See
      https://github.com/ROCm/ROCR-Runtime/issues/117#issuecomment-812758161
      9e209a4a
    • Alex Langford's avatar
      [lldb] Include memory stats in statistics summary (#94671) · 9293fc79
      Alex Langford authored
      The summary already includes other size information, e.g. total debug
      info size in bytes. The only other way I can get this information is by
      dumping all statistics which can be quite large. Adding it to the
      summary seems fair.
      9293fc79
    • Christopher Bate's avatar
      NFC: resolve TODO in LLVM dialect conversions (#91497) · f543dfd1
      Christopher Bate authored
      Relaxes restriction that certain public utility functions only apply
      to the builtin ModuleOp.
      f543dfd1
    • Jay Foad's avatar
      [LLVM] Do not require shell for some tests (#94595) · 878deaee
      Jay Foad authored
      Remove `REQUIRES: shell` from some tests that seem fine without it.
      Tested on Windows and with LIT_USE_INTERNAL_SHELL=1 on Linux.
      878deaee
    • Joseph Huber's avatar
      [Offload] Fix missing `abs` function for test · 2cc16442
      Joseph Huber authored
      Summary:
      We don't have the abs function to link against, just use the builtin.
      2cc16442
    • Med Ismail Bennani's avatar
      [lldb/crashlog] Remove aarch64 requirement on crashlog tests (#94553) · d09231a4
      Med Ismail Bennani authored
      
      
      This PR removes the `target-aarch64` requirement on the crashlog tests
      to exercice them on Intel bots and make image loading single-threaded
      temporarily while implementing a fix for a deadlock issue when loading
      the images in parallel.
      
      Signed-off-by: default avatarMed Ismail Bennani <ismail@bennani.ma>
      d09231a4
    • Haojian Wu's avatar
      [bazel] Port for 649edb8e · 5d0308f3
      Haojian Wu authored
      5d0308f3
    • Fangrui Song's avatar
      [ELF] Keep non-alloc orphan sections at the end · 9ad0175e
      Fangrui Song authored
      https://reviews.llvm.org/D85867 changed the way we assign file offsets
      (alloc sections first, then non-alloc sections).
      
      It also removed a non-alloc special case from `findOrphanPos`.
      Looking at the memory-nonalloc-no-warn.test change, which would be
      needed by #93761, it makes sense to restore the previous behavior: when
      placing non-alloc orphan sections, keep these sections at the end so
      that the section index order matches the file offset order.
      
      This change is cosmetic. In sections-nonalloc.s, GNU ld places the
      orphan `other3` in the middle and the orphan .symtab/.shstrtab/.strtab
      at the end.
      
      Pull Request: https://github.com/llvm/llvm-project/pull/94519
      9ad0175e
    • Stanislav Mekhanoshin's avatar
    • Kazu Hirata's avatar
      [memprof] Use std::unique_ptr instead of std::optional (#94655) · d55e235b
      Kazu Hirata authored
      Changing the type of Frame::SymbolName from std::optional<std::string>
      to std::unique<std::string> reduces sizeof(Frame) from 64 to 32.
      
      The smaller type reduces the cycle and instruction counts by 23% and
      4.4%, respectively, with "llvm-profdata show" modified to deserialize
      all MemProfRecords in a MemProf V2 profile.  The peak memory usage is
      cut down nearly by half.
      d55e235b