1. Nov 13, 2021
    • Matheus Izvekov's avatar
      [clang] retain type sugar in auto / template argument deduction · 4d8fff47
      Matheus Izvekov authored
      
      
      This implements the following changes:
      * AutoType retains sugared deduced-as-type.
      * Template argument deduction machinery analyses the sugared type all the way
      down. It would previously lose the sugar on first recursion.
      * Undeduced AutoType will be properly canonicalized, including the constraint
      template arguments.
      * Remove the decltype node created from the decltype(auto) deduction.
      
      As a result, we start seeing sugared types in a lot more test cases,
      including some which showed very unfriendly `type-parameter-*-*` types.
      
      Signed-off-by: default avatarMatheus Izvekov <mizvekov@gmail.com>
      
      Reviewed By: rsmith
      
      Differential Revision: https://reviews.llvm.org/D110216
      4d8fff47
    • Phoebe Wang's avatar
      [X86][ABI] Change the alignment of f80 in 32-bit calling convention to meet... · e49fcfc7
      Phoebe Wang authored
      [X86][ABI] Change the alignment of f80 in 32-bit calling convention to meet with different data layout
      
      Reviewed By: RKSimon
      
      Differential Revision: https://reviews.llvm.org/D113739
      e49fcfc7
    • Vitaly Buka's avatar
      [asan] More leaks in test · 89fb2c71
      Vitaly Buka authored
      It fails to detect a single leak with GLIBC 2.34.
      89fb2c71
    • Vy Nguyen's avatar
      [lld-macho] Parallelize scanning the symbol tables in export/unexport-ing. · ad932320
      Vy Nguyen authored
      (Split from D113167)
      Benchmarking on one of our large apps which exports a few thousands symbols,
      this showed an improvement of ~17%.
      
      x ./LLD_no_parallel.txt
      + ./LLD_with_parallel.txt
      
          N           Min           Max        Median           Avg        Stddev
      x  10         84.01         89.41         88.64        87.693     1.7424061
      +  10          71.9         74.29         72.63        72.753    0.77734663
      Difference at 95.0% confidence
      	-14.94 +/- 1.26763
      	-17.0367% +/- 1.44553%
      	(Student's t, pooled s = 1.34912)
      
      (wallclock)
      
      Differential Revision: https://reviews.llvm.org/D113820
      ad932320
    • Vitaly Buka's avatar
      [asan] Fix "no matching function" on GCC · 4b768eeb
      Vitaly Buka authored
      4b768eeb
    • Nico Weber's avatar
      [gn build] (semi-manually) port cb0e14ce · a1448693
      Nico Weber authored
      a1448693
    • Vitaly Buka's avatar
      [sanitizer] Fix test linking · afafa883
      Vitaly Buka authored
      afafa883
    • Ben Langmuir's avatar
      [ORC][ORC-RT] Register type metadata from __swift5_types MachO section · 2a739f27
      Ben Langmuir authored
      Similar to how the other swift sections are registered by the ORC
      runtime's macho platform, add the __swift5_types section, which contains
      type metadata. Add a simple test that demonstrates that the swift
      runtime recognized the registered types.
      
      rdar://85358530
      
      Differential Revision: https://reviews.llvm.org/D113811
      2a739f27
    • Craig Topper's avatar
      [RISCV] Fixed duplicate RUN line on float-intrinsics.ll. NFC · 8909dc5e
      Craig Topper authored
      We had two identical RV64I RUN lines. One should be RV32I.
      8909dc5e
    • Josh Learn's avatar
      [clang][objc][codegen] Skip emitting ObjC category metadata when the · 7611e16f
      Josh Learn authored
      category is empty
      
      Currently, if we create a category in ObjC that is empty, we still emit
      runtime metadata for that category. This is a scenario that could
      commonly be run into when using __attribute__((objc_direct_members)),
      which elides the need for much of the category metadata. This is
      slightly wasteful and can be easily skipped by checking the category
      metadata contents during CodeGen.
      
      rdar://66177182
      
      Differential Revision: https://reviews.llvm.org/D113455
      7611e16f
    • Vitaly Buka's avatar
      [sanitizer] Switch dlsym hack to internal_allocator · cb0e14ce
      Vitaly Buka authored
      Since glibc 2.34, dlsym does
        1. malloc 1
        2. malloc 2
        3. free pointer from malloc 1
        4. free pointer from malloc 2
      These sequence was not handled by trivial dlsym hack.
      
      This fixes https://bugs.llvm.org/show_bug.cgi?id=52278
      
      Reviewed By: eugenis, morehouse
      
      Differential Revision: https://reviews.llvm.org/D112588
      cb0e14ce
    • Philip Reames's avatar
      [runtime-unroll] Use incrementing IVs instead of decrementing ones · 37ead201
      Philip Reames authored
      This is one of those wonderful "in theory X doesn't matter, but in practice is does" changes. In this particular case, we shift the IVs inserted by the runtime unroller to clamp iteration count of the loops* from decrementing to incrementing.
      
      Why does this matter?  A couple of reasons:
      * SCEV doesn't have a native subtract node.  Instead, all subtracts (A - B) are represented as A + -1 * B and drops any flags invalidated by such.  As a result, SCEV is slightly less good at reasoning about edge cases involving decrementing addrecs than incrementing ones.  (You can see this in the inferred flags in some of the test cases.)
      * Other parts of the optimizer produce incrementing IVs, and they're common in idiomatic source language.  We do have support for reversing IVs, but in general if we produce one of each, the pair will persist surprisingly far through the optimizer before being coalesced.  (You can see this looking at nearby phis in the test cases.)
      
      Note that if the hardware prefers decrementing (i.e. zero tested) loops, LSR should convert back immediately before codegen.
      
      * Mostly irrelevant detail: The main loop of the prolog case is handled independently and will simple use the original IV with a changed start value.  We could in theory use this scheme for all iteration clamping, but that's a larger and more invasive change.
      37ead201
    • Lawrence D'Anna's avatar
      [lldb] temporarily disable TestPaths.test_interpreter_info on windows · 19cd6f31
      Lawrence D'Anna authored
      I'm disabling this test until the fix is reviewed
      (here https://reviews.llvm.org/D113650/)
      19cd6f31
    • Craig Topper's avatar
      [RISCV] Improve codegen for i32 udiv/urem by constant on RV64. · 02bed66c
      Craig Topper authored
      The division by constant optimization often produces constants that
      are uimm32, but not simm32. These constants require 3 or 4 instructions
      to materialize without Zba.
      
      Since these instructions are often used by a multiply with a LHS
      that needs to be zero extended with an AND, we can switch the MUL
      to a MULHU by shifting both inputs left by 32. Once we shift the
      constant left, the upper 32 bits no longer need to be 0 so constant
      materialization is free to use LUI+ADDIW. This reduces the constant
      materialization from 4 instructions to 3 in some cases while also
      reducing the zero extend of the LHS from 2 shifts to 1.
      
      Differential Revision: https://reviews.llvm.org/D113805
      02bed66c
    • Duncan P. N. Exon Smith's avatar
      lld: const-qualify iterations through VarStreamArray, NFC · 9a2b54af
      Duncan P. N. Exon Smith authored
      No functionality change here; just unblocking a patch to LLVM.
      9a2b54af
    • Duncan P. N. Exon Smith's avatar
      IR: Fix const-correctness of SwitchInst::CaseIterator and CaseHandle · a678c674
      Duncan P. N. Exon Smith authored
      Fix some confusion between the two types of `const` a pointer/iterator
      can have. Users of a SwitchInst::CaseIterator should not (and do not!)
      manually mutate the SwitchInst::CaseHandle that tracks its internal
      state. Change operator*() to return `const CaseHandle&`, remove the
      non-const-qualified operator*(), and const-qualify
      CaseHandle::setValue() and CaseHandle::setSuccessor().
      
      Differential Revision: https://reviews.llvm.org/D113788
      a678c674
    • Duncan P. N. Exon Smith's avatar
      ADT: Avoid repeating iterator adaptor/facade template params, NFC · c3edab8f
      Duncan P. N. Exon Smith authored
      Take advantage of class name injection to avoid redundantly specifying
      template parameters of iterator adaptor/facade base classes.
      
      No functionality change, although the private typedefs changed in a
      couple of cases.
      
        - Added a private typedef HashTableIterator::BaseT, following the
          pattern from r207084 / 3478d4b1, to
          pre-emptively appease MSVC (maybe it's not necessary anymore but
          looks like we do this pretty consistently). Otherwise, I removed
          private
        - Removed private typedefs filter_iterator_impl::BaseT and
          FilterIteratorTest::InputIterator::BaseT since there was only one
          use of each and the definition was no longer interesting.
      c3edab8f
    • Alexey Bataev's avatar
    • Félix Cloutier's avatar
      format_arg attribute does not support nullable instancetype return type · 12ab3e6c
      Félix Cloutier authored
      * The format_arg attribute tells the compiler that the attributed function
        returns a format string that is compatible with a format string that is being
        passed as a specific argument.
      * Several NSString methods return copies of their input, so they would ideally
        have the format_arg attribute. A previous differential (D112670) added
        support for instancetype methods having the format_arg attribute when used
        in the context of NSString method declarations.
      * D112670 failed to account that instancetype can be sugared in certain narrow
        (but critical) scenarios, like by using nullability specifiers. This patch
        resolves this problem.
      
      Differential Revision: https://reviews.llvm.org/D113636
      Reviewed By: ahatanak
      
      Radar-Id: rdar://85278860
      12ab3e6c
    • David Tenty's avatar
      [libcxx][AIX] XFAIL tests enabled by locale.fr_FR.UTF-8 · 4602f52d
      David Tenty authored
      We missed the tests in the earlier XFAIL-ing because the locale.fr_FR.UTF-8
      feature wasn't available, but since an upgrade these are now showing up
      on the CI.
      
      Differential Revision: https://reviews.llvm.org/D113791
      4602f52d
    • Mogball's avatar
      [mlir][ods] Cleanup of Class Codegen helper · 2696a952
      Mogball authored
      Depends on D113331
      
      Reviewed By: jpienaar
      
      Differential Revision: https://reviews.llvm.org/D113714
      2696a952
    • Peter Klausler's avatar
      [flang] Handle ENTRY names in IsPureProcedure() predicate · ece17064
      Peter Klausler authored
      Fortran defines an ENTRY point name as being pure if its enclosing
      subprogram scope defines a pure procedure.
      
      Differential Revision: https://reviews.llvm.org/D113711
      ece17064
    • Mogball's avatar
      8cf674f1
    • Vitaly Buka's avatar
      [asan] Fix GCC warning "left shift count >= width" · 07092ea6
      Vitaly Buka authored
      Fixes PR52385
      07092ea6
    • Jez Ng's avatar
      [lld-macho] Fix symbol relocs handling for LSDAs · 9d0b237c
      Jez Ng authored
      Similar to D113702, but for the LSDAs. Clang seems to emit all LSDA
      relocs as section relocs, but ld -r can turn those relocs into symbol
      ones.
      
      Reviewed By: #lld-macho, oontvoo
      
      Differential Revision: https://reviews.llvm.org/D113721
      9d0b237c
    • Jez Ng's avatar
      [lld-macho] Teach ICF to dedup functions with identical unwind info · d9b6f7e3
      Jez Ng authored
      Dedup'ing unwind info is tricky because each CUE contains a different
      function address, if ICF operated naively and compared the entire
      contents of each CUE, entries with identical unwind info but belonging
      to different functions would never be considered identical. To work
      around this problem, we slice away the function address before
      performing ICF. We rely on `relocateCompactUnwind()` to correctly handle
      these truncated input sections.
      
      Here are the numbers before and after D109944, D109945, and this diff
      were applied, as tested on my 3.2 GHz 16-Core Intel Xeon W:
      
      Without any optimizations:
      
                   base           diff           difference (95% CI)
        sys_time   0.849 ± 0.015  0.896 ± 0.012  [  +4.8% ..   +6.2%]
        user_time  3.357 ± 0.030  3.512 ± 0.023  [  +4.3% ..   +5.0%]
        wall_time  3.944 ± 0.039  4.032 ± 0.031  [  +1.8% ..   +2.6%]
        samples    40             38
      
      With `-dead_strip`:
      
                   base           diff           difference (95% CI)
        sys_time   0.847 ± 0.010  0.896 ± 0.012  [  +5.2% ..   +6.5%]
        user_time  3.377 ± 0.014  3.532 ± 0.015  [  +4.4% ..   +4.8%]
        wall_time  3.962 ± 0.024  4.060 ± 0.030  [  +2.1% ..   +2.8%]
        samples    47             30
      
      With `-dead_strip` and `--icf=all`:
      
                   base           diff           difference (95% CI)
        sys_time   0.935 ± 0.013  0.957 ± 0.018  [  +1.5% ..   +3.2%]
        user_time  3.472 ± 0.022  6.531 ± 0.046  [ +87.6% ..  +88.7%]
        wall_time  4.080 ± 0.040  5.329 ± 0.060  [ +30.0% ..  +31.2%]
        samples    37             30
      
      Unsurprisingly, ICF is now a lot slower, likely due to the much larger
      number of input sections it needs to process. But the rest of the
      linker only suffers a mild slowdown.
      
      Note that the compact-unwind-bad-reloc.s test was expanded because we
      now handle the relocation for CUE's function address in a separate code
      path from the rest of the CUE relocations. The extended test covers both
      code paths.
      
      Reviewed By: #lld-macho, oontvoo
      
      Differential Revision: https://reviews.llvm.org/D109946
      d9b6f7e3
    • Sanjay Patel's avatar
      [AArch64][x86] add tests for swapped cmp+vselect patterns; NFC · 6c32dd4d
      Sanjay Patel authored
      These patterns were noted in the recent D113212 and follow-ups.
      I did not bother to duplicate every test because it should be
      clear if we recognize the swaps from a smaller sample. We have
      complete coverage for the original patterns.
      6c32dd4d
    • wlei's avatar
      [llvm-profgen] Fix bug of setting function entry · aab18100
      wlei authored
      Previously we set `isFuncEntry` flag  to true when the funcName from DWARF is equal to the name in symbol table and we use this flag to ignore reporting callsite sample that's from an intra func branch. However, in HHVM, it appears that the symbol table name is inconsistent with the dwarf info func name, it's likely due to `OptimizeGlobalAliases`.
      
      This change is a workaround in llvm-profgen side to mark the only one range as the function entry and add warnings for the remaining inconsistence.
      
      This also fixed a missing `getCanonicalFnName` for symbol name which caused the mismatching as well.
      
      Reviewed By: hoy, wenlei
      
      Differential Revision: https://reviews.llvm.org/D113492
      aab18100
    • Aaron Puchert's avatar
      Comment Sema: Make most of CommentSema private (NFC) · 59b1e981
      Aaron Puchert authored
      We only need to expose setDecl, copyArray and the actOn* methods.
      59b1e981
    • Aaron Puchert's avatar
      Comment AST: Recognize function-like objects via return type (NFC) · 3010883f
      Aaron Puchert authored
      Instead of pretending that function pointer type aliases or variables
      are functions, and thereby losing the information that they are type
      aliases or variables, respectively, we use the existence of a return
      type in the DeclInfo to signify a "function-like" object.
      
      That seems pretty natural, since it's also the return type (or parameter
      list) from the DeclInfo that we compare the documentation with.
      
      Addresses a concern voiced in D111264#3115104.
      
      Reviewed By: gribozavr2
      
      Differential Revision: https://reviews.llvm.org/D113691
      3010883f
    • Aaron Puchert's avatar
      Comment AST: Find out if function is variadic in DeclInfo::fill · 4e7df1ef
      Aaron Puchert authored
      Then we don't have to look into the declaration again. Also it's only
      natural to collect this information alongside parameters and return
      type, as it's also just a parameter in some sense.
      
      Reviewed By: gribozavr2
      
      Differential Revision: https://reviews.llvm.org/D113690
      4e7df1ef
    • Peter Hawkins's avatar
      Don't define //mlir:MLIRBindingsPythonCore in terms of the NoCAPI and CAPIDeps targets. · 5074a20d
      Peter Hawkins authored
      We noticed that the library structure causes link ordering problems in Google's internal build. However, we don't think the problem is specific to Google's build, it probably can be reproduced anywhere with the right library structure.
      
      In general splitting the Python bindings from their dependencies (the C API targets) creates the possibility that the two libraries might end up in the wrong order on the linker command line. We can avoid this problem happening by reverting the structure of the MLIRBindingsPythonCore to represent its dependencies in the usual way, rather than composing an incomplete `MLIRBindingsPythonCoreNoCAPI` target and their CAPI dependencies. It was probably a mistake to rewrite this particular `cc_library()` rule in terms of the two, since nothing guarantees that the two will be correctly ordered by the linker when both are being linked into the same binary, and it was only an incidental "cleanup...
      5074a20d
    • Jez Ng's avatar
      [reland][lld-macho] Fix symbol relocs handling for compact unwind's functionAddress · ad8df21d
      Jez Ng authored
      Clang seems to emit all functionAddress relocs as section relocs, but
      `ld -r` can turn those relocs into symbol ones. It turns out that we
      weren't handling that case correctly when the symbol was a weak def
      whose definition did not prevail.
      
      Reviewed By: #lld-macho, oontvoo
      
      Differential Revision: https://reviews.llvm.org/D113702
      ad8df21d
    • Jacques Pienaar's avatar
      [mlir][shape] Add value_as_shape op · 153c2983
      Jacques Pienaar authored
      Part of the very first discussion here, but didn't upstream it before as we
      didn't use it yet. Fix that for pending updates. Just adding the op here,
      follow up will add the lowering to codegen.
      153c2983
    • Duncan P. N. Exon Smith's avatar
      Sema: const-qualify ParsedAttr::iterator::operator*() · 46a68c85
      Duncan P. N. Exon Smith authored
      `const`-qualify ParsedAttr::iterator::operator*(), clearing up confusion
      about the two meanings of const for pointers/iterators. Helps unblock
      removal of (non-const) iterator_facade_base::operator->().
      46a68c85
    • Duncan P. N. Exon Smith's avatar
      IR: Avoid duplication of SwitchInst::findCaseValue(), NFC · 8b3e1adf
      Duncan P. N. Exon Smith authored
      Change the non-const version of findCaseValue() to forward to the const
      version.
      8b3e1adf
    • Philip Reames's avatar
      [unroll] Keep unrolled iterations with initial iteration · de2fed61
      Philip Reames authored
      The unrolling code was previously inserting new cloned blocks at the end of the function.  The result of this with typical loop structures is that the new iterations are placed far from the initial iteration.
      
      With unrolling, the general assumption is that the a) the loop is reasonable hot, and b) the first Count-1 copies of the loop are rarely (if ever) loop exiting.  As such, placing Count-1 copies out of line is a fairly poor code placement choice.  We'd much rather fall through into the hot (non-exiting) path.  For code with branch profiles, later layout would fix this, but this may have a positive impact on non-PGO compiled code.
      
      However, the real motivation for this change isn't performance.  Its readability and human understanding.  Having to jump around long distances in an IR file to trace an unrolled loop structure is error prone and tedious.
      de2fed61
    • Peter Klausler's avatar
      [flang] Runtime performance improvements to real formatted input · da25f968
      Peter Klausler authored
      Profiling a basic internal real input read benchmark shows some
      hot spots in the code used to prepare input for decimal-to-binary
      conversion, which is of course where the time should be spent.
      The library that implements decimal to/from binary conversions has
      been optimized, but not the code in the Fortran runtime that calls it,
      and there are some obvious light changes worth making here.
      
      Move some member functions from *.cpp files into the class definitions
      of Descriptor and IoStatementState to enable inlining and specialization.
      
      Make GetNextInputBytes() the new basic input API within the
      runtime, replacing GetCurrentChar() -- which is rewritten in terms of
      GetNextInputBytes -- so that input routines can have the
      ability to acquire more than one input character at a time
      and amortize overhead.
      
      These changes speed up the time to read 1M random reals
      using internal I/O from a character array from 1.29s to 0.54s
      on my machine, which on par w...
      da25f968
    • Keith Smiley's avatar
      [lld-macho] Fix trailing slash in oso_prefix · eb6f9f31
      Keith Smiley authored
      Previously if you passed `-oso_prefix path/to/foo/` with a trailing
      slash at the end, using `real_path` would remove that slash, but that
      slash is necessary to make sure OSO prefix paths end up as valid
      relative paths instead of starting with `/`.
      
      Differential Revision: https://reviews.llvm.org/D113541
      eb6f9f31
    • Duncan P. N. Exon Smith's avatar
      ADT: Fix const-correctness of iterator adaptors · 1b651be0
      Duncan P. N. Exon Smith authored
      This fixes const-correctness of iterator adaptors, dropping non-`const`
      overloads for `operator*()`.
      
      Iterators, like the pointers that they generalize, have two types of
      `const`.
      
      The `const` qualifier on members indicates whether the iterator itself
      can be changed. This is analagous to `int *const`.
      
      The `const` qualifier on return values of `operator*()`, `operator[]()`,
      and `operator->()` controls whether the the pointed-to value can be
      changed. This is analogous to `const int *`.
      
      Since `operator*()` does not (in principle) change the iterator, then
      there should only be one definition, which is `const`-qualified. E.g.,
      iterators wrapping `int*` should look like:
      ```
      int *operator*() const; // always const-qualified, no overloads
      ```
      
      ba7a6b31 changed `iterator_adaptor_base`
      away from this to work around bugs in other iterator adaptors. That was
      already reverted. This patch adds back its test, which combined
      llvm::enumerate() and llvm::make_filter_range(), adds a test for
      iterator_adaptor_base itself, and cleans up the `const`-ness of the
      other iterator adaptors.
      
      This also updates the documented requirements for
      `iterator_facade_base`:
      ```
      /// OLD:
      ///   - const T &operator*() const;
      ///   - T &operator*();
      
      /// New:
      ///   - T &operator*() const;
      ```
      In a future commit we might also clean up `iterator_facade`'s overloads
      of `operator->()` and `operator[]()`. These already (correctly) return
      non-`const` proxies regardless of the iterator's `const` qualifier.
      
      Differential Revision: https://reviews.llvm.org/D113158
      1b651be0