1. Jun 22, 2020
    • Momchil Velikov's avatar
      [LTO] Use StringRef instead of C-style strings in setCodeGenDebugOptions · 75b0bbca
      Momchil Velikov authored
      Fixes an issue with missing nul-terminators and saves us some string
      copying, compared to a version which would insert nul-terminators.
      
      Differential Revision: https://reviews.llvm.org/D82033
      75b0bbca
    • Anatoly Trosinenko's avatar
      [MSP430] Enable some basic support for debug information · a5bd75aa
      Anatoly Trosinenko authored
      This commit technically permits LLVM to emit the debug information for ELF files for MSP430 architecture. Aside from this, it only defines the register numbers as defined by part 10.1 of MSP430 EABI specification (assuming the 1-byte subregisters share the register numbers with corresponding full-size registers).
      
      This commit was basically tested by me with TI-provided GCC 8.3.1 toolchain by compiling an example program with `clang` (please note manual linking may be required due to upstream `clang` not yet handling the `-msim` option necessary to run binaries on the GDB-provided simulator) and then running it and single-stepping with `msp430-elf-gdb` like this:
      
      ```
      $sysroot/bin/msp430-elf-gdb ./test -ex "target sim" -ex "load ./test"
      (gdb) ... traditional GDB commands follow ...
      ```
      
      While this implementation is most probably far from completeness and is considered experimental, it can already help with debugging MSP430 programs as well as finding issues in LLVM debug info support for MSP430 itself.
      
      One of the use cases includes trying to find a point where UBSan check in a trap-on-error mode was triggered.
      
      The expected debug information format is described in the [MSP430 Embedded Application Binary Interface](http://www.ti.com/lit/an/slaa534/slaa534.pdf) specification, part 10.
      
      Differential Revision: https://reviews.llvm.org/D81488
      a5bd75aa
    • Anatoly Trosinenko's avatar
      [DebugInfo] Explicitly permit addr_size = 0x02 when parsing DWARF data · 359fae6e
      Anatoly Trosinenko authored
      Current LLVM implementation uses `MCAsmInfo::CodePointerSize` as addr_size when emitting the DWARF data. llvm-dwarfdump, on the other hand, handles `addr_size`s of 4 and 8 properly and considers all other sizes as an error. This works for most of mainline targets except for MSP430 and AVR.
      
      msp430-gcc v8.3.1 emits DWARF32 with addr_size = 4 (DWARF32 does not imply addr_size = 4, 32 refers to internal offset width of 4 bytes) that is handled by llvm-dwarfdump already. Still, emitting 2-byte target pointers on MSP430 seems correct as well (but not for MSP430X that is supported by msp430-gcc but not by LLVM and has 20-bit address space).
      
      This patch make it possible for MSP430 debug info support to be tested with llvm-dwarfdump.
      
      Differential Revision: https://reviews.llvm.org/D82055
      359fae6e
    • Nathan James's avatar
      [clang-tidy] Improved accuracy of check list updater script · 23063296
      Nathan James authored
       - Added `FixItHint` comments to Check files for the script to mark those checks as offering fix-its when the fix-its are generated in another file.
       - Case insensitive file searching when looking for the file a checker code resides in.
      
      Also regenerated the list, sphinx had no issue generating the docs after this.
      
      Reviewed By: sylvestre.ledru
      
      Differential Revision: https://reviews.llvm.org/D81932
      23063296
    • Florian Hahn's avatar
    • Nathan James's avatar
      c2b22c57
    • Tobias Gysi's avatar
      [mlir] make the bitwidth of device side index computations configurable · d10b1a38
      Tobias Gysi authored
      The patch makes the index type lowering of the GPU to NVVM/ROCDL
      conversion configurable. It introduces a pass option that controls the
      bitwidth used when lowering index computations.
      
      Differential Revision: https://reviews.llvm.org/D80285
      d10b1a38
    • Balázs Kéri's avatar
      [Analyzer][StreamChecker] Add note tags for file opening. · e935a540
      Balázs Kéri authored
      Summary:
      Bug reports of resource leak are now improved.
      If there are multiple resource leak paths for the same stream,
      only one wil be reported.
      
      Reviewers: Szelethus, xazax.hun, baloghadamsoftware, NoQ
      
      Reviewed By: Szelethus, NoQ
      
      Subscribers: NoQ, rnkovacs, xazax.hun, baloghadamsoftware, szepet, a.sidorin, mikhail.ramalho, Szelethus, donat.nagy, dkrupp, gamesh411, Charusso, martong, ASDenysPetrov, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D81407
      e935a540
    • Djordje Todorovic's avatar
      [CSInfo][MIPS] Don't describe parameters loaded by sub/super reg copy · 792786e3
      Djordje Todorovic authored
      When describing parameter value loaded by a COPY instruction, consider
      case where needed Reg value is a sub- or super- register of the COPY
      instruction's destination register. Without this patch, compile process
      will crash with the assertion "TargetInstrInfo::describeLoadedValue
      can't describe super- or sub-regs for copy instructions".
      
      Patch by Nikola Tesic
      
      Differential revision: https://reviews.llvm.org/D82000
      792786e3
    • David Spickett's avatar
      [clang][Driver] Correct tool search path priority · 028571d6
      David Spickett authored
      Summary:
      As seen in:
      https://bugs.llvm.org/show_bug.cgi?id=45693
      
      When clang looks for a tool it has a set of
      possible names for it, in priority order.
      Previously it would look for these names in
      the program path. Then look for all the names
      in the PATH.
      
      This means that aarch64-none-elf-gcc on the PATH
      would lose to gcc in the program path.
      (which was /usr/bin in the bug's case)
      
      This changes that logic to search each name in both
      possible locations, then move to the next name.
      Which is more what you would expect to happen when
      using a non default triple.
      
      (-B prefixes maybe should follow this logic too,
      but are not changed in this patch)
      
      Subscribers: kristof.beyls, cfe-commits
      
      Tags: #clang
      
      Differential Revision: https://reviews.llvm.org/D79988
      028571d6
    • Stephan Herhut's avatar
      [mlir] Add for loop specialization · 4bcd08eb
      Stephan Herhut authored
      Summary:
      We already had a parallel loop specialization pass that is used to
      enable unrolling and consecutive vectorization by rewriting loops
      whose bound is defined as a min of a constant and a dynamic value
      into a loop with static bound (the constant) and the minimum as
      bound, wrapped into a conditional to dispatch between the two.
      This adds the same rewriting for for loops.
      
      Differential Revision: https://reviews.llvm.org/D82189
      4bcd08eb
    • Vassil Vassilev's avatar
      Return false if the identifier is not in the global module index. · 46ea465b
      Vassil Vassilev authored
      This allows clients to use the idiom:
      
      if (GlobalIndex->lookupIdentifier(Name, FoundModules)) {
        // work on the FoundModules
      }
      
      This is also a minor performance improvent for clang.
      
      Differential Revision: https://reviews.llvm.org/D81077
      46ea465b
    • Serguei Katkov's avatar
      [Peeling] Extend the scope of peeling a bit · 29b2c1ca
      Serguei Katkov authored
      Currently we allow peeling of the loops if there is a exiting latch block
      and all other exits are blocks ending with deopt.
      
      Actually we want that exit would end up with deopt unconditionally but
      it is not required that exit itself ends with deopt.
      
      Reviewers: reames, ashlykov, fhahn, apilipenko, fedor.sergeev
      Reviewed By: apilipenko
      Subscribers: hiraditya, zzheng, dantrushin, llvm-commits
      Differential Revision: https://reviews.llvm.org/D81140
      29b2c1ca
    • sameeran joshi's avatar
      [flang]Fix individual tests with lit when building out of tree · fa5d416e
      sameeran joshi authored
      Summary:
      
      Fix individual check tests with lit when building out-of-tree
      `ninja check-flang-<folder>` was not working.
      The CMakeLists.txt was looking for the lit tests in the source directory
      instead of the build directory.
      
      This commit extends @CarolineConcatto previous patch[D81002]
      
      Reviewers: DavidTruby, sscalpone, tskeith, CarolineConcatto, jdoerfert
      
      Reviewed By: DavidTruby
      
      Subscribers: flang-commits, llvm-commits, CarolineConcatto
      
      Tags: #flang, #llvm
      
      Differential Revision: https://reviews.llvm.org/D82120
      fa5d416e
    • Craig Topper's avatar
      [X86] Add an AVX check prefix to bitcast-vector-bool.ll to combine checks... · d3c79d19
      Craig Topper authored
      [X86] Add an AVX check prefix to bitcast-vector-bool.ll to combine checks where AVX1/2/512 are all the same. NFC
      d3c79d19
    • Craig Topper's avatar
      [X86] Add test file that was supposed to go with D81327. · 59d48ead
      Craig Topper authored
      Must have forgotten to git add the file.
      59d48ead
    • Michael Liao's avatar
      [amdgpu] Fix REL32 relocations with negative offsets. · 20a17002
      Michael Liao authored
      Summary: - The offset should be treated as a signed one.
      
      Reviewers: rampitec, arsenm
      
      Subscribers: kzhuravl, jvesely, wdng, nhaehnle, yaxunl, dstuttard, tpr, t-tye, hiraditya, kerbowa, llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D82234
      20a17002
    • Jez Ng's avatar
      [lld-macho] Refactor segment/section creation, sorting, and merging · 3646ee50
      Jez Ng authored
      Summary:
      There were a few issues with the previous setup:
      
      1. The section sorting comparator used a declarative map of section names to
        determine the correct order, but it turns out we need to match on more than
        just names -- in particular, an upcoming diff will sort based on whether the
        S_ZERO_FILL flag is set. This diff changes the sorter to a more imperative but
        flexible form.
      
      2. We were sorting OutputSections stored in a MapVector, which left the
        MapVector in an inconsistent state -- the wrong keys map to the wrong values!
        In practice, we weren't doing key lookups (only container iteration) after the
        sort, so this was fine, but it was still a dubious state of affairs. This diff
        copies the OutputSections to a vector before sorting them.
      
      3. We were adding unneeded OutputSections to OutputSegments and then filtering
        them out later, which meant that we had to remember whether an OutputSegment
        was in a pre- or post-filtered state. This diff only adds the sections to the
        segments if they are needed.
      
      In addition to those major changes, two minor ones worth noting:
      
      1. I renamed all OutputSection variable names to `osec`, to parallel `isec`.
        Previously we were using some inconsistent combination of `osec`, `os`, and
        `section`.
      
      2. I added a check (and a test) for InputSections with names that clashed with
        those of our synthetic OutputSections.
      
      Reviewers: #lld-macho
      
      Subscribers: llvm-commits
      
      Tags: #llvm
      
      Differential Revision: https://reviews.llvm.org/D81887
      3646ee50
    • Craig Topper's avatar
      [X86] Add cooperlake and tigerlake to the enum in cpu_model.c · 90406d62
      Craig Topper authored
      I forgot to do this when I added then to _cpu_indicator_init.
      90406d62
    • Craig Topper's avatar
      [X86] Assign a feature priority to 'tigerlake' so it won't assert when used... · 1d4c8733
      Craig Topper authored
      [X86] Assign a feature priority to 'tigerlake' so it won't assert when used with function multiversioning
      
      Also test cooperlake since it was also just added to function
      multiversioning when it was enabled for __builtin_cpu_is.
      1d4c8733
    • Sanjay Patel's avatar
      [VectorCombine] create class for pass to hold analyses, etc; NFC · 6bdd531a
      Sanjay Patel authored
      This doesn't change anything currently, but it would make sense
      to create a class-level IRBuilder instead of recreating that
      everywhere. As we expand to more optimizations, we will probably
      also want to hold things like the DataLayout or other constant
      refs in here too.
      6bdd531a
    • Craig Topper's avatar
      [X86] Add 'cooperlake' and 'tigerlake' to __builtin_cpu_is. · 42c176c3
      Craig Topper authored
      Cooperlake can be detect by compiler-rt now, but not libgcc yet.
      Tigerlake can't be detected by either. Both names are accepted by
      gcc. Hopefully the detection code will be in place soon.
      42c176c3
    • Craig Topper's avatar
      [X86] Add cooperlake detection to _cpu_indicator_init. · 0e6c9316
      Craig Topper authored
      libgcc has this enum encoding defined for a while, but their
      detection code is missing. I've raised a bug with them so that
      should get fixed soon.
      0e6c9316
    • Nathan James's avatar
      [clang-tidy] Implement storeOptions for checks missing it. · db90d315
      Nathan James authored
      Just adds the storeOptions for Checks that weren't already storing their options.
      
      Reviewed By: aaron.ballman
      
      Differential Revision: https://reviews.llvm.org/D82223
      db90d315
    • Luboš Luňák's avatar
      fix clang/PCH/delayed-pch-instantiate test · 448bbc51
      Luboš Luňák authored
      -target must match between PCH creation and use.
      448bbc51
  2. Jun 21, 2020
    • David Green's avatar
      67121d7b
    • Florian Hahn's avatar
      [DSE,MSSA] Move reachability check to main loop. · 40569db7
      Florian Hahn authored
      As we traverse the CFG backwards, we could end up reaching unreachable
      blocks. For unreachable blocks, we won't have computed post order
      numbers and because DomAccess is reachable, unreachable blocks cannot be
      on any path from it.
      
      This fixes a crash with unreachable blocks.
      40569db7
    • Luboš Luňák's avatar
      add option to instantiate templates already in the PCH · a45f713c
      Luboš Luňák authored
      Add -fpch-instantiate-templates which makes template instantiations be
      performed already in the PCH instead of it being done in every single
      file that uses the PCH (but every single file will still do it as well
      in order to handle its own instantiations). I can see 20-30% build
      time saved with the few tests I've tried.
      
      The change may reorder compiler output and also generated code, but
      should be generally safe and produce functionally identical code.
      There are some rare cases that do not compile with it,
      such as test/PCH/pch-instantiate-templates-forward-decl.cpp. If
      template instantiation bailed out instead of reporting the error,
      these instantiations could even be postponed, which would make them
      work.
      
      Enable this by default for clang-cl. MSVC creates PCHs by compiling
      them using an empty .cpp file, which means templates are instantiated
      while building the PCH and so the .h needs to be self-contained,
      making test/PCH/pch-instantiate-templates-forward-decl.cpp to fail
      with MSVC anyway. So the option being enabled for clang-cl matches this.
      
      Differential Revision: https://reviews.llvm.org/D69585
      a45f713c
    • David Green's avatar
      [CGP] Convert phi types · 730ecb63
      David Green authored
      If a collection of interconnected phi nodes is only ever loaded, stored
      or bitcast then we can convert the whole set to the bitcast type,
      potentially helping to reduce the number of register moves needed as the
      phi's are passed across basic block boundaries. This has to be done in
      CodegenPrepare as it naturally straddles basic blocks.
      
      The alorithm just looks from phi nodes, looking at uses and operands for
      a collection of nodes that all together are bitcast between float and
      integer types. We record visited phi nodes to not have to process them
      more than once. The whole subgraph is then replaced with a new type.
      Loads and Stores are bitcast to the correct type, which should then be
      folded into the load/store, changing it's type.
      
      This comes up in the biquad testcase due to the way MVE needs to keep
      values in integer registers. I have also seen it come up from aarch64
      partner example code, where a complicated set of sroa/inlining produced
      integer phis, where float would have been a better choice.
      
      I also added undef and extract element handling which increased the
      potency in some cases.
      
      This adds it with an option that defaults to off, and disabled for 32bit
      X86 due to potential issues around canonicalizing NaNs.
      
      Differential Revision: https://reviews.llvm.org/D81827
      730ecb63
    • David Green's avatar
      [CGP][AArch64] Convert Phi type tests. NFC · 0ee21cdb
      David Green authored
      0ee21cdb
    • Nikita Popov's avatar
      [ValueTracking, BasicAA] Don't simplify instructions · 37d30307
      Nikita Popov authored
      GetUnderlyingObject() (and by required symmetry
      DecomposeGEPExpression()) will call SimplifyInstruction() on the
      passed value if other checks fail. This simplification is very
      expensive, but has little effect in practice. This patch removes
      the SimplifyInstruction call(), and replaces it with a check for
      single-argument phis (which can occur in canonical IR in LCSSA
      form), which is the only useful simplification case I was able to
      identify.
      
      At O3 the geomean CTMark improvement is -1.7%. The largest
      improvement is SPASS with ThinLTO at -6%.
      
      In test-suite, I see only two tests with a hash difference and
      no code size difference (PAQ8p, Ptrdist), which indicates that
      the simplification only ends up being useful very rarely. (I would
      have liked to figure out which simplification is responsible here,
      but wasn't able to spot it looking at transformation logs.)
      
      The AMDGPU test case that is update was using two selects with
      undef condition, in which case GetUnderlyingObject will return
      the first select operand as the underlying object. This will of
      course not happen with non-undef conditions, so this was not
      testing anything realistic. Additionally this illustrates potential
      unsoundness: While GetUnderlyingObject will pick the first operand,
      the select might be later replaced by the second operand, resulting
      in inconsistent assumptions about the undef value.
      
      Differential Revision: https://reviews.llvm.org/D82261
      37d30307
    • Bruno Ricci's avatar
      Revert "Add --hot-func-list to llvm-profdata show for sample profiles" · 5342dd6b
      Bruno Ricci authored
      This reverts commit 7348b951.
      It is causing Asan failures.
      5342dd6b
    • Sanjay Patel's avatar
      [ValueTracking] improve analysis for fdiv with same operands · 2ad42c26
      Sanjay Patel authored
      (The 'nnan' variant of this pattern is already tested to produce '1.0'.)
      
      https://alive2.llvm.org/ce/z/D4hPBy
      
      define i1 @src(float %x, i32 %y) {
      %0:
        %d = fdiv float %x, %x
        %uge = fcmp uge float %d, 0.000000
        ret i1 %uge
      }
      =>
      define i1 @tgt(float %x, i32 %y) {
      %0:
        ret i1 1
      }
      Transformation seems to be correct!
      2ad42c26
    • Sanjay Patel's avatar
      97c02326
    • Bruno Ricci's avatar
      [clang][test][NFC] Also test for serialization in AST dump tests, part 3/n. · cddc9993
      Bruno Ricci authored
      The outputs between the direct ast-dump test and the ast-dump test after
      deserialization should match modulo a few differences.
      
      For hand-written tests, strip the "<undeserialized declarations>"s and
      the "imported"s with sed.
      
      For tests generated with "make-ast-dump-check.sh", regenerate the output.
      
      Part 3/n.
      cddc9993
    • Bruno Ricci's avatar
      [clang][test][NFC] Also test for serialization in AST dump tests, part 2/n. · ecbf2f5f
      Bruno Ricci authored
      The outputs between the direct ast-dump test and the ast-dump test after
      deserialization should match modulo a few differences.
      
      For hand-written tests, strip the "<undeserialized declarations>"s and
      the "imported"s with sed.
      
      For tests generated with "make-ast-dump-check.sh", regenerate the
      output.
      
      Part 2/n.
      ecbf2f5f
    • Bruno Ricci's avatar
    • Bruno Ricci's avatar
      [clang][utils] Minor tweak to make-ast-dump-check.sh · 0dbeffdd
      Bruno Ricci authored
      Remove the space after the "CHECK:" on each line. This space makes the use
      of FileCheck --match-full-lines impossible.
      0dbeffdd
    • Bruno Ricci's avatar
      [clang][Serialization] Fix the serialization of ConstantExpr. · e7ce0528
      Bruno Ricci authored
      The serialization of ConstantExpr has currently a number of problems:
      
      - Some fields are just not serialized (ConstantExprBits.APValueKind and
        ConstantExprBits.IsImmediateInvocation).
      
      - ASTStmtReader::VisitConstantExpr forgets to add the trailing APValue
        to the list of objects to be destroyed when the APValue needs cleanup.
      
      While we are at it, bring the serialization of ConstantExpr more in-line
      with what is done with the other expressions by doing the following NFCs:
      
      - Get rid of ConstantExpr::DefaultInit. It is better to not initialize
        the fields of an empty ConstantExpr since this will allow msan to
        detect if a field was not deserialized.
      
      - Move the initialization of the fields of ConstantExpr to the constructor;
        ConstantExpr::Create allocates the memory and ConstantExpr::ConstantExpr
        is responsible for the initialization.
      
      Review after commit since this is a straightforward mechanical fix
      similar to the other serialization fixes.
      e7ce0528
    • Bruno Ricci's avatar
      [clang][NFC] Fix typos/wording in the comments of ConstantExpr. · ef3adbfc
      Bruno Ricci authored
      It is "trailing objects" and "tail-allocated storage".
      ef3adbfc