1. Aug 25, 2021
    • Peilin Guo's avatar
      [DAGCombine] Check the legality of the index of EXTRACT_SUBVECTOR · 4c4dbeee
      Peilin Guo authored
      For ISD::EXTRACT_SUBVECTOR, its second operand must be a constant
      multiple of the known-minimum vector length of the result type.
      
      Reviewed By: dmgreen
      
      Differential Revision: https://reviews.llvm.org/D107795
      4c4dbeee
    • Jeremy Morse's avatar
      [DebugInfo][InstrRef] Avoid stack-slot-coloring changing codegen due to DI · cc1e87bf
      Jeremy Morse authored
      Stack slot colouring adds "weight" to slots if a non-dbg-value instruction
      refers to it. This, unfortunately, means that DBG_PHI instructions can have
      an effect on codegen. The fix is very simple, replace isDebugValue with
      isDebugInstr.
      
      The regression test contains a scenario that reproduces this problem; I've
      represented both normal-debug mode and instr-ref debug mode instructions
      in comment lines prefixed with AAAAAA and BBBBBB, and un-comment them with
      sed to test that the two different modes produce the same behaviour.
      
      Differential Revision: https://reviews.llvm.org/D108627
      cc1e87bf
    • River Riddle's avatar
      [mlir][AttrTypeGen] Add support for specifying a "accessor" type of a parameter · c8d9e1ce
      River Riddle authored
      This allows for using a different type when accessing a parameter than the
      one used for storage. This allows for returning parameters by reference,
      enables using more optimized/convient reference results, and more.
      
      Differential Revision: https://reviews.llvm.org/D108593
      c8d9e1ce
    • River Riddle's avatar
      [mlir] Update DialectAsmParser::parseString to use std::string instead of StringRef · 9658b061
      River Riddle authored
      This allows for parsing strings that have escape sequences, which require constructing
      a string (as they can't be represented by looking at the Token contents directly).
      
      Differential Revision: https://reviews.llvm.org/D108589
      9658b061
    • River Riddle's avatar
      [mlir] Move the Operation use iteration utilities to ResultRange · aea3026e
      River Riddle authored
      This allows for iterating and interacting with the uses of a specific subset of
      results as opposed to just the full range.
      
      Differential Revision: https://reviews.llvm.org/D108586
      aea3026e
    • Jean Perier's avatar
      [flang] Implement Posix version of DATE_AND_TIME runtime · b3e392c0
      Jean Perier authored
      Use gettimeofday and localtime_r to implement DATE_AND_TIME intrinsic.
      The Windows version fallbacks to the "no date and time information
      available" defined by the standard (strings set to blanks and values to
      -HUGE).
      
      The implementation uses an ifdef between windows and the rest because
      from my tests, the SFINAE approach leads to undeclared name bogus errors
      with clang 8 that seems to ignore failure to instantiate is not an error
      for the function names (i.e., it understands it should not instantiate
      the version using gettimeofday if it is not there, but still yields an
      error that it is not declared on the spot where it is called in the
      uninstantiated version).
      
      Differential Revision: https://reviews.llvm.org/D108622
      b3e392c0
    • Jan Svoboda's avatar
      [clang][deps] Ensure deterministic order of TU '-fmodule-file=' arguments · b5088cb4
      Jan Svoboda authored
      Translation units with multiple direct modular dependencies trigger a non-deterministic ordering in `clang-scan-deps`. This boils down to usage of `std::unordered_map`, which gets replaced by `std::map` in this patch.
      
      Depends on D103526.
      
      Reviewed By: dexonsmith
      
      Differential Revision: https://reviews.llvm.org/D103807
      b5088cb4
    • Rosie Sumpter's avatar
    • Tres Popp's avatar
    • LLVM GN Syncbot's avatar
      [gn build] Port 48958d02 · b0b26ae4
      LLVM GN Syncbot authored
      b0b26ae4
    • Daniil Fukalov's avatar
      [NFC][AMDGPU] Reduce includes dependencies. · 48958d02
      Daniil Fukalov authored
      1. Splitted out some parts of R600 target to separate modules/headers.
      2. Reduced some include lists in headers.
      3. Found and fixed issue with override `GCNTargetMachine::getSubtargetImpl()`
         and `R600TargetMachine::getSubtargetImpl()` had different return value type
         than base class.
      4. Minor forward declarations cleanup.
      
      Reviewed By: foad
      
      Differential Revision: https://reviews.llvm.org/D108596
      48958d02
    • Jan Svoboda's avatar
      [clang][deps] Use top-level modules as precompiled dependencies · 3b8f536f
      Jan Svoboda authored
      The `ASTReader` populates `Module::PresumedModuleMapFile` only for top-level modules, not submodules. To avoid generating empty `-fmodule-map-file=` arguments, make discovered modules depend on top-level precompiled modules. The granularity of submodules is not important here.
      
      The documentation of `Module::PresumedModuleMapFile` says this field is non-empty only when building from preprocessed source. This means there can still be cases where the dependency scanner generates empty `-fmodule-map-file=` arguments. That's being addressed in separate patch: D108544.
      
      Reviewed By: dexonsmith
      
      Differential Revision: https://reviews.llvm.org/D108647
      3b8f536f
    • serge-sans-paille's avatar
      Have lit preserve SOURCE_DATE_EPOCH · 46c947af
      serge-sans-paille authored
      This environment variable has been standardized for reproducible builds. Setting
      it can help to have reproducible tests too, so keep it as part of the testing
      env when set.
      
      See https://reproducible-builds.org/docs/source-date-epoch/
      
      Differential Revision: https://reviews.llvm.org/D108332
      46c947af
    • Jan Svoboda's avatar
      [clang][deps] Collect precompiled deps from submodules too · 83c633ea
      Jan Svoboda authored
      In this patch, the dependency scanner starts collecting precompiled dependencies from all encountered submodules, not only from top-level modules.
      
      Reviewed By: dexonsmith
      
      Differential Revision: https://reviews.llvm.org/D108540
      83c633ea
    • Florian Mayer's avatar
      [hwasan] do not check if freed pointer belonged to allocator. · 023f18bb
      Florian Mayer authored
      In that case it is very likely that there will be a tag mismatch anyway.
      
      We handle the case that the pointer belongs to neither of the allocators
      by getting a nullptr from allocator.GetBlockBegin.
      
      Reviewed By: hctim, eugenis
      
      Differential Revision: https://reviews.llvm.org/D108383
      023f18bb
    • Konstantin Schwarz's avatar
      [GlobalISel] Do not generate illegal G_SEXTLOADs after legalization · 4b4bc1ea
      Konstantin Schwarz authored
      The sext_inreg_of_load combine did not have the isLegalOrBeforeLegalizer check,
      leading to the generation of potentially illegal G_SEXTLOADs when run after legalization.
      
      Reviewed By: foad
      
      Differential Revision: https://reviews.llvm.org/D108626
      4b4bc1ea
    • Jonas Hahnfeld's avatar
      [CUDA] Fix static device variables with -fgpu-rdc · ea08c4cd
      Jonas Hahnfeld authored
      NVPTX does not allow dots in the identifier, so ptxas errors out with
         fatal   : Parsing error near '.static': syntax error
      because it parses .static as a directive. Avoid this problem by using
      two underscores, similar to what OpenMP does for outlined functions.
      
      Differential Revision: https://reviews.llvm.org/D108456
      ea08c4cd
    • Yi Kong's avatar
      [clang] Don't generate warn-stack-size when the warning is ignored · 5fc4828a
      Yi Kong authored
      8ace1213 introduced a regression for code that explicitly ignores the
      -Wframe-larger-than= warning. Make sure we don't generate the
      warn-stack-size attribute for that case.
      
      Differential Revision: https://reviews.llvm.org/D108686
      5fc4828a
    • Douglas Yung's avatar
      Add "REQUIRES: arm-registered-target" line to test added in D108603. · 323a6bfb
      Douglas Yung authored
      This should fix the test failure on the PS4 build bot.
      323a6bfb
    • Vang Thao's avatar
      [MachineCopyPropagation] Check CrossCopyRegClass for cross-class copys · 549f6a81
      Vang Thao authored
      On some AMDGPU subtargets, copying to and from AGPR registers using another
      AGPR register is not possible. A intermediate VGPR register is needed for AGPR
      to AGPR copy. This is an issue when machine copy propagation forwards a
      COPY $agpr, replacing a COPY $vgpr which results in $agpr = COPY $agpr. It is
      removing a cross class copy that may have been optimized by previous passes and
      potentially creating an unoptimized cross class copy later on.
      
      To avoid this issue, check CrossCopyRegClass if a different register class will
      be needed for the copy. If so then avoid forwarding the copy when the
      destination does not match the desired register class and if the original copy
      already matches the desired register class.
      
      Issue seen while attempting to optimize another AGPR to AGPR issue:
      
      Live-ins: $agpr0
      $vgpr0 = COPY $agpr0
      $agpr1 = V_ACCVGPR_WRITE_B32 $vgpr0
      $agpr2 = COPY $vgpr0
      $agpr3 = COPY $vgpr0
      $agpr4 = COPY $vgpr0
      
      After machine-cp:
      
      $vgpr0 = COPY $agpr0
      $agpr1 = V_ACCVGPR_WRITE_B32 $vgpr0
      $agpr2 = COPY $agpr0
      $agpr3 = COPY $agpr0
      $agpr4 = COPY $agpr0
      
      Machine-cp propagated COPY $agpr0 to replace $vgpr0 creating 3 AGPR to AGPR
      copys. Later this creates a cross-register copy from AGPR->VGPR->AGPR for each
      copy when the prior VGPR->AGPR copy was already optimal.
      
      Reviewed By: lkail, rampitec
      
      Differential Revision: https://reviews.llvm.org/D108011
      549f6a81
    • Lang Hames's avatar
      2a35d59b
    • Lang Hames's avatar
      [ORC] Fix typo in debugging output · fc3b2675
      Lang Hames authored
      fc3b2675
    • Carl Ritson's avatar
      [DAGCombine] Pre-commit test for D108619 · 28ba16c3
      Carl Ritson authored
      28ba16c3
    • Fangrui Song's avatar
      [InstrProfiling] Keep profd non-private for non-renamable comdat functions · 9ab9a959
      Fangrui Song authored
      The NS==0 condition used by D103717 missed a corner case: if the current copy
      does not have a hash suffix (e.g. weak_odr), a copy with value profiling (with a
      different CFG) may exist. This is super rare, but is possible with pre-inlining
      PGO instrumentation (which can make a weak_odr function inlines its callees
      differently, sometimes with value profiling while sometimes without).
      
      If the current copy with private profd is prevailing, the non-prevailing copy
      may get an undefined symbol if a caller inlining the non-prevailing function
      references its profd. If the other copy with non-private profd is prevailing,
      the current copy may cause a "relocation to discarded section" linker error.
      
      The fix is straightforward: just keep non-private profd in such a `DataReferencedByCode` case.
      
      With this change, a stage 2 (`-DLLVM_TARGETS_TO_BUILD=X86 -DLLVM_BUILD_INSTRUMENTED=IR`)
      clang is 0.08% larger (172431496/172286720-1).
      `stat -c %s **/*.o | awk '{s+=$1}END{print s}' is 0.026% larger.
      The majority of D103717's benefits remains.
      
      Reviewed By: xur
      
      Differential Revision: https://reviews.llvm.org/D108432
      9ab9a959
    • Richard Smith's avatar
      PR48030: Fix COMDAT-related linking problem with C++ thread_local static data members. · cd4d6d71
      Richard Smith authored
      Previously when emitting a C++ guarded initializer, we tried to work out what
      the enclosing function would be used for and added it to the COMDAT containing
      the variable if we thought that doing so would be correct. But this was done
      from a context in which we didn't -- and realistically couldn't -- correctly
      infer how the enclosing function would be used.
      
      Instead, add the initialization function to a COMDAT from the code that
      creates it, in the case where it makes sense to do so: when we know that
      the one and only reference to the initialization function is in
      @llvm.global.ctors and that reference is in the same COMDAT.
      
      Reviewed By: rjmccall
      
      Differential Revision: https://reviews.llvm.org/D108680
      cd4d6d71
    • Thomas Lively's avatar
      [WebAssembly] Fix some UB from ca541aa3 · 977eeb0c
      Thomas Lively authored
      977eeb0c
    • Fangrui Song's avatar
      Revert D108432 "[InstrProfiling] Keep profd non-private for non-renamable comdat functions" · 32e2326c
      Fangrui Song authored
      This reverts commit f653beea.
      
      It broke Windows coverage-inline.cpp because link.exe has a limitation
      that external symbols in IMAGE_COMDAT_SELECT_ASSOCIATIVE don't work.
      
      It essentially dropped the previous size optimization for coverage
      because coverage doesn't rename comdat by default.
      Needs more investigation what we should do.
      32e2326c
    • Rob Suderman's avatar
      [mlir][tosa] Quantized tosa.avg_pool2d lowering to linalg · 5541a05d
      Rob Suderman authored
      Includes the quantized version of average pool lowering to linalg dialect.
      This includes a lit test for the transform. It is not 100% correct as the
      multiplier / shift should be done in i64 however this is negligable rounding
      difference.
      
      Reviewed By: NatashaKnk
      
      Differential Revision: https://reviews.llvm.org/D108676
      5541a05d
    • Rob Suderman's avatar
      [mlir][tosa] Table did not apply offset before extract on i8 input · 4ef1770a
      Rob Suderman authored
      Lowering to table was incorrect as it did not apply a 128 offset before
      extracting the value from the table. Fixed and correct tensor length on input
      table.
      
      Reviewed By: NatashaKnk
      
      Differential Revision: https://reviews.llvm.org/D108436
      4ef1770a
    • Matthias Springer's avatar
      [mlir][SCF] Generalize AffineMinSCFCanonicalization to min/max ops · a9cff97f
      Matthias Springer authored
      * Add support for affine.max ops to SCF loop peeling pattern.
      * Add support for affine.max ops to `AffineMinSCFCanonicalizationPattern`.
      * Rename `AffineMinSCFCanonicalizationPattern` to `AffineOpSCFCanonicalizationPattern`.
      * Rename `AffineMinSCFCanonicalization` pass to `SCFAffineOpCanonicalization`.
      
      Differential Revision: https://reviews.llvm.org/D108009
      a9cff97f
    • wren romano's avatar
      [mlir][sparse] Correcting the use of emplace_back · 90e0c657
      wren romano authored
      The emplace commands are variadic and should take all the constructor arguments directly, since they implicitly call the constructor themselves in order to avoid the cost of constructing and then moving/copying temporaries.
      
      Reviewed By: aartbik
      
      Differential Revision: https://reviews.llvm.org/D108670
      90e0c657
    • Heejin Ahn's avatar
      [WebAssembly] Use SSAUpdaterBulk in LowerEmscriptenSjLj · d5244fb1
      Heejin Ahn authored
      We update SSA in two steps in Emscripten SjLj:
      1. Rewrite uses of `setjmpTable` and `setjmpTableSize` variables and
         place `phi`s where necessary, which are updated where we call
         `saveSetjmp`.
      2. Do a whole function level SSA update for all variables, because we
         split BBs where `setjmp` is called and there are possibly variable
         uses that are not dominated by a def.
         (See https://github.com/llvm/llvm-project/blob/955b91c19c00ed4c917559a5d66d14c669dde2e3/llvm/lib/Target/WebAssembly/WebAssemblyLowerEmscriptenEHSjLj.cpp#L1314-L1324)
      
      We have been using `SSAUpdater` to do this, but `SSAUpdaterBulk` class
      was added after this pass was first created, and for the step 2 it looks
      like a better alternative with a possible performance benefit. Not sure
      the author is aware of it, but `SSAUpdaterBulk` seems to have a
      limitation: it cannot handle a use within the same BB as a def but
      before it. For example:
      ```
      ... = %a + 1
      %a = foo();
      ```
      or
      ```
      %a = %a + 1
      ```
      The uses `%a` in RHS should be rewritten with another SSA variable of
      `%a`, most likely one generated from a `phi`. But `SSAUpdaterBulk`
      thinks all uses of `%a` are below the def of `%a` within the same BB.
      (`SSAUpdater` has two different functions of rewriting because of this:
      `RewriteUse` and `RewriteUseAfterInsertions`.) This doesn't affect our
      usage in the step 2 because that deals with possibly non-dominated uses
      by defs after block splitting. But it does in the step 1, which still
      uses `SSAUpdater`.
      
      But this CL also simplifies the step 1 by using `make_early_inc_range`,
      removing the need to advance the iterator before rewriting a use.
      
      This is NFC; the test changes are just the order of PHI nodes.
      
      Reviewed By: dschuff
      
      Differential Revision: https://reviews.llvm.org/D108583
      d5244fb1
    • Rob Suderman's avatar
      [mlir][tosa] Fix conv/depthwise conv padding for quantized values · a7bf9380
      Rob Suderman authored
      When padding quantized operations, the padding needs to equal the zero point
      of the input value. Corrected the pass to change the padding value if quantized.
      
      Reviewed By: NatashaKnk
      
      Differential Revision: https://reviews.llvm.org/D108440
      a7bf9380
    • Heejin Ahn's avatar
      [WebAssembly] Add Wasm SjLj option support for clang · a947b40c
      Heejin Ahn authored
      This adds support for Wasm SjLj in clang. Also this sets the new
      `-mllvm -wasm-enable-eh` option for Wasm EH.
      
      Note there is a little unfortunate inconsistency there: Wasm EH is
      enabled by a clang option `-fwasm-exceptions`, which sets
      `-mllvm -wasm-enable-eh` in the backend options. It also sets
      `-exception-model=wasm` but this is done in the common code.
      
      Wasm SjLj doesn't have a clang-level option like `-fwasm-exceptions`.
      `-fwasm-exceptions` was added because each exception model has its
      corresponding `-f***-exceptions`, but I'm not sure if adding a new
      option like `-fwasm-sjlj` or something is a good idea.
      
      So the current plan is Emscripten sets `-mllvm -wasm-enable-sjlj` if
      Wasm SJLj is enabled in its settings.js, as it does for Emscripten
      EH/SjLj (it sets `-mllvm -enable-emscripten-cxx-exceptions` for
      Emscripten EH and `-mllvm -enable-emscripten-sjlj` for Emscripten SjLj).
      And setting this enables the exception handling feature, and also sets
      `-exception-model=wasm`, but this time this is not done in the common
      code so we do it ourselves.
      
      Also note that other exception models have 1-to-1 correspondance with
      their `-f***-exceptions` flag and their `-exception-model=***` flag, but
      because we use `-exception-model=wasm` also for Wasm SjLj while
      `-fwasm-exceptions` still means Wasm EH, there is also a little
      inconsistency there, but I think it is manageable.
      
      Also this adds various error checking and tests.
      
      Reviewed By: dschuff
      
      Differential Revision: https://reviews.llvm.org/D108582
      a947b40c
    • Ed Maste's avatar
      [clang] allow -fstack-clash-protection on FreeBSD · 6609892a
      Ed Maste authored
      -fstack-clash-protection was added in Clang commit e67cbac8 but was
      enabled only on Linux.  Allow it on FreeBSD as well, as it works fine.
      
      Reviewed By: serge-sans-paille
      
      Differential Revision: https://reviews.llvm.org/D108571
      6609892a
    • Nico Weber's avatar
      [gn build] Manually port dbed061b more · 2847b8b6
      Nico Weber authored
      2847b8b6
    • Heejin Ahn's avatar
      [WebAssembly] Tidy up EH/SjLj options · 77b921b8
      Heejin Ahn authored
      This CL is small, but the description can be a little long because I'm
      trying to sum up the status quo for Emscripten/Wasm EH/SjLj options.
      
      First, this CL adds an option for Wasm SjLj (`-wasm-enable-sjlj`), which
      handles SjLj using Wasm EH. The implementation for this will be added as
      a followup CL, but this adds the option first to do error checking.
      
      This also adds an option for Wasm EH (`-wasm-enable-eh`), which has been
      already implemented. Before we used `-exception-model=wasm` as the same
      meaning as enabling Wasm EH, but after we add Wasm SjLj, it will be
      possible to use Wasm EH instructions for Wasm SjLj while not enabling
      EH, so going forward, to use Wasm EH, `opt` and `llc` will need this
      option. This only affects `opt` and `llc` command lines and does not
      affect Emscripten user interface.
      
      Now we have two modes of EH (Emscripten/Wasm) and also two modes of SjLj
      (also Emscripten/Wasm). The options corresponding to each of are:
      - Emscripten EH: `-enable-emscripten-cxx-exceptions`
      - Emscripten SjLj: `-enable-emscripten-sjlj`
      - Wasm EH: `-wasm-enable-eh -exception-model=wasm`
                 `-mattr=+exception-handling`
      - Wasm SjLj: `-wasm-enable-sjlj -exception-model=wasm`
                   `-mattr=+exception-handling`
      The reason Wasm EH/SjLj's options are a little complicated are
      `-exception-model` and `-mattr` are common LLVM options ane not under
      our control. (`-mattr` can be omitted if it is embedded within the
      bitcode file.)
      
      And we have the following rules of the option composition:
      - Emscripten EH and Wasm EH cannot be turned on at the same itme
      - Emscripten SjLj and Wasm SjLj cannot be turned on at the same time
      - Wasm SjLj should be used with Wasm EH
      
      Which means we now allow these combinations:
      - Emscripten EH + Emscripten SjLj: the current default in `emcc`
      - Wasm EH + Emscripten SjLj:
        This is allowed, but only as an interim step in which we are testing
        Wasm EH but not yet have a working implementation of Wasm SjLj. This
        will error out (D107687) in compile time if `setjmp` is called in a
        function in which Wasm exception is used.
      - Wasm EH + Wasm SjLj:
        This will be the default mode later when using Wasm EH. Currently Wasm
        SjLj implementation doesn't exist, so it doesn't work.
      - Emscripten EH + Wasm SjLj will not work.
      
      This CL moves these error checking routines to
      `WebAssemblyPassConfig::addIRPasses`. Not sure if this is an ideal place
      to do this, but I couldn't find elsewhere. Currently some checking is
      done within LowerEmscriptenEHSjLj, but these checks only run if
      LowerEmscriptenEHSjLj runs so it may not run when Wasm EH is used. This
      moves that to `addIRPasses` and adds some more checks.
      
      Currently LowerEmscriptenEHSjLj pass is responsible for Emscripten EH
      and Emscripten SjLj. Wasm EH transformations are done in multiple
      places, including WasmEHPrepare, LateEHPrepare, and CFGStackify. But in
      the followup CL, LowerEmscriptenEHSjLj pass will be also responsible for
      a part of Wasm SjLj transformation, because WasmSjLj will also be using
      several Emscripten library functions, and we will be sharing more than
      half of the transformation to do that between Emscripten SjLj and Wasm
      SjLj.
      
      Currently we have `-enable-emscripten-cxx-exceptions` and
      `-enable-emscripten-sjlj` but these only work for `llc`, because for
      `llc` we feed these options to the pass but when we run the pass using
      `opt` the pass will be created with no options and the default options
      will be used, which turns both Emscripten EH and Emscripten SjLj on.
      
      Now we have one more SjLj option to care for, LowerEmscriptenEHSjLj pass
      needs a finer way to control these options. This CL removes those
      default parameters and make LowerEmscriptenEHSjLj pass read directly
      from command line options specified. So if we only run
      `opt -wasm-lower-em-ehsjlj`, currently both Emscripten EH and Emscripten
      SjLj will run, but with this CL, none will run unless we additionally
      pass `-enable-emscripten-cxx-exceptions` or `-enable-emscripten-sjlj`,
      or both. This does not affect users; this only affects our `opt` tests
      because `emcc` will not call either `opt` or `llc`. As a result of this,
      our existing Emscripten EH/SjLj tests gained one or both of those
      options in their `RUN` lines.
      
      Reviewed By: dschuff
      
      Differential Revision: https://reviews.llvm.org/D107685
      77b921b8
    • Shimin Cui's avatar
      [GlobalOpt] Fix the assert for null check of global value · cea5ab09
      Shimin Cui authored
      This is to fix the reported assert - https://bugs.llvm.org/show_bug.cgi?id=51608.
      
      Reviewed By: asbirlea
      
      Differential Revision: https://reviews.llvm.org/D108674
      cea5ab09
    • Chenggang Zhao's avatar
      [mlir][docs] A friendlier improvement for the Toy tutorial chapter 4. · 2b2c13e6
      Chenggang Zhao authored
      Add notes for discarding private-visible functions in the Toy tutorial chapter 4.
      
      Reviewed By: mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D108026
      2b2c13e6
    • Jon Chesterfield's avatar
      ba854777