1. Aug 25, 2021
    • LLVM GN Syncbot's avatar
      [gn build] Port 48958d02 · b0b26ae4
      LLVM GN Syncbot authored
      b0b26ae4
    • Daniil Fukalov's avatar
      [NFC][AMDGPU] Reduce includes dependencies. · 48958d02
      Daniil Fukalov authored
      1. Splitted out some parts of R600 target to separate modules/headers.
      2. Reduced some include lists in headers.
      3. Found and fixed issue with override `GCNTargetMachine::getSubtargetImpl()`
         and `R600TargetMachine::getSubtargetImpl()` had different return value type
         than base class.
      4. Minor forward declarations cleanup.
      
      Reviewed By: foad
      
      Differential Revision: https://reviews.llvm.org/D108596
      48958d02
    • Jan Svoboda's avatar
      [clang][deps] Use top-level modules as precompiled dependencies · 3b8f536f
      Jan Svoboda authored
      The `ASTReader` populates `Module::PresumedModuleMapFile` only for top-level modules, not submodules. To avoid generating empty `-fmodule-map-file=` arguments, make discovered modules depend on top-level precompiled modules. The granularity of submodules is not important here.
      
      The documentation of `Module::PresumedModuleMapFile` says this field is non-empty only when building from preprocessed source. This means there can still be cases where the dependency scanner generates empty `-fmodule-map-file=` arguments. That's being addressed in separate patch: D108544.
      
      Reviewed By: dexonsmith
      
      Differential Revision: https://reviews.llvm.org/D108647
      3b8f536f
    • serge-sans-paille's avatar
      Have lit preserve SOURCE_DATE_EPOCH · 46c947af
      serge-sans-paille authored
      This environment variable has been standardized for reproducible builds. Setting
      it can help to have reproducible tests too, so keep it as part of the testing
      env when set.
      
      See https://reproducible-builds.org/docs/source-date-epoch/
      
      Differential Revision: https://reviews.llvm.org/D108332
      46c947af
    • Jan Svoboda's avatar
      [clang][deps] Collect precompiled deps from submodules too · 83c633ea
      Jan Svoboda authored
      In this patch, the dependency scanner starts collecting precompiled dependencies from all encountered submodules, not only from top-level modules.
      
      Reviewed By: dexonsmith
      
      Differential Revision: https://reviews.llvm.org/D108540
      83c633ea
    • Florian Mayer's avatar
      [hwasan] do not check if freed pointer belonged to allocator. · 023f18bb
      Florian Mayer authored
      In that case it is very likely that there will be a tag mismatch anyway.
      
      We handle the case that the pointer belongs to neither of the allocators
      by getting a nullptr from allocator.GetBlockBegin.
      
      Reviewed By: hctim, eugenis
      
      Differential Revision: https://reviews.llvm.org/D108383
      023f18bb
    • Konstantin Schwarz's avatar
      [GlobalISel] Do not generate illegal G_SEXTLOADs after legalization · 4b4bc1ea
      Konstantin Schwarz authored
      The sext_inreg_of_load combine did not have the isLegalOrBeforeLegalizer check,
      leading to the generation of potentially illegal G_SEXTLOADs when run after legalization.
      
      Reviewed By: foad
      
      Differential Revision: https://reviews.llvm.org/D108626
      4b4bc1ea
    • Jonas Hahnfeld's avatar
      [CUDA] Fix static device variables with -fgpu-rdc · ea08c4cd
      Jonas Hahnfeld authored
      NVPTX does not allow dots in the identifier, so ptxas errors out with
         fatal   : Parsing error near '.static': syntax error
      because it parses .static as a directive. Avoid this problem by using
      two underscores, similar to what OpenMP does for outlined functions.
      
      Differential Revision: https://reviews.llvm.org/D108456
      ea08c4cd
    • Yi Kong's avatar
      [clang] Don't generate warn-stack-size when the warning is ignored · 5fc4828a
      Yi Kong authored
      8ace1213 introduced a regression for code that explicitly ignores the
      -Wframe-larger-than= warning. Make sure we don't generate the
      warn-stack-size attribute for that case.
      
      Differential Revision: https://reviews.llvm.org/D108686
      5fc4828a
    • Douglas Yung's avatar
      Add "REQUIRES: arm-registered-target" line to test added in D108603. · 323a6bfb
      Douglas Yung authored
      This should fix the test failure on the PS4 build bot.
      323a6bfb
    • Vang Thao's avatar
      [MachineCopyPropagation] Check CrossCopyRegClass for cross-class copys · 549f6a81
      Vang Thao authored
      On some AMDGPU subtargets, copying to and from AGPR registers using another
      AGPR register is not possible. A intermediate VGPR register is needed for AGPR
      to AGPR copy. This is an issue when machine copy propagation forwards a
      COPY $agpr, replacing a COPY $vgpr which results in $agpr = COPY $agpr. It is
      removing a cross class copy that may have been optimized by previous passes and
      potentially creating an unoptimized cross class copy later on.
      
      To avoid this issue, check CrossCopyRegClass if a different register class will
      be needed for the copy. If so then avoid forwarding the copy when the
      destination does not match the desired register class and if the original copy
      already matches the desired register class.
      
      Issue seen while attempting to optimize another AGPR to AGPR issue:
      
      Live-ins: $agpr0
      $vgpr0 = COPY $agpr0
      $agpr1 = V_ACCVGPR_WRITE_B32 $vgpr0
      $agpr2 = COPY $vgpr0
      $agpr3 = COPY $vgpr0
      $agpr4 = COPY $vgpr0
      
      After machine-cp:
      
      $vgpr0 = COPY $agpr0
      $agpr1 = V_ACCVGPR_WRITE_B32 $vgpr0
      $agpr2 = COPY $agpr0
      $agpr3 = COPY $agpr0
      $agpr4 = COPY $agpr0
      
      Machine-cp propagated COPY $agpr0 to replace $vgpr0 creating 3 AGPR to AGPR
      copys. Later this creates a cross-register copy from AGPR->VGPR->AGPR for each
      copy when the prior VGPR->AGPR copy was already optimal.
      
      Reviewed By: lkail, rampitec
      
      Differential Revision: https://reviews.llvm.org/D108011
      549f6a81
    • Lang Hames's avatar
      2a35d59b
    • Lang Hames's avatar
      [ORC] Fix typo in debugging output · fc3b2675
      Lang Hames authored
      fc3b2675
    • Carl Ritson's avatar
      [DAGCombine] Pre-commit test for D108619 · 28ba16c3
      Carl Ritson authored
      28ba16c3
    • Fangrui Song's avatar
      [InstrProfiling] Keep profd non-private for non-renamable comdat functions · 9ab9a959
      Fangrui Song authored
      The NS==0 condition used by D103717 missed a corner case: if the current copy
      does not have a hash suffix (e.g. weak_odr), a copy with value profiling (with a
      different CFG) may exist. This is super rare, but is possible with pre-inlining
      PGO instrumentation (which can make a weak_odr function inlines its callees
      differently, sometimes with value profiling while sometimes without).
      
      If the current copy with private profd is prevailing, the non-prevailing copy
      may get an undefined symbol if a caller inlining the non-prevailing function
      references its profd. If the other copy with non-private profd is prevailing,
      the current copy may cause a "relocation to discarded section" linker error.
      
      The fix is straightforward: just keep non-private profd in such a `DataReferencedByCode` case.
      
      With this change, a stage 2 (`-DLLVM_TARGETS_TO_BUILD=X86 -DLLVM_BUILD_INSTRUMENTED=IR`)
      clang is 0.08% larger (172431496/172286720-1).
      `stat -c %s **/*.o | awk '{s+=$1}END{print s}' is 0.026% larger.
      The majority of D103717's benefits remains.
      
      Reviewed By: xur
      
      Differential Revision: https://reviews.llvm.org/D108432
      9ab9a959
    • Richard Smith's avatar
      PR48030: Fix COMDAT-related linking problem with C++ thread_local static data members. · cd4d6d71
      Richard Smith authored
      Previously when emitting a C++ guarded initializer, we tried to work out what
      the enclosing function would be used for and added it to the COMDAT containing
      the variable if we thought that doing so would be correct. But this was done
      from a context in which we didn't -- and realistically couldn't -- correctly
      infer how the enclosing function would be used.
      
      Instead, add the initialization function to a COMDAT from the code that
      creates it, in the case where it makes sense to do so: when we know that
      the one and only reference to the initialization function is in
      @llvm.global.ctors and that reference is in the same COMDAT.
      
      Reviewed By: rjmccall
      
      Differential Revision: https://reviews.llvm.org/D108680
      cd4d6d71
    • Thomas Lively's avatar
      [WebAssembly] Fix some UB from ca541aa3 · 977eeb0c
      Thomas Lively authored
      977eeb0c
    • Fangrui Song's avatar
      Revert D108432 "[InstrProfiling] Keep profd non-private for non-renamable comdat functions" · 32e2326c
      Fangrui Song authored
      This reverts commit f653beea.
      
      It broke Windows coverage-inline.cpp because link.exe has a limitation
      that external symbols in IMAGE_COMDAT_SELECT_ASSOCIATIVE don't work.
      
      It essentially dropped the previous size optimization for coverage
      because coverage doesn't rename comdat by default.
      Needs more investigation what we should do.
      32e2326c
    • Rob Suderman's avatar
      [mlir][tosa] Quantized tosa.avg_pool2d lowering to linalg · 5541a05d
      Rob Suderman authored
      Includes the quantized version of average pool lowering to linalg dialect.
      This includes a lit test for the transform. It is not 100% correct as the
      multiplier / shift should be done in i64 however this is negligable rounding
      difference.
      
      Reviewed By: NatashaKnk
      
      Differential Revision: https://reviews.llvm.org/D108676
      5541a05d
    • Rob Suderman's avatar
      [mlir][tosa] Table did not apply offset before extract on i8 input · 4ef1770a
      Rob Suderman authored
      Lowering to table was incorrect as it did not apply a 128 offset before
      extracting the value from the table. Fixed and correct tensor length on input
      table.
      
      Reviewed By: NatashaKnk
      
      Differential Revision: https://reviews.llvm.org/D108436
      4ef1770a
    • Matthias Springer's avatar
      [mlir][SCF] Generalize AffineMinSCFCanonicalization to min/max ops · a9cff97f
      Matthias Springer authored
      * Add support for affine.max ops to SCF loop peeling pattern.
      * Add support for affine.max ops to `AffineMinSCFCanonicalizationPattern`.
      * Rename `AffineMinSCFCanonicalizationPattern` to `AffineOpSCFCanonicalizationPattern`.
      * Rename `AffineMinSCFCanonicalization` pass to `SCFAffineOpCanonicalization`.
      
      Differential Revision: https://reviews.llvm.org/D108009
      a9cff97f
    • wren romano's avatar
      [mlir][sparse] Correcting the use of emplace_back · 90e0c657
      wren romano authored
      The emplace commands are variadic and should take all the constructor arguments directly, since they implicitly call the constructor themselves in order to avoid the cost of constructing and then moving/copying temporaries.
      
      Reviewed By: aartbik
      
      Differential Revision: https://reviews.llvm.org/D108670
      90e0c657
    • Heejin Ahn's avatar
      [WebAssembly] Use SSAUpdaterBulk in LowerEmscriptenSjLj · d5244fb1
      Heejin Ahn authored
      We update SSA in two steps in Emscripten SjLj:
      1. Rewrite uses of `setjmpTable` and `setjmpTableSize` variables and
         place `phi`s where necessary, which are updated where we call
         `saveSetjmp`.
      2. Do a whole function level SSA update for all variables, because we
         split BBs where `setjmp` is called and there are possibly variable
         uses that are not dominated by a def.
         (See https://github.com/llvm/llvm-project/blob/955b91c19c00ed4c917559a5d66d14c669dde2e3/llvm/lib/Target/WebAssembly/WebAssemblyLowerEmscriptenEHSjLj.cpp#L1314-L1324)
      
      We have been using `SSAUpdater` to do this, but `SSAUpdaterBulk` class
      was added after this pass was first created, and for the step 2 it looks
      like a better alternative with a possible performance benefit. Not sure
      the author is aware of it, but `SSAUpdaterBulk` seems to have a
      limitation: it cannot handle a use within the same BB as a def but
      before it. For example:
      ```
      ... = %a + 1
      %a = foo();
      ```
      or
      ```
      %a = %a + 1
      ```
      The uses `%a` in RHS should be rewritten with another SSA variable of
      `%a`, most likely one generated from a `phi`. But `SSAUpdaterBulk`
      thinks all uses of `%a` are below the def of `%a` within the same BB.
      (`SSAUpdater` has two different functions of rewriting because of this:
      `RewriteUse` and `RewriteUseAfterInsertions`.) This doesn't affect our
      usage in the step 2 because that deals with possibly non-dominated uses
      by defs after block splitting. But it does in the step 1, which still
      uses `SSAUpdater`.
      
      But this CL also simplifies the step 1 by using `make_early_inc_range`,
      removing the need to advance the iterator before rewriting a use.
      
      This is NFC; the test changes are just the order of PHI nodes.
      
      Reviewed By: dschuff
      
      Differential Revision: https://reviews.llvm.org/D108583
      d5244fb1
    • Rob Suderman's avatar
      [mlir][tosa] Fix conv/depthwise conv padding for quantized values · a7bf9380
      Rob Suderman authored
      When padding quantized operations, the padding needs to equal the zero point
      of the input value. Corrected the pass to change the padding value if quantized.
      
      Reviewed By: NatashaKnk
      
      Differential Revision: https://reviews.llvm.org/D108440
      a7bf9380
    • Heejin Ahn's avatar
      [WebAssembly] Add Wasm SjLj option support for clang · a947b40c
      Heejin Ahn authored
      This adds support for Wasm SjLj in clang. Also this sets the new
      `-mllvm -wasm-enable-eh` option for Wasm EH.
      
      Note there is a little unfortunate inconsistency there: Wasm EH is
      enabled by a clang option `-fwasm-exceptions`, which sets
      `-mllvm -wasm-enable-eh` in the backend options. It also sets
      `-exception-model=wasm` but this is done in the common code.
      
      Wasm SjLj doesn't have a clang-level option like `-fwasm-exceptions`.
      `-fwasm-exceptions` was added because each exception model has its
      corresponding `-f***-exceptions`, but I'm not sure if adding a new
      option like `-fwasm-sjlj` or something is a good idea.
      
      So the current plan is Emscripten sets `-mllvm -wasm-enable-sjlj` if
      Wasm SJLj is enabled in its settings.js, as it does for Emscripten
      EH/SjLj (it sets `-mllvm -enable-emscripten-cxx-exceptions` for
      Emscripten EH and `-mllvm -enable-emscripten-sjlj` for Emscripten SjLj).
      And setting this enables the exception handling feature, and also sets
      `-exception-model=wasm`, but this time this is not done in the common
      code so we do it ourselves.
      
      Also note that other exception models have 1-to-1 correspondance with
      their `-f***-exceptions` flag and their `-exception-model=***` flag, but
      because we use `-exception-model=wasm` also for Wasm SjLj while
      `-fwasm-exceptions` still means Wasm EH, there is also a little
      inconsistency there, but I think it is manageable.
      
      Also this adds various error checking and tests.
      
      Reviewed By: dschuff
      
      Differential Revision: https://reviews.llvm.org/D108582
      a947b40c
    • Ed Maste's avatar
      [clang] allow -fstack-clash-protection on FreeBSD · 6609892a
      Ed Maste authored
      -fstack-clash-protection was added in Clang commit e67cbac8 but was
      enabled only on Linux.  Allow it on FreeBSD as well, as it works fine.
      
      Reviewed By: serge-sans-paille
      
      Differential Revision: https://reviews.llvm.org/D108571
      6609892a
    • Nico Weber's avatar
      [gn build] Manually port dbed061b more · 2847b8b6
      Nico Weber authored
      2847b8b6
    • Heejin Ahn's avatar
      [WebAssembly] Tidy up EH/SjLj options · 77b921b8
      Heejin Ahn authored
      This CL is small, but the description can be a little long because I'm
      trying to sum up the status quo for Emscripten/Wasm EH/SjLj options.
      
      First, this CL adds an option for Wasm SjLj (`-wasm-enable-sjlj`), which
      handles SjLj using Wasm EH. The implementation for this will be added as
      a followup CL, but this adds the option first to do error checking.
      
      This also adds an option for Wasm EH (`-wasm-enable-eh`), which has been
      already implemented. Before we used `-exception-model=wasm` as the same
      meaning as enabling Wasm EH, but after we add Wasm SjLj, it will be
      possible to use Wasm EH instructions for Wasm SjLj while not enabling
      EH, so going forward, to use Wasm EH, `opt` and `llc` will need this
      option. This only affects `opt` and `llc` command lines and does not
      affect Emscripten user interface.
      
      Now we have two modes of EH (Emscripten/Wasm) and also two modes of SjLj
      (also Emscripten/Wasm). The options corresponding to each of are:
      - Emscripten EH: `-enable-emscripten-cxx-exceptions`
      - Emscripten SjLj: `-enable-emscripten-sjlj`
      - Wasm EH: `-wasm-enable-eh -exception-model=wasm`
                 `-mattr=+exception-handling`
      - Wasm SjLj: `-wasm-enable-sjlj -exception-model=wasm`
                   `-mattr=+exception-handling`
      The reason Wasm EH/SjLj's options are a little complicated are
      `-exception-model` and `-mattr` are common LLVM options ane not under
      our control. (`-mattr` can be omitted if it is embedded within the
      bitcode file.)
      
      And we have the following rules of the option composition:
      - Emscripten EH and Wasm EH cannot be turned on at the same itme
      - Emscripten SjLj and Wasm SjLj cannot be turned on at the same time
      - Wasm SjLj should be used with Wasm EH
      
      Which means we now allow these combinations:
      - Emscripten EH + Emscripten SjLj: the current default in `emcc`
      - Wasm EH + Emscripten SjLj:
        This is allowed, but only as an interim step in which we are testing
        Wasm EH but not yet have a working implementation of Wasm SjLj. This
        will error out (D107687) in compile time if `setjmp` is called in a
        function in which Wasm exception is used.
      - Wasm EH + Wasm SjLj:
        This will be the default mode later when using Wasm EH. Currently Wasm
        SjLj implementation doesn't exist, so it doesn't work.
      - Emscripten EH + Wasm SjLj will not work.
      
      This CL moves these error checking routines to
      `WebAssemblyPassConfig::addIRPasses`. Not sure if this is an ideal place
      to do this, but I couldn't find elsewhere. Currently some checking is
      done within LowerEmscriptenEHSjLj, but these checks only run if
      LowerEmscriptenEHSjLj runs so it may not run when Wasm EH is used. This
      moves that to `addIRPasses` and adds some more checks.
      
      Currently LowerEmscriptenEHSjLj pass is responsible for Emscripten EH
      and Emscripten SjLj. Wasm EH transformations are done in multiple
      places, including WasmEHPrepare, LateEHPrepare, and CFGStackify. But in
      the followup CL, LowerEmscriptenEHSjLj pass will be also responsible for
      a part of Wasm SjLj transformation, because WasmSjLj will also be using
      several Emscripten library functions, and we will be sharing more than
      half of the transformation to do that between Emscripten SjLj and Wasm
      SjLj.
      
      Currently we have `-enable-emscripten-cxx-exceptions` and
      `-enable-emscripten-sjlj` but these only work for `llc`, because for
      `llc` we feed these options to the pass but when we run the pass using
      `opt` the pass will be created with no options and the default options
      will be used, which turns both Emscripten EH and Emscripten SjLj on.
      
      Now we have one more SjLj option to care for, LowerEmscriptenEHSjLj pass
      needs a finer way to control these options. This CL removes those
      default parameters and make LowerEmscriptenEHSjLj pass read directly
      from command line options specified. So if we only run
      `opt -wasm-lower-em-ehsjlj`, currently both Emscripten EH and Emscripten
      SjLj will run, but with this CL, none will run unless we additionally
      pass `-enable-emscripten-cxx-exceptions` or `-enable-emscripten-sjlj`,
      or both. This does not affect users; this only affects our `opt` tests
      because `emcc` will not call either `opt` or `llc`. As a result of this,
      our existing Emscripten EH/SjLj tests gained one or both of those
      options in their `RUN` lines.
      
      Reviewed By: dschuff
      
      Differential Revision: https://reviews.llvm.org/D107685
      77b921b8
    • Shimin Cui's avatar
      [GlobalOpt] Fix the assert for null check of global value · cea5ab09
      Shimin Cui authored
      This is to fix the reported assert - https://bugs.llvm.org/show_bug.cgi?id=51608.
      
      Reviewed By: asbirlea
      
      Differential Revision: https://reviews.llvm.org/D108674
      cea5ab09
    • Chenggang Zhao's avatar
      [mlir][docs] A friendlier improvement for the Toy tutorial chapter 4. · 2b2c13e6
      Chenggang Zhao authored
      Add notes for discarding private-visible functions in the Toy tutorial chapter 4.
      
      Reviewed By: mehdi_amini
      
      Differential Revision: https://reviews.llvm.org/D108026
      2b2c13e6
    • Jon Chesterfield's avatar
      ba854777
    • Thomas Lively's avatar
      [WebAssembly] Fix up out-of-range BUILD_VECTOR lane constants · ca541aa3
      Thomas Lively authored
      Fixes PR51605 in which a DAG combine and legalization sequence generated
      out-of-range constants in BUILD_VECTOR lanes. In the v16i8 case, the constants
      were 255, which would be in range if DAG ISel used unsigned constants, but it is
      out of range because DAG ISel uses signed constants.
      
      Differential Revision: https://reviews.llvm.org/D108669
      ca541aa3
    • Vitaly Buka's avatar
      2d743af4
    • Vitaly Buka's avatar
      [msan] Don't EXPECT_POISONED beyond the we_wordv · 4c699b1c
      Vitaly Buka authored
      Partially reverts commit 629411d7.
      
      EXPECT_POISONED argument is outside of the allocation so we can't
      assume the state of shadow there.
      4c699b1c
    • Richard Smith's avatar
      Extend diagnostic for out of date AST input file. · df7b6b91
      Richard Smith authored
      If the size has changed, list the old and new sizes; if the mtime has
      changed, list the old and new mtimes (as raw time_t values).
      df7b6b91
    • Matthias Springer's avatar
      [mlir][linalg] Replace AffineMinSCFCanonicalizationPattern with SCF reimplementation · 2de2dbef
      Matthias Springer authored
      Use the new canonicalization pattern in the SCF dialect.
      
      Differential Revision: https://reviews.llvm.org/D107732
      2de2dbef
    • Vitaly Buka's avatar
      [msan] Fix wordexp after D108646 · 629411d7
      Vitaly Buka authored
      I introduced this bug reformating the patch before commit.
      629411d7
    • Amara Emerson's avatar
      Revert "[AArch64][GlobalISel] Don't contract cross-bank copies into truncating stores." · 2ed8053d
      Amara Emerson authored
      This reverts commit 67bf3ac7.
      
      The reason is that this change is now superseded by 04fb9b72 which fixes the
      underlying problem in the selector. Now it's fine to generate truncating FP stores
      since the selector code will just generate subreg copies to handle them.
      2ed8053d
    • Aart Bik's avatar
      [mlir][sparse] enable a few vectorized runs in integration tests · c5735fad
      Aart Bik authored
      Recent changes outside sparse compiler exposed the requirement of running a
      new pass (lower-affine) but this only became apparent with private testing.
      By adding some vectorized runs to integration test, we will detect the need
      for such changes earlier and also widen codegen coverage of course.
      
      Reviewed By: gussmith23
      
      Differential Revision: https://reviews.llvm.org/D108667
      c5735fad
    • Amara Emerson's avatar
      [AArch64][GlobalISel] Fix incorrect handling of fp truncating stores. · 04fb9b72
      Amara Emerson authored
      When the tablegen patterns fail to select a truncating scalar FPR store,
      our manual selection code also failed to handle it silently, trying to
      generate an invalid copy. Fix this by adding support in the manual code
      to generate a proper subreg copy before selecting a non-truncating store.
      04fb9b72