1. Jun 13, 2021
  2. Jun 12, 2021
    • Matheus Izvekov's avatar
      [clang] NRVO: Improvements and handling of more cases. · 1e50c3d7
      Matheus Izvekov authored
      
      
      This expands NRVO propagation for more cases:
      
      Parse analysis improvement:
      * Lambdas and Blocks with dependent return type can have their variables
        marked as NRVO Candidates.
      
      Variable instantiation improvements:
      * Fixes crash when instantiating NRVO variables in Blocks.
      * Functions, Lambdas, and Blocks which have auto return type have their
        variables' NRVO status propagated. For Blocks with non-auto return type,
        as a limitation, this propagation does not consider the actual return
        type.
      
      This also implements exclusion of VarDecls which are references to
      dependent types.
      
      Signed-off-by: default avatarMatheus Izvekov <mizvekov@gmail.com>
      
      Reviewed By: Quuxplusone
      
      Differential Revision: https://reviews.llvm.org/D99696
      1e50c3d7
    • Florian Hahn's avatar
    • Shashij gupta's avatar
      [MLIR] Simplify affine.if ops with trivial conditions · 466e5aba
      Shashij gupta authored
      
      
      The commit simplifies affine.if ops :
      The affine if operation gets removed if the condition is universally true or false and then/else block is merged with the parent block.
      
      Signed-off-by: default avatarShashij Gupta <shashij.gupta@polymagelabs.com>
      
      Reviewed By: bondhugula, pr4tgpt
      
      Differential Revision: https://reviews.llvm.org/D104015
      466e5aba
    • Florian Hahn's avatar
      b4583a5a
    • Kristina Bessonova's avatar
      [lit] Attempt for fix tests failing because of 'warning: non-portable path to file' · 8e627979
      Kristina Bessonova authored
      This is an attempt to fix clang test failures due to 'nonportable-include-path'
      warnings on Windows when a path to llvm-project's base directory contains some
      uppercase letters (excluding a drive letter).
      
      The issue originates from 2 problems:
      * discovery.py loads site config in lower case causing all the paths
      based on __file__ and requested within the config file to be in lowercase as well,
      * neither os.path.abspath() nor os.path.realpath() (both used to obtain paths of
      config files, sources, object directories, etc) do not return paths in the correct
      case for Windows (at least consistently for all python versions).
      
      As os.path library doesn't seem to provide any relaible way to restore
      the case for paths on Windows, this patch proposes to use pathlib.resolve().
      pathlib is a part of Python 3.4 while llvm lit requires Python 3.6.
      
      Reviewed By: Meinersbur
      
      Differential Revision: https://reviews.llvm.org/D103014
      8e627979
    • Florian Hahn's avatar
      Revert "[X86FixupLEAs] Transform the sequence LEA/SUB to SUB/SUB" · 5cd66420
      Florian Hahn authored
      This reverts commit 1b748faf because it
      breaks building the llvm-test-suite with -verify-machineinstrs on X86:
      http://green.lab.llvm.org/green/job/test-suite-verify-machineinstrs-x86_64-O3/9585/
      
      Running llc -verify-machineinstr on X86 crashes on the IR below:
      
          target datalayout = "e-m:o-p270:32:32-p271:32:32-p272:64:64-i64:64-f80:128-n8:16:32:64-S128"
      
          %struct.widget = type { i32, i32, i32, i32, i32*, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, [16 x [16 x i16]], [6 x [32 x i32]], [16 x [16 x i32]], [4 x [12 x [4 x [4 x i32]]]], [16 x i32], i8**, i32*, i32***, i32**, i32, i32, i32, i32, %struct.baz*, %struct.wobble.1*, i32, i32, i32, i32, i32, i32, %struct.quux.2*, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, [3 x i32], i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32***, i32***, i32****, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, [3 x [2 x i32]], [3 x [2 x i32]], i32, i32, i64, i64, %struct.zot.3, %struct.zot.3, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32 }
          %struct.baz = type { i32, i32, i32, i32, i32, i32, i32, i32, i32, %struct.snork*, %struct.wombat.0*, %struct.wobble*, i32, i32*, i32*, i32*, i32, i32*, i32*, i32*, i32 (%struct.widget*, %struct.eggs*)*, i32, i32, i32, i32 }
          %struct.snork = type { %struct.spam*, %struct.zot, i32 (%struct.wombat*, %struct.widget*, %struct.snork*)* }
          %struct.spam = type { i32, i32, i32, i32, i8*, i32 }
          %struct.zot = type { i32, i32, i32, i32, i32, i8*, i32* }
          %struct.wombat = type { i32, i32, i32, i32, i32, i32, i32, i32, void (i32, i32, i32*, i32*)*, void (%struct.wombat*, %struct.widget*, %struct.zot*)* }
          %struct.wombat.0 = type { [4 x [11 x %struct.quux]], [2 x [9 x %struct.quux]], [2 x [10 x %struct.quux]], [2 x [6 x %struct.quux]], [4 x %struct.quux], [4 x %struct.quux], [3 x %struct.quux] }
          %struct.quux = type { i16, i8 }
          %struct.wobble = type { [2 x %struct.quux], [4 x %struct.quux], [3 x [4 x %struct.quux]], [10 x [4 x %struct.quux]], [10 x [15 x %struct.quux]], [10 x [15 x %struct.quux]], [10 x [5 x %struct.quux]], [10 x [5 x %struct.quux]], [10 x [15 x %struct.quux]], [10 x [15 x %struct.quux]] }
          %struct.eggs = type { [1000 x i8], [1000 x i8], [1000 x i8], i32, i32, i32, i32, i32, i32, i32, i32 }
          %struct.wobble.1 = type { i32, [2 x i32], i32, i32, %struct.wobble.1*, %struct.wobble.1*, i32, [2 x [4 x [4 x [2 x i32]]]], i32, i64, i64, i32, i32, [4 x i8], [4 x i8], i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32 }
          %struct.quux.2 = type { i32, i32, i32, i32, i32, %struct.quux.2* }
          %struct.zot.3 = type { i64, i16, i16, i16 }
      
          define void @blam(%struct.widget* %arg, i32 %arg1) local_unnamed_addr {
          bb:
            %tmp = load i32, i32* undef, align 4
            %tmp2 = sdiv i32 %tmp, 6
            %tmp3 = sdiv i32 undef, 6
            %tmp4 = load i32, i32* undef, align 4
            %tmp5 = icmp eq i32 %tmp4, 4
            %tmp6 = select i1 %tmp5, i32 %tmp3, i32 %tmp2
            %tmp7 = getelementptr inbounds [4 x [4 x i32]], [4 x [4 x i32]]* undef, i64 0, i64 0, i64 0
            %tmp8 = zext i16 undef to i32
            %tmp9 = zext i16 undef to i32
            %tmp10 = load i16, i16* undef, align 2
            %tmp11 = zext i16 %tmp10 to i32
            %tmp12 = zext i16 undef to i32
            %tmp13 = zext i16 undef to i32
            %tmp14 = zext i16 undef to i32
            %tmp15 = load i16, i16* undef, align 2
            %tmp16 = zext i16 %tmp15 to i32
            %tmp17 = zext i16 undef to i32
            %tmp18 = sub nsw i32 %tmp8, %tmp9
            %tmp19 = shl nsw i32 undef, 1
            %tmp20 = add nsw i32 %tmp19, %tmp18
            %tmp21 = sub nsw i32 %tmp11, %tmp12
            %tmp22 = shl nsw i32 undef, 1
            %tmp23 = add nsw i32 %tmp22, %tmp21
            %tmp24 = sub nsw i32 %tmp13, %tmp14
            %tmp25 = shl nsw i32 undef, 1
            %tmp26 = add nsw i32 %tmp25, %tmp24
            %tmp27 = sub nsw i32 %tmp16, %tmp17
            %tmp28 = shl nsw i32 undef, 1
            %tmp29 = add nsw i32 %tmp28, %tmp27
            %tmp30 = sub nsw i32 %tmp20, %tmp29
            %tmp31 = sub nsw i32 %tmp23, %tmp26
            %tmp32 = shl nsw i32 %tmp30, 1
            %tmp33 = add nsw i32 %tmp32, %tmp31
            store i32 %tmp33, i32* undef, align 4
            %tmp34 = mul nsw i32 %tmp31, -2
            %tmp35 = add nsw i32 %tmp34, %tmp30
            store i32 %tmp35, i32* undef, align 4
            %tmp36 = select i1 %tmp5, i32 undef, i32 undef
            br label %bb37
      
          bb37:                                             ; preds = %bb
            %tmp38 = load i32, i32* undef, align 4
            %tmp39 = ashr i32 %tmp38, %tmp6
            %tmp40 = load i32, i32* undef, align 4
            %tmp41 = sdiv i32 %tmp39, %tmp40
            store i32 %tmp41, i32* undef, align 4
            ret void
          }
      5cd66420
    • Florian Hahn's avatar
      Revert "[X86FixupLEAs] Sub register usage of LEA dest should block LEA/SUB optimization" · e087b4f1
      Florian Hahn authored
      This reverts commit f35bcea1 because it
      depends on 1b748faf, which breaks
      building the llvm-test-suite with -verify-machineinstrs on X86.
      
      See 154adc0f135cff3f8a8861c335d2b88c8049d098 for more details.
      e087b4f1
    • madhur13490's avatar
      [AMDGPU][IndirectCalls] Fix register usage propagation for indirect/external calls · c27e8141
      madhur13490 authored
      This patch computes max SGPRs and VGPRs used by module
      in presence of indirect calls and makes that
      as register requirement for functions/kernels
      which makes indirect calls.
      
      This patch also refactors code AMDGPUSubTarget.cpp
      which add a "base" variants of getMaxNumSGPRs which
      is used by MachineFunction and new Function version.
      
      Reviewed By: arsenm
      
      Differential Revision: https://reviews.llvm.org/D103636
      c27e8141
    • spupyrev's avatar
      A post-processing for BFI inference · 0a0800c4
      spupyrev authored
      The current implementation for computing relative block frequencies does
      not handle correctly control-flow graphs containing irreducible loops. This
      results in suboptimally generated binaries, whose perf can be up to 5%
      worse than optimal.
      
      To resolve the problem, we apply a post-processing step, which iteratively
      updates block frequencies based on the frequencies of their predesessors.
      This corresponds to finding the stationary point of the Markov chain by
      an iterative method aka "PageRank computation". The algorithm takes at
      most O(|E| * IterativeBFIMaxIterations) steps but typically converges faster.
      
      It is turned on by passing option `use-iterative-bfi-inference`
      and applied only for functions containing profile data and irreducible loops.
      
      Tested on SPEC06/17, where it is helping to get correct profile counts for one of
      the binaries (403.gcc). In prod binaries, we've seen a speedup of up to 2%-5%
      for binaries containing functions with hot irreducible loops.
      
      Reviewed By: hoy, wenlei, davidxl
      
      Differential Revision: https://reviews.llvm.org/D103289
      0a0800c4
    • Michael Kruse's avatar
      [Flang][test] Fix Windows buildbot. · dbc26296
      Michael Kruse authored
      Commit 1b241b9b /
      patch https://reviews.llvm.org/D104130 introduced an new test which
      calls a UNIX shell script. Add
      REQUIRES: shell
      to not run it on Windows.
      dbc26296
    • Stephen Neuendorffer's avatar
      [mlir] make normalizeAffineFor public · 984e270a
      Stephen Neuendorffer authored
      Previously this was just a static method.
      984e270a
    • Adrian Prantl's avatar
    • Alexander Shaposhnikov's avatar
      [lld][MachO] Fix function starts section · b9095f5e
      Alexander Shaposhnikov authored
      Sort the addresses stored in FunctionStarts section.
      Previously we were encoding potentially large numbers (due to unsigned overflow).
      
      Test plan: make check-all
      
      Differential revision: https://reviews.llvm.org/D103662
      b9095f5e
    • Jez Ng's avatar
      [lld-macho] Fix debug build · 5de7467e
      Jez Ng authored
      D103977 broke a bunch of stuff as I had only tested the release build
      which eliminated asserts.
      
      I've retained the asserts where possible, but I also removed a bunch
      instead of adding a whole lot of verbose ConcatInputSection casts.
      5de7467e
    • Uday Bondhugula's avatar
      [MLIR] Execution engine python binding support for shared libraries · c8b8e8e0
      Uday Bondhugula authored
      Add support to Python bindings for the MLIR execution engine to load a
      specified list of shared libraries - for eg. to use MLIR runtime
      utility libraries.
      
      Differential Revision: https://reviews.llvm.org/D104009
      c8b8e8e0
    • Kai Luo's avatar
      [AIX][compiler-rt] Fix cmake build of libatomic for cmake-3.16+ · 6393164c
      Kai Luo authored
      cmake-3.16+ for AIX changes the default behavior of building a `SHARED` library which breaks AIX's build of libatomic, i.e., cmake-3.16+ builds `SHARED` as an archive of dynamic libraries. To fix it, we have to build `libatomic.so.1` as `MODULE` which keeps `libatomic.so.1` as an normal dynamic library.
      
      Reviewed By: jsji
      
      Differential Revision: https://reviews.llvm.org/D103786
      6393164c