1. Nov 18, 2021
    • Alex Zinenko's avatar
      [mlir] Fix wrong variable name in Linalg OpDSL · bca003de
      Alex Zinenko authored
      The name seems to have been left over from a renaming effort on an unexercised
      codepaths that are difficult to catch in Python. Fix it and add a test that
      exercises the codepath.
      
      Reviewed By: gysit
      
      Differential Revision: https://reviews.llvm.org/D114004
      bca003de
    • owenca's avatar
      e852cc0d
    • Louis Dionne's avatar
      [runtimes][NFC] Remove filenames at the top of the license notice · eb8650a7
      Louis Dionne authored
      We've stopped doing it in libc++ for a while now because these names
      would end up rotting as we move things around and copy/paste stuff.
      This cleans up all the existing files so as to stop the spreading
      as people copy-paste headers around.
      eb8650a7
    • Peyton, Jonathan L's avatar
      [OpenMP][libomp] Allow users to specify KMP_HW_SUBSET in any order · a0afb9d0
      Peyton, Jonathan L authored
      Remove restriction forcing users to specify the KMP_HW_SUBSET value in
      topology order. This patch sorts the user KMP_HW_SUBSET value before
      trying to apply it. For example: 1s,4c,2t is equivalent to 2t,1s,4c
      
      Differential Revision: https://reviews.llvm.org/D112027
      a0afb9d0
    • Lawrence D'Anna's avatar
      [lldb] remove usage of distutils, fix python path on debian/ubuntu · 63270710
      Lawrence D'Anna authored
      distutils is deprecated and will be removed, so we shouldn't be
      using it.
      
      We were using it to compute LLDB_PYTHON_RELATIVE_PATH.
      
      Discussing a similar issue
      [at python.org](https://bugs.python.org/issue41282), Filipe Laíns said:
      
          If you are relying on the value of distutils.sysconfig.get_python_lib()
          as you shown in your system, you probably don't want to. That
          directory (dist-packages) should be for Debian provided packages
          only, so moving to sysconfig.get_path() would be a good thing,
          as it has the correct value for user installed packages on your
          system.
      
      So I propose using a relative path from `sys.prefix` to
      `sysconfig.get_path("platlib")` instead.
      
      On Mac and windows, this results in the same paths as we had before,
      which are `lib/python3.9/site-packages` and `Lib\site-packages`,
      respectively.
      
      On ubuntu however, this will change the path from
      `lib/python3/dist-packages` to `lib/python3.9/site-packages`.
      
      This change seems to be correct, as Filipe said above, `dist-packages`
      belongs to the distribution, not us.
      
      Reviewed By: labath
      
      Differential Revision: https://reviews.llvm.org/D114106
      63270710
    • Louis Dionne's avatar
    • Yitzhak Mandelbaum's avatar
    • Nico Weber's avatar
      [clang] Fix typo in 36873fb7 · 3623163a
      Nico Weber authored
      3623163a
    • Nico Weber's avatar
      [clang] Try to fix test more after ae98182c · 36873fb7
      Nico Weber authored
      We need to use the td-based marshalling instead of doing this manually,
      else the setting gets lost on the way to codegen in most build configs.
      36873fb7
    • Jonathan Peyton's avatar
    • Philip Reames's avatar
      [SCEVAA] Avoid forming malformed pointer diff expressions · ad69402f
      Philip Reames authored
      This solves the same crash as in D104503, but with a different approach.
      
      The test case test_non_dom demonstrates a case where scev-aa crashes today. (If exercised either by -eval-aa or -licm.) The basic problem is that SCEV-AA expects to be able to compute a pointer difference between two SCEVs for any two pair of pointers we do an alias query on. For (valid, but out of scope) reasons, we can end up asking whether expressions in different sub-loops can alias each other. This results in a subtraction expression being formed where neither operand dominates the other.
      
      The approach this patch takes is to leverage the "defining scope" notion we introduced for flag semantics to detect and disallow the formation of the problematic SCEV. This ends up being relatively straight forward on that new infrastructure. This change does hint that we should probably be verifying a similar property for all SCEVs somewhere, but I'll leave that to a follow on change.
      
      Differential Revision: D114112
      ad69402f
    • Michael Liao's avatar
      Fix -Wparentheses warnings. NFC. · b861c360
      Michael Liao authored
      b861c360
    • Jonas Paulsson's avatar
      [SystemZ] [Sanitizer] Bugfixes in internal_clone(). · 4c32e3d9
      Jonas Paulsson authored
      The __flags variable needs to be of type 'long' in order to get sign extended
      properly.
      
      internal_clone() uses an svc (Supervisor Call) directly (as opposed to
      internal_syscall), and therefore needs to take care to set errno and return
      -1 as needed.
      
      Review: Ulrich Weigand
      4c32e3d9
    • Simon Pilgrim's avatar
      [X86] splitVector - only extract lower half subvector from splats · 5f99f771
      Simon Pilgrim authored
      If we're splitting a source vector that is a splat (with no undefs), just extract (for free) the lower half subvector and use it for both halfs.
      5f99f771
    • Nico Weber's avatar
      [clang] Try to fix test after ae98182c · a11d27f4
      Nico Weber authored
      The test assumes an integrated assembler, so use a triple where
      that's the default.
      a11d27f4
    • Pavel Labath's avatar
      [lldb] Port PlatformWindows, PlatformOpenBSD and PlatformRemoteGDBServer to... · 1b468f1c
      Pavel Labath authored
      [lldb] Port PlatformWindows, PlatformOpenBSD and PlatformRemoteGDBServer to GetSupportedArchitectures
      1b468f1c
    • Nico Weber's avatar
      [clang] Address review comments on https://reviews.llvm.org/D113707 · b1ad813b
      Nico Weber authored
      - Drop a needless `l` size suffix on a mov instruction in AT&T mode
      - Move varying bits of test flags to front
      - Add a comment about MS mode test
      b1ad813b
    • Michael Jones's avatar
      [libc] fix strtof/d/ld NaN parsing · 47d0c83e
      Michael Jones authored
      Fix the fact that previously strtof/d/ld would only accept a NaN as
      having parentheses if the thing in the parentheses was a valid number,
      now it will accept any combination of letters and numbers, but will only
      put valid numbers in the mantissa.
      
      Reviewed By: sivachandra
      
      Differential Revision: https://reviews.llvm.org/D113790
      47d0c83e
    • Simon Pilgrim's avatar
      3020608b
    • Simon Pilgrim's avatar
      [X86] LowerRotate - improve vXi8 rotate-by-scalar lowering with direct use of... · e76032c1
      Simon Pilgrim authored
      [X86] LowerRotate - improve vXi8 rotate-by-scalar lowering with direct use of (extended) shift-by-scalar helpers.
      
      If we're rotating vXi8 by a splatted amount, then unpack to vXi16, perform a SHL by the (extended) scalar, and then pack the results.
      
      This is a vector equivalent to the "rotl(x,y) -> (((aext(x) << bw) | zext(x)) << (y & (bw-1))) >> bw" style expansion we do for scalars in LowerFunnelShift.
      
      I think we can usefully use this for other vector types and vector funnel-shifts in the future, depending how we expand beyond D113192 for matching rotations/funnel-shifts for more type/ops.
      e76032c1
    • Mike Rice's avatar
      [OpenMP] Add version macro support for 5.1 and 5.2 · 69f35f89
      Mike Rice authored
      Differential Revision: https://reviews.llvm.org/D114102
      69f35f89
    • Stanislav Mekhanoshin's avatar
      [InstCombine] Generalize complex OR patterns to AND · 6d3db280
      Stanislav Mekhanoshin authored
      For every pattern with only NOT, OR, and AND operations there is
      always a symmetrical attern with AND and OR swapped.
      
      This adds 2 transformations: https://reviews.llvm.org/D113526
      
      ```
      (~(a & b) | c) & (~(a & c) | b) --> ~((b ^ c) & a)
      (~(a & b) | c) & ~(a & c) --> ~((b | c) & a)
      ```
      
      ```
      ----------------------------------------
      define i4 @src(i4 %a, i4 %b, i4 %c) {
      %0:
        %and1 = and i4 %b, %a
        %not1 = xor i4 %and1, 15
        %and2 = and i4 %a, %c
        %not2 = xor i4 %and2, 15
        %or = or i4 %not2, %b
        %r = and i4 %or, %not1
        ret i4 %r
      }
      =>
      define i4 @tgt(i4 %a, i4 %b, i4 %c) {
      %0:
        %or = or i4 %b, %c
        %and = and i4 %or, %a
        %r = xor i4 %and, 15
        ret i4 %r
      }
      Transformation seems to be correct!
      
      ----------------------------------------
      define i4 @src(i4 %a, i4 %b, i4 %c) {
      %0:
        %and1 = and i4 %a, %b
        %not1 = xor i4 %and1, 15
        %or1 = or i4 %not1, %c
        %and2 = and i4 %a, %c
        %not2 = xor i4 %and2, 15
        %or2 = or i4 %not2, %b
        %and3 = and i4 %or1, %or2
        ret i4 %and3
      }
      =>
      define i4 @tgt(i4 %a, i4 %b, i4 %c) {
      %0:
        %xor = xor i4 %b, %c
        %and = and i4 %xor, %a
        %not = xor i4 %and, 15
        ret i4 %not
      }
      Transformation seems to be correct!
      ```
      
      Differential Revision: https://reviews.llvm.org/D113526
      6d3db280
    • Nico Weber's avatar
      [llvm-objcopy] Fix some comment typos · 1718fe46
      Nico Weber authored
      1718fe46
    • Nico Weber's avatar
      [clang] Make -masm=intel affect inline asm style · ae98182c
      Nico Weber authored
      With this,
      
        void f() {  __asm__("mov eax, ebx"); }
      
      now compiles with clang with -masm=intel.
      
      This matches gcc.
      
      The flag is not accepted in clang-cl mode. It has no effect on
      MSVC-style `__asm {}` blocks, which are unconditionally in intel
      mode both before and after this change.
      
      One difference to gcc is that in clang, inline asm strings are
      "local" while they're "global" in gcc. Building the following with
      -masm=intel works with clang, but not with gcc where the ".att_syntax"
      from the 2nd __asm__() is in effect until file end (or until a
      ".intel_syntax" somewhere later in the file):
      
        __asm__("mov eax, ebx");
        __asm__(".att_syntax\nmovl %ebx, %eax");
        __asm__("mov eax, ebx");
      
      This also updates clang's intrinsic headers to work both in
      -masm=att (the default) and -masm=intel modes.
      The official solution for this according to "Multiple assembler dialects in asm
      templates" in gcc docs->Extensions->Inline Assembly->Extended Asm
      is to write every inline asm snippet twice:
      
          bt{l %[Offset],%[Base] | %[Base],%[Offset]}
      
      This works in LLVM after D113932 and D113894, so use that.
      
      (Just putting `.att_syntax` at the start of the snippet works in some but not
      all cases: When LLVM interpolates in parameters like `%0`, it uses at&t or
      intel syntax according to the inline asm snippet's flavor, so the `.att_syntax`
      within the snippet happens to late: The interpolated-in parameter is already
      in intel style, and then won't parse in the switched `.att_syntax`.)
      
      It might be nice to invent a `#pragma clang asm_dialect push "att"` /
      `#pragma clang asm_dialect pop` to be able to force asm style per snippet,
      so that the inline asm string doesn't contain the same code in two variants,
      but let's leave that for a follow-up.
      
      Fixes PR21401 and PR20241.
      
      Differential Revision: https://reviews.llvm.org/D113707
      ae98182c
    • Keith Smiley's avatar
      [llvm-objcopy][MachO] Add llvm-strip support for newer load commands · 68311f21
      Keith Smiley authored
      Previously llvm-strip would fail because of unknown commands.
      
      Fixes https://bugs.llvm.org/show_bug.cgi?id=50044
      
      Differential Revision: https://reviews.llvm.org/D113734
      68311f21
    • Louis Dionne's avatar
      [libc++] Refactor tests for trivially copyable atomics · 3e957e5d
      Louis Dionne authored
      - Replace irrelevant synopsis by a comment
      - Use a .verify.cpp test instead of .compile.fail.cpp
      - Remove unnecessary includes in one of the tests (was a copy-paste error)
      
      Differential Revision: https://reviews.llvm.org/D114094
      3e957e5d
    • Nico Weber's avatar
      [x86/asm] Let EmitMSInlineAsmStr() handle variants too · bf834b26
      Nico Weber authored
      This is preparation for D113707, where I want to make `-masm=intel`
      emit `asm inteldialect` instructions.
      
      `{movq %rbx, %rax|mov rax, rbx}` is supposed to evaluate to the bit
      between { and | for att and to the bit between | and } for intel.
      Since intel will become `asm inteldialect`, which alls EmitMSInlineAsmStr(),
      EmitMSInlineAsmStr() has to support variants as well.
      
      (clang translates `{...|...}` to `$(...$|...$)`. I'm not sure why
      it doesn't just send along only the first `...` or the second `...`
      to LLVM, but given the notes in PR23933 let's not do a big
      reorganization in this codepath.)
      
      Differential Revision: https://reviews.llvm.org/D113932
      bf834b26
    • Craig Topper's avatar
      [RISCV] Lower vector CTLZ_ZERO_UNDEF/CTTZ_ZERO_UNDEF by converting to FP and... · 0274be28
      Craig Topper authored
      [RISCV] Lower vector CTLZ_ZERO_UNDEF/CTTZ_ZERO_UNDEF by converting to FP and extracting the exponent.
      
      If we have a large enough floating point type that can exactly
      represent the integer value, we can convert the value to FP and
      use the exponent to calculate the leading/trailing zeros.
      
      The exponent will contain log2 of the value plus the exponent bias.
      We can then remove the bias and convert from log2 to leading/trailing
      zeros.
      
      This doesn't work for zero since the exponent of zero is zero so we
      can only do this for CTLZ_ZERO_UNDEF/CTTZ_ZERO_UNDEF. If we need
      a value for zero we can use a vmseq and a vmerge to handle it.
      
      We need to be careful to make sure the floating point type is legal.
      If it isn't we'll continue using the integer expansion. We could split the vector
      and concatenate the results but that needs some additional work and evaluation.
      
      Differential Revision: https://reviews.llvm.org/D111904
      0274be28
    • Nico Weber's avatar
      [x86/asm] Make variants work when converting at&t inline asm input to intel asm output · 103cc914
      Nico Weber authored
      `asm` always has AT&T-style input (`asm inteldialect` has Intel-style asm
      input), so EmitGCCInlineAsmStr() always has to pick the same variant since it
      cares about the input asm string, not the output asm string.
      
      For PowerPC, that default variant is 1. For other targets, it's 0.
      
      Without this, the included test case errors out with
      
          error: unknown use of instruction mnemonic without a size suffix
                   mov rax, rbx
      
      since it picks the intel branch and then tries to interpret it as AT&T
      when selecting intel-style output with `-x86-asm-syntax=intel`.
      
      Differential Revision: https://reviews.llvm.org/D113894
      103cc914
    • Kadir Cetinkaya's avatar
      [clangd] Dont include file version in task name · e76e5729
      Kadir Cetinkaya authored
      This will drop file version information from span names, reducing
      overall cardinality and also effect logging when skipping actions in scheduler.
      
      Differential Revision: https://reviews.llvm.org/D113390
      e76e5729
    • Keith Smiley's avatar
      693b0202
    • Peter Klausler's avatar
      [flang] Deal with negative character lengths in semantics · 78d60094
      Peter Klausler authored
      Fortran defines LEN(X) = 0 after CHARACTER(LEN=-1)::X so
      apply MAX(0, ...) to character length expressions.
      
      Differential Revision: https://reviews.llvm.org/D114030
      78d60094
    • DianQK's avatar
      Fix the side effect of outlined function when the register is implicit use and... · 1e9fa0b1
      DianQK authored
      Fix the side effect of outlined function when the register is implicit use and implicit-def in the same instruction.
      
      This is the diff associated with {D95267}, and we need to mark $x0 as live whether or not $x0 is dead.
      
      The compiler also needs to mark register $x0 as live in for the following case.
      
      ```
      $x1 = ADDXri $sp, 16, 0
      BL @spam, csr_darwin_aarch64_aapcs, implicit-def dead $lr, implicit $sp, implicit $x0, implicit killed $x1, implicit-def $sp, implicit-def $x0
      ```
      
      This change fixes an issue where the wrong registers were used when -machine-outliner-reruns>0.
      As an example:
      
      ```
      lang=c
      typedef struct {
          double v1;
          double v2;
      } D16;
      
      typedef struct {
          D16 v1;
          D16 v2;
      } D32;
      
      typedef long long LL8;
      typedef struct {
          long long v1;
          long long v2;
      } LL16;
      typedef struct {
          LL16 v1;
          LL16 v2;
      } LL32;
      
      typedef struct {
          LL32 v1;
          LL32 v2;
      } LL64;
      
      LL8 needx0(LL8 v0, LL8 v1);
      
      void bar(LL64 v1, LL32 v2, LL16 v3, LL32 v4, LL8 v5, D16 v6, D16 v7, D16 v8);
      
      LL8 foo(LL8 v0, LL64 v1, LL32 v2, LL16 v3, LL32 v4, LL8 v5, D16 v6, D16 v7, D16 v8)
      {
        LL8 result = needx0(v0, 0);
        bar(v1, v2, v3, v4, v5, v6, v7, v8);
        return result + 1;
      }
      ```
      
      As you can see from the `foo` function, we should not modify the value of `x0` until we call `needx0`.
      This code is compiled to give the following instruction MIR code.
      
      ```
      $sp = frame-setup SUBXri $sp, 256, 0
      frame-setup STPDi killed $d13, killed $d12, $sp, 16
      frame-setup STPDi killed $d11, killed $d10, $sp, 18
      frame-setup STPDi killed $d9, killed $d8, $sp, 20
      
      frame-setup STPXi killed $x26, killed $x25, $sp, 22
      frame-setup STPXi killed $x24, killed $x23, $sp, 24
      frame-setup STPXi killed $x22, killed $x21, $sp, 26
      frame-setup STPXi killed $x20, killed $x19, $sp, 28
      ...
      $x1 = MOVZXi 0, 0
      BL @needx0, csr_darwin_aarch64_aapcs, implicit-def dead $lr, implicit $sp, implicit $x0, implicit $x1, implicit-def $sp, implicit-def $x0
      ...
      ```
      
      Since there are some other instruction sequences that duplicate `foo`, after the first execution of Machine Outliner you will get:
      ```
      $sp = frame-setup SUBXri $sp, 256, 0
      frame-setup STPDi killed $d13, killed $d12, $sp, 16
      frame-setup STPDi killed $d11, killed $d10, $sp, 18
      frame-setup STPDi killed $d9, killed $d8, $sp, 20
      
      $x7 = ORRXrs $xzr, $lr, 0
      BL @OUTLINED_FUNCTION_0, implicit-def $lr, implicit $sp, implicit-def $lr, implicit $sp, implicit $xzr, implicit $x7, implicit $x19, implicit $x20, implicit $x21, implicit $x22, implicit $x23, implicit $x24, implicit $x25, implicit $x26
      $lr = ORRXrs $xzr, $x7, 0
      ...
      BL @OUTLINED_FUNCTION_1, implicit-def $lr, implicit $sp, implicit-def $lr, implicit-def $sp, implicit-def $x0, implicit-def $x1, implicit $sp
      ...
      ```
      
      For the first time we outlined the following sequence:
      ```
      frame-setup STPXi killed $x26, killed $x25, $sp, 22
      frame-setup STPXi killed $x24, killed $x23, $sp, 24
      frame-setup STPXi killed $x22, killed $x21, $sp, 26
      frame-setup STPXi killed $x20, killed $x19, $sp, 28
      ```
      and
      ```
      $x1 = MOVZXi 0, 0
      BL @needx0, csr_darwin_aarch64_aapcs, implicit-def dead $lr, implicit $sp, implicit $x0, implicit $x1, implicit-def $sp, implicit-def $x0
      ```
      
      When we execute the outline again, we will get:
      ```
      $x0 = ORRXrs $xzr, $lr, 0 <---- here
      BL @OUTLINED_FUNCTION_2_0, implicit-def $lr, implicit $sp, implicit-def $sp, implicit-def $lr, implicit $sp, implicit $xzr, implicit $d8, implicit $d9, implicit $d10, implicit $d11, implicit $d12, implicit $d13, implicit $x0
      $lr = ORRXrs $xzr, $x0, 0
      
      $x7 = ORRXrs $xzr, $lr, 0
      BL @OUTLINED_FUNCTION_0, implicit-def $lr, implicit $sp, implicit-def $lr, implicit $sp, implicit $xzr, implicit $x7, implicit $x19, implicit $x20, implicit $x21, implicit $x22, implicit $x23, implicit $x24, implicit $x25, implicit $x26
      $lr = ORRXrs $xzr, $x7, 0
      ...
      BL @OUTLINED_FUNCTION_1, implicit-def $lr, implicit $sp, implicit-def $lr, implicit-def $sp, implicit-def $x0, implicit-def $x1, implicit $sp
      ```
      
      When calling `OUTLINED_FUNCTION_2_0`, we used `x0` to save the `lr` register.
      The reason for the above error appears to be that:
      ```
      BL @OUTLINED_FUNCTION_1, implicit-def $lr, implicit $sp, implicit-def $lr, implicit-def $sp, implicit-def $x0, implicit-def $x1, implicit $sp
      ```
      should be:
      ```
      BL @OUTLINED_FUNCTION_1, implicit-def $lr, implicit $sp, implicit-def $lr, implicit-def $sp, implicit-def $x0, implicit-def $x1, implicit $sp, implicit $x0
      ```
      
      When processing the same instruction with both `implicit-def $x0` and `implicit $x0` we should keep `implicit $x0`.
      A reproducible demo is available at: [https://github.com/DianQK/reproduce_outlined_function_use_live_x0](https://github.com/DianQK/reproduce_outlined_function_use_live_x0).
      
      Reviewed By: jinlin
      
      Differential Revision: https://reviews.llvm.org/D112911
      1e9fa0b1
    • Ben Langmuir's avatar
      [JITLink] Allow duplicate symbol names for locals · 52737735
      Ben Langmuir authored
      Local symbols can have the same name. I ran into this with JITLink
      while working with an object file that had been run through `strip -S`
      that had many "func.eh" symbols, but it can also happen using `ld -r`.
      
      rdar://85352156
      
      Differential Revision: https://reviews.llvm.org/D114042
      52737735
    • Lawrence D'Anna's avatar
      [lldb] build failure for LLDB_PYTHON_EXE_RELATIVE_PATH on greendragon · f07ddbc6
      Lawrence D'Anna authored
      see: https://green.lab.llvm.org/green/view/LLDB/job/lldb-cmake/38387/console
      
      ```
      Could not find a relative path to sys.executable under sys.prefix
      tried: /usr/local/opt/python/bin/python3.7
      tried: /usr/local/opt/python/bin/../Frameworks/Python.framework/Versions/3.7/bin/python3.7
      sys.prefix: /usr/local/Cellar/python/3.7.1/Frameworks/Python.framework/Versions/3.7
      ```
      
      It was unable to find LLDB_PYTHON_EXE_RELATIVE_PATH because it was not resolving
      the real path of sys.prefix.
      
      caused by: https://reviews.llvm.org/D113650
      f07ddbc6
    • Jean Perier's avatar
      [flang] Check ArrayRef base for contiguity in IsSimplyContiguousHelper · 394d6fcf
      Jean Perier authored
      Previous code was returning true for `x(:)` where x is a pointer without
      the contiguous attribute.
      In case the array ref is a whole array section, check the base for contiguity
      to solve the issue.
      
      Differential Revision: https://reviews.llvm.org/D114084
      394d6fcf
    • Arthur Eubanks's avatar
      [gn build] Add missed comma · e1ef1406
      Arthur Eubanks authored
      e1ef1406
    • Arthur Eubanks's avatar
      [NewPM] Add option to prevent rerunning function pipeline on functions in CGSCC adaptor · e3e25b51
      Arthur Eubanks authored
      In a CGSCC pass manager, we may visit the same function multiple times
      due to SCC mutations. In the inliner pipeline, this results in running
      the function simplification pipeline on a function multiple times even
      if it hasn't been changed since the last function simplification
      pipeline run.
      
      We use a newly introduced analysis to keep track of whether or not a
      function has changed since the last time the function simplification
      pipeline has run on it. If we see this analysis available for a function
      in a CGSCCToFunctionPassAdaptor, we skip running the function passes on
      the function. The analysis is queried at the end of the function passes
      so that it's available after the first time the function simplification
      pipeline runs on a function. This is a per-adaptor option so it doesn't
      apply to every adaptor.
      
      The goal of this is to improve compile times. However, currently we
      can't turn this on by default at least for the ...
      e3e25b51
    • Alexey Bataev's avatar
    • Kazu Hirata's avatar