1. Nov 18, 2021
    • Nico Weber's avatar
      [clang] Address review comments on https://reviews.llvm.org/D113707 · b1ad813b
      Nico Weber authored
      - Drop a needless `l` size suffix on a mov instruction in AT&T mode
      - Move varying bits of test flags to front
      - Add a comment about MS mode test
      b1ad813b
    • Michael Jones's avatar
      [libc] fix strtof/d/ld NaN parsing · 47d0c83e
      Michael Jones authored
      Fix the fact that previously strtof/d/ld would only accept a NaN as
      having parentheses if the thing in the parentheses was a valid number,
      now it will accept any combination of letters and numbers, but will only
      put valid numbers in the mantissa.
      
      Reviewed By: sivachandra
      
      Differential Revision: https://reviews.llvm.org/D113790
      47d0c83e
    • Simon Pilgrim's avatar
      3020608b
    • Simon Pilgrim's avatar
      [X86] LowerRotate - improve vXi8 rotate-by-scalar lowering with direct use of... · e76032c1
      Simon Pilgrim authored
      [X86] LowerRotate - improve vXi8 rotate-by-scalar lowering with direct use of (extended) shift-by-scalar helpers.
      
      If we're rotating vXi8 by a splatted amount, then unpack to vXi16, perform a SHL by the (extended) scalar, and then pack the results.
      
      This is a vector equivalent to the "rotl(x,y) -> (((aext(x) << bw) | zext(x)) << (y & (bw-1))) >> bw" style expansion we do for scalars in LowerFunnelShift.
      
      I think we can usefully use this for other vector types and vector funnel-shifts in the future, depending how we expand beyond D113192 for matching rotations/funnel-shifts for more type/ops.
      e76032c1
    • Mike Rice's avatar
      [OpenMP] Add version macro support for 5.1 and 5.2 · 69f35f89
      Mike Rice authored
      Differential Revision: https://reviews.llvm.org/D114102
      69f35f89
    • Stanislav Mekhanoshin's avatar
      [InstCombine] Generalize complex OR patterns to AND · 6d3db280
      Stanislav Mekhanoshin authored
      For every pattern with only NOT, OR, and AND operations there is
      always a symmetrical attern with AND and OR swapped.
      
      This adds 2 transformations: https://reviews.llvm.org/D113526
      
      ```
      (~(a & b) | c) & (~(a & c) | b) --> ~((b ^ c) & a)
      (~(a & b) | c) & ~(a & c) --> ~((b | c) & a)
      ```
      
      ```
      ----------------------------------------
      define i4 @src(i4 %a, i4 %b, i4 %c) {
      %0:
        %and1 = and i4 %b, %a
        %not1 = xor i4 %and1, 15
        %and2 = and i4 %a, %c
        %not2 = xor i4 %and2, 15
        %or = or i4 %not2, %b
        %r = and i4 %or, %not1
        ret i4 %r
      }
      =>
      define i4 @tgt(i4 %a, i4 %b, i4 %c) {
      %0:
        %or = or i4 %b, %c
        %and = and i4 %or, %a
        %r = xor i4 %and, 15
        ret i4 %r
      }
      Transformation seems to be correct!
      
      ----------------------------------------
      define i4 @src(i4 %a, i4 %b, i4 %c) {
      %0:
        %and1 = and i4 %a, %b
        %not1 = xor i4 %and1, 15
        %or1 = or i4 %not1, %c
        %and2 = and i4 %a, %c
        %not2 = xor i4 %and2, 15
        %or2 = or i4 %not2, %b
        %and3 = and i4 %or1, %or2
        ret i4 %and3
      }
      =>
      define i4 @tgt(i4 %a, i4 %b, i4 %c) {
      %0:
        %xor = xor i4 %b, %c
        %and = and i4 %xor, %a
        %not = xor i4 %and, 15
        ret i4 %not
      }
      Transformation seems to be correct!
      ```
      
      Differential Revision: https://reviews.llvm.org/D113526
      6d3db280
    • Nico Weber's avatar
      [llvm-objcopy] Fix some comment typos · 1718fe46
      Nico Weber authored
      1718fe46
    • Nico Weber's avatar
      [clang] Make -masm=intel affect inline asm style · ae98182c
      Nico Weber authored
      With this,
      
        void f() {  __asm__("mov eax, ebx"); }
      
      now compiles with clang with -masm=intel.
      
      This matches gcc.
      
      The flag is not accepted in clang-cl mode. It has no effect on
      MSVC-style `__asm {}` blocks, which are unconditionally in intel
      mode both before and after this change.
      
      One difference to gcc is that in clang, inline asm strings are
      "local" while they're "global" in gcc. Building the following with
      -masm=intel works with clang, but not with gcc where the ".att_syntax"
      from the 2nd __asm__() is in effect until file end (or until a
      ".intel_syntax" somewhere later in the file):
      
        __asm__("mov eax, ebx");
        __asm__(".att_syntax\nmovl %ebx, %eax");
        __asm__("mov eax, ebx");
      
      This also updates clang's intrinsic headers to work both in
      -masm=att (the default) and -masm=intel modes.
      The official solution for this according to "Multiple assembler dialects in asm
      templates" in gcc docs->Extensions->Inline Assembly->Extended Asm
      is to write every inline asm snippet twice:
      
          bt{l %[Offset],%[Base] | %[Base],%[Offset]}
      
      This works in LLVM after D113932 and D113894, so use that.
      
      (Just putting `.att_syntax` at the start of the snippet works in some but not
      all cases: When LLVM interpolates in parameters like `%0`, it uses at&t or
      intel syntax according to the inline asm snippet's flavor, so the `.att_syntax`
      within the snippet happens to late: The interpolated-in parameter is already
      in intel style, and then won't parse in the switched `.att_syntax`.)
      
      It might be nice to invent a `#pragma clang asm_dialect push "att"` /
      `#pragma clang asm_dialect pop` to be able to force asm style per snippet,
      so that the inline asm string doesn't contain the same code in two variants,
      but let's leave that for a follow-up.
      
      Fixes PR21401 and PR20241.
      
      Differential Revision: https://reviews.llvm.org/D113707
      ae98182c
    • Keith Smiley's avatar
      [llvm-objcopy][MachO] Add llvm-strip support for newer load commands · 68311f21
      Keith Smiley authored
      Previously llvm-strip would fail because of unknown commands.
      
      Fixes https://bugs.llvm.org/show_bug.cgi?id=50044
      
      Differential Revision: https://reviews.llvm.org/D113734
      68311f21
    • Louis Dionne's avatar
      [libc++] Refactor tests for trivially copyable atomics · 3e957e5d
      Louis Dionne authored
      - Replace irrelevant synopsis by a comment
      - Use a .verify.cpp test instead of .compile.fail.cpp
      - Remove unnecessary includes in one of the tests (was a copy-paste error)
      
      Differential Revision: https://reviews.llvm.org/D114094
      3e957e5d
    • Nico Weber's avatar
      [x86/asm] Let EmitMSInlineAsmStr() handle variants too · bf834b26
      Nico Weber authored
      This is preparation for D113707, where I want to make `-masm=intel`
      emit `asm inteldialect` instructions.
      
      `{movq %rbx, %rax|mov rax, rbx}` is supposed to evaluate to the bit
      between { and | for att and to the bit between | and } for intel.
      Since intel will become `asm inteldialect`, which alls EmitMSInlineAsmStr(),
      EmitMSInlineAsmStr() has to support variants as well.
      
      (clang translates `{...|...}` to `$(...$|...$)`. I'm not sure why
      it doesn't just send along only the first `...` or the second `...`
      to LLVM, but given the notes in PR23933 let's not do a big
      reorganization in this codepath.)
      
      Differential Revision: https://reviews.llvm.org/D113932
      bf834b26
    • Craig Topper's avatar
      [RISCV] Lower vector CTLZ_ZERO_UNDEF/CTTZ_ZERO_UNDEF by converting to FP and... · 0274be28
      Craig Topper authored
      [RISCV] Lower vector CTLZ_ZERO_UNDEF/CTTZ_ZERO_UNDEF by converting to FP and extracting the exponent.
      
      If we have a large enough floating point type that can exactly
      represent the integer value, we can convert the value to FP and
      use the exponent to calculate the leading/trailing zeros.
      
      The exponent will contain log2 of the value plus the exponent bias.
      We can then remove the bias and convert from log2 to leading/trailing
      zeros.
      
      This doesn't work for zero since the exponent of zero is zero so we
      can only do this for CTLZ_ZERO_UNDEF/CTTZ_ZERO_UNDEF. If we need
      a value for zero we can use a vmseq and a vmerge to handle it.
      
      We need to be careful to make sure the floating point type is legal.
      If it isn't we'll continue using the integer expansion. We could split the vector
      and concatenate the results but that needs some additional work and evaluation.
      
      Differential Revision: https://reviews.llvm.org/D111904
      0274be28
    • Nico Weber's avatar
      [x86/asm] Make variants work when converting at&t inline asm input to intel asm output · 103cc914
      Nico Weber authored
      `asm` always has AT&T-style input (`asm inteldialect` has Intel-style asm
      input), so EmitGCCInlineAsmStr() always has to pick the same variant since it
      cares about the input asm string, not the output asm string.
      
      For PowerPC, that default variant is 1. For other targets, it's 0.
      
      Without this, the included test case errors out with
      
          error: unknown use of instruction mnemonic without a size suffix
                   mov rax, rbx
      
      since it picks the intel branch and then tries to interpret it as AT&T
      when selecting intel-style output with `-x86-asm-syntax=intel`.
      
      Differential Revision: https://reviews.llvm.org/D113894
      103cc914
    • Kadir Cetinkaya's avatar
      [clangd] Dont include file version in task name · e76e5729
      Kadir Cetinkaya authored
      This will drop file version information from span names, reducing
      overall cardinality and also effect logging when skipping actions in scheduler.
      
      Differential Revision: https://reviews.llvm.org/D113390
      e76e5729
    • Keith Smiley's avatar
      693b0202
    • Peter Klausler's avatar
      [flang] Deal with negative character lengths in semantics · 78d60094
      Peter Klausler authored
      Fortran defines LEN(X) = 0 after CHARACTER(LEN=-1)::X so
      apply MAX(0, ...) to character length expressions.
      
      Differential Revision: https://reviews.llvm.org/D114030
      78d60094
    • DianQK's avatar
      Fix the side effect of outlined function when the register is implicit use and... · 1e9fa0b1
      DianQK authored
      Fix the side effect of outlined function when the register is implicit use and implicit-def in the same instruction.
      
      This is the diff associated with {D95267}, and we need to mark $x0 as live whether or not $x0 is dead.
      
      The compiler also needs to mark register $x0 as live in for the following case.
      
      ```
      $x1 = ADDXri $sp, 16, 0
      BL @spam, csr_darwin_aarch64_aapcs, implicit-def dead $lr, implicit $sp, implicit $x0, implicit killed $x1, implicit-def $sp, implicit-def $x0
      ```
      
      This change fixes an issue where the wrong registers were used when -machine-outliner-reruns>0.
      As an example:
      
      ```
      lang=c
      typedef struct {
          double v1;
          double v2;
      } D16;
      
      typedef struct {
          D16 v1;
          D16 v2;
      } D32;
      
      typedef long long LL8;
      typedef struct {
          long long v1;
          long long v2;
      } LL16;
      typedef struct {
          LL16 v1;
          LL16 v2;
      } LL32;
      
      typedef struct {
          LL32 v1;
          LL32 v2;
      } LL64;
      
      LL8 needx0(LL8 v0, LL8 v1);
      
      void bar(LL64 v1, LL32 v2, LL16 v3, LL32 v4, LL8 v5, D16 v6, D16 v7, D16 v8);
      
      LL8 foo(LL8 v0, LL64 v1, LL32 v2, LL16 v3, LL32 v4, LL8 v5, D16 v6, D16 v7, D16 v8)
      {
        LL8 result = needx0(v0, 0);
        bar(v1, v2, v3, v4, v5, v6, v7, v8);
        return result + 1;
      }
      ```
      
      As you can see from the `foo` function, we should not modify the value of `x0` until we call `needx0`.
      This code is compiled to give the following instruction MIR code.
      
      ```
      $sp = frame-setup SUBXri $sp, 256, 0
      frame-setup STPDi killed $d13, killed $d12, $sp, 16
      frame-setup STPDi killed $d11, killed $d10, $sp, 18
      frame-setup STPDi killed $d9, killed $d8, $sp, 20
      
      frame-setup STPXi killed $x26, killed $x25, $sp, 22
      frame-setup STPXi killed $x24, killed $x23, $sp, 24
      frame-setup STPXi killed $x22, killed $x21, $sp, 26
      frame-setup STPXi killed $x20, killed $x19, $sp, 28
      ...
      $x1 = MOVZXi 0, 0
      BL @needx0, csr_darwin_aarch64_aapcs, implicit-def dead $lr, implicit $sp, implicit $x0, implicit $x1, implicit-def $sp, implicit-def $x0
      ...
      ```
      
      Since there are some other instruction sequences that duplicate `foo`, after the first execution of Machine Outliner you will get:
      ```
      $sp = frame-setup SUBXri $sp, 256, 0
      frame-setup STPDi killed $d13, killed $d12, $sp, 16
      frame-setup STPDi killed $d11, killed $d10, $sp, 18
      frame-setup STPDi killed $d9, killed $d8, $sp, 20
      
      $x7 = ORRXrs $xzr, $lr, 0
      BL @OUTLINED_FUNCTION_0, implicit-def $lr, implicit $sp, implicit-def $lr, implicit $sp, implicit $xzr, implicit $x7, implicit $x19, implicit $x20, implicit $x21, implicit $x22, implicit $x23, implicit $x24, implicit $x25, implicit $x26
      $lr = ORRXrs $xzr, $x7, 0
      ...
      BL @OUTLINED_FUNCTION_1, implicit-def $lr, implicit $sp, implicit-def $lr, implicit-def $sp, implicit-def $x0, implicit-def $x1, implicit $sp
      ...
      ```
      
      For the first time we outlined the following sequence:
      ```
      frame-setup STPXi killed $x26, killed $x25, $sp, 22
      frame-setup STPXi killed $x24, killed $x23, $sp, 24
      frame-setup STPXi killed $x22, killed $x21, $sp, 26
      frame-setup STPXi killed $x20, killed $x19, $sp, 28
      ```
      and
      ```
      $x1 = MOVZXi 0, 0
      BL @needx0, csr_darwin_aarch64_aapcs, implicit-def dead $lr, implicit $sp, implicit $x0, implicit $x1, implicit-def $sp, implicit-def $x0
      ```
      
      When we execute the outline again, we will get:
      ```
      $x0 = ORRXrs $xzr, $lr, 0 <---- here
      BL @OUTLINED_FUNCTION_2_0, implicit-def $lr, implicit $sp, implicit-def $sp, implicit-def $lr, implicit $sp, implicit $xzr, implicit $d8, implicit $d9, implicit $d10, implicit $d11, implicit $d12, implicit $d13, implicit $x0
      $lr = ORRXrs $xzr, $x0, 0
      
      $x7 = ORRXrs $xzr, $lr, 0
      BL @OUTLINED_FUNCTION_0, implicit-def $lr, implicit $sp, implicit-def $lr, implicit $sp, implicit $xzr, implicit $x7, implicit $x19, implicit $x20, implicit $x21, implicit $x22, implicit $x23, implicit $x24, implicit $x25, implicit $x26
      $lr = ORRXrs $xzr, $x7, 0
      ...
      BL @OUTLINED_FUNCTION_1, implicit-def $lr, implicit $sp, implicit-def $lr, implicit-def $sp, implicit-def $x0, implicit-def $x1, implicit $sp
      ```
      
      When calling `OUTLINED_FUNCTION_2_0`, we used `x0` to save the `lr` register.
      The reason for the above error appears to be that:
      ```
      BL @OUTLINED_FUNCTION_1, implicit-def $lr, implicit $sp, implicit-def $lr, implicit-def $sp, implicit-def $x0, implicit-def $x1, implicit $sp
      ```
      should be:
      ```
      BL @OUTLINED_FUNCTION_1, implicit-def $lr, implicit $sp, implicit-def $lr, implicit-def $sp, implicit-def $x0, implicit-def $x1, implicit $sp, implicit $x0
      ```
      
      When processing the same instruction with both `implicit-def $x0` and `implicit $x0` we should keep `implicit $x0`.
      A reproducible demo is available at: [https://github.com/DianQK/reproduce_outlined_function_use_live_x0](https://github.com/DianQK/reproduce_outlined_function_use_live_x0).
      
      Reviewed By: jinlin
      
      Differential Revision: https://reviews.llvm.org/D112911
      1e9fa0b1
    • Ben Langmuir's avatar
      [JITLink] Allow duplicate symbol names for locals · 52737735
      Ben Langmuir authored
      Local symbols can have the same name. I ran into this with JITLink
      while working with an object file that had been run through `strip -S`
      that had many "func.eh" symbols, but it can also happen using `ld -r`.
      
      rdar://85352156
      
      Differential Revision: https://reviews.llvm.org/D114042
      52737735
    • Lawrence D'Anna's avatar
      [lldb] build failure for LLDB_PYTHON_EXE_RELATIVE_PATH on greendragon · f07ddbc6
      Lawrence D'Anna authored
      see: https://green.lab.llvm.org/green/view/LLDB/job/lldb-cmake/38387/console
      
      ```
      Could not find a relative path to sys.executable under sys.prefix
      tried: /usr/local/opt/python/bin/python3.7
      tried: /usr/local/opt/python/bin/../Frameworks/Python.framework/Versions/3.7/bin/python3.7
      sys.prefix: /usr/local/Cellar/python/3.7.1/Frameworks/Python.framework/Versions/3.7
      ```
      
      It was unable to find LLDB_PYTHON_EXE_RELATIVE_PATH because it was not resolving
      the real path of sys.prefix.
      
      caused by: https://reviews.llvm.org/D113650
      f07ddbc6
    • Jean Perier's avatar
      [flang] Check ArrayRef base for contiguity in IsSimplyContiguousHelper · 394d6fcf
      Jean Perier authored
      Previous code was returning true for `x(:)` where x is a pointer without
      the contiguous attribute.
      In case the array ref is a whole array section, check the base for contiguity
      to solve the issue.
      
      Differential Revision: https://reviews.llvm.org/D114084
      394d6fcf
    • Arthur Eubanks's avatar
      [gn build] Add missed comma · e1ef1406
      Arthur Eubanks authored
      e1ef1406
    • Arthur Eubanks's avatar
      [NewPM] Add option to prevent rerunning function pipeline on functions in CGSCC adaptor · e3e25b51
      Arthur Eubanks authored
      In a CGSCC pass manager, we may visit the same function multiple times
      due to SCC mutations. In the inliner pipeline, this results in running
      the function simplification pipeline on a function multiple times even
      if it hasn't been changed since the last function simplification
      pipeline run.
      
      We use a newly introduced analysis to keep track of whether or not a
      function has changed since the last time the function simplification
      pipeline has run on it. If we see this analysis available for a function
      in a CGSCCToFunctionPassAdaptor, we skip running the function passes on
      the function. The analysis is queried at the end of the function passes
      so that it's available after the first time the function simplification
      pipeline runs on a function. This is a per-adaptor option so it doesn't
      apply to every adaptor.
      
      The goal of this is to improve compile times. However, currently we
      can't turn this on by default at least for the ...
      e3e25b51
    • Alexey Bataev's avatar
    • Kazu Hirata's avatar
    • Martin Storsjö's avatar
      [OpenMP] Silence build warnings when built with MinGW · 9b2b5498
      Martin Storsjö authored
      There's an attempt to upstream this change in
      https://github.com/intel/ittapi/pull/25 too.
      
      Differential Revision: https://reviews.llvm.org/D114069
      9b2b5498
  2. Nov 17, 2021