1. Dec 08, 2020
    • Arthur Eubanks's avatar
      [test] Fix store_cost.ll under NPM · a820261b
      Arthur Eubanks authored
      The NPM processes loops in forward program order, whereas the legacy PM
      processes them in reverse program order. No reason to test both PMs
      here, so just stick to the NPM.
      a820261b
    • Jonas Devlieghere's avatar
      [lldb] Include thread id in the reproducer trace (NFC) · 33e3b07a
      Jonas Devlieghere authored
      Include the current thread ID in the reproducer trace during
      capture/recording.
      33e3b07a
    • Vitaly Buka's avatar
      [NFC][MSan] Round up OffsetPtr in PoisonMembers · 6e614b0c
      Vitaly Buka authored
      getFieldOffset(layoutStartOffset)  is expected to point to the first trivial
      field or the one which follows non-trivial. So it must be byte aligned already.
      However this is not obvious without assumptions about callers.
      This patch will avoid the need in such assumptions.
      
      Depends on D92727.
      
      Differential Revision: https://reviews.llvm.org/D92728
      6e614b0c
    • Arthur Eubanks's avatar
      [test] Fix widen-iv.ll under NPM · deac8b1f
      Arthur Eubanks authored
      The -loop-flatten legacy pass preserves loop analyses. The legacy PM
      will check all passes that preserve loop analyses that they preserve
      LCSSA. This implicitly involves running -loop-simplify. The test
      shouldn't depend on verify flags being set in order to run
      -loop-simplify, so explicitly add it. The new PM ends up not running it
      otherwise.
      deac8b1f
    • Kai Luo's avatar
      [DAGCombine][PowerPC] Simplify nabs by using legal `smin` operation · 44bd8ea1
      Kai Luo authored
      Convert `0 - abs(x)` to `smin (x, -x)` if `smin` is a legal operation.
      
      Verification: https://alive2.llvm.org/ce/z/vpquFR
      
      Reviewed By: RKSimon
      
      Differential Revision: https://reviews.llvm.org/D92637
      44bd8ea1
    • Esme-Yi's avatar
      [PowerPC] Correct the bit-width definition for some imm operand in td. · 49599cb1
      Esme-Yi authored
      Summary: The imm operands of some instructions are not defined accurately in td.
      This is a small patch to correct these definitions.
      
      Reviewed By: steven.zhang
      
      Differential Revision: https://reviews.llvm.org/D91603
      49599cb1
    • Richard Smith's avatar
      Fix assertion failure due to incorrect dependence bits on a DeclRefExpr · 590e1465
      Richard Smith authored
      that can only be set correctly after instantiating the initializer for a
      variable.
      590e1465
    • Fangrui Song's avatar
      [test] Rewrite split-debug.c · 29295e21
      Fangrui Song authored
      Use generic ELF target triples.
      Add missing coverage: -gsplit-dwarf=split -g -fsplit-dwarf-inlining
      Reorganize and add comments.
      Test -gno-pubnames
      29295e21
    • Arthur Eubanks's avatar
      [test] Fix LoopFusion tests under NewPM · 689b8e91
      Arthur Eubanks authored
      The legacy pass depended on -loop-simplify running. The NPM does not
      allow for a non-analysis pass to depend on another non-analysis pass.
      689b8e91
    • Jessica Paquette's avatar
      [AArch64][GlobalISel] Refactor G_BRCOND selection · d49f6491
      Jessica Paquette authored
      `selectCompareBranch` was hard to understand.
      
      Also, it was being needlessly pessimistic with the `ProduceNonFlagSettingCondBr`
      case. It assumed that everything in `selectCompareBranch` would emit a TB(N)Z
      or C(B)NZ. That's not true; the G_FCMP + G_BRCOND case would never emit those
      instructions, and the G_ICMP + G_BRCOND case was capable of emitting an integer
      compare + Bcc.
      
      - Refactor `selectCompareBranch` into separate functions based off of what is
      feeding the G_BRCOND's condition.
      
      - Move G_BRCOND selection code from `select` to `selectCompareBranch`.
      
      - Remove duplicated constraint code from the code originally in `select`;
        `emitTestBit` already handles that, so no need to constrain twice.
      
      - Factor out the G_FCMP + G_BRCOND case into `selectCompareBranchFedByFCmp`.
      
      - Split the G_ICMP + G_BRCOND case into an optimization function,
      `tryOptCompareBranchFedByICmp` and a general selection function,
      `selectCompareBranchFedByICmp`.
      
      - Reduce the number of things passed to `tryOptAndIntoCompareBranch`.
      
      - Improve documentation.
      
      - Give some variables more descriptive names.
      
      Other than improving the code generation for functions with
      speculative_load_hardening by getting the logic correct, this is NFC.
      
      Differential Revision: https://reviews.llvm.org/D92582
      d49f6491
    • Valentin Churavy's avatar
      [VNCoercion] Disallow coercion between different ni addrspaces · 700cf7dc
      Valentin Churavy authored
      
      
      I'm not sure if it would be legal by the IR reference to introduce
      an addrspacecast here, since the IR reference is a bit vague on
      the exact semantics, but at least for our usage of it (and I
      suspect for many other's usage) it is not. For us, addrspacecasts
      between non-integral address spaces carry frontend information that the
      optimizer cannot deduce afterwards in a generic way (though we
      have frontend specific passes in our pipline that do propagate
      these). In any case, I'm sure nobody is using it this way at
      the moment, since it would have introduced inttoptrs, which
      are definitely illegal.
      
      Fixes PR38375
      
      Co-authored-by: default avatarKeno Fischer <keno@alumni.harvard.edu>
      
      Reviewed By: reames
      
      Differential Revision: https://reviews.llvm.org/D50010
      700cf7dc
    • Yaxun (Sam) Liu's avatar
      Fix lit test failure due to 0b81d9 · efc063b6
      Yaxun (Sam) Liu authored
      These lit tests now requires amdgpu-registered-target since they
      use clang driver and clang driver passes an LLVM option which
      is available only if amdgpu target is registered.
      
      Change-Id: I2df31967409f1627fc6d342d1ab5cc8aa17c9c0c
      efc063b6
    • Douglas Yung's avatar
      Fixup test in path to use C:\ instead of D:\ which may be mapped to a removable. · ccc5160d
      Douglas Yung authored
      Our internal build bot hit a failure in llvm/test/tools/llvm-symbolizer/pdb/missing_pdb.test
      because the test was checking for an error message that is emitted when a pdb file is
      missing. But when the drive is mapped to a removalable drive (such as a DVD drive) in
      Windows, you get a different error message which causes the test to fail.
      
      This fixes the test by changing the drive the missing pdb is expected to be on to C:\
      instead of D:\ as that is the drive historically used to install Windows and thus
      if present should be a hard drive.
      
      Reviewed By: rnk
      
      Differential Revision: https://reviews.llvm.org/D92787
      ccc5160d
    • Richard Smith's avatar
      Fix deserialization cycle in preferred_name attribute. · a64c26a4
      Richard Smith authored
      This is really just a workaround for a more fundamental issue in the way
      we deserialize attributes. See PR48434 for details.
      
      Also fix tablegen code generator to produce more correct indentation to
      resolve buildbot issues with -Werror=misleading-indentation firing
      inside the generated code.
      a64c26a4
    • Yaxun (Sam) Liu's avatar
      [AMDGPU] add -mcode-object-version=n · 0b81d9a9
      Yaxun (Sam) Liu authored
      Add option -mcode-object-version=n to control code object version for
      AMDGPU.
      
      Differential Revision: https://reviews.llvm.org/D91310
      0b81d9a9
    • Yaxun (Sam) Liu's avatar
      [clang][AMDGPU] remove mxnack and msramecc options · 5cae7080
      Yaxun (Sam) Liu authored
      Remove mxnack and msramecc options since they
      are deprecated by --offload-arch.
      
      This is part of https://reviews.llvm.org/D60620
      5cae7080
    • Yaxun (Sam) Liu's avatar
      [HIP] fix bundle entry ID for -- · 4bed1d9b
      Yaxun (Sam) Liu authored
      Canonicalize triple used in fat binary. Change from
      amdgcn-amd-amdhsa to amdgcn-amd-amdhsa-.
      
      This is part of https://reviews.llvm.org/D60620
      4bed1d9b
    • Mehdi Amini's avatar
      Add Python binding for MLIR Type Attribute · e56f398d
      Mehdi Amini authored
      Differential Revision: https://reviews.llvm.org/D92711
      e56f398d
    • Mehdi Amini's avatar
      Customize exception thrown from mlir.Operation.create() python bindings · e15ae454
      Mehdi Amini authored
      The default exception handling isn't very user friendly and does not
      point accurately to the issue. Instead we can indicate which of the
      operands isn't valid and provide contextual information in the error
      message.
      
      Differential Revision: https://reviews.llvm.org/D92710
      e15ae454
    • Yaxun (Sam) Liu's avatar
      [clang][AMDGPU] rename sram-ecc as sramecc · 40ad476a
      Yaxun (Sam) Liu authored
      As backend renamed sram-ecc to sramecc, this patch makes
      corresponding change in clang.
      
      Differential Revision: https://reviews.llvm.org/D86217
      40ad476a
    • Jessica Paquette's avatar
      [AArch64][GlobalISel] Narrow 128-bit regs to 64-bit regs in emitTestBit · 195a7af0
      Jessica Paquette authored
      When we have a 128-bit register, emitTestBit would incorrectly narrow to 32
      bits always. If the bit number was > 32, then we would need a TB(N)ZX. This
      would cause a crash, as we'd have the wrong register class. (PR48379)
      
      This generalizes `narrowExtReg` into `moveScalarRegClass`.
      
      This also allows us to remove `widenGPRBankRegIfNeeded` entirely, since
      `selectCopy` correctly handles SUBREG_TO_REG etc.
      
      This does create some codegen changes (since `selectCopy` uses the `all`
      regclass variants). However, I think that these will likely be optimized away,
      and we can always improve the `selectCopy` code. It looks like we should
      revisit `selectCopy` at this point, and possibly refactor it into at least one
      `emit` function.
      
      Differential Revision: https://reviews.llvm.org/D92707
      195a7af0
    • Philip Reames's avatar
      Teach isKnownNonEqual how to recurse through invertible multiplies · 26568853
      Philip Reames authored
      Build on the work started in 8f076291, and add the multiply case. In the process, more clearly describe the requirement for the operation we're looking through.
      
      Differential Revision: https://reviews.llvm.org/D92726
      26568853
    • Jann Horn's avatar
      [clang] Fix noderef for AddrOf on MemberExpr · 6dad7ec5
      Jann Horn authored
      Committing on behalf of thejh (Jann Horn).
      
      As part of this change, one existing test case has to be adjusted
      because it accidentally stripped the NoDeref attribute without
      getting caught.
      
      Depends on D92140
      
      Differential Review: https://reviews.llvm.org/D92141
      6dad7ec5
    • peter klausler's avatar
      [flang] Improve initializer semantics, esp. for component default values · 641ede93
      peter klausler authored
      This patch plugs many holes in static initializer semantics, improves error
      messages for default initial values and other component properties in
      parameterized derived type instantiations, and cleans up several small
      issues noticed during development.  We now do proper scalar expansion,
      folding, and type, rank, and shape conformance checking for component
      default initializers in derived types and PDT instantiations.
      The initial values of named constants are now guaranteed to have been folded
      when installed in the symbol table, and are no longer folded or
      scalar-expanded at each use in expression folding.  Semantics documentation
      was extended with information about the various kinds of initializations
      in Fortran and when each of them are processed in the compiler.
      
      Some necessary concomitant changes have bulked this patch out a bit:
      * contextual messages attachments, which are now produced for parameterized
        derived type instantiations so that the user can figure out which
        instance caused a problem with a component, have been added as part
        of ContextualMessages, and their implementation was debugged
      * several APIs in evaluate::characteristics was changed so that a FoldingContext
        is passed as an argument rather than just its intrinsic procedure table;
        this affected client call sites in many files
      * new tools in Evaluate/check-expression.cpp to determine when an Expr
        actually is a single constant value and to validate a non-pointer
        variable initializer or object component default value
      * shape conformance checking has additional arguments that control
        whether scalar expansion is allowed
      * several now-unused functions and data members noticed and removed
      * several crashes and bogus errors exposed by testing this new code
        were fixed
      * a -fdebug-stack-trace option to enable LLVM's stack tracing on
        a crash, which might be useful in the future
      
      TL;DR: Initialization processing does more and takes place at the right
      times for all of the various kinds of things that can be initialized.
      
      Differential Review: https://reviews.llvm.org/D92783
      641ede93
    • Leonard Chan's avatar
      [clang] Fix noderef for array member of deref expr · 155fca3c
      Leonard Chan authored
          Committing on behalf of thejh (Jann Horn).
      
          Given an attribute((noderef)) pointer "p" to the struct
      
          struct s { int a[2]; };
          ensure that the following expressions are treated the same way by the
          noderef logic:
      
          p->a
          (*p).a
          Until now, the first expression would be treated correctly (nothing is
          added to PossibleDerefs because CheckMemberAccessOfNoDeref() bails out
          on array members), but the second expression would incorrectly warn
          because "*p" creates a PossibleDerefs entry.
      
          Handle this case the same way as for the AddrOf operator.
      
          Differential Revision: https://reviews.llvm.org/D92140
      155fca3c
    • Mitch Phillips's avatar
      Revert "[test] Fix asan/TestCases/Linux/globals-gc-sections-lld.cpp with... · 1d03a54d
      Mitch Phillips authored
      Revert "[test] Fix asan/TestCases/Linux/globals-gc-sections-lld.cpp with -fsanitize-address-globals-dead-stripping"
      
      This reverts commit 14080876.
      
      Reason: Broke the upstream bots - discussed offline.
      1d03a54d
    • Erik Pilkington's avatar
      [clang] Add support for attribute 'swift_async' · 5a28e1d9
      Erik Pilkington authored
      This attributes specifies how (or if) a given function or method will be
      imported into a swift async method. rdar://70111252
      
      Differential revision: https://reviews.llvm.org/D92742
      5a28e1d9
    • Erik Pilkington's avatar
      [clang] Add a new nullability annotation for swift async: _Nullable_result · 9cd2413f
      Erik Pilkington authored
      _Nullable_result generally like _Nullable, except when being imported into a
      swift async method. rdar://70106409
      
      Differential revision: https://reviews.llvm.org/D92495
      9cd2413f
    • Mehdi Amini's avatar
      Set the target branch for `arc land` to main · 234d88ab
      Mehdi Amini authored
      234d88ab
    • wlei's avatar
      [CSSPGO][llvm-profgen] Context-sensitive profile data generation · 1f05b1a9
      wlei authored
      This stack of changes introduces `llvm-profgen` utility which generates a profile data file from given perf script data files for sample-based PGO. It’s part of(not only) the CSSPGO work. Specifically to support context-sensitive with/without pseudo probe profile, it implements a series of functionalities including perf trace parsing, instruction symbolization, LBR stack/call frame stack unwinding, pseudo probe decoding, etc. Also high throughput is achieved by multiple levels of sample aggregation and compatible format with one stop is generated at the end. Please refer to: https://groups.google.com/g/llvm-dev/c/1p1rdYbL93s for the CSSPGO RFC.
      
      This change supports context-sensitive profile data generation into llvm-profgen. With simultaneous sampling for LBR and call stack, we can identify leaf of LBR sample with calling context from stack sample . During the process of deriving fall through path from LBR entries, we unwind LBR by replaying all the calls and returns (including implicit calls/returns due to inlining) backwards on top of the sampled call stack. Then the state of call stack as we unwind through LBR always represents the calling context of current fall through path.
      
      we have two types of virtual unwinding 1) LBR unwinding and 2) linear range unwinding.
      Specifically, for each LBR entry which can be classified into call, return, regular branch, LBR unwinding will replay the operation by pushing, popping or switching leaf frame towards the call stack and since the initial call stack is most recently sampled, the replay should be in anti-execution order, i.e. for the regular case, pop the call stack when LBR is call, push frame on call stack when LBR is return. After each LBR processed, it also needs to align with the next LBR by going through instructions from previous LBR's target to current LBR's source, which we named linear unwinding. As instruction from linear range can come from different function by inlining, linear unwinding will do the range splitting and record counters through the range with same inline context.
      
      With each fall through path from LBR unwinding, we aggregate each sample into counters by the calling context and eventually generate full context sensitive profile (without relying on inlining) to driver compiler's PGO/FDO.
      
      A breakdown of noteworthy changes:
      - Added `HybridSample` class as the abstraction perf sample including LBR stack and call stack
      * Extended `PerfReader` to implement auto-detect whether input perf script output contains CS profile, then do the parsing. Multiple `HybridSample` are extracted
      * Speed up by aggregating  `HybridSample` into `AggregatedSamples`
      * Added VirtualUnwinder that consumes aggregated  `HybridSample` and implements unwinding of calls, returns, and linear path that contains implicit call/return from inlining. Ranges and branches counters are aggregated by the calling context.
 Here calling context is string type, each context is a pair of function name and callsite location info, the whole context is like `main:1 @ foo:2 @ bar`.
      * Added PorfileGenerater that accumulates counters by ranges unfolding or branch target mapping, then generates context-sensitive function profile including function body, inferring callee's head sample, callsite target samples, eventually records into ProfileMap.

      * Leveraged LLVM build-in(`SampleProfWriter`) writer to support different serialization format with no stop
      - `getCanonicalFnName` for callee name and name from ELF section
      - Added regression test for both unwinding and profile generation
      
      Test Plan:
      ninja & ninja check-llvm
      
      Reviewed By: hoy, wenlei, wmi
      
      Differential Revision: https://reviews.llvm.org/D89723
      1f05b1a9
    • Vitaly Buka's avatar
      [CodeGen][MSan] Don't use offsets of zero-sized fields · 3e1cb0db
      Vitaly Buka authored
      Such fields will likely have offset zero making
      __sanitizer_dtor_callback poisoning wrong regions.
      E.g. it can poison base class member from derived class constructor.
      
      Differential Revision: https://reviews.llvm.org/D92727
      3e1cb0db
    • Alex Zinenko's avatar
      [OpenMPIRBuilder] introduce createStaticWorkshareLoop · c102c783
      Alex Zinenko authored
      Introduce a function that creates a statically-scheduled workshare loop
      out of a canonical loop created earlier by the OpenMPIRBuilder. This
      basically amounts to injecting runtime calls to the preheader and the
      after block and updating the trip count. Static scheduling kind is
      currently hardcoded and needs to be extracted from the runtime library
      into common TableGen definitions.
      
      Differential Revision: https://reviews.llvm.org/D92476
      c102c783
    • Michael Kruse's avatar
      [Polly][CodeGen] Remove use of ScalarEvolution. · 6249bfee
      Michael Kruse authored
      ScalarEvolution::getSCEV cannot be used during codegen. ScalarEvolution
      assumes a stable IR and control flow which is under construction during
      Polly's CodeGen. In particular, it uses DominatorTree for compute the
      backedge taken count. However the DominatorTree is not updated during
      codegen.
      
      In this case, SCEV was used to determine the base pointer of an array
      access. Replace it by our own function. Polly generates only GEP and
      BitCasts for array acceses, i.e. it is sufficient to handle these to to
      find the base pointer.
      
      Fixes llvm.org/PR48422
      6249bfee
    • Amy Huang's avatar
      [CodeView] Fix inline sites that are missing code offsets. · 399bc48e
      Amy Huang authored
      When an inline site has a starting code offset of 0, we sometimes
      don't emit the starting offset.
      
      Bug: https://bugs.llvm.org/show_bug.cgi?id=48377
      
      Differential Revision: https://reviews.llvm.org/D92590
      399bc48e
    • Nico Weber's avatar
      docs: Add pointer to cmake caches for PGO · b570f82f
      Nico Weber authored
      Also add a link to end-user PGO documentation.
      
      Differential Revision: https://reviews.llvm.org/D92768
      b570f82f
    • Richard Smith's avatar
      Add new 'preferred_name' attribute. · 98f76adf
      Richard Smith authored
      This attribute permits a typedef to be associated with a class template
      specialization as a preferred way of naming that class template
      specialization. This permits us to specify that (for example) the
      preferred way to express 'std::basic_string<char>' is as 'std::string'.
      
      The attribute is applied to the various class templates in libc++ that have
      corresponding well-known typedef names.
      
      Differential Revision: https://reviews.llvm.org/D91311
      98f76adf
    • Amara Emerson's avatar
    • Nathan James's avatar
      [llvm][NFC] Made RefCountBase constructors protected · a61d5084
      Nathan James authored
      Matches ThreadSafeRefCountBase and forces the class to be inherited.
      a61d5084
    • Nathan James's avatar
      [llvm] Add asserts in (ThreadSafe)?RefCountedBase destructors · dc361d5c
      Nathan James authored
      Added a trivial destructor in release mode and in debug mode a destructor that asserts RefCount is indeed zero.
      This ensure people aren't manually (maybe accidentally) destroying these objects like in this contrived example.
      ```lang=c++
      {
        std::unique_ptr<SomethingRefCounted> Object;
        holdIntrusiveOwnership(Object.get());
        // Object Destructor called here will assert.
      }
      ```
      
      Reviewed By: dblaikie
      
      Differential Revision: https://reviews.llvm.org/D92480
      dc361d5c
    • Derek Schuff's avatar
      [WebAssembly] Add Object and ObjectWriter support for wasm COMDAT sections · 0a391060
      Derek Schuff authored
      Allow sections to be placed into COMDAT groups, in addtion to functions and data
      segments.
      
      Also make section symbols unnamed, which allows sections with identical names
      (section names are independent of their section symbols, but previously we
      gave the symbols the same name as their sections, which results in collisions
      when sections are identically-named).
      
      Differential Revision: https://reviews.llvm.org/D92691
      0a391060