1. May 15, 2023
    • Chenle Yu's avatar
      [OpenMP] Implement task record and replay mechanism · 36d4e4c9
      Chenle Yu authored
      This patch implements the "task record and replay" mechanism.  The idea is to be able to store tasks and their dependencies in the runtime so that we do not pay the cost of task creation and dependency resolution for future executions. The objective is to improve fine-grained task performance, both for those from "omp task" and "taskloop".
      
      The entry point of the recording phase is __kmpc_start_record_task, and the end of record is triggered by __kmpc_end_record_task.
      
      Tasks encapsulated between a record start and a record end are saved, meaning that the runtime stores their dependencies and structures, referred to as TDG, in order to replay them in subsequent executions. In these TDG replays, we start the execution by scheduling all root tasks (tasks that do not have input dependencies), and there will be no involvement of a hash table to track the dependencies, yet tasks do not need to be created again.
      
      At the beginning of __kmpc_start_record_task, we must check if a TDG has already been recorded. If yes, the function returns 0 and starts to replay the TDG by calling __kmp_exec_tdg; if not, we start to record, and the function returns 1.
      
      An integer uniquely identifies TDGs. Currently, this identifier needs to be incremented manually in the source code. Still, depending on how this feature would eventually be used in the library, the caller function must do it; also, the caller function needs to implement a mechanism to skip the associated region, according to the return value of __kmpc_start_record_task.
      
      Reviewed By: tianshilei1992
      
      Differential Revision: https://reviews.llvm.org/D146642
      36d4e4c9
    • Nikita Popov's avatar
      [KnownBitsTest] Align with ConstantRange test infrastructure (NFC) · 0f1fb626
      Nikita Popov authored
      Align the way we perform exhaustive tests for KnownBits with what
      we do for ConstantRange. Test each case separately by specifying
      a function on KnownBits and one on APInts. Additionally, specify
      a callback that determines which cases are supposed to be optimal,
      rather than only correct. Unlike the ConstantRange case there is
      a well-defined, unique notion of optimality for KnownBits.
      
      If a failure occurs, print out the inputs, computed result and
      exact result. Adjust the printing function to produce the output
      in a format that is meaningful for KnownBits, i.e. print the
      actual known bits, using ? to signify unknowns and ! to signify
      conflicts.
      0f1fb626
    • Erich Keane's avatar
      Update __cplusplus for C++23, add C++23 diag group alias. · ba9eaf59
      Erich Keane authored
      This came up during the C++26 flag discussion, so split this out into a
      separate patch.
      ba9eaf59
    • Sameer Sahasrabuddhe's avatar
      [LLVM][Uniformity] Propagate temporal divergence explicitly · b0f0dd25
      Sameer Sahasrabuddhe authored
      At a cycle C with divergent exits, UA was using a naive traversal of the exiting
      edges to locate blocks that may use values defined inside C. But this traversal
      fails when it encounters a cycle. This is now replaced with a much simpler
      propagation that iterates over every instruction in C and checks any uses that
      are outside C. But such an iteration can be expensive when C is very large; the
      original strategy may need to be reconsidered if there is a regression in
      compilation times.
      
      Also fixed lit tests that should have originally caught the missed propagation
      of temporal divergence.
      
      Reviewed By: foad
      
      Differential Revision: https://reviews.llvm.org/D149646
      b0f0dd25
    • Manupa Karunaratne's avatar
      [MLIR][ROCDL] add gpu to rocdl erf support · 8c3a8d17
      Manupa Karunaratne authored
      This commit adds lowering of lib func
      call to support erf in rocdl.
      
      Reviewed By: ThomasRaoux
      
      Differential Revision: https://reviews.llvm.org/D150355
      8c3a8d17
    • Alex Zinenko's avatar
      [mlir] allow repeated payload in structured.fuse_into_containing · 1365ff74
      Alex Zinenko authored
      Structured fusion proceeds by iteratively finding the next suitable
      producer to be fused into the loop. Therefore, it shouldn't matter if
      the same producer is listed multiple times (e.g., it is used as multiple
      operands). Adjust the implementation of the transform op to support this
      case.
      
      Also fix the checking code in the interpreter to actually respect the
      TransformOpInterface indication that repeated payload is allowed, it
      seems to have been accidentally dropped in one of the refactorings.
      
      Reviewed By: nicolasvasilache
      
      Differential Revision: https://reviews.llvm.org/D150561
      1365ff74
    • Kyle Huey's avatar
      [X86] Use the CFA as the DWARF frame base for better variable locations around calls. · d421f522
      Kyle Huey authored
      Prior to this patch, for the DWARF frame base LLVM uses the frame pointer
      register if available, otherwise the stack pointer register. If the stack
      pointer register is being used and a call or other code modifies the stack
      pointer during the body of the function this results in the locations being
      wrong and the debugger displaying the wrong values for variables.
      
      By using DW_OP_call_frame_cfa in these situations the emitted location for
      the variable will automatically handle changes in the stack pointer.
      The CFA needs to be adjusted for the offset between the frame pointer/stack
      pointer to allow the variable locations themselves to remain unchanged by
      this patch.
      
      Reviewed By: #debug-info, scott.linder, jryans
      
      Differential Revision: https://reviews.llvm.org/D143463
      d421f522
    • Florian Hahn's avatar
      [AArch64] Add test case where widening mull could be used. · f516ad61
      Florian Hahn authored
      Extra test using mull for D150482.
      f516ad61
    • khei4's avatar
      [ConstantFold] use StoreSize for VectorType folding · f5dbbf49
      khei4 authored
      Differential Revision: https://reviews.llvm.org/D150515
      Reviewed By: nikic
      f5dbbf49
    • Nikolas Klauser's avatar
      Revert "[libc++][PSTL] Implement std::transform" · 627e6efa
      Nikolas Klauser authored
      This reverts commit cbd9e545.
      
      The wrong patch was landed.
      627e6efa
    • Francesco Petrogalli's avatar
      [unittests][llvm-exegesis] Remove build warnings [NFCI] · 930421e1
      Francesco Petrogalli authored
      Remove the warning caused by a missing field initializer.
      
      The field is `StartAtCycle` of `struct MCWriteProcResEntry`.
      
      It has been set to the default `StartAtCycle = 0`.
      
      Differential Revision: https://reviews.llvm.org/D150569
      930421e1
    • Nikolas Klauser's avatar
      [libc++][PSTL] Implement std::transform · cbd9e545
      Nikolas Klauser authored
      Reviewed By: ldionne, #libc
      
      Spies: libcxx-commits
      
      Differential Revision: https://reviews.llvm.org/D149615
      cbd9e545
    • Matthias Springer's avatar
      [mlir][scf][bufferize] Fix bug in WhileOp analysis verification · ae8cb643
      Matthias Springer authored
      Block arguments and yielded values are not equivalent if there are not enough block arguments. This fixes #59442.
      
      Differential Revision: https://reviews.llvm.org/D145575
      ae8cb643
    • Matthias Springer's avatar
      [mlir][bufferization] Add option to dump alias sets · bb9d1b55
      Matthias Springer authored
      This is useful for debugging.
      
      Differential Revision: https://reviews.llvm.org/D143314
      bb9d1b55
    • Jan Kuhle's avatar
      clang-format: [JS] terminate import sorting on `export type X = Y` · e1f34b73
      Jan Kuhle authored
      Contributed by @jankuehle!
      
      https://reviews.llvm.org/D150116 introduced a bug. `export type X = Y` was considered an export declaration and took part in import sorting. This is not correct. With this change `export type X = Y` properly terminates import sorting.
      
      Reviewed By: krasimir
      
      Differential Revision: https://reviews.llvm.org/D150563
      e1f34b73
    • Matthias Springer's avatar
      [mlir][bufferization] Improve findValueInReverseUseDefChain signature · 1f479c1e
      Matthias Springer authored
      Instead of passing traversal options as a long list of arguments, store them in a TraversalConfig object and pass that object.
      
      Differential Revision: https://reviews.llvm.org/D143927
      1f479c1e
    • Jay Foad's avatar
      [AMDGPU] Simplify liveins in some MIR tests · 03b97e0c
      Jay Foad authored
      We can use the following 16-VGPR tuple directly instead of splitting it
      into smaller parts:
      
      $vgpr240_vgpr241_vgpr242_vgpr243_vgpr244_vgpr245_vgpr246_vgpr247_vgpr248_vgpr249_vgpr250_vgpr251_vgpr252_vgpr253_vgpr254_vgpr255
      03b97e0c
    • Manna, Soumi's avatar
      Fix build error caused by https://reviews.llvm.org/D149718 · 292a6c1c
      Manna, Soumi authored
      The patch(https://reviews.llvm.org/D149718) broke buildbot
      
      ../../clang/include/clang/Sema/ParsedAttr.h:705:18: error: explicitly defaulted move assignment operator is implicitly deleted [-Werror,-Wdefaulted-function-deleted]
        AttributePool &operator=(AttributePool &&pool) = default;
                       ^
      ../../clang/include/clang/Sema/ParsedAttr.h:674:21: note: move assignment operator of 'AttributePool' is implicitly deleted because field 'Factory' is of reference type 'clang::AttributeFactory &'
        AttributeFactory &Factory;
                          ^
      1 error generated.
      
      This patch fixes the build error.
      292a6c1c
    • Nikita Popov's avatar
      [Pipelines] Don't skip GlobalDCE in ThinLTO pre-link · 3060ee0c
      Nikita Popov authored
      GlobalDCE will only remove functions with available externally
      linkage if they are unreferenced. As such, I don't believe there
      is any problem with running this pass as part of the ThinLTO pre-link
      pipeline. It will only remove functions that are completely dead in
      that module, and I don't think there is any benefit to keeping them
      around for the post-link phase.
      
      There is no compile-time impact from the additional pass.
      
      This is a followup to one of the side discussions in D146776.
      
      Differential Revision: https://reviews.llvm.org/D149446
      3060ee0c
    • Piotr Sobczak's avatar
      [ValueTracking] Fix computeKnownFPClass with canonicalize · 7322d354
      Piotr Sobczak authored
      Update code that assumes llvm.canonicalize only handles scalars,
      by adding a call to getScalarType().
      This is fine, as the intrinsic is trivially vectorizable.
      
      Introduced in D147870, and uncovered by D148065.
      
      Differential Revision: https://reviews.llvm.org/D150556
      7322d354
    • Matthias Springer's avatar
      [mlir][IR][tests] Fix incorrect API usage in RewritePatterns · e219e66e
      Matthias Springer authored
      Incorrect API usage was detected by D144552.
      
      Differential Revision: https://reviews.llvm.org/D145167
      e219e66e
    • Matthias Springer's avatar
      [mlir][bufferization] Fix unknown ops in BufferViewFlowAnalysis · 38bef476
      Matthias Springer authored
      If an op is unknown to the analysis, it must be treated conservatively: assume that every operand aliases with every result.
      
      Differential Revision: https://reviews.llvm.org/D150546
      38bef476
    • Haojian Wu's avatar
      [clangd] Fix fixAll not shown when there is only one unused-include and... · d0e89116
      Haojian Wu authored
      [clangd] Fix fixAll not shown when there is only one unused-include and missing-include diagnostics.
      
      Discovered during the review in https://reviews.llvm.org/D149437#inline-1444851.
      
      Differential Revision: https://reviews.llvm.org/D149822
      d0e89116
    • Alejandro Álvarez Ayllón's avatar
      [clang][parser] Fix namespace dropping after malformed declarations · b321738f
      Alejandro Álvarez Ayllón authored
      * After a malformed top-level declaration
      * After a malformed templated class method declaration
      
      In both cases, when there is a malformed declaration, any following
      namespace is dropped from the AST. This can trigger a cascade of
      confusing diagnostics that may hide the original error. An example:
      ```
      // Start #include "SomeFile.h"
      template <class T>
      void Foo<T>::Bar(void* aRawPtr) {
          (void)(aRawPtr);
      }
      // End #include "SomeFile.h"
      
      int main() {}
      ```
      We get the original error, plus 19 others from the standard library.
      With this patch, we only get the original error.
      
      clangd can also benefit from this patch, as namespaces following the
      malformed declaration is now preserved. i.e.
      ```
      
      MACRO_FROM_MISSING_INCLUDE("X")
      
      namespace my_namespace {
          //...
      }
      ```
      Before this patch, my_namespace is not visible for auto-completion.
      
      Differential Revision: https://reviews.llvm.org/D150258
      b321738f
    • Joseph Huber's avatar
      [libc] Cache ownership of the shared buffer in the port · 45b899b9
      Joseph Huber authored
      This patch adds another variable to cache cases where we know that we
      own the buffer. This allows us to skip the atomic load on the inbox
      because we already know its state. This is legal immediately after
      opening a port, or when sending immediately after a recieve. This
      caching nets a significant (~17%) speedup for the basic open, send,
      recieve combination.
      
      Reviewed By: JonChesterfield
      
      Differential Revision: https://reviews.llvm.org/D150516
      45b899b9
    • Joseph Huber's avatar
      [libc] Make the bump pointer explicitly return null on buffer oveerrun · 9417d9fc
      Joseph Huber authored
      We use a simple bump ptr in the `libc` tests. If we run out of data we
      can currently return other static memory and have weird failure cases.
      We should fail more explicitly here by returning a null pointer instead.
      
      Reviewed By: sivachandra
      
      Differential Revision: https://reviews.llvm.org/D150529
      9417d9fc
    • Martin Braenne's avatar
      [clang][dataflow] Don't analyze templated declarations. · 1a42f795
      Martin Braenne authored
      Attempting to analyze templated code doesn't have a good cost-benefit ratio. We
      have so far done a best-effort attempt at this, but maintaining this support has
      an ongoing high maintenance cost because the AST for templates can violate a lot
      of the invariants that otherwise hold for the AST of concrete code. As just one
      example, in concrete code the operand of a UnaryOperator '*' is always a prvalue
      (https://godbolt.org/z/s3e5xxMd1), but in templates this isn't true
      (https://godbolt.org/z/6W9xxGvoM).
      
      Further rationale for not analyzing templates:
      
      * The semantics of a template itself are weakly defined; semantics can depend
        strongly on the concrete template arguments. Analyzing the template itself (as
        opposed to an instantiation) therefore has limited value.
      
      * Analyzing templates requires a lot of special-case code that isn't necessary
        for concrete code because dependent types are hard to deal with and the AST
        violates invariants that otherwise hold for concrete code (see above).
      
      * There's precedent in that neither Clang Static Analyzer nor the flow-sensitive
        warnings in Clang (such as uninitialized variables) support analyzing
        templates.
      
      Reviewed By: gribozavr2, xazax.hun
      
      Differential Revision: https://reviews.llvm.org/D150352
      1a42f795
    • Florian Hahn's avatar
      [VPlan] Use VPRecipeWithIRFlags for VPReplicateRecipe, retire poison map · 701f7230
      Florian Hahn authored
      Update VPReplicateRecipe to use VPRecipeWithIRFlags for IR flag
      handling. Retire separate MayGeneratePoisonRecipes map.
      
      Depends on D149082.
      
      Reviewed By: Ayal
      
      Differential Revision: https://reviews.llvm.org/D150027
      701f7230
    • Jay Foad's avatar
    • Ivan Chikish's avatar
      [X86] LowerRotate: prefer unpack-based algorithm · 579812c0
      Ivan Chikish authored
      Splitting and improving from the https://reviews.llvm.org/D146357
      
      When running tests for LowerShift, I discovered some poor codegen in rotate and funnel shift tests. This patch attempts to address some of them.
      
      Using unpack for splitting and using double-bitwidth shifts may improve performance according to https://uica.uops.info tests.
      
          No cross-lane shuffles
          No dirtying double-width registers
          Massive improvement for AVX2 rotates in some cases (var_funnnel_v8i16, var_funnnel_v16i16) — because unpack is currently only used for vXi8 vectors.
      
      Differential Revision: https://reviews.llvm.org/D149071
      579812c0
    • Jacob Crawley's avatar
      [flang][hlfir] lower hlfir.any into fir runtime call · b5d1ea9d
      Jacob Crawley authored
      Depends on: D150272
      
      Differential Revision: https://reviews.llvm.org/D150451
      b5d1ea9d
    • Jacob Crawley's avatar
      [flang] lower any intrinsic to hlfir.any operation · d7b19b0e
      Jacob Crawley authored
      Carries out the lowering of the any intrinsic into HLFIR
      
      Depends on: D149964
      
      Differential Revision: https://reviews.llvm.org/D150272
      d7b19b0e
    • Jacob Crawley's avatar
      [flang] add hlfir.any intrinsic · 622281a7
      Jacob Crawley authored
      Adds a HLFIR operation for the ANY intrinsic according to the
      design set out in flang/docs/HighLevel.md
      
      Differential Revision: https://reviews.llvm.org/D149964
      622281a7
    • Peter Smith's avatar
      [LLD][ELF] Add missing program header parsing to OVERLAY · e16af8a2
      Peter Smith authored
      In D72756 the change to add INPUT_SECTION_FLAGS inadvertantly
      removed the line to parse the program header assignment information for
      OutputSections within an OVERLAY.
      
      This change adds back the missing line and adds a test for it.
      
      Differential Revision: https://reviews.llvm.org/D150445
      e16af8a2
    • Tobias Hieta's avatar
      [docs] Add Python coding standard to documentation · 83768e66
      Tobias Hieta authored
      As discussed on the forums:
      
      https://discourse.llvm.org/t/rfc-document-and-standardize-python-code-style/
      
      Reviewed By: jhenderson, JDevlieghere
      
      Differential Revision: https://reviews.llvm.org/D143852
      83768e66
    • Francesco Petrogalli's avatar
      [TableGen][SubtargetEmitter] Add the StartAtCycles field in the WriteRes class. · 4bfe4108
      Francesco Petrogalli authored
      Conditions that need to be met:
      
      1. count(StartAtCycle) == count(ReservedCycles);
      2. For each i: StartAtCycles[i] < ReservedCycles[i];
      3. For each i: StartAtCycles[i] >= 0;
      4. If left unspecified, the elements are set to 0.
      
      Differential Revision: https://reviews.llvm.org/D150310
      4bfe4108
    • Matthias Springer's avatar
      [mlir][transform] Use TrackingListener-aware iterator for getPayloadOps · 0e37ef08
      Matthias Springer authored
      Instead of returning an `ArrayRef<Operation *>`, return at iterator that skips ops that were erased/replaced while iterating over the payload ops.
      
      This fixes an issue in conjuction with TrackingListener, where a tracked op was erased during the iteration. Elements may not be removed from an array while iterating over it; this invalidates the iterator.
      
      When ops are erased/removed via `replacePayloadOp`, they are not immediately removed from the mappings data structure. Instead, they are set to `nullptr`. `nullptr`s are not enumerated by `getPayloadOps`. At the end of each transformation, `nullptr`s are removed from the mapping data structure.
      
      Differential Revision: https://reviews.llvm.org/D149847
      0e37ef08
    • Guillaume Chatelet's avatar
      [libc] Add optimized memset for RISCV · c5dede88
      Guillaume Chatelet authored
      This patch adds two versions of `memset` optimized for architectures where unaligned accesses are either illegal or extremely slow.
      It is currently enabled for RISCV 64 and RISCV 32 but it could be used for ARM 32 architectures as well.
      
      Here is the before / after output of libc.benchmarks.memory_functions.opt_host --benchmark_filter=BM_Memset on a quad core Linux starfive RISCV 64 board running at 1.5GHz.
      
      Before
      ```
      Run on (4 X 1500 MHz CPU s)
      CPU Caches:
        L1 Instruction 32 KiB (x4)
        L1 Data 32 KiB (x4)
        L2 Unified 2048 KiB (x1)
      ------------------------------------------------------------------------
      Benchmark              Time             CPU   Iterations UserCounters...
      ------------------------------------------------------------------------
      BM_Memset/0/0        506 ns          252 ns      2883584 bytes_per_cycle=0.238312/s bytes_per_second=340.908M/s items_per_second=3.96043M/s __llvm_libc::memset,memset Google A
      BM_Memset/1/0        296 ns          189 ns      2900992 bytes_per_cycle=0.234589/s bytes_per_second=335.583M/s items_per_second=5.29382M/s __llvm_libc::memset,memset Google B
      BM_Memset/2/0       2110 ns         1049 ns       678912 bytes_per_cycle=0.24687/s bytes_per_second=353.151M/s items_per_second=953.527k/s __llvm_libc::memset,memset Google D
      BM_Memset/3/0        397 ns          254 ns      3055616 bytes_per_cycle=0.238479/s bytes_per_second=341.147M/s items_per_second=3.93224M/s __llvm_libc::memset,memset Google L
      BM_Memset/4/0       1119 ns          621 ns      1079296 bytes_per_cycle=0.244925/s bytes_per_second=350.368M/s items_per_second=1.61047M/s __llvm_libc::memset,memset Google M
      BM_Memset/5/0        605 ns          349 ns      1644544 bytes_per_cycle=0.241364/s bytes_per_second=345.274M/s items_per_second=2.8614M/s __llvm_libc::memset,memset Google Q
      BM_Memset/6/0        472 ns          271 ns      2310144 bytes_per_cycle=0.238615/s bytes_per_second=341.341M/s items_per_second=3.68799M/s __llvm_libc::memset,memset Google S
      BM_Memset/7/0        262 ns          143 ns      3956736 bytes_per_cycle=0.225812/s bytes_per_second=323.026M/s items_per_second=7.0087M/s __llvm_libc::memset,memset Google U
      BM_Memset/8/0        454 ns          261 ns      2940928 bytes_per_cycle=0.238883/s bytes_per_second=341.725M/s items_per_second=3.82716M/s __llvm_libc::memset,memset Google W
      BM_Memset/9/0       8768 ns         5998 ns       115712 bytes_per_cycle=0.249196/s bytes_per_second=356.478M/s items_per_second=166.724k/s __llvm_libc::memset,uniform 384 to 4096
      ```
      
      After
      ```
      BM_Memset/0/0        117 ns         69.5 ns      9761792 bytes_per_cycle=0.935152/s bytes_per_second=1.30639G/s items_per_second=14.3834M/s __llvm_libc::memset,memset Google A
      BM_Memset/1/0       97.8 ns         58.5 ns     13002752 bytes_per_cycle=0.892814/s bytes_per_second=1.24725G/s items_per_second=17.0848M/s __llvm_libc::memset,memset Google B
      BM_Memset/2/0        326 ns          163 ns      5156864 bytes_per_cycle=1.54408/s bytes_per_second=2.15706G/s items_per_second=6.1192M/s __llvm_libc::memset,memset Google D
      BM_Memset/3/0        132 ns         65.4 ns     11455488 bytes_per_cycle=0.876411/s bytes_per_second=1.22433G/s items_per_second=15.2803M/s __llvm_libc::memset,memset Google L
      BM_Memset/4/0        222 ns          120 ns      6405120 bytes_per_cycle=1.44398/s bytes_per_second=2.01722G/s items_per_second=8.30758M/s __llvm_libc::memset,memset Google M
      BM_Memset/5/0        119 ns         79.2 ns      8930304 bytes_per_cycle=1.13327/s bytes_per_second=1.58317G/s items_per_second=12.6189M/s __llvm_libc::memset,memset Google Q
      BM_Memset/6/0        123 ns         64.0 ns     11609088 bytes_per_cycle=1.008/s bytes_per_second=1.40817G/s items_per_second=15.6365M/s __llvm_libc::memset,memset Google S
      BM_Memset/7/0       85.9 ns         52.1 ns     12423168 bytes_per_cycle=0.641164/s bytes_per_second=917.192M/s items_per_second=19.1937M/s __llvm_libc::memset,memset Google U
      BM_Memset/8/0        114 ns         67.1 ns     10347520 bytes_per_cycle=0.911968/s bytes_per_second=1.274G/s items_per_second=14.9015M/s __llvm_libc::memset,memset Google W
      BM_Memset/9/0       1326 ns          785 ns       907264 bytes_per_cycle=1.89716/s bytes_per_second=2.6503G/s items_per_second=1.27348M/s __llvm_libc::memset,uniform 384 to 4096
      ```
      
      Again not as good as current glibc but it's a first step in the right direction.
      ```
      BM_Memset/0/0        108 ns         53.6 ns     12894208 bytes_per_cycle=1.02858/s bytes_per_second=1.4369G/s items_per_second=18.668M/s glibc::memset,memset Google A
      BM_Memset/1/0       84.6 ns         47.6 ns     14284800 bytes_per_cycle=1.00197/s bytes_per_second=1.39974G/s items_per_second=21.0256M/s glibc::memset,memset Google B
      BM_Memset/2/0        160 ns         85.8 ns      8927232 bytes_per_cycle=3.30805/s bytes_per_second=4.62129G/s items_per_second=11.6596M/s glibc::memset,memset Google D
      BM_Memset/3/0       78.9 ns         53.6 ns     13326336 bytes_per_cycle=1.14058/s bytes_per_second=1.59338G/s items_per_second=18.674M/s glibc::memset,memset Google L
      BM_Memset/4/0       99.2 ns         60.8 ns     11460608 bytes_per_cycle=2.54751/s bytes_per_second=3.55884G/s items_per_second=16.4587M/s glibc::memset,memset Google M
      BM_Memset/5/0       93.0 ns         56.1 ns     12219392 bytes_per_cycle=1.73379/s bytes_per_second=2.42207G/s items_per_second=17.8157M/s glibc::memset,memset Google Q
      BM_Memset/6/0       89.4 ns         47.2 ns     14692352 bytes_per_cycle=1.34846/s bytes_per_second=1.88377G/s items_per_second=21.1795M/s glibc::memset,memset Google S
      BM_Memset/7/0       84.0 ns         50.0 ns     14468096 bytes_per_cycle=0.911198/s bytes_per_second=1.27293G/s items_per_second=19.994M/s glibc::memset,memset Google U
      BM_Memset/8/0       93.4 ns         52.8 ns     13063168 bytes_per_cycle=1.06642/s bytes_per_second=1.48977G/s items_per_second=18.9524M/s glibc::memset,memset Google W
      BM_Memset/9/0        438 ns          241 ns      2853888 bytes_per_cycle=6.1185/s bytes_per_second=8.54744G/s items_per_second=4.15064M/s glibc::memset,uniform 384 to 4096
      ```
      
      Reviewed By: sivachandra
      
      Differential Revision: https://reviews.llvm.org/D150433
      c5dede88
    • Guillaume Chatelet's avatar
      [NFC] Refactor GlobalVariable Ctor · ce9b89f8
      Guillaume Chatelet authored
      Reuse logic from other ctor and remove code duplication.
      
      Reviewed By: courbet
      
      Differential Revision: https://reviews.llvm.org/D150453
      ce9b89f8
    • Christian Ulmann's avatar
      [IR] Drop const in DILocation::getMergedLocation · 794b58b4
      Christian Ulmann authored
      This commit removes constness from DILocation::getMergedLocation and
      fixes all its users accordingly.
      
      Having constness on the parameters forced the return type to be const
      as well, which does force usage of `const_cast` when the location needs
      to be used in metadata nodes.
      
      Reviewed By: ftynse
      
      Differential Revision: https://reviews.llvm.org/D149942
      794b58b4