1. Jun 13, 2021
    • Kristina Bessonova's avatar
      [ARM][NEON] Combine base address updates for vld1Ndup intrinsics · f6b9836b
      Kristina Bessonova authored
      Reviewed By: dmgreen
      
      Differential Revision: https://reviews.llvm.org/D103836
      f6b9836b
    • Luo, Yuanke's avatar
      [X86] Check immediate before get it. · 5be314f7
      Luo, Yuanke authored
      For CMP imm instruction, when the operand 1 is symbol address we should
      check if it is immediate first. Here is the example code.
      `CMP64mi32 $noreg, 8, killed renamable $rcx, @d, $noreg, @a, implicit-def
      $eflags`
      Many thanks to Craig, Topper for the test case to reproduce this issue.
      
      Differential Revision: https://reviews.llvm.org/D104037
      5be314f7
    • Luo, Yuanke's avatar
      Revert "[X86] Check immediate before get it." · 1e72b9d5
      Luo, Yuanke authored
      This reverts commit 9eb2f723.
      1e72b9d5
    • Shoaib Meenai's avatar
      [runtimes] Fix umbrella component targets · aa93603f
      Shoaib Meenai authored
      When we're building the runtimes for multiple platform targets, we
      create umbrella build targets for each distribution component, but those
      targets didn't have any dependencies and were just no-ops. Make the
      umbrella target depend on the sub-targets for each platform to fix this,
      which is consistent with the behavior of the umbrella targets for each
      runtime, and also consistent with the behavior when we've only specified
      the default target.
      aa93603f
    • David Blaikie's avatar
      llvm-objcopy: fix section size truncation/extension when dumping sections · 02c71830
      David Blaikie authored
      Since this only comes up with inputs containing sections at least 4GB
      large (I guess I could use a bzero section or something, so the input
      file doesn't have to be 4GB, but even then the output file would have to
      be 4GB, right?) I've skipped testing this. If there's a nice way to test
      this without needing 4GB inputs or output files.
      
      The subtlety here is demonstrated by this code:
      
      struct t { operator uint64_t(); };
      static_assert(std::is_same_v<int, decltype(std::declval<bool>() ? 0 : std::declval<t>())>);
      static_assert(std::is_same_v<uint64_t, decltype(std::declval<bool>() ? 0 : std::declval<uint64_t>())>);
      
      Because of this difference, the original source code was getting an int
      type (truncating the actual size) and then extending it again, resulting
      in bogus values (I haven't thought through this hard enough to explain
      why the resulting value was 0xffff... - sign extension, possible UB, but
      in any case it's the wrong answer - in this particular case I was
      looking at that resulted in a size so large that we couldn't open a file
      large enough to write to and ended up with a rather vague:
      
      error: 'file_name.o': Invalid argument
      02c71830
    • Luo, Yuanke's avatar
      [X86] Check immediate before get it. · 9eb2f723
      Luo, Yuanke authored
      For CMP imm instruction, when the operand 1 is symbol address we should
      check if it is immediate first. Here is the example code.
      `CMP64mi32 $noreg, 8, killed renamable $rcx, @d, $noreg, @a, implicit-def
      $eflags`
      Many thanks to Craig, Topper for the test case to reproduce this issue.
      
      Differential Revision: https://reviews.llvm.org/D104037
      9eb2f723
    • Lang Hames's avatar
      [ORC-RT] Split Simple-Packed-Serialization code into its own header. · 49f4a58d
      Lang Hames authored
      This will simplify integration of this code into LLVM -- The
      Simple-Packed-Serialization code can be copied near-verbatim, but
      WrapperFunctionResult will require more adaptation.
      49f4a58d
    • Mehdi Amini's avatar
      Simplify getArgAttrDict/getResultAttrDict by removing unnecessary checks · 152c9871
      Mehdi Amini authored
      There is a slight change in behavior: if the arg dictionnary is empty
      then we return this empty dictionnary instead of a null attribute.
      This is more consistent with accessing it through:
      
        ArrayAttr args_attr = func_op.getAllArgAttrs();
        args_attr[num].cast<DictionnaryAttr>() ...
      
      Differential Revision: https://reviews.llvm.org/D104189
      152c9871
    • Roman Lebedev's avatar
    • Mehdi Amini's avatar
      Use dyn_cast_or_null instead of dyn_cast in FunctionLike::verifyTrait (NFC) · 8bc1ce0f
      Mehdi Amini authored
      This is making the verifier more tolerant to cases where a "null"
      Attribute would be inserted in the array of func arguments/results
      attributes.
      8bc1ce0f
    • Ian McIntyre's avatar
      [llvm-objcopy] Exclude empty sections in IHexWriter output · 58992787
      Ian McIntyre authored
      IHexWriter was evaluating a section's physical address when deciding if
      that section should be written to an output. This approach does not
      account for a zero-sized section that has the same physical address as a
      sized section. The behavior varies from GNU objcopy, and may result in a
      HEX file that does not include all program sections.
      
      The IHexWriter now excludes zero-sized sections when deciding what
      should be written to the output. This affects the contents of the
      writer's `Sections` collection; we will not try to insert multiple
      sections that could have the same physical address. The behavior seems
      consistent with GNU objcopy, which always excludes empty sections,
      no matter the address.
      
      The new test case evaluates the IHexWriter behavior when provided a
      variety of empty sections that overlap or append a filled section. See
      the input file's comments for more information. Given that test input,
      and the change to the IHexWriter, GNU objcopy and llvm-objcopy produce
      the same output.
      
      Reviewed By: jhenderson, MaskRay, evgeny777
      
      Differential Revision: https://reviews.llvm.org/D101332
      58992787
    • Xun Li's avatar
      [CHR] Don't run ControlHeightReduction if any BB has address taken · fae7deba
      Xun Li authored
      This patch is to address https://bugs.llvm.org/show_bug.cgi?id=50610.
      In computed goto pattern, there are usually a list of basic blocks that are all targets of indirectbr instruction, and each basic block also has address taken and stored in a variable.
      CHR pass could potentially clone these basic blocks, which would generate a cloned version of the indirectbr and clonved version of all basic blocks in the list.
      However these basic blocks will not have their addresses taken and stored anywhere. So latter SimplifyCFG pass will simply remove all tehse cloned basic blocks, resulting in incorrect code.
      To fix this, when searching for scopes, we skip scopes that contains BBs with addresses taken.
      Added a few test cases.
      
      Reviewed By: aeubanks, wenlei, hoy
      
      Differential Revision: https://reviews.llvm.org/D103867
      fae7deba
    • Craig Topper's avatar
      [X86] Add ISD::FREEZE and ISD::AssertAlign to the list of opcodes that don't... · c997867d
      Craig Topper authored
      [X86] Add ISD::FREEZE and ISD::AssertAlign to the list of opcodes that don't guarantee upper 32 bits are zero.
      
      The freeze issue was reported here
      https://llvm.discourse.group/t/bug-or-feature-freeze-instruction/3639
      
      I don't have a test for AssertAlign. I just noticed it was missing
      and assume it should be similar to the other two Asserts.
      
      Reviewed By: RKSimon
      
      Differential Revision: https://reviews.llvm.org/D104178
      c997867d
    • Saleem Abdulrasool's avatar
      Revert "Revert "DirectoryWatcher: add an implementation for Windows"" · 76f1baa7
      Saleem Abdulrasool authored
      This reverts commit 0ec1cf13.
      
      Restore the implementation with some minor tweaks:
      - Use std::unique_ptr for the path instead of std::vector
        * Stylistic improvement as the buffer is already heap allocated, this
          just makes it clearer.
      - Correct the notification buffer allocation size
        * Memory usage fix: we were allocating 4x the computed size
      - Correct the passing of the buffer size to RDC
        * Memory usage fix: we were reporting 1/4th of the size
      - Convert the operation event to auto-reset
        * Bug Fix: we never reset the event
      - Remove `FILE_NOTIFY_CHANGE_LAST_ACCESS` from RDC events
        * Memory usage fix: we never needed this notification
      - Fold events for the notification action
        * Stylistic improvement to be clear how the events map
      - Update comment
        * Stylistic improvement to be clear what the RAII controls
      - Fix the race condition that was uncovered previously
        * We would return from the construction before the watcher thread
          began execution.  The test would then proceed to begin execution,
          and we would miss the initial notifications.  We now ensure that the
          watcher thread is initialized before we return.  This ensures that
          we do not miss the initial notifications.
      
      Running the test on a SSD was able to uncover the access pattern.  This
      now seems to pass reliably where it was previously flaky locally.
      76f1baa7
  2. Jun 12, 2021
    • Matheus Izvekov's avatar
      [clang] NRVO: Improvements and handling of more cases. · 1e50c3d7
      Matheus Izvekov authored
      
      
      This expands NRVO propagation for more cases:
      
      Parse analysis improvement:
      * Lambdas and Blocks with dependent return type can have their variables
        marked as NRVO Candidates.
      
      Variable instantiation improvements:
      * Fixes crash when instantiating NRVO variables in Blocks.
      * Functions, Lambdas, and Blocks which have auto return type have their
        variables' NRVO status propagated. For Blocks with non-auto return type,
        as a limitation, this propagation does not consider the actual return
        type.
      
      This also implements exclusion of VarDecls which are references to
      dependent types.
      
      Signed-off-by: default avatarMatheus Izvekov <mizvekov@gmail.com>
      
      Reviewed By: Quuxplusone
      
      Differential Revision: https://reviews.llvm.org/D99696
      1e50c3d7
    • Florian Hahn's avatar
    • Shashij gupta's avatar
      [MLIR] Simplify affine.if ops with trivial conditions · 466e5aba
      Shashij gupta authored
      
      
      The commit simplifies affine.if ops :
      The affine if operation gets removed if the condition is universally true or false and then/else block is merged with the parent block.
      
      Signed-off-by: default avatarShashij Gupta <shashij.gupta@polymagelabs.com>
      
      Reviewed By: bondhugula, pr4tgpt
      
      Differential Revision: https://reviews.llvm.org/D104015
      466e5aba
    • Florian Hahn's avatar
      b4583a5a
    • Kristina Bessonova's avatar
      [lit] Attempt for fix tests failing because of 'warning: non-portable path to file' · 8e627979
      Kristina Bessonova authored
      This is an attempt to fix clang test failures due to 'nonportable-include-path'
      warnings on Windows when a path to llvm-project's base directory contains some
      uppercase letters (excluding a drive letter).
      
      The issue originates from 2 problems:
      * discovery.py loads site config in lower case causing all the paths
      based on __file__ and requested within the config file to be in lowercase as well,
      * neither os.path.abspath() nor os.path.realpath() (both used to obtain paths of
      config files, sources, object directories, etc) do not return paths in the correct
      case for Windows (at least consistently for all python versions).
      
      As os.path library doesn't seem to provide any relaible way to restore
      the case for paths on Windows, this patch proposes to use pathlib.resolve().
      pathlib is a part of Python 3.4 while llvm lit requires Python 3.6.
      
      Reviewed By: Meinersbur
      
      Differential Revision: https://reviews.llvm.org/D103014
      8e627979
    • Florian Hahn's avatar
      Revert "[X86FixupLEAs] Transform the sequence LEA/SUB to SUB/SUB" · 5cd66420
      Florian Hahn authored
      This reverts commit 1b748faf because it
      breaks building the llvm-test-suite with -verify-machineinstrs on X86:
      http://green.lab.llvm.org/green/job/test-suite-verify-machineinstrs-x86_64-O3/9585/
      
      Running llc -verify-machineinstr on X86 crashes on the IR below:
      
          target datalayout = "e-m:o-p270:32:32-p271:32:32-p272:64:64-i64:64-f80:128-n8:16:32:64-S128"
      
          %struct.widget = type { i32, i32, i32, i32, i32*, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, [16 x [16 x i16]], [6 x [32 x i32]], [16 x [16 x i32]], [4 x [12 x [4 x [4 x i32]]]], [16 x i32], i8**, i32*, i32***, i32**, i32, i32, i32, i32, %struct.baz*, %struct.wobble.1*, i32, i32, i32, i32, i32, i32, %struct.quux.2*, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, [3 x i32], i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32***, i32***, i32****, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, [3 x [2 x i32]], [3 x [2 x i32]], i32, i32, i64, i64, %struct.zot.3, %struct.zot.3, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32 }
          %struct.baz = type { i32, i32, i32, i32, i32, i32, i32, i32, i32, %struct.snork*, %struct.wombat.0*, %struct.wobble*, i32, i32*, i32*, i32*, i32, i32*, i32*, i32*, i32 (%struct.widget*, %struct.eggs*)*, i32, i32, i32, i32 }
          %struct.snork = type { %struct.spam*, %struct.zot, i32 (%struct.wombat*, %struct.widget*, %struct.snork*)* }
          %struct.spam = type { i32, i32, i32, i32, i8*, i32 }
          %struct.zot = type { i32, i32, i32, i32, i32, i8*, i32* }
          %struct.wombat = type { i32, i32, i32, i32, i32, i32, i32, i32, void (i32, i32, i32*, i32*)*, void (%struct.wombat*, %struct.widget*, %struct.zot*)* }
          %struct.wombat.0 = type { [4 x [11 x %struct.quux]], [2 x [9 x %struct.quux]], [2 x [10 x %struct.quux]], [2 x [6 x %struct.quux]], [4 x %struct.quux], [4 x %struct.quux], [3 x %struct.quux] }
          %struct.quux = type { i16, i8 }
          %struct.wobble = type { [2 x %struct.quux], [4 x %struct.quux], [3 x [4 x %struct.quux]], [10 x [4 x %struct.quux]], [10 x [15 x %struct.quux]], [10 x [15 x %struct.quux]], [10 x [5 x %struct.quux]], [10 x [5 x %struct.quux]], [10 x [15 x %struct.quux]], [10 x [15 x %struct.quux]] }
          %struct.eggs = type { [1000 x i8], [1000 x i8], [1000 x i8], i32, i32, i32, i32, i32, i32, i32, i32 }
          %struct.wobble.1 = type { i32, [2 x i32], i32, i32, %struct.wobble.1*, %struct.wobble.1*, i32, [2 x [4 x [4 x [2 x i32]]]], i32, i64, i64, i32, i32, [4 x i8], [4 x i8], i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32, i32 }
          %struct.quux.2 = type { i32, i32, i32, i32, i32, %struct.quux.2* }
          %struct.zot.3 = type { i64, i16, i16, i16 }
      
          define void @blam(%struct.widget* %arg, i32 %arg1) local_unnamed_addr {
          bb:
            %tmp = load i32, i32* undef, align 4
            %tmp2 = sdiv i32 %tmp, 6
            %tmp3 = sdiv i32 undef, 6
            %tmp4 = load i32, i32* undef, align 4
            %tmp5 = icmp eq i32 %tmp4, 4
            %tmp6 = select i1 %tmp5, i32 %tmp3, i32 %tmp2
            %tmp7 = getelementptr inbounds [4 x [4 x i32]], [4 x [4 x i32]]* undef, i64 0, i64 0, i64 0
            %tmp8 = zext i16 undef to i32
            %tmp9 = zext i16 undef to i32
            %tmp10 = load i16, i16* undef, align 2
            %tmp11 = zext i16 %tmp10 to i32
            %tmp12 = zext i16 undef to i32
            %tmp13 = zext i16 undef to i32
            %tmp14 = zext i16 undef to i32
            %tmp15 = load i16, i16* undef, align 2
            %tmp16 = zext i16 %tmp15 to i32
            %tmp17 = zext i16 undef to i32
            %tmp18 = sub nsw i32 %tmp8, %tmp9
            %tmp19 = shl nsw i32 undef, 1
            %tmp20 = add nsw i32 %tmp19, %tmp18
            %tmp21 = sub nsw i32 %tmp11, %tmp12
            %tmp22 = shl nsw i32 undef, 1
            %tmp23 = add nsw i32 %tmp22, %tmp21
            %tmp24 = sub nsw i32 %tmp13, %tmp14
            %tmp25 = shl nsw i32 undef, 1
            %tmp26 = add nsw i32 %tmp25, %tmp24
            %tmp27 = sub nsw i32 %tmp16, %tmp17
            %tmp28 = shl nsw i32 undef, 1
            %tmp29 = add nsw i32 %tmp28, %tmp27
            %tmp30 = sub nsw i32 %tmp20, %tmp29
            %tmp31 = sub nsw i32 %tmp23, %tmp26
            %tmp32 = shl nsw i32 %tmp30, 1
            %tmp33 = add nsw i32 %tmp32, %tmp31
            store i32 %tmp33, i32* undef, align 4
            %tmp34 = mul nsw i32 %tmp31, -2
            %tmp35 = add nsw i32 %tmp34, %tmp30
            store i32 %tmp35, i32* undef, align 4
            %tmp36 = select i1 %tmp5, i32 undef, i32 undef
            br label %bb37
      
          bb37:                                             ; preds = %bb
            %tmp38 = load i32, i32* undef, align 4
            %tmp39 = ashr i32 %tmp38, %tmp6
            %tmp40 = load i32, i32* undef, align 4
            %tmp41 = sdiv i32 %tmp39, %tmp40
            store i32 %tmp41, i32* undef, align 4
            ret void
          }
      5cd66420
    • Florian Hahn's avatar
      Revert "[X86FixupLEAs] Sub register usage of LEA dest should block LEA/SUB optimization" · e087b4f1
      Florian Hahn authored
      This reverts commit f35bcea1 because it
      depends on 1b748faf, which breaks
      building the llvm-test-suite with -verify-machineinstrs on X86.
      
      See 154adc0f135cff3f8a8861c335d2b88c8049d098 for more details.
      e087b4f1
    • madhur13490's avatar
      [AMDGPU][IndirectCalls] Fix register usage propagation for indirect/external calls · c27e8141
      madhur13490 authored
      This patch computes max SGPRs and VGPRs used by module
      in presence of indirect calls and makes that
      as register requirement for functions/kernels
      which makes indirect calls.
      
      This patch also refactors code AMDGPUSubTarget.cpp
      which add a "base" variants of getMaxNumSGPRs which
      is used by MachineFunction and new Function version.
      
      Reviewed By: arsenm
      
      Differential Revision: https://reviews.llvm.org/D103636
      c27e8141
    • spupyrev's avatar
      A post-processing for BFI inference · 0a0800c4
      spupyrev authored
      The current implementation for computing relative block frequencies does
      not handle correctly control-flow graphs containing irreducible loops. This
      results in suboptimally generated binaries, whose perf can be up to 5%
      worse than optimal.
      
      To resolve the problem, we apply a post-processing step, which iteratively
      updates block frequencies based on the frequencies of their predesessors.
      This corresponds to finding the stationary point of the Markov chain by
      an iterative method aka "PageRank computation". The algorithm takes at
      most O(|E| * IterativeBFIMaxIterations) steps but typically converges faster.
      
      It is turned on by passing option `use-iterative-bfi-inference`
      and applied only for functions containing profile data and irreducible loops.
      
      Tested on SPEC06/17, where it is helping to get correct profile counts for one of
      the binaries (403.gcc). In prod binaries, we've seen a speedup of up to 2%-5%
      for binaries containing functions with hot irreducible loops.
      
      Reviewed By: hoy, wenlei, davidxl
      
      Differential Revision: https://reviews.llvm.org/D103289
      0a0800c4
    • Michael Kruse's avatar
      [Flang][test] Fix Windows buildbot. · dbc26296
      Michael Kruse authored
      Commit 1b241b9b /
      patch https://reviews.llvm.org/D104130 introduced an new test which
      calls a UNIX shell script. Add
      REQUIRES: shell
      to not run it on Windows.
      dbc26296
    • Stephen Neuendorffer's avatar
      [mlir] make normalizeAffineFor public · 984e270a
      Stephen Neuendorffer authored
      Previously this was just a static method.
      984e270a
    • Adrian Prantl's avatar
    • Alexander Shaposhnikov's avatar
      [lld][MachO] Fix function starts section · b9095f5e
      Alexander Shaposhnikov authored
      Sort the addresses stored in FunctionStarts section.
      Previously we were encoding potentially large numbers (due to unsigned overflow).
      
      Test plan: make check-all
      
      Differential revision: https://reviews.llvm.org/D103662
      b9095f5e
    • Jez Ng's avatar
      [lld-macho] Fix debug build · 5de7467e
      Jez Ng authored
      D103977 broke a bunch of stuff as I had only tested the release build
      which eliminated asserts.
      
      I've retained the asserts where possible, but I also removed a bunch
      instead of adding a whole lot of verbose ConcatInputSection casts.
      5de7467e
    • Uday Bondhugula's avatar
      [MLIR] Execution engine python binding support for shared libraries · c8b8e8e0
      Uday Bondhugula authored
      Add support to Python bindings for the MLIR execution engine to load a
      specified list of shared libraries - for eg. to use MLIR runtime
      utility libraries.
      
      Differential Revision: https://reviews.llvm.org/D104009
      c8b8e8e0
    • Kai Luo's avatar
      [AIX][compiler-rt] Fix cmake build of libatomic for cmake-3.16+ · 6393164c
      Kai Luo authored
      cmake-3.16+ for AIX changes the default behavior of building a `SHARED` library which breaks AIX's build of libatomic, i.e., cmake-3.16+ builds `SHARED` as an archive of dynamic libraries. To fix it, we have to build `libatomic.so.1` as `MODULE` which keeps `libatomic.so.1` as an normal dynamic library.
      
      Reviewed By: jsji
      
      Differential Revision: https://reviews.llvm.org/D103786
      6393164c
    • Adrian Prantl's avatar
      Allow signposts to take advantage of deferred string substitution · 4fc93a3a
      Adrian Prantl authored
      One nice feature of the os_signpost API is that format string
      substitutions happen in the consumer, not the logging
      application. LLVM's current Signpost class doesn't take advantage of
      this though and instead always uses a static "Begin/End %s" format
      string.
      
      This patch uses variadic macros to allow the API to be used as
      intended. Unfortunately, the primary use-case I had in mind (the
      LLDB_SCOPED_TIMER() macro) does not get much better from this, because
      __PRETTY_FUNCTION__ is *not* a macro, but a static string, so
      signposts created by LLDB_SCOPED_TIMER() still use a static "%s"
      format string. At least LLDB_SCOPED_TIMERF() works as intended.
      
      This reapplies the previsously reverted patch with support for
      platforms where signposts are unavailable.
      
      Differential Revision: https://reviews.llvm.org/D103575
      4fc93a3a
    • Jez Ng's avatar
      [lld-macho] Have dead-stripping work with literal sections · 464d3dc3
      Jez Ng authored
      Literal sections are not atomically live or dead. Rather,
      liveness is tracked for each individual literal they contain. CStrings
      have their liveness tracked via a `live` bit in StringPiece, and
      fixed-width literals have theirs tracked via a BitVector.
      
      The live-marking code now needs to track the offset within each section
      that is to be marked live, in order to identify the literal at that
      particular offset.
      
      Numbers for linking chromium_framework on my 3.2 GHz 16-Core Intel Xeon W
      with both `-dead_strip` and `--deduplicate-literals`, with and without this diff
      applied:
      
      ```
          N           Min           Max        Median           Avg        Stddev
      x  20          4.32          4.44         4.375         4.372    0.03105174
      +  20           4.3          4.39          4.36        4.3595   0.023277502
      No difference proven at 95.0% confidence
      ```
      This gives us size savings of about 0.4%.
      
      Reviewed By: #lld-macho, thakis
      
      Differential Revision: https://reviews.llvm.org/D103979
      464d3dc3
    • Jez Ng's avatar
      [lld-macho][nfc] Have InputSection ctors take some parameters · 681cfeb5
      Jez Ng authored
      This is motivated by an upcoming diff in which the
      WordLiteralInputSection ctor sets itself up based on the value of its
      section flags. As such, it needs to be passed the `flags` value as part
      of its ctor parameters, instead of having them assigned after the fact
      in `parseSection()`. While refactoring code to make that possible, I
      figured it would make sense for the other InputSections to also take
      their initial values as ctor parameters.
      
      Reviewed By: #lld-macho, thakis
      
      Differential Revision: https://reviews.llvm.org/D103978
      681cfeb5
    • Jez Ng's avatar
      [lld-macho][nfc] Move liveness-tracking fields into ConcatInputSection · 7f2ba39b
      Jez Ng authored
      These fields currently live in the parent InputSection class,
      but they should be specific to ConcatInputSection, since the other
      InputSection classes (that contain literals) aren't atomically live or
      dead -- rather their component string/int literals should have
      individual liveness states. (An upcoming diff will add liveness bits for
      StringPieces and fixed-sized literals.)
      
      I also factored out some asserts for isCoalescedWeak() in MarkLive.cpp.
      We now avoid putting coalesced sections in the `inputSections` vector,
      so we don't have to check/assert against it everywhere.
      
      Reviewed By: #lld-macho, thakis
      
      Differential Revision: https://reviews.llvm.org/D103977
      7f2ba39b
    • Jez Ng's avatar
      [lld-macho] Deduplicate fixed-width literals · 5d88f2dd
      Jez Ng authored
      Conceptually, the implementation is pretty straightforward: we put each
      literal value into a hashtable, and then write out the keys of that
      hashtable at the end.
      
      In contrast with ELF, the Mach-O format does not support variable-length
      literals that aren't strings. Its literals are either 4, 8, or 16 bytes
      in length. LLD-ELF dedups its literals via sorting + uniq'ing, but since
      we don't need to worry about overly-long values, we should be able to do
      a faster job by just hashing.
      
      That said, the implementation right now is far from optimal, because we
      add to those hashtables serially. To parallelize this, we'll need a
      basic concurrent hashtable (only needs to support concurrent writes w/o
      interleave reads), which shouldn't be to hard to implement, but I'd like
      to punt on it for now.
      
      Numbers for linking chromium_framework on my 3.2 GHz 16-Core Intel Xeon W:
      
            N           Min           Max        Median           Avg        Stddev
        x  20          4.27          4.39         4.315        4.3225   0.033225703
        +  20          4.36          4.82          4.44        4.4845    0.13152846
        Difference at 95.0% confidence
                0.162 +/- 0.0613971
                3.74783% +/- 1.42041%
                (Student's t, pooled s = 0.0959262)
      
      This corresponds to binary size savings of 2MB out of 335MB, or 0.6%.
      It's not a great tradeoff as-is, but as mentioned our implementation can
      be signficantly optimized, and literal dedup will unlock more
      opportunities for ICF to identify identical structures that reference
      the same literals.
      
      Reviewed By: #lld-macho, gkm
      
      Differential Revision: https://reviews.llvm.org/D103113
      5d88f2dd
    • Adrian Prantl's avatar
      Revert "Allow signposts to take advantage of deferred string substitution" · b90f9bea
      Adrian Prantl authored
      I forgot to make the LLDB macro conditional on Linux.
      
      This reverts commit 541ccd1c.
      b90f9bea
    • Andrew Litteken's avatar
      [IRSim] Strip out the findSimilarity call from the constructor · f6dea2e7
      Andrew Litteken authored
      Both doInitialize and runOnModule were running the entire analysis
      due to the actual work being done in the constructor. Strip it out here
      and only get the similarity during runOnModule.
      
      Author: lanza
      Reviewers: AndrewLitteken, paquette, plofti
      
      Differential Revision: https://reviews.llvm.org/D92524
      f6dea2e7
    • Adrian Prantl's avatar
      Disambiguate usage of struct mach_header and other MachO definitions. · 635b7213
      Adrian Prantl authored
      Unfortunately the Darwin signpost header also pulls in the system
      MachO header and so we need to make sure to use the LLVM versions of
      those definitions.
      635b7213
    • Adrian Prantl's avatar
      Allow signposts to take advantage of deferred string substitution · 541ccd1c
      Adrian Prantl authored
      One nice feature of the os_signpost API is that format string
      substitutions happen in the consumer, not the logging
      application. LLVM's current Signpost class doesn't take advantage of
      this though and instead always uses a static "Begin/End %s" format
      string.
      
      This patch uses variadic macros to allow the API to be used as
      intended. Unfortunately, the primary use-case I had in mind (the
      LLDB_SCOPED_TIMER() macro) does not get much better from this, because
      __PRETTY_FUNCTION__ is *not* a macro, but a static string, so
      signposts created by LLDB_SCOPED_TIMER() still use a static "%s"
      format string. At least LLDB_SCOPED_TIMERF() works as intended.
      
      Differential Revision: https://reviews.llvm.org/D103575
      541ccd1c
    • Alexander Shaposhnikov's avatar
      [llvm-objcopy][MachO] Do not strip symbols with the flag REFERENCED_DYNAMICALLY set · 0276cc74
      Alexander Shaposhnikov authored
      Do not strip symbols having the flag REFERENCED_DYNAMICALLY set.
      
      Test plan: make check-all
      
      Differential revision: https://reviews.llvm.org/D104092
      0276cc74