- Oct 30, 2020
-
-
Jonas Devlieghere authored
Recognize the __apple_ sections as debug info sections and make sure they're included in the --show-sections-sizes output. Differential revision: https://reviews.llvm.org/D90433
-
Keith Smiley authored
This is to enable --allow-unused-duplicates=false. This prefix appears to be outdated and intentionally unused. Reviewed By: rupprecht Differential Revision: https://reviews.llvm.org/D90427
-
Stella Laurenzo authored
* Removes index based insertion. All insertion now happens through the insertion point. * Introduces thread local context managers for implicit creation relative to an insertion point. * Introduces (but does not yet use) binding the Context to the thread local context stack. Intent is to refactor all methods to take context optionally and have them use the default if available. * Adds C APIs for mlirOperationGetParentOperation(), mlirOperationGetBlock() and mlirBlockGetTerminator(). * Removes an assert in PyOperation creation that was incorrectly constraining. There is already a TODO to rework the keepAlive field that it was guarding and without the assert, it is no worse than the current state. Differential Revision: https://reviews.llvm.org/D90368
-
Wouter van Oortmerssen authored
Differential Revision: https://reviews.llvm.org/D90428
-
Nico Weber authored
Make check_clang_tidy.py not just pass -format-style=none by default but a full -config={}. Without this, with a build dir outside of the llvm root dir and a .clang-tidy config further up that contains CheckOptions: - key: modernize-use-default-member-init.UseAssignment value: 1 these tests would fail: Clang Tools :: clang-tidy/checkers/cppcoreguidelines-prefer-member-initializer-modernize-use-default-member-init.cpp Clang Tools :: clang-tidy/checkers/modernize-use-default-member-init-bitfield.cpp Clang Tools :: clang-tidy/checkers/modernize-use-default-member-init.cpp After this change, they pass fine, despite the unrelated .clang-tidy file further up. -
Krzysztof Parzyszek authored
-
https://reviews.llvm.org/D88483Jim Ingham authored
The test can be cleaned up a bit, but this should be good to see why the Debian bot is failing...
-
Aaron Puchert authored
This fixes the issue pointed out in D84604#2363134. For now we exclude static members completely, we'll take them into account later.
-
Arthur Eubanks authored
-
Scott Linder authored
Make all of the "AMDGPU Machine Code GFX*" columns in the Memory Model table a consistent width of 32-characters. Best viewed with something like --word-diff Differential Revision: https://reviews.llvm.org/D89977
-
Scott Linder authored
Mostly NFC, but some changes are "bug fixes" rather than just e.g. formatting changes or typo corrections. - Fix typo "competing" -> "completing". - Document why waintcnt is added to stores and not loads for sequentially consistent ordering. - Lowercase some mentions of `buffer_gl{0,1}_inv`. - Make mentions of `*cnt(0)` consistently include the `(0)` count. - Remove some mentions of instructions for incorrect address spaces. For example, remove mention of `flat_load` from `load atomic acquire workgroup global`. - Re-flow some text to get all the target columns to fit in a 32-character wide column. Makes a future NFC patch to make these columns both 32-character wide more straightforward. Modified cherry-pick of patch by Tony Tye Reviewed By: t-tye Differential Revision: https://reviews.llvm.org/D89596 -
Kostya Kortchinsky authored
Mitch expressed a preference to not have `#ifdef`s in platform agnostic code, this change tries to accomodate this. I am not attached to the method this CL proposes, so if anyone has a suggestion, I am open. We move the platform specific member of the mutex into its own platform specific class that the main `Mutex` class inherits from. Functions are implemented in their respective platform specific compilation units. For Fuchsia, we use the sync APIs, as those are also the ones being used in Scudo. Differential Revision: https://reviews.llvm.org/D90351
-
Joachim Meyer authored
This is very similar to 7f1e6fcf, just fixing a left-over. With this, it should be possible to use both, -x cuda and -fopenmp in the same invocation, enabling to use both OpenMP, targeting CPU, and CUDA, targeting the GPU. Reviewed By: jdoerfert Differential Revision: https://reviews.llvm.org/D90415
-
Jim Ingham authored
-
Thomas Raoux authored
Fix semantic in the distribute integration test based on offline feedback. This exposed a bug in block distribution, we need to make sure the id is multiplied by the stride of the vector. Fix the transformation and unit test. Differential Revision: https://reviews.llvm.org/D89291
-
Craig Topper authored
This combine makes two calls to SimplifyDemandedBits, one for the LHS and one for the RHS. If the LHS call returns true, we don't make the RHS call. When SimplifyDemandedBits makes a change, it will add the nodes around the change to the DAG combiner worklist. If the simplification happens on the first recursion step, the N will get added to the worklist. But if the simplification happens deeper in the recursion, then N will not be revisited until the next time the DAG combiner runs. This patch explicitly addes N to the worklist anytime a Simplification is made. Without this we might miss additional simplifications on the LHS or never simplify the RHS. Special care also needs to be taken to not add N if it has been CSEd by the simplification. There are similar examples in DAGCombiner and the X86 target, but I don't have a test for it for RISC-V. I've also returned SDValue(N, 0) instead of SDValue() so DAGCombiner knows a change was made and will update its Statistic variable. The test here was constructed so that 2 simplifications happen to the LHS. Without this fix one happens in the post type legalization DAG combine and the other happens after LegalizeDAG. This prevents the RHS from ever being simplified causing the left and right shift to clear the upper 32 bits of the RHS to be left behind. Differential Revision: https://reviews.llvm.org/D90339
-
Craig Topper authored
-
Jim Ingham authored
we should be exporting one by one. Differential Revision: https://reviews.llvm.org/D78972
-
Jim Ingham authored
The intention is not to allow stop-hook commands to query the user, so this is correct. It also works around a deadlock in switching to the Python Session to execute python based commands in the stop hook when the Debugger stdin is backed by a FILE *. Differential Revision: https://reviews.llvm.org/D90332
-
Christian Sigg authored
Do not use the pass yet, except in a test. Reviewed By: herhut Differential Revision: https://reviews.llvm.org/D89937
-
Christian Sigg authored
For the synchronous case, destroy the stream after synchronization. Sneak in a unrelated change to report why the gpu.wait conversion pattern didn't match. Reviewed By: herhut Differential Revision: https://reviews.llvm.org/D89933
-
Nikita Popov authored
Make the existing VECREDUCE based code more generic, but expressing it in terms of the neutral value of the base opcode instead.
-
Stefanos Baziotis authored
Differential Revision: https://reviews.llvm.org/D89739
-
Christian Sigg authored
This is a roll-forward of rGec7780eb, now that the remaining gpu.launch_func have been converted to custom form in rGb22f1110. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D90420
-
Nikita Popov authored
Use -0.0 instead of 0.0 as the start value. The previous use of 0.0 was fine for all existing uses of this function though, as it is always generated with fast flags right now, and thus nsz.
-
Ilya Bukonkin authored
-
Florian Hahn authored
Some architectures do not have general vector select instructions (e.g. AArch64). But some cmp/select patterns can be vectorized using other instructions/intrinsics. One example is using min/max instructions for certain patterns. This patch updates the cost calculations for selects in the SLP vectorizer to consider using min/max intrinsics. This patch does not change SLP vectorizer's codegen itself to actually generate those intrinsics, but relies on the backends to lower the vector cmps & selects. This keeps things simple on the SLP side and works well in practice for AArch64. This exposes additional SLP vectorization opportunities in some benchmarks on AArch64 (-O3 -flto). Metric: SLP.NumVectorInstructions Program base slp diff test-suite...ications/JM/ldecod/ldecod.test 502.00 697.00 38.8% test-suite...ications/JM/lencod/lencod.test 1023.00 1414.00 38.2% test-suite...-typeset/consumer-typeset.test 56.00 65.00 16.1% test-suite...6/464.h264ref/464.h264ref.test 804.00 822.00 2.2% test-suite...006/453.povray/453.povray.test 3335.00 3357.00 0.7% test-suite...CFP2000/177.mesa/177.mesa.test 2110.00 2121.00 0.5% test-suite...:: External/Povray/povray.test 2378.00 2382.00 0.2% Reviewed By: RKSimon, samparker Differential Revision: https://reviews.llvm.org/D89969
-
Thomas Lively authored
This commit removes unused FileCheck prefixes from WebAssembly test files to avoid causing test failures once FileCheck disallows unused prefixes by default. See D90281 and the corresponding llvm-dev thread for context. Reviewed By: aardappel Differential Revision: https://reviews.llvm.org/D90416
-
Nikita Popov authored
The neutral value for FADD is -0.0, not 0.0, so this is what we need to pad vectors with.
-
Nikita Popov authored
The neutral value is -0.0, not 0.0. This doesn't matter for "fast" reductions due to nsz, but does matter for reassoc-only and seq reductions. Change tests to mostly use -0.0 where the neutral value was intended, and add some additional test coverage in some places. Also update LangRef to use the right value.
-
Christian Sigg authored
This should fix the reason for the failures after ec7780eb. I will roll forward in a separate change. Reviewed By: antiagainst Differential Revision: https://reviews.llvm.org/D90410
-
Tony authored
- AMDGPUUsage.rst: Correct AMD GPU DWARF address space table address sizes which are in bits and not bytes. - clang/.../Options.td: Improve description of AMD GPU options. - Re-generate ClangComamndLineReference.rst from clang/.../Options.td . Differential Revision: https://reviews.llvm.org/D90364
-
Alex Orlov authored
Use LLVM/utils/remote-exec.py to run compiler-rt tests remotely on the target. Reviewed By: vvereschaka Differential Revision: https://reviews.llvm.org/D90054
-
Marcel Hlopko authored
This preprocessor define was meant to be used to conditionally include VCSVersion.inc. However, the define was always set, and it was the content of the header that was conditionally generated. Therefore HAVE_VCS_VERSION_INC should be cleaned up. Reviewed By: gribozavr2, MaskRay Differential Revision: https://reviews.llvm.org/D84623
-
Nikita Popov authored
-
Adhemerval Zanella authored
On aarch64 with kernel 4.12.13 the test sporadically fails with RSS at start: 1564, after mmap: 103964, after mmap+set label: 308768, \ after fixed map: 206368, after another mmap+set label: 308768, after \ munmap: 206368 release_shadow_space.c.tmp: [...]/release_shadow_space.c:80: int \ main(int, char **): Assertion `after_fixed_mmap <= before + delta' failed. It seems on some executions the memory is not fully released, even after munmap. And it also seems that ASLR is hurting it by adding some fragmentation, by disabling it I could not reproduce the issue in multiple runs.
-
Peyton, Jonathan L authored
Patch by Nawrin Sultana Differential Revision: https://reviews.llvm.org/D90403
-
Dávid Bolvanský authored
One step closer to fix PR47644. Differential Revision: https://reviews.llvm.org/D89645
-
Paul-Antoine Arras authored
This diff adds support for LLVM bitcode objects to llvm-libtool-darwin. Test plan: make check-all Differential revision: https://reviews.llvm.org/D88722
-
Utkarsh Saxena authored
With every incremental change, one needs to check-in new model upstream. This also significantly increases the size of the git repo with every new model. Testing and comparing the old and previous model is also not possible as we run only a single model at any point. One solution is to have a "staging" decision forest which can be injected into clangd without pushing it to upstream. Compare the performance of the staging model with the live model. After a couple of enhancements have been done to staging model, we can then replace the live model upstream with the staging model. This reduces upstream churn and also allows us to compare models with current baseline model. This is done by having a callback in CodeCompleteOptions which is called only when we want to use a decision forest ranking model. This allows us to inject different completion model internally. Differential Revision: https://reviews.llvm.org/D90014
-