- Jun 17, 2022
-
-
Walter Erquinigo authored
This is the final functional patch to support intel pt decoding per cpu. It works by doing the following: - First, all context switches are split by tid and sorted in order. This produces a list of continuous executes per thread per core. - Then, all intel pt subtraces are split by PSB boundaries and assigned to individual thread continuous executions on the same core by doing simple TSC-based comparisons. - With this, we have, per thread, a sorted list of continuous executions each one with a list of intel pt subtraces. Up to this point, this is really fast because no instructions were actually decoded. - Then, each thread can be decoded by traversing their continuous executions and intel pt subtraces. An advantage of having these continuous executions is that we can identify if a continuous exexecution doesn't have intel pt data, and thus has a gap in it. We can later to more sofisticated comparisons to identify if within a continuous execution there are gaps. I'm adding a test as well. Differential Revision: https://reviews.llvm.org/D126394
-
Walter Erquinigo authored
- Add the logic that parses all cpu context switch traces and produces blocks of continuous executions, which will be later used to assign intel pt subtraces to threads and to identify gaps. This logic can also identify if the context switch trace is malformed. - The continuous executions blocks are able to indicate when there were some contention issues when producing the context switch trace. See the inline comments for more information. - Update the 'dump info' command to show information and stats related to the multicore decoding flow, including timing about context switch decoding. - Add the logic to conver nanoseconds to TSCs. - Fix a bug when returning the context switches. Now they data returned makes sense and even empty traces can be returned from lldb-server. - Finish the necessary bits for loading and saving a multi-core trace bundle from disk. - Change some size_t to uint64_t for compatibility with 32 bit systems. Tested by saving a trace session of a program that sleeps 100 times, it was able to produce the following 'dump info' text: ``` (lldb) trace load /tmp/trace3/trace.json (lldb) thread trace dump info Trace technology: intel-pt thread #1: tid = 4192415 Total number of instructions: 1 Memory usage: Total approximate memory usage (excluding raw trace): 2.51 KiB Average memory usage per instruction (excluding raw trace): 2573.00 bytes Timing for this thread: Timing for global tasks: Context switch trace decoding: 0.00s Events: Number of instructions with events: 0 Number of individual events: 0 Multi-core decoding: Total number of continuous executions found: 2499 Number of continuous executions for this thread: 102 Errors: Number of TSC decoding errors: 0 ``` Differential Revision: https://reviews.llvm.org/D126267 -
Paul Robinson authored
-
Dávid Bolvanský authored
-
Philip Reames authored
If we're writing to an undef vector (i.e. implicit_def), we can change the value of bits outside the requested write without consequence. This allows us to avoid a VSETVLI just for narrowing the value written. Differential Revision: https://reviews.llvm.org/D127880
-
Paul Robinson authored
-
Michael Jones authored
On windows size_t != unsigned long. Differential Revision: https://reviews.llvm.org/D127989
-
Greg Clayton authored
The bug was introduced when the AddressRange class was no longer able to modify the End address directly and the entire range of the .text address range that contained the trailing empty symbol was replaced. There was no unit test for this, so it wasn't caught. I fixed the bug and added a unit test for it. The effects of this bug are serious as the AddressOffsetSize in the header would be incorrectly calculated and an invalid GSYM would be created. Differential Revision: https://reviews.llvm.org/D127811
-
Peter Klausler authored
Overflow detection in the folding of int/nint/ceiling is incorrectly signalling overflow when a negative argument yields a zero result. Differential Revision: https://reviews.llvm.org/D127785
-
Alexey Bataev authored
If the root scalar is mapped to to the smallest bit width, the vector is truncated and the types between original buildvector and extracted value mismatched. For extract, we emit sext/zext instructions, for shuffles we can reuse oringal vector instead of the truncated one. Differential Revision: https://reviews.llvm.org/D127974
-
Mark de Wever authored
-
Philip Reames authored
Apparently the parser/verifier is more lax than it should be. The typo'd names should have been rejected.
-
Jay Foad authored
Differential Revision: https://reviews.llvm.org/D127955
-
Jay Foad authored
This includes: - New llvm.amdgcn.image.msaa.load.* intrinsics - NSA changes, because MIMG-NSA is now limited to 3 dwords - Split CD forms of IMAGE_SAMPLE instructions out into separate test files since they are no longer supported in GFX11 Differential Revision: https://reviews.llvm.org/D127837
-
Jay Foad authored
Differential Revision: https://reviews.llvm.org/D127671
-
David Green authored
-
Arthur Eubanks authored
Reviewed By: #opaque-pointers, rnk, nikic Differential Revision: https://reviews.llvm.org/D126309
-
Peter Klausler authored
There's a few (3) cases where Fortran allows two distinct symbols to have the same name in the same scope. Module file output copes with only two of them. The third involves a separate module procedure that isn't separate: both the procedure and its declared interface appear in the same (sub)module. Fix to ensure that the interface is included in the module file output, so that the module file reader doesn't suffer a bogus error about a "separate module procedure without an interface". Differential Revision: https://reviews.llvm.org/D127784
-
Craig Topper authored
This removes one of the uses of ForceTailUndisturbed.
-
Peter Klausler authored
Previous one was returning a bogus error status about a bad WAIT statement ID number. Differential Revision: https://reviews.llvm.org/D127979
-
Michael Jones authored
Previously, any line buffered write of size 0 would cause an error. The variable used to track the index of the last newline started at the size of the write - 1, which underflowed. Now it's handled properly, and a test has been added to prevent regressions. Reviewed By: sivachandra, lntue Differential Revision: https://reviews.llvm.org/D127914
-
Michael Jones authored
The hex converter handles the %x and %X conversions. Reviewed By: sivachandra Differential Revision: https://reviews.llvm.org/D126082
-
Thomas Raoux authored
contraction op can have mixed type, add support for this case to the pattern lowering contraction op to outerproduct. Differential Revision: https://reviews.llvm.org/D127926
-
Arthur Eubanks authored
Reviewed By: pcc Differential Revision: https://reviews.llvm.org/D127876
-
Thomas Raoux authored
Support the case where convolution does float extension of the inputs. Differential Revision: https://reviews.llvm.org/D127925
-
Philip Reames authored
-
Alex Brachet authored
-
Mircea Trofin authored
This requires DominatorTree be updated, which we do in the ml inliner case, but not in the default case, and the cost of doing so is noticeable to compile time for the latter[1]. So the patch only affects the ML inliner. [1] https://llvm-compile-time-tracker.com/compare.php?from=9fc0aa45e3312944431ba7e1ca0cec99c613992b&to=7af461b1ce0d9138211ef5f883f35d5b9ddf47be&stat=wall-time Differential Revision: https://reviews.llvm.org/D127899
-
Craig Topper authored
We may eventually need tail undisturbed patterns, but we will need a policy operand on the ISD node to communicate it.
-
Adrian Tong authored
This helps handling a case where the BUILD_VECTOR has i16 element type and i32 constant operands t2: v8i16 = setcc t8, t17, setult:ch t3: v8i16 = BUILD_VECTOR Constant:i32<1>, ... t4: v8i16 = and t2, t3 t5: v8i16 = add t8, t4 This can be turned into t5: v8i16 = sub t8, t2, and allows us to remove t3 and t4 from the DAG. Differential Revision: https://reviews.llvm.org/D127354 -
Nikolas Klauser authored
This reverts commit 147f74b6.
-
- Jun 16, 2022
-
-
Philip Reames authored
This change just moves some code around, and extracts out a helper function expected to be useful when reusing the demanded field logic in the forward dataflow.
-
Ahsan Saghir authored
This patch fixes the load and store quadword instructions on PowerPC to use correct offset and base address. Reviewed By: #powerpc, nemanjai, lkail Differential Revision: https://reviews.llvm.org/D126807
-
Philip Reames authored
-
Alex Brachet authored
Fuchsia's dynamic linker does not and will never support IFUNC's. Differential revision: https://reviews.llvm.org/D127933
-
Craig Topper authored
If the merge operand isn't undef we need to be using tail undisturbed. Turns out all of our uses of riscv_slidedown_vl use undef so this doesn't affect any tests.
-
Corentin Jabot authored
LLVM had 2 methods to convert a number to an hexa string, this remove one of them. Differential Revision: https://reviews.llvm.org/D127958
-
Kevin P. Neal authored
-
Philip Reames authored
-
Philip Reames authored
The motivating case, and the only one actually enabled by this patch, is a load or store followed by another op with the same SEW/LMUL ratio. As an example, consider: define void @test1(ptr %in, ptr %out) { entry: %0 = load <8 x i16>, ptr %in, align 2 %1 = sext <8 x i16> %0 to <8 x i32> store <8 x i32> %1, ptr %out, align 4 ret void } Without this patch, we get: vsetivli zero, 8, e16, mf4, ta, mu vle16.v v8, (a0) vsetvli zero, zero, e32, mf2, ta, mu vsext.vf2 v9, v8 vse32.v v9, (a1) ret Whereas with the patch we get: vsetivli zero, 8, e32, mf2, ta, mu vle16.v v8, (a0) vsext.vf2 v9, v8 vse32.v v9, (a1) ret We have rewritten the first vsetvli and thus removed the second one. As is strongly hinted by the code structure and todos, I am planning on communing this with all (or most all?) of the cases from isCompatible used in the forward data flow. This will be done in a series of following changes - some NFC reworks, and some reviewed optimization extensions. Differential Revision: https://reviews.llvm.org/D127780
-