teach GCC how to use CMPXCHG on 128 bits
On x86-64, the reference counted pointer implementation ends up being a 128-bit scalar (`__int128` provided by GCC, Clang and ICC). The MOV instruction on x86-64 is only guaranteed atomic up to 64 bits, a quadword. To get atomic operations on 128 bits we need to leverage CMPXCHG. The result is a counter- intuitive degenerate CMPXCHG to implement an atomic read. A further complication is that GCC version >= 7.1 will not emit CMPXCHG for __atomic_compare_exchange* even when you pass the command-line flag '-mcx16'. To get around this, we use the old __sync built-ins, despite that they are considered deprecated. For more information, see https://gcc.gnu.org/bugzilla/show_bug.cgi?id=80878. We don't consider this to fully close Github #32, because a further optimisation would be to consider simply doing the initial read non-atomically. In a sense, it doesn't matter if we perform a torn read here because we only commit it in a CMPXCHG, which would simply fail if the dependent read was torn. Related to Github #32 "mcx16 hacks"
parent
0e03a0bc
Please register or sign in to comment