Note that the four bottom sorts (RustSort, StdSort, Pdqsort, IPS4o) are not stable. GlidesortN is glidesort using n elements of memory, Glidesort is the default memory setting and Glidesort1024 uses a fixed 1024 element buffer.
Arm is a 2021 Apple M1, Amd is a 2018 Threadripper 2950x.
Those are some awfully large array sizes. Not something I ever had the patience for to sit out, or optimize for, as it didn't seem like the most common real-world scenario.
One thing glidesort appears to do extremely well is to branchless merge till the very end, this gets quite tricky when memory constrained. It does not seem worth it intuitively, but judging from the performance of rotate mergesort, it probably is.
No, it is running serially. IPS4o is inherently more cache efficient due to having a very wide partition operator, making it better on machines with smaller caches, and better in general for strings (where no matter how cache-local your algorithm is, the string pointer indirection mostly destroys it).