5. SIMD/vector register size and memory system
We've seen that the essential difference between SIMD and vector is the number of instructions (opcodes). We have also seen that vector extensions would benefit from larger vector registers. Three examples show that larger registers lead us to reconsider the memory system.
The AVX-512 instructions VMOVDQA, VMOVAPS, VMOVAPD for aligned loads and stores, and the equivalent instructions for non-aligned accesses, transfer 512 bits between the SIMD registers and the L1 data cache. Intel CPUs with the AVX-512 extension have cache lines of 64 bytes (512 bits). A CPU-to-L1 cache transfer therefore transfers an entire line. This means that the use of prefetch (hardware, or hardware plus software) is essential to avoid a cache miss on each access. The L1D cache...
You do not have access to this resource.
Exclusive to subscribers. 97% yet to be discovered!
Already subscribed?
Log in!
Ongoing reading
SIMD/vector register size and memory system