Rust SIMD on the GPU
VectorWare, a company building the first GPU-native software stack, previously brought Rust threads to the GPU by mapping each std::thread to a GPU warp. While that enabled many concurrent threads, it did not use the parallel lanes within each warp. On CPUs, the abstraction for within-thread parallelism is SIMD, where a single instruction operates on a vector of data elements.
Rust's portable SIMD (core::simd) provides a generic Simd<T,N> type that abstracts over architecture-specific intrinsics such as _mm256_add_ps on x86-64 or vaddq_f32 on Arm, allowing one implementation to target multiple ISAs. Because portable SIMD lives in core rather than std, it also does not require the std support VectorWare had already brought to the GPU.
The key insight is that a GPU warp executes in SIMT (Single Instruction, Multiple Thread), which is effectively a wide vector unit. A Simd<i16,32> gives one i16 element to each of the warp's 32 lanes, and adding two such vectors compiles to a single warp instruction where every lane adds its elements in parallel. The post demonstrates that on a CPU the same operation compiles to vpaddw with 32-bit zmm registers, highlighting the direct mapping.
This milestone is a step toward VectorWare's vision of enabling developers to write complex, high-performance GPU applications using familiar Rust abstractions, now including portable SIMD for both CPU and GPU targets.