Amd · Behavioral
Optimize the Memory Accesses of a GPU Matrix Transpose
TrueInterview
September 26, 2026 · 1 min read
Examine and optimize a GPU kernel that transposes a matrix. When a matrix is stored in row-major order, a simple thread assignment typically yields contiguous memory accesses for either the reads or the writes across adjacent lanes, leaving the other direction with a stride.
Constraints and Assumptions
The actual ROCm kernel source is not provided. Work with the following standalone model: assume input array A holds R rows and C columns in row-major layout, and output array B is sized C rows by R columns, where . Describe a tiled approach that explicitly uses shared on-chip scratchpad memory. Tile sizes must stay within the thread count and memory constraints of the target GPU.
Questions to Clarify
Which index changes across neighboring threads in a warp? What coalescing rules and scratchpad memory bank layout are relevant? Do the matrix dimensions divide evenly by the tile dimensions? Are the input and output buffers stored in separate memory allocations?
What a Solid Response Covers
Analysis of read and write address patterns, tiled staging with barrier synchronization, handling of edge cases, attention to bank conflicts, and performance evaluation.
Follow-Up Questions
Why does merely reordering threads often shift the stride issue to the opposite access side? Why is it necessary for every participating thread to arrive at the barrier? Is adding a one-element padding always the best approach for every GPU architecture?
Overview: Examine row-major transpose memory access patterns and apply tiled shared-memory staging to achieve coalesced reads and writes, ensuring barrier correctness, edge handling, and device-aware bank conflict checks.
This prompt is drawn from a shared AMD Software Engineer interview experience.
Loading comments…