Algorithm · Applied Intuition · Hard
Parallel reductions on CUDA GPUs can produce different answers across executions when the order in which partial values are combined is not fixed. This matters especially for floating-point addition, because floating-point addition is not strictly associative. Design a deterministic reduction method whose evaluation order is known in advance. Use prefix-sum/scan techniques as the basis for the design. Answer each item: Within a single CUDA block, how would you organize the…
Checking your access…