Occupancy Study
Key Insight
Measuring Streaming Multiprocessor (SM) occupancy shows the ratio of active warps to the maximum supported warps on the hardware. Sweeping block size in a custom CUDA kernel demonstrates how resource limits, such as register allocation and shared memory size, constrain parallelism and affect the GPU's ability to hide memory latency.