Hand-counted FLOPs
Key Insight
Counting FLOPs by hand reveals where the computational weight of a transformer model actually lies. By calculating the cost of matrix multiplications in the attention and MLP blocks, we see that the vast majority of operations are simple linear projections. This mathematical exercise sets the baseline for analyzing hardware efficiency and resource utilization.