Tiled Matrix Multiplication

Step through C = A x B as cache-sized tiles are fetched, reused, and written.

Progress 0 / 0 Ready
Thread-owned page
cache line: 64B
fetched now active tile written now partial C completed C not used now

A matrix

M x K

B matrix

K x N

C matrix

M x N

Outer-product tile micro-step

T x T tile update

A sub-vector

B sub-vector

rank-1 product grid

C tile after this k

Memory layout heatmap

linearized arrays
read write/update same cache line not touched

A memory

B memory

C memory

This step

Next step

Tiled loop order

for ii in 0..M step T:
  for jj in 0..N step T:
    for kk in 0..K step T:
      for i in ii..ii+T-1:
        for j in jj..jj+T-1:
          for k in kk..kk+T-1:
            C[i,j] += A[i,k] * B[k,j]