Tiled Matrix Multiplication
Step through C = A x B as cache-sized tiles are fetched, reused, and written.
Progress
0 / 0
Ready
cache line: 64B
A matrix
M x KB matrix
K x NC matrix
M x NMemory layout heatmap
linearized arrays
read
write/update
same cache line
not touched
A memory
B memory
C memory
This step
Next step
Tiled loop order
for ii in 0..M step T:
for jj in 0..N step T:
for kk in 0..K step T:
for i in ii..ii+T-1:
for j in jj..jj+T-1:
for k in kk..kk+T-1:
C[i,j] += A[i,k] * B[k,j]