Tensor parallelism · TP
Splitting a single layer's math across multiple GPUs so they jointly compute one forward/backward pass.
Current numbers
~42 GBaggregate KV-cache per Llama 3 70B sequence at 128K across participating ranks; TP=8 ideal head-sharded illustration ~5.25 GB/rank before overhead
~18×NVLink5 scale-up (1.8 TB/s/GPU bidirectional, ~900 GB/s/dir; 130 TB/s NVL72 rack) over ~400G scale-out NIC (~50 GB/s/dir) — keep TP/EP inside scale-up