#01 Performance
Build a Model Sharding System for Serving 70B+ Parameter Models
Premium
Design a model sharding system that splits a 70B-parameter model's 140GB of FP16 weights across 8 A100 GPUs using tensor and pipeline parallelism, sustaining 40 tokens/s per replica with under 20GB of KV cache headroom per shard.