Two tools to help you predict latency and estimate GPU requirements. Use the tabs below 👇
Estimate GPUs needed from max concurrency.
Formula: GPUs = ceil(max_concurrency × target_stream_tps / (per_gpu_tps × util_cap))
GPUs = ceil(max_concurrency × target_stream_tps / (per_gpu_tps × util_cap))
Optional: derive Per-GPU TPS from a measured cluster
Estimate total response and first-token (TTFT) latency using your fitted model.
⚙️ Validation: TP × DP ≤ 8; only 8× GPU setups are valid.