🧮 LLM GPU Planner

Two tools to help you predict latency and estimate GPU requirements. Use the tabs below 👇

Estimate GPUs needed from max concurrency.

Formula: GPUs = ceil(max_concurrency × target_stream_tps / (per_gpu_tps × util_cap))

0.1 0.95

Optional: derive Per-GPU TPS from a measured cluster