# Comparison of GPUs for LLM Inference Units are Tera (10^9). ```table GPU Year RAM B/W FP16 FP8 INT8 FP4 INT4 AMD 6900 XT 2020 24G 0.51 0.04 - - - - Nvidia 3090 2020 24G 0.94 0.14 - 0.28 - 0.57 AMD 7900 XTX 2022 24G 0.96 0.12 - 0.12 - 0.25 Nvidia 4090 2022 24G 1.00 0.33 0.66 0.66 - 1.32 AMD 9070 XT 2025 16G 0.64 0.20 0.39 0.39 - 0.78 Nvidia 5090 2025 32G 1.79 0.42 0.83 0.83 1.67 ? ``` The TOPS (Tera Operations per Second) is significant for prefill. The B/W (Bandwidth in GB/s) is significant for token generation. There may be other architectural differences that are not accounted for. Software support is better for Nvidia cards (AMD really sucks for support in general). Of course for best-value, you will need up-to-date second-hand pricing.