How many GPUs is 1M/B/T tokens?

原始链接: https://cedana.com/resources/tokens-to-gpus/

Time to serve this volumeOne dayOne month (30 days)One year (365 days)ModelLlama 3.1 8BLlama 3.3 70Bgpt-oss 120B (5.1B active)DeepSeek V4 Flash (13B active)MiniMax M3 428BDeepSeek R1 / V3 (37B active)GLM-5.3 (40B active)DeepSeek V4 Pro (49B active)Custom model:: 70b total / 70b active / 70 gb weights / ref h100 / moderate confidenceGPU typeReference GPU for the modelA100V100H100B200Sets the GPU for the headline, the calculation, and the sensitivity diagram. All GPU types stay in the chart.Token mixAll output tokens (upper bound)Chat: 3 input to 1 outputRAG or agent: 10 input to 1 outputAn input token costs less GPU time than an output token.Advanced assumptionsThe first row is output tokens per second for one GPU of the reference type at full load. Input cost is the GPU time for one input token divided by the time for one output token.Reset to defaults »

相关文章

原文

:: 70b total / 70b active / 70 gb weights / ref h100 / moderate confidence

Sets the GPU for the headline, the calculation, and the sensitivity diagram. All GPU types stay in the chart.

An input token costs less GPU time than an output token.

Advanced assumptions

The first row is output tokens per second for one GPU of the reference type at full load. Input cost is the GPU time for one input token divided by the time for one output token.

联系我们 contact @ memedata.com