NVIDIA Reasoning Model text

Nemotron 3 Ultra Pricing & Token Costs

Official API rate card for Nemotron 3 Ultra provided by NVIDIA. All token pricing is normalized per 1 Million tokens ($/1M).

Token Pricing Rates ($/1M Tokens)

Input Tokens 0.4230 per 1M tokens
Output Tokens 2.6100 per 1M tokens
Prompt Cache Read Not Supported per 1M tokens
Prompt Cache Write Not Supported per 1M tokens

Technical Specifications

Provider
NVIDIA
Context Window
256k tokens (256,000)
Max Output Tokens
8k tokens (8,192)
Reasoning Architecture
Yes
Supported Modalities
text

Estimated Inference Costs

Example cost calculation for standard application workloads assuming a 3:1 input-to-output token ratio.

100,000 requests (500 in / 150 out) 0.06
1,000,000 requests (1,000 in / 300 out) 1.21
← Compare Nemotron 3 Ultra on Master Pricing Table