Back to models
Try in Agents
Model · NVIDIA
Nemotron 3 Ultra
nvidia/nemotron-3-ultra
NVIDIA's 550B open reasoning model (55B active), a hybrid Mamba-Transformer MoE built for long-running agent workflows
reasoningcaching
Specs
Context
1M
Max output
16K
Provider
NVIDIA
Input
$0.600/ 1M tokens
Output
$2.40/ 1M tokens
Cache read
$0.120/ 1M tokens
Try it
Chat with Nemotron 3 UltraTerramind API
⌘+Enter to run
Code
import { streamText } from "ai";
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
const terramind = createOpenAICompatible({
name: "terramind",
baseURL: "https://terramind.com/api/v1",
apiKey: process.env.TERRAMIND_API_KEY!,
});
const result = streamText({
model: terramind("nemotron-3-ultra"),
prompt: "Why is the sky blue?",
});
for await (const chunk of result.textStream) {
process.stdout.write(chunk);
}
Want to run this in your app? Grab an API key · Full chat reference
More by NVIDIA
| Model | Context | In | Out | Capabilities |
|---|---|---|---|---|
| 256K | $0.150 | $0.650 | reasoning | |
| 262K | $0.050 | $0.200 | reasoningcaching |