Back to models
Model · NVIDIA
NVIDIA

Nemotron 3.5 Lightning

nvidia/nemotron-3.5-lightning
Try in Agents

NVIDIA's 30B hybrid Mamba-Transformer MoE (3B active) — a very fast, very cheap reasoning model for agents and tool calls

reasoningcaching
Specs
Context
262K
Max output
33K
Provider
NVIDIA
Input
$0.050/ 1M tokens
Output
$0.200/ 1M tokens
Cache read
$0.010/ 1M tokens
Try it
Chat with Nemotron 3.5 LightningTerramind API
⌘+Enter to run
Code
import { streamText } from "ai";
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";

const terramind = createOpenAICompatible({
  name: "terramind",
  baseURL: "https://terramind.com/api/v1",
  apiKey: process.env.TERRAMIND_API_KEY!,
});

const result = streamText({
  model: terramind("nemotron-3.5-lightning"),
  prompt: "Why is the sky blue?",
});

for await (const chunk of result.textStream) {
  process.stdout.write(chunk);
}

Want to run this in your app? Grab an API key · Full chat reference

More by NVIDIA
ModelContextInOutCapabilities
NVIDIAnvidia/nemotron-3-super256K$0.150$0.650
reasoning
NVIDIAnvidia/nemotron-3-ultra1M$0.600$2.40
reasoningcaching