SGLang
Nemotron 3.5 Lightning cookbook support was merged on 2026-08-11.
This reference covers the open Lightning checkpoint and its confirmed companion repositories. It separates vendor deployment guidance from measured RunInfra evidence.
NVIDIA Nemotron 3.5 Lightning 30B-A3B is the vendor reference for nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16, as of Aug 12, 2026. Scale: 30B Total / 3B Active; Context: Up to 1M tokens (for single H100 deployment, we use 256K), as of Aug 12, 2026. RunInfra serves this model on the Model APIs with measured numbers published at Nemotron 3.5 Lightning 30B on Model APIs; no measured optimization package is published for it, and every figure below belongs to its named source, cited and dated.
RunInfra Model APIs: $0.05 input, $0.01 cached input, $0.15 output per million tokens.
| Group | Fact | Source-cited display value | Source and date |
|---|---|---|---|
| Identity | Release identity | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | huggingface.coRetrieved |
| Identity | Release | Nemotron 3.5 Lightning; August 11, 2026 | blogs.nvidia.comRetrieved |
| Identity | Family position | an expansion of the Nemotron 3 model family, described by the vendor as the highest-efficiency model in its class for long-running agentic AI workloads | blogs.nvidia.comRetrieved |
| Identity | NIM API id | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | docs.api.nvidia.comRetrieved |
| Architecture | Scale | 30B Total / 3B Active | huggingface.coRetrieved |
| Architecture | Architecture | Mixture-of-Experts Hybrid (Mamba + Transformer) | huggingface.coRetrieved |
| Architecture | Layer design | interleaved Mamba-2 and MoE layers, along with select Attention layers | huggingface.coRetrieved |
| Architecture | Training design | Nemotron-3-Lightning + Multi-Token Prediction (MTP) | huggingface.coRetrieved |
| Architecture | Pre-training | pre-trained with over 20T tokens | huggingface.coRetrieved |
| Context | Context | Up to 1M tokens (for single H100 deployment, we use 256K) | huggingface.coRetrieved |
| Context | Maximum output | Not published by the vendor. | huggingface.coRetrieved |
| Modalities | Modality | text | huggingface.coRetrieved |
| Modalities | Reasoning | Configurable on/off via chat template (enable_thinking=True/False) | huggingface.coRetrieved |
| Modalities | Structured use | tool calling, instruction following, structured outputs | huggingface.coRetrieved |
| Modalities | Languages | English (and coding languages), Spanish, French, German, Italian, Japanese | huggingface.coRetrieved |
| License | License | OpenMDW License Agreement, version 1.1 | openmdw.aiRetrieved |
| License | License distinction | This is not the NVIDIA Open Model License. | openmdw.aiRetrieved |
| Pricing | NVIDIA API price: NVIDIA pricing | Not published by NVIDIA. | docs.api.nvidia.comRetrieved |
| Pricing | NVIDIA API price: OpenRouter paid input | $0.05 per 1M input tokens | openrouter.aiRetrieved |
| Pricing | NVIDIA API price: OpenRouter paid output | $0.20 per 1M output tokens | openrouter.aiRetrieved |
| Pricing | NVIDIA API price: OpenRouter free endpoint | Free endpoint listed. | openrouter.aiRetrieved |
| Availability | Reference weights | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | huggingface.coRetrieved |
| Availability | Companion repositories | NVFP4, NVFP4-DSpark, and NVFP4-DFlash repositories are available. | huggingface.coRetrieved |
| Availability | Single-device guidance | 1x H100 80GB (or 1x A100 80GB) for 256K context | huggingface.coRetrieved |
| Availability | Full-context guidance | 8x H100 - TP8 + expert parallel for 1M context | huggingface.coRetrieved |
| Availability | Hardware families | Blackwell (GB200, B200), Hopper (H100, H200), Ampere (A100) | huggingface.coRetrieved |
| Availability | Reference weights | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 | huggingface.coRetrieved |
| Availability | Base weights | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Base-BF16 | huggingface.coRetrieved |
| Availability | Production quantized weights | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 | huggingface.coRetrieved |
| Availability | Speculative decoding draft variant | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark | huggingface.coRetrieved |
| Availability | Speculative decoding draft variant | nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DFlash | huggingface.coRetrieved |
"This release follows Nemotron 3 Nano and reflects NVIDIA's commitment to continually improving open models for greater accuracy and speed."
blogs.nvidia.comRetrieved
These results are vendor-claimed, not independently measured by RunInfra.
| Benchmark | Vendor-claimed value | Source and date |
|---|---|---|
| MMLU Pro | 81.94 | huggingface.coRetrieved |
| GPQA Diamond | 75.44 | huggingface.coRetrieved |
| SWE-bench Verified | 51.56 | huggingface.coRetrieved |
| HLE | 11.72 | huggingface.coRetrieved |
Listed rows have a cited upstream support signal. They are not RunInfra measurements.
Nemotron 3.5 Lightning cookbook support was merged on 2026-08-11.
BF16 recipes were merged on 2026-08-12.
The vendor pins a specific vLLM container version on the model card and a merged pull request fixed the model family's multi-token prediction path; no release note names 3.5 Lightning yet.
RunInfra serves this model on the Model APIs, where its measured serving numbers are published. No measured optimization package is published for it yet; when one is, the record will state throughput, latency, memory, quality, serving conditions, and reproducible evidence.
As of Aug 12, 2026, Release identity: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16.
As of Aug 12, 2026, Reference weights: nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16.
As of Aug 12, 2026, Single-device guidance: 1x H100 80GB (or 1x A100 80GB) for 256K context.
Use a workspace API key and pay for input, cached input, and output tokens.
View Model APIs