Models / NVIDIA
Chat

NVIDIA Nemotron 3.5 Lightning

Fast, customizable open model for high-volume always-on agent workloads

About model

Nemotron 3.5 Lightning is NVIDIA's speed-focused open model for always-on agents, built to handle the high-volume, repetitive, domain-specific steps that make up most of an agent's work. A hybrid Mixture-of-Experts model with 30B total parameters and 3B active, distilled from NVIDIA's frontier Nemotron 3 Ultra and developed with the Nemotron Coalition, it pairs multi-token prediction and DFlash speculative decoding with a context window of up to 1M tokens. NVIDIA projects up to 4x higher throughput and up to 30% faster task completion on high-volume agentic workflows. Trained for popular agent harnesses and designed for post-training, so organizations can adapt it to their own tools, workflows, and policies for domain-specific accuracy. Available on Together AI.

Throughput (NVIDIA-Projected)

Up to 4x

With up to 30% faster task completion on high-volume agent workflows

Total Parameters (3B active)

30B

Hybrid MoE keeping per-step inference cost at small-model levels

Context Window

1M

Long-running, multi-turn agent sessions without losing state

Model key capabilities
  • Fast Task Completion: NVIDIA-projected up to 4x higher throughput and up to 30% faster completion on high-volume, specialized agent steps
  • Built for Agent Harnesses: Trained for popular agent harnesses, with reliability-focused behavior for multi-step workflows
  • Designed for Customization: Open model trained on open datasets, built to be post-trained on an organization's own tools, workflows, and policies for domain accuracy
  • Production-Ready Infrastructure: 99.9% SLA, available on Together AI dedicated infrastructure
  • Model card

    Architecture Overview:
    • Hybrid Mixture-of-Experts architecture with 30B total parameters and 3B active per token
    • Multi-token prediction with DFlash speculative decoding for faster generation
    • Context window of up to 1M tokens for long-running, multi-turn agent sessions
    • Text input, text output

    Training Methodology:
    • Distilled from NVIDIA's frontier Nemotron 3 Ultra
    • Developed with the Nemotron Coalition of AI labs and trained on open datasets
    • Trained for agentic tasks and popular agent harnesses, targeting the high-volume specialized steps in multi-model agent systems

    Performance Characteristics:
    • NVIDIA-projected up to 4x higher throughput and up to 30% faster task completion on high-volume agentic workflows
    • NVIDIA reports strong preliminary results on agent productivity and reliability benchmarks, including PinchBench and AA-Omniscience Non-Hallucination

  • Prompting

    Together AI API Access:
    • Access Nemotron 3.5 Lightning via Together AI APIs using the endpoint nvidia/nemotron-3.5-lightning
    • Authenticate using your Together AI API key in request headers
    • Text in, text out, with a context window of up to 1M tokens for long agent sessions
    • Suited to serving the high-volume, specialized steps in multi-model agent systems
    • Available on Together AI dedicated infrastructure

  • Applications & use cases

    Always-On Agent Systems:
    • Serve the high-volume, repetitive steps in agent pipelines while larger models handle frontier reasoning
    • Run long-lived personal and productivity agents that manage email, calendars, projects, and bookings
    • Keep multi-turn agent state across sessions with the 1M-token context window

    Financial Services Workflows:
    • Extract data from documents and prepare structured summaries at scale
    • Check policy rules and monitor risk signals as dedicated agent steps
    • Serve high request volumes with the model's 3B-active-parameter cost profile

    Security Operations:
    • Enrich alerts, classify incidents, and correlate indicators as always-on agent tasks
    • Query logs and validate controls through structured agent workflows
    • Prepare structured findings for analysts from continuous monitoring streams

    Telecom & Retail Operations:
    • Triage network alarms and answer billing questions with specialized agents
    • Enrich product catalogs and resolve inventory and fulfillment exceptions
    • Handle order, return, and loyalty questions at high volume

Related models
  • Model provider
    NVIDIA
  • Type
    Chat
  • Speed
    High
  • Deployment
    Dedicated
  • Parameters
    30B
  • Activated parameters
    3B
  • Context length
    1M
  • Input modalities
    Text
  • Output modalities
    Text
  • Released
    August 10, 2026
  • Category
    Chat