Two leaders. One platform.
Together AI and NVIDIA combine Together AI’s fastest, research-optimized AI platform with the NVIDIA full-stack AI Factory platform to help teams move from experimentation to production at scale.

The AI Native Cloud, built on the NVIDIA accelerated computing platform
The most advanced AI workloads demand more than bare-metal compute alone. Together AI brings inference optimization, developer experience, and enterprise trust directly onto NVIDIA accelerated computing to help your team ship faster — without managing the underlying infrastructure.
Purpose-built for AI-native companies that need reliability, speed, and control at every stage of the model lifecycle.
NVIDIA provides the best full-stack AI factory platform spanning accelerated computing, open models, networking, storage, reference architectures, and software — designed to deliver high performance and efficient AI training and inference at scale.
Everything your AI team needs to move faster
Through extreme co-design, Together AI and NVIDIA provide the lowest cost-per-token on leading open models, unlocking higher performance on the NVIDIA accelerated computing platform.
High-performance inference
Together Inference runs on latest-gen NVIDIA accelerated compute with custom kernels, adaptive speculative decoding, and intelligent quantization to deliver industry-leading tokens/sec at scale.
Private & dedicated deployments
Run models on dedicated NVIDIA AI infrastructure in Together's global data centers. Zero Data Retention policy and SOC II compliance for privacy-sensitive production workloads.
200+ models, one API
Access the broadest catalog of open models — including Nemotron, DeepSeek, Kimi, MiniMax, Qwen, GLM, Gemma, GPT OSS, and more — all running on NVIDIA AI infrastructure via a single OpenAI-compatible endpoint.
Built for every AI workload.
Serve 10T+ tokens per day with sub-100ms time to first token (TTFT). Using our cutting-edge research paired with NVIDIA inference software, NVIDIA TensorRT, and NVIDIA Dynamo to maximize throughput and achieve lower cost per token.
Access NVIDIA Nemotron models for reasoning, agentic AI applications, and speech recognition — optimized and served via Together AI’s inference API.
Fine-tune your models to your domain. Domain-specific fine-tuning, reinforcement learning (RLHF), and direct preference optimization (DPO) on the NVIDIA accelerated computing platform. From LoRA adapters to full-model updates, you retain control of your weights and deployment.



Production ready from day one
Privacy-sensitive U.S. organizations and global enterprises choose Together AI because security and compliance aren’t afterthoughts, they’re built into every layer of the stack.
- SOC II Type 2 compliant
Independently audited security controls covering availability, confidentiality, and processing integrity.
- Zero Data Retention
Prompts and completions are never stored, logged, or used for training. Your data stays yours.
- Dedicated infrastructure
Isolated NVIDIA AI infrastructure for customers who need multi-tenant environments, network isolation, and custom SLAs.
- Global data centers
Multi-region deployments across North America, Europe, and Asia with data residency and zero-trust controls.

