Deepgram
Deploy the latest Deepgram speech-to-text on Together AI. Accurate, real-time transcription co-located with your LLM and TTS for end-to-end voice latency under 500ms.

Why Deepgram on Together AI?
Designed for production workloads that need consistent performance and operational control.
Production-grade real-time transcription
Deepgram provides accurate, cost-effective speech-to-text, in real time or batch, cloud or self-hosted. It is the listening layer for voice agents, contact centers, and transcription pipelines.
One co-located voice stack
With native Deepgram integration, transcription runs on the same platform as your model and speech output. You manage one API, one vendor, and one bill instead of stitching separate STT, LLM, and TTS providers together.
Enterprise-ready from day one
SOC 2 Type II certified, HIPAA compliant, and deployed on US-based infrastructure. Full model ownership with no data retention by default.
Meet the Deepgram family
Explore top-performing models across text, image, video, code, and voice.
Deployment options
Run models using different deployment options depending on latency needs, traffic patterns, and infrastructure control.
A fully managed real-time or batch inference API with access to dozens of the most popular AI models.
Best for
Reserved token capacity with SLA guarantees. Priced in PTUs, a normalized throughput unit.
Best for
An inference endpoint backed by reserved, isolated compute resources and Together AI inference research.
Best for
Run inference with your own engine and model on fully-managed, scalable infrastructure.
Best for