Mistral
7.3B model surpassing Llama 2 13B, nearing CodeLlama 7B on code, with GQA for speed and SWA for efficient long-sequence handling.
7.3B model surpassing Llama 2 13B, nearing CodeLlama 7B on code, with GQA for speed and SWA for efficient long-sequence handling.

Launching soon
Want to be notified when the model is available on Together AI?
You're on the list
We'll email you when the endpoint goes live. Until then, 200+ models are ready to call.