Skip to main content

Overview

Together AI provides access to 100+ open-source AI models with blazing-fast inference, competitive pricing, and support for the latest community models. Perfect for developers who want to leverage open-source models at scale. Base URL: https://api.together.xyz

Supported Features

  • ✅ Chat Completions
  • ✅ Completions
  • ✅ Streaming
  • ✅ Embeddings
  • ✅ Function Calling
  • ✅ Vision (select models)
  • ✅ Image Generation
  • ❌ Fine-tuning (via Together platform)

Quick Start

Chat Completions

Streaming

Meta Llama

Mistral & Mixtral

Qwen

Image Generation

Embeddings

Together AI excels at:
  • Open-source models - Access 100+ community models
  • Fast inference - Optimized infrastructure
  • Latest models - Quick addition of new releases
  • Cost-effective - Competitive pricing
  • Developer-friendly - Simple API, great docs

Configuration Options

Advanced Features

Function Calling

Vision Models

Image Generation

Embeddings

Completions (Legacy)

Fallback Configuration

Fallback to OpenAI:

Load Balancing

Balance across different Llama models:

Error Handling

Best Practices

  1. Choose right model size - Balance cost vs capability
  2. Use Turbo models - Optimized for speed
  3. Enable streaming - Better user experience
  4. Leverage function calling - Available on many models
  5. Try vision models - For multimodal tasks
  6. Use embeddings - For semantic search
  7. Monitor costs - Different models have different pricing
  8. Test models - Performance varies by use case

Model Categories

By Size

  • Large (100B+): Best quality, higher cost
  • Medium (30-100B): Balanced performance
  • Small (7-30B): Fast, cost-effective

By Type

  • Chat/Instruct: Conversational models
  • Code: Specialized for coding
  • Vision: Multimodal capabilities
  • MoE: Mixture of Experts for efficiency

Pricing

Together AI offers competitive pricing for open models:

Together AI Pricing

View detailed pricing for all Together AI models

Anyscale

Another open models platform

Groq

Ultra-fast inference

Function Calling

Advanced function calling

Load Balancing

Balance across models