WHAT IT IS
Together API is the inference interface offered by Together AI, introduced in 2022, for running open-weights models on managed clusters. Beyond inference it also offers services such as fine-tuning. It addresses the case where a team wants to call open models through an API without building its own GPU cluster.
Why it matters
It let open-weights models be called much like closed APIs, narrowing the gap between open weights and a usable service.
Key specs
- Capabilities
- Inference and fine-tuning for open models
- Interface
- OpenAI-style calling convention
Related concepts
Inference Optimization & Serving
Training happens once; inference happens a billion times a day — and serving is torn between fast first tokens and high throughput, which usually pull against each other
Training & Inference Infrastructure
Memory decides how large a model you can train, communication how long it takes — raw compute is rarely the bottleneck
Comparable products
Replicate
2019Run community-uploaded models through an API
Ollama
2023Run open models locally with one command
Hugging Face Hub
2016The open clearing house for models and datasets
vLLM
2023A high-throughput inference and serving engine for LLMs