WHAT IT IS
Ollama is an open-source tool released in 2023 that lets a user download and run open language models locally with a single command. It bundles quantized weights, the inference runtime and model management, built on implementations such as llama.cpp. It addresses the fiddly environment and dependency setup of deploying open models on a local machine.
Why it matters
It folded weights, runtime and model management into one command, turning "run a model on my machine" from an engineering task into an everyday action.
Key specs
- Backend
- Built on implementations such as llama.cpp
- Platforms
- macOS, Linux, Windows
- Distribution
- Quantized weights managed with the model
Related concepts
Inference Optimization & Serving
Training happens once; inference happens a billion times a day — and serving is torn between fast first tokens and high throughput, which usually pull against each other
Model Compression
Make a model smaller, faster and cheaper with almost no accuracy loss — but you can usually have only two of the three at once
Comparable products
vLLM
2023A high-throughput inference and serving engine for LLMs
Hugging Face Hub
2016The open clearing house for models and datasets
Together API
2022An inference API for open models
LangChain
2022Chain models, tools and retrieval together