Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО
Ollama is an open-source tool released in 2023 that lets a user download and run open language models locally with a single command. It bundles quantized weights, the inference runtime and model management, built on implementations such as llama.cpp. It addresses the fiddly environment and dependency setup of deploying open models on a local machine.
Почему стоит запомнить
It folded weights, runtime and model management into one command, turning "run a model on my machine" from an engineering task into an everyday action.
Ключевые характеристики
- Backend
- Built on implementations such as llama.cpp
- Platforms
- macOS, Linux, Windows
- Distribution
- Quantized weights managed with the model
Связанные концепции
Оптимизация инференса и обслуживание
Обучение случается один раз, вывод — миллиарды раз в день; быстрый первый токен и высокая пропускная способность обычно тянут в разные стороны
Сжатие моделей
Сделать модель меньше, быстрее и дешевле почти без потери точности — но три сразу удаётся редко
Аналоги
vLLM
2023Высокопроизводительный движок вывода и обслуживания LLM
Hugging Face Hub
2016Открытый узел моделей и наборов данных
Together API
2022API вывода для открытых моделей
LangChain
2022Связывание моделей, инструментов и поиска