Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE C'EST
Ollama is an open-source tool released in 2023 that lets a user download and run open language models locally with a single command. It bundles quantized weights, the inference runtime and model management, built on implementations such as llama.cpp. It addresses the fiddly environment and dependency setup of deploying open models on a local machine.
Pourquoi il compte
It folded weights, runtime and model management into one command, turning "run a model on my machine" from an engineering task into an everyday action.
Caractéristiques clés
- Backend
- Built on implementations such as llama.cpp
- Platforms
- macOS, Linux, Windows
- Distribution
- Quantized weights managed with the model
Concepts liés
Optimisation et service d’inférence
L’entraînement a lieu une fois, l’inférence des milliards de fois par jour — et le premier token s’oppose souvent au débit
Compression de modèles
Rendre un modèle plus petit, plus rapide et moins cher sans perdre en précision — rarement les trois à la fois
Produits comparables
vLLM
2023Moteur d’inférence et de service à haut débit pour LLM
Hugging Face Hub
2016Le carrefour des modèles et jeux de données ouverts
Together API
2022Une API d’inférence pour les modèles ouverts
LangChain
2022Enchaîner modèles, outils et recherche