Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS ES IST
Ollama is an open-source tool released in 2023 that lets a user download and run open language models locally with a single command. It bundles quantized weights, the inference runtime and model management, built on implementations such as llama.cpp. It addresses the fiddly environment and dependency setup of deploying open models on a local machine.
Warum es wichtig ist
It folded weights, runtime and model management into one command, turning "run a model on my machine" from an engineering task into an everyday action.
Wichtige Eckdaten
- Backend
- Built on implementations such as llama.cpp
- Platforms
- macOS, Linux, Windows
- Distribution
- Quantized weights managed with the model
Verwandte Konzepte
Inferenz-Optimierung und Serving
Training passiert einmal, Inferenz milliardenfach am Tag – und erstes Token und Durchsatz stehen oft im Konflikt
Modellkompression
Ein Modell kleiner, schneller und günstiger machen, ohne Genauigkeit zu verlieren – aber meist nur zwei von drei zugleich
Vergleichbare Produkte
vLLM
2023Hochdurchsatz-Engine für Inferenz und Serving von LLMs
Hugging Face Hub
2016Die Sammelstelle für offene Modelle und Datensätze
Together API
2022Eine Inferenz-API für offene Modelle
LangChain
2022Modelle, Werkzeuge und Retrieval verketten