Running Ollama on Azure Kubernetes Service
Summary
This article explains how to run Ollama on Azure Kubernetes Service with a GPU-enabled node pool. It walks through the Terraform setup, GPU scheduling requirements, and the Nvidia device plugin needed to make the deployment work. It also shows how to deploy the Ollama Helm chart and expose the service through a public hostname for testing. The piece closes by comparing this simpler setup with more advanced GPU options such as Nvidia GPU Operator and Triton Inference Server.
Classifications
industries
HealthTech
applications
Accounting and Taxes
AskAI Classifications
Labels
AI Software
Developer Tools
MLOps