
ARCHITECT
FEATURED ANALYSIS
Running a self-hosted LLM in Kubernetes with vLLM
BY
Matt Kereczman, LINBIT
SOURCE
Cloud Native Computing Foundation
DATE
READ
1 min read
Running large language model (LLM) workloads in-house is one of several patterns teams adopt alongside managed API services. Managed API services are convenient and well suited to many workloads. Self-hosting is a …
Running large language model (LLM) workloads in-house is one of several patterns teams adopt alongside managed API services. Managed API services are convenient and well suited to many workloads. Self-hosting is a complementary option that some…