> For the complete documentation index, see [llms.txt](https://docs.cloudeka.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.cloudeka.ai/deka-gpu/deka-gpu-autoscaling/keda-autoscalling/example-autoscaling-vllm-with-keda-based-on-gpu-kv-cache-usage/introduction.md).

# Introduction

This section describes how to configure KEDA (Kubernetes Event-Driven Autoscaler) to automatically scale vLLM deployments based on GPU KV cache utilization. Autoscaling ensures that resources are used efficiently and that the system can handle changes in workload dynamically. The primary metric used for autoscaling in this setup is GPU KV cache usage.
