Quick Start
Get up and running with Kthena in minutes! This guide will walk you through deploying your first AI model. We'll install a model from Hugging Face and perform inference using a simple curl command.
No GPUs available? Follow the GPU-Free Quick Start to evaluate Kthena end to end with a mock inference backend on a CPU-only cluster.
For new deployments, use ModelServing together with ModelServer and ModelRoute. The deprecated ModelBooster example is retained below for existing users.
Prerequisites
- Kthena installed on your Kubernetes cluster (see Installation)
- Access to a Kubernetes cluster with
kubectlconfigured - Pod in Kubernetes can access the internet
- volcano is installed.
ModelServing
You can flexibly configure your own self-hosted LLM through ModelServing.
Model Serving Controller is a component of Kthena that provides a flexible and customizable way to deploy LLMs. It allows you to configure your own LLM through ModelServing CRD. ModelServing supports deploying large language models (LLMs) based on roles, with support for gang scheduling and network topology scheduling. It also provides fundamental features such as scaling and rolling updates.
Here is an example of deploying the PD-disaggregation Qwen/Qwen3-0.6B model on GPU using ModelServing. For the complete walkthrough, including the matching ModelServer and ModelRoute resources, see Prefill-Decode Disaggregation with ModelServing (vLLM, NIXL & LMCache).
Step 1: Create a ModelServing Resource Object:
kubectl apply -f https://raw.githubusercontent.com/volcano-sh/kthena/refs/heads/main/examples/model-serving/gpu-pd-disaggregation.yaml
Step 2: Wait for ModelServing to be Ready
After all Pods awaiting deployment have started running, you can run the following command to see the result:
kubectl get pod -owide -l modelserving.volcano.sh/name=vllm-qwen-06b
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
vllm-qwen-06b-0-decode-0-0 1/1 Running 0 2m <pod-ip> <node> <none> <none>
vllm-qwen-06b-0-prefill-0-0 1/1 Running 0 2m <pod-ip> <node> <none> <none>
------------------------------------------
kubectl get modelserving vllm-qwen-06b -o jsonpath='{.status.conditions}' | jq '.'
[
{
"lastTransitionTime": "2025-09-29T08:11:16Z",
"message": "Some groups is progressing: [0]",
"reason": "GroupProgressing",
"status": "False",
"type": "Progressing"
},
{
"lastTransitionTime": "2025-09-29T08:11:21Z",
"message": "All Serving groups are ready",
"reason": "AllGroupsReady",
"status": "True",
"type": "Available"
}
]
Step 3: Send Inference Request
Before you can chat with the LLM, create the matching ModelServer and ModelRoute resources. You can refer to the ModelServer configuration and ModelRoute configuration in the vLLM GPU PD guide.
Then you can use the following command to send a request:
export MODEL="Qwen/Qwen3-0.6B"
export ROUTER_IP=$(kubectl get svc kthena-router -n kthena-system -o jsonpath='{.spec.clusterIP}')
curl -v http://$ROUTER_IP:80/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"'"$MODEL"'","messages":[{"role":"user","content":"Hello"}]}'
ModelBooster
ModelBooster is deprecated in v1.1; use ModelServing, ModelServer, and ModelRoute. Removal is no earlier than v1.5. For new deployments, follow the ModelServing steps above. See the deprecation details for existing deployments.
The ModelBooster CRD creates and manages the underlying serving resources automatically. This example remains available for existing users during the deprecation period.
Step 1: Create a ModelBooster Resource
Create the example model in your namespace (replace <your-namespace> with your actual namespace):
kubectl apply -n <your-namespace> -f https://raw.githubusercontent.com/volcano-sh/kthena/refs/heads/main/examples/model-booster/Qwen2.5-0.5B-Instruct.yaml
Content of the Model:
# ModelBooster is deprecated in v1.1; use ModelServing, ModelServer, and ModelRoute.
# Removal is no earlier than v1.5.
apiVersion: workload.serving.volcano.sh/v1alpha1
kind: ModelBooster
metadata:
name: demo
spec:
backend:
name: "backend1"
type: "vLLM"
modelURI: "hf://Qwen/Qwen2.5-0.5B-Instruct"
cacheURI: "hostpath://tmp/cache"
replicas: 1
env:
- name: "HF_ENDPOINT" # Optional: set to https://hf-mirror.com if you have network issues
value: "https://huggingface.co"
workers:
- type: "server"
image: "vllm/vllm-openai-cpu:latest" # This model will run on CPU, for more details visit https://docs.vllm.ai/en/stable/getting_started/installation/cpu.html#pre-built-images
replicas: 1
pods: 1
config:
served-model-name: "Qwen2.5-0.5B-Instruct"
dtype: float16
max-model-len: 2048
max-num-batched-tokens: 512
max-num-seqs: 1
block-size: 128
gpu-memory-utilization: 0.35
enable-prefix-caching: ""
enforce-eager: ""
resources:
limits:
cpu: "4"
memory: "8Gi"
requests:
cpu: "2"
memory: "4Gi"
Step 2: Wait for Model to be Ready
Wait model condition Active to become true. You can check the status using:
kubectl get modelBooster demo -o jsonpath='{.status.conditions}' -n <your-namespace>
And the status section should look like this when the model is ready:
[
{
"lastTransitionTime": "2025-09-05T02:14:16Z",
"message": "Model initialized",
"reason": "ModelCreating",
"status": "True",
"type": "Initialized"
},
{
"lastTransitionTime": "2025-09-05T02:18:46Z",
"message": "Model is ready",
"reason": "ModelAvailable",
"status": "True",
"type": "Active"
}
]
Step 3: Perform Inference
You can now perform inference using the model. Here's an example of how to send a request:
curl -X POST http://<model-route-ip>/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "demo",
"messages": [
{
"role": "user",
"content": "Where is the capital of China?"
}
],
"stream": false
}'
Use the following command to get the <model-route-ip>:
kubectl get svc kthena-router -o jsonpath='{.spec.clusterIP}' -n kthena-system
This IP can only be used inside the cluster. If you want to chat from outside the cluster, you can use the EXTERNAL-IP
of kthena-router after you bind it.