licenseMIT statusalpha PRswelcome runs onKubernetes

Self-healing LLM inference, fixed by pull request.

llm-platform is an open-source stack that serves open-weight models with vLLM on Kubernetes. Its AIOps agent turns a firing alert into a diagnosed, reviewable fix PR. The AI gets no cluster write access. Your Git workflow stays the approval step.

$ git clone https://github.com/hritikmunde/llm-platform
kubectl logs deploy/agent
==================================================
ALERT RECEIVED  status=firing  alertname=VLLMDown  deployment=vllm
==================================================
---- GATHERED CONTEXT ----
--- pod vllm-… (phase=Running) [lastTerminated=Error] ---
ValueError: User-specified max_model_len (999999) is
greater than the derived max_model_len (32768) …
---- DIAGNOSIS ----
  root_cause: max_model_len exceeds the model's
              max position embeddings (32768)
  action    : lower_max_model_len
---- PR ----
  PR opened: github.com/hritikmunde/llm-platform/pull/8
==================================================

The real incident behind PR #8, abridged

Get started

From zero to a self-healing cluster

Terraform defines the whole platform, and ArgoCD deploys it from Git. You run three steps.

// prerequisites

  • AWS account with GPU quota (Running On-Demand G and VT instances)
  • Bedrock model access for Claude Haiku
  • aws, terraform, kubectl, helm
  • A GitHub token the agent can use to open PRs on your fork
bash
# 1. provision VPC + EKS + GPU node group
$ git clone https://github.com/hritikmunde/llm-platform && cd llm-platform/terraform
$ terraform init && terraform apply
$ aws eks update-kubeconfig --region us-east-1 --name llm-platform

# 2. install ArgoCD + the monitoring stack
$ kubectl create namespace argocd
$ kubectl apply -n argocd -f https://raw.githubusercontent.com/argoproj/argo-cd/stable/manifests/install.yaml
$ helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
$ kubectl create namespace monitoring
$ helm install monitoring prometheus-community/kube-prometheus-stack -n monitoring

# 3. let ArgoCD deploy the platform from Git
$ cd .. && kubectl apply -f argocd-app.yaml

The GPU node bills by the hour. terraform destroy tears everything down, and because all state lives in Git, the next apply rebuilds it exactly. Full walkthrough in the README.

Real output, not a mockup

vLLM crashed. The agent wrote the fix.

We deliberately pushed an invalid --max-model-len. vLLM refused to start and VLLMDown fired. The agent read the pod logs, asked Bedrock for a diagnosis, and opened this PR. A maintainer merged it and ArgoCD rolled it out.

Merged

agent: fix VLLMDown (lower max-model-len 999999 → 4096) #8

agent-fix-lower_max_model_len-1785182951 → master

View on GitHub ↗

// Automated remediation by the SRE agent

alert
VLLMDown
root cause
User-specified max_model_len (999999) exceeds the model's maximum position embeddings (32768), causing vLLM initialization to fail.
action
lower_max_model_len
reason
Reducing it to fit the model's limits lets the pod start successfully.
manifests/vllm.yaml
args:
- "--model"
- "Qwen/Qwen2.5-0.5B-Instruct"
- "--max-model-len"
- - "999999"
+ - "4096"
- "--gpu-memory-utilization"
Diagnosed via AWS Bedrock (Claude Haiku) · edit made by deterministic code +1 −1

How it works

The self-healing loop

The parts are standard cloud-native tools. The agent connects them, and Git sits between the agent and the cluster.

01

Detect

Prometheus rules catch VLLMDown and CrashLoopBackOff. Alertmanager sends a webhook to the agent.

02

Gather

The agent pulls pod state and recent logs with a read-only ServiceAccount.

03

Diagnose

Bedrock returns {root_cause, action, reason}, choosing from a fixed list of actions.

04

Propose

Deterministic code edits the manifest and opens a PR that includes the diagnosis.

05

Reconcile

A human merges. ArgoCD syncs the cluster and the platform recovers.

safety · bounded autonomy

The LLM chooses. It never writes.

The model returns one action from a vetted list. Plain code makes the edit, so you can predict every possible diff in advance.

lower_gpu_memory_utilizationlower_max_model_lenrollback_deploymentno_action_needed
safety · human gate + isolation

Nothing ships without a merge

The PR is the approval step, and its description holds the full reasoning. Diagnosis runs on Bedrock, outside the cluster, so the agent stays up even when the served model goes down.

Extend it

Teach it a new fix in ~10 lines

Each capability is a small, reviewable addition. Contributing a new failure mode is the easiest way to get involved.

  1. Name the action in ALLOWED_ACTIONS. That list is all the LLM can choose from.
  2. Write the edit in apply_fix_to_manifest(). Keep it one bounded, deterministic change.
  3. Add the signal, if the failure needs one, in manifests/alert-rules.yaml.
Read the contributing guide →
agent/app.py
# The LLM only CHOOSES one of these.
ALLOWED_ACTIONS = [
    "lower_gpu_memory_utilization",
    "lower_max_model_len",
    "rollback_deployment",
    "no_action_needed",
+   "raise_memory_limit",
]

def apply_fix_to_manifest(text, action):
    # ...existing actions...
+   if action == "raise_memory_limit":
+       m = re.search(r'(memory:\s*")(\d+)(Gi")', text)
+       if m:
+           new = str(int(m.group(2)) * 2)
+           return (text[:m.start()] + m.group(1) + new
+                   + m.group(3) + text[m.end():],
+                   f"raise memory {m.group(2)}Gi -> {new}Gi")

Architecture

Infrastructure for AI, and AI for infrastructure

A tainted GPU node group serves the model. A CPU node group runs observability, the chat UI, the agent, and ArgoCD. All of it is defined in Terraform and Kubernetes manifests.

Architecture diagram of llm-platform on Amazon EKS
Cluster topology on Amazon EKSclick to enlarge
vLLMOpenAI-compatible serving
Amazon EKSCPU + GPU (g4dn / T4) nodes
TerraformVPC, cluster, IAM, ECR
ArgoCDGitOps auto-sync & self-heal
Prometheus + GrafanaTTFT, tokens/sec, GPU cache
Python agentFlask · k8s API · PyGithub
AWS BedrockClaude Haiku via IRSA
Open WebUIChat frontend

Project status

Where it stands today

The core loop works end to end on real infrastructure. Here's what's solid and what isn't yet.

● Works today

  • vLLM serving on EKS GPU nodes, deployed by ArgoCD
  • Prometheus metrics and alerts for pod-down and crash-loop
  • Alert → Bedrock diagnosis → auto-opened, merge-ready PR
  • Automatic fixes for bad GPU memory and context-length config

● Known limitations

  • AWS only for now (EKS + Bedrock)
  • Single model, single GPU replica
  • rollback_deployment is diagnosed but not yet automated
  • No tagged releases yet. Track master

Community

Get involved

This is an early project, so every issue and PR has a big influence on where it goes.