← Back to portfolio AI · Infrastructure · Kubernetes · LLM

AI Ops Assistant — Infrastructure Incident Diagnostics

A Kubernetes-deployed AI assistant that analyses infrastructure failure scenarios and returns structured root cause analysis — summary, probable causes, verification commands, and remediation steps. Built with FastAPI, NGINX, and a locally hosted LLM via LM Studio. Tested against 25+ real-world failure patterns.

25+Failure scenarios tested
4Output sections per diagnosis
100%Local — no external API
k8sFully containerised deployment
FastAPI Kubernetes (kind) NGINX LM Studio Local LLM Docker Python REST API
Live demo

AI-assisted troubleshooting workflow

What it does

From raw failure input to structured diagnosis

Infrastructure engineers typically spend significant time manually correlating logs, error codes, and symptoms before identifying a root cause. This assistant automates the first-pass diagnostic step — you submit a structured failure description, and the system returns an actionable analysis in seconds.

The output is structured into four consistent sections, making it usable in runbooks, incident reports, and escalation summaries.

── Incident Diagnosis Output ────────────────────── SUMMARY Service unavailable due to pod crashloop — likely caused by a misconfigured environment variable pointing to an unreachable DB host. PROBABLE CAUSES 1. DATABASE_HOST env var set to localhost instead of service FQDN 2. PostgreSQL service not running or not exposed on expected port 3. Network policy blocking inter-namespace traffic VERIFICATION COMMANDS kubectl describe pod api-pod -n production kubectl logs api-pod --previous kubectl get svc -n database kubectl exec -it api-pod -- curl postgres-svc:5432 REMEDIATION Update DATABASE_HOST to postgres-svc.database.svc.cluster.local Redeploy the pod. Verify readiness probe passes before routing traffic.
Test coverage

Failure categories tested

The assistant was validated across 25+ scenarios spanning four failure domains:

Pod CrashLoopBackOff Image pull errors OOMKilled containers DNS resolution failure Network policy block Service unreachable DB connection refused API 502 / 503 errors Flux reconciliation error Resource quota exceeded Misconfigured env vars Missing secrets Probe misconfiguration Node not ready Ingress routing failure
Architecture

System overview

Architecture diagram for the AI Ops Assistant project

Frontend Interface

User-facing dashboard for submitting structured failure descriptions and viewing the diagnostic output — summary, causes, commands, and remediation steps.

FastAPI Backend

Exposes /api/analyze endpoint. Receives JSON payloads describing failure scenarios, constructs structured prompts, forwards to the LLM, and returns parsed responses.

Kubernetes Deployment

Both the frontend and backend run as containerised Kubernetes workloads (kind cluster), with Deployments and Services wiring the application flow together.

Local LLM via LM Studio

A locally hosted language model serves inference — no external API calls, no data leaving the environment. Chosen for privacy, cost, and latency control.

NGINX Reverse Proxy

Routes frontend requests to the FastAPI backend, handling path-based proxying within the Kubernetes cluster cleanly.

Prompt Engineering

Prompts are structured to enforce consistent four-part output — ensuring the LLM always returns actionable, formatted diagnostics rather than freeform text.

Build process

How I built it

  1. Local LLM setup: Installed LM Studio, selected a model suitable for structured technical output, and validated that it could produce consistent formatted responses.
  2. Prompt design: Engineered prompts that force four-section output — summary, probable causes, verification commands, remediation — regardless of input format.
  3. FastAPI backend: Built the /api/analyze endpoint to accept JSON failure payloads, inject them into the structured prompt, call the LM Studio API, and return parsed output.
  4. Frontend: Built a simple interface for submitting scenarios and displaying structured diagnostic sections clearly.
  5. Containerisation: Packaged both frontend and backend as Docker images with appropriate Dockerfiles for Kubernetes compatibility.
  6. Kubernetes deployment: Wrote Deployment and Service manifests for both services; added NGINX as a reverse proxy to route frontend requests to the API.
  7. Testing: Ran 25+ failure scenarios covering network failures, misconfigurations, service issues, and API errors — validating output quality and consistency.
Takeaway

What this demonstrates

  • Practical application of LLM tooling to a real infrastructure problem — not a chatbot, a diagnostic tool
  • Kubernetes deployment of a multi-service application with proper networking and proxying
  • Prompt engineering skills: enforcing structured, operationally useful output from a local model
  • FastAPI API design: typed endpoints, structured request/response, clear separation of concerns
  • Testing discipline: 25+ structured scenarios across four failure domains with validated output quality
← Portfolio Hire me →