AI-assisted troubleshooting workflow
From raw failure input to structured diagnosis
Infrastructure engineers typically spend significant time manually correlating logs, error codes, and symptoms before identifying a root cause. This assistant automates the first-pass diagnostic step — you submit a structured failure description, and the system returns an actionable analysis in seconds.
The output is structured into four consistent sections, making it usable in runbooks, incident reports, and escalation summaries.
Failure categories tested
The assistant was validated across 25+ scenarios spanning four failure domains:
System overview
Frontend Interface
User-facing dashboard for submitting structured failure descriptions and viewing the diagnostic output — summary, causes, commands, and remediation steps.
FastAPI Backend
Exposes /api/analyze endpoint. Receives JSON payloads describing failure scenarios, constructs structured prompts, forwards to the LLM, and returns parsed responses.
Kubernetes Deployment
Both the frontend and backend run as containerised Kubernetes workloads (kind cluster), with Deployments and Services wiring the application flow together.
Local LLM via LM Studio
A locally hosted language model serves inference — no external API calls, no data leaving the environment. Chosen for privacy, cost, and latency control.
NGINX Reverse Proxy
Routes frontend requests to the FastAPI backend, handling path-based proxying within the Kubernetes cluster cleanly.
Prompt Engineering
Prompts are structured to enforce consistent four-part output — ensuring the LLM always returns actionable, formatted diagnostics rather than freeform text.
How I built it
- Local LLM setup: Installed LM Studio, selected a model suitable for structured technical output, and validated that it could produce consistent formatted responses.
- Prompt design: Engineered prompts that force four-section output — summary, probable causes, verification commands, remediation — regardless of input format.
- FastAPI backend: Built the
/api/analyzeendpoint to accept JSON failure payloads, inject them into the structured prompt, call the LM Studio API, and return parsed output. - Frontend: Built a simple interface for submitting scenarios and displaying structured diagnostic sections clearly.
- Containerisation: Packaged both frontend and backend as Docker images with appropriate Dockerfiles for Kubernetes compatibility.
- Kubernetes deployment: Wrote Deployment and Service manifests for both services; added NGINX as a reverse proxy to route frontend requests to the API.
- Testing: Ran 25+ failure scenarios covering network failures, misconfigurations, service issues, and API errors — validating output quality and consistency.
What this demonstrates
- Practical application of LLM tooling to a real infrastructure problem — not a chatbot, a diagnostic tool
- Kubernetes deployment of a multi-service application with proper networking and proxying
- Prompt engineering skills: enforcing structured, operationally useful output from a local model
- FastAPI API design: typed endpoints, structured request/response, clear separation of concerns
- Testing discipline: 25+ structured scenarios across four failure domains with validated output quality