Skip to content

2026 · Ongoing

Kubernetes LLM Incident Analyser

An incident-analysis system that collects Kubernetes logs, extracts relevant failure evidence, and generates structured incident reports containing likely root causes, affected components, and remediation suggestions.

  • Python
  • FastAPI
  • Kubernetes
  • Docker
  • Pydantic
  • LLM APIs

Overview

A system designed to investigate whether LLMs can assist with the initial phase of Kubernetes incident analysis.

The pipeline collects pod logs, removes irrelevant noise, extracts potential failure evidence and asks an LLM to produce a validated structured incident report.

Output includes a likely root cause, the affected component, the evidence that supports the conclusion, and suggested remediation.

Approach

Raw Kubernetes logs are noisy. Most of the work is not the model call — it is deciding what the model is allowed to see.

Pipeline from cluster to validated report.

text
Pods
↓
Log collection
↓
Noise reduction
↓
Evidence extraction
↓
LLM analysis
↓
Pydantic validation
↓
Incident report

The model is constrained to a strict schema. If the response does not validate, it is rejected rather than surfaced, because a confidently wrong incident report is worse than no report.

Related writing