Skip to content

Projects

Things I've built to understand systems beyond the tutorial.

These projects are not intended to be isolated demos. I use them to explore infrastructure, deployment, networking, observability, reliability and the trade-offs involved in operating software.

Production AI Platform

2026 · Ongoing

  • AWS
  • EKS
  • Terraform
  • Kubernetes
  • Helm
  • Argo CD
  • GitHub Actions
  • Prometheus
  • OpenTelemetry

A production-style platform for deploying and operating AI applications on Kubernetes, with infrastructure as code, GitOps delivery, observability, autoscaling, and reliability controls.

  • Infrastructure as code
  • Kubernetes application architecture
  • GitOps deployments
  • Secrets and configuration
  • Autoscaling
  • Observability
  • Failure recovery
  • SLOs
  • Infrastructure cost
  • Python
  • FastAPI
  • Kubernetes
  • Docker
  • Pydantic
  • LLM APIs

An incident-analysis system that collects Kubernetes logs, extracts relevant failure evidence, and generates structured incident reports containing likely root causes, affected components, and remediation suggestions.

  • Log collection and noise reduction
  • Failure evidence extraction
  • Schema-validated LLM output
  • Structured incident reports
  • Human-reviewable remediation

AI Inference Gateway

2026 · Ongoing

  • FastAPI
  • Redis
  • Kubernetes
  • KEDA
  • OpenTelemetry

An OpenAI-compatible inference gateway exploring request routing, rate limiting, caching, autoscaling, observability, and the operational characteristics of serving AI workloads.

  • OpenAI-compatible API
  • Authentication and rate limiting
  • Request routing and queues
  • Semantic caching
  • Autoscaling on queue depth
  • Latency and token throughput
  • Cost visibility
  • Kubernetes
  • Prometheus
  • Grafana
  • Loki
  • Chaos Engineering

A controlled environment for creating infrastructure failures and observing how applications and monitoring systems respond.

  • Pod termination
  • Resource exhaustion
  • Broken dependencies
  • Node failure
  • Increased network latency
  • Bad deployments
  • Application crashes