Build and operate the self-hosted AI inference platform — Kubernetes, NVIDIA H100 infrastructure, vLLM, observability, GitOps, and production reliability.
Context
We are building a self-hosted LLM inference platform running on NVIDIA H100 GPUs and bare-metal Kubernetes. The platform provides vLLM inference workers and an OpenAI-compatible API for internal services.
There is no managed cloud Kubernetes layer — the engineer will work directly with the infrastructure and own the platform from cluster bootstrap through production operations.
What you will work on * Bootstrap and maintain self-hosted Kubernetes clusters with HA control plane, etcd backup/restore and upgrades. * Set up and maintain NVIDIA GPU infrastructure, including GPU Operator, DCGM and the CUDA/container runtime stack. * Configure GPU sharing and scheduling using MPS or time-slicing. * Deploy and operate vLLM inference workers with zero-downtime rolling updates and graceful connection draining. * Build Prometheus/Grafana monitoring and alerting for inference SLIs such as TTFT, queue depth, KV-cache utilisation and rejection rate. * Implement IaC for bare-metal provisioning and cluster bootstrap using Terraform or Pulumi. * Set up GitOps workflows with ArgoCD or Flux for controlled rollouts and rollbacks. * Configure API Gateway, token-aware rate limiting, NetworkPolicy and mTLS. * Build automation and runbooks for GPU-specific production incidents.
Must-have * 5+ years in DevOps, Platform Engineering or SRE with production Kubernetes. * Hands-on experience with self-hosted Kubernetes (kubeadm, RKE2 or k3s), including bootstrap, upgrades, etcd and HA. * Practical experience with NVIDIA GPUs in Kubernetes, GPU Operator/Device Plugin and the CUDA driver stack. * Strong understanding of GPU sharing and scheduling: MPS, time-slicing and multi-process workloads. * Strong Linux/system administration skills, including NUMA, huge pages, CPU management and bare-metal troubleshooting. * Experience with zero-downtime deployments of long-running or stateful workloads. * Strong Prometheus and Grafana skills, including recording and alerting rules. * Experience with Terraform or Pulumi. * Bash/Python scripting for automation, health checks and monitoring. * Basic understanding of LLM inference metrics and concepts: TTFT, token throughput, KV-cache and batching.
Nice-to-have * Experience with vLLM, Triton Inference Server, KServe or similar. * Production GitOps experience with ArgoCD or FluxCD. * Experience with Kong, Envoy, APISIX or similar API gateways. * Experience with Vault or External Secrets Operator. * CKA, CKS or NVIDIA certifications. * Experience benchmarking GPU workloads under mixed-model load.
Ways of working * Remote, distributed engineering team. * Production-focused environment with ownership of the infrastructure and inference stack. * The role involves building the platform from the ground up and keeping it reliable in production.
Important: As this is a Germany-based project, we are primarily seeking candidates based in Western Ukraine, with Vinnytsia and Lviv being our preferred locations. Other secure western-region locations can be considered case by case.