Cluster Engineer

Build and operate the distributed systems that power InterpretAI's AI infrastructure, from reliable cluster orchestration to high-throughput platform services.

EngineeringSF OfficeFull-time

About InterpretAI

InterpretAI builds infrastructure for evaluating, interpreting, and improving AI systems. We care about reliability, clear engineering judgment, and systems that keep working under real-world load.

About The Role

We are looking for a Cluster Engineer to help design, build, and operate the distributed systems behind our AI platform. This role sits close to the core infrastructure: scheduling workloads, improving reliability, managing compute resources, and building the platform primitives that let research and product teams move quickly.

You will work across backend services, orchestration layers, observability, and deployment systems. The ideal candidate is comfortable reasoning about distributed systems, debugging production issues, and turning operational pain into durable engineering improvements.

What You Will Do

  • Build and maintain distributed infrastructure for AI workloads and platform services.
  • Improve reliability, scalability, and performance across cluster-backed systems.
  • Design tools and abstractions that make compute easier for internal teams to use.
  • Debug complex production issues across services, queues, schedulers, and infrastructure.
  • Strengthen observability, alerting, and incident response practices.
  • Partner with engineering and research teams to support new workloads safely.

What We Are Looking For

  • Strong experience with distributed systems, backend infrastructure, or platform engineering.
  • Comfort working with orchestration systems, service networking, queues, storage, and observability tooling.
  • Ability to debug complex systems using logs, metrics, traces, and careful reasoning.
  • Experience operating production systems where reliability and performance matter.
  • Clear communication and a bias toward simple, maintainable systems.
  • Familiarity with AI/ML infrastructure, GPU clusters, Kubernetes, or workload schedulers is a plus.

Why Join Us

You will help shape the infrastructure foundation for a company working on hard AI systems problems. This is a high-ownership role with room to define architecture, improve developer velocity, and build systems that matter.