Build and operate the distributed systems that power InterpretAI's AI infrastructure, from reliable cluster orchestration to high-throughput platform services.
InterpretAI builds infrastructure for evaluating, interpreting, and improving AI systems. We care about reliability, clear engineering judgment, and systems that keep working under real-world load.
We are looking for a Cluster Engineer to help design, build, and operate the distributed systems behind our AI platform. This role sits close to the core infrastructure: scheduling workloads, improving reliability, managing compute resources, and building the platform primitives that let research and product teams move quickly.
You will work across backend services, orchestration layers, observability, and deployment systems. The ideal candidate is comfortable reasoning about distributed systems, debugging production issues, and turning operational pain into durable engineering improvements.
You will help shape the infrastructure foundation for a company working on hard AI systems problems. This is a high-ownership role with room to define architecture, improve developer velocity, and build systems that matter.