AI Fleet Platform Software Engineer
Cerebras Systems, Inc. · Sunnyvale, CA; Toronto, CAN · hybrid
You will build and operate software that manages large fleets of AI clusters. You will create services, integrations, operational tools, and user-facing applications that help operators monitor health, capacity, performance, and incidents. You will lead projects from design through production, automate operational workflows, and improve reliability as the fleet grows.
Responsibilities- Build and operate software for managing large fleets of AI clusters
- Provide operators with actionable views of cluster health, capacity, performance, and issues
- Develop services and integrations across infrastructure systems
- Automate incident investigation and service-restoration workflows
- Design reliable systems that withstand component and site failures
- Gather platform-user needs and make practical product and engineering decisions
- Lead projects from design through production and use operational feedback to improve them
Requirements- 12+ years of industry experience building and operating production software for distributed systems or large-scale infrastructure
- Strong Go or Python skills
- Experience designing services and APIs
- Expertise in control planes, fleet management systems, or operational platforms
- Experience with Linux, containers, Kubernetes, and distributed-system failures
- Experience designing for asynchronous work, retries, and partial failures
- Experience with event streaming, workflow automation, or time-series telemetry
- Strong judgment in reliability, security, and observability
- Ability to lead ambiguous projects and collaborate across engineering and operations teams