Senior Infrastructure Engineer
deep-infra · Palo Alto, USA
About DeepInfra
DeepInfra is building the foundation for companies to use modern AI in production — simply, reliably, and at scale. Our team has deep experience building large systems that serve hundreds of millions of users, and we're bringing that same level of rigor to a rapidly evolving AI inference space. Our mission is to make advanced AI available to people and teams everywhere.
We're an early, tight-knit team where you can influence product direction, try bold ideas, and drive meaningful work forward quickly. If you want to join a fast-growing company at a defining moment, we'd love to talk.
DeepInfra is backed by leading investors including A.Capital, Felicis, 500 Global, Georges Harik, Samsung Next, Supermicro, Upper90, Peak6, SVAngel, and NVIDIA.
Why This Role Matters
Demand for AI inference is growing fast, and DeepInfra runs on GPU infrastructure we own and operate. Every new site, rack, and network link we bring online translates directly into capacity our customers can build on — so how quickly, reliably, and cost-effectively we expand our physical footprint is central to how we grow.
This role spans the full lifecycle of that expansion. You'll find and evaluate data center space, negotiate with colocation and hardware vendors, design and build the networks that tie our clusters together, and run the projects that take a site from signed contract to serving production traffic. You'll work closely with Engineering and our co-founders, and you'll be as comfortable reviewing a colo contract as you are on the data center floor with a console cable.
What You'll Do
- Design, build, and operate high-performance data center networks — including spine/leaf fabrics, BGP, transit and peering, and network security — across multiple colocation sites.
- Own site bring-ups end to end: low- and high-level designs, capacity planning, rack and power layouts, cabling plans, hardware installation, turn-up, and handoff to production.
- Operate and improve the underlying infrastructure that our compute platform runs on, with strong Linux fundamentals, monitoring, and automation.
- Run infrastructure projects across multiple vendors and sites, managing timelines, dependencies, shipment and logistics coordination, and risk.
- Travel to data centers to install, connect, and troubleshoot hardware alongside remote-hands and vendor teams.
- Serve as the primary technical point of contact for data center vendors and for customers whose infrastructure we manage.
- Build repeatable playbooks, standards, and automation that make every new site faster to stand up than the last.
What You Bring
- Experience: 10+ years building and operating infrastructure for 24/7 production environments, including significant hands-on experience with data center infrastructure and networking.
- Data Center Networking: Deep experience designing and operating high-performance data center networks, including spine/leaf or Clos architectures, BGP, and network security. Experience with EVPN/VXLAN, transit, and peering is strongly preferred.
- Site Deployment: Experience bringing colocation or data center infrastructure online, including network design, rack layouts, cabling, hardware installation, turn-up, and coordination with on-site or remote-hands teams.
- Infrastructure Operations: Strong Linux fundamentals and the ability to troubleshoot across networking, compute, storage, and the physical layer.
- Automation: Strong scripting and automation skills using Python, Bash, Ansible, or similar tools.
- Project Ownership: Proven ability to drive complex infrastructure projects across multiple vendors, dependencies, and locations from planning through production.
- Vendor Management: Experience working directly with data center, network, or hardware vendors. Experience with capacity planning, procurement, or commercial negotiations is a strong plus.
Preferred Qualifications
- Experience deploying GPU infrastructure, including high-speed interconnects such as InfiniBand or RoCE and high-density power and cooling environments.
- On-prem private cloud experience with KVM, Proxmox, OpenStack, Ceph, or similar technologies.
- Kubernetes and containers in production, along with observability tooling such as Prometheus and Grafana.
- Experience managing internet resources at scale, including ASNs, IP prefixes, IRR records, and routing policy with ISPs and IXPs.
- Experience owning hardware procurement, colocation contracts, capacity planning, or infrastructure vendor negotiations.
- Source, evaluate, and secure data center and colocation space across power, cooling, density, connectivity, and cost.
- Lead commercial negotiations with data center, network, and hardware vendors across pricing, contract terms, SLAs, and expansion options.
How We Work
Three traits define the people who thrive here, and this role leans on all three.
Initiative. We take ownership and step in where we can add value. Whether it's starting something new, improving what exists, or helping move ideas forward, we aim to be proactive and thoughtful in how we contribute.
Drive. We're energized by hard problems. Building AI infrastructure is complex, and we lean into that. We care about doing things well, moving fast, and continuously improving — because solving meaningful challenges is what motivates us.
Grit. Things don't always work on the first try — and that's expected. You'll balance negotiations, network design, and hands-on site work at once, often across multiple vendors, time zones, and shifting delivery timelines. We stay persistent, adapt quickly, and learn as we go.