Lead Devops Engineer
Way2B1 · Florida, Florida, United States
Our Company and the Role
This is a hands-on Lead DevOps Engineer role responsible for designing, operating, and evolving a highly available, multi-tenant platform on AWS. You will work closely with software engineering to deploy, operate, and scale production systems while driving improvements in reliability, automation, and performance.
You will lead a small DevOps team of 1–2 engineers while remaining primarily hands-on. You will serve as the team's technical lead, set standards and priorities, mentor and develop engineers, and take ownership of the most difficult infrastructure and production problems.
You will also help introduce and operationalize AI/LLM capabilities within the platform.
This role reports to the Chief Technology Officer (CTO) and is part of the engineering leadership team.
What You'll Be Responsible For
- Design, build, and operate scalable, highly available infrastructure in AWS
- Own and evolve infrastructure as code (Terraform) across all environments
- Operate and optimize Aurora PostgreSQL, including replication, failover, and performance tuning
- Operate ECS (Fargate), ECR, and containerized services
- Operate Kafka-based event streaming systems
- Manage Auto Scaling Groups and EC2-based workloads
- Design and maintain CI/CD pipelines using Buildkite
- Build automation to eliminate manual operational work
- Manage and secure secrets and access using Vault, AWS Secrets Manager, and IAM
- Partner with engineering teams to improve system reliability and performance
- Lead, mentor, and manage a small DevOps team
- Set technical direction, standards, and priorities for the DevOps function
- Drive cost optimization across AWS infrastructure
- Operate systems behind Cloudflare, including WAF, CDN, and traffic management
Production Reliability & Incident Ownership
- Own production incident response end-to-end, including triage, mitigation, and coordination
- Lead high-severity outage response under pressure
- Serve as an escalation point for complex production issues
- Drive root cause analysis (RCA) and enforce follow-up actions
- Continuously improve system resilience and recovery mechanisms
Observability & System Insight
- Design and operate end-to-end observability across metrics, logs, and tracing
- Build high-signal monitoring, alerting, and dashboards
- Define and enforce SLIs, SLOs, and alerting standards
- Reduce alert fatigue and improve the signal-to-noise ratio
AI / LLM Systems — Emerging Area
- Experience using AI/agentic developer tools, such as Claude Code, Cursor, or similar tools, to accelerate DevOps workflows and improve engineering efficiency
- You don't have to be an expert with AI yet, but you must have the desire to learn quickly and become proficient
- This is an area we're heavily investing in as a company
What You Bring
- Deep experience operating production systems on AWS, including ECS/Fargate, EC2, networking, and IAM
- Expert-level Terraform experience managing infrastructure at scale
- Strong experience with containerized applications and distributed systems, such as Kafka
- Experience operating multi-tenant, highly available systems
- Proven ownership of production on-call and experience resolving critical incidents
- Strong systems fundamentals, including Linux, networking, and debugging
- Strong scripting ability in Bash, Python, or an equivalent language
- Experience designing and operating CI/CD systems
- Strong understanding of security best practices, including IAM and secrets management
- Demonstrated technical leadership and experience mentoring engineers
- Experience managing engineers, or clear readiness and interest in taking on direct people management
- Strong communication and prioritization skills
Nice to Have
- Experience operating multi-region or globally distributed systems
- Experience working with Cloudflare at scale
- Experience optimizing high-throughput or event-driven systems
- Experience leading a small infrastructure, platform, SRE, or DevOps team
- Experience operating or integrating LLM/AI services in production environments, including tracing and evaluation with OpenTelemetry, LangSmith, Langfuse, or equivalent tools
- Experience managing the performance, cost, and reliability of LLM workloads, including latency, token usage, rate limiting, and fallbacks
Oh Yeah, and We Can Offer You
- A tight-knit team of motivated, dedicated individuals who work together without ego
- Extensive access to and engagement with leadership
- An Agile working methodology
- Competitive salary
- Access to new technologies
- 401(k)
- Cell phone reimbursement
- Health, dental, and vision benefits