AYN

MLOps Engineer

ai71 Careers Site · Abu Dhabi, UAE

About AI71:

AI71 is an industry leader in artificial intelligence, delivering innovative solutions that empower developers, businesses and governments to solve complex challenges. AI71 builds secure, enterprise-ready applications powered by cutting-edge technology—tailored for knowledge workers and sector-specific needs. AI71 bridges the gap between advanced AI and real-world impact. Guided by a strong commitment to research and responsibility, we create transformative solutions that drive progress and empower communities.

The Role:

As an MLOps Engineer you set the ML infrastructure and reliability strategy across AI71's platform, including how LLMs and other deep learning models are deployed, fine-tuned, and served at scale. You own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive multi-quarter ML infrastructure strategy. You are a force multiplier.

What You'll Do:

- Define ML infrastructure architecture across the platform: model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP)

- Set direction for ML system reliability: monitoring, latency / throughput / availability targets, and incident response across research and production environments.

- Mentor senior MLOps engineers; raise the operational bar across multiple teams.

- Drive cross-team initiatives that improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).

- Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.

- Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.

What You'll Bring:

- 10+ years of MLOps, ML infrastructure, or machine learning engineering with history of architectural ownership.

- Proven track record architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.

- Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong Python proficiency

- Mentorship record — engineers you have grown now operate independently at higher levels.

- Deep comfort architecting ML systems that run in both managed SaaS and on-premises / disconnected air-gapped environments.

- Kubernetes at architectural depth — GPU scheduling, multi-tenancy, operators, and the failure modes of distributed workloads on shared clusters.

- Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams.

Strong Preference:

- Ownership of production reliability at platform level: SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.

- Architecture-level experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.

- Deep GPU systems knowledge: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.

- Model optimization strategy at portfolio level: quantization (FP8, AWQ, GPTQ), speculative decoding, with measurable cost or latency outcomes across multiple systems.

- Experience in regulated or security-constrained environments — compliance-driven architecture, model governance, lineage, audit, and secrets management.

- On-prem / air-gap ML delivery architecture experience at scale.

- Track record maturing MLOps practice in a growing organization: standards, platform abstractions, and paved paths that outlived your involvement.

- Bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.

Nice to Have:

- Conference speaking, technical writing, or industry thought leadership.

- Open-source contributions to inference, serving, or ML infrastructure projects, particularly maintainer-level involvement.

- C/C++ or CUDA kernel experience for performance-critical paths.

- Arabic language skills.

Why AI71:

- Mission-Driven Work: Work on cutting-edge AI applications with a talented and passionate team, solving real-world challenges in critical sectors.

- Unparalleled Opportunity: This is a chance to innovate and solve real-world challenges using AI at a company with unique access to world-leading models and resources.

- Career Growth: We offer competitive compensation, benefits, and significant career growth opportunities as a foundational member of the team.

- World-Class Environment: Enjoy a flexible working environment and the latest tools & technologies needed to do your best work.

Apply on the employer’s site