Software Development Engineer II, AWS SageMaker AI Training Job
Amazon · Bellevue, Washington, United States
SageMaker AI Training Job team is looking for a Software Development Engineer!
This is a great opportunity to join the SageMaker Training Job team, which lies at the very core of SageMaker. SageMaker Training is a set of managed services that allow customers to train machine learning models using large datasets on managed infrastructure. As the industry leader, SageMaker training is the fastest, easiest, and most cost-effective platform for data scientists to train their models. Learn more at https://docs.aws.amazon.com/sagemaker/latest/dg/how-it-works-training.html.
You will be responsible for building and maintaining mission-critical systems in our best-in-class machine learning platform, and scale those systems to support training jobs that run on hundreds of thousands of machines with less than 0.1% failure rate.
You will constantly experiment with new technologies and ideas in order to innovate on our platform and help ensure SageMaker remains best-in-class.
You will be expected to perform at the highest levels of engineering and operational excellence: building highly resilient and scalable systems, writing clear and effective documents, actively contribute to discussions on technical direction and strategy, and raise the bar on all fronts.
Most importantly, you will get to work with a team of highly talented engineers who have all answered the call above and strive for new heights every day.
Key job responsibilities
As a Software Development Engineer on the SageMaker AI team, you will:
- Design, build, test, and operate services that orchestrate foundation-model data preparation, training, evaluation, and deployment as reliable, contract-validated workflows.
- Own delivery of individual components end-to-end — from design and implementation through deployment, monitoring, and on-call operations.
- Build and extend compute-backend integrations and job launchers — submitting, monitoring, and recovering large-scale training jobs across SageMaker (Training/Processing), and AWS Batch.
- Improve the platform's resiliency and operability for long-running distributed jobs — checkpoint/resume, fault detection and recovery, retries, and observability (metrics, logging, experiment tracking).
- Contribute to the SDK, workflow orchestration, and schema/contract layer that teams use to declare and run jobs, and to the CDK infrastructure that deploys the platform.
- Integrate containerized training and evaluation frameworks (e.g., PyTorch/FSDP, verl, NeMo/Megatron) into the platform's task and recipe model.
- Contribute to design and architecture discussions, write clear technical designs, and uphold engineering best practices (code review, testing, operational readiness).
- Collaborate with ML scientists and internal customers to translate training requirements into reliable, self-service platform capabilities, and help onboard and mentor interns and new engineers as you grow.