Software Development Engineer, EC2 UltraServer Availability
Amazon · Seattle, WA · Software Development
About this role
Amazon is hiring a mid-level Software Engineer based in Seattle, WA. The posting calls out experience with AWS, Machine Learning, Embedded Systems, EC2. Compensation is listed at $143,700–$194,400 per year.
- Role
- Software Engineer
- Function
- software engineering
- Level
- mid
- Track
- Individual contributor
- Employment
- Full-time
- Location
- Seattle, WA
- Department
- Software Development
- Posted
- Apr 10, 2026
More roles at Amazon
Job description
from Amazon careersThe Software Development Engineer II will design, build, and maintain cloud-based repair and recovery workflows for NVIDIA GB200 / GB300 UltraServers, orchestrating repair and recovery operations from impairment detection through completed recovery. This role requires expertise in AWS services, system architecture, and cross-functional collaboration with Capacity Management, Hardware Engineering, and Datacenter Operations to manage AI/ML infrastructure. Key job responsibilities The Software Development Engineer (SDE II) on the EC2 UltraServer Availability team is responsible for ensuring high availability of customer GB200 and GB300 UltraServers by orchestrating complex repair and recovery workflows. Following are the core responsibilities System Design Architecture * Design and architect solutions that are cross-functional to Capacity Management, Hardware Engineering, and Datacenter Operations * Work in environments where the technology strategy is defined but the solution design is not * Build solutions that are stable, logical, testable, and efficient with the ability to independently make trade-off decisions * Investigate and develop design concepts to frame solution sets at an application and product level Software Development * Build cloud-based solutions using AWS native services for scaling infrastructure frameworks * Write high-quality, maintainable code with proper testing and code reviews * Develop and maintain the repair and recovery workflows for GB200…