View All Jobs 132913

Sr. Availability Engineer , Colocation Design Standards (codes)

Lead global colocation availability improvements through incident management and strategic planning
Sydney
Senior
17 hours agoBe an early applicant
Amazon

Amazon

A global e-commerce giant offering a vast array of products, cloud services, and digital streaming content.

Senior Availability Engineer

AWS Infrastructure Services owns the design, planning, delivery, and operation of all AWS global infrastructure. In other words, we're the people who keep the cloud running. We support all AWS data centers and all of the servers, storage, networking, power, and cooling equipment that ensure our customers have continual access to the innovation they rely on. We work on the most challenging problems, with thousands of variables impacting the supply chain — and we're looking for talented people who want to help.

You'll join a diverse team of software, hardware, and network engineers, supply chain specialists, security experts, operations managers, and other vital roles. You'll collaborate with people across AWS to help us deliver the highest standards for safety and security while providing seemingly infinite capacity at the lowest possible cost for our customers. And you'll experience an inclusive culture that welcomes bold ideas and empowers you to own them to completion.

The Senior Availability Engineer is responsible for managing and improving Colocation infrastructure availability, incident response, and technical risk management across global infrastructure design and operations. This role requires collaboration with various teams and stakeholders at all levels to ensure system reliability and implement strategic improvements. As technical leader facing uncertainty, you will decompose complex problems into straightforward and actionable solutions.

Key job responsibilities:

  • Develop global lessons learnt strategies through incident management processes, including emergency (FOC) calls participation, 48-hour report contribution, and implementation of lessons learned using 5 whys methodology.
  • Review and technical inspect Root Cause Analysis (RCA) and lead discussions for the corrective actions related to site/equipment failures. Directly support operational issues escalations including some on-call event support.
  • Contribute to regular Colocation availability reporting for Senior Leadership.
  • Define Corrective global program implementation through strategic action plans, workflows development, and mentoring of regional teams.
  • Provide strategic technical direction for the SPOF (Single Point of Failure) program, including availability strategies, approval of remediation designs, and global Risk Reduction Library.
  • Participate in Infra Project reviews and Approvals (CIMAP/HIMAPS, including accepted deviation endorsement and documentation review.
  • Handle technical escalations on first-of-a-kind design features, projects and product deployments. Provide strategic availability recommendations for sites globally.
  • Drive process improvement initiatives, including Colo Regional Engineering training and mentoring activities, and contribute to the Infrastructure Risk Prioritization Scoring.
  • 25% international travel.

About the team

About AWS

Diverse Experiences

AWS values diverse experiences. Even if you do not meet all of the qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you from applying.

Why AWS?

Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that's why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

Inclusive Team Culture

AWS values curiosity and connection. Our employee-led and company-sponsored affinity groups promote inclusion and empower our people to take pride in what makes us unique. Our inclusion events foster stronger, more collaborative teams. Our continual innovation is fueled by the bold ideas, fresh perspectives, and passionate voices our teams bring to everything we do.

Mentorship & Career Growth

We're continuously raising our performance bar as we strive to become Earth's Best Employer. That's why you'll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional.

Work/Life Balance

We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud.

+ Show Original Job Post
























Sr. Availability Engineer , Colocation Design Standards (codes)
Sydney
Engineering
About Amazon
A global e-commerce giant offering a vast array of products, cloud services, and digital streaming content.