Jobs / Software Development Engineer (SDE) / Microsoft Senior Site Reliability Engineer Jobs in Hyderabad
Microsoft logo

Microsoft

Microsoft Senior Site Reliability Engineer Jobs in Hyderabad

location_on Hyderabad | Hybrid
work 12+ years Experience
payments Competitive Salary (Not disclosed)
schedule Full Time

Role Overview

Job Overview

Microsoft is hiring a Senior Site Reliability Engineer for its Azure Specialized AI Infrastructure team in India. The position is based in Hyderabad, Telangana, and is a full-time Individual Contributor role with a work schedule requiring three days per week in the office. The role focuses on building and operating reliable infrastructure for high-performance computing and generative AI workloads.

The engineer will help maintain large-scale distributed systems that support modern AI applications and machine learning models. The primary focus is reliability, scalability, security, performance, automation, and operational excellence across AI infrastructure. This includes working with compute, storage, networking, GPUs, InfiniBand, containers, and cloud infrastructure.


Key Responsibilities

  • Improve the reliability, scalability, and security of infrastructure supporting high-performance computing and AI workloads.
  • Lead incident response activities, perform root cause analysis, and implement improvements that reduce downtime and strengthen service availability.
  • Investigate performance issues across compute, storage, networking, GPUs, InfiniBand, and related infrastructure components.
  • Build and maintain automation for deployment, monitoring, predictive analysis, and infrastructure management.
  • Support containerized environments using technologies such as Kubernetes and Docker.
  • Provide technical guidance on cloud and AI infrastructure practices while collaborating with cross-functional engineering teams.
  • Advocate for customers by focusing on service quality and reliability for live AI workloads.
  • Track developments in AI infrastructure and recommend useful technologies or approaches for adoption.
  • Contribute to continuous improvement of operational processes and reliability practices.


Required Skills

The position requires extensive professional software engineering experience along with strong expertise in service operations, monitoring, and infrastructure reliability. Candidates should have substantial experience developing and supporting infrastructure services for AI or cloud platforms.

A strong understanding of distributed systems, cloud infrastructure, monitoring, incident management, and reliability engineering is important. The role requires engineers who can diagnose complex infrastructure issues, identify bottlenecks, coordinate incident response, and implement durable solutions.


Preferred Skills

Experience with large-scale cloud or distributed systems is preferred. Familiarity with Azure, Kubernetes, Docker, and the broader container ecosystem can strengthen a candidate's profile.

Additional relevant experience includes working with large supercomputers, AI platforms, GPUs, InfiniBand, or comparable high-performance technologies. Publications or certifications related to cloud or AI infrastructure are also considered a plus.


Education

The posting accepts several education and experience combinations. Required qualifications include a Master's Degree in Computer Science, Information Technology, or a related field with the stated technical experience, or a Bachelor's Degree in one of these areas with the required experience. Equivalent experience is also accepted.


Experience

The role requires 12+ years of professional software engineering experience, including at least 8 years focused on service operations, monitoring, and reliability improvement for infrastructure.

Candidates also need at least 5 years of hands-on experience developing and supporting infrastructure services for AI or cloud platforms, plus at least 1 year of experience with incident management and reliability engineering in cloud or AI environments.

Preferred experience includes 3+ years working with large-scale cloud or distributed systems, experience with distributed systems or cloud platforms, and exposure to large supercomputers, AI platforms, GPUs, or InfiniBand.


Required Technologies

Key technologies and technical areas associated with this position include:

  • Azure
  • Kubernetes
  • Docker
  • Containers
  • Cloud platforms
  • Distributed systems
  • AI infrastructure
  • High-performance computing
  • GPUs
  • InfiniBand
  • Monitoring and infrastructure automation


Soft Skills

Strong incident leadership, problem solving, collaboration, and customer advocacy are important for this position. The engineer should be able to remain focused during operational issues, coordinate with other teams, communicate technical findings, and turn incidents into long-term reliability improvements.

A learning mindset is also valuable because the team works in a rapidly changing AI infrastructure environment. Candidates should be comfortable researching emerging technologies, evaluating their practical value, and recommending improvements.


Benefits of Working in this Role

This role provides an opportunity to work on infrastructure supporting high-performance computing and generative AI workloads. Engineers can gain experience with large-scale distributed systems, cloud platforms, specialized hardware, containers, reliability engineering, and modern AI infrastructure.


Work Mode

The Microsoft posting specifies a work site requiring three days per week in the office. Travel is listed as less than 25%.


Location

The position is based in Hyderabad, Telangana, India.


Who Should Apply

This opportunity is designed for experienced Site Reliability Engineers, infrastructure engineers, cloud engineers, software engineers, and reliability specialists with extensive professional experience.

It is particularly suitable for candidates who have worked on large-scale cloud or distributed systems and have hands-on experience with AI infrastructure, high-performance computing, Kubernetes, Docker, GPUs, InfiniBand, or related technologies.


Career Growth

The role can deepen expertise in Site Reliability Engineering, AI infrastructure, distributed systems, cloud operations, high-performance computing, automation, and service reliability. It can also strengthen technical leadership through cross-functional collaboration, incident ownership, customer advocacy, and infrastructure innovation.


Application Advice

Tailor your resume to emphasize your total professional software engineering experience and your years working in service operations, monitoring, and reliability engineering. Clearly describe infrastructure services you have developed or supported for cloud or AI environments.

Highlight hands-on experience with Azure, Kubernetes, Docker, distributed systems, GPUs, InfiniBand, high-performance computing, or large-scale AI platforms where applicable. Include examples of incident response, root cause analysis, performance optimization, automation, monitoring, and reliability improvements.

If you hold relevant cloud or AI infrastructure certifications or have technical publications, include them because the posting identifies these as additional strengths. Also demonstrate customer advocacy, cross-team collaboration, and your ability to evaluate and adopt useful infrastructure technologies.

The job requires Microsoft and customer or government security screening requirements, including a Microsoft Cloud Background Check upon hire or transfer and periodically thereafter. Candidates should consider this requirement when applying.


Technical Ecosystem

Eligibility Criteria

school

Education

Master's Degree in Computer Science, Information Technology, or a related field with the specified technical experience, or Bachelor's Degree in the same fields with the required experience; equivalent experience is accepted.

work_history

Experience

12+ years

Top Picks for You

smart_toy

SoftoBot

Beta

Your job search assistant

Ask about roles, locations, remote work, or fresher jobs.