Jobs / Software Development Engineer (SDE) / NVIDIA Hiring Site Reliability Engineer in Bengaluru
NVIDIA logo

NVIDIA

NVIDIA Hiring Site Reliability Engineer in Bengaluru

location_on Bangalore | On-site
work Fresher
payments Competitive Salary (Not disclosed)
schedule Full Time

Role Overview

Job Overview

NVIDIA is hiring a Site Reliability Engineer in Bengaluru for a full-time engineering opportunity focused on reliability, scalability, automation, cloud infrastructure, and modern distributed systems. The role is suitable for an early-career candidate who has a strong technical foundation and wants to build practical experience in Site Reliability Engineering, DevOps, cloud platforms, and platform operations.

As part of the role, you will support engineering initiatives that help enterprise systems remain reliable and scalable. You will work with distributed systems, cloud-native infrastructure, databases, monitoring platforms, automation, and incident management processes. The position also provides exposure to AI-powered engineering workflows and modern tools used to improve developer productivity.

This opportunity is well suited to candidates who enjoy understanding how complex systems operate and who are comfortable learning from experienced engineers. Personal projects, internships, coursework, open-source contributions, or technical competitions can help demonstrate practical interest in cloud infrastructure and reliability engineering.


Key Responsibilities

  • Contribute to reliability and scalability initiatives for enterprise systems and AI-focused products and services.
  • Assist with distributed system development and learn modern architectural approaches used in cloud environments.
  • Support automation for database operations such as provisioning, scaling, backup, and failover.
  • Help develop monitoring dashboards, alerts, and automation scripts for system performance and reliability.
  • Participate in incident response, issue triage, resolution activities, and post-incident reviews.
  • Work with Cloud, Platform, Security, and AI/ML teams on reliability and operational improvements.
  • Operate and troubleshoot Kubernetes-based and other cloud-native environments under established engineering practices.
  • Explore AI-assisted development approaches, including coding agents and LLM-powered tools, to improve engineering workflows.
  • Apply version control, automation, documentation, and collaborative development practices throughout assigned work.


Required Skills

Candidates should have foundational programming knowledge in at least one language such as Python, TypeScript, JavaScript, or Go. The role also requires a basic understanding of cloud platforms, with exposure to AWS, Azure, or Google Cloud Platform.

Knowledge of containerization is important, particularly Docker and Kubernetes. Candidates should understand basic Linux or Unix concepts, networking fundamentals, and Git-based version control.

An understanding of observability concepts is valuable. Candidates should be familiar with the purpose of logs, metrics, and traces and have an interest in tools such as OpenTelemetry, Prometheus, or Grafana.

The position also requires basic relational database knowledge. Candidates should understand SQL, indexing, and simple query optimization and have exposure to databases such as PostgreSQL or MySQL.

Infrastructure-as-code is another important area. Coursework or exposure to Terraform, AWS CDK, or CloudFormation is useful, while candidates without direct experience should be prepared to learn these technologies.


Preferred Skills

Practical projects can make an early-career application stronger. Useful experience includes building cloud infrastructure, automating repetitive operational tasks, creating CI/CD pipelines, deploying containers, or experimenting with DevOps and SRE practices.

Open-source contributions, hackathons, coding competitions, and participation in technical communities can also demonstrate initiative and a willingness to solve unfamiliar problems.

Exposure to AI and machine learning is an advantage, particularly projects involving simple machine learning models, LLM APIs, or AI-assisted developer tools such as GitHub Copilot or Cursor.


Education

A BS degree in Computer Science or a related technical discipline such as Physics or Mathematics is specified. Equivalent practical experience may also be considered. Relevant academic coursework in operating systems, networking, databases, programming, cloud computing, distributed systems, or automation can help demonstrate readiness for the role.


Experience

The job description does not specify a professional experience range. The requirements are framed around foundational technical proficiency and practical exposure, making the role suitable for early-career applicants. Internships, coursework, personal projects, open-source work, and academic assignments involving cloud, programming, automation, DevOps, or SRE concepts can be relevant.


Required Technologies

The technology areas mentioned include Python, TypeScript, JavaScript, Go, AWS, Azure, Google Cloud Platform, Docker, Kubernetes, Terraform, AWS CDK, CloudFormation, Linux, Unix, Git, OpenTelemetry, Prometheus, Grafana, PostgreSQL, MySQL, SQL, and infrastructure-as-code.

The role also includes exposure to CI/CD, observability, distributed systems, database automation, cloud-native infrastructure, networking, monitoring, logging, metrics, tracing, and AI-assisted engineering. LLM-powered tools, GitHub Copilot, Cursor, and LLM APIs are relevant areas for candidates interested in modern AI-enabled development workflows.


Soft Skills

Site Reliability Engineering requires curiosity and a structured approach to problem solving. Candidates should be comfortable investigating unfamiliar systems, asking questions, learning from senior engineers, and turning technical problems into practical next steps.

Communication and teamwork are important because the role involves collaboration with Cloud, Platform, Security, and AI/ML teams. Ownership, initiative, adaptability, and a willingness to learn are valuable qualities for working in a fast-moving engineering environment.


Benefits of Working in this Role

The position provides practical exposure to cloud infrastructure, distributed systems, databases, Kubernetes, observability, automation, incident response, and DevOps practices. Early-career engineers can develop a broad understanding of how production systems are designed, monitored, operated, and improved.

The role also offers an opportunity to explore AI-assisted engineering practices and collaborate with multiple technical teams. These experiences can help build a strong foundation for long-term careers in SRE, DevOps, cloud engineering, platform engineering, and infrastructure development.


Work Mode

The provided job description does not specify an onsite, hybrid, or remote work arrangement. Candidates should confirm the current work model with NVIDIA during the application process.


Location

This Site Reliability Engineer opportunity is based in Bengaluru, Karnataka, India. It is listed under NVIDIA's Engineering job category and is a full-time position.


Who Should Apply

This role is suitable for candidates with a computer science or related technical background who want to start or grow a career in Site Reliability Engineering. Applicants should have foundational programming skills and an interest in cloud platforms, containers, Linux, networking, databases, automation, and observability.

Fresh graduates can strengthen their applications by presenting relevant academic projects, internships, certifications, GitHub work, hackathon projects, or personal cloud deployments. Candidates should focus on demonstrating what they built, how they solved problems, and what they learned rather than simply listing technologies.


Career Growth

Experience in this role can provide a foundation for careers such as Site Reliability Engineer, DevOps Engineer, Cloud Engineer, Platform Engineer, Infrastructure Engineer, or Reliability-focused Software Engineer. Continued development in Kubernetes, cloud architecture, infrastructure-as-code, observability, distributed systems, automation, and incident management can support progression toward more advanced engineering responsibilities.


Application Advice

Tailor your resume to the reliability and infrastructure focus of the position. Highlight Python, TypeScript, JavaScript, or Go projects, along with practical work involving AWS, Azure, GCP, Docker, Kubernetes, Terraform, Git, SQL, or monitoring tools.

If you have a personal project, describe the architecture and operational problem you solved. For example, explain how you deployed an application with containers, automated infrastructure, created monitoring dashboards, or implemented a CI/CD workflow. Open-source contributions and technical competition experience can also demonstrate initiative.

Be prepared to explain basic cloud concepts, containerization, Linux commands, networking fundamentals, SQL queries, Git workflows, observability, infrastructure-as-code, and the purpose of SRE practices. Candidates should also be ready to discuss how they would investigate an incident and improve system reliability.


Technical Ecosystem

Eligibility Criteria

school

Education

Bachelor of Science degree in Computer Science or a related technical field such as Physics or Mathematics, or equivalent practical experience.

work_history

Experience

Fresher

Top Picks for You

smart_toy

SoftoBot

Beta

Your job search assistant

Ask about roles, locations, remote work, or fresher jobs.