Site Reliability Engineer
Insight Global
5 - 7 years
Bengaluru
Posted: 28/06/2026
Job Description
Position: Site Reliability Engineer
Location: Hybrid 3x's/week in Bengaluru, Karnataka 560064, India
Required Skills & Experience
- ~35 years of experience in a Site Reliability Engineer or similar role (58 years total IT experience)
- Strong experience working in Azure cloud environments
- Hands-on experience with monitoring and observability tools (e.g., Elastic, Prometheus, Grafana, or similar)
- Experience supporting production applications (application-focused SRE vs. infrastructure-only)
- Working knowledge of Kubernetes (AKS) for monitoring, alerting, or administration
- Ability to troubleshoot and debug applications, including reading and understanding code
- Familiarity with .NET/C# application environments
- Experience with databases (SQL and/or NoSQL such as Cosmos DB, PostgreSQL, etc.)
Nice to Have Skills & Experience
- Experience in multi-cloud or hybrid-cloud environments (Azure + AWS/GCP)
- Exposure to EKS or non-Azure Kubernetes environments
- Experience supporting single-page applications (e.g., Angular)
- Background in automation and scripting
- Experience working in global or distributed teams
Job Description
We are seeking a Site Reliability Engineer (SRE) to support a modern, cloud-based application platform. This role will focus on improving reliability, scalability, and observability across core systems while helping transition teams from reactive support to proactive engineering practices.
The ideal candidate will bring solid experience in cloud environments, application monitoring, and automation, along with the ability to collaborate closely with development teams and contribute to the maturation of SRE practices.
Key Responsibilities
- Support and enhance monitoring, alerting, and observability frameworks across production environments
- Troubleshoot production issues and participate in root cause analysis to reduce recurrence
- Support cloud-based applications and contribute to platform reliability improvements
- Partner with engineering teams to identify and resolve performance, scalability, and resiliency gaps
- Automate operational tasks and workflows to reduce manual intervention and improve efficiency
- Support and monitor containerized environments (Kubernetes/AKS)
- Contribute to incident response processes and help drive improvements in response time and downtime reduction
- Assist in establishing and adhering to SRE best practices and standards
Services you might be interested in
We Search & Apply Jobs for You!
Our team scans through 1000s of opportunities and applies to roles best suited to your profile
Save 100+ hours and focus on what matters - cracking interviews and landing offers.
