Login Sign Up

Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering

Deloitte

4 - 8 years

Bengaluru

Posted: 08/08/2026

Job Description

Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering Job requisition ID : 108935 Location : Bengaluru Entity : Deloitte Touche Tohmatsu India LLP Senior Team Lead |Engineering, AI & Data -Engineering |Site Reliability Engineering Location: Bangalore The team Engineering helps Reimagine and re-engineer mission-critical operations and processes; Leverage engineering-led design, deep industry knowledge, and AI and data-driven insights to transform the technology platforms at the heart of business. Working alongside team, we empower and drive mission-critical solutions whether we need to modernize existing systems or implement new technology products and platforms. Through innovation, we improve financial performance, accelerate new digital businesses and fuel growth. Learn more about Engineering, AI and Data Your work profile We are looking for a highly skilled Site Reliability Engineer (SRE) to manage and scale mission-critical, production-grade distributed systems running on Google Cloud Platform (GCP). The ideal candidate will focus on reliability, automation, observability, and operational excellence while minimizing toil and improving system availability. Maintaining and improving 4 Nines of uptime to 5 Nines with engineering efforts. This role requires deep technical expertise in cloud-native technologies, Kubernetes, infrastructure automation, Linux administration and TCP/IP fundamentals, programming in one language and strong troubleshooting capabilities for distributed systems. The candidate needs to participate in the overall lifecycle management of mission critical banking services with a 24x7 operations mode in an rotational on-call basis. The job requires the candidate to have strong troubleshooting skills in a distributed environment spread across multiple cloud environments. The bare minimum ask would be to maintain high level of agility, learnability and adaptability in different scenarios. An engineer with a zeal to learn fast and having a bias for action would be the best fit for the role. Reliability & Operations Own end-to-end production systems reliability, availability, scalability, cost and performance. Drive measurable improvements in MTTR, MTTA, and incident response practices using automation and runbook additions and process enhancements. Participate in 24x7 on-call rotations and handle high-severity incidents and document the learnings on ongoing basis. Establish and manage SLI, SLO, SLA, Error Budgets, and operational metrics for mission critical services and partner with engineering teams with full accountability for upholding the SLOs. Partner with the various engineering, operations and cloud management teams to deliver highly reliable service in a timely manner. Cloud & Infrastructure Design, deploy, and manage infrastructure on Google Cloud Platform (GCP). Work extensively on: GKE (Kubernetes Engine), Compute, networking, IAM, Load Balancers, TLS Certs, BigQuery, Pub/Sub, cloud logging enhancement, metrics and logs analysis Implement and manage infrastructure using Terraform (Infrastructure as Code). Kubernetes & Containers Deploy and manage containerized workloads using Kubernetes (GKE). Troubleshoot issues related to: Pods, nodes, networking, storage, services on an ongoing basis Manage deployments using Helm, YAML, and rollout strategies (Canary/Blue-Green). Automation & CI/CD Build and maintain CI/CD pipelines using:Jenkins (pipeline-based, Groovy / Shell / Python scripting) Strong experience in using GitHub as a PowerUser Develop automation using Python and Shell scripting. Reduce operational toil through automation initiatives. Observability & Monitoring Implement and manage monitoring systems using:Dynatrace, Grafana, logs and metrics explorer Work with logs, metrics, and traces for deep observability to identify trends and arrest problems proactively. Define alerting strategies based on system behaviour and SLOs and create runbooks. Work alongside operations teams to identify, fix the production incidents and own the problem resolution. Work with engineering teams to isolate infra and application issues and set up right tooling for debugging production incidents. System & Application Troubleshooting Perform deep troubleshooting for: Distributed systems, Microservices-based architectures on containerised workloads, Java and Golang applications Strong debugging of: Application issues, Infrastructure issues,Network-related problems. Plan and execute continuous improvement Identify and eliminate repetitive manual tasks. Drive reliability engineering practices and culture. (DRY Dont Repeat Yourself) Collaborate with development teams to improve system design and resilience. Key skills required Education: Any bachelors or masters degree 48 years of relevant and progressive experience in SRE / DevOps / Cloud Engineering Hands-on experience managing production-grade systems (24x7 environments) Experience in high-scale distributed systems Exposure to banking/financial domain (optional but valuable) Understanding of security and compliance practices Experience with deployment strategies: Canary, Blue Green Technical Skills Cloud & Platform Strong expertise in Google Cloud Platform (GCP): GKE, VPC, IAM, Load Balancing, LB, Certs, KMS, logs and metrics exploration, BigQuery, Pub/Sub Good understanding of cloud architecture and landing zones Infrastructure as Code Strong hands-on experience with Terraform Ability to write and debug Terraform code from scratch Containers & Orchestration Deep expertise in: Kubernetes (GKE), Docker Strong troubleshooting experience in Kubernetes environments CI/CD & Automation Hands-on experience with: Jenkins (pipeline-based CI/CD),GitHub Strong scripting skills: Python (preferred),Shell scripting Experience with automation frameworks and tooling Observability Experience with:Dynatrace / Grafana Log, metrics, and trace-based monitoring Programming & Debugging Working knowledge of:Java and/or Golang applications Strong debugging skills across application and infrastructure layers Linux & Networking Strong Linux fundamentals Deep understanding of TCP/IP networking Ability to debug network issues in distributed systems Reliability Engineering Skills Solid understanding of: SLI, SLO, SLA, Error Budgets Demonstrable and Proven Experience improving:MTTR, MTTA Experience handling incident management lifecycle Soft Skills Strong analytical and troubleshooting mindset Excellent communication and stakeholder management Ability to work in high-pressure production environments Ownership-driven and proactive approach Ideal Candidate Profile in summary would be like : Strong GCP + Kubernetes + Terraform core Hands-on production troubleshooting expert Good at automation + reducing toil Deep understanding of SRE principles Comfortable in 24x7 production environments

About Company

Deloitte is a global professional services firm that provides a wide range of services, including audit and assurance, consulting, tax, risk management, and financial advisory. With a presence in over 150 countries and a network of member firms, Deloitte serves clients across various industries, helping them solve complex business challenges, improve operations, and innovate. Known for its expertise in management consulting, technology solutions, and strategy, Deloitte is one of the Big Four accounting firms and is recognized for its commitment to quality, integrity, and making an impact in the marketplace.

Services you might be interested in

We Search & Apply Jobs for You!

Our team scans through 1000s of opportunities and applies to roles best suited to your profile

Save 100+ hours and focus on what matters - cracking interviews and landing offers.