Site Reliability Engineer
PwC India
2 - 5 years
Pune City
Posted: 29/06/2026
Job Description
Opportunity
We are looking for SREs who want to define what reliability means for the next generation of industrial software. Defining SLIs/SLOs, building observability platforms, and establishing incident management processes.
Responsibilitie
- sDefine and implement SLI/SLO frameworks for complex engineering systems across manufacturing and industrial client
- sDesign and deploy observability platforms using Prometheus, Grafana, and Datado
- gEstablish incident management processes and lead blameless post-mortem
- sImplement chaos engineering practices to proactively identify system weaknesse
- sDrive toil elimination through automation and platform improvement
- sBuild reliability engineering capabilities within the practice and client organisation
s
Essential Skil
- lsSLI/SLO definition and implementation at enterprise sca
- leObservability: Prometheus, Grafana, Datadog, New Rel
- icIncident management and post-mortem facilitati
- onChaos engineering: Gremlin, Chaos Monkey, Litm
- usPython testing for reliability validation and automated runboo
- ksAutomation and scripting: Python, Go, Ba
- shCloud platforms: AWS, Azure, G
CP
Experie
nce3+ years in SRE or Production Engineering roles with experience in enterprise or industrial environme
Services you might be interested in
We Search & Apply Jobs for You!
Our team scans through 1000s of opportunities and applies to roles best suited to your profile
Save 100+ hours and focus on what matters - cracking interviews and landing offers.
