hireejobs
Hyderabad Jobs
Banglore Jobs
Chennai Jobs
Delhi Jobs
Ahmedabad Jobs
Mumbai Jobs
Pune Jobs
Vijayawada Jobs
Gurgaon Jobs
Noida Jobs
Oil & Gas Jobs
Banking Jobs
Construction Jobs
Top Management Jobs
IT - Software Jobs
Medical Healthcare Jobs
Purchase / Logistics Jobs
Sales
Ajax Jobs
Designing Jobs
ASP .NET Jobs
Java Jobs
MySQL Jobs
Sap hr Jobs
Software Testing Jobs
Html Jobs
IT Jobs
Logistics Jobs
Customer Service Jobs
Airport Jobs
Banking Jobs
Driver Jobs
Part Time Jobs
Civil Engineering Jobs
Accountant Jobs
Safety Officer Jobs
Nursing Jobs
Civil Engineering Jobs
Hospitality Jobs
Part Time Jobs
Security Jobs
Finance Jobs
Marketing Jobs
Shipping Jobs
Real Estate Jobs
Telecom Jobs

Site Reliability Engineer

1.00 to 10.00 Years   Pune, BangaloreBangalore, Noida, Chennai, Hyderabad, Gurugram, Mumbai City, Delhi   06 Aug, 2026
Job LocationPune, BangaloreBangalore, Noida, Chennai, Hyderabad, Gurugram, Mumbai City, Delhi
EducationNot Mentioned
SalaryNot Disclosed
IndustryIT Services & Consulting
Functional AreaIT Operations / EDP / MIS
EmploymentTypeFull-time

Job Description

    Key Responsibilities:- Define, measure, and enforce service level objectives, service level indicators, and error budgets for production systems in collaboration with engineering and product teams- Build, maintain, and improve monitoring, alerting, and observability frameworks across production infrastructure using tools such as Prometheus, Grafana, Datadog, ELK Stack, or similar- Respond to and lead incident management activities including detection, triage, escalation, resolution, and post-incident review for high-severity production incidents- Conduct thorough post-mortem and root cause analysis for production incidents and drive implementation of corrective and preventive actions to eliminate recurring failures- Design and implement automation frameworks to eliminate toil, improve operational efficiency, and reduce manual intervention in repetitive infrastructure and deployment tasks- Collaborate with software engineering teams to embed reliability practices including chaos engineering, fault injection, and resilience testing into the software development lifecycle- Manage and optimize cloud infrastructure on AWS, Azure, or GCP for high availability, fault tolerance, and cost efficiency- Oversee capacity planning, performance benchmarking, and traffic forecasting to ensure production systems can scale to meet business demand- Drive adoption of DevOps and SRE best practices including CI/CD, infrastructure as code, configuration management, and change management processes- Maintain comprehensive runbooks, operational documentation, and on-call playbooks to ensure consistent and effective incident response across the SRE team

Keyskills :
kubernetesincident managementpythonawsterraformobservabilityprometheusgrafana

Site Reliability Engineer Related Jobs

© 2019 Hireejobs All Rights Reserved