hireejobs
Hyderabad Jobs
Banglore Jobs
Chennai Jobs
Delhi Jobs
Ahmedabad Jobs
Mumbai Jobs
Pune Jobs
Vijayawada Jobs
Gurgaon Jobs
Noida Jobs
Oil & Gas Jobs
Banking Jobs
Construction Jobs
Top Management Jobs
IT - Software Jobs
Medical Healthcare Jobs
Purchase / Logistics Jobs
Sales
Ajax Jobs
Designing Jobs
ASP .NET Jobs
Java Jobs
MySQL Jobs
Sap hr Jobs
Software Testing Jobs
Html Jobs
IT Jobs
Logistics Jobs
Customer Service Jobs
Airport Jobs
Banking Jobs
Driver Jobs
Part Time Jobs
Civil Engineering Jobs
Accountant Jobs
Safety Officer Jobs
Nursing Jobs
Civil Engineering Jobs
Hospitality Jobs
Part Time Jobs
Security Jobs
Finance Jobs
Marketing Jobs
Shipping Jobs
Real Estate Jobs
Telecom Jobs

Site Reliability Engineering (SRE) Lead AWS

Fresher   Hyderabad, All India   11 Jan, 2026
Job LocationHyderabad, All India
EducationNot Mentioned
SalaryNot Disclosed
IndustryIT Services & Consulting
Functional AreaNot Mentioned
EmploymentTypeFull-time

Job Description

    Job DescriptionRole OverviewAs the Site Reliability Engineering (SRE) Lead, you will be responsible for owning the reliability strategy for mission-critical systems. You will lead a team of engineers to ensure high availability, scalability, and performance. Your role will involve combining technical expertise with leadership skills to drive operational excellence and cultivate a culture of reliability across engineering teams.Key Responsibilities- Define and implement SRE best practices organization-wide.- Demonstrate expertise in production support, resilience engineering, disaster recovery (DCR), automation, and cloud operations.- Mentor and guide a team of SREs to foster growth and technical excellence.- Collaborate with senior stakeholders to align reliability goals with business objectives.- Establish SLIs, SLOs, and SLAs for critical services and ensure adherence.- Drive initiatives to enhance system resilience and decrease operational toil.- Design systems that can detect and remediate issues without manual intervention, focusing on self-healing systems and runbook automation.- Utilize tools like Gremlin, Chaos Monkey, and AWS FIS to simulate outages and enhance fault tolerance.- Act as the primary point of escalation for critical production issues, leading major incident responses, root cause analyses, and postmortems.- Conduct detailed post-incident investigations to identify root causes, document findings, and share learnings to prevent recurrence.- Implement preventive measures and continuous improvement processes.- Advocate for monitoring, logging, and ing strategies using tools such as Prometheus, Grafana, ELK, and AWS CloudWatch.- Develop real-time dashboards to visualize system health and reliability metrics.- Configure intelligent ing based on anomaly detection and thresholds.- Combine metrics, logs, and traces to enable root cause analysis and reduce Mean Time to Resolution (MTTR).- Possess knowledge of AIOps or ML-based anomaly detection for proactive reliability management.- Collaborate closely with development teams to integrate reliability into application design and deployment.- Foster a culture of shared responsibility for uptime and performance across engineering teams.- Demonstrate strong interpersonal and communication skills for technical and non-technical audiences.Qualifications Required- Minimum 15 years of experience in software engineering, systems engineering, or related fields, with team management experience.- Minimum 5 years of hands-on experience in AWS environments.- Deep expertise in AWS services such as EC2, ECS/EKS, RDS, S3, Lambda, VPC, FIS, and CloudWatch.- Hands-on experience with Infrastructure as Code using Terraform, Ansible, and CloudFormation.- Advanced knowledge of monitoring and observability tools.- Excellent leadership, communication, and stakeholder management skills.Additional Company Details(No additional company details provided in the job description.) Job DescriptionRole OverviewAs the Site Reliability Engineering (SRE) Lead, you will be responsible for owning the reliability strategy for mission-critical systems. You will lead a team of engineers to ensure high availability, scalability, and performance. Your role will involve combining technical expertise with leadership skills to drive operational excellence and cultivate a culture of reliability across engineering teams.Key Responsibilities- Define and implement SRE best practices organization-wide.- Demonstrate expertise in production support, resilience engineering, disaster recovery (DCR), automation, and cloud operations.- Mentor and guide a team of SREs to foster growth and technical excellence.- Collaborate with senior stakeholders to align reliability goals with business objectives.- Establish SLIs, SLOs, and SLAs for critical services and ensure adherence.- Drive initiatives to enhance system resilience and decrease operational toil.- Design systems that can detect and remediate issues without manual intervention, focusing on self-healing systems and runbook automation.- Utilize tools like Gremlin, Chaos Monkey, and AWS FIS to simulate outages and enhance fault tolerance.- Act as the primary point of escalation for critical production issues, leading major incident responses, root cause analyses, and postmortems.- Conduct detailed post-incident investigations to identify root causes, document findings, and share learnings to prevent recurrence.- Implement preventive measures and continuous improvement processes.- Advocate for monitoring, logging, and ing strategies using tools such as Prometheus, Grafana, ELK, and AWS CloudWatch.- Develop real-time dashboards to visualize system health and reliability metrics.- Configure intelligent ing based on anomaly detection and thresholds.- Combine metrics, logs, and traces to enable root cause analysis and reduce Mean Time to Resolution (MTTR).- Possess knowledge of AIOps or ML-based an

Keyskills :
AWSAnsibleIncident managementDisaster RecoverySRETerraform

Site Reliability Engineering (SRE) Lead AWS Related Jobs

© 2019 Hireejobs All Rights Reserved