Lead deployment, updates, and operational support for cloud environments. Ensure availability, scalability, performance, and reliability; define SLAs/SLOs/SLIs; drive IaC and automation; improve monitoring, observability, and security; perform RCAs and cost optimizations; collaborate with cross-functional teams to deliver reliable services.
This role is for one of the Weekday's clients
Salary range: Rs 3000000 - Rs 4500000 (ie INR 30-45 LPA)
Min Experience: 7+ years
Location: Chennai
JobType: full-time
The Senior SRE is responsible for deployment, updates, and operational support for environments hosting our leading client’s cloud-based solutions. This role ensures operational excellence, a seamless client experience, and continuous improvement across infrastructure and delivery processes. The ideal candidate combines strong technical capabilities with the ability to lead delivery through influence and hands-on engineering expertise.
RequirementsKey Responsibilities
- Manage deployments, upgrades, maintenance, and operational support for cloud environments.
- Ensure high availability, scalability, performance, and reliability of production systems.
- Define, monitor, and improve SLAs, SLOs, and SLIs.
- Drive automation initiatives and Infrastructure as Code (IaC) adoption.
- Perform Root Cause Analysis (RCA) and implement preventive actions.
- Optimize cloud infrastructure, operational efficiency, and costs.
- Enhance monitoring, observability, security, and deployment processes.
- Collaborate with Engineering, Project Management, Customer Success, and cross-functional teams to deliver reliable services.
- Strong hands-on experience with AWS cloud platforms.
- Expertise in Kubernetes for container orchestration and cluster management.
- Experience with Terraform for Infrastructure as Code (IaC).
- Proficiency in Ansible for configuration management and automation.
- Hands-on experience with Helm for Kubernetes application deployments.
- Experience managing MariaDB and MongoDB databases in production environments.
- Strong understanding of CI/CD pipelines, deployment automation, and DevOps practices.
- Experience with monitoring, observability, logging, and alerting tools (e.g., Prometheus, Grafana, ELK, CloudWatch, Azure Monitor).
- Good knowledge of Linux administration, networking, DNS, load balancing, and cloud security best practices.
- Experience troubleshooting production environments, conducting Root Cause Analysis (RCA), and improving platform reliability.
- Understanding of SRE principles, including SLAs, SLOs, and SLIs.
- Scripting experience using Bash, Python, or Shell for automation.
- Excellent problem-solving, communication, and stakeholder management skills.
- Experience working in Site Reliability Engineering, DevOps, or Cloud Operations roles.
- Experience managing large-scale, production cloud environments.
- Ability to thrive in a fast-paced, customer-focused environment.
- Strong analytical mindset with a proactive approach to continuous improvement.
AWS, Kubernetes, Site Reliability Engineering
Good-to-have skillsHelm Charts, IaC, monitoring
Similar Jobs
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Operate and improve the reliability, availability, and performance of large-scale GeForce NOW services. Participate in incident triage and on-call rotations, build automation and tooling, enhance observability (metrics/logs/traces), drive SLO/SRI practices, run postmortems, and design/operate Kubernetes-based services across cloud and datacenter environments.
Top Skills:
AWSAzureBashContainerizationElk/OpensearchGCPGoGrafanaKubernetesMicroservicesOpentelemetryPrometheusPython
Cloud • Security • Software • Cybersecurity
Design, implement, and maintain reliable, scalable infrastructure for large distributed content delivery systems. Define and measure SLIs/SLOs, monitor availability and performance, troubleshoot incidents, and implement corrective actions. Develop automation to reduce manual work, participate in design reviews, and collaborate with product and engineering teams to improve system reliability and performance.
Top Skills:
AdbmsBashCloud ComputingDatadogGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux
Cloud • Security • Software • Cybersecurity
Senior SRE responsible for monitoring, analyzing, and improving availability, performance, and reliability of Akamai's Mapping Service. Define KPIs, build tooling to prevent recurrence, collaborate with product engineers on scalable designs, troubleshoot incidents, and use data analysis and network diagnostics to recommend improvements.
Top Skills:
CC++GrafanaJavaMonitoring/Alerting/Logging ToolsNetwork Diagnostics ToolsPerlPythonShellSQLUnix/Linux
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


