Build, automate, and operate a global cloud platform: develop automation in Go/Python, manage large-scale EKS clusters (Karpenter), author Terraform and Helm IaC, lead incident response and post-mortems, define SLIs/SLOs, implement observability (Datadog/Prometheus/Grafana) and PagerDuty on-call, and develop secure self-service tools to meet SOC2. Night-shift role based in Ahmedabad, India.
Position:Senior Site Reliability Engineer (SRE)Job Description:
The Senior Site Reliability Engineer (SRE) will build, automate, and operate Hanwha Vision's global cloud platform. We treat operations as a software engineering problem. Your goal is to eliminate manual tasks and replace them with stable, automated systems.
Key Responsibilities
- Automation: Write clean code (Go or Python) to automate cloud operations and deployment pipelines.
- Kubernetes Engineering: Build and manage large-scale Amazon EKS clusters, including networking, security, and scaling (using Karpenter).
- Infrastructure as Code: Write and maintain modular Terraform and Helm templates.
- Incident Management: Lead troubleshooting for critical system outages. Write clear, blameless post-mortems to prevent issues from happening again.
- Observability: Define SLIs/SLOs. Set up monitoring, logging, and on-call alerts using Datadog and PagerDuty.
- Security & Compliance: Build secure, self-service tools (like automated JIT AWS access) to meet SOC2 requirements.
Required Technical Skills
- 10+ years of professional experience in SRE, DevOps, or Systems/Infrastructure Engineering.
- Coding: Strong programming skills in Go or Python.
- Containers: Production experience running and scaling Kubernetes (EKS preferred).
- IaC: Expert-level knowledge of Terraform.
- Monitoring: Experience with Datadog, Prometheus, Grafana, or PagerDuty.
- Database/Messaging (Preferred): Basic understanding of DynamoDB, Kakfa/MSK, or caching layers.
- Languages: Strong written English proficiency is required to collaborate with global teams.
Certification : Preferred - AWS certified Solution Architect, AWS certified DevOps.
Remarks- Ready to work on Night shift only.
Similar Jobs
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Operate and improve the reliability, availability, and performance of large-scale GeForce NOW services. Participate in incident triage and on-call rotations, build automation and tooling, enhance observability (metrics/logs/traces), drive SLO/SRI practices, run postmortems, and design/operate Kubernetes-based services across cloud and datacenter environments.
Top Skills:
AWSAzureBashContainerizationElk/OpensearchGCPGoGrafanaKubernetesMicroservicesOpentelemetryPrometheusPython
Cloud • Security • Software • Cybersecurity
Design, implement, and maintain reliable, scalable infrastructure for large distributed content delivery systems. Define and measure SLIs/SLOs, monitor availability and performance, troubleshoot incidents, and implement corrective actions. Develop automation to reduce manual work, participate in design reviews, and collaborate with product and engineering teams to improve system reliability and performance.
Top Skills:
AdbmsBashCloud ComputingDatadogGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux
Cloud • Security • Software • Cybersecurity
Senior SRE responsible for monitoring, analyzing, and improving availability, performance, and reliability of Akamai's Mapping Service. Define KPIs, build tooling to prevent recurrence, collaborate with product engineers on scalable designs, troubleshoot incidents, and use data analysis and network diagnostics to recommend improvements.
Top Skills:
CC++GrafanaJavaMonitoring/Alerting/Logging ToolsNetwork Diagnostics ToolsPerlPythonShellSQLUnix/Linux
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


