iLink Digital Logo

iLink Digital

Senior Site Reliability Engineer

Reposted 6 Days Ago
Be an Early Applicant
In-Office
Chennai, Tamil Nadu, IND
Senior level
In-Office
Chennai, Tamil Nadu, IND
Senior level
Lead SRE role managing multi-cloud (AWS/Azure) production systems, Kubernetes clusters, and PostgreSQL. Build IaC with Terraform/Ansible, implement observability and incident management (Prometheus/Grafana/Datadog/ELK), define SLOs/SLIs, drive vulnerability remediation, run on-call rotations, and collaborate on capacity planning, DR, and post-mortems to improve reliability.
The summary above was generated by AI

About The Company:


iLink Digital is a Global Software Solution Provider and Systems Integrator, delivers next-generation technology solutions to help clients solve complex business challenges, improve organizational effectiveness, increase business productivity, realize sustainable enterprise value and transform your business inside-out. iLink integrates software systems and develops custom applications, components, and frameworks on the latest platforms for IT departments, commercial accounts, application services providers (ASP) and independent software vendors (ISV). iLink solutions are used in a broad range of industries and functions, including healthcare, telecom, government, oil and gas, education, and life sciences. iLink’s expertise includes Cloud Computing & Application Modernization, Data Management & Analytics, Enterprise Mobility, Portal, collaboration & Social Employee Engagement, Embedded Systems and User Experience design etc.

 

What makes iLink's offerings unique is the fact that we use pre-created frameworks, designed to accelerate software development and implementation of business processes for our clients. iLink has over 60 frameworks (solution accelerators), both industry-specific and horizontal, that can be easily customized and enhanced to meet your current business challenges.



Requirements
  • 6–10 years of experience in SRE, DevOps, infrastructure and production support engineering roles.
  • Proven experience managing multi-cloud environments (AWS + Azure).
  • Demonstrated experience handling P1/P2 production incidents in cloud environments.
  • Familiarity with Prometheus, Grafana, Datadog, or Splunk
  • Design, deploy, and manage Kubernetes clusters for production workloads at scale.
  • Architect and maintain PostgreSQL databases — performance tuning, HA setup, backup/restore strategies.
  • Build and manage cloud infrastructure on AWS and Azure using Terraform and Ansible.
  • Lead vulnerability management programs — identify, prioritize, and remediate security risks across the stack.
  • Define and enforce SLOs, SLIs, and error budgets; drive reliability improvements across services.
  • Implement IaC best practices, automate provisioning pipelines, and reduce manual toil.
  • Collaborate with development teams on capacity planning, disaster recovery, and incident post-mortems.
  • Build and maintain monitoring, alerting, and observability frameworks (Prometheus, Grafana, ELK, etc.).
  • Lead end-to-end incident management — detection, triage, escalation, resolution, and communication.
  • Serve as an on-call engineer; manage and respond to alerts and production incidents effectively.
  • Conduct blameless post-mortems and implement action items to prevent recurrence.
  • Monitor system health using dashboards and alerting tools; proactively identify degradation risks.
  • Collaborate with Dev, QA, and infrastructure teams to identify and reduce toil and failure points.
  • Support Kubernetes workloads and assist in troubleshooting cluster-level issues.
  • Work across AWS and Azure environments for incident containment and recovery.
  • Maintain and improve runbooks, playbooks, and incident response documentation.
  • Strong understanding of networking, security, and distributed systems.
  • Excellent communication skills for cross-team collaboration and post-mortem documentation.
  • Experience with Helm, ArgoCD, or GitOps workflows.


Benefits
  • Competitive salaries
  • Medical Insurance
  • Employee Referral Bonuses
  • Performance Based Bonuses
  • Flexible Work Options & Fun Culture
  • Robust Learning & Development Programs
  • In-House Technology Training


iLink Digital Chennai, Tamil Nadu, IND Office

9th Floor, Symbyont Co-workspace RMZ Millenia Business Park Phase 2, Campus 5, Kandanchavadi, Perungudi, – , Chennai, India, 600096

Similar Jobs

2 Days Ago
In-Office
Chennai, Tamil Nadu, IND
Senior level
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead the design, implementation, and operation of AI Ops and SRE solutions in public cloud environments. Build RAG pipelines, agentic workflows, and AI-powered enterprise applications; automate infrastructure with Terraform and GitHub Actions; manage Kubernetes; lead incident response; ensure reliability, scalability, and performance; establish AI evaluation and monitoring frameworks; and mentor engineering teams.
Top Skills: AksAWSAzureCi/CdEksGCPGenerative AiGithub ActionsGkeInfrastructure As CodeKubernetesNode.jsPythonRagTerraformVector Search
2 Days Ago
In-Office
Chennai, Tamil Nadu, IND
Senior level
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Develop and operate reliable cloud infrastructure and applications, automate cloud operations and deployments, implement infrastructure procedures, and provide 24/7 rotating on-call support. The role requires enterprise cloud administration across public-cloud platforms, programming, CI/CD, Kubernetes or Terraform experience, and knowledge of cloud security, networking, monitoring, logging, and Linux/Windows systems.
Top Skills: AWSAzureCi/CdGCPKubernetesLinuxNode.jsPythonTerraformWindows
3 Days Ago
In-Office
Senior level
Senior level
Software
Support reliability and modernization of Kubernetes-based microservices platforms. Build Datadog observability solutions including dashboards, alerts, APM, metrics, logging, and tracing; integrate monitoring with AWS and CI/CD pipelines; and automate operational tasks using Python. Manage agents, integrations, API keys, access controls, and secure configurations while leading maintenance and platform improvements that enhance reliability, scalability, and performance.
Top Skills: APIsApmAWSCi/CdDatadogDynatraceElasticGoGrafanaJavaJavaScriptKubernetesNew RelicNode.jsPrometheusPythonSplunk Observability

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account