CloudifyOps Logo

CloudifyOps

Cloud Engineer

Posted One Month Ago
Be an Early Applicant
In-Office
Chennai, Tamil Nadu, IND
Mid level
In-Office
Chennai, Tamil Nadu, IND
Mid level
Own cloud infrastructure and production reliability across AWS and Kubernetes environments. Responsibilities include on-call incident management, root-cause analysis, monitoring and observability improvements, Terraform provisioning, CI/CD health monitoring, application performance analysis, and cluster troubleshooting. The role also contributes to an early-stage AI-powered pipeline monitoring tool through prototyping and experimentation. Clear client communication, precise documentation, and strong ownership are essential.
The summary above was generated by AI

Culture at CloudifyOps :

Working at CloudifyOps is a rewarding experience! Great people, a work environment that thrives on creativity, and the opportunity to take on roles beyond a defined job description are just some of the reasons you should work with us.

About the Role :

We’re looking for someone who genuinely wants to understand why systems fail, not just respond to alerts. This role sits at the crossroads of cloud infrastructure and production reliability. You’ll own monitoring, handle on-call, and be the person who digs in when things go wrong. At the same time, we’re building an AI-powered pipeline monitoring tool and need someone curious enough to contribute to shaping it, not just watching over it.


What you’ll do:

  • Handle the on-call rotation and own incidents end-to-end triage, mitigation, escalation where needed, and clean resolution. You don’t pass the baton and disappear.

  • Write clear, structured RCAs after every significant incident  what happened, when, why, and what changes going forward. These go to clients, so they need to work for both an engineer and a non-technical reader.

  • Maintain and improve the monitoring stack across environments dashboards, alerting rules, log pipelines, and distributed traces. Treat noisy alerts as a problem to fix, not something to mute.

  • Provision and manage cloud infrastructure on AWS using Terraform. This is a hands-on role not just reviewing what others set up.

  • Work with Kubernetes across multiple environments, debugging pod and node issues.

  • Monitor CI/CD pipeline health via Jenkins and support teams using Rancher for workload and cluster management.

  • Track application performance using APM tooling and JVM metrics: spot anomalies, investigate degradation, and flag systemic issues before they become incidents.

  • Contribute to the AI monitoring tool initiative: prototype, test, iterate. This is early-stage work and needs someone willing to figure things out, not just execute a finished design.

Tech Stack:

  • Cloud & Infrastructure  :  AWS · Kubernetes (K8s) · Terraform · Linux

  • Observability & Metrics :  Prometheus · Grafana · APM (Datadog / New Relic / Kfuse) · JVM Metrics & GC Analysis · ELK / EFK Stack · Distributed Tracing

  • CI/CD & Platform :  Jenkins · ArgoCD · Rancher · Git · Docker

  • Good to Have(Not Mandatory) : Python / Bash scripting · OpenTelemetry · Zenduty / OpsGenie · ML / AI basics

Expectations:


On-call here is real.Incidents happen outside business hours and when they do, it’s your responsibility to pick them up and drive them forward. That’s not unusual for this type of role but we want to be direct about it upfront.

Client expectations are high.  You’ll produce RCAs, incident timelines, and status communications that clients read closely. Your writing needs to be clear, structured, and free of vagueness. “We investigated and fixed the issue” isn’t good enough. What was the issue, why did it happen, what was the business impact, and what prevents recurrence.

We expect precision regarding your own work.  After a change, an incident, or a deployment, you should be able to clearly explain what you did and why without being prompted. Ownership doesn’t end when the alert clears.


Who we’re looking for:


  • The ideal candidate should have 2.5 years to 5 years of work experience.

  • Strong fundamentals.  You understand how distributed systems actually behave under load, not just that a dashboard went red. You can read logs, metrics, and traces together.

  • Ownership without prompting.  If you find a gap in monitoring coverage, you close it. If an RCA feels incomplete, you go back and make it precise. You don’t wait to be asked.

  • Writes clearly under pressure.  During an incident, your updates should help not add noise. After one, your documentation should be good enough that anyone picking it up six months later understands what happened.

  • Curious about what comes next.  The AI tooling initiative needs someone interested in figuring it out, not just waiting for a ticket. Some comfort with experimentation and ambiguity goes a long way here.


Equal opportunity employer 

CloudifyOps is proud to be an equal opportunity employer with a global culture that embraces diversity. We are committed to providing an environment free of unfair discrimination and harassment. We do not discriminate based on age, race, color, sex, religion, national origin, disability, pregnancy, marital status, sexual orientation, gender reassignment, veteran status, or other protected category.



CloudifyOps Chennai, Tamil Nadu, IND Office

OMR Service Road, 3rd Phase, No. 1, Santhosh Nagar, Kandhanchavadi, Perungudi,, Chennai, Chennai, India, 600096

Similar Jobs

Yesterday
In-Office or Remote
2 Locations
Senior level
Senior level
Automotive
Designs, implements, and manages non-relational databases on Google Cloud, including MongoDB, Firestore, Bigtable, Memorystore, and Neo4j. Responsibilities include Terraform-based infrastructure, CI/CD, observability with Dynatrace, performance optimization, migrations, high availability and disaster recovery, security, cost optimization, and developer enablement. The role also establishes SRE practices, database monitoring, IAM and OAuth2 controls, AI agent-to-data tracing, and production support.
Top Skills: BigtableCi/CdCloud FunctionsCloud RunCycodeDevOpsDqlDynatraceFirestoreGitGoogle Cloud Platform (Gcp)IamMemorystore For RedisMemorystore For ValkeyMiroMongoDBMongodb AtlasNeo4JNeo4J AuradbOauth2PythonQuery InsightsSystem InsightsTerraformVs CodeWorkload Identity
Yesterday
In-Office
Chennai, Tamil Nadu, IND
Senior level
Senior level
Biotech • Pharmaceutical
Design and automate secure, scalable AWS infrastructure for laboratory and digital applications. Build Infrastructure as Code modules, CI/CD pipelines, container platforms, observability solutions, networking, security controls, and cost optimization practices. Support cloud migration, incident analysis, service reliability, documentation, and responsible AI-assisted engineering across enterprise teams.
Top Skills: Amazon EcsAmazon EksAmazon QApi GatewayArgocdAWSAws CdkAzure DevopsBashChatgpt EnterpriseCheckovClaudeCloudFormationCloudwatchCursorDockerEc2ElkEventbridgeFargateFluxGitGithub ActionsGithub CopilotGitlab CiGrafanaGuarddutyHelmIamInspectorJenkinsKubernetesLambdaLinuxOpaOpensearchOpentelemetryPrometheusPythonRdsRoute 53S3Security HubShellStep FunctionsTerraformVpc
Yesterday
In-Office
Chennai, Tamil Nadu, IND
Senior level
Senior level
HR Tech • Professional Services
Manage and optimize AWS and Azure cloud infrastructure across cloud and on-premise environments. Responsibilities include infrastructure automation, monitoring, security, performance tuning, capacity planning, cost optimization, migration support, database administration, Microsoft infrastructure services, troubleshooting, and on-call support. The role requires strong scripting, Terraform or CloudFormation, cloud observability, AWS and Azure administration, MSSQL and MySQL, and infrastructure migration experience. AWS Solutions Architect certification is required.
Top Skills: Active DirectoryAWSAws CloudwatchAzureAzure AdAzure Ad ConnectCloudFormationDhcpDnsDockerGrafanaMssqlMySQLNagiosNew RelicPowershellPrtgPuppetPythonRed HatSaltShellSolarwindsTerraformTomcat

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account