Design, build, operate, and scale reliable cloud-native infrastructure for partner organizations. Manage Kubernetes clusters, CI/CD pipelines, observability platforms, infrastructure as code, and production incident response. Implement SRE practices including SLIs and SLOs, automate operational workflows using Go or Python, support 24/7 on-call coverage, and explore AI, AIOps, and GPU infrastructure. Contribute to technical documentation, presentations, and open-source projects.
Note: If you have already applied to CloudRaft in the last 90 days, we already have your CV/resume on file. Multiple applications from the same candidate will not be considered.
About CloudRaft
CloudRaft is a premier cloud-native consulting and engineering company that helps ambitious startups and digital-first organizations build, scale, and operate mission-critical platforms. We partner with innovators at the forefront of artificial intelligence, developer productivity, observability, digital commerce, and enterprise software—enabling them to accelerate growth with resilient, scalable, and production-ready cloud infrastructure.
Our experience spans organizations developing AI safety and governance platforms, AI Cloud, AI agent ecosystems, developer tooling, observability solutions, digital health products, customer engagement platforms, and technology-driven franchise networks. By combining deep expertise in Platform Engineering, Kubernetes, DevOps, Observability, and Cloud Native technologies, CloudRaft helps high-growth companies move faster, operate more reliably, and focus on building category-defining products.
Job Description
We are looking for passionate Site Reliability Engineers (SREs) to join our growing team. In this role, you will take end-to-end ownership of designing, building, operating, and scaling mission-critical infrastructure for our partners. You will be responsible for ensuring reliability, performance, security, and operational excellence while driving automation, improving system efficiency, and implementing innovative solutions. Working at the intersection of software engineering and operations, you will help create resilient platforms that enable fast-growing organizations to scale with confidence.
Responsibilities
- Manage and maintain Kubernetes clusters across cloud platforms, including OpenShift, Amazon EKS, Azure AKS, and Google GKE.
- Implement and manage CI/CD pipelines using tools such as Jenkins, GitHub Actions, Argo CD, or GitLab CI/CD.
- Design and maintain observability stacks with tools including Prometheus, Grafana, Loki, OpenTelemetry, and related technologies. Be part of the team who support open source projects like Prometheus, Thanos, Mimir, CloudNativePG, Istio and more.
- Optimize system performance and resolve production issues. Be part of the on call roster to provide 24x7 coverage for the critical production systems.
- Implement SRE principles, including Service Level Indicators (SLIs) and Service Level Objectives (SLOs), to uphold system reliability.
- Automate infrastructure and operational tasks using programming languages such as Go or Python, and Infrastructure as Code (IaC) tools like Terraform.
- Apply agentic AIto automate the SDLC lifecycle, AIOps and automation.
- Learn about emerging technologies, including AI, GPU Infrastructure
- Contribute to knowledge sharing through technical writing and presentations.
Qualifications
- Bachelor’s degree in Computer Science, Information Technology, or a related field.
- 2-5 years of experience in SRE, Platform Engineering, or DevOps Engineer.
- Strong expertise in Kubernetes, cloud-native technologies, on-premise and major cloud platforms (AWS, Azure, GCP).
- Proficiency in programming languages such as Python or Go or Node.js.
- Familiarity with CI/CD tools and modern deployment practices.
- Proficiency in one or more open source observability stacks and Infrastructure as Code (Terraform/Pulumi).
- CKA/CKAD Certified (Brownie points!)
- Excellent problem-solving abilities and communication skills.
- Inclination toward open-source contributions is advantageous.
Benefits :
- Competitive salary
- Premium health insurance and various health & wellness benefits from a leading insurance provider through Plum
- Opportunity to work on the latest AI stack and GPU infrastructure
- Collaborative and supportive work environment full of learning
- Chance to take a front seat where you lead and deliver
Similar Jobs
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Establish SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems. Operate Kubernetes and GCP environments, implement Infrastructure as Code, optimize capacity and performance, and drive vulnerability remediation, penetration-testing support, and cloud platform security.
Top Skills:
Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonShellTerraform
Fintech • Real Estate • Software
Leads the architecture, reliability, security, and operational excellence of cloud infrastructure systems. Owns SLOs, monitoring, incident response, on-call practices, and preventative reliability improvements. Drives medium-to-large SRE projects, collaborates across engineering and product, manages project risks, and mentors engineers. Requires deep experience with AWS, Linux, infrastructure as code, Kubernetes, PostgreSQL, observability, CI/CD, cloud security, and production AI/LLM integrations.
Top Skills:
Api GatewaysAWSAws CdkAws Well-Architected FrameworkCi/CdCloudFormationDistributed TracingEmbeddingsGitopsGoIamKubernetesLinuxLlmsLoggingMetricsMicroservicesObservabilityPostgresPythonRagSecrets ManagementService MeshTerraform
Information Technology • Productivity • Software • Manufacturing
Own reliability for major AWS production domains by defining SLOs, building observability and automation, managing capacity and self-healing, and leading complex incident response. Design Terraform modules and progressive delivery pipelines, operate ECS, EKS, Lambda, and PostgreSQL workloads, and reduce operational toil. Establish security and compliance controls, apply governed AI to operations, mentor SRE engineers, and standardize reliability practices across global teams.
Top Skills:
AWSCi/CdCloudwatchDnsDockerEcs FargateEksGenerative AiGithub ActionsIamIso 27001KubernetesLambdaLlmsOpenobserveOpentelemetryPagerdutyPythonRds PostgresqlSoc 2TerraformVpc
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.



