Drivetrain Logo

Drivetrain

Site Reliability Engineer - SRE

Reposted One Month Ago
Be an Early Applicant
Remote
Hiring Remotely in India
Senior level
Remote
Hiring Remotely in India
Senior level
The Senior Site Reliability Engineer ensures the availability, performance, and security of Drivetrain's SaaS platform, managing multi-cloud infrastructure, optimizing CI/CD pipelines, and driving automation.
The summary above was generated by AI
Drivetrain is on a mission to empower businesses to make better decisions. Our financial planning & decision-making platform helps companies scale and achieve their targets predictably.

Drivetrain is a remote-first company headquartered in the San Francisco Bay Area. Founded in 2021 by a couple of ex-Googlers, Drivetrain is a fast-growing company on a trajectory for success with backing from leading venture capital firms.

Drivetrain provides a great culture for its employees to thrive in and be happy. 

💜 Remote-friendly: Drivetrain brings together the best and the brightest, no matter where they are and provides them a great degree of autonomy. We trust our people.
🗣️ Open & transparent:  We know that when our creators have access to all the information they need, their best work will emerge.
👏 Idea-friendly:  We provide an environment to explore new ideas, to take risks, to make mistakes, and to learn, so you can succeed. Anyone in the company can come up with great ideas and become a catalyst for positive change. We let the best ideas win.
👥 Customer-centric:  We follow a product-led growth strategy, continuously  learning from our customers and collaborating to build the amazing software that Drivetrain is.

As a Senior Site Reliability Engineer at Drivetrain, you will be a cornerstone of our engineering organization, ensuring our fast-growing SaaS platform remains highly available, performant, and secure. At this stage of our growth, scaling infrastructure efficiently while maintaining the rigorous security and reliability standards required for financial data is paramount. You will take ownership of our multi-cloud infrastructure, drive automation, champion observability, and collaborate closely with development teams to build a culture of reliability from code commit to production.

Key Responsibilities

Cloud Infrastructure & Orchestration

  • Multi-Cloud Management: Architect, manage, and continuously optimize highly available cloud infrastructure across both AWS and GCP. Balance workload demands to ensure maximum cost-efficiency, scalability, and strict security compliance across both platforms.

  • Advanced Kubernetes Orchestration: Lead the design, deployment, and management of scalable Kubernetes clusters. Utilize configuration management tools like Kustomize to enforce standardized, repeatable, and automated deployment configurations across all environments.

  • Service Mesh & Security Integration: Implement and maintain service mesh technologies (e.g., Istio, Linkerd) to secure, control, and observe service-to-service communication. Drive container security best practices, including image scanning, runtime protection, and strict RBAC enforcement.

CI/CD & Automation

  • Pipeline Engineering: Architect, maintain, and optimize robust CI/CD pipelines using Git and Jenkins. Focus on reducing deployment friction, accelerating release velocity, and enforcing automated testing and security gates.

  • Infrastructure as Code (IaC): Treat infrastructure as software. Write, review, and maintain Terraform modules to provision and manage cloud resources predictably and safely.

  • Operational Automation: Aggressively reduce operational toil. Develop robust Python scripts and tooling to automate routine maintenance, data backups, scaling operations, and system recovery processes.

Observability & Reliability

  • Comprehensive Monitoring: Design and enhance our observability stack to provide deep, real-time insights into system health. Manage and scale tools including Prometheus, Grafana, ELK/EFK stack, AWS CloudWatch, and GCP Operations Suite.

  • Reliability Engineering: Spearhead reliability initiatives critical to a scaling SaaS platform. Drive rigorous capacity planning exercises to stay ahead of growth.

  • Incident Management & SLOs: Own the incident response lifecycle. Facilitate blameless postmortems to extract actionable learnings. Define, track, and enforce SLIs, SLOs, and SLAs, ensuring the platform consistently meets its reliability guarantees.

Collaboration & Leadership

  • DevOps Culture: Act as an embedded reliability advocate. Collaborate closely with software engineers early in the development lifecycle to ensure applications are designed for deployability, scalability, and resilience.

  • Continuous Improvement: Proactively identify system bottlenecks and architectural weaknesses. Contribute to process improvements, build internal developer tooling, and maintain comprehensive documentation to elevate team productivity and system understanding.

Required Proficiency & Qualifications
  • Experience: 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, preferably within a fast-paced SaaS environment.

  • Cloud Platforms: Deep, proven proficiency in AWS (EC2, EKS, RDS, VPC, IAM, S3) AND GCP (GKE, Compute Engine, Cloud SQL, IAM, Cloud Storage). Ability to navigate and optimize multi-cloud architectures.

  • Containerization: Expert-level knowledge of Docker and Kubernetes, including advanced deployment strategies and lifecycle management.

  • Automation/IaC: Strong programming skills in Python and extensive experience with Terraform.

  • Observability: Hands-on expertise building dashboards and alerting systems using Prometheus, Grafana, and log aggregation stacks (ELK/EFK).

  • Networking & Security: Solid understanding of cloud networking (VPC peering, load balancing, DNS) and zero-trust security principles in a containerized environment.

Sounds exciting? Apply at [email protected]. It may just be the next best decision you’ve ever made!

Similar Jobs

24 Days Ago
Remote or Hybrid
Senior level
Senior level
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Establish SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems. Operate Kubernetes and GCP environments, implement Infrastructure as Code, optimize capacity and performance, and drive vulnerability remediation, penetration-testing support, and cloud platform security.
Top Skills: Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonShellTerraform
An Hour Ago
In-Office or Remote
India
Junior
Junior
Cloud • Security • Software • Cybersecurity
Deploy and maintain observability platforms and internal tooling for Akamai security products. Improve reliability, scalability, monitoring, alerting, log aggregation, and automated remediation across cloud and Kubernetes environments. Collaborate with support, operations, and engineering teams to troubleshoot complex issues, guide service performance improvements, manage GitOps and CI/CD workflows, and participate in on-call rotations for service restoration.
Top Skills: ArgocdAWSAzureBashCi/CdClickhouseDockerGithub ActionsGitopsKafkaKqlKubernetesLinodeLinuxPythonSQL
3 Days Ago
Remote
India
Senior level
Senior level
HR Tech • Legal Tech • Software • Consulting
Own and evolve Mitratech’s AWS DevOps and platform engineering practice, including multi-platform CI/CD, Terraform and AWS CDK infrastructure as code, AWS compute, networking, databases, security controls, and AI/analytics operations. Design reusable pipelines, enforce quality and security gates, manage multi-account environments, improve automation reliability, mentor engineers, and participate in on-call support.
Top Skills: Account Factory For TerraformAlbAmazon LinuxAnsibleAuroraAWSAws CdkAws ConfigAws OrganizationsAzureBedrockBitbucket PipelinesCfn-GuardCheckovCloudtrailControl TowerDirect ConnectEc2Ec2 Image BuilderEcrEcs FargateEksEventbridgeGithub ActionsGuarddutyHashicorp VaultIam Access AnalyzerJenkinsLambdaLinearbLinuxMacieNlbOpentofuPackerPrivatelinkPythonQuicksightRdsRedshiftRhelRoute 53S3Secrets ManagerSecurity HubSleuthSnsSqsSsm Patch ManagerStep FunctionsTerraformTransit GatewayTypescriptUbuntuVpcVpn

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account