GoGuardian Logo

GoGuardian

Site Reliability Engineer II

Posted 4 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in India
Junior
Remote
Hiring Remotely in India
Junior
Build and maintain scalable AWS cloud infrastructure, observability and monitoring systems, deployment pipelines, and automation. Participate in on-call rotations, incident response, post-mortems, and root-cause analysis. Support product engineering teams with infrastructure troubleshooting and operational best practices while applying security and compliance controls. The role requires experience with Terraform, CI/CD, Linux, shell scripting, cloud services, managed Kubernetes, databases, and debugging application code.
The summary above was generated by AI
What We Do
 
At GoGuardian, we’re helping build a future where all learners are ready and inspired to solve the world’s greatest challenges. Our award-winning system of learning solutions is purpose-built for K-12 and trusted by school leaders to promote effective teaching and equitable engagement while helping empower educators to keep students safe. 
 
What It’s Like to Work at GoGuardian

We are an outcomes-focused learning company with a steadfast focus on improving learning environments, one classroom at a time. Working with us means joining a remote team of diverse, committed, mission-driven employees who are inspired by our vision, dedicated to our customers, and ready to roll up their sleeves. Guardians put their heads together to solve problems, learn together from experiments that fail, and stand together by their work with full accountability. We balance our diligence with an inclusive culture that invites everyone to bring their whole self to work. Join us and learn why “I love the people here” is one of the most frequent comments we hear from Guardians.

The Role

We’re looking for a Site Reliability Engineer (SRE) II to help build, maintain, and scale the infrastructure that powers our core products and services. In this role, you’ll work alongside engineering teams to support operational excellence, optimize system performance, and ensure high availability across production environments. This position sits on Tech Foundation, a team that manages core cloud infrastructure, shared data services, and developer tooling to empower our product teams to deliver software efficiently and securely. The ideal candidate has practical experience with cloud infrastructure, automation, and core reliability practices, with a strong desire to solve complex operational challenges in a collaborative environment.

_________________________________________________________________________________________

What You'll Do

  • Build, maintain, and support scalable cloud infrastructure to ensure high availability for core products.
  • Maintain and enhance observability and monitoring frameworks to deliver accurate alerts and improve incident detection.
  • Participate in on-call rotations and support incident response, helping conduct post-mortems and trace RCAs to mitigate future issues.
  • Maintain and improve deployment pipelines and automation scripts to support operational safety and engineering velocity.
  • Collaborate with product development teams to provide day-to-day infrastructure support and assist with operational best practices.
  • Apply established security standards and compliance controls across all managed cloud infrastructure.

Who You Are

  • 2+ years of professional experience in Site Reliability Engineering, Infrastructure, or DevOps roles supporting production SaaS applications.
  • Working knowledge of AWS core services (including EC2, VPC, S3) along with exposure to Serverless frameworks or managed Kubernetes environments like EKS.
  • Hands-on experience writing and maintaining Infrastructure as Code (IaC) using Terraform.
  • Familiarity with supporting, troubleshooting, or monitoring data layers such as MongoDB, Redshift, or OpenSearch.
  • Experience working with CI/CD tools and deployment workflows using systems like Jenkins, AWS CodeBuild/CodePipeline, or GitHub Actions.
  • Solid understanding of Linux operating system fundamentals and Unix shell scripting.
  • Ability to read and debug code written in JavaScript/TypeScript, Python, or Go to assist in troubleshooting application-level errors.
  • Strong communication and interpersonal skills, with a collaborative approach to solving technical problems across teams.
  • Eager to take initiative in a fast-paced, ever-changing, dynamic environment.
  • Fueled by the opportunity to truly impact the education landscape.
  • Something else? Tell us! We want to learn more about you…

Please share this with your friends or co-workers who may be interested in working at GoGuardian! We have multiple openings and are always looking for talented people. 

 
GoGuardian is an equal opportunity employer and makes employment decisions on the basis of merit and business needs. GoGuardian does not discriminate against employees, applicants, interns or volunteers on the basis of race, religion, color, national origin, ancestry, physical disability, mental disability, medical condition, pregnancy, marital status, sex, age, sexual orientation, military and veteran status, registered domestic partner status, genetic information, gender, gender identity, gender expression, or any other characteristic protected by applicable law.
 
GoGuardian's Job Applicant Privacy Policy is located here
 
#BI-Remote

Similar Jobs

Yesterday
Remote
India
Mid level
Mid level
Information Technology • Productivity • Software • Manufacturing
Build and operate reliable AWS-based SaaS platforms through observability, infrastructure automation, CI/CD, container and serverless operations, incident response, security controls, and self-healing systems. The role owns services end to end, improves SLOs and incident metrics, maintains Terraform and GitHub Actions automation, participates in 24/7 on-call, and mentors engineers while applying governed AI-assisted engineering practices.
Top Skills: Amazon CloudwatchAmazon Ec2Amazon Ecs FargateAmazon EksAmazon Rds PostgresqlAmazon S3Amazon VpcAWSAws IamAws LambdaBashCi/CdDnsGitGithub ActionsGrafanaOpenobserveOpentelemetryPagerdutyPrometheusPythonTerraform
2 Days Ago
Remote
India
Senior level
Senior level
Software
The Site Reliability Engineer will design scalable infrastructure, automated deployment pipelines, monitoring, and alerting systems. Responsibilities include troubleshooting production incidents, participating in on-call rotations, maintaining systems, improving security and compliance, optimizing reliability and efficiency, and mentoring junior engineers. The role requires collaboration with development and cross-functional teams, technical communication, infrastructure automation, and expertise in distributed systems, cloud platforms, and networking.
Top Skills: AnsibleAWSAzureChefDockerGCPGoGrafanaJavaKubernetesNagiosPrometheusPuppetPythonRubyTerraform
2 Days Ago
In-Office or Remote
India
Junior
Junior
Cloud • Security • Software • Cybersecurity
Build and improve reliable, scalable distributed content delivery systems. Define SLIs and SLOs, enhance monitoring and alerting, analyze performance data, resolve complex incidents, automate operational tasks, and participate in architecture reviews. The role requires scripting, Oracle SQL analysis, Unix/Linux expertise, and experience with observability tools including Prometheus, Grafana, ADBMS, and Datadog. Collaboration with product and cross-functional teams is central to ensuring high availability, performance, and resilience.
Top Skills: AdbmsBashCloud ComputingDatadogDevOpsGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account