Akamai Technologies Logo

Akamai Technologies

Site Reliability Engineer II

Posted 3 Hours Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Junior
In-Office or Remote
Hiring Remotely in India
Junior
Deploy and maintain observability platforms and internal tooling for Akamai security products. Improve reliability, scalability, monitoring, alerting, log aggregation, and automated remediation across cloud and Kubernetes environments. Collaborate with support, operations, and engineering teams to troubleshoot complex issues, guide service performance improvements, manage GitOps and CI/CD workflows, and participate in on-call rotations for service restoration.
The summary above was generated by AI

Are you passionate about enhancing the reliability and efficiency of systems and services?

Do you thrive on leading and inspiring team to achieve operational excellence and deliver outstanding results?

Join Our Zero Trust Security Team

At SIA Enterprise, protective measures are developed utilizing Akamai's real-time cloud security intelligence to enhance digital safety effectively.

Partner with the best

As a Site Reliability Engineer to ensure Akamai's security products operate efficiently. Maintain system reliability, enhance monitoring, alerting, and log aggregation. Develop advanced tools while leveraging technologies like Kubernetes, Kafka, and ClickHouse. Collaborate with top cloud platforms, including AWS, Azure, and Linode, to drive continuous improvements and optimize performance.

As a Site Reliability Engineer II, you will be responsible for:

  • Deploying and maintaining our observability platform and internal tooling
  • Partnering across teams to ensure the reliability, scalability and usability of our products and services
  • Providing guidance to engineers and developers to increase confidence that their services are performing as expected
  • Collaborating with our support, operations, and engineering teams to investigate and troubleshoot complex problems
  • Improving monitoring and analysis platforms to ensure rapid error detection and remediation, including developing automated remediation
  • Participating in on-call rotations, guiding restoration and repair of service-impacting issues

Do what you love

To be successful in this role you will:

  • Have 2+ years of relevant experience and a Bachelor's degree in Computer Science or related field
  • Demonstrate a proven ability to design and implement a comprehensive monitoring and observability strategy for complex software products
  • Utilize ArgoCD to implement GitOps-based deployments while overseeing Kubernetes application delivery workflows effectively and efficiently.
  • Create and manage CI/CD workflows leveraging GitHub Actions to automate processes, including building, testing, and deploying pipelines effectively.
  • Have prior experience in defining and implementing frameworks and tools
  • Demonstrate expertise with Kubernetes, Docker, and Third Party clouds such as AWS, Azure, or Linode.
  • Have experience working on Linux-based infrastructure, Bash/Python
  • Demonstrate expertise in problem-solving and troubleshooting across network, system, application, and database layers, including network protocols, Linux OS, and SQL/KQL.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Similar Jobs

12 Days Ago
Remote
India
Mid level
Mid level
Information Technology • Productivity • Software • Manufacturing
Build and operate reliable AWS-based SaaS platforms through observability, infrastructure automation, CI/CD, container and serverless operations, incident response, security controls, and self-healing systems. The role owns services end to end, improves SLOs and incident metrics, maintains Terraform and GitHub Actions automation, participates in 24/7 on-call, and mentors engineers while applying governed AI-assisted engineering practices.
Top Skills: Amazon CloudwatchAmazon Ec2Amazon Ecs FargateAmazon EksAmazon Rds PostgresqlAmazon S3Amazon VpcAWSAws IamAws LambdaBashCi/CdDnsGitGithub ActionsGrafanaOpenobserveOpentelemetryPagerdutyPrometheusPythonTerraform
13 Days Ago
Remote
India
Senior level
Senior level
Software
The Site Reliability Engineer will design scalable infrastructure, automated deployment pipelines, monitoring, and alerting systems. Responsibilities include troubleshooting production incidents, participating in on-call rotations, maintaining systems, improving security and compliance, optimizing reliability and efficiency, and mentoring junior engineers. The role requires collaboration with development and cross-functional teams, technical communication, infrastructure automation, and expertise in distributed systems, cloud platforms, and networking.
Top Skills: AnsibleAWSAzureChefDockerGCPGoGrafanaJavaKubernetesNagiosPrometheusPuppetPythonRubyTerraform
13 Days Ago
In-Office or Remote
India
Junior
Junior
Cloud • Security • Software • Cybersecurity
Build and improve reliable, scalable distributed content delivery systems. Define SLIs and SLOs, enhance monitoring and alerting, analyze performance data, resolve complex incidents, automate operational tasks, and participate in architecture reviews. The role requires scripting, Oracle SQL analysis, Unix/Linux expertise, and experience with observability tools including Prometheus, Grafana, ADBMS, and Datadog. Collaboration with product and cross-functional teams is central to ensuring high availability, performance, and resilience.
Top Skills: AdbmsBashCloud ComputingDatadogDevOpsGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account