Akamai Technologies Logo

Akamai Technologies

Senior Site Reliability Engineer

Reposted Yesterday
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Designs, develops, and operates infrastructure, observability platforms, and internal tooling for Akamai’s Compute products. Responsibilities include proactive troubleshooting, automation, systems programming, Kubernetes and production-system operations, reliability and scalability improvements, cross-team collaboration, and guidance on service performance. The role requires Linux administration, networking knowledge, CI/CD, infrastructure as code, configuration management, cloud storage exposure, and scripting expertise.
The summary above was generated by AI

Do you like collaborating across teams to solve complex problems?

Do you enjoy solving large scale distributed content delivery challenges?

Join our highly skilled Compute Site Reliability team

Our team designs, develops, and manages applications and infrastructure that support Akamai's Compute products and services. We specialize in creating solutions that help improve observability and enforce SLAs across all internal teams.

Partner with the best

As a Site Reliability Engineer Senior, you will collaborate across operations teams and application development teams. Together, you will be creating tooling and software that monitors and improves the reliability of our systems. You'll work with a diverse range of technologies as we release new applications and modernize existing tooling

As a Site Reliability Engineer Senior, you will be responsible for:

  • Solving complex problems in a timely and accurate manner through proactive troubleshooting, automation and systems programming
  • Deploying and maintaining our observability platform and internal tooling
  • Partnering across teams to ensure the reliability, scalability and usability of our products and services
  • Providing guidance to engineers and developers to increase confidence that their services are performing as expected
  • Collaborating with our support, operations, and engineering teams to investigate and troubleshoot complex problems

Do what you love

To be successful in this role you will:

  • Have a Bachelor's degree in Computer Science, Engineering
  • Have 6 years of experience in Site Reliability Engineering or a related engineering role, with Linux system administration expertise.
  • Have an understanding of networking fundamentals, including TCP/IP, DNS, routing/switching, and storage concepts.
  • Have hands-on experience with containerized environments and Kubernetes, including operating and troubleshooting production systems.
  • Have knowledge of CI/CD and DevOps practices, with hands-on experience using tools such as Jenkins, Git, Prometheus, and Grafana.
  • Have experience with Infrastructure as Code and configuration management using Terraform, Ansible, SaltStack, or similar tools; exposure to cloud storage systems
  • Have automation/scripting skills using Python, Bash, Go, Rust, or similar languages.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Similar Jobs

24 Days Ago
Remote or Hybrid
Senior level
Senior level
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Establish SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems. Operate Kubernetes and GCP environments, implement Infrastructure as Code, optimize capacity and performance, and drive vulnerability remediation, penetration-testing support, and cloud platform security.
Top Skills: Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonShellTerraform
10 Days Ago
In-Office or Remote
Senior level
Senior level
Software
Owns Kubernetes-based development, CI, pre-production, and customer-facing production environments. Responsibilities include SRE operations, incident response, on-call support, Helm and CI/CD lifecycle management, infrastructure automation, observability, security, disaster recovery, stateful platform services, and multi-region deployments. The role also provides technical consultation, creates operational documentation, validates upgrades, and mentors engineers.
Top Skills: AnsibleApi GatewaysArgo CdAWSCluster ApiDockerFluxGithub ActionsGitopsGoGrafanaHelmKafkaKeycloakKindKubernetesKyvernoMetal3OpaOpenstackOpentelemetryPostgresPrometheusPytestPythonTemporalTerraform
10 Days Ago
In-Office or Remote
Senior level
Senior level
Internet of Things • Mobile • Retail
Own platform reliability, DevOps automation, CI/CD, Kubernetes deployments, observability, logging, cloud infrastructure governance, and incident response. Lead Kafka and streaming-platform operations, capacity and disaster-recovery planning, access controls, cost governance, and reliability improvements. Provide advanced production troubleshooting for microservices and mentor Tier 1 and Tier 2 support teams.
Top Skills: AlertmanagerApache AirflowApache FlinkAws MskAzureAzure Container Registry (Acr)Azure Event HubsAzure Kubernetes Service (Aks)Azure MonitorCi/CdConfluent CloudConfluent KafkaFluent BitGithub ActionsGrafanaHelmJavaJfrogKubernetesOpensearchPostgresPrometheusPythonReactSpring BootThanos

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account