Akamai Technologies Logo

Akamai Technologies

Senior Site Reliability Engineer

Reposted 4 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Senior SRE responsible for monitoring, analyzing, and improving availability, performance, and reliability of Akamai's Mapping Service. Define KPIs, build tooling to prevent recurrence, collaborate with product engineers on scalable designs, troubleshoot incidents, and use data analysis and network diagnostics to recommend improvements.
The summary above was generated by AI

Do you like collaborating across teams to solve complex problems?

Do you enjoy solving large scale distributed systems problems?

Join the Mapping SRE team

As Akamai Mapping SREs, we manage the reliability, performance, and scalability of a global system routing trillions of daily client requests and tens of terabits of traffic per second. We combine core SRE principles with data science to analyze massive datasets, define critical KPIs, build advanced monitoring infrastructure, and troubleshoot complex production issues. Our unique intersection of skills allows us to precisely measure and continuously optimize how our mapping architecture impacts customer performance.

Partner with the best

As a Site Reliability Engineer, you will collaborate with cross-functional teams to optimize the performance, availability, and reliability of Akamai’s core Mapping Service. In this high-impact role, you will define critical KPIs, advance our monitoring and alerting infrastructure, and architect automated operational responses. Operating at the intersection of systems engineering and data science, you will apply statistical analysis and cutting-edge machine learning to diagnose and solve the internet’s most complex content delivery challenges.

Note: This is a highly strategic engineering role requiring deep, independent analytical skills—not a standard QA, DevOps, or operational position.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Drive Observability: Co-design, manage, and track product SLIs/SLOs to proactively monitor, investigate, and analyze system performance and availability.
  • Leverage Data Insights: Apply advanced analytical skills and statistical insights to identify mapping bottlenecks, resolve reliability challenges, and engineer long-term solutions.
  • Innovate Tooling: Build and deploy internal tools that automate proactive performance tracking and accelerate independent incident diagnosis.
  • Advocate for Reliability: Partner with product engineers to champion scalable, resilient, and highly supportable system architectures.
  • Influence Strategy: Provide data-driven insights to guide executive-level decision-making and identify high-impact technology investments.
  • Resolve Incidents: Collaborate with internal engineering teams to swiftly troubleshoot, root-cause, and resolve complex customer escalations.

Do what you love

To be successful in this role you will:

  • Master’s or PhD in Computer Science or a highly analytical equivalent field.
  • 5+ years of experience in Site Reliability Engineering (SRE) or a related engineering role.
  • Deep mastery of Unix/Linux internals, computer networking protocols, and distributed system design.
  • Professional fluency operating within a command-line UNIX/Linux computing environment.
  • Data & Observability

  • Strong background in statistical data analysis, SQL database querying, and data integrity troubleshooting.
  • Proven ability to transform complex datasets into actionable strategic roadmaps.
  • Practical knowledge of enterprise observability, logging, and alerting systems like Grafana. 
  • Software Engineering & Leadership

  • Coding proficiency in a major backend or scripting language (e.g., Python).
  • Self-motivated communicator capable of articulating complex systems to non-technical stakeholders while managing multiple timelines.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Similar Jobs

2 Days Ago
In-Office or Remote
2 Locations
Senior level
Senior level
Cloud • Enterprise Web • Hardware • Information Technology • Internet of Things • Robotics • Semiconductor
Build, automate, and operate a global cloud platform: develop automation in Go/Python, manage large-scale EKS clusters (Karpenter), author Terraform and Helm IaC, lead incident response and post-mortems, define SLIs/SLOs, implement observability (Datadog/Prometheus/Grafana) and PagerDuty on-call, and develop secure self-service tools to meet SOC2. Night-shift role based in Ahmedabad, India.
Top Skills: Amazon EksAWSCachingDatadogDynamoDBGoGrafanaHelmKafkaKarpenterKubernetesMskPagerdutyPrometheusPythonTerraform
19 Days Ago
In-Office or Remote
India
Senior level
Senior level
Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
Operate and improve the reliability, availability, and performance of large-scale GeForce NOW services. Participate in incident triage and on-call rotations, build automation and tooling, enhance observability (metrics/logs/traces), drive SLO/SRI practices, run postmortems, and design/operate Kubernetes-based services across cloud and datacenter environments.
Top Skills: AWSAzureBashContainerizationElk/OpensearchGCPGoGrafanaKubernetesMicroservicesOpentelemetryPrometheusPython
20 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
Design, implement, and maintain reliable, scalable infrastructure for large distributed content delivery systems. Define and measure SLIs/SLOs, monitor availability and performance, troubleshoot incidents, and implement corrective actions. Develop automation to reduce manual work, participate in design reviews, and collaborate with product and engineering teams to improve system reliability and performance.
Top Skills: AdbmsBashCloud ComputingDatadogGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account