Akamai Technologies Logo

Akamai Technologies

Senior Site Reliability Engineer

Reposted Yesterday
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Lead reliability and performance investigations across Akamai's global edge and media/web delivery platforms. Troubleshoot distributed systems across application, network, platform, and OS layers; design observability (SLIs/SLOs, telemetry, dashboards, alerts); analyze performance and bottlenecks; develop automation, internal tools, and AI-assisted diagnostics; partner with engineering and operations on scalable fixes; and provide escalation and off-hours support during critical incidents.
The summary above was generated by AI

Do you like collaborating across teams to solve complex problems?

Do you enjoy solving large scale distributed content delivery challenges?

Join our critical Edge Reliability Engineering Team!

Site Reliability Engineers at Akamai leverage software engineering, systems expertise, and operational skills to deliver reliable global services. The ERE team ensures performance, resilience, and availability of Akamai's media and web delivery platform while addressing distributed systems challenges. This role acts as the top technical escalation point for critical customer-impacting issues, connecting engineering, operations, and global support teams effectively.

Partner with the best

In this role, you will balance deep system diagnostics with progressive automation. You will serve as a premier technical authority for our platform and a champion for reducing systemic operational toil.

As a Senior Site Reliability Engineer, you will be responsible for:

  • Leading complex reliability and performance investigations across Akamai's global edge, media delivery, and web delivery platforms.
  • Troubleshooting critical distributed systems issues spanning application, platform, network, and operating system layers, serving as the highest technical escalation point.
  • Partnering with Engineering, Product, Support, and Network teams to identify root causes and deliver scalable, long-term solutions that improve platform reliability.
  • Designing and improving observability through SLIs, SLOs, KPIs, telemetry, dashboards, and alerts to identify and address customer-impacting issues.
  • Analyzing platform performance, traffic patterns, and system bottlenecks to improve scalability, resilience, and overall service reliability.
  • Developing automation, internal tools, AI-assisted diagnostics, and self-service workflows to streamline operations, reduce manual effort, and accelerate incident response.
  • Enhancing operational excellence through reliability-centered architecture reviews, post-incident analysis, continuous improvements, and offering off-hours support during critical incidents as needed.

Do what you love

To be successful in this role you will:

  • Possess Bachelors in CS/Engineering or a related field with 6 years of industry experience in large-scale SRE/Systems Infrastructure roles.
  • Have logical reasoning skills diagnosing complex performance bottlenecks, data integrity anomalies, and system failure modes in distributed environments.
  • Have understanding of internet technologies and foundational networking concepts, including caching, proxies, TLS, TCP/IP, DNS, and HTTP/HTTPS architectures.
  • Have foundation in Linux/Unix administration, diagnostic tools, and low-level environment troubleshooting.
  • Be able to retrieve data, analyze telemetry streams, and troubleshoot platform data integrity issues through SQL queries.
  • Have experience developing automation tools using languages like Python, Bash, or Go.
  • Demonstrate expertise in AI models and focus on implementing agentic workflows to reduce operational inefficiencies effectively.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Similar Jobs

5 Days Ago
Remote
Shri Bhrigukshetra, BLR, Uttar Pradesh, IND
Senior level
Senior level
Fintech • Analytics
Senior SRE responsible for service availability, performance, and scalability. Build automation and IaC, improve observability and reliability, participate in on-call rotations, incident response, postmortems, cloud migration enablement, and partner with development teams to improve release velocity.
Top Skills: AWSAzureBigpandaCi/CdDatadogDockerDynatraceEntraidGitKubernetesPythonShellTerraform
12 Days Ago
Remote
India
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Cybersecurity
Design, build, operate, and scale cloud-native infrastructure and Kubernetes clusters across major clouds. Implement CI/CD, observability, IaC, SLIs/SLOs, automate with Go/Python, participate in 24x7 on-call, and contribute to open-source and technical knowledge sharing.
Top Skills: AiopsAmazon EksArgo CdAzure AksCloudnativepgGithub ActionsGitlab Ci/CdGoGoogle GkeGpuGrafanaIstioJenkinsKubernetesLokiMimirNode.jsOpenshiftOpentelemetryPrometheusPulumiPythonTerraformThanos
22 Days Ago
In-Office or Remote
2 Locations
Senior level
Senior level
Cloud • Enterprise Web • Hardware • Information Technology • Internet of Things • Robotics • Semiconductor
Build, automate, and operate a global cloud platform: develop automation in Go/Python, manage large-scale EKS clusters (Karpenter), author Terraform and Helm IaC, lead incident response and post-mortems, define SLIs/SLOs, implement observability (Datadog/Prometheus/Grafana) and PagerDuty on-call, and develop secure self-service tools to meet SOC2. Night-shift role based in Ahmedabad, India.
Top Skills: Amazon EksAWSCachingDatadogDynamoDBGoGrafanaHelmKafkaKarpenterKubernetesMskPagerdutyPrometheusPythonTerraform

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account