Jefferies Logo

Jefferies

Associate - SRE - Platform Engineering

Posted Yesterday
Be an Early Applicant
In-Office
Pune, Maharashtra
Mid level
In-Office
Pune, Maharashtra
Mid level
Build and operate reliable, scalable platforms supporting post-trade processing. Responsibilities include monitoring, incident response, troubleshooting, observability, capacity planning, automation, production support, and release and change management. The role develops dashboards and alerts with Grafana, Prometheus, and OpenTelemetry; supports Kafka-based messaging and distributed systems; analyzes logs, metrics, and traces; and improves resilience by reducing operational toil.
The summary above was generated by AI

Associate Platform Reliability Engineer (SRE)

Location: Mumbai / Pune 

Role Overview

We are seeking a highly motivated Associate Platform Reliability Engineer (SRE) to join our global Platform Reliability Engineering team. This is a hands-on engineering role focused on the reliability, scalability, and operational excellence of critical front-to-back platforms supporting post-trade processing.

The ideal candidate will have a strong software engineering foundation, production support experience, and a passion for automation, observability, and reliability engineering. You will work closely with development, infrastructure, and business teams to improve system resilience, enhance operational visibility, reduce manual intervention, and deliver highly available services.

 

Key Responsibilities

  • Proactively monitor platform health and drive improvements in reliability, performance, availability, and operational efficiency.
  • Perform incident triage, troubleshooting, communication, and post-incident reviews to minimize business impact and prevent recurrence.
  • Collaborate with engineering, infrastructure, and business stakeholders to design and implement scalable and resilient solutions.
  • Build and enhance deployment and observability capabilities, including dashboards, alerts, and service health monitoring using Grafana, Prometheus, and OpenTelemetry.
  • Analyze logs, metrics, and distributed traces to proactively identify system bottlenecks and reliability issues.
  • Support enterprise messaging and event-driven architectures, including Kafka-based platforms and integrations.
  • Apply capacity planning and availability management best practices to improve platform resilience.
  • Develop automation to reduce operational toil, minimize manual intervention, and improve service efficiency.
  • Participate in production support, problem management, release management, and change management activities.

Required Qualifications

  • Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline.
  • 3+ years of experience in Site Reliability Engineering (SRE), Platform Reliability Engineering (PRE), DevOps, Production Support, or Application Support.
  • Strong programming and scripting experience in one or more languages such as Python, Go, or Java.
  • Solid understanding of software engineering principles, data structures, algorithms, and system design.
  • Strong working knowledge of Linux/Unix and Windows Server environments.
  • Experience with modern monitoring and observability practices and tooling.
  • Strong understanding of observability and reliability concepts, including metrics, logs, traces, SLIs, SLOs, and alerting.
  • Good understanding of event-driven architectures and enterprise messaging platforms such as Kafka and MQ.
  • Experience troubleshooting distributed production systems, including APIs, middleware components, and message flows.
  • Understanding incident management, problem management, root cause analysis, and operational support processes.
  • Familiarity with source control, CI/CD pipelines, Infrastructure as Code (IaC), and DevOps practices.
  • Strong verbal and written communication skills with the ability to engage both technical and business stakeholders.
  • Self-motivated, detail-oriented, and capable of working independently in a fast-paced environment.

 

Preferred Qualifications

Observability & Monitoring: Grafana, Prometheus, OpenTelemetry and Loki

DevOps & Automation: Git, Ansible and CI/CD Frameworks

Container & Platform Technologies: Docker, Kubernetes

Data & Messaging Platforms: Kafka, Redis, MQ

Cloud Technologies: AWS

About Us

Jefferies is a leading global, full-service investment banking and capital markets firm that provides advisory, sales and trading, research, and wealth and asset management services. With more than 40 offices around the world, we offer insights and expertise to investors, companies, and governments.

At Jefferies, we believe that diversity fosters creativity, innovation and thought leadership through the infusion of new ideas and perspectives. We have made a commitment to building a culture that provides opportunities for all employees regardless of our differences and supports a workforce that is reflective of the communities where we work and live. As a result, we are able to pool our collective insights and intelligence to provide fresh and innovative thinking for our clients.

Jefferies is an equal employment opportunity employer, and takes affirmative action to ensure that all qualified applicants will receive consideration for employment without regard to race, creed, color, national origin, ancestry, religion, gender, pregnancy, age, physical or mental disability, marital status, sexual orientation, gender identity or expression, veteran or military status, genetic information, reproductive health decisions, or any other factor protected by applicable law. We are committed to hiring the most qualified applicants and complying with all federal, state, and local equal employment opportunity laws. As part of this commitment, Jefferies will extend reasonable accommodations to individuals with disabilities, as required by applicable law.

Similar Jobs

6 Days Ago
In-Office
Mid level
Mid level
Financial Services
Build and operate reliable, scalable post-trade platforms through monitoring, incident response, troubleshooting, automation, observability, and capacity planning. Develop dashboards, alerts, and service health monitoring with Grafana, Prometheus, and OpenTelemetry. Support Kafka-based event-driven systems, production releases, change management, and root-cause analysis while collaborating with engineering, infrastructure, and business stakeholders to improve resilience and reduce operational toil.
Top Skills: AnsibleAWSCi/CdDockerGitGoGrafanaInfrastructure As CodeJavaKafkaKubernetesLinuxLokiMqOpentelemetryPrometheusPythonRedisUnixWindows Server
An Hour Ago
Hybrid
Entry level
Entry level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Monitor cloud infrastructure and production services in a 24x7 operations environment. Respond to incidents, troubleshoot servers and applications, manage escalations, coordinate issue resolution, meet SLAs, document runbooks, conduct shift handoffs, and prepare operational reports. The role requires AWS and Linux expertise, infrastructure monitoring experience, incident management knowledge, and familiarity with networking, security, Docker, and Kubernetes.
Top Skills: Amazon Ec2Amazon S3Amazon VpcAWSAws IamDockerDynatraceElastic Load BalancingFtpHTTPHttpsIptablesKubernetesLinuxSelinuxSftpSmtpSplunkSshSsl/TlsTcp/IpUdpVi/Vim
An Hour Ago
Hybrid
Senior level
Senior level
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Designs, automates, and optimizes enterprise Windows platforms using SCCM, Ansible, PowerShell, and Infrastructure as Code. Responsibilities include OS deployment, patching, software distribution, compliance enforcement, server build pipelines, CI/CD integration, self-service delivery, troubleshooting, capacity planning, and platform modernization. The role leads Wintel engineering workstreams, establishes automation standards, collaborates with infrastructure and security teams, and supports scalable, secure, and highly available services.
Top Skills: Active DirectoryAnsibleAWSAzureCertificate ServicesCi/CdCis BenchmarksConfluenceDhcpDnsGitGitopsGroup PolicyHyper-VIntuneItilJIRAKanbanLoad BalancingMicrosoft Endpoint Configuration ManagerPkiPowershellQualysSccmScrumSplunkTcp/IpTerraformVMwareWindows ServerWsus

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account