JPMorganChase Logo

JPMorganChase

Site Reliability Engineer II

Posted 49 Minutes Ago
Be an Early Applicant
Hybrid
Bengaluru, Bengaluru Urban, Karnataka
Junior
Hybrid
Bengaluru, Bengaluru Urban, Karnataka
Junior
Supports reliability and incident management for network services by troubleshooting infrastructure, building Python, Shell, and Ansible automation, reducing operational toil, and improving observability and SRE practices. The role works across routing, switching, firewalls, load balancers, proxies, SD-WAN, cloud infrastructure, and operating systems while applying authorized AI tools to incident analysis and reliability improvements.
The summary above was generated by AI
Play a key role in ensuring system reliability at one of the world’s most iconic and largest financial institutions. 
As a Site Reliability Engineer II at JPMorgan Chase within the Enterprise Technology - Infrastructure Platforms team, you will use technology to solve business problems and leverage software engineering best practices as we strive towards excellence. This role often works independently to execute small to medium projects, but you’ll also have the opportunity to collaborate with cross functional teams to continually improve your level of knowledge about JPMorgan Chase’s business and relevant technologies. 
Job responsibilities
  • Contribute to problem management activities: evidence collection, timeline building, contributing to RCA and action items
  • Build and enhance automation using Python, Shell scripting, and Ansible (e.g., basic health checks, data collection scripts, configuration validation
  • Uses enterprise-authorized AI capabilities within the work environment to speed up incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Participate in incident management for network services: monitoring, triage, troubleshooting, mitigation support, and escalation.
  • Recognizes toil within the role and proactively works towards eliminating it through systems engineering or updating application code
  • Support and troubleshoot core networking domains: Routing and Switching, Firewalls, Load Balancers, Proxies and SD-WAN, SDA, and broader software-defined networking (SND) concepts
  • Applies enterprise-authorized AI capabilities within the work environment to identify recurring toil and reliability risks from operational signals, prioritizing reuse-first improvements and measurable SLO outcomes.
  • Demonstrate SRE mindset: learn and apply concepts such as reliability, toil reduction, and operational readiness; understand NFRs and introductory FMEA concepts
  • Supports the adoption of site reliability engineering best practices within your team
  • Should complete SRE Bar Raiser Program
 
Required qualifications, capabilities, and skills

 

  • Formal training or certification on Site Reliability concepts and 2+ years applied experience  
  • Experience or strong exposure to network operations/engineering and incident/problem management practices.
  • Hands-on automation skills in Python and/or Shell, and Ansible for configuration/operational tasks.
  • Practical troubleshooting across routing/switching and at least one of: firewall, load balancer, proxy.
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows (e.g., troubleshooting support and runbook drafting) with strong validation habits and awareness of data sensitivity. 
  • Ability to assess AI-assisted operational recommendations for correctness and risk, and apply appropriate controls to maintain resiliency, security, and auditability.
  • Experience maintaining a cloud-based infrastructure
  • Familiar with site reliability concepts, principles, and practices
  • Familiar with observability such as white and black box monitoring, service level objective alerting, and telemetry collection
  • Familiarity with containers or a common server OS such as Linux and Windows
  • Emerging knowledge of continuous integration and continuous delivery practices and related tooling
 
 
 
Preferred qualifications, capabilities, and skills
 
  • Exposure to Cisco ACI / Fabrics.
  • Certifications such as CCNA and/or other vendor certifications.
  • Experience in a financial institution environment

Similar Jobs

2 Days Ago
In-Office
Mid level
Mid level
Artificial Intelligence • Machine Learning • Software • Analytics
Build, operate, and improve reliable production infrastructure across AWS and Kubernetes. Automate infrastructure with Terraform, develop CI/CD and GitOps workflows, and use Python, Go, or scripting to reduce operational toil. Monitor systems through metrics, logs, traces, and alerts; participate in on-call, incident response, root-cause analysis, and remediation. Partner with application, platform, and security teams to improve scalability, reliability, SLIs, SLOs, and operational efficiency.
Top Skills: Amazon CloudwatchAmazon EksArgocdAWSBashCi/CdDatadogGithub ActionsGitlab CiGitopsGoGrafanaJenkinsKubernetesLinuxPrometheusPythonTerraform
11 Days Ago
In-Office
Entry level
Entry level
Financial Services
Entry-level Site Reliability Engineer role supporting CME’s Globex trading platform. Responsibilities include learning observability, monitoring, alerting, automation, disaster recovery, resiliency testing, cloud migration, and production reliability practices. The engineer will write basic scripts, participate in blameless post-mortems, collaborate with product teams, and eventually join an on-call rotation under senior-engineer support. Strong foundational programming, problem-solving, communication, teamwork, and willingness to learn are required.
Top Skills: AWSAzureBashC++DockerGoogle Cloud Platform (Gcp)GrafanaHTTPJavaKubernetesLinuxPrometheusPythonSplunkTcp/Ip
18 Days Ago
In-Office
Entry level
Entry level
Artificial Intelligence • Fintech • Machine Learning • Financial Services
Own and scale AWS production infrastructure, improve CI/CD automation, lead incident response and RCA reporting, strengthen security and observability, manage cloud costs, and build cross-region disaster recovery and failover capabilities. The role requires operating Kubernetes clusters, Terraform infrastructure, monitoring and logging platforms, database and messaging systems, and cloud security controls.
Top Skills: Amazon EksAmazon RdsAWSAws Secrets ManagerBashDjangoElkGrafanaIdentity ProvidersKubernetesOpensearchPrometheusPythonRabbitMQSsoTeleportTerraformVaultVictoriametrics

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account