US Bank Logo

US Bank

Software Engineering Manager - SRE

Reposted 8 Days Ago
Be an Early Applicant
In-Office
Chennai, Tamil Nadu, IND
Expert/Leader
In-Office
Chennai, Tamil Nadu, IND
Expert/Leader
Lead SRE teams to own reliability for distributed systems, APIs, microservices and data pipelines. Define SLOs/SLIs, run incident response, capacity planning, resiliency testing, and automation to reduce toil. Manage people, on-call operations, cross-functional collaboration, platform enablement, and operational governance for production readiness and continuous reliability improvement.
The summary above was generated by AI

At U.S. Bancorp India, we’re on a journey to do our best. We believe it takes all of us to bring our shared ambition to life, and each person is unique in their potential. A career with U.S. Bancorp India gives you a wide, ever-growing range of opportunities to discover what makes you thrive at every stage of your career. Try new things, learn new skills and discover what you excel at—all from Day One.

Job Description
GCC SRE Manager:

Job Description – SRE Manager (Offshore GCC, Chennai)

Role Overview

The SRE Manager is responsible for leading multiple Site Reliability Engineering teams in a global capability center (GCC) environment, ensuring the availability, scalability, resiliency, and operational excellence of critical platforms and services. This role combines strong people leadership, deep technical expertise, and hands-on operational rigor, managing multiple SRE leads and their respective teams while remaining closely engaged in platform reliability strategy and production engineering decisions.

The position plays a critical role in driving service reliability, production stability, observability, automation, performance engineering, and incident response maturity, while aligning offshore execution with global engineering, platform, and business continuity strategies. The role is expected to provide technical direction on system behavior under load, failure handling, release safety, and operational readiness across distributed systems.

Key Responsibilities

Reliability Engineering & Service Operations

  • Own reliability outcomes for distributed systems, APIs, microservices, data pipelines, and critical production platforms, with accountability for availability, latency, throughput, and saturation
  • Define and operationalize SLOs, SLIs, error budgets, alert thresholds, and service health indicators to improve customer experience and engineering accountability
  • Lead production readiness reviews, capacity planning, performance testing, failover validation, chaos/resiliency testing, and disaster recovery preparedness
  • Drive standardization of monitoring, telemetry, distributed tracing, logging, synthetic checks, and runbook practices across services and platforms
  • Partner with engineering teams to reduce operational toil through automation, self-healing workflows, auto-remediation, and reliability-focused platform improvements

Performance Engineering

  • Proactive Performance Optimization: Continuously identify, prioritize, and remediate performance bottlenecks before they impact customers, production stability, or service scalability
  • Capacity Planning & Traffic Forecasting: Analyze current and historical usage trends, traffic patterns, incident data, and load profiles to forecast compute, storage, network, database, and platform resource needs
  • Performance Testing & Baselines: Define load, stress, endurance, and spike testing practices while establishing measurable baselines for latency, throughput, error rate, saturation, and resource utilization
  • Application Profiling & Diagnostics: Drive use of profiling, tracing, APM, logs, telemetry, and performance analytics to identify code, database, network, and infrastructure bottlenecks
  • Collaborative Performance Improvements: Partner with development, architecture, database, infrastructure, and platform teams to improve responsiveness, throughput, scalability, caching strategies, connection handling, query efficiency, and backend resource utilization
  • Performance Regression Prevention: Embed performance checks, release guardrails, and CI/CD validation into delivery processes to detect degradation before production deployment

Incident Management & Operational Excellence

  • Lead major incident response, escalation management, and technical triage for high-severity production events, ensuring rapid mitigation, stakeholder communication, and service recovery in high-pressure, time-critical environments
  • Establish strong practices for root cause analysis, problem management, failure mode analysis, and durable corrective/preventive actions
  • Drive operational governance using metrics such as MTTR, MTTD, change failure rate, incident recurrence, alert noise, and service error budget consumption
  • Partner with infrastructure, application, database, and network teams to proactively identify scaling risks, dependency bottlenecks, and single points of failure

People & Team Management

  • Directly manage multiple SRE leads and senior reliability engineers, providing technical coaching, operational guidance, and performance leadership across teams
  • Build and scale high-performing SRE teams in the GCC environment with strong focus on production ownership, operational engineering, and systems thinking
  • Drive team capability in on-call operations, debugging, incident command, automation development, and platform diagnostics
  • Foster a culture of blameless incident analysis, engineering accountability, and continuous reliability improvement
  • Manage staffing, on-call coverage, skill distribution, and hiring aligned to platform complexity, production demand, and business criticality

Cross-Functional Collaboration

  • Partner with software engineering, infrastructure, database, network, security, and platform teams to improve system stability, deployment safety, and operational readiness
  • Translate business criticality and customer impact into technical reliability priorities, architecture guardrails, recovery objectives, and measurable engineering outcomes
  • Work effectively in a distributed/global operating model, ensuring seamless coordination with onshore engineers, command centers, platform owners, and leadership teams during both steady-state and incident scenarios

Automation & Platform Enablement

  • Promote and govern infrastructure as code, configuration management, CI/CD reliability, release guardrails, policy enforcement, and automated rollback/recovery patterns
  • Enable engineering teams with standardized tooling for observability, deployment validation, incident response, debugging, performance diagnostics, and service dependency analysis
  • Drive platform modernization through reusable automation, reliability frameworks, production diagnostics, and engineering patterns that improve resiliency and reduce mean time to recovery

Basic Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience
  • 12+ years of experience in software engineering, site reliability engineering, production operations, platform engineering, or infrastructure engineering
  • 5+ years of experience leading reliability or production engineering teams, including managing leads or senior engineers in technically complex environments

Preferred Skills & Experience

  • Proven experience managing multiple SRE, production engineering or platform operations, performance engineering teams in a matrix/global setup
  • Strong experience working in offshore/onshore operating models, preferably within Banking and Financial Services
  • Hands-on knowledge of cloud platforms, Kubernetes, Linux systems, networking fundamentals, distributed systems, and infrastructure automation
  • Experience with observability platforms, telemetry pipelines, APM, distributed tracing, incident tooling, configuration management, and service governance
  • Strong understanding of CI/CD pipelines, scripting languages, infrastructure as code, release engineering, automated operational workflows, release guardrails, and production readiness validation
  • Ability to review architecture and operational designs for scalability, fault tolerance, recovery, performance bottlenecks, and dependency constraints
  • Experience leading performance engineering practices, including load testing, stress testing, capacity modeling, performance baselining, traffic forecasting, resource utilization analysis, and performance regression prevention for mission-critical platforms
  • Strong knowledge of application profiling, query optimization, caching strategies, connection pooling, and runtime diagnostics to identify and resolve application, database, and infrastructure bottlenecks
  • Strong stakeholder management and communication skills across global engineering teams and senior leadership

Leadership Competencies

  • Ability to balance technical depth, operational discipline, architecture awareness, and people leadership
  • Strong decision-making skills with a focus on business continuity, failure risk, service reliability, and engineering trade-offs, with the ability to remain calm and effective in high-pressure operational situations
  • Proven ability to influence engineering design and operational practices without direct authority across platform and application teams
  • High ownership mindset with focus on resilience engineering, operational excellence, and predictable service behavior in production

Location & Work Model

  • Location: Chennai (Offshore GCC)
  • Hybrid work model with in-office collaboration expected 3+ days per week

#LI-DNI

If there’s anything we can do to accommodate a disability during any portion of the application or hiring process, please refer to our disability accommodations for applicants.

Posting may be closed earlier due to high volume of applicants.

This is an U.S. Bancorp India posting. U.S. Bancorp India is a part of the U.S. Bank family.

Similar Jobs

Yesterday
In-Office
Chennai, Tamil Nadu, IND
Senior level
Senior level
Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Define, package, price and drive go-to-market for telco cloud and HCP-based managed services. Lead service lifecycle, enable sales and delivery, engage customers and partners, prioritize tools/resources, and optimize profitability and operational scalability.
Top Skills: CnfHybrid CloudHyperscaler Cloud PlatformsMulti-CloudOrchestration And Automation LayersPlatform-As-A-ServiceTelco CloudVnfZero-Touch Operations
Yesterday
Hybrid
Chennai, Tamil Nadu, IND
Junior
Junior
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Processes and fulfills transactional, batch, and self-service data requests within service-level agreements. Validates inputs, system setups, data extraction, cleansing, outputs, and reports using Ab Initio on UNIX. Coordinates with customer engagement and sales teams, resolves data concerns and complaints, ensures quality and turnaround targets, follows documented procedures, and identifies process improvements. The role requires SQL and Python expertise, attention to detail, problem-solving ability, and willingness to work US shifts.
Top Skills: Ab InitioAutosysMS OfficePythonSQLUnix
Yesterday
In-Office
Chennai, Tamil Nadu, IND
Mid level
Mid level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Field-based pharmaceutical sales role responsible for achieving territory sales budgets through HCP engagement, sales planning, analytics, retailer and hospital coordination, KOL management, compliance with SOPs (including adverse event reporting), executing promotional campaigns, mentoring trainees, and using technology for field operations and reporting.

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account