Leads a Site Reliability and Performance Engineering team supporting business-critical B2B and B2C eCommerce platforms across on-premises and public-cloud environments. Responsibilities include team management, technical mentoring, monitoring and observability, performance and load testing, CI/CD automation, infrastructure management, operational issue resolution, cloud adoption, documentation, compliance, and reliability improvements. The role requires 24x7 operational support, including weekend and public holiday coverage.
Role Summary
As a Manager of Information Technology at Staples, you will collaborate with a business-critical team of engineers responsible for the B2B and B2C sites performance and availability of one of the top eCommerce companies in the United States. You will be a key contributor to the success of our Public Cloud Adoption initiative. This program will drive critical technology and tangible business value utilizing the latest cloud technologies. We are looking for a highly motivated and experienced Site Reliability and Performance Engineering leader who wants to grow their career and work with cutting-edge tools and technologies. The candidate must have a proven track record of supporting B2B, B2C sites and their integrations, both on-premises and in the public cloud, with demonstrated expertise in related technologies.
Duties & Responsibilities
Oversee the day-to-day operations of the Site Reliability and Performance Engineering team.
Set clear team goals, supervise, and manage the team.
Provide technical leadership and mentoring to team members.
Engage and collaborate with cross-functional Product, Engineering, Security, Operations, Infrastructure teams and Vendors to improve MTTD and MTTR
Design, develop, and implement infrastructure & application monitoring to ensure optimal platform availability and performance
Design and execute performance testing strategies including load, stress, and capacity planning using tools such as JMeter, Locust, and LoadRunner.
Automate performance testing within CI/CD pipelines to ensure continuous validation.
Research, analyze and recommend approaches for solving challenging operational issues
Develop and maintain robust knowledge documentation for the Site Reliability Engineering team and its partners
Proactively perform analysis and identify opportunities to innovate, automate, improve efficiency, and achieve cost savings
Foster innovation by encouraging new ideas and technologies within the team.
Ensure compliance with company standards and industry best practices.
Periodically review and assess the team's performance, providing feedback and facilitating professional growth.
Requirements
Basic Qualifications
- Bachelor’s degree in Computer Science or related field with continuous and progressive experience
- Minimum of 8 years of related experience working with these technologies:
- Application Performance Management and Monitoring tools such as New Relic, AppDynamics, SiteSpect, and Datadog
- Content Delivery: Akamai
- Infrastructure monitoring tools like Zabbix, and Prometheus
- Databases eg: MongoDB, Oracle, Couchbase, Redis, MySQL
- Frameworks such as Dust/Angular, Nodejs, Springboot
- Log Analytics tools like Splunk, and ELK/Elastic
- Digital experience tools like Fullstory
- Performance Testing tools such as JMeter, Loadrunner, etc.
- Performance tuning experience with Tomcat, Node.js and Spring Boot.
- Strong understanding of non-functional requirements, performance testing processes, and defect tracking.
- 8+ years of experience with Cloud Technologies, at least half of which should be on the Microsoft Azure platform
- Strong hands-on experience with infrastructure and services (systems, network, cloud technology, provisioning, storage, etc)
- Must have strong experience with programming in one or more scripting languages (Python, Azure CLI, or Powershell)
- Hands-on experience with tool sets related to automation, orchestration, and managing infrastructure (Terraform, Puppet, Ansible, or Jenkins)
- Experience with configuring, deploying, and administering infrastructure and application monitoring tools that assist in troubleshooting performance and stability issues in a cloud environment.
- As SRE and EIRE are global operational functions providing 24x7 support, weekend and public holiday coverage is an inherent expectation of these roles.
- Eligible coverage will be offset through compensatory time off, aligned with company policy.
Preferred Qualifications
Master’s degree in Computer Science Software Engineering or a related field.
Certifications in project management or specific software development methodologies.
Experience in working with cross-functional teams and stakeholders at high organizational levels.
Similar Jobs
Fintech
Lead SRE teams to own reliability for distributed systems, APIs, microservices and data pipelines. Define SLOs/SLIs, run incident response, capacity planning, resiliency testing, and automation to reduce toil. Manage people, on-call operations, cross-functional collaboration, platform enablement, and operational governance for production readiness and continuous reliability improvement.
Top Skills:
Auto-RemediationChaos/Resiliency TestingCi/CdCloud PlatformsConfiguration ManagementDistributed TracingIncident Management ToolsInfrastructure As CodeKubernetesLinuxLoggingNetworkingObservabilityRelease EngineeringScripting LanguagesSynthetic ChecksTelemetry
Healthtech • Information Technology • Telehealth
Lead and grow a Chennai-based SRE team, drive reliability, observability, automation, and operational readiness across Linux, cloud, Kubernetes/EKS environments; own IaC, monitoring, alerting, incident response, onboarding, and Agile delivery while reducing operational toil.
Top Skills:
AlertmanagerAmazon EksAnsibleBashCi/CdGoGrafanaIcingaJavaKubernetesLinuxNew RelicOpensearchPrometheusPuppetPythonRubyTerraform
Digital Media • Information Technology • News + Entertainment
Designs, implements, and maintains secure enterprise and data center network infrastructure. Manages Fortinet, Palo Alto, and F5 firewalls and load balancers; supports routing, switching, high availability, disaster recovery, monitoring, incident response, compliance, documentation, and automation. Collaborates with infrastructure teams and vendors, resolves complex production issues, performs root cause analysis, and mentors junior engineers.
Top Skills:
AnsibleBgpF5 Big-IpF5 DnsF5 GtmF5 LtmFortinetGitHsrpIgmpLacpMlagMulticastOspfPalo Alto FirewallsPimPort-ChannelPythonRest ApisStpVpcVrfVrrp
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


.png)
