Designs, monitors, and maintains reliable, scalable WebMethods integration platforms and infrastructure. Responsibilities include improving incident detection and recovery, implementing monitoring, supporting failover and redundancy, troubleshooting operational issues, performing upgrades and migrations, documenting processes, and automating efficiency improvements. The role collaborates with product, engineering, security, operations, infrastructure teams, and vendors while providing 24/7 operational support, including weekend and public holiday coverage.
We are seeking a skilled Site Reliability Engineer (SRE) with expertise in webMethods to join our team. This role bridges the gap between software development and IT operations, focusing on automation, reliability, and performance of integration platforms. You will be responsible for ensuring the availability, scalability, and resilience of systems powered by WebMethods and other enterprise technologies.
Duties & Responsibilities
Engage and collaborate with cross-functional Product, Engineering, Security, Operations, Infrastructure teams and Vendors to improve MTTD and MTTR
Design, develop, and implement infrastructure & application monitoring to ensure optimal platform availability and performance
Research, analyze and recommend approaches for solving challenging operational issues
Maintain fault-tolerant webMethods integrations and infrastructure.
Support automated failover, load balancing, and redundancy strategies.
Develop and maintain robust knowledge documentation for the Site Reliability Engineering team and its partners
Proactively perform analysis and identify opportunities to innovate, automate, improve efficiency, and achieve cost savings
Perform webMethods version upgrades and environment migrations
Requirements
Basic Qualifications
- Bachelor’s degree in Computer Science or related field with continuous and progressive experience
- Minimum of 4 years of related experience working with some of these technologies:
- Hands-on experience with WebMethods Integration Server, Broker, and MWS.
- Experience with Jenkins and CI/CD.
- Experience with Apache ActiveMQ.
- Strong knowledge of Linux operating systems.
- Strong understanding of distributed systems and cloud platforms (Azure, GCP).
- Experience working with agile methodologies – Scrum, Kanban & SAFe (Scaled Agile Framework) principles.
- Excellent troubleshooting, communication, and documentation skills.
- Application Performance Management and Monitoring tools such as New Relic, AppDynamics, SiteSpect, and Datadog
- Infrastructure monitoring tools like Zabbix, and Prometheus
- Databases eg: MongoDB, Oracle, Couchbase, Redis, MySQL
- WebMethods Suite of product version 10.x and 11.x
- Log Analytics tools like Splunk, and ELK/Elastic
- As SRE and EIRE are global operational functions providing 24x7 support, weekend and public holiday coverage is an inherent expectation of these roles.
- Eligible coverage will be offset through compensatory time off, aligned with company policy.
Preferred Qualifications
Awareness of AI/ML applications in observability and incident response. Familiarity with LLMs and AI-driven automation tools. Understanding of AI-enhanced anomaly detection and predictive analytics
Similar Jobs
Digital Media • Information Technology • News + Entertainment
Designs, implements, and maintains secure enterprise and data center network infrastructure. Manages Fortinet, Palo Alto, and F5 firewalls and load balancers; supports routing, switching, high availability, disaster recovery, monitoring, incident response, compliance, documentation, and automation. Collaborates with infrastructure teams and vendors, resolves complex production issues, performs root cause analysis, and mentors junior engineers.
Top Skills:
AnsibleBgpF5 Big-IpF5 DnsF5 GtmF5 LtmFortinetGitHsrpIgmpLacpMlagMulticastOspfPalo Alto FirewallsPimPort-ChannelPythonRest ApisStpVpcVrfVrrp
Digital Media • Information Technology • News + Entertainment
Build, maintain, and improve a software automation platform, including workflows, CI/CD pipelines, and Go- or Python-based tools. Collaborate with stakeholders to identify automation opportunities, troubleshoot incidents, improve reliability, and participate in on-call support. The role also involves system design, technical documentation, mentoring junior staff, software releases, performance analysis, and cross-functional collaboration with quality assurance.
Top Skills:
ArgocdAWSCi/CdDockerGoInfrastructure As CodeJenkinsKubernetesPythonTerraform
Digital Media • Information Technology • News + Entertainment
Builds, configures, maintains, and tests cybersecurity systems and infrastructure. Conducts security assessments, vulnerability testing, audits, and technical analysis across networks, applications, operating systems, and network devices. Supports security architecture, develops policies and documentation, recommends vulnerability remediation, and troubleshoots network stacks at scale. Requires knowledge of secure routing, DNS/DNSSEC, SDN, automation, high-availability architectures, and varied network platforms. Variable schedules, including nights and weekends, may be required.
Top Skills:
Antivirus SystemsCmtsCybersecurityDnsDnssecDocsisFirewallsIntrusion Detection SystemsNetwork AutomationNetwork MonitoringNetwork SecurityOperating SystemsOptical DevicesPonRoutersSdnSecure RoutingSwitchesTcpVcmtsVulnerability Assessment
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

