Lead and manage major and high-priority incidents from detection to closure, coordinate cross-functional technical teams, run incident bridge calls, communicate status to stakeholders, conduct post-incident reviews, maintain incident documentation and dashboards, drive problem management and continual service improvement, and ensure ITIL-compliant processes while supporting change activities.
This role is for one of the Weekday's clients
Min Experience: 2+ years
Location: Chennai, Tamil Nadu, India
JobType: full-time
Requirements
Key Responsibilities
- Manage the complete lifecycle of Major and High-Priority Incidents from identification through resolution and closure.
- Lead incident bridge calls and coordinate with application, infrastructure, cloud, network, database, and support teams to expedite resolution.
- Ensure timely incident response, escalation, communication, and resolution in accordance with defined SLAs.
- Monitor incident queues and prioritize issues based on business impact and urgency.
- Provide regular status updates to business stakeholders, customers, and senior leadership during critical incidents.
- Coordinate technical teams to identify workarounds and restore services as quickly as possible.
- Conduct post-incident reviews (PIRs), document root causes, corrective actions, and lessons learned.
- Track recurring incidents and collaborate with Problem Management teams to drive permanent resolutions.
- Maintain incident documentation, dashboards, and reports for operational reviews and management reporting.
- Ensure compliance with ITIL Incident Management processes and organizational governance.
- Support change implementation activities and assess potential impacts during planned maintenance windows.
- Identify opportunities for automation and continual service improvement to enhance operational efficiency.
Required Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- 2–5 years of experience in Incident Management, IT Operations, Service Delivery, or Production Support.
- Strong understanding of IT Service Management (ITSM) processes and ITIL best practices.
- Experience managing Major Incidents in enterprise production environments.
- Excellent communication, coordination, and stakeholder management skills.
- Ability to work effectively under pressure in a fast-paced operational environment.
- Willingness to work 24x7 rotational shifts, including weekends and public holidays.
Required Technical Skills
- ITSM Tools: ServiceNow, BMC Remedy, Jira Service Management, or equivalent.
- Monitoring Tools: Splunk, Dynatrace, AppDynamics, Datadog, Grafana, SolarWinds, or similar.
- Ticketing and Incident Tracking Systems.
- Microsoft Office Suite (Excel, PowerPoint, Word).
- Basic understanding of cloud platforms (AWS, Azure, GCP) is desirable.
- Familiarity with Linux/Windows environments, networking, databases, and enterprise applications is an advantage.
Preferred Qualifications
- Experience supporting large-scale enterprise production environments.
- Exposure to cloud-based applications, microservices, APIs, and DevOps environments.
- Knowledge of Change Management, Problem Management, and Release Management processes.
- Experience working with global teams across multiple time zones.
Preferred Certifications
- ITIL Foundation Certification (Preferred)
- ITIL Intermediate (Good to have)
- Microsoft Azure Fundamentals / AWS Cloud Practitioner (Good to have)
Must-have skills
ITSM, Incident Management
Similar Jobs
Payments
Own the end-to-end major incident lifecycle: act as incident commander for Sev-1 events, improve MTTR and on-call practices, run blameless postmortems, drive reliability metrics and roadmaps, manage incident communications and status updates, and define incident operating models, runbooks, and severity frameworks.
Top Skills:
AlertingDatadogLoggingObservabilityOpsgeniePagerdutyPaging PlatformsPci-DssStatuspageTracing
Agency • Information Technology
Manage L1.5 incident troubleshooting and resolution for infrastructure and applications, monitor alerts and SLAs, initiate and run technical bridges for major incidents, coordinate vendors and technical teams, maintain incident records and reports, perform proactive problem management, and provide shift/on-call support with handovers.
Top Skills:
JIRAMonitoring ToolsMS OfficeScripting/Programming LanguagesServicenowSQL
Food • Logistics
Lead and coordinate major IT incident response: run incident bridge calls, assign ownership, communicate status, ensure SLA adherence, document root cause and remediation, drive automation and process improvements, and provide 24x coverage on a rotating basis.
Top Skills:
Active DirectoryCloudDatadogDhcpDynatraceFirewallsItilLanLinuxNagiosNew RelicRoutingScriptingSplunkSQLTcp/IpWanWindows
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.



