Astreya Logo

Astreya

IT Infrastructure Operations Engineer II

Reposted 6 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in India
Senior level
Remote
Hiring Remotely in India
Senior level
Provide advanced L2 support for enterprise server and network infrastructure: troubleshoot Dell PowerEdge and Cisco devices, perform firmware/BIOS updates, manage hardware break/fix, improve monitoring and alerting, mentor L1 engineers, maintain runbooks/diagrams, participate in change management and post-mortems, and provide 24x7 on-call support in a global environment.
The summary above was generated by AI

About the Job

We are looking for an experienced L2 IT Infrastructure Operations Engineer to provide advanced technical support for our enterprise server and network infrastructure. This mid-level position bridges the gap between frontline support and expert-level engineering, handling escalated incidents, performing complex troubleshooting, and contributing to operational excellence. The ideal candidate will possess hands-on experience with Dell PowerEdge servers, Cisco networking equipment, and enterprise monitoring solutions. You will mentor L1 engineers, participate in change management activities, and collaborate with cross-functional teams to ensure high availability and performance of critical infrastructure in a 24x7 global environment.

Key Responsibilities

  • Provide advanced troubleshooting and fault isolation for escalated server and network incidents, utilizing iDRAC, Redfish, and Cisco CLI tools to diagnose and resolve complex issues.
  • Execute firmware, BIOS, and driver updates on Dell PowerEdge servers following standardized procedures, ensuring minimal service disruption and maintaining system stability.
  • Perform IOS/NX-OS firmware and software updates on Cisco routers and switches, adhering to change management protocols and conducting post-update validation.
  • Manage hardware break/fix procedures for server infrastructure, coordinating with Dell support for warranty claims, parts ordering, and scheduling on-site technician dispatch.
  • Conduct regular network health audits and performance analysis, identifying potential bottlenecks and recommending optimization measures to prevent service degradation.
  • Collaborate with the SRE team to enhance monitoring dashboards and refine alerting thresholds, ensuring proactive detection of infrastructure instability or security events.
  • Mentor and provide technical guidance to L1 engineers, conducting knowledge transfer sessions and assisting with complex ticket resolution to build team capability.
  • Participate in blameless post-mortems following major incidents, contributing to root cause analysis and implementing preventative actions to improve system reliability.
  • Maintain and update operational runbooks, network diagrams, and technical documentation to reflect current configurations and best practices.
  • Support hardware lifecycle management activities including equipment provisioning, asset tracking, and coordination with vendors for hardware returns and repairs.
  • Provide 24x7 on-call support for critical escalations, ensuring rapid response to high-priority incidents affecting production systems.
  • Collaborate with the FTE IT Team Lead on capacity planning activities, providing data-driven insights on infrastructure utilization trends and growth projections.

Required Skills

  • Related field Experience with 5+ years of hands-on experience in enterprise IT infrastructure operations.
  • Strong proficiency with Dell PowerEdge server administration, including hardware troubleshooting, iDRAC/Redfish management, and firmware lifecycle management.
  • Solid experience with Cisco networking equipment (routers, switches), including IOS/NX-OS configuration, troubleshooting, and upgrade procedures.
  • Working knowledge of monitoring and logging tools, with ability to create dashboards, configure alerts, and analyze performance metrics for proactive issue detection.
  • Excellent problem-solving abilities with demonstrated experience in incident management, root cause analysis, and implementing corrective actions in production environments.
  • Industry certifications such as, Dell Server certifications, or ITIL Foundation; ability to work rotating shifts in a 24x7 global support model.

Tools Required

  • Server & Hardware Tools: Dell iDRAC, Lifecycle Controller, OpenManage, RAID/PERC utilities for server provisioning, firmware baselining, and remote management.
  • OS Deployment Tools: PXE boot infrastructure, iDRAC Virtual Media, Windows Server & Linux ISOs with hardening and automation scripts.
  • Network Tools: Cisco IOS CLI, PoE management, VLAN/QoS configuration tools, network monitoring, and bandwidth/latency testing utilities.
  • Automation & Operations Tools: Ansible, Python, CMDB systems, configuration backup tools, and documentation/diagramming platforms for global 24x7 operations.

Similar Jobs

6 Days Ago
Remote
India
Senior level
Senior level
Information Technology
Provide L2 advanced support and escalation for enterprise server and network infrastructure. Troubleshoot Dell PowerEdge and Cisco equipment using iDRAC/Redfish and CLI, perform firmware and IOS/NX-OS updates, manage hardware break/fix and vendor coordination, run network health audits, enhance monitoring with SREs, mentor L1 engineers, maintain runbooks and diagrams, participate in post-mortems, and provide 24x7 on-call support for production systems.
Top Skills: AnsibleBandwidth/Latency Testing UtilitiesCisco CliCisco IosCisco Nx-OsCmdbConfiguration Backup ToolsDell PoweredgeIdracIdrac Virtual MediaLifecycle ControllerLinuxLogging ToolsNetwork MonitoringOpenmanagePoe ManagementPxePythonQosRaid/PercRedfishVlanWindows Server
An Hour Ago
In-Office or Remote
Senior level
Senior level
Artificial Intelligence • Cloud • Software • Big Data Analytics
The Senior Solutions Architect will design and develop scalable data solutions, manage Hadoop clusters, and collaborate with clients to resolve complex data challenges.
Top Skills: AnsibleApache HiveBashChefHadoopJavaNifiPerlPuppetPythonScalaSpark
An Hour Ago
In-Office or Remote
Senior level
Senior level
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Lead and grow engineering teams building large-scale AI/ML platform infrastructure (model serving, training, feature stores). Drive technical strategy, deliver complex backend systems to production, partner with product and stakeholders, uphold high engineering and operational standards, and coach engineers for career growth.
Top Skills: Backend PlatformsDistributed SystemsFeature StoresGpu-Accelerated Model ServingLlmsTraining And Data Access Infrastructure

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account