Nexcess
Teams at Nexcess
Recently posted jobs
Software
The Platform SRE maintains reliable, secure, and modern hosting infrastructure. Responsibilities include reducing technical debt, managing fleet-wide patching, building infrastructure automation and CI/CD pipelines, defining SLIs and SLOs, improving observability, responding to production incidents, performing root cause analysis, and implementing remediation. The role partners with engineering teams on scalability and operational readiness, supports compliance, maintains runbooks, mentors engineers, and leads platform reliability initiatives across cloud, managed hosting, and hybrid environments.
Software
Coordinates incident, problem, change, and service transition management across engineering, infrastructure, security, and operations teams. Leads incident response activities, post-mortems, corrective-action tracking, reliability reporting, observability initiatives, SLO and SLI development, change governance, CMDB accuracy, and operational readiness. Uses Jira Service Management workflows and automation to improve platform stability, reduce response and resolution times, govern changes, and communicate reliability trends to technical and executive stakeholders.
