Leads a senior software engineering team building and operating enterprise-scale infrastructure services using AI and agentic workflows. Owns architecture, production systems, reliability, security, observability, cost optimization, and incident response. Sets engineering standards and governance for AI-enabled software development, establishes validation and security guardrails, writes and reviews production code, mentors engineers, and partners with infrastructure stakeholders to deliver measurable operational outcomes.
As a Director of Software Engineering in Autonomous Infrastructure and AI4Reliability team you will lead a small, senior team designing, building, and operating tools and services which will manage and heal our infrastructure using AI and agentic flows. You will be a hands-on technical leader, owning production systems end-to-end, setting architecture and engineering standards, and fostering the growth of your team. You’ll work AI-native across the software development lifecycle, ensuring correctness, security, reliability, and cost-effectiveness.
Job Responsibilities:
- Lead and grow a small team of senior engineers, fostering a high-performance, high-ownership culture
- Own the design, delivery, and operation of enterprise-scale infrastructure services from architecture through production
- Write and review production code, maintaining a hands-on approach and setting the bar for engineering quality
- Establish AI-native engineering practices with robust validation standards to ensure speed never compromises correctness
- Participate in an on-call rotation and act as an escalation point for production incidents, building operability and observability from the start
- Analyze and optimize systems for scalability, efficiency, reliability, and performance
- Decompose ambiguous problems into clear, executable work for both engineers and AI agents
- Sets direction and governance for agentic AI-enabled engineering and SDLC/TLM automation within a technical area to drive measurable improvements in speed, quality, and operational outcomes (e.g., AI-orchestrated delivery workflows, release readiness controls, automated test modernization, and incident triage acceleration), while establishing guardrails for validation, security, resiliency, traceability, and reuse across teams.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation and support capacity unlock initiatives at scale.
- Define success criteria, measure impact, and hold the team accountable to outcomes
- Partner with stakeholders across the infrastructure organization, resolving technical disagreements and proactively raising risks
Required Qualifications, Capabilities, and Skills:
- Hands-on software engineering background with production coding experience in an industry-standard language (e.g., Python, Go, Java, C++, Rust)
- Experience running production systems at scale, including on-call ownership, incident response, and designing for reliability and operability
- Experience leading or mentoring engineers and setting technical direction
- Strong systems thinking, including interfaces, contracts, failure modes, and interactions at scale
- Ability to direct AI tools for real engineering work, with sound judgment on where AI applies and where human expertise is required
- Security-first mindset, integrating risk judgment from design through production
- Clear, direct communication with engineers, stakeholders, and peers
- Outcome orientation, focused on impact, reliability, and cost
- Comfort operating with ambiguity and greenfield scope
- Inclusive, collaborative leadership style, able to attract, grow, and retain strong senior engineers
Preferred Qualifications, Capabilities, and Skills:
- Track record of reducing operational toil and cost through automation and better engineering
- Experience adopting AI-native engineering practices at team or organizational scale
- Experience leading adoption of agentic AI-enabled engineering practices (using enterprise-authorized tools within the work environment) across teams, including defining operating expectations (human-in-the-loop validation, quality gates), measuring outcomes, and ensuring secure handling of sensitive inputs/outputs.
- Strong understanding of responsible AI use and control expectations in engineering workflows, including data sensitivity, resiliency/security implications, and governance; ability to influence leaders on safe scaling patterns and reuse.
- Prior experience in regulated or large-scale enterprise environments in the payment industries
- Experience with greenfield builds and establishing engineering culture
Similar Jobs
Cloud • Security • Software • Cybersecurity • Automation
The Staff Data Analyst will transform data into insights for GitLab’s Support Engineering organization. Responsibilities include capacity and budget modeling, utilization scorecards, support economics analysis, performance indexes, AI adoption ROI dashboards, predictive churn analytics, customer-schema standardization, executive reporting, and ad hoc analysis. The role partners with Support Engineering, go-to-market, Business Operations, and Product Engineering in a remote environment.
Top Skills:
Artificial IntelligenceLarge Language ModelsLookerMachine LearningSQLTableauZendesk Explore
Automotive
Designs, develops, and supports scalable SAP solutions using ABAP, RAP, CDS, OData, and S/4HANA extensibility. Responsibilities include API-led integrations, Fiori/UI5 backend services, BTP extensions, testing, debugging, performance optimization, code reviews, architecture discussions, and Agile delivery. The role requires maintaining custom SAP developments across S/4HANA and ECC while following clean core principles, enterprise security standards, and SAP best practices.
Top Skills:
Abap Development Tools (Adt)Api ManagementCloud Application Programming Model (Cap)Core Data Services (Cds)EclipseGitIdocsObject-Oriented AbapOdata V2Odata V4Rest ApisRestful Abap Programming Model (Rap)RfcsSap AbapSap Btp Abap EnvironmentSap Business Technology Platform (Btp)Sap EccSap FioriSap GatewaySap Integration SuiteSap S/4HanaSoapUi5Web Services
Automotive
Leads global supplier warranty recovery operations, analyzing claims, recovery trends, costs, and performance metrics. Manages supplier engagement, overdue recoveries, invoices, reconciliations, forecasts, dashboards, and monthly reporting. Partners with Finance, Warranty, Purchasing, Quality, Manufacturing, and Data Analytics teams to identify recovery opportunities, monitor financial targets, resolve bottlenecks, and recommend corrective actions to reduce warranty costs.
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


