Ensure reliability, scalability, and performance of the enterprise data platform. Automate IaC, CI/CD, and deployments; manage observability, governance, FinOps, and incident response; collaborate with engineering and analytics teams to improve platform reliability and developer experience.
Role Overview
We are looking for a Data Platform Reliability Engineer to ensure the reliability, scalability, and performance of our enterprise data platform. This role blends Site Reliability Engineering (SRE) and DevOps practices to support modern cloud-based data ecosystems built on Google BigQuery, lakehouse architectures, and distributed data pipelines.
You will play a key role in building highly resilient, governable, observable, and cost-efficient data platforms, while enabling engineering teams to operate at scale with automation and best practices.
Key Responsibilities
- Automate infrastructure provisioning and operations using Infrastructure as Code (IaC)
- Implement and manage CI/CD pipelines for data and platform deployments
- Implement and manage Data Governance tools such as Collibra
- Improve system resilience through capacity planning, performance tuning, and fault tolerance design
- Optimize cloud usage and costs through FinOps best practices
- Collaborate with engineering and analytics teams to improve platform reliability and developer experience
- Drive security, compliance, and access control best practices
- Own platform reliability and availability for enterprise data systems (SLAs, SLOs, error budgets)
- Monitor and manage data cloud infrastructure, ingestion frameworks, and transformation workflows
- Lead incident management, root cause analysis (RCA), and postmortems
Required Skills & Experience
- 3–5 years of experience in SRE, DevOps, or platform engineering roles
- Strong experience with cloud platforms (GCP preferred; AWS/Azure is a plus)
- Experience with CI/CD tools (GitHub Actions, Jenkins, etc.)
- Proficiency in Python, Bash, and SQL
- Experience with Infrastructure as Code tools (Terraform preferred)
- Strong understanding of monitoring and observability tools (logs, metrics, tracing)
- Experience managing production incidents and on-call rotations
- Exposure to Kubernetes / containerization
- Understanding of data governance and security practices
- Exposure to AI/ML or GenAI tools for automation and operational efficiency
Preferred Skills
- Experience with data observability tools/frameworks
- Knowledge of Collibra or similar Data Governance tools
- Knowledge of Looker / Tableau / Power BI or similar BI tools
- Understanding of cost optimization (FinOps) in cloud data platforms
What You Will Bring
- Strong ownership mindset and bias for automation
- Ability to troubleshoot complex distributed systems under pressure
- Passion for improving system reliability and performance at scale
- Excellent collaboration and communication skills
- Willingness to participate in a 24/7 support/on-call rotation
Why Join Us
- Build and scale next-gen cloud data platforms powering enterprise-wide analytics
- Join Pearson, a global leader transforming lives through learning and innovation
- Work hands-on with GenAI, AI-powered data engineering, and intelligent automation
- Shape the future of self-service data, data products, and observability at scale
- Collaborate with high-performing global teams solving real-world, high-impact problems
Accelerate your career in a fast-evolving, innovation-driven data ecosystem
#LI-P1
Similar Jobs
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Provide Workday HCM and Recruiting production support: resolve cases, perform light configurations, run EIB data loads, maintain SOPs, monitor integrations, collect SOX evidence, test changes, and drive process improvements across global stakeholders.
Top Skills:
Ai TechnologiesEibExcelTicketing SystemsWorkday AbsenceWorkday HcmWorkday IntegrationsWorkday RecruitingWorkday Time Tracking
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
The Corporate Account Executive is responsible for driving new business, engaging customers, running sales processes, collaborating with teams, and managing customer success.
Top Skills:
Cloud-Native Security PlatformsCybersecurityEndpoint Detection And Response (Edr) TechnologiesSaaS
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead the design, implementation, and operation of AI Ops and SRE solutions in public cloud environments. Build RAG pipelines, agentic workflows, and AI-powered enterprise applications; automate infrastructure with Terraform and GitHub Actions; manage Kubernetes; lead incident response; ensure reliability, scalability, and performance; establish AI evaluation and monitoring frameworks; and mentor engineering teams.
Top Skills:
AksAWSAzureCi/CdEksGCPGenerative AiGithub ActionsGkeInfrastructure As CodeKubernetesNode.jsPythonRagTerraformVector Search
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


