Build and maintain Azure-based data pipelines for governance and analytics. Develop Databricks ETL/ELT workflows using PySpark, SQL, and Delta Lake; automate data quality, lineage, monitoring, access auditing, cost analysis, retention tracking, policy enforcement, and metadata integration with Microsoft Purview. Maintain curated governance metric tables, optimize performance, troubleshoot pipeline issues, and collaborate with engineering, QA, and reporting teams while meeting security and privacy standards.
Job title : Data Engineer
Experience: 6+ years
Location: Chennai
Shift: US Eastern Time ( 5:00 PM – 2:00 AM )
Requirements
We are seeking a hands-on Data Engineer to develop, optimize, and maintain automated
data pipelines supporting data governance and analytics initiatives. This role will focus on
building production-ready workflows for ingestion, transformation, quality checks, lineage
capture, access auditing, cost usage analysis, retention tracking, and metadata integration,
primarily using Azure Databricks, Azure Data Lake, and Microsoft Purview.
Experience: 6+ years in data engineering, with strong Azure and Databricks experience
Pipeline Development – Design, build, and deploy robust ETL/ELT pipelines in Databricks
(PySpark, SQL, Delta Lake) to ingest, transform, and curate governance and operational
metadata from multiple sources landed in Databricks.
Granular Data Quality Capture – Implement profiling logic to capture issue-level metadata
(source table, column, timestamp, severity, rule type) to support drill-down from dashboards
into specific records and enable targeted remediation.
Governance Metrics Automation – Develop data pipelines to generate metrics for
dashboards covering data quality, lineage, job monitoring, access & permissions, query cost,
usage & consumption, retention & lifecycle, policy enforcement, sensitive data mapping, and
governance KPIs.
Microsoft Purview Integration – Automate asset onboarding, metadata enrichment,
classification tagging, and lineage extraction for integration into governance reporting.
Data Retention & Policy Enforcement – Implement logic for retention tracking and policy
compliance monitoring (masking, RLS, exceptions).
Job & Query Monitoring – Build pipelines to track job performance, SLA adherence, and
query costs for cost and performance optimization.
Metadata Storage & Optimization – Maintain curated Delta tables for governance metrics,
structured for efficient dashboard consumption.
Testing & Troubleshooting – Monitor pipeline execution, optimize performance, and resolve
issues quickly.
Collaboration – Work closely with the lead engineer, QA, and reporting teams to validate
metrics and resolve data quality issues.
Security & Compliance – Ensure all pipelines meet organizational governance, privacy, and
security standards.
Similar Jobs
Digital Media • Information Technology • News + Entertainment
Designs, develops, and manages advanced data architectures, ingestion frameworks, and pipelines. Ensures data quality, lineage, integrity, accessibility, privacy, and regulatory compliance across AWS, Databricks, Kubernetes, and Teradata environments. Builds APIs, database views, and data extracts; selects appropriate storage platforms; applies transformation rules; monitors data quality; and collaborates with cross-functional teams to optimize data processing and resolve issues. Requires independent judgment and availability for variable schedules, including nights and weekends.
Top Skills:
Apache AirflowAPIsAWSDatabricksKubernetesPythonTeradata
Digital Media • Information Technology • News + Entertainment
Design, implement, and administer SAP Analytics Cloud (SAC) planning and reporting solutions. Lead SAC modelling, story creation, SAC admin, and migrations from BPC to SAC. Collaborate with data, product, and functional teams to define data acquisitions, virtual models in HANA, pipelines, and security. Provide support, training, and stabilization efforts while managing multiple projects and mentoring team members.
Top Skills:
AbapAWSAzureEpmGCPSap Analytics Cloud (Sac)Sap AoeSap BpcSap BtpSap BwSap DatasphereSap HanaSap S/4 Hana
Big Data • Information Technology • Analytics • Business Intelligence
Design, develop, and maintain scalable ETL pipelines and data infrastructure for large-scale structured and unstructured datasets. Integrate and govern data across sources, optimize pipeline performance, and deploy cloud-based solutions. Collaborate with data scientists, analysts, and engineers to deliver accessible datasets. Implement automation, testing, security, compliance, monitoring, and continuous improvements across data engineering workflows.
Top Skills:
Amazon RedshiftApache KafkaSparkAWSAzure SynapseGoogle BigqueryGoogle Cloud PlatformHadoopJavaAzureMongoDBMySQLPostgresPythonScalaSQL
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


