Build and maintain Azure-based data pipelines for governance and analytics. Develop Databricks ETL/ELT workflows using PySpark, SQL, and Delta Lake; automate data quality, lineage, monitoring, access auditing, cost analysis, retention tracking, policy enforcement, and metadata integration with Microsoft Purview. Maintain curated governance metric tables, optimize performance, troubleshoot pipeline issues, and collaborate with engineering, QA, and reporting teams while meeting security and privacy standards.
Job title : Data Engineer
Experience: 6+ years
Location: Chennai
Shift: US Eastern Time ( 5:00 PM – 2:00 AM )
Requirements
We are seeking a hands-on Data Engineer to develop, optimize, and maintain automated
data pipelines supporting data governance and analytics initiatives. This role will focus on
building production-ready workflows for ingestion, transformation, quality checks, lineage
capture, access auditing, cost usage analysis, retention tracking, and metadata integration,
primarily using Azure Databricks, Azure Data Lake, and Microsoft Purview.
Experience: 6+ years in data engineering, with strong Azure and Databricks experience
Pipeline Development – Design, build, and deploy robust ETL/ELT pipelines in Databricks
(PySpark, SQL, Delta Lake) to ingest, transform, and curate governance and operational
metadata from multiple sources landed in Databricks.
Granular Data Quality Capture – Implement profiling logic to capture issue-level metadata
(source table, column, timestamp, severity, rule type) to support drill-down from dashboards
into specific records and enable targeted remediation.
Governance Metrics Automation – Develop data pipelines to generate metrics for
dashboards covering data quality, lineage, job monitoring, access & permissions, query cost,
usage & consumption, retention & lifecycle, policy enforcement, sensitive data mapping, and
governance KPIs.
Microsoft Purview Integration – Automate asset onboarding, metadata enrichment,
classification tagging, and lineage extraction for integration into governance reporting.
Data Retention & Policy Enforcement – Implement logic for retention tracking and policy
compliance monitoring (masking, RLS, exceptions).
Job & Query Monitoring – Build pipelines to track job performance, SLA adherence, and
query costs for cost and performance optimization.
Metadata Storage & Optimization – Maintain curated Delta tables for governance metrics,
structured for efficient dashboard consumption.
Testing & Troubleshooting – Monitor pipeline execution, optimize performance, and resolve
issues quickly.
Collaboration – Work closely with the lead engineer, QA, and reporting teams to validate
metrics and resolve data quality issues.
Security & Compliance – Ensure all pipelines meet organizational governance, privacy, and
security standards.
Similar Jobs
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Owns the architecture, engineering standards, and technology roadmap for a secure healthcare data platform across AWS, Azure, and Databricks. Designs ingestion, lakehouse, interoperability, governance, privacy, analytics, and operational reporting solutions. Establishes standards for Spark, Python, SQL, data quality, observability, metadata, CI/CD, infrastructure as code, and disaster recovery. Leads healthcare data interoperability, security architecture, proofs of concept, performance assessments, technical roadmaps, and executive decision support.
Top Skills:
Amazon EventbridgeAmazon S3Apache IcebergSparkAthenaAuto LoaderAWSAws LambdaAws Secrets ManagerAws Step FunctionsAzureCi/CdClinical NlpCloudwatchCptDatabricksDatabricks SqlDatabricks WorkflowsDelta LakeEcsEksFhir R4Hl7 V2Icd-10Infrastructure As CodeJSONKmsLlmsLoincMlopsNdjsonOmop CdmPysparkPythonRxnormSnomed CtSnowflakeSQLSreTrinoUnity CatalogXML
Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Analyze business and user requirements across financial services value streams, define technical requirements, develop cloud solution options, and manage business and IT stakeholders. The role requires understanding existing framework capabilities, complex data models, cross-functional solution design, clear documentation, team coordination across regions, and familiarity with the SDLC from design through testing and implementation.
Top Skills:
Google Cloud Platform
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Provide L1-L3 production support for enterprise MongoDB Atlas environments, including monitoring, scaling, backups, upgrades, performance tuning, security, networking, disaster recovery, and incident resolution. Build Terraform-based infrastructure automation, support multi-cloud deployments across Azure, AWS, and GCP, partner with application and engineering teams, coordinate vendor escalations, and participate in 24/7 on-call coverage.
Top Skills:
AWSAzureGCPGitGithub ActionsGoJavaKubernetesMongodb AtlasPowershellPythonShellSQL ServerTerraform
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


