EXL Logo

EXL

Senior Databricks Engineer

Reposted 26 Days Ago
Be an Early Applicant
Hybrid
Chennai, Tamil Nadu, IND
Senior level
Hybrid
Chennai, Tamil Nadu, IND
Senior level
Design, build, and optimize Databricks-based lakehouse solutions for financial crime use cases. Manage Databricks workspaces, clusters, Unity Catalog, and jobs. Develop scalable batch and streaming PySpark/Spark SQL pipelines (Kafka/Event Hubs/Kinesis) using Delta Lake, DLT, and Medallion architecture. Ensure data governance, performance tuning, error handling, and integration with cloud storage (ADLS/S3/GCS) and secret management.
The summary above was generated by AI

We are looking for a skilled and passionate Senior Databricks Engineer to design, build, and optimize enterprise-scale data lakehouse solutions on the Databricks platform. The successful candidate will be responsible for creating Databricks pipeline delivering Financial Crime platforms covering Anti-Money Laundering (AML), Know Your Customer (KYC), Customer Risk Assessment (CRA), Sanctions Screening, Transaction Monitoring, Fraud Detection, and Regulatory Reporting

Responsibilities

Databricks Platform Engineering

  • Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments.
  • Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage.
  • Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance.
  • Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager).
  • Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management.

Delta Lake & Lakehouse Architecture

  • Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM).
  • Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization.
  • Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations.
  • Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing.
  • Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).

Data Pipeline Development (PySpark / SQL)

  • Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake.
  • Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables.
  • Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning.
  • Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity.
  • Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines.

Qualifications

Education

  • Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or related field.

Experience

  • 6-8 years of total experience in data engineering or software engineering.
  • 4+ years of dedicated hands-on experience with the Databricks platform in production environments.
  • Strong background in big data engineering, cloud data platforms, and distributed computing.

EXL Chennai, Tamil Nadu, IND Office

Chennai, India

Similar Jobs

5 Days Ago
In-Office
Saidapet, Chennai, Tamil Nadu, IND
Senior level
Senior level
Big Data • Information Technology • Analytics • Business Intelligence
Designs and maintains scalable Databricks and Spark data pipelines, including ETL/ELT workflows, Delta Lake transformations, data quality checks, and performance optimization. Works with cloud platforms, data warehousing, modeling, CI/CD, monitoring, and production support. Troubleshoots large-scale data processing issues and collaborates with engineers, data scientists, analysts, and business teams. Preferred experience includes Unity Catalog, orchestration tools, streaming technologies, infrastructure automation, and Databricks certifications.
Top Skills: SparkAWSAws GlueAzureAzure Data FactoryCi/CdDatabricksDelta LakeGCPGitInfrastructure As CodeKafkaPysparkPythonSpark Structured StreamingSQLUnity Catalog
One Month Ago
Hybrid
Chennai, Tamil Nadu, IND
Senior level
Senior level
Big Data • Information Technology
Design, build and maintain scalable automation frameworks and test suites for cloud/big-data systems. Implement automated and manual tests across UI, services, APIs and performance; deploy via CI/CD; collaborate with cross-functional teams and provide technical guidance.
Top Skills: AirflowDatabricksHiveMachine LearningNoSQLPostgresPysparkPythonSQL
15 Hours Ago
Hybrid
Chennai, Tamil Nadu, IND
Expert/Leader
Expert/Leader
Digital Media • Information Technology • News + Entertainment
Provides senior technical leadership across data engineering, platform engineering, and application development initiatives. Responsibilities include architecture and design governance, engineering standards, modernization, cloud and AI adoption, production incident escalation, reliability improvement, and mentorship of engineers. The role drives technical strategy, system scalability, performance, security, operational stability, and delivery excellence across enterprise platforms.
Top Skills: Ai/Ml PlatformsApache AirflowApache HiveSparkAWSCi/CdData LakesData WarehousingDatabricksDevOpsJavaKafkaObservability And Monitoring ToolsPysparkPythonSQL

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account