Build and maintain scalable Databricks ETL/ELT pipelines using Python, PySpark, and SQL. Integrate databases, Amazon S3, files, and REST APIs; develop Delta Lake models using Medallion Architecture; manage Databricks Jobs and Workflows; optimize Spark performance; implement data quality, monitoring, error handling, and dimensional modeling; and collaborate with cross-functional teams on production-ready data solutions.
Position Overview
We are looking for a Data Engineer with hands-on experience in building scalable data pipelines and data engineering solutions on the Databricks Lakehouse Platform. The ideal candidate should have strong expertise in Python, PySpark, SQL, Databricks, AWS, and REST API integrations for data ingestion, managing large volumes of data, and data export
We are looking for a Data Engineer with hands-on experience in building scalable data pipelines and data engineering solutions on the Databricks Lakehouse Platform. The ideal candidate should have strong expertise in Python, PySpark, SQL, Databricks, AWS, and REST API integrations for data ingestion, managing large volumes of data, and data export
ShyftLabs is a growing data product company that was founded in early 2020 and works primarily with Fortune 500 companies. We deliver digital solutions built to help accelerate the growth of businesses in various industries, by focusing on creating value through innovation.
Job Responsibilities:
Design, develop, and maintain scalable ETL/ELT pipelines using Databricks,
PySpark, and SQL.
● Integrate data from multiple sources, including databases, Amazon S3, files, and REST APIs.
● Build data pipelines with Databricks Unity Catalog.
● Implement business logic, data transformations, and dimensional data models.
● Create, schedule, monitor, and optimize Databricks Jobs and Workflows.
● Design and manage Delta Lake tables using Medallion Architecture (Bronze, Silver,Gold).
● Ensure data quality through validations, error handling, logging, and monitoring.
● Optimize Spark workloads for performance, scalability, and reliability.
● Collaborate with cross-functional teams to deliver production-ready data solutions.
PySpark, and SQL.
● Integrate data from multiple sources, including databases, Amazon S3, files, and REST APIs.
● Build data pipelines with Databricks Unity Catalog.
● Implement business logic, data transformations, and dimensional data models.
● Create, schedule, monitor, and optimize Databricks Jobs and Workflows.
● Design and manage Delta Lake tables using Medallion Architecture (Bronze, Silver,Gold).
● Ensure data quality through validations, error handling, logging, and monitoring.
● Optimize Spark workloads for performance, scalability, and reliability.
● Collaborate with cross-functional teams to deliver production-ready data solutions.
Basic Qualification:
Strong expertise in Python, PySpark, and Advanced SQL.
● Hands-on experience with the Databricks Lakehouse Platform.
● Good understanding of Unity Catalog, Delta Lake, Databricks Workflows/Jobs,
Clusters, Notebooks, Repos, and Medallion Architecture.
● Experience integrating with REST APIs for data ingestion and data export.
● Strong knowledge of ETL/ELT development, batch processing, incremental loading,
and data transformation.
● Experience with data modeling (Star Schema, Snowflake Schema, Fact & Dimension
tables, SCD concepts).
● Understanding of data warehousing concepts and best practices.
● Experience working with structured and semi-structured data (CSV, JSON, Parquet,
Delta).
● Knowledge of partitioning, file optimization, Spark performance tuning, and query
optimization.
● Experience with Git and CI/CD best practices
● Hands-on experience with the Databricks Lakehouse Platform.
● Good understanding of Unity Catalog, Delta Lake, Databricks Workflows/Jobs,
Clusters, Notebooks, Repos, and Medallion Architecture.
● Experience integrating with REST APIs for data ingestion and data export.
● Strong knowledge of ETL/ELT development, batch processing, incremental loading,
and data transformation.
● Experience with data modeling (Star Schema, Snowflake Schema, Fact & Dimension
tables, SCD concepts).
● Understanding of data warehousing concepts and best practices.
● Experience working with structured and semi-structured data (CSV, JSON, Parquet,
Delta).
● Knowledge of partitioning, file optimization, Spark performance tuning, and query
optimization.
● Experience with Git and CI/CD best practices
Preferred Qualifications:
4+ years of experience in Data Engineering with 2+ years of hands-on Databricks
experience.
● Experience with Auto Loader, Spark Declarative pipelines, Kafka, Airflow, or dbt is a plus.
● Databricks certification is an added advantage.
experience.
● Experience with Auto Loader, Spark Declarative pipelines, Kafka, Airflow, or dbt is a plus.
● Databricks certification is an added advantage.
We are proud to offer a competitive salary alongside a strong insurance package. We pride ourselves on the growth of our employees, offering extensive learning and development resources.
Similar Jobs
Blockchain • Database • Analytics
Lead the design, development, and optimization of scalable Databricks ETL/ELT pipelines. Integrate data from databases, S3, files, and REST APIs; implement Delta Lake medallion architecture, Unity Catalog, dimensional models, data quality controls, and Spark performance tuning. Schedule and monitor Databricks workflows, collaborate cross-functionally, and guide engineers while driving technical decisions.
Top Skills:
AirflowAmazon S3SparkAuto LoaderAWSCi/CdDatabricksDatabricks Lakehouse PlatformDbtDelta LakeGitKafkaPysparkPythonRest ApisSQLUnity Catalog
Edtech • HR Tech • Information Technology • Professional Services
Designs, builds, and configures scalable data solutions using Databricks. Responsibilities include developing robust data pipelines, optimizing data workflows, supporting large-scale data processing, integrating data, managing data quality, troubleshooting technical issues, and delivering enterprise-grade analytics solutions. The role also requires collaboration with teams and stakeholders while serving at a team lead or consultant level.
Top Skills:
Databricks Data EngineeringDatabricks Unified Data Analytics Platform
Digital Media • Information Technology • News + Entertainment
Designs and customizes software applications, manages releases, gathers requirements, and improves functionality and user experience. Collaborates with stakeholders, QA, and cross-functional teams to deliver reliable solutions. Tracks performance metrics, maintains technical documentation, researches industry trends, and leads development, prototyping, and design sessions. Mentors junior engineers and exercises independent judgment while supporting variable work schedules.
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.


