Designs and optimizes scalable ETL/ELT pipelines using Databricks, PySpark, SQL, GCP, BigQuery, and Delta Lake. Builds agentic workflows for data validation, troubleshooting, and automation; develops reusable frameworks, quality controls, monitoring, exception handling, and production processes. The role also requires performance optimization, testing, deployment, production support, and collaboration with business, data engineering, and platform teams.
We are looking for a Data Engineer with strong Databricks and GCP experience to build scalable ETL/ELT pipelines and productionize agentic workflows for data engineering automation.
The mandatory requirements are 5+ years of hands-on Databricks and PySpark experience, advanced SQL skills, strong GCP and BigQuery knowledge, and experience with Delta Lake and lakehouse architectures.
RECRUITMENT PROCESS
1. Application — Share a few details about your experience and background.
2. Coding Challenge — If applicable, complete it in the coding language you are most comfortable with.
3. Video Interview — Record a short video introduction in English.
4. Technical Interview or Hiring Manager Interview — Discuss your experience and fit for the role.
5. Offer — If it’s a match, you’ll receive an offer.
WHAT YOU’LL GAIN
- Remote work — Work from where you feel most productive.
- Local presence in India — Work in a structured and compliant environment aligned with Indian regulations.
- Competitive compensation in INR — Receive compensation in INR, plus support for learning, education, and wellness.
- Exciting projects — Work with modern technologies for global clients and fast-growing companies.
MUST HAVES
- 5+ years of strong hands-on experience with Databricks and PySpark.
- Advanced SQL and data-processing skills.
- Hands-on experience with GCP, particularly BigQuery.
- Experience with Delta Lake and modern data lake/lakehouse architectures.
- Strong understanding of ETL/ELT, data pipeline design, performance optimization, and data quality.
- Experience building reliable, scalable, production-grade data solutions.
- Strong analytical and troubleshooting skills.
- Understanding of software engineering practices, including testing, version control, deployment, monitoring, and production support.
- Upper-intermediate English level.
NICE TO HAVES
- Experience developing or integrating AI/agentic workflows, AI agents, or workflow automation solutions.
- Experience applying AI to automate data engineering, validation, troubleshooting, or operational processes.
- Familiarity with orchestration and automation frameworks.
- Experience developing reusable data engineering frameworks and platform components.
- Exposure to productionizing AI-enabled solutions with appropriate validation, monitoring, and human oversight.
WHAT YOU WILL DO
- Design, develop, and optimize scalable ETL/ELT pipelines using Databricks, PySpark, and SQL.
- Build and operationalize agentic workflows that automate data validation, issue identification, troubleshooting, and workflow execution.
- Integrate agentic capabilities with Databricks, GCP, BigQuery, and Delta Lake environments while supporting new business requirements and datasets.
- Build reusable frameworks and components that can support multiple data engineering and business use cases.
- Implement data quality checks, monitoring, validation, exception handling, and production controls.
- Optimize PySpark and SQL workloads and support testing, deployment, productionization, and ongoing enhancement of data and agentic solutions.
- Troubleshoot complex production issues and collaborate with business, data engineering, and platform teams to identify additional automation opportunities.
Similar Jobs
Automotive
Develop and support Oracle HCM Cloud technical solutions for HR and Payroll operations. Responsibilities include building integrations, HCM Extracts, BI Publisher and OTBI reports, HDL/HSDL loads, Fast Formulas, data conversions, troubleshooting, testing, deployments, and production support. The role translates business requirements into technical solutions, supports Oracle Cloud updates, prepares documentation, and collaborates with functional teams, stakeholders, and vendors.
Top Skills:
Bi PublisherFast FormulasHcm ExtractsHdlHsdlOracle Hcm CloudOracle Integration CloudOtbiRest ApisSoap ApisSQL
Automotive
Designs, develops, and supports scalable SAP solutions using ABAP, RAP, CDS, OData, and S/4HANA extensibility. Responsibilities include API-led integrations, Fiori/UI5 backend services, BTP extensions, testing, debugging, performance optimization, code reviews, architecture discussions, and Agile delivery. The role requires maintaining custom SAP developments across S/4HANA and ECC while following clean core principles, enterprise security standards, and SAP best practices.
Top Skills:
Abap Development Tools (Adt)Api ManagementCloud Application Programming Model (Cap)Core Data Services (Cds)EclipseGitIdocsObject-Oriented AbapOdata V2Odata V4Rest ApisRestful Abap Programming Model (Rap)RfcsSap AbapSap Btp Abap EnvironmentSap Business Technology Platform (Btp)Sap EccSap FioriSap GatewaySap Integration SuiteSap S/4HanaSoapUi5Web Services
Automotive
Design, implement, maintain, and optimize Ford’s on-premises and GCP-based Dassault Systemes 3DEXPERIENCE infrastructure. Automate provisioning, configuration, and deployments with Terraform and Ansible; manage cloud resources, servers, networking, databases, and load balancing; troubleshoot Linux-based applications and platform incidents; implement security controls; optimize performance; document infrastructure; and collaborate with application, database, and security engineering teams.
Top Skills:
AnsibleApache Http ServerApache TomcatBashCi/CdCloud Load BalancingCloud SqlCloud StorageCompute EngineContainersGoogle Cloud Platform (Gcp)HaproxyInfrastructure As Code (Iac)Java Virtual Machine (Jvm)LinuxPythonServerless ComputingTerraformVpcWindows
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

