Design, build, and optimize scalable ETL/ELT data pipelines using PySpark and Python. Model and tune MongoDB for large datasets, develop data ingestion and validation frameworks, integrate REST APIs, monitor and troubleshoot pipelines, and collaborate with analysts and data scientists while applying data governance and CI/CD practices.
Job Title: Data Engineer (MongoDB, PySpark & Python)
Experience
Experience
5–8 Years
As per business requirement
We are looking for an experienced Data Engineer with strong expertise in MongoDB, PySpark, and Python to design, develop, and optimize scalable data pipelines. The ideal candidate should have experience working with large datasets, NoSQL databases, distributed data processing, and cloud-based data platforms.
- Design, build, and maintain scalable ETL/ELT data pipelines using PySpark and Python.
- Develop data ingestion frameworks to process structured, semi-structured, and unstructured data.
- Work extensively with MongoDB for data modeling, querying, indexing, aggregation, and performance optimization.
- Optimize Spark jobs for high-performance processing of large datasets.
- Build reusable data transformation and validation frameworks.
- Develop REST API integrations and automate data ingestion using Python.
- Monitor, troubleshoot, and optimize data pipelines for reliability and performance.
- Collaborate with business analysts, data scientists, and application teams to deliver data solutions.
- Implement data quality, governance, and security best practices.
- Participate in code reviews and follow CI/CD and Agile development practices.
- Strong experience in Python programming.
- Hands-on experience with PySpark and Spark SQL.
- Strong knowledge of MongoDB, including:
- CRUD Operations
- Aggregation Framework
- Indexing
- Replication
- Sharding
- Performance Tuning
- CRUD Operations
- Good understanding of data structures and algorithms.
- Experience in developing ETL/ELT pipelines.
- Strong SQL skills.
- Experience with Git version control.
- Knowledge of Linux/Unix commands.
- Experience working with JSON, XML, and Parquet data formats.
- Experience with Databricks.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Knowledge of Apache Kafka or other streaming technologies.
- Experience with orchestration tools such as Apache Airflow.
- Understanding of Delta Lake and Lakehouse architecture.
- Familiarity with CI/CD pipelines.
Similar Jobs
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Lead P2P reporting and analytics for Procurement and Accounts Payable: design dashboards, perform complex Snowflake/SQL analyses, build Tableau executive visualizations, automate reporting (Google Apps Script, Snowflake Tasks), partner with stakeholders, drive projects end-to-end, mentor junior analysts, and apply AI/automation to improve reporting efficiency and anomaly detection.
Top Skills:
ChatgptCoupaGoogle Apps ScriptGoogle SheetsMicrosoft AccessMicrosoft CopilotExcelMicrosoft PowerpointMicrosoft VisioMicrosoft WordNetSuitePythonRSnowflakeSnowflake TasksSQLTableauTableau CloudTableau Server
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Lead architecture and technology strategy for enterprise-scale Java-based microservices and Vue.js front-ends. Define technical roadmaps, standards, cloud-native and DevOps adoption, CI/CD and observability, ensure scalability/security/availability, mentor teams, support production troubleshooting, and drive modernization initiatives.
Top Skills:
SparkApi GatewaysArtifactoryConfluenceDockerGitGoogle Cloud Platform (Gcp)GrafanaHarnessJavaJenkinsKafkaKubernetesLinuxMavenNoSQLPrometheusRelational DatabasesSonarqubeSplunkSpring BootSpring FrameworkVue
Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
The role involves developing executive-level relationships, managing end-to-end customer engagement, and demonstrating effective solution-based sales processes in complex sales campaigns with enterprise customers.
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.



