Photon Logo

Photon

SPARK Data Onboarding Engineer - Chennai

Reposted 6 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in India
Senior level
Remote
Hiring Remotely in India
Senior level
Design, build, and optimize PySpark applications and ETL pipelines to process large-scale datasets from relational, NoSQL, file, and streaming sources. Ensure data quality, implement error handling, and collaborate with analysts, scientists, and architects to deliver performant data solutions.
The summary above was generated by AI

Job Title: PySpark Data Engineer

Summary:

We are seeking a skilled PySpark Data Engineer to join our team and drive the development of robust data processing and transformation solutions within our data platform. You will be responsible for designing, implementing, and maintaining PySpark-based applications to handle complex data processing tasks, ensure data quality, and integrate with diverse data sources. The ideal candidate possesses strong PySpark development skills, experience with big data technologies, and the ability to work in a fast-paced, data-driven environment.

Key Responsibilities: Data Engineering Development:

  • Design, develop, and test PySpark-based applications to process, transform, and analyze large-scale datasets from various sources, including relational databases, NoSQL databases, batch files, and real-time data streams.
  • Implement efficient data transformation and aggregation using PySpark and relevant big data frameworks.
  • Develop robust error handling and exception management mechanisms to ensure data integrity and system resilience within Spark jobs.
  • Optimize PySpark jobs for performance, including partitioning, caching, and tuning of Spark configurations.

Data Analysis and Transformation:

  • Collaborate with data analysts, data scientists, and data architects to understand data processing requirements and deliver high-quality data solutions.
  • Analyze and interpret data structures, formats, and relationships to implement effective data transformations using PySpark.
  • Work with distributed datasets in Spark, ensuring optimal performance for large-scale data processing and analytics.

Data Integration and ETL:

  • Design and implement ETL (Extract, Transform, Load) processes to ingest and integrate data from various sources, ensuring consistency, accuracy, and performance.
  • Integrate PySpark applications with data sources such as SQL databases, NoSQL databases, data lakes, and streaming platforms

Qualifications and Skills:

  • Bachelor's degree in Computer Science, Information Technology, or a related field.
  • 5+ years of hands-on experience in big data development, preferably with exposure to data-intensive applications.
  • Strong understanding of data processing principles, techniques, and best practices in a big data environment.
  • Proficiency in PySpark, Apache Spark, and related big data technologies for data processing, analysis, and integration.
  • Experience with ETL development and data pipeline orchestration tools (e.g., Apache Airflow, Luigi).
  • Strong analytical and problem-solving skills, with the ability to translate business requirements into technical solutions.
  • Excellent communication and collaboration skills to work effectively with data analysts, data architects, and other team members.

Photon Chennai, Tamil Nadu, IND Office

DLF IT Park 1/124 Mount Poonamallee Road Sivaji Gardens Manapakkam , Chennai, India, 600089

Similar Jobs

52 Minutes Ago
Easy Apply
Remote
India
Easy Apply
Mid level
Mid level
Enterprise Web • Mobile • Professional Services • Software
Build and improve production LLM agents and AI features through evaluation, experimentation, observability, prompting, context engineering, tool use, and workflow design. Develop backend services, APIs, data models, and feedback pipelines that make agent behavior reliable and measurable. Investigate performance issues, run staged rollouts and production replays, define quality standards, and partner with Product, Data Science, and Sales to deliver business outcomes with appropriate privacy, security, and human-oversight safeguards.
Top Skills: Agentic WorkflowsAPIsBackend ServicesBraintrustClaude CodeCursorData ModelsDatadog Llm ObservabilityEvaluation HarnessesGithub CopilotLangchainLanggraphLangsmithLlm-As-JudgeLlm-Based SystemsMcp
52 Minutes Ago
Easy Apply
Remote
India
Easy Apply
Mid level
Mid level
Enterprise Web • Mobile • Professional Services • Software
Own the AI product roadmap from discovery through launch and iteration. Define model evaluation standards, failure modes, rollback plans, and agent-autonomy boundaries. Prototype and ship production AI features, influence architecture, and apply prompting and context design. Collaborate with Design, Research, Engineering, Sales, Marketing, and Customer Success while tracking quality, latency, cost, adoption, and growth metrics.
Top Skills: AIBraintrustClaudeCodexCursorLangsmithLlmsOpenclaw
53 Minutes Ago
Easy Apply
Remote
India
Easy Apply
Mid level
Mid level
Enterprise Web • Mobile • Professional Services • Software
Build and own full-stack product features from ambiguous problems through deployment. Partner with Product and Design, make UX decisions, integrate LLM and agent capabilities into trustworthy product experiences, talk with users, instrument releases, and iterate based on usage. The role requires strong product judgment, user empathy, independent execution, and comfort using AI coding tools while collaborating with AI engineers on model behavior.
Top Skills: Claude CodeCursorGithub CopilotLlm AgentsLlm ApisPrompting

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account