Lead design and build of data pipelines and ETL using Databricks/Spark, integrate Fenergo APIs, ensure data quality and compliance, analyze data, and develop Power BI dashboards for KYC/AML reporting while collaborating with stakeholders and maintaining governance documentation.
Responsibilities
API Integration & Data Extraction
- Lead the integration of Fenergo APIs to extract relevant KYC and AML data, ensuring seamless connectivity and data flow between systems
- Design, develop, and maintain scalable data pipelines and ETL processes to support data ingestion from various sources, including databases, APIs, and flat files
- Ensure robust data extraction processes that maintain data quality and compliance with regulatory requirements
Data Processing & Pipeline Development
- Utilize Databricks and Apache Spark to design and implement robust data processing pipelines, ensuring high data quality and performance
- Work with DataFrames for transforming data and implementing the Medallion Architecture
- Execute SQL queries for data extraction, manipulation, and complex data operations
- Join datasets and add fields to reports to provide comprehensive analytical insights
- Leverage AI tools in Databricks to assist with data workflows and optimization
- Use notebooks as data transformation pipelines for efficient data processing
Data Analysis & Interpretation
- Analyze and interpret complex data sets to identify trends, patterns, and anomalies that can inform business decisions related to client and investor lifecycle management
- Understand and navigate the data model to ensure accurate data representation and reporting
- Conduct regular data quality assessments and audits to ensure data integrity and compliance with industry standards
- Perform root cause analysis to swiftly identify data issues and collaborate with relevant teams to implement effective solutions
Data Visualisation & Reporting
- Architect and develop interactive dashboards and reports in Power BI, translating complex data into clear, actionable insights for clients, leadership, and stakeholders
- Develop and maintain dashboards and reports to provide insights into key performance indicators (KPIs) and operational metrics for KYC/AML processes
- Create visual representations that highlight critical data points for regular reporting to clients and senior management
- Ensure reports meet the needs of both technical and non-technical stakeholders
Collaboration & Stakeholder Management
- Collaborate with cross-functional teams, including IT, Compliance, Risk Management, business analysts, and senior management, to gather data requirements and deliver strategic insights
- Engage with clients and internal stakeholders to understand their reporting needs and ensure alignment with business objectives
- Work closely with KYC/AML operations teams to ensure data solutions support compliance and regulatory requirements
- Act as a bridge between technical teams and business users, translating complex data concepts into actionable business insights
Documentation & Governance
- Maintain comprehensive documentation of data processes, API integrations, data flows, data management processes, and reporting solutions for future reference and compliance
- Document data governance practices and ensure adherence to data quality best practices
- Ensure all data handling complies with regulatory standards and internal policies
Continuous Improvement & Problem-Solving
- Recommend long-term product solutions to enhance data quality, accessibility, and usability
- Identify opportunities for process optimization and automation in data workflows
- Stay up-to-date with industry trends and best practices in data engineering, analysis, and management
- Proactively identify and resolve data-related issues, ensuring timely and accurate reporting
- Demonstrate creativity and insightfulness in developing dynamic approaches to complex data challenges
Quality Assurance
- Ensure data integrity throughout all pipelines and reporting mechanisms
- Implement data validation and quality control measures
- Monitor data processes and implement control mechanisms to ensure reliability
Skills
Core Data Engineering Skills (Required):
- Proficiency in Databricks and Apache Spark for data processing and pipeline development
- Strong knowledge of Power BI for data visualization and reporting, with ability to create executive-level dashboards
- Expert-level proficiency in SQL for data querying, manipulation, and complex analytical operations
- Experience with programming languages such as Python or R for data analysis and automation
- Strong understanding of data warehousing concepts and ETL processes
- Knowledge of data modeling concepts and best practices for data management
- Understanding of the Medallion Architecture and data lakehouse principles
- Experience working with DataFrames for data transformation
- Ability to leverage AI tools in Databricks to optimize data workflows
API & Integration Skills (Required):
- Strong experience in API integration for data extraction and system connectivity
- Ability to ensure seamless data flow between multiple systems
Cloud & Infrastructure (Required):
- Experience with cloud platforms, particularly Azure or AWS
- Knowledge of Git connection to Databricks for version control
- Experience with AWS/Azure and Databricks integration/mounting
- Understanding of data governance and data quality best practices
Additional Technical Skills (Preferred):
- Databricks administration skills
- Familiarity with machine learning concepts and their application in data analysis
- Experience with graph data models
- Understanding of data governance and compliance standards (GDPR, AML regulations, etc.)
- Knowledge of secure data handling practices
Photon Chennai, Tamil Nadu, IND Office
DLF IT Park 1/124 Mount Poonamallee Road Sivaji Gardens Manapakkam , Chennai, India, 600089
Similar Jobs
Enterprise Web • Mobile • Professional Services • Software
Build and improve production LLM agents and AI features through evaluation, experimentation, observability, prompting, context engineering, tool use, and workflow design. Develop backend services, APIs, data models, and feedback pipelines that make agent behavior reliable and measurable. Investigate performance issues, run staged rollouts and production replays, define quality standards, and partner with Product, Data Science, and Sales to deliver business outcomes with appropriate privacy, security, and human-oversight safeguards.
Top Skills:
Agentic WorkflowsAPIsBackend ServicesBraintrustClaude CodeCursorData ModelsDatadog Llm ObservabilityEvaluation HarnessesGithub CopilotLangchainLanggraphLangsmithLlm-As-JudgeLlm-Based SystemsMcp
Enterprise Web • Mobile • Professional Services • Software
Own the AI product roadmap from discovery through launch and iteration. Define model evaluation standards, failure modes, rollback plans, and agent-autonomy boundaries. Prototype and ship production AI features, influence architecture, and apply prompting and context design. Collaborate with Design, Research, Engineering, Sales, Marketing, and Customer Success while tracking quality, latency, cost, adoption, and growth metrics.
Top Skills:
AIBraintrustClaudeCodexCursorLangsmithLlmsOpenclaw
Enterprise Web • Mobile • Professional Services • Software
Build and own full-stack product features from ambiguous problems through deployment. Partner with Product and Design, make UX decisions, integrate LLM and agent capabilities into trustworthy product experiences, talk with users, instrument releases, and iterate based on usage. The role requires strong product judgment, user empathy, independent execution, and comfort using AI coding tools while collaborating with AI engineers on model behavior.
Top Skills:
Claude CodeCursorGithub CopilotLlm AgentsLlm ApisPrompting
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

