Design, deploy, monitor, and maintain production machine learning systems and end-to-end ML pipelines. Build scalable MLOps frameworks, automate deployments and retraining, manage cloud-native infrastructure and CI/CD, implement IaC, containerization, monitoring, and governance, and collaborate with cross-functional teams to ensure reliable, compliant enterprise AI platforms.
What You ‘ll Do
You will join our high performance Data & AI team and play a key role in designing, deploying, monitoring, and maintaining enterprise-grade machine learning solutions. You will bridge the gap between Data Science, AI Engineering, and Cloud Operations by building scalable MLOps platforms, automating ML workflows, and ensuring reliable production AI systems.
- Design, build, deploy, and maintain machine learning models in production environments.
- Develop and manage end-to-end machine learning pipelines covering data ingestion, feature engineering, model training, validation, deployment, monitoring, and retraining.
- Build scalable MLOps frameworks that enable efficient model lifecycle management across enterprise AI platforms.
- Implement model versioning, experiment tracking, governance, and reproducibility best practices.
- Automate model deployment, retraining, rollback, and release workflows.
- Design and manage cloud-native infrastructure supporting enterprise AI and machine learning workloads.
- Develop and maintain CI/CD pipelines for machine learning applications and AI services.
- Implement Infrastructure as Code (IaC) using Terraform, ARM Templates, Bicep, or equivalent technologies.
- Deploy and manage containerized AI applications using Docker and Kubernetes.
- Monitor model performance, prediction quality, data drift, concept drift, system health, and resource utilization.
- Troubleshoot production issues related to ML pipelines, model serving, infrastructure, and deployment workflows.
- Implement logging, monitoring, alerting, and observability solutions for AI platforms.
- Optimize model serving performance, scalability, latency, and infrastructure efficiency.
- Collaborate with Data Scientists, AI Engineers, Software Developers, DevOps Engineers, and Business Stakeholders to operationalize machine learning solutions.
- Support AI governance, model security, compliance, audit readiness, and enterprise AI standards.
- Document MLOps processes, deployment architectures, operational runbooks, and engineering best practices.
- Participate in architecture reviews and continuously improve AI platform capabilities using emerging cloud-native technologies.
What We Seek In You
- 3+ years of experience in MLOps, Machine Learning Engineering, DevOps, Cloud Engineering, or AI Platform Engineering.
- Bachelor's degree in Computer Science, Artificial Intelligence, Data Science, Information Technology, Engineering, or a related discipline.
- Strong programming expertise in:
- Python
- SQL
- Bash / Shell Scripting
- PowerShell (preferred)
- Strong understanding of:
- Machine Learning Lifecycle Management
- Model Deployment
- Model Serving
- Feature Engineering Concepts
- Experiment Tracking
- Model Governance
- Hands-on experience with MLOps platforms including:
- MLflow
- Azure Machine Learning
- AWS SageMaker
- Google Vertex AI
- Git
- GitHub
- GitLab
- Azure DevOps
- Strong expertise in cloud-native DevOps and infrastructure technologies including:
- Docker
- Kubernetes
- CI/CD Pipelines
- Terraform
- Microsoft Azure
- Amazon Web Services (AWS)
- Google Cloud Platform (GCP)
- Experience working with enterprise databases and data platforms including:
- SQL Server
- PostgreSQL
- MySQL
- Data Lakes
- Data Warehouses
- ETL / ELT Pipelines
- Hands-on experience implementing monitoring and observability solutions using:
- Prometheus
- Grafana
- Azure Monitor
- Cloud-native monitoring platforms
- Strong understanding of IAM, cloud security, infrastructure security, and enterprise governance practices.
- Experience deploying and supporting production-scale Machine Learning systems.
- Strong analytical thinking, troubleshooting, and problem-solving capabilities.
- Excellent communication, documentation, and stakeholder management skills.
- Ability to collaborate effectively across Data Science, AI Engineering, Cloud Infrastructure, and DevOps teams.
- Strong ownership mindset with the ability to manage multiple priorities and deliver high-quality AI platforms.
Preferred Qualifications
- Experience with Generative AI and Large Language Models (LLMs).
- Knowledge of Retrieval-Augmented Generation (RAG) and enterprise knowledge retrieval solutions.
- Experience with:
- OpenAI
- Azure OpenAI
- Claude
- Gemini
- Other foundation model platforms
- Familiarity with AI orchestration frameworks including:
- LangChain
- Semantic Kernel
- AutoGen
- CrewAI
- Experience with distributed data processing and orchestration platforms including:
- Apache Airflow
- Prefect
- Apache Spark
- Databricks
- Knowledge of Vector Databases including:
- Pinecone
- Weaviate
- FAISS
- ChromaDB
- Exposure to Responsible AI, Explainable AI, AI Governance, and Model Explainability frameworks.
- Experience working in Manufacturing, Automotive, Healthcare, Financial Services, Supply Chain, or Enterprise AI domains is highly preferred.
Life At Next
At our core, we're driven by the mission of tailoring growth for our customers by enabling them to transform their aspirations into tangible outcomes. We're dedicated to empowering them to shape their futures and achieve ambitious goals. To fulfil this commitment, we foster a culture defined by agility, innovation, and an unwavering commitment to progress. Our organizational framework is both streamlined and vibrant, characterized by a hands-on leadership style that prioritizes results and fosters growth.
Perks Of Working With Us
- Clear objectives to ensure alignment with our mission, fostering your meaningful contribution.
- Abundant opportunities for engagement with customers, product managers, and leadership.
- You'll be guided by progressive paths while receiving insightful guidance from managers through ongoing feedforward sessions.
- Cultivate and leverage robust connections within diverse communities of interest. Choose your mentor to navigate your current endeavors and steer your future trajectory.
- Embrace continuous learning and upskilling opportunities through Nexversity.
- Enjoy the flexibility to explore various functions, develop new skills, and adapt to emerging technologies. Embrace a hybrid work model promoting work-life balance.
- Access comprehensive family health insurance coverage, prioritizing the well-being of your loved ones.
- Embark on accelerated career paths to actualize your professional aspirations.
Who we are?
We enable high growth enterprises build hyper personalized solutions to transform their vision into reality. With a keen eye for detail, we apply creativity, embrace new technology and harness the power of data and AI to co-create solutions tailored made to meet unique needs for our customers.
Join our passionate team and tailor your growth with us!
TVS Next Chennai, Tamil Nadu, IND Office
Chennai, India
Similar Jobs
Fintech • Financial Services
Design, develop, and deploy GenAI/LLM and NLP solutions (training, fine-tuning, inference). Build RAG pipelines with vector DBs, implement cloud MLOps and CI/CD, ensure model governance, monitor performance, collaborate cross-functionally, and mentor junior engineers.
Top Skills:
AWSAws BedrockAws LambdaAzureFaissGCPGitHugging Face TransformersMlopsOpenaiPineconePythonPyTorchRetrieval-Augmented Generation (Rag)SagemakerTensorFlow
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead offshore engineering teams delivering full stack and Generative AI solutions. Provide technical leadership, mentor engineers, and ensure delivery quality.
Top Skills:
Github CopilotLangchainLanggraphMongoDBNode.jsPythonReactRestful Apis
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead design, implementation, and operation of contact center and Genesys IVR systems; configure monitoring (Splunk, Zabbix, Grafana, Dynatrace); integrate AI-driven routing and analytics; manage CI/CD (Jenkins, GitHub Actions); support cloud and on-prem deployments; mentor engineers and participate in Agile processes.
Top Skills:
AnsibleAWSAzureChatbot PlatformsChefConversational AiDynatraceGenesys IvrGitGithub ActionsGCPGrafanaJenkinsLinux/UnixNlpPerlPowershellPythonPyTorchSplunkSplunk ObservabilitySplunk On-CallSQLTensorFlowTerraformZabbix
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

