Maintain and optimize production LLM/VLM services and microservices (Flask/FastAPI), manage async queues, deploy and monitor Uvicorn/Gunicorn hosts, integrate LLM routing tools, design prompt strategies, build feedback/evaluation pipelines, expose secure REST APIs, and track token usage, latency, and errors to ensure SaaS-grade AI performance.
Role Overview
We are looking for an AI Engineer to maintain and enhance the AI-driven backbone of the Sootra platform. This role involves ensuring production stability of LLM/VLM pipelines, optimizing model interactions, maintaining APIs and queues, and building feedback loops that continuously improve AI outputs.
Responsibilities
- Maintain and optimize LLM- and VLM-powered services for content generation, compliance scoring, and campaign testing.
- Manage and scale Flask/FastAPI microservices, ensuring high uptime and low latency.
- Maintain Dramatiq queues for async AI workflows, campaign generation, and pipeline orchestration.
- Deploy, monitor, and debug Uvicorn/Gunicorn-based hosting in production environments.
- Integrate with OpenRouter and equivalent LLM routing tools to balance cost, latency, and quality.
- Design and refine prompt engineering strategies for reliability, context-awareness, and compliance.
- Build and maintain feedback pipelines for AI model evaluation (human-in-the-loop scoring, automated quality checks, reinforcement).
- Expose and maintain REST APIs for AI services, ensuring secure, versioned endpoints.
- Collaborate with backend/frontend teams to keep microservice architecture aligned and maintainable.
- Track token consumption, latency, and error rates to ensure production-grade performance.
Required Skills
- Programming: Strong in Python, with experience in production-grade codebases.
- Frameworks: Flask (for APIs), FastAPI (optional), Uvicorn/Gunicorn for async hosting.
- Queues/Workers: Dramatiq (or Celery/RQ equivalent) for background jobs.
- AI/ML: Hands-on with LLMs and VLMs, including prompt engineering, fine-tuning, and evaluation.
- AI Infrastructure: Familiar with OpenRouter or equivalent LLM/VLM routing & fallback tools.
- Architecture: Experience designing and maintaining microservice architectures.
- APIs: Strong experience with REST API design (auth, rate limiting, documentation).
- Production: Dockerized deployments, CI/CD pipelines, logging/monitoring, error handling.
- Feedback Loops: Building structured evaluation/feedback systems for AI model performance.
- Cloud: AWS/GCP experience preferred (deployment, monitoring, scaling).
Experience
- 3–5 years as an AI Engineer or Python Backend Engineer working with production systems.
- Prior work with SaaS platforms, LLM/VLM integrations, or AI-first products is highly valued.
Demonstrated ability to maintain AI pipelines in production, not just prototypes.
Similar Jobs
Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Performs private equity, hedge fund, and debt fund accounting for global clients. Responsibilities include maintaining fund structures, recording investor commitments and capital activity, booking transactions, finalizing accounts, preparing NAV workbooks and reports, calculating management fees, preferred returns, carried interest, and performance ratios, and supporting investor and client reporting. The role also handles fund instruments such as bank debt, TRS, MBS, and CLOs while working UK/US night shifts in a hybrid Mumbai-based environment.
Cloud • Security • Software • Cybersecurity • Automation
Build and deploy internal AI-powered solutions across Sales, Marketing, and Customer Support. Diagnose workflows, identify bottlenecks, determine whether AI is appropriate, and own initiatives from discovery through deployment. Develop integrations, prototypes, agentic architectures, prompts, guardrails, and evaluation frameworks while measuring business, flow, adoption, and ROI outcomes. Partner with stakeholders, assess AI tools and models, document reusable patterns, and advance GitLab’s AI-first transformation in a remote environment.
Top Skills:
Agentic AiAnthropicGitlab Ci/CdGitlab Duo Agent PlatformGleanGraphQLJavaScriptLarge Language Models (Llms)MakeMarketoN8NOpenaiPrompt EngineeringPythonRelevance AiRest ApisRetrieval-Augmented Generation (Rag)SalesforceTypescriptWorkatoZendesk
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Design, deploy, automate, and support large-scale cloud networking across AWS, Azure, and GCP. Develop infrastructure-as-code, support Kubernetes networking and firewalls, establish monitoring and capacity strategies, troubleshoot complex network issues, and create technical roadmaps. Provide engineering leadership, mentor network development engineers, improve reliability and eliminate operational toil, maintain service ownership, and participate in on-call support and occasional travel.
Top Skills:
AnsibleAWSAzureBgpCheck PointChefCloudFormationDnsFortinetGCPGoJavaKubernetesLinuxPalo Alto NetworksPythonRubySpaceliftTcp/IpTerraformTls
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.



