OneMagnify Logo

OneMagnify

AI Ops Engineer

Posted 6 Days Ago
Be an Early Applicant
In-Office
Chennai, Tamil Nadu, IND
Mid level
In-Office
Chennai, Tamil Nadu, IND
Mid level
Build and operate secure, scalable cloud infrastructure for AI and generative AI applications. Responsibilities include MLOps and LLMOps pipelines, CI/CD automation, infrastructure as code, container orchestration, observability, reliability engineering, security governance, and developer enablement. The role supports AI engineers, software engineers, and analytical modelers while managing production deployments, model versioning, evaluation, monitoring, drift detection, and hallucination mitigation.
The summary above was generated by AI

OneMagnify is an AI native, platform-enabled B2B digital agency operating at the intersection of data, technology, and creativity. We help complex organizations drive measurable business outcomes by building smarter customer experiences and delivering highly integrated solutions across digital, media, and technology. By combining deep industry expertise with advanced analytics and artificial intelligence, we enable our clients to make better decisions, move faster, and compete more effectively in dynamic markets.

We are seeking a motivated and talented AI Operations Engineer with strong cloud infrastructure and AI operations expertise to build, automate, and secure the platform that powers advanced artificial intelligence solutions.

The Impact You’ll Have:

In this role, you will support quantitative analytics and AI engineering teams by designing, automating, and operating end-to-end cloud and AI infrastructure. Responsibilities include managing CI/CD pipelines, containerized microservices, observability platforms, and governance controls to ensure AI applications and models run safely, reliably, and at scale in production environments.

What you’ll do:

  • Cloud Platform Engineering: Architect and operate highly available, multi-service AI infrastructure on the cloud, managing the full lifecycle of compute, storage, networking, and security for AI workloads.
  • AI-Ops / LLMOps Pipelines: Establish robust MLOps and LLMOps pipelines covering the end-to-end lifecycle of Generative AI tools — including model deployment, version control, prompt and artifact tracking, automated evaluation, and continuous monitoring for performance, drift, and hallucination mitigation.
  • CI/CD Automation: Design and maintain automated build, test, and deployment pipelines for full-stack AI applications, ensuring seamless and secure continuous integration and delivery across front-end, back-end, and AI components.
  • Infrastructure as Code: Manage cloud infrastructure using Infrastructure as Code (e.g., Terraform, Cloud Build, Kubernetes manifests) to deliver reproducible, auditable, and scalable environments.
  • Containerization & Orchestration: Build and operate containerized microservices (Docker/Kubernetes), managing scaling, rolling deployments, resource optimization, and service resilience for AI workloads.
  • Observability & Reliability: Implement comprehensive monitoring, logging, tracing, alerting, and SRE practices to ensure platform reliability, availability, and performance of AI applications in production.
  • Security & Governance: Embed security across the platform — managing user identities and controlling access rights, safeguarding sensitive credentials, network policies, and data protection — ensuring all AI workloads meet Ford's strict data privacy, security, and compliance standards.
  • Developer Enablement: Work closely with AI engineers, software engineers, and analytical modelers to provide self-service tooling, environments, and automated workflows that remove friction from development to production.

What you’ll need:

  • Education: Master's or Bachelor's degree in Computer Science, Software Engineering, Cloud Computing, Data Engineering, or a related technical discipline.
  • DevOps/Cloud Engineering Experience: 3–5 years of overall experience in cloud engineering, DevOps, or SRE, with at least 1–2 years of dedicated, hands-on experience deploying and operating AI, ML, and Generative AI applications in production.
  • AI/MLOps Engineering: Proven track record of implementing CI/CD for AI workloads, containerization (Docker/Kubernetes), and cloud infrastructure management (Terraform, Cloud Build, GKE).
  • Automation & Reliability: Demonstrated experience with infrastructure automation, incident response, and building observable, self-healing production systems.
  • Cloud Platform: Extensive hands-on experience with a leading cloud provider (e.g., Google Cloud Platform), including Cloud Run, Cloud Build, GKE, GCS, BigQuery, IAM, VPC networking, and Secret Manager.
  • DevOps / SRE Practices: Strong proficiency in CI/CD tooling, GitOps, containerization (Docker), orchestration (Kubernetes), Infrastructure as Code (Terraform), and cloud-native monitoring and logging.
  • AI-Ops / LLMOps: Proficiency in tools and platforms for model deployment, prompt/model versioning, evaluation, tracing, and monitoring of LLM and GenAI outputs.
  • Scripting & Automation: Hands-on experience with a programming/scripting language such as Python, or Bash for automating infrastructure and operational tasks, and for building tooling that serves collaborators.
  • Software Engineering Fundamentals: Solid understanding of full-stack application architecture and modern deployment patterns for AI tools, with the ability to integrate front-end, back-end, and AI services reliably.
  • Security & Compliance: Familiarity with cloud security best practices, identity and access management, secrets management, and compliance standards in regulated environments.
  • Analytics Workflow Understanding: Awareness of the typical workflows of data scientists and modelers (data wrangling, feature engineering, model validation) so you can build reliable platforms and pipelines that serve them.

Future-Ready Skills (Nice to Have):

  • Experience in integrated marketing, digital agency, marketing services, or consulting environments preferred.
  • Previous exposure to the Banking, Financial Services, or Credit Analytics industries. Experience with Machine Learning engineering and model serving frameworks, or relevant cloud certifications (e.g., Google Cloud Professional DevOps Engineer / Cloud Architect).

Benefits

We offer a comprehensive benefits package including Medical Insurance, PF, Gratuity, paid holidays, and more.

We are an equal opportunity employer

We believe that Innovative ideas and solutions start with unique perspectives. That’s why we’re committed to providing every employee a workplace that’s free of discrimination and intolerance. We’re proud to be an equal opportunity employer and actively search for like-minded people to join our team.

OneMagnify Chennai, Tamil Nadu, IND Office

Chennai, India

Similar Jobs

One Month Ago
In-Office
Chennai, Tamil Nadu, IND
Senior level
Senior level
Automotive
Build and operate AWS-based MLOps platforms for autonomous-driving workloads. Maintain highly available multi-zone environments, distributed multi-GPU Ray training clusters, and infrastructure for compute, storage, networking, and security. Develop Airflow and MLflow pipelines, support GitHub-based CI/CD for ML code, models, and infrastructure, troubleshoot platform issues, and ensure workflows are reproducible, traceable, and auditable.
Top Skills: Amazon EksApache AirflowAWSCi/CdDevOpsGitKubernetesMlflowMlopsPythonRayTerraform
Senior level
Information Technology • Software • Consulting
The Senior AI / ML Ops Engineer will design and operate Agentic AI systems, implement MLOps practices, and ensure production automation is secure and reliable.
Top Skills: AiopsApi IntegrationsCi/CdDevOpsDynatraceLangchainLanggraphLangsmithMlopsPython
47 Seconds Ago
Hybrid
Chennai, Tamil Nadu, IND
Senior level
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Coordinates release readiness and lifecycle management for 60+ Veeva RIM integrations. Tracks dependencies, risks, remediation, testing, validation, stakeholder communications, cutover activities, and operational handoffs. Partners with product, regulatory, technical, quality, vendor, and consuming application teams to assess Veeva RIM data model impacts and ensure integration changes meet Agile, SDLC, and CSV expectations.
Top Skills: AgileAPIsCsvData MappingETLMiddlewareSafeScrumSdlcSsoVeeva RimVeeva Vault

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account