UOB Logo

UOB

Manager, GenAI L3 Support Engineer, Group Technology & Ops

Posted 13 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in Eastern Region, East
Mid level
Remote
Hiring Remotely in Eastern Region, East
Mid level
Own L3 production incident response for GenAI applications, including triage, debugging, hotfixes, root-cause analysis, and reliability improvements. Harden RAG pipelines, improve observability, manage Elasticsearch and Redis operations, maintain runbooks, and coordinate with infrastructure and platform teams. Implement secure data handling, guardrails, deployment controls, and resilient production patterns. Lead post-incident reviews and drive corrective actions with business stakeholders.
The summary above was generated by AI
Company: 1011 United Overseas Bank Ltd

About UOB

United Overseas Bank Limited (UOB) is a leading bank in ASEAN with a global network in Southeast Asia, Asia Pacific, Europe and North America. Operating through our head office in Singapore and banking subsidiaries in China, Indonesia, Malaysia, Thailand and Vietnam, we have a global network of about 430 branches and offices in 19 markets. At the heart of UOB is our culture, shaped by the UOB Way and anchored on our four values – Honourable, Enterprising, United and Committed. For more than 90 years, these values have guided how we do right by our customers, collaborate with one another and create long-term value for the communities we operate in. As One Bank, we are committed to helping our colleagues build sustainable careers grounded in purpose, supported by strong values, and enriched with meaningful opportunities to grow.


Job Description

  • Own L3 incident response for GenAI use cases in production. Triage, deep debug, fix or work with the vendor to resolve defects and improve reliability
  • Review, reproduce and root cause complex bugs across LangChain flows, Elasticsearch queries, Redis caches and safety layers such as Guardrails AI and LlamaGuard
  • Implement safe changes and hotfixes. Build, test and deploy code and config updates through Jenkins with proper approvals and rollbacks
  • Partner with business users to understand issue impact and edge cases. Translate vague problem reports into actionable steps and clear acceptance criteria
  • Harden use case pipelines. Add input validation, timeouts, retries, circuit breakers and fallback strategies to reduce customer facing impact
  • Improve RAG quality for stability. Tune chunking, retrieval parameters, query rewriting and caching to reduce latency and errors
  • Maintain and evolve runbooks, playbooks and knowledge articles. Keep them current and prove they work with regular game days
  • Drive observability for the use cases. Define golden signals, add OpenTelemetry traces, wire up Prometheus metrics and curate Grafana dashboards
  • Manage production hygiene. Track error budgets, SLOs and SLAs. Push for defect burn down and change quality
  • Coordinate with platform and infra teams when incidents touch OpenShift, GPUs, vLLM or NVIDIA Enterprise AI services
  • Perform safe data operations. Validate indices, manage Redis eviction strategies and handle backfills and reindex tasks with minimal risk
  • Champion secure by default practices. Enforce secrets handling, PII redaction and prompt or output guardrails
  • Lead post incident reviews with blameless RCAs and concrete corrective actions. Close the loop with business stakeholders
  • Proactively surface reliability risks and propose code or architecture changes that remove recurring failure modes

Job Requirement

  • 3 to 4 years of software engineering or SRE with at least one year supporting AI or data intensive services in production
  • Strong Python. Comfortable reading and fixing LangChain code, FastAPI or Flask services, async patterns and task queues
  • Hands on with LangChain in real projects. Chains, tools, retrievers, memory and callback handlers
  • Elasticsearch proficiency. Query DSL, relevance tuning, index lifecycle, scaling, snapshots and Kibana for triage
  • Redis proficiency. Caching patterns, pub or sub, streams, eviction, persistence and troubleshooting latency or timeouts
  • Safety and guardrails experience. Guardrails AI configuration, schema and validator design, LlamaGuard or similar classifiers in the loo
  • CI or CD with Jenkins. Declarative pipelines, approvals, artifacts, environment promotion and rollback strategy
  • Observability depth. Prometheus metrics, Grafana dashboards, OpenTelemetry traces and logs.
  • Able to decide which signals matter and set actionable alerts
  • Solid Git. Branching, pull requests, code reviews and release tagging. Comfortable with feature flags and canary or blue green patterns
  • Working knowledge of Kubernetes and OpenShift as a strong plus. Debugging pods, logs, events and basic resource tuning
  • Familiarity with LLM serving stacks such as vLLM or NVIDIA Enterprise AI is a plus. Know how to read model server logs, timeouts and token throughput
  • Data handling discipline. Understanding of PII, masking, prompt redaction and safe logging
  • Clear communicator who can write crisp incident updates, RCAs and user facing notes

Additional Requirements

Be a Part of the UOB Family

UOB is an equal opportunity employer. UOB does not discriminate on the basis of a candidate's age, race, gender, color, religion, sexual orientation, physical or mental disability, or other non-merit factors. All employment decisions at UOB are based on business needs, job requirements and qualifications. If you require any assistance or accommodations to be made for the recruitment process, please inform us when you submit your online application.

Apply now and make a Difference

Similar Jobs

11 Hours Ago
Remote
Entry level
Entry level
Fashion • Retail
Provides customer service and sales support in a Mountain Warehouse store. Responsibilities include processing transactions, maintaining product displays and stock levels, supporting promotions, meeting sales targets, answering customer questions, and contributing to store operations. The role requires teamwork, reliability, enthusiasm for retail, strong communication, and a commitment to creating a welcoming shopping experience while following company policies and health and safety standards.
14 Hours Ago
In-Office or Remote
Northern Province, IND
Expert/Leader
Expert/Leader
Fintech • Financial Services
Design and deliver end-to-end eGRC solutions by translating Risk & Compliance requirements into technical architectures, lead implementations on GRC platforms (ServiceNow, Archer, SAI360), guide development teams, maintain architecture artifacts, and ensure delivery, security, and sustainability across multiple initiatives.
Top Skills: APIsArcherAWSAzureData PipelinesGCPSai360Servicenow Irm
Senior level
Agency
Provides regional technical leadership on risk management, compliance, safeguarding, fraud prevention, audits, partner oversight, and corrective actions across multiple country programs. Advises teams on policies and donor requirements, analyzes systemic risks and trends, develops practical tools and guidance, reviews agreements and proposals, supports investigations and audit closure, and escalates sensitive or high-risk issues. The role also mentors country compliance staff and promotes consistent application of global standards.

What you need to know about the Chennai Tech Scene

To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account