Develop, deploy, and maintain LLM pipelines and RAG QA/search systems; design and optimize prompts and multi-agent LLM architectures; operate multi‑GPU/cluster inference; build evaluation pipelines for model quality, bias, and hallucination; collaborate with product and CS teams to integrate conversational AI.
Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million people in 100+ countries for our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-asset products. Binance offerings range from trading and finance to education, research, payments, institutional services, Web3 features, and more. We leverage the power of digital assets and blockchain to build an inclusive financial ecosystem to advance the freedom of money and improve financial access for people around the world.
We are seeking a highly skilled professional to join our team, focusing on advancing through innovative AI solutions.
The successful candidate will develop and refine Large Language Models (LLMs) to extract actionable insights, improve business decision-making, and optimize prompt design for more accurate outputs. Additionally, the role includes creating scalable and robust LLM/RAG frameworks tailored to customer service scheduling, fostering innovation and maintaining a competitive market edge.
This role is 100% Remote, Work from Home based.
Responsibilities
- Own the full LLM pipeline from data preparation to production real case usage.
- Design, iterate and optimize prompts (zero-/few-shot, chain-of-thought, tool-calling, etc.) to maximize model utility and safety across products and languages.
- Build and maintain Retrieval-Augmented Generation (RAG) QA/search systems that connect to multi-source knowledge bases.
- Familiar with vLLM/SGLang inference architectures and have proven experience deploying and operating LLM services on multi‑GPU or cluster environments.
- Design, implement and operate multi‑agent LLM architectures (e.g. LangGraph, CrewAI, AutoGen) including task decomposition, agent orchestration, memory sharing and tool‑calling workflows.
- Develop evaluation pipelines (automatic metrics & human feedback) to measure prompt and model quality, bias, and hallucination rates.
- Collaborate with product and CS teams to integrate AI models into conversational Chatbot in different scenarios.
- Track cutting-edge research, author tech blogs, and keep improve current architecture.
Requirements
- Master’s Degree or higher in Computer Science, Data Science or related field..
- At least 2 years of deep-learning/NLP experience, including 1+ year practical LLM work (SFT, DPO, RAG, quantization, inference optimization, etc.).
- Demonstrated prompt engineering & tuning expertise (few-shot design, structured prompting, prefix-/p-tuning, reward re-ranking, safety filtering).
- Practical experience building and deploying multi‑agent LLM workflows, with understanding of agent‑orchestrator patterns, shared memory, long‑horizon planning and guard‑rail design.
- Proficient in both English and Chinese communication for efficient cross team collaboration
Why Binance
• Shape the future with the world’s leading blockchain ecosystem
• Collaborate with world-class talent in a user-centric global organization with a flat structure
• Tackle unique, fast-paced projects with autonomy in an innovative environment
• Thrive in a results-driven workplace with opportunities for career growth and continuous learning
• Competitive salary and company benefits
• Work-from-home arrangement (the arrangement may vary depending on the work nature of the business team)
Binance is committed to being an equal opportunity employer. We believe that having a diverse workforce is fundamental to our success.
By submitting a job application, you confirm that you have read and agree to our Candidate Privacy Notice.
Similar Jobs
Professional Services • Real Estate • Consulting
Support cost management for data centre projects including estimating, quantity take-offs, tender support, contract administration, cost tracking, monthly reporting, progress valuations, change control, value engineering, and maintaining audit-ready cost records under guidance of a Cost Manager.
Top Skills:
Costx
Professional Services • Real Estate • Consulting
Lead mechanical engineering for MEP systems on regulated pharmaceutical/biotech projects. Translate client technical requirements into design deliverables, review contractor documentation, prepare URS and FAT/SAT protocols, support C&Q activities, ensure GMP, SHE and EHS compliance, and provide technical guidance through design, procurement, FAT/SAT and commissioning.
Top Skills:
Black UtilitiesBmsC&QClean UtilitiesDqEmsFatFire SuppressionGmpHvacIqOqPressure VesselsSatUrs
Professional Services • Real Estate • Consulting
Serve as the primary client sustainability contact, defining and assuring KPIs across energy, embodied/operational carbon, water, and waste. Support Sustainability Delivery Plans, review contractor compliance, coordinate stage reviews and verification (LEED/BREEAM), and compile KPI tracking data to ensure client standards and stage-gate objectives are met.
Top Skills:
BreeamCarbon TrackingEhs StandardsEmbodied Carbon AnalysisEnergy ModelingGlobal Engineering Sustainability ChecklistKpi Tracking DashboardLeedOperational Carbon AnalysisSustainability Delivery PlanWaste TrackingWater Tracking
What you need to know about the Chennai Tech Scene
To locals, it's no secret that South India is leading the charge in big data infrastructure. While the environmental impact of data centers has long been a concern, emerging hubs like Chennai are favored by companies seeking ready access to renewable energy resources, which provide more sustainable and cost-effective solutions. As a result, Chennai, along with neighboring Bengaluru and Hyderabad, is poised for significant growth, with a projected 65 percent increase in data center capacity over the next decade.

