Senior Algorithm Engineer – Customer Service Chatbot AI Agent
Shopee
Description
The Engineering and Technology team is at the core of the Shopee platform development. The team is made up of a group of passionate engineers from all over the world, striving to build the best systems with the most suitable technologies. Our engineers do not merely solve problems at hand; We build foundations for a long-lasting future. We don't limit ourselves on what we can or can't do; we take matters into our own hands even if it means drilling down to the bottom layer of the computing platform. Shopee's hyper-growing business scale has transformed most "innocent" problems into huge technical challenges, and there is no better place to experience it first-hand if you love technologies as much as we do.
About the Team:
The Marketplace Intelligence and Data team's mission is to build sustainable and efficient data and intelligence products to facilitate Shopee's business development. The team is responsible for Shopee e-commerce data warehouse construction, merchant and operation data product construction, all-link traffic data, product algorithms, including product release, control, information optimization, SPU library and its comparison business, marketing algorithms, including Merchandising, Product Selection, Recommendation Algorithm, Evaluation Algorithms, User Profiling, and in addition, basic AI capabilities, such as Machine Translation, Speech Algorithm, Image Algorithm, and Real-person Authentication.
The Customer Service Chatbot team is building the next-generation multilingual and multi-scenario customer service systems for consumers and merchants. The team focuses on applying large language models (LLMs), AI agents, reinforcement learning (RL), and other advanced technologies to real-world customer service scenarios. Through Post-Training, Agentic RL, Harness, and Loop Engineering, we continuously improve the model’s understanding, decision-making, planning, and execution capabilities.
We have already deployed a new intelligent customer service system based on an AI Agent architecture. Going forward, we aim to leverage Harness and Loop Engineering to align with human-level capabilities, evolving the system from a “question-answering chatbot” into an “intelligent customer service agent that understands context, plans and executes solutions, and self-improves,” reaching and exceeding the performance of human agents.
Job Description:
- Working around AI Agent architecture and Harness / Loop Engineering / Post-Training / Agentic RL, you will participate in or lead the following:
- AI Agent Architecture and Harness (Agent Architecture and Runtime Framework)
- Design and iterate on AI Agent architectures (Planning / Tool Use / Memory / Multi-Agent coordination, etc.), and build toolsets, context engineering, SOP/knowledge integration, and other Harness components to fully utilize foundation model capabilities in customer service scenarios.
- Improve Agent system efficiency through model layering, parallel scheduling, context compression, and cache reuse, reducing end-to-end latency and inference cost.
- Loop Engineering (Data and Training Self-Evolution Loop)
- Build a self-evolution loop: “real online feedback (CSAT, escalation to human, human rewrites, etc.) → automated evaluation → training data generation → Post-Training / RL → gray-scale validation,” enabling the system to continuously improve with real traffic.
- With the goal of “reaching and surpassing human customer service,” establish an automated evaluation and LLM-as-a-Judge framework to measure performance using both offline metrics and business metrics.
- Post-Training and Agentic RL
- Design post-training strategies for customer service scenarios, including SFT / DPO / GRPO / PPO, constructing preference and reward signals based on explicit and implicit feedback (user ratings, CSAT, escalations, human rewrites, etc.).
- Design Agentic RL objectives and training methods around key Agent decision points (action selection, tool calling, clarification/follow-up strategies, etc.), exploring the application of RLHF / RLAIF in dialogue policies.
- Cutting-Edge Practices
- Continuously track the latest research advances in Post-Training, RL, Agents, and Harness / Loop Engineering, and drive their adoption in business applications.
Requirements:
- Master’s degree or above in Computer Science, Artificial Intelligence, or a related field
- At least 3 years of Algorithm Engineering / Machine Learning experience are welcome
- Strong computer science and mathematics foundation, with solid algorithm and data structure skills.
- Systematic knowledge of machine learning and deep learning fundamentals, with a deep understanding of Transformer and LLM architectures.
- Substantial research or project experience in at least one or two of the following areas:
- Post-Training (SFT / DPO / GRPO, etc.);
- Reinforcement Learning / Agentic RL;
- Agent direction (Planning / Reasoning / Tool Use / Memory / Multi-Agent, etc.);
- Harness / Loop Engineering (Agent runtime frameworks, automated evaluation, and data self-evolution loops).
- Proficient in Python and familiar with mainstream deep learning frameworks such as PyTorch / TensorFlow.
- Interest in cross-disciplinary problems spanning research, engineering, and business, with willingness to validate and refine methods in real-world scenarios.
Preferred Qualifications:
- Experience working on Agent or LLM projects at leading AI teams (OpenAI, Google DeepMind, Anthropic, Meta, ByteDance Seed, Alibaba Tongyi, Tencent Hunyuan, Baidu Wenxin, etc.).
- Experience designing and scaling AI Agent systems, or practical experience improving Agent inference efficiency (model distillation, model routing, parallel scheduling, context engineering, etc.).
- Publications at top-tier conferences/journals in NLP, LLM, RL, Agent, or related fields (e.g., ACL, EMNLP, NeurIPS, ICLR, ICML), or complete research work.
- Hands-on experience participating in LLM Post-Training / Reinforcement Learning (including Agentic RL) projects.
- Practical experience applying RLHF / Agentic RL in dialogue/customer service/Agent scenarios.
- Familiarity with distributed training and large model optimization (e.g., DeepSpeed, FSDP, Megatron, etc.).