招聘城市:伦敦
…an intern to work on evaluation and reliability infrastructure for a real-world LLM agent system in the UA performance marketing field. The agent performs multi-step reasoning, retrieves context, selects tools, executes actions, handles user confirmations, and interacts with external services.
The goal of this internship is to build transferable expertise in agent evaluation engineering: evaluating tool use, measuring trajectory quality, designing benchmarks, analyzing traces, comparing model and prompt variants, and improving the reliability of agentic AI systems.
This role is ideal for someone interested in future opportunities in LLM agent evaluation, AI safety evaluation, research engineering, LLMOps, or applied AI infrastructure.Research the state-of-the-art agentic workflow evaluation frameworks in the industry and in the research field.
Apply the theory to build automated evaluation pipelines that can run agent scenarios, capture execution artifacts, score results, and detect regressions.
Evaluate tool-use behavior, including whether the agent…
招聘城市:上海
…Intern)LLM-大模型强化学习算法工程师 前端|后端|系统集成|软件开发 实习 本科 5 20000以上
发布时间 2026-08-20 16:50:27
职位有效期 2026-11-29 23:59:59
专业 人工智能;信息安全;光电信息工程;数学;数理统计;新一代电子信息技术(含量子技术等);测试专用专业;计算金融交叉试验班;通信工程
工作地址 上海市市辖区松江区
职位要求 岗位要求计算机、人工智能、机器学习、数学等相关专业本科高年级、硕士或博士在读。熟悉机器学习和深度学习基础,了解大模型训练、微调或 post-training 基本流程。对强化学习有基础理解,熟悉 PPO、GRPO、reward modeling、policy optimization、credit assignment 等概念者优先。熟练使用 Python,熟悉 PyTorch;具备良好的工程实现和实验分析能力。对 AI Agent…
招聘城市:广东省深圳
…持续抬升基础模型的能力上限。设计并实现奖励信号与自动评测器(如代码判题、数学答案校验、指令约束校验、Agent 任务成功率),保证 reward 正确性锚定在模型之外、不产生自我强化误差。参与高质量训练数据构建:数据合成与蒸馏、去重去污染、质量打分、失败样本挖掘、偏好数据构建,以及按「有效监督 token / 考点分布 / 失败机制」的定向配比策略。打通并优化 mid-training / SFT / RL 的训推链路(如 verl / SGLang / vLLM / Megatron),定位并修复训推不一致、熵坍缩、显存与吞吐瓶颈等工程问题。跟踪 LLM post-training、RL for LLM、reasoning models、Agentic RL、agent evaluation 等方向的前沿论文与开源框架,并快速复现或落地。岗位要求计算机、人工智能、机器学习、数学等相关专业硕士或…
招聘城市:上海
…数学答案校验、指令约束校验、Agent 任务成功率),保证 reward 正确性锚定在模型之外、不产生自我强化误差。参与高质量训练数据构建:数据合成与蒸馏、去重去污染、质量打分、失败样本挖掘、偏好数据构建,以及按「有效监督 token / 考点分布 / 失败机制」的定向配比策略。打通并优化 mid-training / SFT / RL 的训推链路(如 verl / SGLang / vLLM / Megatron),定位并修复训推不一致、熵坍缩、显存与吞吐瓶颈等工程问题。跟踪 LLM post-training、RL for LLM、reasoning models、Agentic RL、agent evaluation 等方向的前沿论文与开源框架,并快速复现或落地。
单位信息
单位名称 深圳虾皮信息科技有限公司
单位性质 三资企业 单位行业 信息传输、软件和信息技术服务业
标签
隶属单位 下属单位 单位官微
单位图片…
招聘城市:新加坡
岗位职责:
Business Unit
What the Role Entails
We are seeking highly motivated Research Interns to work on cutting-edge problems in multimodal foundation models and multimodal agents.
The intern will contribute to advancing models that understand, generate, and act across multiple modalities (e.g., vision, language, video, audio, and GUI environments), with applications in embodied intelligence, computer-use agents, and world models.
You will collaborate closely with a team of researchers and engineers to design novel algorithms, build large-scale training pipelines, and publish at top venues (e.g., CVPR, ICCV, NeurIPS, ICLR, ACL).
Responsibilities
- Conduct original research on multimodal foundation models or multimodal agents
- Implement and experiment with large-scale models (training, fine-tuning, evaluation)
- Design new model architectures, objectives, or data pipelines
- Work with large multimodal datasets (image, video, text, UI trajectories, etc.)
- Contribute to papers, technical reports, and open-source projects
- Collaborate with cross-functional teams…
招聘城市:新加坡
…ICCV, NeurIPS, ICLR, ACL, etc.) .
Experience training large models or working with distributed systems.
Experience with multimodal datasets and evaluation benchmarks.
Fluency in English and Mandarin to collaborate effectively with international teams and HQ stakeholders.
Must-Have ExperienceAt least one of the following areas:
Vision-language models (VLMs)
Large language models (LLMs)
Video understanding and generation
Reinforcement learning or imitation learning
Transformer architectures and scaling laws
Multimodal alignment (contrastive learning, instruction tuning)
Agent training (RLHF/RLAIF, planning, tool use)
Synthetic data generation or simulation environments
Experience with long-context training or memory mechanisms
Equal Employment Opportunity at Tencent
As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.
Work Location: Singapore-CapitaSky
岗位要求:
岗位要求…
招聘城市:新加坡
…with at least one of:
- Vision–language models
- Large language models
- Video understanding/generation
- Reinforcement learning or imitation learning
- Strong problem-solving and research skills
Publications at top conferences (CVPR, ICCV, NeurIPS, ICLR, ACL, etc.)
Experience training large models or working with distributed systems
Experience with multimodal datasets and evaluation benchmarks
Familiarity with:
- Transformer architectures and scaling laws
- Multimodal alignment (contrastive learning, instruction tuning)
- Agent training (RLHF/RLAIF, planning, tool use)
- Synthetic data generation or simulation environments
- Experience with long-context training or memory mechanisms
Equal Employment Opportunity at Tencent
As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.
Work Location: Singapore-CapitaSky
岗位要求:
岗位要求详情请见上方岗位职责内说明