牛大妈在校招职位搜索Research Intern Reinforcement 有 21 条结果

招聘城市:新加坡
…groups on technology and operational platforms, as well as the construction and operation of R&D management and data centers, TEG provides users with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.
What the Role Entails
Research directions include but are not limited to:
- RL Algorithms for Reasoning Models:Design robust RL training recipes (PPO / GRPO / GSPO variants) for large-scale reasoning models. Tackle training instability, reward hacking, and policy collapse in long-horizon and async settings. Explore how to bridge the gap between RL post-training and genuine reasoning capability improvement.
- RL for Autonomous Agents: Build RL pipelines for long-horizon terminal agents and tool-use agents. Investigate credit assignment, exploration strategies…
招聘城市:新加坡
岗位职责:
Business Unit
What the Role Entails
Responsibilities:
1. Conduct research on RL algorithms for multimodal models, including diffusion models for image, video, and 3D generation, autoregressive models for multimodal understanding, and potentially unified multimodal frameworks.
2. Design and develop RL infrastructure and reward modeling strategies to enable efficient large-scale training, improve training stability, and mitigate reward hacking and related failure modes.
3. Explore next-generation RL paradigms that more directly and effectively learn from environment feedback.
Who We Look For
Requirements:
1. Currently enrolled as a PhD student in Computer Science or a closely related field.
2. Demonstrated strong research capability, with publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR, ICCV, ECCV, SIGGRAPH.
3. Strong hands-on programming skills, with solid experience in deep learning system implementation, model training and inference optimization, CPU/GPU acceleration, and distributed training and inference.
4. Prior experience with…
招聘城市:新加坡
…As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.
What the Role Entails
Job Description1. Conduct research on RL algorithms for multimodal models, including diffusion models for image, video, and 3D generation, autoregressive models for multimodal understanding, and potentially unified multimodal frameworks.
2. Design and develop RL infrastructure and reward modeling strategies to enable efficient large-scale training, improve training stability, and mitigate reward hacking and related failure modes.
3. Explore next-generation RL paradigms that more directly and effectively learn from environment feedback.
Who We Look For
1. Currently enrolled as a PhD student in Computer Science or a closely related field.
2. Demonstrated strong research capability, with publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR…
招聘城市:新加坡
…with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.
What the Role Entails
Conduct research on RL algorithms for multimodal models, including diffusion models for image, video, and 3D generation, autoregressive models for multimodal understanding, and potentially unified multimodal frameworks
Design and develop RL infrastructure and reward modeling strategies to enable efficient large-scale training, improve training stability, and mitigate reward hacking and related failure modes.
Explore nextgeneration RL paradigms that more directly and effectively learn from environment feedback.
Who We Look For
Currently enrolled as a PhD student in Computer Science or a closely related field
Demonstrated strong research capability, with publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR…
招聘城市:新加坡
…with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.
What the Role Entails
Conduct research on RL algorithms for multimodal models, including diffusion models for image, video, and 3D generation, autoregressive models for multimodal understanding, and potentially unified multimodal frameworks
Design and develop RL infrastructure and reward modeling strategies to enable efficient large-scale training, improve training stability, and mitigate reward hacking and related failure modes.
Explore nextgeneration RL paradigms that more directly and effectively learn from environment feedback.
Who We Look For
Currently enrolled as a PhD student in Computer Science or a closely related field
Demonstrated strong research capability, with publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR…
招聘城市:新加坡
…with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.
What the Role Entails
Conduct research on RL algorithms for multimodal models, including diffusion models for image, video, and 3D generation, autoregressive models for multimodal understanding, and potentially unified multimodal frameworks
Design and develop RL infrastructure and reward modeling strategies to enable efficient large-scale training, improve training stability, and mitigate reward hacking and related failure modes.
Explore nextgeneration RL paradigms that more directly and effectively learn from environment feedback.
Who We Look For
Currently enrolled as a PhD student in Computer Science or a closely related field
Demonstrated strong research capability, with publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR…
招聘城市:新加坡
…with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.
What the Role Entails
Conduct research on RL algorithms for multimodal models, including diffusion models for image, video, and 3D generation, autoregressive models for multimodal understanding, and potentially unified multimodal frameworks
Design and develop RL infrastructure and reward modeling strategies to enable efficient large-scale training, improve training stability, and mitigate reward hacking and related failure modes.
Explore nextgeneration RL paradigms that more directly and effectively learn from environment feedback.
Who We Look For
Currently enrolled as a PhD student in Computer Science or a closely related field
Demonstrated strong research capability, with publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR…
招聘城市:新加坡
…with a full range of customer services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.
What the Role Entails
Conduct research on RL algorithms for multimodal models, including diffusion models for image, video, and 3D generation, autoregressive models for multimodal understanding, and potentially unified multimodal frameworks
Design and develop RL infrastructure and reward modeling strategies to enable efficient large-scale training, improve training stability, and mitigate reward hacking and related failure modes.
Explore nextgeneration RL paradigms that more directly and effectively learn from environment feedback.
Who We Look For
Currently enrolled as a PhD student in Computer Science or a closely related field
Demonstrated strong research capability, with publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR…
招聘城市:新加坡
…services. As the operator of the largest networking, devices, and data center in Asia,TEG also leads the Tencent Technology Committee in strengthening infrastructure R&D through internal and distributed open source collaboration, constructing new platforms and supporting business innovation.
What the Role Entails
1. Conduct research on RL algorithms for multimodal models, including diffusion models for image, video, and 3D generation, autoregressive models for multimodal understanding, and potentially unified multimodal frameworks.
2. Design and develop RL infrastructure and reward modeling strategies to enable efficient large-scale training, improve training stability, and mitigate reward hacking and related failure modes.
3. Explore next-generation RL paradigms that more directly and effectively learn from environment feedback.
Who We Look For
1. Currently enrolled as a PhD student in Computer Science or a closely related field.
2. Demonstrated strong research capability, with publications in top-tier conferences such as ICML, NeurIPS, ICLR, CVPR…

腾讯(tencent) Research Intern 107048

兼职 新加坡
招聘城市:新加坡
…and publish at top venues (e.g., CVPR, ICCV, NeurIPS, ICLR, ACL).
Responsibilities
- Conduct original research on multimodal foundation models or multimodal agents
- Implement and experiment with large-scale models (training, fine-tuning, evaluation)
- Design new model architectures, objectives, or data pipelines
- Work with large multimodal datasets (image, video, text, UI trajectories, etc.)
- Contribute to papers, technical reports, and open-source projects
- Collaborate with cross-functional teams on research prototypes
Who We Look For
Currently pursuing a PhD or Master’s in Computer Science, AI, Machine Learning, or related fields
Strong background in deep learning and machine learning fundamentals
Solid programming skills in Python and PyTorch/JAX
Experience with at least one of:
- Vision–language models
- Large language models
- Video understanding/generation
- Reinforcement learning or imitation learning
- Strong problem-solving and research skills
Publications at top conferences (CVPR, ICCV, NeurIPS, ICLR, ACL, etc.)
Experience training large models or working with…
招聘城市:贝尔维尤
…models. Research areas include but are not limited to Reinforcement Learning Algorithms, Reward Modeling, and World Models. We will conduct large-scale experiments of RL algorithms in scenarios such as complex reasoning and autonomous agents, deliver impactful algorithms for real world applications, and publish influential research papers.
Who We Look For
Requirements & QualificationsThe ideal intern candidates are those who
Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or related fields from a top university,
are self-motivated and excited about developing novel techniques,
have research experiences in natural language processing or machine learning,
are proficient in Python programming and experienced in developing with deep learning frameworks such as PyTorch.
have good publication track records and history of creativity and intellectual flexibility,
have excellent communication and teamwork skills, capable of collaborating with cross-functional teams to drive project success and innovation.
Intern duration: 3 months (with the possibility of extension)…
招聘城市:伦敦
岗位职责:
Business Unit
LIGHTSPEED STUDIOS is made up of passionate players who advance the art & science of game development through great stories, great gameplay, and advanced technology. We are focused on bringing next generation experiences to gamers who want to enjoy them anywhere, anytime, across multiple genres and devices.
About the Hiring Team
Lightspeed Tech Center is a R&D department under Lightspeed Studios which develop PUBG Mobile and other high-quality games. Our Tech Center leads the research, exploration, and discovery of innovative technologies and provides technical services for all games during all phases of life cycle, including engine, audio, QA, AI, next generation game, technical cooperation, etc.
What the Role Entails
1. Assist in researching the model's ability to understand generated content, including parsing semantics, objects, relationships, and spatial structures.
2. Help implement state tracking and evaluate consistency modeling for generated videos.
3. Participate in exploring…
招聘城市:东京
…games. Our Tech Center leads the research, exploration, and discovery of innovative technologies and provides technical services for all games during all phases of life cycle, including engine, audio, QA, AI, next generation game, technical cooperation, etc.
仕事内容/What the Role Entails
仕事内容:
1. ゲーム分野における強化学習(RL)アルゴリズムの実装・開発を行う。
2. 研究環境や実務環境におけるタスクに対して、RLモデルの訓練を行う。
3. 研究目標の達成に向けて、チームメンバーと積極的に協力する。
Job Responsibilities:
1. Explore and develop novel reinforcement learning (RL) algorithms to address challenges in game contexts.
2. Train RL models on innovative and practical tasks.
3. Collaborate proactively with team members to achieve research goals.
応募資格/Who We Look For
応募条件:
1. PythonおよびPyTorchなどの主要…
招聘城市:新加坡
岗位职责:
Business Unit
What the Role Entails
1.Conduct research and development on multimodal processing algorithms and models, including but not limited to image and video understanding and generation, as well as alignment and integration of multimodal information with textual data;
2. Design and optimize existing algorithms to improve performance and accuracy, ensuring a high-quality user experience;
3. Conduct in-depth research on and stay up to date with cutting-edge technologies in multimodal learning, NLP, CV, and related fields, and promptly apply new technologies to products.
4. Explore and develop reinforcement learning algorithms and frameworks, including but not limited to policy optimization, reward modeling
Who We Look For
1. Bachelor’s degree or above in Computer Science, Information Engineering, Pattern Recognition, Artificial Intelligence, or related fields;
2. Proficient in fundamental algorithms and applications related to computer vision and image processing; familiar with at least one deep learning…
招聘城市:伦敦
…unlock the true potential of their games.
What the Role Entails
This role focuses on advancing AI Agent architecture, reasoning, and autonomy through applied research and experimentation. Interns will work on cutting-edge projects involving foundation models, planning algorithms, and multi-agent systems, contributing to both theoretical innovation and practical applications.
Key Responsibilities
• Conduct research on AI foundation models to enhance reasoning and data-driven analytical capabilities.
• Design and develop advanced planning and decision-making algorithms that enable Agents to autonomously perform complex tasks and improve operational efficiency.
• Investigate tool-use mechanisms to allow Agents to effectively integrate and utilize diverse data analysis and system tools for greater adaptability in real-world environments.
• Explore and implement end-to-end reinforcement learning frameworks and multi-agent collaboration algorithms to improve coordination and optimization across intelligent systems.
Who We Look For
Qualifications
• Master’s or PhD student in Computer Science, Mathematics, Statistics…
招聘城市:新加坡
…and resources to our network of developers and partner studios around the world to help them unlock the true potential of their games.
What the Role Entails
We are hiring internships for the data & AI group. Join our team to build AI bots in games using reinforcement learning and latest genAI technologies. Support the production of scalable and optimised AIService for AI bots. Run experiments to test the performance of deployed models, and identifies and resolves bugs that arise in the process. There is also opportunities of research new approaches to improve the performance.
Who We Look For
- Currently enrolled in a Master's or Ph.D. program in a quantitative discipline such as statistics, machine learning, computer science, math, etc.
- Expertise in one or more of the following areas with practical experience: deep learning, reinforcement learning, LLM.
- Experience in video game, reinforcement learning, and development with GenAI is preferred…
招聘城市:奥克兰
…or scenario-based implementation is a plus.
Basic understanding of development, testing, and operational workflows; strong product mindset and analytical skills.
Exceptional creativity and systematic design ability, capable of independently drafting requirements documentation and coordinating cross-team communication.
Bonus points:
AI Technical Proficiency: Solid grasp of AI fundamentals and mainstream algorithms, with deep expertise in at least one AI domain (e.g., generative models, reinforcement learning) and its practical limitations.
AI Agent Expertise: Hands-on experience researching AI agent technologies, familiarity with mainstream frameworks, and awareness of deployment challenges.
Gaming AI Experience: Prior work applying AI in gaming contexts (e.g., smart NPCs, dynamic environment generation, player behavior prediction).
Equal Employment Opportunity at Tencent
As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported…
招聘城市:新加坡
…Identify opportunities for technical enhancement through keen data insights, and continuously optimize recommendation strategies and models to improve core business metrics.
Explore and conduct research on cutting-edge recommendation technologies.Be responsible for the development and performance optimization of overseas advertising recommendation algorithms. You will focus on one of the following specific areas:
Model & Algorithm: Optimize estimation models for precise ranking (pCTR/pCVR/pLTV) and recall(LTR). This includes but is not limited to research in ultra-long sequence modeling, multimodal large models for recommendation, multi-scenario & multi-task learning, and scaling up complex cross networks.
Bidding Mechanism: Optimize algorithmic strategies for cost attainment, volume scaling, and cold start. This includes but is not limited to applying reinforcement learning and generative bidding techniques.
Industry-Specific Algorithms: Optimize models and strategies for key industries such as gaming and e-commerce. This includes but is not limited to exploring AI-powered intelligent…
招聘城市:新加坡
…or Go.
Networking: TCP/IP, DNS, HTTP basics.
Core Competencies:
- Analytical problem-solving and passion for infrastructure technologies.
- Ability to learn quickly in a fast-paced environment.
- Bilingual Fluency in English & Chinese to deal with both HQ and International stakeholders (written and verbal)
- Basic Mandarin communication skills are required to collaborate with China-based teams and access internal resources.
- Experience with cloud platforms (Tencent Cloud, AWS, or Azure).
- Familiarity with IaC tools (Terraform, Ansible) or observability stacks (ELK, Prometheus).
Knowledge of containerization (Docker/Kubernetes).
Experience with at least one of:
- Vision–language models
- Large language models
- Video understanding/generation
- Reinforcement learning or imitation learning
- Strong problem-solving and research skills
Publications at top conferences (CVPR, ICCV, NeurIPS, ICLR, ACL, etc.)
Experience training large models or working with distributed systems
Experience with multimodal datasets and evaluation benchmarks
Familiarity with:
- Transformer architectures and scaling laws
- Multimodal alignment (contrastive learning, instruction tuning)
- Agent…
招聘城市:帕罗奥多
…Darktide, Dune: Awakening from Funcom, V Rising from Stunlock Studios and many more. To learn more about Level Infinite, visit levelinfinite.com, and follow on Twitter, Facebook, Instagram and YouTube.
We are hiring an Applied Machine Learning Intern to join our algorithm team and work on the next-generation opinion system driven by cutting-edge LLMs.
Responsibilities:
Assist in designing, training, and evaluating NLP models (e.g., LLMs, transformers, BERT, GPT).
Preprocess and analyze large-scale text datasets.
Implement and optimize ML pipelines for real-world applications.
Stay updated with the latest advancements in NLP research.
Collaborate with cross-functional teams to integrate models into production systems.
Who We Look For
Expertise in deep learning and/or reinforcement learning with hands-on experience.
Pursuing a degree in Computer Science, AI, Linguistics, or related fields.
Familiarity with Python, PyTorch/TensorFlow, Hugging Face, and NLP libraries.
Basic understanding of deep learning, transformers…