牛大妈在校招职位搜索Video Generation Foundation Mo 有 22 条结果

招聘城市:新加坡
…rate videos.
Lead video pre-training at the 10-100 billion token scale, defining data mixture and curriculum learning strategies.
Explore Scaling Laws, define scientific scale-up paths, and continuously improve key model capabilities.
Track industry state-of-the-art , lead comparative experiments, drive technical roadmap decisions, and publish academic papers.
Who We Look For
Ph.D. in AI-related fields with first-author papers at top-tier conferences.
Proficient in the principles and engineering implementation of diffusion models and autoregressive generation.
Experience training video/image generation models from scratch; highly proficient in PyTorch and large-scale distributed training.
Deep understanding of the design trade-offs in Video VAE/Tokenizers, with 3+ years of relevant research experience.
Publications related to video generation or diffusion models.
Experience in core industry product R&D or hands-on experience with the latest technologies is preferred.
Experience leading end-to-end video foundation model…
招聘城市:上海
…可控性与生产效率
研究游戏场景下内容生成的关键问题,包括 style control、long-form generation、character-conditioned generation、multimodal consistency、interactive generation 与 user-steerable creation
研究如何将 foundation model 与产品工作流高效结合,使内容能力既能支撑专业生产,也能服务 AI 原生的 UGC 体验
与产品、设计、工程团队紧密协作,将研究成果转化为可在线迭代的生成能力与创作工具
职位要求
我们希望你具备
扎实的机器学习、深度学习与统计基础,对生成式模型的训练、后训练与评估有系统理解
熟悉以下一个或多个方向:LLM post-training、diffusion / autoregressive generation、multimodal models、story generation、audio / speech / image / video generation、controllable generation、creative AI tooling 具备较强的 research taste,能够围绕开放问题定义研究目标,设计严谨实验,并…

腾讯(tencent) Research Intern 107048

兼职 新加坡
招聘城市:新加坡
…ACL).
Responsibilities
- Conduct original research on multimodal foundation models or multimodal agents
- Implement and experiment with large-scale models (training, fine-tuning, evaluation)
- Design new model architectures, objectives, or data pipelines
- Work with large multimodal datasets (image, video, text, UI trajectories, etc.)
- Contribute to papers, technical reports, and open-source projects
- Collaborate with cross-functional teams on research prototypes
Who We Look For
Currently pursuing a PhD or Master’s in Computer Science, AI, Machine Learning, or related fields
Strong background in deep learning and machine learning fundamentals
Solid programming skills in Python and PyTorch/JAX
Experience with at least one of:
- Vision–language models
- Large language models
- Video understanding/generation
- Reinforcement learning or imitation learning
- Strong problem-solving and research skills
Publications at top conferences (CVPR, ICCV, NeurIPS, ICLR, ACL, etc.)
Experience training large models or working with distributed systems
Experience with multimodal datasets and evaluation benchmarks
Familiarity with…
招聘城市:帕罗奥多
…About the Position
We are seeking an exceptional Research Intern to join our team in building the next generation of video world models. While traditional generative models focus on creating passive video (text-to-video), our mission is to build "World Models"—foundation models that understand physics, causality, and dynamics directly from large-scale data, and can be explored and interacted in real-time. You will work at the frontier of generative AI research, enabling the model to "dream" and interact with complex virtual worlds.
Who We Look For
Requirements:
1, Currently pursuing a PhD (or Master’s degree with strong research track record) in Computer Science, Machine Learning, or a related field.
2, Strong proficiency in Python and a deep learning framework (PyTorch or JAX). Experience in large-scale machine learning systems is a great plus.
3, Deep understanding of Generative Models (Diffusion, Transformers, VAEs, Auto-regressive models).
4…
招聘城市:新加坡
…QA, AI, next generation game, technical cooperation, etc.
What the Role Entails
We're looking for a Senior Researcher in multi-modality to help shape the next generation of multimodal foundation models and agentic AI within gaming scenarios. This role focuses on AI research and applications for in-game contexts — abstracting research problems from real game business scenarios, solving them, and deploying solutions that serve a wide range of gaming use cases. You'll work on multimodal understanding, post-training, and agent research, driving breakthroughs that advance both the scientific frontier and Tencent's game products at scale.Pioneer new research directions in multimodal understanding, post-training, reasoning, grounding, and agent planning, grounded in real gaming scenarios.
Abstract research problems from game business scenarios, solve them, and land solutions that serve diverse in-game applications.
Advance the understanding capabilities of multimodal large models (VLMs/MLLMs) across images, video, text, and…
招聘城市:贝尔维尤
…particular focus on innovative breakthroughs in large foundation models. The lab's long-term ambition is to drive the development of Artificial General Intelligence (AGI), and ultimately, Artificial Superintelligence (ASI). We are seeking research interns who are interested in developing novel speech/music/audio/vision/language processing techniques and large multimodal models for our Seattle area office located at Bellevue WA for the year 2026.
Every research intern will work with researchers on a research project aimed at attacking one of the core problems by inventing cutting edge techniques. We encourage discussions and collaborations between researchers and interns. Interns are also encouraged to publish the results from the internship. Our projects span a wide range of areas, including developing more effective multimodal pretraining and post-training strategies for audio, speech, music, image, and video understanding and generation. We aim to enable fully duplex conversations, design more efficient large-model architectures…
招聘城市:贝尔维尤
…particular focus on innovative breakthroughs in large foundation models. The lab's long-term ambition is to drive the development of Artificial General Intelligence (AGI), and ultimately, Artificial Superintelligence (ASI). We are seeking research interns who are interested in developing novel speech/music/audio/vision/language processing techniques and large multimodal models for our Seattle area office located at Bellevue WA for the year 2026.
Every research intern will work with researchers on a research project aimed at attacking one of the core problems by inventing cutting edge techniques. We encourage discussions and collaborations between researchers and interns. Interns are also encouraged to publish the results from the internship. Our projects span a wide range of areas, including developing more effective multimodal pretraining and post-training strategies for audio, speech, music, image, and video understanding and generation. We aim to enable fully duplex conversations, design more efficient large-model architectures…
招聘城市:帕罗奥多
foundational "World Models" that inherently understand physics, causality, action spaces, and complex dynamics directly from internet-scale data. Our goal is to train models that can simulate and "dream" complex virtual worlds, allowing users and agents to explore and interact with them in real time.
This is not a purely theoretical role.
Training interactive world models at this scale requires pushing the limits of modern computing power. We operate at the intersection of cutting-edge generative AI research and high-performance machine learning systems. We are looking for "full-stack" hacker-researchers—visionary thinkers who are also elite engineers, capable of co-designing novel neural architectures and engineering the highly optimized infrastructure required to train them across large-scale computing cluster.What You Will Do
Architect & Scale Foundation Models: Design, train, and scale state-of-the-art interactive world models (combining Diffusion, Autoregressive Transformers, VAEs, LLMs, VLMs) on massive video
招聘城市:帕罗奥多
foundational model algorithm design, optimization related to pre-training, SFT, and RL, model capability evaluation, and exploration of downstream application scenarios.
2. Analyze R&D challenges scientifically, identify performance bottlenecks, and develop solutions based on first principles to accelerate the development and iteration of world models, ensuring competitiveness and leadership.
3. Explore diverse paradigms for world model implementation, research next-generation model architectures, and push the boundaries of world model capabilities.
Who We Look For
1. Bachelor’s degree or higher (preferred) in Computer Science, Artificial Intelligence, Mathematics, or a related field.
2. Solid foundation in deep learning algorithms and proven experience in large model R&D. Candidates with experience in Diffusion Models and Autoregressive Models, publications in top-tier conferences, or practical experience in text-to-image/text-to-video generation are preferred.
3. Familiarity with implementation details of deep learning networks and operators, model tuning for training/inference…
招聘城市:帕罗奥多
foundational model algorithm design, optimization related to pre-training, SFT, and RL, model capability evaluation, and exploration of downstream application scenarios.
2.Analyze R&D challenges scientifically, identify performance bottlenecks, and develop solutions based on first principles to accelerate the development and iteration of world models, ensuring competitiveness and leadership.
3.Explore diverse paradigms for world model implementation, research next-generation model architectures, and push the boundaries of world model capabilities.
Who We Look For
1.Bachelor’s degree or higher (preferred) in Computer Science, Artificial Intelligence, Mathematics, or a related field.
2.Solid foundation in deep learning algorithms and proven experience in large model R&D. Candidates with experience in Diffusion Models and Autoregressive Models, publications in top-tier conferences, or practical experience in text-to-image/text-to-video generation are preferred.
3.Familiarity with implementation details of deep learning networks and operators, model tuning for training/inference…
招聘城市:新加坡
…the next generation of intelligent systems. The intern will engage in cutting-edge R&D involving the integration and understanding of diverse data types such as image, text, audio, and video. You will work alongside top researchers to explore advanced techniques like multimodal pre-training, long-video interaction, and visual reasoning, bridging theoretical innovation with real-world applications. This is a unique opportunity to develop impactful solutions, contribute to open-source or academic outputs, and push the frontier of AI at Tencent Cloud.ResponsibilitiesCore Research & Development: Conduct advanced research on technical solutions for Multimodal Large Models (MLLMs), focusing on the perception, understanding, and interaction of mixed modalities including image, text, audio, and video. Explore cutting-edge topics such as native multimodal pre-training schemes, long-video interactive understanding, visual reasoning, and multimodal agents.
Innovation & Breakthroughs: Keep pace with state-of-the-art (SOTA) trends in the multimodal large model field…
招聘城市:新加坡
…as image generation, multi-modal large models, and few-shot learning.
Based on inhouse products and business needs, improve the performance and experience of AI painting, text generation, and video generation through: Prompt optimization / Generation model R&D / Adapter development / Performance acceleration which also includes resolving algorithm bottlenecks when applying models in real business scenarios.
Address the industrial deployment of multimodal generative models and actively explore model design and optimization in an R&D context
Who We Look For
PhD (preferably fulltime) in Computer Science, Artificial Intelligence, Mathematics, or related fields.
Solid foundation in computer vision or machine learning algorithms; candidates with publications in top conferences or journals are preferred.
Proficient in machine learning and deep learning fundamentals, and familiar with mainstream AIGC frameworks, including GAN, VAE, VQGAN, Diffusion models, etc.
Familiar with generation model extensions such as ControlNet, LoRA, and Text Inversion.
Familiar with multi-modal models like CLIP…
招聘城市:新加坡
…as image generation, multi-modal large models, and few-shot learning.
Based on inhouse products and business needs, improve the performance and experience of AI painting, text generation, and video generation through: Prompt optimization / Generation model R&D / Adapter development / Performance acceleration which also includes resolving algorithm bottlenecks when applying models in real business scenarios.
Address the industrial deployment of multimodal generative models and actively explore model design and optimization in an R&D context
Who We Look For
PhD (preferably fulltime) in Computer Science, Artificial Intelligence, Mathematics, or related fields.
Solid foundation in computer vision or machine learning algorithms; candidates with publications in top conferences or journals are preferred.
Proficient in machine learning and deep learning fundamentals, and familiar with mainstream AIGC frameworks, including GAN, VAE, VQGAN, Diffusion models, etc.
Familiar with generation model extensions such as ControlNet, LoRA, and Text Inversion.
Familiar with multi-modal models like CLIP…
招聘城市:新加坡
…as image generation, multi-modal large models, and few-shot learning.
Based on inhouse products and business needs, improve the performance and experience of AI painting, text generation, and video generation through: Prompt optimization / Generation model R&D / Adapter development / Performance acceleration which also includes resolving algorithm bottlenecks when applying models in real business scenarios.
Address the industrial deployment of multimodal generative models and actively explore model design and optimization in an R&D context
Who We Look For
PhD (preferably fulltime) in Computer Science, Artificial Intelligence, Mathematics, or related fields.
Solid foundation in computer vision or machine learning algorithms; candidates with publications in top conferences or journals are preferred.
Proficient in machine learning and deep learning fundamentals, and familiar with mainstream AIGC frameworks, including GAN, VAE, VQGAN, Diffusion models, etc.
Familiar with generation model extensions such as ControlNet, LoRA, and Text Inversion.
Familiar with multi-modal models like CLIP…
招聘城市:新加坡
…as image generation, multi-modal large models, and few-shot learning.
Based on inhouse products and business needs, improve the performance and experience of AI painting, text generation, and video generation through: Prompt optimization / Generation model R&D / Adapter development / Performance acceleration which also includes resolving algorithm bottlenecks when applying models in real business scenarios.
Address the industrial deployment of multimodal generative models and actively explore model design and optimization in an R&D context
Who We Look For
PhD (preferably fulltime) in Computer Science, Artificial Intelligence, Mathematics, or related fields.
Solid foundation in computer vision or machine learning algorithms; candidates with publications in top conferences or journals are preferred.
Proficient in machine learning and deep learning fundamentals, and familiar with mainstream AIGC frameworks, including GAN, VAE, VQGAN, Diffusion models, etc.
Familiar with generation model extensions such as ControlNet, LoRA, and Text Inversion.
Familiar with multi-modal models like CLIP…
招聘城市:新加坡
…research and development of large-scale video world models, including the design and construction of training datasets, foundational model algorithm design, optimization related to pre-training, SFT, and RL, model capability evaluation, and exploration of downstream application scenarios.
Analyze R&D challenges scientifically, identify performance bottlenecks, and develop solutions based on first principles to accelerate the development and iteration of world models, ensuring competitiveness and leadership.
Explore diverse paradigms for world model implementation, research next-generation model architectures, and push the boundaries of world model capabilities.
Who We Look For
Bachelor’s degree or higher (preferred) in Computer Science, Artificial Intelligence, Mathematics, or a related field.
Solid foundation in deep learning algorithms and proven experience in large model R&D. Candidates with experience in Diffusion Models and Autoregressive Models, publications in top-tier conferences, or practical experience in text-to-image/text-to-video generation are preferred.
Familiarity with implementation…
招聘城市:新加坡
…as image generation, multi-modal large models, and few-shot learning.
Based on inhouse products and business needs, improve the performance and experience of AI painting, text generation, and video generation through: Prompt optimization / Generation model R&D / Adapter development / Performance acceleration which also includes resolving algorithm bottlenecks when applying models in real business scenarios.
Address the industrial deployment of multimodal generative models and actively explore model design and optimization in an R&D context
Who We Look For
PhD (preferably fulltime) in Computer Science, Artificial Intelligence, Mathematics, or related fields.
Solid foundation in computer vision or machine learning algorithms; candidates with publications in top conferences or journals are preferred.
Proficient in machine learning and deep learning fundamentals, and familiar with mainstream AIGC frameworks, including GAN, VAE, VQGAN, Diffusion models, etc.
Familiar with generation model extensions such as ControlNet, LoRA, and Text Inversion.
Familiar with multi-modal models like CLIP…
招聘城市:新加坡
…research and development of large-scale video world models, including the design and construction of training datasets, foundational model algorithm design, optimization related to pre-training, SFT, and RL, model capability evaluation, and exploration of downstream application scenarios.
Analyze R&D challenges scientifically, identify performance bottlenecks, and develop solutions based on first principles to accelerate the development and iteration of world models, ensuring competitiveness and leadership.
Explore diverse paradigms for world model implementation, research next-generation model architectures, and push the boundaries of world model capabilities.
Who We Look For
Bachelor’s degree or higher (preferred) in Computer Science, Artificial Intelligence, Mathematics, or a related field.
Solid foundation in deep learning algorithms and proven experience in large model R&D. Candidates with experience in Diffusion Models and Autoregressive Models, publications in top-tier conferences, or practical experience in text-to-image/text-to-video generation are preferred.
Familiarity with implementation…
招聘城市:新加坡
…training datasets, foundational model algorithm design, optimization related to pre-training, SFT, and RL, model capability evaluation, and exploration of downstream application scenarios.
Analyze R&D challenges scientifically, identify performance bottlenecks, and develop solutions based on first principles to accelerate the development and iteration of world models, ensuring competitiveness and leadership.
Explore diverse paradigms for world model implementation, research next-generation model architectures, and push the boundaries of world model capabilities.
Who We Look For
Bachelor’s degree or higher (preferred) in Computer Science, Artificial Intelligence, Mathematics, or a related field.
Solid foundation in deep learning algorithms and proven experience in large model R&D. Candidates with experience in Diffusion Models and Autoregressive Models, publications in top-tier conferences, or practical experience in text-to-image/text-to-video generation are preferred.
3.Familiarity with implementation details of deep learning networks and operators, model tuning for training/inference, CPU/GPU…
招聘城市:新加坡
…as image generation, multi-modal large models, and few-shot learning.
Based on inhouse products and business needs, improve the performance and experience of AI painting, text generation, and video generation through: Prompt optimization / Generation model R&D / Adapter development / Performance acceleration which also includes resolving algorithm bottlenecks when applying models in real business scenarios.
Address the industrial deployment of multimodal generative models and actively explore model design and optimization in an R&D context
Who We Look For
PhD (preferably fulltime) in Computer Science, Artificial Intelligence, Mathematics, or related fields.
Solid foundation in computer vision or machine learning algorithms; candidates with publications in top conferences or journals are preferred.
Proficient in machine learning and deep learning fundamentals, and familiar with mainstream AIGC frameworks, including GAN, VAE, VQGAN, Diffusion models, etc.
Familiar with generation model extensions such as ControlNet, LoRA, and Text Inversion.
Familiar with multi-modal models like CLIP…