Student Researcher - (Seed Model - Multimodal Interaction & World Model) β 2026 Start (PhD)
- πΊπΈ United States
- On-site
- Internship
- 2 months ago
- AI
Not enough detail in this posting to match
The Seed Multimodal Interaction and World Model team is dedicated to developing models that boast human-level multimodal understanding and interaction capabilities. The team also aspires to advance the exploration and development of multimodal assistant products.
We are looking for talented individuals to join us for an internship. PhD internships at Our Company provide students with the opportunity to actively contribute to our products and research, as well as to the organization's future plans and emerging technologies.
Our dynamic internship experience blends hands-on learning, enriching community-building and professional development events, and collaboration with industry experts.
Applications will be reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume (Start date, End date).
Responsibilities
- Design and implement reinforcement learning (RL) training systems for large-scale multimodal foundation models
- Develop unified modeling frameworks that integrate video, audio, and language, with a focus on visual latent reasoning
- Explore RL-based approaches to bridge understanding and generation for multimodal visual reasoning
- Collaborate with researchers to evaluate models on tasks involving world modeling, reasoning, and instruction-conditioned generation
Minimum Qualifications
- Currently pursuing a PhD in Software Development, Computer Science, Computer Engineering, or a related technical discipline
- Publications in accredited venues, such as CVPR, ECCV, ICCV, NeurIPS, ICLR, ICML, or other leading conferences in AI and ML
- Strong research background in at least one of the following: reinforcement learning, multimodal learning, video understanding, or vision-language modeling
- Must obtain work authorization in the country of employment at the time of hire, and maintain ongoing work authorization during employment
Preferred Qualifications
- Experience with reinforcement learning in multimodal or interactive environments
- Familiarity with video generation or diffusion-based generative models
- Experience with large-scale model training (e.g., distributed training, curriculum learning, or memory-augmented transformers)
- Solid programming and engineering skills, with experience building training or evaluation pipelines for ML models
As a condition of employment, all successful candidates must be able to establish authorization to work in the United States. For this position, the Company does not provide sponsorship or any immigration-related benefits.
Student Researcher - (Seed Model - Multimodal Interaction & World Model) β 2026 Start (PhD) Β· ByteDance