Arrow Research search
Back to AAAI

AAAI 2025

Open-World Multimodal Understanding and Generation with Efficiently Finetuned Foundation Models

Conference Paper New Faculty Highlights Artificial Intelligence

Abstract

With the astonishing ability of different pretrained foundation models (e.g., large language models (LLMs), vision-language models, diffusion models), today’s AI research and development tendency has been revolutionized. In this talk, I will answer two questions: Q1: How can we efficiently train or fine-tune foundation models? Q2: How can we build strong open-world multimodal understanding and generation models with these pretrained foundation models?

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
AAAI Conference on Artificial Intelligence
Archive span
1980-2026
Indexed papers
28718
Paper id
221209385351655495
v2026.09.13