2nd MUCG @ ECCV 2026

Multimodal Large Language Models for Unified Comprehension and Generation

🗓️ September 8, 2026 📍 Malmömässan A, Malmö, Sweden
Learn More

Welcome to the ECCV 2026 Workshop on Multimodal Large Language Models for Unified Comprehension and Generation. This workshop aims to consolidate emerging research on unified multimodal intelligence, with a focus on systems that understand, generate, and act across modalities within a coherent framework.

Recent multimodal AI has evolved from vision-language understanding toward broader multimodal foundation models spanning image, video, audio, 3D, and generation. At the same time, the community is moving from modular pipelines toward unified tokenization, hybrid autoregressive-diffusion designs, shared representations, and synergistic learning between understanding and generation.

Our goal is to bring together researchers from academia and industry to discuss architectural designs, tokenization strategies, training objectives, evaluation protocols, and practical challenges for building general-purpose multimodal systems.

Topics and Themes

We welcome technical, position, and perspective papers related to unified multimodal modeling. Topics of interest include, but are not limited to:

01

Unified Multimodal Understanding

Captioning, VQA, retrieval, grounding, segmentation, reasoning, visual document understanding, long-video understanding, and cross-modal knowledge extraction.

02

Unified Multimodal Content Generation

Text-to-image/video, image-to-image, controllable generation, sequential image generation, visual editing, and cross-modal generative modeling.

03

Unified MLLM Understanding and Generation

Architectures and objectives that jointly model comprehension and generation, including autoregressive, diffusion, flow-based, and hybrid paradigms.

04

Synergistic Learning

How understanding and generation, or different modalities and tasks, can mutually enhance each other during pre-training, instruction tuning, and post-training.

05

Benchmarking and Evaluation

Evaluation protocols and benchmarks for unified multimodal systems, including realism, controllability, generalization, reasoning, and fair comparison.

06

Broader Directions

Reinforcement learning for unified modeling, multimodal chain-of-thought, joint vision-language/audio/3D models, cross-task transfer, and efficient training.

Submission Instructions

The workshop accepts two submission tracks to encourage broader participation:

Regular Archival Submissions
Up to 14 pages

Regular archival papers may be up to 14 pages long, including figures and tables, and should use Springer LNCS formatting. Additional pages containing only cited references are permitted. Submissions must follow the ECCV 2026 template and be uploaded through OpenReview. The review process is double-blind and managed by the workshop organizers and program committee. Conflicts of interest will be handled according to the ECCV 2026 Submission Policy.

Non-Archival Submissions
At least 4 pages

Non-archival submissions are intended for work that is already published or that authors prefer not to include in the proceedings. Eligible papers include those already peer-reviewed at major CV/ML conferences or journals. Previously reviewed or published papers will not be re-reviewed; acceptance is based on topical fit and poster-board availability. Unpublished submissions in this track will undergo double-blind review, following the same review process as regular archival submissions. Authors should submit the paper or a link to the paper via the workshop's OpenReview page.

All accepted papers will be presented as posters.
Best Paper Awards will be selected based on reviewer scores and committee evaluation.
Submissions are handled through the official ECCV 2026 MUCG OpenReview site.
View Important Dates View Schedule

Important Dates (AoE)

Archival submission deadlineJul 01 Jul 05
Archival notificationJul 18 Jul 20 Jul 22
Non-archival submission deadlineJul 25
Non-archival notificationAug 7
Registration deadlineAug 10
Camera-ready deadline for all accepted papersAug 15
Workshop (during ECCV 2026)Sep 8

Schedule

Date September 8, 2026
Venue Malmömässan A, Malmö, Sweden

The workshop will be held in a hybrid format, supporting both onsite and online participation. The program consists of invited keynote talks, oral presentations, and poster presentations of accepted papers.

Time Schedule Speaker / Presentation
08:50 – 09:00 Introduction and Opening Remarks Organizers
09:00 – 09:40 Keynote Talk 1
Mike Z. Shou
Show-o Series: One Single Transformer to Unify Multimodal Understanding and Generation
View abstractHide abstract

Exciting models have been developed in multimodal video understanding and generation, such as video LLM and video diffusion model. One emerging pathway to the ultimate intelligence is to create one single foundation model that can do both understanding and generation. In this work, we present Show-o, one single transformer that handles both multimodal understanding and generation. Unlike fully autoregressive models, Show-o is the first to unify autoregressive and discrete diffusion modeling, flexibly supporting a wide range of vision-language tasks including visual question-answering, text-to-image generation, text-guided inpainting/extrapolation, and mixed-modality generation of any input/output format, all within one single transformer. Across various benchmarks, Show-o demonstrates comparable or superior performance, shedding light for building the next-generation foundation model. At the end, I will share our latest edition, Show-o2, which further integrates video understanding and generation into this single transformer.

09:40 – 10:20 Keynote Talk 2
Yuren Cong
Unified Representation for Unified Models
View abstractHide abstract

Unified multimodal models aim to integrate visual understanding and generation within a single model, yet their visual representations often remain fragmented by task-specific encoders and inherited representation biases. In this talk, I present our exploration of unified visual representations, from TUNA to TUNA-2. We begin with TUNA, which unifies understanding and generation in a shared continuous visual space by cascading a reconstruction-oriented VAE with a semantic representation encoder. This design eliminates the representation mismatch of decoupled approaches. Building on this, we ask a more fundamental question: if a unified model is trained end-to-end, are pretrained vision encoders necessary at all? TUNA-2 progressively removes both the VAE and representation encoder, learning directly from raw pixels through simple patch embeddings. To address the increased difficulty and redundancy of pixel-space learning, we introduce masking-based feature learning to encourage more robust representations. At scale, this encoder-free design remains competitive in generation while yielding stronger fine-grained visual understanding. Together, these works suggest a path toward truly unified multimodal models: moving from unifying pretrained representations to learning unified representations natively from shared multimodal objectives.

10:20 – 10:50 Coffee Break ˙✧˖°☕ ༘ ⋆。˚
10:50 – 11:30 Keynote Talk 3
Rogerio Schmidt Feris
Granite Multimodal Models: from Pixels to Prosody
View abstractHide abstract

I will begin this talk by introducing Granite Vision, our vision-language model that specializes in visual document understanding and enables efficient content extraction from tables, charts, infographics, plots, diagrams, sketches, and more. In particular, I will discuss the notion of code-guided data generation for multimodal comprehension. Next, I will expand to speech and audio, discussing the challenges of building a unified model for simultaneous comprehension of multimodal video and generation of prosodic speech for sports AI commentary. Our work was deployed at the US Open and Wimbledon tennis tournaments to provide commentary for matches that would otherwise lack coverage. I will conclude by discussing our current research directions and future opportunities in this space.

11:30 – 12:00
Oral Presentations
Session 1 · 2 papers
  1. Paper #22FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
  2. Paper #35Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
12:00 – 13:40 Lunch Break 🥤🥗🍔🍗🍟🥓
13:40 – 14:20 Keynote Talk 4
Chen Change Loy
A Camera-Centric Framework for Unified Multimodal Understanding and Generation
View abstractHide abstract

Unified multimodal models promise to understand and generate in a single architecture, but most of today’s designs treat the camera in oversimplified ways, assuming fixed viewpoints or predefined fields of view. This limits their ability to handle real-world scenarios where perspectives shift and contexts are dynamic. In this talk, I will introduce Puffin, our new multimodal framework that brings the camera dimension into the picture. Puffin combines autoregressive and diffusion modeling to interpret and generate scenes from arbitrary viewpoints with continuous fields of view. A key idea behind Puffin is to treat the camera as language, allowing the model to “think with camera.” This aligns visual cues with photographic terms and grounds reasoning in physical context, making the model more spatially aware. Puffin is trained on a large dataset of four million vision–language–camera triplets, with both global camera parameters and pixel-level camera maps. This enables precise and flexible control over spatial generation. I’ll show how Puffin outperforms specialized baselines in controllable generation and camera understanding, and how, with instruction tuning, it generalizes to diverse tasks like spatial imagination, world exploration, and photography guidance.

14:20 – 14:50
Oral Presentations
Session 2 · 2 papers
  1. Paper #26Representation Forcing for Bottleneck-Free Unified Multimodal Models
  2. Paper #41Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation
14:50 – 15:00 Transition / Poster Setup 📰
15:00 – 16:00 Poster Session + Coffee Break Accepted authors (MalmöMässan Exhibit Hall, No. 394–423)
16:00 – 16:40 Keynote Talk 5
Jaemin Cho
Unify the Interface, Not Just the Parameters: A Shared Symbolic Canvas for Human–AI Visual Co-Creation
View abstractHide abstract

Visual creation is a long-horizon, communicative process: we plan, sketch, get feedback, and revise, over many rounds and often with many collaborators. Today’s generative models keep getting stronger, yet iterating with them still hurts: text cannot name a specific element, untouched content drifts across rounds, and every edit re-renders everything. What is missing is a shared, unified communication medium—a grounded symbolic canvas that humans, understanding models, and generation models can all read and write, the way code, diffs, and tests already serve coding agents. I will trace my work building pieces of this interface, from LLM-based layout planning to localized critique and revision, and close with a vision for models that learn the visual creation process itself.

16:40 – 17:20 Keynote Talk 6
Mohit Bansal
Multimodal Unification, Communication, and Composable Generalization
17:20 – 17:30 Closing Remarks + Best Paper Award Organizers

Invited Speakers

Mike Z. Shou
Mike Z. ShouNational University of Singapore
Mohit Bansal
Mohit BansalUniversity of North Carolina Chapel Hill
Chen Change Loy
Chen Change LoyNanyang Technological University
Rogerio Schmidt Feris
Rogerio Schmidt FerisMIT-IBM Watson AI Lab
Jaemin Cho
Jaemin ChoAllen Institute for AI / Johns Hopkins University
Yuren Cong
Yuren CongMeta

Organizing Committee

Shengqiong Wu
Shengqiong WuOrganizerUniversity of Oxford
Jinheng Xie
Jinheng XieOrganizerNational University of Singapore
Haozhe Liu
Haozhe LiuOrganizerNVIDIA Research
Tian Ye
Tian YeOrganizerHKUST(GZ) & NVIDIA Research
Yanguang Zhao
Yanguang ZhaoExecutorNational University of Singapore
Enze Xie
Enze XieOrganizerNVIDIA Research / MIT HAN Lab
Sivan Doveh
Sivan DovehOrganizerStanford University
Anna Kukleva
Anna KuklevaOrganizerMeta & Max Planck Institute for Informatics
Jehanzeb Mirza
Jehanzeb MirzaOrganizerXero
Mingchen Zhuge
Mingchen ZhugeOrganizerKAUST
Hao Fei
Hao FeiOrganizerUniversity of Oxford

Accepted Papers

32 papers

Congratulations to all accepted authors. Select a paper title to view its OpenReview page.

  1. Combining Finetuning and RAG for Structured Output Prediction on a Multimodal High Velocity Retail Dataset
  2. FCMBench-Video: Benchmarking Document Video Intelligence
  3. From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
  4. Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs
  5. Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing
  6. SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
  7. Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
  8. A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
  9. FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
  10. X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis
  11. Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization
  12. Scene-Graph Guided Conditioning for Text-to-Image Generation
  13. Representation Forcing for Bottleneck-Free Unified Multimodal Models
  14. SchemixQA and CoRe-VLM: A Benchmark and Collaborative Refinement Framework for Technical Schematic VQA
  15. Test-Time Hallucination Control in Large Vision-Language Models
  16. Weakly Supervised Motion Learning for Co-speech Gesture Video Generation
  17. The Plan, Not the Decoder: Diagnosing and Repairing Compositional Failure in Reasoning-Augmented Text-to-Image Generation
  18. Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
  19. AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models
  20. ComfyClaw: Self-Evolving Skill Harnesses for Image Generation Workflows
  21. Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
  22. CHIMERA: Adaptive Cache Injection and Semantic Anchor Prompting for Zero-shot Image Morphing with Morphing-oriented Metrics
  23. Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation
  24. DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA
  25. From Visual Widgets to UI Code: Efficient Tool-Grounded Generation
  26. SeoUL: Semantic-Centric Alignment for Universal Cross-Domain Retrieval with Multimodal Large Language Models
  27. MathGen: Revealing the Illusion of Mathematical Competence through Text-to-Image Generation
  28. Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models
  29. DuET: Dual Expert Trajectories for Diffusion Image Editing
  30. 3Dify-Anything: A Unified Model for Multimodal 3D Generation via Context Pre-Training
  31. Query-Adaptive Diversity for Long-Form Video Question Answering
  32. In-Context Audio Control of Video Diffusion Transformers

Diversity and Inclusion

We are committed to promoting diversity and inclusion across the organizing committee, invited speakers, program committee, accepted papers, and audience. We will proactively encourage submissions from underrepresented groups and institutions, provide inclusive wording in the call for papers, support mentoring opportunities during poster sessions and panels, and ensure virtual access for remote or resource-constrained participants.

Contact

Questions? Please contact the workshop organizers (shengqiongwu@gmail.com).