Welcome to the ECCV 2026 Workshop on Multimodal Large Language Models for Unified Comprehension and Generation. This workshop aims to consolidate emerging research on unified multimodal intelligence, with a focus on systems that understand, generate, and act across modalities within a coherent framework.
Recent multimodal AI has evolved from vision-language understanding toward broader multimodal foundation models spanning image, video, audio, 3D, and generation. At the same time, the community is moving from modular pipelines toward unified tokenization, hybrid autoregressive-diffusion designs, shared representations, and synergistic learning between understanding and generation.
Our goal is to bring together researchers from academia and industry to discuss architectural designs, tokenization strategies, training objectives, evaluation protocols, and practical challenges for building general-purpose multimodal systems.
Topics and Themes
We welcome technical, position, and perspective papers related to unified multimodal modeling. Topics of interest include, but are not limited to:
Unified Multimodal Understanding
Captioning, VQA, retrieval, grounding, segmentation, reasoning, visual document understanding, long-video understanding, and cross-modal knowledge extraction.
Unified Multimodal Content Generation
Text-to-image/video, image-to-image, controllable generation, sequential image generation, visual editing, and cross-modal generative modeling.
Unified MLLM Understanding and Generation
Architectures and objectives that jointly model comprehension and generation, including autoregressive, diffusion, flow-based, and hybrid paradigms.
Synergistic Learning
How understanding and generation, or different modalities and tasks, can mutually enhance each other during pre-training, instruction tuning, and post-training.
Benchmarking and Evaluation
Evaluation protocols and benchmarks for unified multimodal systems, including realism, controllability, generalization, reasoning, and fair comparison.
Broader Directions
Reinforcement learning for unified modeling, multimodal chain-of-thought, joint vision-language/audio/3D models, cross-task transfer, and efficient training.
Submission Instructions
The workshop accepts two submission tracks to encourage broader participation:
Regular archival papers may be up to 14 pages long, including figures and tables, and should use Springer LNCS formatting. Additional pages containing only cited references are permitted. Submissions must follow the ECCV 2026 template and be uploaded through OpenReview. The review process is double-blind and managed by the workshop organizers and program committee. Conflicts of interest will be handled according to the ECCV 2026 Submission Policy.
Non-archival submissions are intended for work that is already published or that authors prefer not to include in the proceedings. Eligible papers include those already peer-reviewed at major CV/ML conferences or journals. Previously reviewed or published papers will not be re-reviewed; acceptance is based on topical fit and poster-board availability. Unpublished submissions in this track will undergo double-blind review, following the same review process as regular archival submissions. Authors should submit the paper or a link to the paper via the workshop's OpenReview page.
Important Dates (AoE)
Schedule
The workshop will be held in a hybrid format, supporting both onsite and online participation. The program consists of invited keynote talks, oral presentations, and poster presentations of accepted papers.
| Time | Schedule | Speaker / Presentation |
|---|---|---|
| 08:50 – 09:00 | Introduction and Opening Remarks | Organizers |
| 09:00 – 09:40 | Keynote Talk 1 |
Mike Z. Shou
Show-o Series: One Single Transformer to Unify Multimodal Understanding and Generation
View abstractHide abstractExciting models have been developed in multimodal video understanding and generation, such as video LLM and video diffusion model. One emerging pathway to the ultimate intelligence is to create one single foundation model that can do both understanding and generation. In this work, we present Show-o, one single transformer that handles both multimodal understanding and generation. Unlike fully autoregressive models, Show-o is the first to unify autoregressive and discrete diffusion modeling, flexibly supporting a wide range of vision-language tasks including visual question-answering, text-to-image generation, text-guided inpainting/extrapolation, and mixed-modality generation of any input/output format, all within one single transformer. Across various benchmarks, Show-o demonstrates comparable or superior performance, shedding light for building the next-generation foundation model. At the end, I will share our latest edition, Show-o2, which further integrates video understanding and generation into this single transformer. |
| 09:40 – 10:20 | Keynote Talk 2 |
Yuren Cong
Unified Representation for Unified Models
View abstractHide abstractUnified multimodal models aim to integrate visual understanding and generation within a single model, yet their visual representations often remain fragmented by task-specific encoders and inherited representation biases. In this talk, I present our exploration of unified visual representations, from TUNA to TUNA-2. We begin with TUNA, which unifies understanding and generation in a shared continuous visual space by cascading a reconstruction-oriented VAE with a semantic representation encoder. This design eliminates the representation mismatch of decoupled approaches. Building on this, we ask a more fundamental question: if a unified model is trained end-to-end, are pretrained vision encoders necessary at all? TUNA-2 progressively removes both the VAE and representation encoder, learning directly from raw pixels through simple patch embeddings. To address the increased difficulty and redundancy of pixel-space learning, we introduce masking-based feature learning to encourage more robust representations. At scale, this encoder-free design remains competitive in generation while yielding stronger fine-grained visual understanding. Together, these works suggest a path toward truly unified multimodal models: moving from unifying pretrained representations to learning unified representations natively from shared multimodal objectives. |
| 10:20 – 10:50 | Coffee Break | ˙✧˖°☕ ༘ ⋆。˚ |
| 10:50 – 11:30 | Keynote Talk 3 |
Rogerio Schmidt Feris
Granite Multimodal Models: from Pixels to Prosody
View abstractHide abstractI will begin this talk by introducing Granite Vision, our vision-language model that specializes in visual document understanding and enables efficient content extraction from tables, charts, infographics, plots, diagrams, sketches, and more. In particular, I will discuss the notion of code-guided data generation for multimodal comprehension. Next, I will expand to speech and audio, discussing the challenges of building a unified model for simultaneous comprehension of multimodal video and generation of prosodic speech for sports AI commentary. Our work was deployed at the US Open and Wimbledon tennis tournaments to provide commentary for matches that would otherwise lack coverage. I will conclude by discussing our current research directions and future opportunities in this space. |
| 11:30 – 12:00 |
Oral Presentations
|
|
| 12:00 – 13:40 | Lunch Break | 🥤🥗🍔🍗🍟🥓 |
| 13:40 – 14:20 | Keynote Talk 4 |
Chen Change Loy
A Camera-Centric Framework for Unified Multimodal Understanding and Generation
View abstractHide abstractUnified multimodal models promise to understand and generate in a single architecture, but most of today’s designs treat the camera in oversimplified ways, assuming fixed viewpoints or predefined fields of view. This limits their ability to handle real-world scenarios where perspectives shift and contexts are dynamic. In this talk, I will introduce Puffin, our new multimodal framework that brings the camera dimension into the picture. Puffin combines autoregressive and diffusion modeling to interpret and generate scenes from arbitrary viewpoints with continuous fields of view. A key idea behind Puffin is to treat the camera as language, allowing the model to “think with camera.” This aligns visual cues with photographic terms and grounds reasoning in physical context, making the model more spatially aware. Puffin is trained on a large dataset of four million vision–language–camera triplets, with both global camera parameters and pixel-level camera maps. This enables precise and flexible control over spatial generation. I’ll show how Puffin outperforms specialized baselines in controllable generation and camera understanding, and how, with instruction tuning, it generalizes to diverse tasks like spatial imagination, world exploration, and photography guidance. |
| 14:20 – 14:50 |
Oral Presentations
|
|
| 14:50 – 15:00 | Transition / Poster Setup | 📰 |
| 15:00 – 16:00 | Poster Session + Coffee Break | Accepted authors (MalmöMässan Exhibit Hall, No. 394–423) |
| 16:00 – 16:40 | Keynote Talk 5 |
Jaemin Cho
Unify the Interface, Not Just the Parameters: A Shared Symbolic Canvas for Human–AI Visual Co-Creation
View abstractHide abstractVisual creation is a long-horizon, communicative process: we plan, sketch, get feedback, and revise, over many rounds and often with many collaborators. Today’s generative models keep getting stronger, yet iterating with them still hurts: text cannot name a specific element, untouched content drifts across rounds, and every edit re-renders everything. What is missing is a shared, unified communication medium—a grounded symbolic canvas that humans, understanding models, and generation models can all read and write, the way code, diffs, and tests already serve coding agents. I will trace my work building pieces of this interface, from LLM-based layout planning to localized critique and revision, and close with a vision for models that learn the visual creation process itself. |
| 16:40 – 17:20 | Keynote Talk 6 |
Mohit Bansal
Multimodal Unification, Communication, and Composable Generalization
|
| 17:20 – 17:30 | Closing Remarks + Best Paper Award | Organizers |
Invited Speakers
Organizing Committee
Accepted Papers
Congratulations to all accepted authors. Select a paper title to view its OpenReview page.
- Combining Finetuning and RAG for Structured Output Prediction on a Multimodal High Velocity Retail Dataset
- FCMBench-Video: Benchmarking Document Video Intelligence
- From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
- Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs
- Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing
- SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
- Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
- A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval
- FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
- X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis
- Spatially-Grounded Text-to-Video Generation via Inference-Time Gradient-Free Optimization
- Scene-Graph Guided Conditioning for Text-to-Image Generation
- Representation Forcing for Bottleneck-Free Unified Multimodal Models
- SchemixQA and CoRe-VLM: A Benchmark and Collaborative Refinement Framework for Technical Schematic VQA
- Test-Time Hallucination Control in Large Vision-Language Models
- Weakly Supervised Motion Learning for Co-speech Gesture Video Generation
- The Plan, Not the Decoder: Diagnosing and Repairing Compositional Failure in Reasoning-Augmented Text-to-Image Generation
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models
- ComfyClaw: Self-Evolving Skill Harnesses for Image Generation Workflows
- Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
- CHIMERA: Adaptive Cache Injection and Semantic Anchor Prompting for Zero-shot Image Morphing with Morphing-oriented Metrics
- Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation
- DoCoG: Mask-based Multi-Type Grounded Chain-of-Thought for Document QA
- From Visual Widgets to UI Code: Efficient Tool-Grounded Generation
- SeoUL: Semantic-Centric Alignment for Universal Cross-Domain Retrieval with Multimodal Large Language Models
- MathGen: Revealing the Illusion of Mathematical Competence through Text-to-Image Generation
- Revisiting Classifier-Free Guidance Methods in Latent Diffusion Models
- DuET: Dual Expert Trajectories for Diffusion Image Editing
- 3Dify-Anything: A Unified Model for Multimodal 3D Generation via Context Pre-Training
- Query-Adaptive Diversity for Long-Form Video Question Answering
- In-Context Audio Control of Video Diffusion Transformers
Diversity and Inclusion
We are committed to promoting diversity and inclusion across the organizing committee, invited speakers, program committee, accepted papers, and audience. We will proactively encourage submissions from underrepresented groups and institutions, provide inclusive wording in the call for papers, support mentoring opportunities during poster sessions and panels, and ensure virtual access for remote or resource-constrained participants.
Contact
Questions? Please contact the workshop organizers (shengqiongwu@gmail.com).
















