Aerial view of Malmo waterfront, Turning Torso, and Oresund Bridge

ECCV 2026 · Malmö, Sweden · Sep 8-12

Efficient and trustworthy AI methods for image generation and multimodal understanding.

Three accepted papers and one tutorial on efficient, trustworthy image generation and multimodal understanding.

Explore

The ECCV 2026 Crew

Meet our encoders

Five ENCODE Lab researchers contributing papers, presentations, and the tutorial.

Discover ENCODE Lab

Accepted Papers

Three ECCV 2026 papers from ENCODE Lab.

LISA qualitative image-generation results and speedup comparison from Figure 1 of the paper Paper Figure 1
01 ECCV 2026

LISA: Locality-Informed Speculative Decoding for Accelerating Autoregressive Image Generation

Speculative decoding for autoregressive image generation

Ying Li, Siyong Jian, Zhaode Wang, Zhiwen Chen, Chengfei Lv, and Huan Wang

EVAR edge deployment, generation samples, and pruning comparison from Figure 1 of the paper Paper Figure 1
02 ECCV 2026

EVAR: Edge Visual Autoregressive Models via Principled Pruning

Principled pruning for edge visual autoregressive models

Zefang Wang, Ying Li, Yanyu Li, Mingluo Su, Simin Xu, Guanzhong Tian, and Huan Wang

ECCV 2026 Tutorial

Efficient MLLM Inference via Approximate and Exact Computing.

Multimodal large language models are powerful but expensive to serve. This tutorial presents a systematic view of efficient inference through two complementary lenses: approximate computing that removes model and data redundancy while preserving practical utility, and exact computing that improves systems and hardware execution without changing model outputs.

Date
Sep 8, 2026
Time
9AM - 12PM
Location
Malmömässan, K3 Room
Full Tutorial Page

Three main topics

01

Model Compression

Quantization, sparsity, and low-rank decomposition for reducing model redundancy in MLLM inference.

02

Token Efficiency

Training and training-free compression methods for long multimodal inputs and autoregressive generation.

03

System-Level Design

Exact computing methods and inference frameworks that improve throughput and latency without changing the computed result.

Morning program

MLLM Inference Efficiency.

Three technical sessions, followed by questions and closing remarks.

  1. Opening RemarksHuan Wang
  2. Approximate Computing I: Model CompressionSathya Narayanan Ravi
  3. Approximate Computing II: Token EfficiencyBo Li
  4. Coffee Break and DiscussionsInformal discussion
  5. Exact Computing: System-Level OptimizationChenyang Zhao
  6. Q&A and Closing RemarksHuan Wang