Buch, Englisch, 698 Seiten, Format (B × H): 155 mm x 235 mm
19th European Conference, Malmö, Sweden, September 8–12, 2026, Proceedings, Part LXV
Buch, Englisch, 698 Seiten, Format (B × H): 155 mm x 235 mm
Reihe: Lecture Notes in Computer Science
ISBN: 978-3-032-37188-1
Verlag: Springer
The multi-volume set of LNCS books with volume numbers 17001 up to 17083 constitutes the refereed proceedings of the 19th European Conference on Computer Vision, ECCV 2026, held in Malmö, Sweden, during September 8–12, 2026.
The 2866 papers presented in these proceedings were carefully reviewed and selected from a total of 10,473 submissions. They deal with topics such as computer vision; machine learning; deep neural networks; reinforcement learning; object recognition; image classification; image processing; object detection; semantic segmentation; human pose estimation; 3d reconstruction; stereo vision; computational photography; neural networks; image coding; image reconstruction; motion estimation.
Zielgruppe
Research
Autoren/Hrsg.
Fachgebiete
- Technische Wissenschaften Elektronik | Nachrichtentechnik Nachrichten- und Kommunikationstechnik Signalverarbeitung
- Mathematik | Informatik EDV | Informatik Informatik Mensch-Maschine-Interaktion
- Mathematik | Informatik EDV | Informatik Informatik Künstliche Intelligenz Maschinelles Lernen
- Mathematik | Informatik EDV | Informatik Informatik Bildsignalverarbeitung
- Technische Wissenschaften Elektronik | Nachrichtentechnik Elektronik
Weitere Infos & Material
Rethinking Attention Reallocation for Multimodal Emotion Recognition.- Ego-Human Motion Prediction with 3D-Aware LLM.- COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models.- GAP-Track: Bridging the Resolution Gap for Cross-Resolution RGBT Tracking.- Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation.- CARE: Causally-Aligned Reasoning Exploration for Medical Large Language Models.- WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion.- Small Vision-Language Models are Smart Compressors for Long Video Understanding.- WaterGen: Decoupling Scene and Medium in Underwater Image Generation.- TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning.- Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability.- NaVLM-PVC: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs.- Under One Sun: Multi-Object Generative Perception of Materials and Illumination.- VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment.- VLA-Hijack: A Transferable Patch Attack against Vision-Language-Action Models via Visual Proprioception Hijacking.- Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs.- Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker.- H-SFP: Hierarchical Federated Learning with Decoupled Split-Model Prototyping.- Training-Free Refinement of Flow Matching with Divergence-based Sampling.- Reward Modeling for Computer-Using Agent from Video Execution.- Ctrl-Z Sampling: Scaling Diffusion Sampling with Controlled Random Zigzag Explorations.- Geometry-Aware Spatio-Temporal Context Modeling for 4D Occupancy Forecasting.- Wavelet-Driven Cross-Domain Consistency for Mixed-Supervised 3D Tumor Segmentation.- ECTraj: Enhanced Consistency Training for Multi-Agent Trajectory Prediction.- EmbodiedVAE: Disentangled Video VAE for Efficient and Controllable Embodied Manipulation.- RAE-NWM: Navigation World Model in Dense Visual Representation Space.- Show Me Examples: Inferring Visual Concepts from Image Sets.- PointLAM: Local Attentive Mamba for Efficient Point-based 3D Object Detection.- SEM-ROVER: Semantic Voxel-Guided Diffusion for Large-Scale Driving Scene Generation.- SWSL: Semantic-aware Weakly Supervised Learning for 3D Motion Generation using 2D Motion Data.- Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment.- One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models.- Let ViT Speak: Generative Language-Image Pre-training.- Repurposing Geometric Foundation Models for Multi-view Diffusion.- LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving.- Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR.




