Paper at a Glance
Paper Title: Improved Baselines with Visual Instruction Tuning
Authors: Haotian Liu, Chunyuan Li, Yuheng Li, Yong Jae Lee
Affiliation: University of Wisconsin-Madison, Microsoft Research
Published in: CVPR 2024
Link to Paper: https://openaccess.thecvf.com//content/CVPR2024/html...
Paper at a Glance
Paper Title: Stacked Intelligent Metasurfaces for Multi-Modal Semantic Communications
Authors: Guojun Huang, Jiancheng An, Lu Gan, Dusit Niyato, Mérouane Debbah, and Tie Jun Cui
Affiliation: University of Electronic Science and Technology of China, Nanyang Technological Univers...
Paper at a Glance
Paper Title: When Tokens Talk Too Much: A Survey of Multimodal Long-Context Token Compression across Images, Videos, and Audios
Authors: Kele Shao, Keda Tao, Kejia Zhang, Sicheng Feng, Mu Cai, Yuzhang Shang, Haoxuan You, Can Qin, Yang Sui, and Huan Wang.
Affiliation: A collabor...
Paper at a Glance
Paper Title: Scale-space flow for end-to-end optimized video compression
Authors: Eirikur Agustsson, David Minnen, Nick Johnston, Johannes Ballé, Sung Jin Hwang, George Toderici
Affiliation: Google Research, Perception Team
Published in: IEEE/CVF Conference on Computer Vision a...
Paper at a Glance
Paper Title: Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
Authors: Alex Kendall, Yarin Gal, Roberto Cipolla
Affiliation: University of Cambridge
Published in: Conference on Computer Vision and Pattern Recognition (CVPR), 2018
Link to Pa...