Paper Archive

Browse and export your curated research paper collection

197
Archived Days
1958
Total Papers
7.9
Avg Score
9
Categories

Export Archive Data

Download your archived papers in various formats

JSON: Complete data with analysis • CSV: Tabular data for analysis • Markdown: Human-readable reports • BibTeX: Academic citations
Browse by Date

Papers for March 22, 2026

10 papers found

Xianjin Wu, Dingkang Liang, Tianrui Feng, Kui Xia, Yumeng Zhang, Xiaofan Li, Xiao Tan, Xiang Bai 3/19/2026 arxiv

computer vision

While Multimodal Large Language Models demonstrate impressive semantic capabilities, they often suffer from spatial blindness, struggling with fine-grained geometric reasoning and physical dynamics. Existing solutions typically rely on explicit 3D modalities or complex geometric scaffolding, which a...

Keywords: VEGA-3D, video diffusion, latent world simulator, implicit 3D priors, MLLM, spatial reasoning, embodied manipulation, 3D scene understanding

Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Jeffrey Hu, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao, Josef Bengtson, Chenliang Zhou, Weihao Xia, Cristina Nader Vasconcelos. Cengiz Oztireli 3/19/2026 arxiv

computer vision

The ability to render scenes at adjustable fidelity from a single model, known as level of detail (LoD), is crucial for practical deployment of 3D Gaussian Splatting (3DGS). Existing discrete LoD methods expose only a limited set of operating points, while concurrent continuous LoD approaches enable...

Keywords: Matryoshka Gaussian Splatting, 3D Gaussian Splatting, Level of Detail, stochastic budget training, continuous LoD, real-time rendering, prefix rendering, model compression

Yuqing Wang, Chuofan Ma, Zhijie Lin, Yao Teng, Lijun Yu, Shuai Wang, Jiaming Han, Jiashi Feng, Yi Jiang, Xihui Liu 3/19/2026 arxiv

computer vision

Visual generation with discrete tokens has gained significant attention as it enables a unified token prediction paradigm shared with language models, promising seamless multimodal architectures. However, current discrete generation methods remain limited to low-dimensional latent tokens (typically ...

Keywords: CubiD, discrete diffusion, high-dimensional tokens, visual generation, representation discretization, masking, ImageNet-256, multimodal

Bryce Grant, Xijia Zhao, Peng Wang 3/19/2026 arxiv

robotics

Vision-Language-Action (VLA) models combine perception, language, and motor control in a single architecture, yet how they translate multimodal inputs into actions remains poorly understood. We apply activation injection, sparse autoencoders (SAEs), and linear probes to six models spanning 80M--7B p...

Keywords: Vision-Language-Action, VLA, activation injection, sparse autoencoders, linear probes, motor programs, Action Atlas, robotics

Haitian Li, Haozhe Xie, Junxiang Xu, Beichen Wen, Fangzhou Hong, Ziwei Liu 3/19/2026 arxiv

computer vision

Reconstructing articulated 3D objects from a single image requires jointly inferring object geometry, part structure, and motion parameters from limited visual evidence. A key difficulty lies in the entanglement between motion cues and object structure, which makes direct articulation regression uns...

Keywords: MonoArt, monocular, articulated 3D reconstruction, progressive structural reasoning, canonicalization, part representations, motion-aware embeddings, PartNet-Mobility

Huaide Jiang, Yash Chaudhary, Yuping Wang, Zehao Wang, Raghav Sharma, Manan Mehta, Yang Zhou, Lichao Sun, Zhiwen Fan, Zhengzhong Tu, Jiachen Li 3/19/2026 arxiv

machine learning

There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language instructions; and Object-Goal Navigation (OGN), where agents navigate to a specified target object. However, existing work primarily evaluates model performanc...

Keywords: NavTrust, embodied_navigation, robustness, VLN, OGN, RGB-D corruptions, instruction corruption, benchmark

Chenyang Gu, Mingyuan Zhang, Haozhe Xie, Zhongang Cai, Lei Yang, Ziwei Liu 3/19/2026 arxiv

machine learning

Prior motion generation largely follows two paradigms: continuous diffusion models that excel at kinematic control, and discrete token-based generators that are effective for semantic conditioning. To combine their strengths, we propose a three-stage framework comprising condition feature extraction...

Keywords: MoTok, diffusion-based discrete motion tokenizer, motion generation, semantic conditioning, kinematic control, HumanML3D, MaskControl, Perception-Planning-Control

Xinyao Zhang, Wenkai Dong, Yuxin Song, Bo Fang, Qi Zhang, Jing Wang, Fan Chen, Hui Zhang, Haocheng Feng, Yu Lu, Hang Zhou, Chun Yuan, Jingdong Wang 3/19/2026 arxiv

computer vision

Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely on injecting explicit external priors (e.g., VLM features or structural conditions) to mitigate these issues, this relia...

Keywords: SAMA, instruction-guided video editing, semantic anchoring, motion alignment, factorized pretraining, cube inpainting, speed perturbation, tube shuffle

Nobuo Yoshii, Xinran Nicole Han, Ryo Kawahara, Todd Zickler, Ko Nishino 3/19/2026 arxiv

computer vision

We introduce Multi-Object Generative Perception (MultiGP), a generative inverse rendering method for stochastic sampling of all radiometric constituents -- reflectance, texture, and illumination -- underlying object appearance from a single image. Our key idea to solve this inherently ambiguous radi...

Keywords: inverse rendering, reflectance, illumination, texture, diffusion, axial attention, ControlNet, generative perception

Yang Fu, Yike Zheng, Ziyun Dai, Henghui Ding 3/19/2026 arxiv

computer vision

Video object removal aims to eliminate dynamic target objects and their visual effects, such as deformation, shadows, and reflections, while restoring seamless backgrounds. Recent diffusion-based video inpainting and object removal methods can remove the objects but often struggle to erase these eff...

Keywords: video_object_removal, VOR_dataset, EffectErase, video_inpainting, reciprocal_learning, insertion_removal_consistency, task_aware_region_guidance, diffusion_based_methods
Loading...

Preparing your export...