Paper Archive

Browse and export your curated research paper collection

298
Archived Days
2963
Total Papers
7.7
Avg Score
9
Categories

Export Archive Data

Download your archived papers in various formats

JSON: Complete data with analysis | CSV: Tabular data for analysis | Markdown: Human-readable reports | BibTeX: Academic citations
Browse by Date

Papers for July 11, 2026

10 papers found

Jiangwei Ren, Xingyu Jiang, Zijie Song, Wei Xu, Hongkai Lin, Dingkang Liang, Xiang Bai 7/9/2026 arxiv

computer vision

Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pioneering methods rely on massive dense annotations that are impractical in underwater settings. In this paper, we propose ...

Keywords: underwater 3D geometry, semi-supervised learning, cross-domain adaptation, point cloud reconstruction, depth estimation

Fabio Tosi, Luca Bartolomei, Matteo Poggi, Stefano Mattoccia 7/9/2026 arxiv

computer vision

Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond the reach of embedded and mobile platforms. Lightweight alternatives exist, but have been developed almost exclusively wi...

Keywords: monocular depth estimation, zero-shot learning, lightweight models, knowledge distillation, mobile computing

Cheng-De Fan, Chun-Wei Tuan Mu, Chen-Wei Chang, Chin-Yang Lin, Kun-Ru Wu, Yu-Chee Tseng, Yu-Lun Liu 7/9/2026 arxiv

computer vision

Recovering high-quality video from sparse event streams is a challenging task. Regression methods often blur textures, while existing generative models struggle with long-term stability. We propose LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-ba...

Keywords: video reconstruction, event-based video, video diffusion models, frame interpolation, temporal coherence, zero-shot generalization

Weijian Chen, Weibo Yao, Yuhang Zhang, Xiaolin Tang, Guo Wang, Weijun Zhang, Xitong Gao, Yihao Chen, Hongde Qin, Lu Qi 7/9/2026 arxiv

computer vision

Scaling 3D Gaussian Splatting (3DGS) to large outdoor scenes is costly in both data acquisition and computation. Adopting panoramic images with equirectangular projection (ERP) can reduce capture effort via their full $360^{\circ}$ field of view, yet the resulting omnipresent visibility invalidates ...

Keywords: 3D Gaussian Splatting, panoramic images, PanoLOG, sky-sphere modeling, G^2PS

Zhekai Chen, Chengqi Duan, Kaiyue Sun, Bohao Li, Yuqing Wang, Manyuan Zhang, Xihui Liu 7/9/2026 arxiv

machine learning

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they...

Keywords: UniClawBench, proactive agents, real-world tasks, benchmark, large language models, multimodal understanding

Hongyu Liu, Chun Wang, Feng Gao, Xuanhua He, Yue Ma, Ziyu Wan, Yong Zhang, Xiaoming Wei, Qifeng Chen 7/9/2026 arxiv

computer vision

We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from error accumulation and weakened motion dynamics during long autoregr...

Keywords: video generation, autoregressive models, self-distillation, post-training, few-step AR

Haoran Feng, Ruiyang Zhang, Longyi Zhang, Dizhe Zhang, Lu Qi 7/9/2026 arxiv

computer vision

In this work, we present Canvas360, a two-stage framework for in-context panoramic generation that combines geometry-aware pretraining with downstream task-specific fine-tuning. To address the lack of large-scale, high-quality training data tailored to in-context panoramic tasks, we propose Canvas36...

Keywords: in-context panoramic generation, geometry-aware pretraining, Canvas360Dataset, parallel depth generation, FAED metric

Xinyan Chen, Ziyu Guo, Renrui Zhang, Dongzhi Jiang, Hongsheng Li 7/9/2026 arxiv

computer vision

Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct from previous Chain-of-Thought (CoT): reasoning can unfold through temporally connected frames, known...

Keywords: OpenCoF, video generation, reasoning, Chain-of-Frame, Wan-CoF, visual reasoning, textual reasoning

David GonzΓ‘lez-MartΓ­nez, Shiwei Liu 7/9/2026 arxiv

machine learning

Low-rank factorization is widely used to compress neural networks, but modern models are often not naturally amenable to aggressive factorization without significant accuracy loss. Existing training-time low-rank regularizers can improve compressibility, but they often require SVDs of large weight m...

Keywords: SLORR, low-rank regularization, neural network compression, Hoyer sparsity, nuclear norm, GPU-friendly, accuracy preservation

Yunchao Yao, Zhuxiu Xu, Tianqi Zhang, Zixian Liu, Sikai Li, Zhenyu Wei, Feng Chen, Dihong Huang, Kechang Wan, Chenyang Ma, Shuqi Zhao, Shenghua Gao, Masayoshi Tomizuka, Yi Ma, Mingyu Ding 7/9/2026 arxiv

robotics

Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. However, existing benchmarks remain limited in task and data diversity, embodime...

Keywords: DexVerse, dexterous manipulation, benchmark, multi-task, multi-embodiment, visuomotor generalization
Loading...

Preparing your export...