Paper Archive

Browse and export your curated research paper collection

298
Archived Days
2963
Total Papers
7.7
Avg Score
9
Categories

Export Archive Data

Download your archived papers in various formats

JSON: Complete data with analysis | CSV: Tabular data for analysis | Markdown: Human-readable reports | BibTeX: Academic citations
Browse by Date

Papers for July 3, 2026

10 papers found

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Qingyan Bai, Ka Leong Cheng, Yue Yu, Yixuan Li, Yihao Meng, Zichen Liu, Yanhong Zeng, Yujun Shen, Qifeng Chen 7/2/2026 arxiv

computer vision

We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain mot...

Keywords: WorldDirector, video world model, dynamic memory, LLM, 3D trajectory, camera movement, semantic motion orchestration, visual generation

Qiaowei Miao, Kehan Li, Yawei Luo, Yi Yang 7/2/2026 arxiv

computer vision

Generative diffusion models excel at synthesizing high-quality images, videos, and 3D content under multimodal control. However, arbitrary user-defined modality-to-4D (X-to-4D) generation remains challenging due to the high cost of constructing diverse datasets and the limited scalability of existin...

Keywords: Align4D, X-to-4D generation, video-3D pairs, Object Distance Alignment, Motion-Geometry Joint Alignment, Asynchronous Optimization, X4D dataset

Haofei Xu, Rundi Wu, Philipp Henzler, Nikolai Kalischek, Michael Oechsle, Fabian Manhardt, Marc Pollefeys, Andreas Geiger, Federico Tombari, Michael Niemeyer 7/2/2026 arxiv

computer vision

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage pre-trained latent diffusion models. In this work, we show that such architectural overhead and intricate loss formulati...

Keywords: monocular geometry estimation, diffusion models, ViT, DINOv3, 3D reconstruction

Josh Hills, Ida Caspary, Asa Cooper Stickland 7/2/2026 arxiv

machine learning

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt-injected agent can distribute attacks across pull requests (PRs) and time its payload for the PR wi...

Keywords: AI control, distributed attacks, monitoring, iterative VibeCoding, stateful link-tracker, monitor ensemble

Matteo Boglioni, Thibault Rousset, Siva Reddy, Marius Mosbach, Verna Dankers 7/2/2026 arxiv

machine learning

LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged as a promising solution, with state-of-the-art(SOTA) methods often following a localize-first, unlearn-second paradigm th...

Keywords: LLM unlearning, LACUNA testbed, parameter-level localization, resurfacing attacks, privacy

Yanjun Zhao, Ruizhong Qiu, Tianxin Wei, Yuanchen Bei, Zhining Liu, Lingjie Chen, Ismini Lourentzou, Hanghang Tong, Jingrui He 7/2/2026 arxiv

machine learning

Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a...

Keywords: LLM, long-context reasoning, ReContext, evidence replay, associative memory

Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu, Jiacheng Shao, Pengfei Chen, Jiannan Ge, Kaiwen Duan, Qi Tian 7/2/2026 arxiv

computer vision

Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken utterance to its respective character. In this paper, we advance this field throu...

Keywords: speaker recognition, long-form TV dramas, DramaSR-532K, DramaSR-LRM, multimodal aggregation, large reasoning model

Shuai Tian, Yupeng Zheng, Yuhang Zheng, Songen Gu, Yujie Zang, Yuxing Qin, Weize Li, Haoran Li, Wenchao Ding, Dongbin Zhao 7/2/2026 arxiv

computer vision

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observations. Existing visual-tactile policies usually feed tactile observations directly into action prediction, but rarely mod...

Keywords: Visual-Tactile World Action Model, contact-rich manipulation, Asymmetric Mixture-of-Transformers, contact-gated attention, robotic assembly

Ling Xu, Chuyu Han, Borui Li, Hao Wu, Shiqi Jiang, Ting Cao, Chuanyou Li, Sheng Zhong, Shuai Wang 7/2/2026 arxiv

computer vision

Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, and robot-side glue code, especially on heterogeneous edge devices. Existing inference runtimes are de...

Keywords: embodied AI, inference runtime, heterogeneous robots, VLA models, WAM models, C++, modular architecture, multi-rate execution

Gil Harari, Yoel Zimmermann, Ola Tangen Kulseng, Laura Zichi, Chuin Wei Tan, Marc L. Descoteaux, Boris Kozinsky 7/2/2026 arxiv

machine learning

Machine learning interatomic potentials (MLIPs) have become a hallmark of AI for scientific simulation. While efforts on new architectures and datasets have led to increasingly accurate and general models, the choice of optimizer for training has largely remained unexplored, defaulting to Adam and i...

Keywords: machine learning, interatomic potentials, optimizers, SOAP, Muon, MLIPs, NequIP, Allegro
Loading...

Preparing your export...