Paper Archive

Browse and export your curated research paper collection

298
Archived Days
2963
Total Papers
7.7
Avg Score
9
Categories

Export Archive Data

Download your archived papers in various formats

JSON: Complete data with analysis | CSV: Tabular data for analysis | Markdown: Human-readable reports | BibTeX: Academic citations
Browse by Date

Papers for July 4, 2026

10 papers found

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang, Qingyan Bai, Ka Leong Cheng, Yue Yu, Yixuan Li, Yihao Meng, Zichen Liu, Yanhong Zeng, Yujun Shen, Qifeng Chen 7/2/2026 arxiv

computer vision

We present WorldDirector, a highly controllable video world model framework designed for persistent dynamic object memory and unrestricted viewpoint exploration. Unlike existing world models that entangle physical dynamics with pixel rendering and rely on continuous visual observation to sustain mot...

Keywords: WorldDirector, dynamic world model, persistent memory, LLM, 3D trajectory coordination, video generation

Qiaowei Miao, Kehan Li, Yawei Luo, Yi Yang 7/2/2026 arxiv

computer vision

Generative diffusion models excel at synthesizing high-quality images, videos, and 3D content under multimodal control. However, arbitrary user-defined modality-to-4D (X-to-4D) generation remains challenging due to the high cost of constructing diverse datasets and the limited scalability of existin...

Keywords: Align4D, X-to-4D generation, video, 3D, diffusion models, Object Distance Alignment, Motion-Geometry Joint Alignment, Asynchronous Optimization

Haofei Xu, Rundi Wu, Philipp Henzler, Nikolai Kalischek, Michael Oechsle, Fabian Manhardt, Marc Pollefeys, Andreas Geiger, Federico Tombari, Michael Niemeyer 7/2/2026 arxiv

computer vision

State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage pre-trained latent diffusion models. In this work, we show that such architectural overhead and intricate loss formulati...

Keywords: monocular geometry estimation, diffusion models, 3D reconstruction, pixel-space Diffusion Transformer, DINOv3

Josh Hills, Ida Caspary, Asa Cooper Stickland 7/2/2026 arxiv

machine learning

As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt-injected agent can distribute attacks across pull requests (PRs) and time its payload for the PR wi...

Keywords: AI control, distributed attacks, persistent-state, monitoring, evasion, Iterative VibeCoding, stateful link-tracker, monitor ensemble
View Paper

Matteo Boglioni, Thibault Rousset, Siva Reddy, Marius Mosbach, Verna Dankers 7/2/2026 arxiv

machine learning

LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc removal methods. Unlearning has emerged as a promising solution, with state-of-the-art(SOTA) methods often following a localize-first, unlearn-second paradigm th...

Keywords: LLM unlearning, data privacy, parameter localization, resurfacing attacks, AI security

Wentao Zhang, Liliana Hotsko, Woojeong Kim, Pengyu Nie, Stuart Shieber, Yuntian Deng 7/2/2026 arxiv

computer vision

Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON, or ranking search results by intent, and are increasingly outsourced to large language model APIs at the cost of locality, reproducibility, and price. We propose...

Keywords: fuzzy-function programming, Program-as-Weights, neural artifacts, AI tool building, AI ethics, explainability

Yanjun Zhao, Ruizhong Qiu, Tianxin Wei, Yuanchen Bei, Zhining Liu, Lingjie Chen, Ismini Lourentzou, Hanghang Tong, Jingrui He 7/2/2026 arxiv

machine learning

Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applications. Although recent LLMs support increasingly long context windows, they often fail to use relevant evidence that is already present in the input, revealing a...

Keywords: LLM, long-context reasoning, RECONTEXT, evidence replay, associative memory

Dengyang Jiang, Mengmeng Wang, Harry Yang, Jingdong Wang 7/2/2026 arxiv

computer vision

Representation alignment has become an effective way to accelerate diffusion transformer training and improve generation quality. Recent self-alignment methods, such as SRA and Self-Flow, further remove the dependency on external pretrained encoders by constructing alignment within the diffusion mod...

Keywords: data augmentation, self-supervision, diffusion models, attention separation, image generation

Yuxuan Li, Lingxi Xie, Xinyue Huo, Jihao Qiu, Jiacheng Shao, Pengfei Chen, Jiannan Ge, Kaiwen Duan, Qi Tian 7/2/2026 arxiv

computer vision

Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies on \textbf{speaker recognition}, the task of accurately attributing each spoken utterance to its respective character. In this paper, we advance this field throu...

Keywords: speaker recognition, long-form TV dramas, multimodal, reasoning model, DramaSR-532K, DramaSR-LRM

Shuai Tian, Yupeng Zheng, Yuhang Zheng, Songen Gu, Yujie Zang, Yuxing Qin, Weize Li, Haoran Li, Wenchao Ding, Dongbin Zhao 7/2/2026 arxiv

computer vision

Contact-rich manipulation requires policies to react to local deformation, pressure, slip, and friction, yet these cues are temporally sparse and often invisible in visual observations. Existing visual-tactile policies usually feed tactile observations directly into action prediction, but rarely mod...

Keywords: Visual-Tactile World Action Model, contact-rich manipulation, tactile deformation dynamics, contact-gated attention, robotics, machine learning
Loading...

Preparing your export...