Paper Archive

Browse and export your curated research paper collection

358
Archived Days
3528
Total Papers
7.4
Avg Score
12
Categories

Export Archive Data

Download your archived papers in various formats

JSON: Complete data with analysis | CSV: Tabular data for analysis | Markdown: Human-readable reports | BibTeX: Academic citations
Browse by Date

Papers for September 17, 2026

10 papers found

[object Object], [object Object], [object Object] 9/16/2026 huggingface

reinforcement learning

Current Mixture-of-Agents (MoA) paradigms generally treat query routing and agent fine-tuning as separate processes, limiting their ability to respond to evolving agent capabilities. This disconnect prevents routing strategies from adapting to evolving agent capabilities during post-training and pre...

Keywords: reinforcement learning, fine-tuning

[object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object] 9/16/2026 huggingface

natural language processing

Speech retrieval has become increasingly important as spoken content continues to grow across meetings, lectures, podcasts, and videos. Existing benchmarks and models have advanced semantic search over spoken content, but largely focus on what is said while overlooking who says it. In many real-worl...

[object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object] 9/16/2026 huggingface

natural language processing

Large audio-language models have shown promise for anti-fraud detection by directly processing speech and reasoning over fraud-related evidence. Their deployment, however, requires predictions to follow a predefined label space and a structured decision protocol consisting of service-scenario identi...

Keywords: fine-tuning, detection, classification

[object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object] 9/16/2026 huggingface

reinforcement learning

Coding agents are typically evaluated with desired behavior specified through issues or instructions. In practical web development, however, agents may need to infer behavior from working software and implement it in an incomplete application. We introduce ProgramDistill, a benchmark evaluating codi...

Keywords: gpt

[object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object] 9/16/2026 huggingface

computer vision

In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated fro...

Keywords: reinforcement learning

Meng'en Qin, Yinchen Liu, Mingxuan Cui, Youlu Xing 9/16/2026 arxiv

computer vision

Visual signals require compact yet sufficient representations for robust downstream prediction. Convolutional sparse coding (CSC) provides an explicit mechanism for suppressing redundant components while preserving signal content, but its sparsity coefficient is typically fixed and manually selected...

Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin 9/16/2026 arxiv

natural language processing

Direct preference alignment methods are widely used to align large language models (LLMs) with human preferences because of their computational and memory efficiency. However, likelihood displacement motivates alternative ways to extract information from preference pairs with small likelihood margin...

Keywords: fine-tuning

Elizabeth Pavlova, Hidenori Tanaka 9/16/2026 arxiv

reinforcement learning

Emergent coordinated behaviors of AI agents are starting to present critical safety risks. A key phenomenon driving these behaviors is the rapid formation and spread of beliefs about the world, and mechanistic understanding is crucial for collective alignment. To this end, we introduce the Flag Game...

Sara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid 9/16/2026 arxiv

computer vision

Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that c...

Keywords: segmentation

Bardienus P. Duisterhof, Kaifeng Zhang, Adam Hung, Bowen Wen, Stan Birchfield, Yunzhu Li, Deva Ramanan, Jeffrey Ichnowski 9/16/2026 arxiv

reinforcement learning

World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods typically require robot action labels to learn action-condit...

Keywords: transformer
Loading...

Preparing your export...