Paper Archive

Browse and export your curated research paper collection

358
Archived Days
3528
Total Papers
7.4
Avg Score
12
Categories

Export Archive Data

Download your archived papers in various formats

JSON: Complete data with analysis | CSV: Tabular data for analysis | Markdown: Human-readable reports | BibTeX: Academic citations
Browse by Date

Papers for September 23, 2026

10 papers found

[object Object], [object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface

natural language processing

AI agents are beginning to automate research and development across the AI stack, from improving training efficiency to optimizing inference. A natural next step is to improve the research efficiency of the agents themselves. When an AI research agent's own code is the object of optimization, each a...

Keywords: machine learning

[object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface

natural language processing

Reinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely Completen...

Keywords: reinforcement learning

[object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface

reinforcement learning

Deep learning systems now mediate military decisions to use force, yet their internal logic resists inspection, their evaluation practices are gameable, and their deployment fractures accountability across dispersed stakeholders. The ethical challenge posed by these systems is fundamentally epistemi...

Keywords: deep learning

[object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface

machine learning

LLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generativ...

[object Object], [object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface

natural language processing

Large language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores only through local decoding noise, it tends to prod...

Keywords: reinforcement learning

Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers 9/22/2026 arxiv

natural language processing

Agents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improvin...

Keywords: gpt

Yanshuo Bai, Kanji Tanaka 9/22/2026 arxiv

computer vision

Thermal Visual Place Recognition (Thermal VPR) maps camera observations to metric poses within a mapped environment, serving as a prerequisite for autonomous navigation. However, thermal VPR suffers from severe environmental dependence, heavy online retraining overheads, and an inability to model dy...

Keywords: backpropagation, fine-tuning

Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen 9/22/2026 arxiv

natural language processing

Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation. However, their practical deployment remains limited by inefficient inference, largely due to the absence of effective Key-Value (KV) caching...

Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian, Yan Xia, Furu Wei 9/22/2026 arxiv

natural language processing

A multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often constrained by a central orchestrator's capacity to allocate tasks and coordi...

Keywords: gpt

Bao Tang, Jiahao Guo, Haoxiang Cao, Wenyu Liu, Changqian Yu, Kun Gai, Xinggang Wang 9/22/2026 arxiv

computer vision

Vector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challen...

Loading...

Preparing your export...