Browse and export your curated research paper collection
[object Object], [object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface
natural language processingAI agents are beginning to automate research and development across the AI stack, from improving training efficiency to optimizing inference. A natural next step is to improve the research efficiency of the agents themselves. When an AI research agent's own code is the object of optimization, each a...
[object Object], [object Object], [object Object], [object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface
natural language processingReinforcement learning has become a central component of large language model (LLM) post-training, yet token-level credit lacks a generally accepted mathematical definition, leaving its relationship to commonly used training signals unclear. We formulate three regularity conditions, namely Completen...
[object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface
reinforcement learningDeep learning systems now mediate military decisions to use force, yet their internal logic resists inspection, their evaluation practices are gameable, and their deployment fractures accountability across dispersed stakeholders. The ethical challenge posed by these systems is fundamentally epistemi...
[object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface
machine learningLLM-as-a-judge enables evaluation across diverse tasks, but inference cost and confidence reliability become critical at scale. We study whether a decision-only judge can provide an economical first pass and identify when stronger evaluation is needed. Comparing jev-as-a-judge with sixteen generativ...
[object Object], [object Object], [object Object], [object Object], [object Object] 9/22/2026 huggingface
natural language processingLarge language models increasingly tackle hard reasoning problems by spending more test-time compute, yet the dominant strategy remains naive repeated sampling: draw many independent solutions and hope one is correct. Because such sampling explores only through local decoding noise, it tends to prod...
Trang Nguyen, Eulrang Cho, Bingqing Chen, Tim Dettmers 9/22/2026 arxiv
natural language processingAgents often work on complex problems that require millions of tokens of context, which necessitates compacting across sessions due to limited context windows. We develop CliffCompaction, an autocompaction technique that reduces cost by up to 50% under a bounded context while maintaining or improvin...
Yanshuo Bai, Kanji Tanaka 9/22/2026 arxiv
computer visionThermal Visual Place Recognition (Thermal VPR) maps camera observations to metric poses within a mapped environment, serving as a prerequisite for autonomous navigation. However, thermal VPR suffers from severe environmental dependence, heavy online retraining overheads, and an inability to model dy...
Quan Nguyen-Tri, Mukul Ranjan, Zhiqiang Shen 9/22/2026 arxiv
natural language processingDiffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation. However, their practical deployment remains limited by inefficient inference, largely due to the absence of effective Key-Value (KV) caching...
Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian, Yan Xia, Furu Wei 9/22/2026 arxiv
natural language processingA multi-agent system can reduce latency on complex tasks by executing work concurrently. Several pioneering harness frameworks support multi-agent systems. However, the scalability of current multi-agent harnesses is often constrained by a central orchestrator's capacity to allocate tasks and coordi...
Bao Tang, Jiahao Guo, Haoxiang Cao, Wenyu Liu, Changqian Yu, Kun Gai, Xinggang Wang 9/22/2026 arxiv
computer visionVector Quantization (VQ) is fundamental to discrete visual tokenizers that power modern autoregressive and masked image generation models. While recent shared-projection codebook methods have substantially advanced codebook utilization, training stability remains a critical and underexplored challen...
Preparing your export...