覆盖北京时间日期:2026-07-20、2026-07-21。聚焦 Agent、LLM reasoning/planning/tool use/memory、RAG、coding agents、evaluation/benchmark 与训练/推理基础设施。
Top Picks
#1 · Total 22/25
Siddharth Mishra-Sharma · 2026-07-20 · arXiv export API
一句话结论:Neural simulation-based inference enables parameter estimation for complex models, but typically requires the user to specify a simulator encoding a fixed model structure. We present a frame…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Neural simulation-based inference enables parameter estimation for complex models, but typically requires the user to specify a simulator encoding a fixed model structure. We present a framework for joint model selection and parameter estimation that combines…
实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 3Actionability 5
arXiv · PDF
#2 · Total 20/25
Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu, Atri Rudra, Beidi Chen · 2026-07-21 · arXiv export API
一句话结论:Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specif…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism.…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 4Evidence 2Actionability 5
arXiv · PDF
#3 · Total 22/25
Zhuohang Fan, Beichen Zhang, Yuanfa Li, Changqiao Wu, Wei Liu, Jian Luan · 2026-07-20 · arXiv export API
一句话结论:Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-h…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-horizon interaction trajectories collected from element-rich and rapidl…
实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 4Evidence 4Actionability 5
arXiv · PDF
#4 · Total 20/25
Ehsan Futuhi, Nathan R. Sturtevant · 2026-07-20 · arXiv export API
一句话结论:Large Language Models (LLMs) have had a remarkable impact across many areas of machine learning. However, recent studies have shown that they struggle to reliably solve planning problems. At…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Large Language Models (LLMs) have had a remarkable impact across many areas of machine learning. However, recent studies have shown that they struggle to reliably solve planning problems. At the same time, theoretical results have shown that transformers, the…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 3Substance 5Evidence 2Actionability 5
arXiv · PDF
#5 · Total 17/25
Riccardo Poiani, Martino Bernasconi, Andrea Celli · 2026-07-20 · arXiv export API
一句话结论:Reinforcement Learning is a cornerstone technique for modern large reasoning models. Usually, for difficult tasks such as code generation and theorem proving, the agent is evaluated by gener…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Reinforcement Learning is a cornerstone technique for modern large reasoning models. Usually, for difficult tasks such as code generation and theorem proving, the agent is evaluated by generating $K$ responses rather than sampling a single response, and perfor…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 4Novelty 4Substance 4Evidence 2Actionability 3
arXiv · PDF
#6 · Total 21/25
Kemal Devrim Kafadar, Eren Özaltun, Mahmud Efnan Şanlı, Feyza Orak, Emirhan Gazi, Kubilay Kağan Kömürcü · 2026-07-20 · arXiv export API
一句话结论:Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments. To maintain op…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments. To maintain operation during these intermittent communication failures, agents can e…
实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 4Evidence 3Actionability 5
arXiv · PDF
#7 · Total 22/25
Kaiyuan Tang, Maizhe Yang, Chaoli Wang · 2026-07-21 · arXiv export API
一句话结论:Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical. How…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical. However, conventional compressors often struggle to preserve fine structu…
实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 4Novelty 4Substance 5Evidence 4Actionability 5
arXiv · PDF
#8 · Total 16/25
Hany Hamed, Abhishek Naik, Colin Bellinger, A. Rupam Mahmood · 2026-07-20 · arXiv export API
一句话结论:Transfer-oriented reinforcement learning requires evaluating algorithms along dimensions that go beyond standard sample efficiency. We focus on two dimensions: practical efficiency, which as…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Transfer-oriented reinforcement learning requires evaluating algorithms along dimensions that go beyond standard sample efficiency. We focus on two dimensions: practical efficiency, which asks whether conclusions about algorithm suitability change under wall-c…
实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 4Novelty 4Substance 4Evidence 1Actionability 3
arXiv · PDF
#9 · Total 19/25
Damien Teney, Liangze Jiang, Hemanth Saratchandran, Simon Lucey · 2026-07-20 · arXiv export API
一句话结论:Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be key for p…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be key for pushing AI beyond merely scaling current designs. *Method.* We present…
实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 3Novelty 3Substance 5Evidence 4Actionability 4
arXiv · PDF
#10 · Total 14/25
Gaoyue Zhou, Zichen Jeff Cui, Ada Langford, Bowen Tan, Yann LeCun, Lerrel Pinto · 2026-07-21 · arXiv export API
一句话结论:Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a sin…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global token, or rely on visual backbones trained from scratch, sa…
实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 3Novelty 3Substance 3Evidence 1Actionability 4
arXiv · PDF
#11 · Total 18/25
Keru Wang, Yixin Deng, Yao Lyu, Stephen Redmond, Shengbo Eben Li · 2026-07-20 · arXiv export API
一句话结论:Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evalua…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting…
实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 3Novelty 4Substance 5Evidence 3Actionability 3
arXiv · PDF
#12 · Total 18/25
Yasheng Sun, Bohan Li, Youqi Tao, Jürgen Schmidhuber · 2026-07-20 · arXiv export API
一句话结论:A hallmark of neurodegenerative diseases such as Alzheimer's and Parkinson's is the aberrant aggregation of proteins into amyloid fibrils, and small molecules that selectively bind to these…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:A hallmark of neurodegenerative diseases such as Alzheimer's and Parkinson's is the aberrant aggregation of proteins into amyloid fibrils, and small molecules that selectively bind to these fibrils hold promise as diagnostics, imaging probes, and therapeutics.…
实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 3Novelty 3Substance 5Evidence 4Actionability 3
arXiv · PDF
#13 · Total 16/25
Vishal Pandey, Gopal Singh · 2026-07-20 · arXiv export API
一句话结论:Multi-provider LLM gateways reverse proxies that route, load-balance, and rate-limit requests across foundation-model APIs have become critical production infrastructure. Yet the failure mod…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Multi-provider LLM gateways reverse proxies that route, load-balance, and rate-limit requests across foundation-model APIs have become critical production infrastructure. Yet the failure modes specific to this architectural layer remain undocumented, scattered…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 3Novelty 3Substance 5Evidence 2Actionability 3
arXiv · PDF
#14 · Total 17/25
Zijian Zhao, Sen Li · 2026-07-20 · arXiv export API
一句话结论:Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL). A central design question is how many neighboring…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL). A central design question is how many neighboring agents\footnote{In this paper, "neighbors" refer not only to physical…
实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 3Substance 3Evidence 1Actionability 5
arXiv · PDF
#15 · Total 22/25
Peiyu Zang, Bosen Xie, Ruoxiang Xu, Yongqiang Cai · 2026-07-20 · arXiv export API
一句话结论:Physics-Informed Neural Networks (PINNs) solve PDEs by incorporating physical constraints into neural-network training, but large-scale problems are limited by automatic-differentiation memo…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Physics-Informed Neural Networks (PINNs) solve PDEs by incorporating physical constraints into neural-network training, but large-scale problems are limited by automatic-differentiation memory overhead and inefficient execution of grid-based PDE operators. We…
实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 4Novelty 4Substance 5Evidence 4Actionability 5
arXiv · PDF