覆盖北京时间日期:2026-07-13、2026-07-14。聚焦 Agent、LLM reasoning/planning/tool use/memory、RAG、coding agents、evaluation/benchmark 与训练/推理基础设施。
Top Picks
#1 · Total 23/25
Xuefeng Li, Pengfei Liu · 2026-07-13 · arXiv export API
一句话结论:Mixture-of-Experts (MoE) models scale capacity without proportional compute cost and have become a key architecture for frontier large language models (LLMs). Yet domain-specific post-traini…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Mixture-of-Experts (MoE) models scale capacity without proportional compute cost and have become a key architecture for frontier large language models (LLMs). Yet domain-specific post-training inherits an expert pool shaped by mixed-domain pre-training: a subs…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 3Substance 5Evidence 5Actionability 5
arXiv · PDF
#2 · Total 18/25
Zhe Xiao, Longfei Li, Xu He, Haoying Wu, Zixing Zhang, Mingyu Liu · 2026-07-13 · arXiv export API
一句话结论:Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is challenging. This process requires accurate visual-to-sy…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is challenging. This process requires accurate visual-to-symbolic construction of circuit structure from images and correct multi…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 3Substance 3Evidence 2Actionability 5
arXiv · PDF
#3 · Total 24/25
Tatiana Pelc, Gila Kamhi, Asaf Avrahamy, Adi Fledel-Alon · 2026-07-13 · arXiv export API
一句话结论:In the rapidly evolving landscape of information retrieval systems, the ability to adapt and improve through user feedback is paramount. This study introduces a novel methodology for refinin…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:In the rapidly evolving landscape of information retrieval systems, the ability to adapt and improve through user feedback is paramount. This study introduces a novel methodology for refining the performance of a primary Retrieval Augmented Generation (RAG) sy…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 5Actionability 5
arXiv · PDF
#4 · Total 24/25
Chenglin Yu, Hongquan Gui, Ying Yu, Hongxia Yang, Ming Li · 2026-07-13 · arXiv export API
一句话结论:LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and de…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, including logs, context snapshots, checkpoints, and debug traces. We introduce AgentFootprint, a cross-framework benchmark o…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 5Actionability 5
arXiv · PDF
#5 · Total 24/25
Yue Fang, Zhibang Yang, Fangkai Yang, Xiaoting Qin, Liqun Li, Qingwei Lin · 2026-07-13 · arXiv export API
一句话结论:Large language model (LLM) agents increasingly rely on external tools served by shared providers and accessed by heterogeneous downstream agents. Existing approaches improve tool use on the…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Large language model (LLM) agents increasingly rely on external tools served by shared providers and accessed by heterogeneous downstream agents. Existing approaches improve tool use on the agent side through parameter updates, prompt refinement, or agent-side…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 5Actionability 5
arXiv · PDF
#6 · Total 23/25
Tianjing Zeng, Yuntao Hong, Zhongjun Ding, Dandan Liu, Yinan Mei, Yunxiang Su · 2026-07-13 · arXiv export API
一句话结论:Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software engineering, it operates in an open, ambiguous, and continuously evolving environment. These characteristics call for a da…
实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 4Actionability 5
arXiv · PDF
#7 · Total 22/25
Wencheng Ye, Yi Bin, Yujuan Ding, Hongye Fang, Zheng Wang, Xing Xu · 2026-07-13 · arXiv export API
一句话结论:Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounde…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning. To understand this fragility, we examine the internal dyn…
实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 4Evidence 4Actionability 5
arXiv · PDF
#8 · Total 22/25
Xinchen Liu, Hang Zhou, Yingjie Zong, Yuchuan Tian, Liuyang Song, Shuo Zhang · 2026-07-13 · arXiv export API
一句话结论:Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification. At…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification. At the same time, frontier and open models are becoming structurally spec…
实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 3Actionability 5
arXiv · PDF
#9 · Total 22/25
Bingteng Sun, Hao Yin, Yiling Chen, Renjie Xiao, Lei Xie, Shanyou Wang · 2026-07-13 · arXiv export API
一句话结论:Zero-dimensional reduced-order models (0D ROMs) are central to multi-dimensional design workflows for high-end complex equipment. However, the planning process currently relies on manual exp…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Zero-dimensional reduced-order models (0D ROMs) are central to multi-dimensional design workflows for high-end complex equipment. However, the planning process currently relies on manual expertise, limiting topological exploration and prolonging iterations. Ev…
实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 3Actionability 5
arXiv · PDF
#10 · Total 21/25
Xiaojian Liu, Han Xu, Jianqiang Xia, Zhixuan Li, Ke Xu, Yiwei Dai · 2026-07-13 · arXiv export API
一句话结论:Reinforcement learning with verifiable rewards (RLVR) optimizes LLMs using sparse verifiable final-answer rewards. This sparse anchor reliably verifies whether a trajectory succeeds but prov…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Reinforcement learning with verifiable rewards (RLVR) optimizes LLMs using sparse verifiable final-answer rewards. This sparse anchor reliably verifies whether a trajectory succeeds but provides no direct feedback on the reasoning path that produced it. Before…
实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 4Evidence 3Actionability 5
arXiv · PDF
#11 · Total 20/25
Yibo Hu, Ren Wang · 2026-07-14 · arXiv export API
一句话结论:As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so ev…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 4Evidence 2Actionability 5
arXiv · PDF
#12 · Total 18/25
Ke Xu, Han Xu, Xinran Chen, Yuqian Wang, Zhixuan Li, Xiaojian Liu · 2026-07-13 · arXiv export API
一句话结论:Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expo…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we call the…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 3Evidence 2Actionability 4
arXiv · PDF
#13 · Total 18/25
Xiuwei Chen, Quanlin Chen, Wentao Hu, Zisheng Chen, Kun Xiang, Zehua Ma · 2026-07-13 · arXiv export API
一句话结论:Recent multimodal large language models (MLLMs) have made remarkable progress on fine-grained perception tasks under the "Thinking with Images" (TwI) paradigm by iteratively performing vario…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Recent multimodal large language models (MLLMs) have made remarkable progress on fine-grained perception tasks under the "Thinking with Images" (TwI) paradigm by iteratively performing various visual tool operations. However, this paradigm relies heavily on fr…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 3Evidence 2Actionability 4
arXiv · PDF
#14 · Total 14/25
Bojie Li, Noah Shi · 2026-07-13 · arXiv export API
一句话结论:There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one. Both share a hidden limit: they are internal. Every extra tok…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one. Both share a hidden limit: they are internal. Every extra token comes from the same frozen weights and the same prompt, so neither…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 4Novelty 3Substance 2Evidence 2Actionability 3
arXiv · PDF
#15 · Total 24/25
Runhui Huang, Qihui Zhang, Zhe Liu, Yu Gao, Jie Wu, Hengshuang Zhao · 2026-07-14 · arXiv export API
一句话结论:In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verifi…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 5Actionability 5
arXiv · PDF