覆盖北京时间日期:2026-07-12、2026-07-13。聚焦 Agent、LLM reasoning/planning/tool use/memory、RAG、coding agents、evaluation/benchmark 与训练/推理基础设施。
Top Picks
#1 · Total 24/25
Kaiji Zhou, Ales Leonardis, Yue Feng · 2026-07-11 · arXiv export API
一句话结论:Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call API…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs based on coarse-grained matching between tasks and the functions of…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 5Actionability 5
arXiv · PDF
#2 · Total 20/25
Katherine Swinea, Kshitiz Aryal, Lopamudra Praharaj, Maanak Gupta · 2026-07-11 · arXiv export API
一句话结论:Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a need for scalable and adaptive secu…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a need for scalable and adaptive security testing approaches. While recent adoptions of Large Language Mode…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 4Evidence 2Actionability 5
arXiv · PDF
#3 · Total 24/25
Nirjhar Das, Md. Al-Mamun Provath · 2026-07-11 · arXiv export API
一句话结论:We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and ac…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 5Actionability 5
arXiv · PDF
#4 · Total 20/25
Michał Mazuryk, Fleur Dolmans, Louis Gehringer, Ina Klaric, Jia-Huei Ju, Mohammad Aliannejadi · 2026-07-04 · arXiv export API
一句话结论:Recent work has suggested that adding irrelevant documents to the input of retrieval-augmented generation (RAG) systems can improve question-answering performance, a phenomenon referred to a…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Recent work has suggested that adding irrelevant documents to the input of retrieval-augmented generation (RAG) systems can improve question-answering performance, a phenomenon referred to as the Power of Noise. This motivated investigations into the role of n…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 4Evidence 2Actionability 5
arXiv · PDF
#5 · Total 19/25
Kangwei Xu, Bing Li, Ulf Schlichtmann · 2026-07-11 · arXiv export API
一句话结论:As chip complexity increases and time-to-market pressures grow, front-end design has become a critical bottleneck in chip development. Recently, Large Language Models (LLMs) have shown great…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:As chip complexity increases and time-to-market pressures grow, front-end design has become a critical bottleneck in chip development. Recently, Large Language Models (LLMs) have shown great potential in Electronic Design Automation (EDA). Beyond specification…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 3Evidence 2Actionability 5
arXiv · PDF
#6 · Total 22/25
Cláudio Lúcio do Val Lopes, Lucca Machado da Silva · 2026-07-11 · arXiv export API
一句话结论:Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing to balance anomaly interdiction with customer friction. To overcome th…
实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 4Novelty 4Substance 5Evidence 4Actionability 5
arXiv · PDF
#7 · Total 23/25
Shravan Murlidaran, Miguel P. Eckstein · 2026-07-11 · arXiv export API
一句话结论:Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descrip…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 3Substance 5Evidence 5Actionability 5
arXiv · PDF
#8 · Total 21/25
Yuan Cao, Haiqian Yang · 2026-07-11 · arXiv export API
一句话结论:Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they sh…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: the representational frame within which t…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 5Evidence 2Actionability 5
arXiv · PDF
#9 · Total 18/25
Chengkai Zhu, Ziao Tang, Guocheng Zhen, Yimeng Cao, Yusheng Zhao, Ranyiliu Chen · 2026-07-11 · arXiv export API
一句话结论:Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information processing, underpinning quantum communication, computation, and error correctio…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information processing, underpinning quantum communication, computation, and error correction. Formalizing its coding theorems requires connecting finite-block pr…
实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 5Novelty 4Substance 3Evidence 1Actionability 5
arXiv · PDF
#10 · Total 21/25
Yujie Pang, Zudong Li · 2026-07-11 · arXiv export API
一句话结论:Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization bu…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often introduce high inference latency and GPU-memory cost, while vi…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 4Novelty 3Substance 5Evidence 5Actionability 4
arXiv · PDF
#11 · Total 20/25
Sanjid Hasan, Md. Abdur Rahman · 2026-07-11 · arXiv export API
一句话结论:Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Beng…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause of this failure as the model…
实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 4Novelty 4Substance 5Evidence 3Actionability 4
arXiv · PDF
#12 · Total 19/25
Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou · 2026-07-11 · arXiv export API
一句话结论:The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual represent…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich…
实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 3Novelty 4Substance 5Evidence 3Actionability 4
arXiv · PDF
#13 · Total 17/25
Hannah M. Liu, Rhea Saxena, Shiv Asthana · 2026-07-11 · arXiv export API
一句话结论:The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them. In this pape…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them. In this paper, we introduce the TrustX Agent Risk Classification Framework, a stru…
实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 4Novelty 3Substance 4Evidence 1Actionability 5
arXiv · PDF
#14 · Total 20/25
Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen · 2026-07-11 · arXiv export API
一句话结论:Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordab…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion necessary f…
实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 3Novelty 3Substance 5Evidence 5Actionability 4
arXiv · PDF
#15 · Total 16/25
Lulin Liu, Nuo Chen, Yan Wang, Bangya Liu, Wenyan Cong, Hezhen Hu · 2026-07-11 · arXiv export API
一句话结论:Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-t…
为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。
方法要点:Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sour…
实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。
局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。
Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。
Relevance 3Novelty 3Substance 4Evidence 2Actionability 4
arXiv · PDF