FJAIResearch StudioResearch online

AI Paper Daily

AI 论文日报 · 2026-07-13

覆盖北京时间日期:2026-07-12、2026-07-13。聚焦 Agent、LLM reasoning/planning/tool use/memory、RAG、coding agents、evaluation/benchmark 与训练/推理基础设施。

执行摘要

candidate count33
new included count33
selected count15

Top 3:Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation;VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents;Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

Top Picks

#1 · Total 24/25

Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation

Kaiji Zhou, Ales Leonardis, Yue Feng · 2026-07-11 · arXiv export API

一句话结论:Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call API…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs based on coarse-grained matching between tasks and the functions of…

实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 5Evidence 5Actionability 5

#2 · Total 20/25

VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents

Katherine Swinea, Kshitiz Aryal, Lopamudra Praharaj, Maanak Gupta · 2026-07-11 · arXiv export API

一句话结论:Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a need for scalable and adaptive secu…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a need for scalable and adaptive security testing approaches. While recent adoptions of Large Language Mode…

实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 4Evidence 2Actionability 5

#3 · Total 24/25

Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026

Nirjhar Das, Md. Al-Mamun Provath · 2026-07-11 · arXiv export API

一句话结论:We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyramid-style questions from incrementally revealed text and ac…

实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 5Evidence 5Actionability 5

#4 · Total 20/25

The Powerless Noise: How Experimental Settings Shape the Reported Power of Noise

Michał Mazuryk, Fleur Dolmans, Louis Gehringer, Ina Klaric, Jia-Huei Ju, Mohammad Aliannejadi · 2026-07-04 · arXiv export API

一句话结论:Recent work has suggested that adding irrelevant documents to the input of retrieval-augmented generation (RAG) systems can improve question-answering performance, a phenomenon referred to a…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Recent work has suggested that adding irrelevant documents to the input of retrieval-augmented generation (RAG) systems can improve question-answering performance, a phenomenon referred to as the Power of Noise. This motivated investigations into the role of n…

实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 4Evidence 2Actionability 5

#5 · Total 19/25

LLM for EDA in Front-End Design: Challenges and Opportunities

Kangwei Xu, Bing Li, Ulf Schlichtmann · 2026-07-11 · arXiv export API

一句话结论:As chip complexity increases and time-to-market pressures grow, front-end design has become a critical bottleneck in chip development. Recently, Large Language Models (LLMs) have shown great…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:As chip complexity increases and time-to-market pressures grow, front-end design has become a critical bottleneck in chip development. Recently, Large Language Models (LLMs) have shown great potential in Electronic Design Automation (EDA). Beyond specification…

实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 3Evidence 2Actionability 5

#6 · Total 22/25

Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection

Cláudio Lúcio do Val Lopes, Lucca Machado da Silva · 2026-07-11 · arXiv export API

一句话结论:Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collapse'', defaulting to the majority class and failing to balance anomaly interdiction with customer friction. To overcome th…

实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 4Novelty 4Substance 5Evidence 4Actionability 5

#7 · Total 23/25

Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models

Shravan Murlidaran, Miguel P. Eckstein · 2026-07-11 · arXiv export API

一句话结论:Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple scenes (MS-COCO) that do not showcase complex human interactions or behaviors, only a handful of non-curated human descrip…

实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 3Substance 5Evidence 5Actionability 5

#8 · Total 21/25

Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI

Yuan Cao, Haiqian Yang · 2026-07-11 · arXiv export API

一句话结论:Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they sh…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research tasks. These are powerful capabilities, but they share a structural limitation: the representational frame within which t…

实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 5Evidence 2Actionability 5

#9 · Total 18/25

Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory

Chengkai Zhu, Ziao Tang, Guocheng Zhen, Yimeng Cao, Yusheng Zhao, Ranyiliu Chen · 2026-07-11 · arXiv export API

一句话结论:Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information processing, underpinning quantum communication, computation, and error correctio…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information processing, underpinning quantum communication, computation, and error correction. Formalizing its coding theorems requires connecting finite-block pr…

实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 3Evidence 1Actionability 5

#10 · Total 21/25

PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers

Yujie Pang, Zudong Li · 2026-07-11 · arXiv export API

一句话结论:Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization bu…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often introduce high inference latency and GPU-memory cost, while vi…

实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 4Novelty 3Substance 5Evidence 5Actionability 4

#11 · Total 20/25

Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR

Sanjid Hasan, Md. Abdur Rahman · 2026-07-11 · arXiv export API

一句话结论:Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Beng…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Lightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause of this failure as the model…

实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 4Novelty 4Substance 5Evidence 3Actionability 4

#12 · Total 19/25

Scalable Visual Pretraining for Language Intelligence

Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou · 2026-07-11 · arXiv export API

一句话结论:The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual represent…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms of knowledge are conveyed through visual representations, where figures, typeset equations, and page layouts carry rich…

实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 3Novelty 4Substance 5Evidence 3Actionability 4

#13 · Total 17/25

TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems

Hannah M. Liu, Rhea Saxena, Shiv Asthana · 2026-07-11 · arXiv export API

一句话结论:The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them. In this pape…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them. In this paper, we introduce the TrustX Agent Risk Classification Framework, a stru…

实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 4Novelty 3Substance 4Evidence 1Actionability 5

#14 · Total 20/25

4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception

Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen · 2026-07-11 · arXiv export API

一句话结论:Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordab…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millimeter-wave radar has emerged as a robust and affordable sensor, yet its sparse returns make radar-camera fusion necessary f…

实验/证据:Evidence 5/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 3Novelty 3Substance 5Evidence 5Actionability 4

#15 · Total 16/25

OpenLongTail: Generative Scaling of Long-Tail Driving Data

Lulin Liu, Nuo Chen, Yan Wang, Bangya Liu, Wenyan Cong, Hezhen Hu · 2026-07-11 · arXiv export API

一句话结论:Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-t…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sour…

实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 3Novelty 3Substance 4Evidence 2Actionability 4

版本更新提醒

本次默认不重复收录历史已覆盖论文;未发现摘要层面足以单独列出的重大版本更新。

今日未纳入但可观察论文

数据源失败或不确定性说明

附录:检索式/过滤规则/去重状态摘要

Categories: cs.AI, cs.CL, cs.CV, cs.LG, stat.ML, cs.IR;按 published/updated 转北京时间过滤;去重读取 seen_papers.json 并扫描既有报告。