# AI 论文日报 2026-07-12（覆盖北京时间 2026-07-11、2026-07-12；官方源最新可见批次回退）

- 运行时间：2026-07-12 14:01 BJT / 2026-07-12 06:01 UTC
- candidate count: 599
- new included count: 542
- selected count: 12
- 评分维度：Relevance、Novelty、Substance、Evidence、Actionability（0–5）。

## 执行摘要
本次检查 599 条 arXiv 官方候选。北京时间 7/11–7/12 为周末/发布空窗，目标自然日没有新的 arXiv submitted/updated 记录，因此按最新可见 release 批次做 continuity fallback；历史去重后新论文 542 篇，精选 12 篇。

**Top 3**
1. [CausalDS: Benchmarking Causal Reasoning in Data-Science Agents](https://arxiv.org/abs/2607.08093) — 直接服务代码智能体与自动化开发评测。
2. [Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions](https://arxiv.org/abs/2607.03233) — 对 agent 规划、工具使用或长程任务很有启发。
3. [Reliable and Developer-Aligned Evaluation of Agents for Software Engineering](https://arxiv.org/abs/2607.06713) — 直接服务代码智能体与自动化开发评测。

## Top Picks

### 1. [CausalDS: Benchmarking Causal Reasoning in Data-Science Agents](https://arxiv.org/abs/2607.08093)
- **作者/机构**：Andrej Leban, Yuekai Sun
- **日期**：published(BJT) 2026-07-09；updated(BJT) 2026-07-09
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.08093
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 agent, agents, reasoning, tool, coding, benchmark, evaluation, llm；总分 24/25。
- **方法要点**：Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic causal reasoning benchmarks without realistic data analysis or data analysis benchmarks without a principled causal data-generating structure. Furthermore, existing causal evaluation datasets are often restricted to curated examp…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 5/5；Substance 5/5；Evidence 4/5；Actionability 5/5。

### 2. [Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions](https://arxiv.org/abs/2607.03233)
- **作者/机构**：Eduardo Almeida Palmieri, Mohamed Chahine Ghanem, Dipo Dunsin, Zubair Baig, Ed de Quincey, Kim-Kwang Raymond Choo
- **日期**：published(BJT) 2026-07-03；updated(BJT) 2026-07-03
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.03233
- **一句话结论**：对 agent 规划、工具使用或长程任务很有启发。
- **为什么重要**：核心信号 agent, reasoning, tool, rag, retrieval, benchmark, evaluation, llm；总分 23/25。
- **方法要点**：The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for modern intelligence, cybersecurity, and cyber investigation. Large language models (LLMs) and agentic AI systems, capable of tool use, multi-step reasoning, and iterative intelligence generation, have emerged as promising solutions, yet evaluation frameworks have not kept pace with reported …
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 4/5；Substance 5/5；Evidence 4/5；Actionability 5/5。

### 3. [Reliable and Developer-Aligned Evaluation of Agents for Software Engineering](https://arxiv.org/abs/2607.06713)
- **作者/机构**：Razvan Mihai Popescu
- **日期**：published(BJT) 2026-07-08；updated(BJT) 2026-07-08
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.06713
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 agent, agents, rag, coding, benchmark, evaluation, llm, large language model；总分 23/25。
- **方法要点**：Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments. Despite their accelerated adoption, existing evaluation techniques are limited due to their fragmented nature and distorted projection of true model capabilities, often obtained from hypothetical syntactic scenarios…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 4/5；Substance 5/5；Evidence 4/5；Actionability 5/5。

### 4. [AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation](https://arxiv.org/abs/2607.06624)
- **作者/机构**：Andrey Podivilov, Vadim Lomshakov, Sergey Savin, Matvei Startsev, Roman Pozharskiy, Maksim Parshin 等
- **日期**：published(BJT) 2026-07-07；updated(BJT) 2026-07-07
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.06624
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 agent, agents, tool, coding, code, benchmark, evaluation, llm；总分 23/25。
- **方法要点**：We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs for…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 4/5；Substance 5/5；Evidence 4/5；Actionability 5/5。

### 5. [TF-Engram: A Train-Free Engram with SSD-Backed Memory for Large Language Models](https://arxiv.org/abs/2607.07388)
- **作者/机构**：Yutang Ma, Kecheng Huang, Xikun Jiang, Zili Shao
- **日期**：published(BJT) 2026-07-08；updated(BJT) 2026-07-08
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.07388
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 memory, rag, retrieval, coding, evaluation, llm, large language model, inference；总分 21/25。
- **方法要点**：Large Language Models (LLMs) store factual knowledge and domain-specific patterns implicitly in dense Transformer parameters, making knowledge expansion costly through pretraining, fine-tuning, retrieval augmentation, or longer contexts. Engram-style memory offers a compact hidden-state injection pathway, but existing GPU-resident designs often rely on hash-based compression, causing unrelated phrases to collide in shared slot…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 3/5；Substance 5/5；Evidence 3/5；Actionability 5/5。

### 6. [PERFOPT-Bench: Evaluating Coding Agents on Software Performance Optimization](https://arxiv.org/abs/2607.07744)
- **作者/机构**：Yingyun Cui, Yi Xie, Piaohong Wang, Jiawei Ma, Bo Liu, Liangliang Cao
- **日期**：published(BJT) 2026-07-08；updated(BJT) 2026-07-08
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.07744
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 agent, agents, coding, code, benchmark, evaluation, llm, dataset；总分 24/25。
- **方法要点**：Coding-agent benchmarks have largely measured whether agents can produce functionally correct patches, but production software also demands measurable speedups on real execution targets. Performance optimization is a distinct agentic task: agents must profile executions, diagnose cross-layer bottlenecks, edit code without breaking correctness, and verify that gains are reproducible rather than measurement artifacts. We introdu…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 5/5；Substance 5/5；Evidence 4/5；Actionability 5/5。

### 7. [DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks](https://arxiv.org/abs/2607.07946)
- **作者/机构**：Wenqi Huang, Charley Lee, Leonard Tng, Serena Ge
- **日期**：published(BJT) 2026-07-09；updated(BJT) 2026-07-09
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.07946
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 agent, agents, coding, code, benchmark, evaluation, llm, training；总分 23/25。
- **方法要点**：DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks for evaluating coding agents. Most public agentic coding benchmarks follow SWE-bench in mining merged fixes from public GitHub repositories, which creates two problems: the fixes and their discussion were likely seen during pretraining, so a high score can reflect recall rather than problem-solving; and each task is graded by the tests that shipped…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 4/5；Substance 5/5；Evidence 4/5；Actionability 5/5。

### 8. [Harnessing Code Agents for Automatic Software Verification](https://arxiv.org/abs/2607.06341)
- **作者/机构**：Shuangxiang Kan, Shuanglong Kan, Sebastian Ertel
- **日期**：published(BJT) 2026-07-07；updated(BJT) 2026-07-07
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.06341
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 agent, agents, memory, feedback, rag, code, llm, large language model；总分 21/25。
- **方法要点**：Formal verification offers the strongest guarantee of software correctness, but it does not scale: the proofs demanded by interactive theorem provers such as Coq require enormous expert effort. Large language models (LLMs) promise to generate these proofs automatically, yet existing approaches wire a fixed, human-designed proof strategy into the system and constrain the model to follow it (retrieving premises and predicting ta…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 3/5；Substance 5/5；Evidence 3/5；Actionability 5/5。

### 9. [The Poisoned Chalice of LLM Evaluation Report](https://arxiv.org/abs/2607.07481)
- **作者/机构**：Jonathan Katzy, Ali Al-Kaswan, Razvan Mihai Popescu, Zhou Yang
- **日期**：published(BJT) 2026-07-08；updated(BJT) 2026-07-08
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.07481
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 rag, code, benchmark, evaluation, llm, large language model, inference, training；总分 23/25。
- **方法要点**：Large language models are increasingly used to evaluate and support software engineering tasks, yet the validity of these evaluations is often undermined by uncertainty about whether benchmark instances were seen during pretraining. This can lead to data contamination, which may inflate performance and result in misleading conclusions about model capability. Despite this, the training corpora of many modern models are only par…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 5/5；Substance 4/5；Evidence 4/5；Actionability 5/5。

### 10. [Detecting Vulnerability-Inducing Commits via Multi-Stage Reasoning with LLM-Based Agents](https://arxiv.org/abs/2607.05772)
- **作者/机构**：Liyou Chen, Hailong Sun, Xiang Gao, Yue Pan
- **日期**：published(BJT) 2026-07-07；updated(BJT) 2026-07-07
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.05772
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 agent, agents, reasoning, rag, code, llm, dataset, workflow；总分 23/25。
- **方法要点**：Detecting vulnerability-inducing commits (VICs) at submission time is critical for improving the security and reliability of software systems. However, this task is highly challenging because it requires reasoning about the semantic impact of code changes from heterogeneous information sources, including code diffs, commit messages, and the surrounding contextual code. Existing approaches often struggle to fully capture these …
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 4/5；Substance 5/5；Evidence 4/5；Actionability 5/5。

### 11. [TTHE: Test-Time Harness Evolution](https://arxiv.org/abs/2607.08124)
- **作者/机构**：Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo 等
- **日期**：published(BJT) 2026-07-09；updated(BJT) 2026-07-09
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.08124
- **一句话结论**：直接服务代码智能体与自动化开发评测。
- **为什么重要**：核心信号 agent, agents, tool, coding, evaluation, llm, training, workflow；总分 21/25。
- **方法要点**：The behavior of an LLM agent is determined not only by the underlying model, but also by its harness: the executable program that constructs context, invokes tools, verifies intermediate results, and recovers from failures. Existing approaches optimize such harnesses before deployment, searching training or development data for a fixed agent workflow that is then frozen at test time. This limits adaptation when the test distri…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 3/5；Substance 5/5；Evidence 3/5；Actionability 5/5。

### 12. [AgentTether: Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operation](https://arxiv.org/abs/2607.06273)
- **作者/机构**：Chenyu Zhao, Shenglin Zhang, Wenwei Gu, Yongqian Sun, Dan Pei, Chetan Bansal 等
- **日期**：published(BJT) 2026-07-07；updated(BJT) 2026-07-07
- **来源**：arXiv export latest-release fallback；PDF: https://arxiv.org/pdf/2607.06273
- **一句话结论**：对 agent 规划、工具使用或长程任务很有启发。
- **为什么重要**：核心信号 agent, agents, tool, memory, feedback, llm, large language model, training；总分 20/25。
- **方法要点**：Large language model (LLM) agents are increasingly used for multi-step, stateful tool-use tasks, yet production reliability remains limited. Unlike static software repair, agent repair must recover dynamic trajectories whose early decisions can propagate into later errors and external state changes. Existing automatic remedies address only part of this problem: blind retry adds no diagnosis, outcome feedback says whether a run…
- **实验/证据**：基于摘要中的 benchmark / dataset / experiment / result 等线索给 Evidence 分；建议阅读全文确认 baseline、ablation 与失败案例。
- **局限/风险**：本日报为元数据与摘要级快速筛选；fallback 条目不声称目标日首次提交。
- **Lucian 可采取的下一步**：抽取任务定义、数据构造、指标和 baseline，若贴合 agent/coding/RAG/eval 则做 1 周小复现。
- **评分表**：Relevance 5/5；Novelty 3/5；Substance 5/5；Evidence 2/5；Actionability 5/5。

## 版本更新提醒
未发现此前已覆盖论文的重大版本更新需要单独提醒。

## 今日未纳入但可观察论文
- [Field Order Should Not Matter: Permutation-Invariant Embedding Model Fine-Tuning for Structured Metadata Retrieval](https://arxiv.org/abs/2606.30473)：总分 23/25，相关但优先级低于 Top Picks。
- [RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications](https://arxiv.org/abs/2607.06411)：总分 23/25，相关但优先级低于 Top Picks。
- [Latent Programming Horizons in Coding Agents](https://arxiv.org/abs/2607.05188)：总分 23/25，相关但优先级低于 Top Picks。
- [Token-Flow Firewall: Semantic Runtime Auditing for Persistent AI Agents](https://arxiv.org/abs/2607.08395)：总分 22/25，相关但优先级低于 Top Picks。
- [ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping](https://arxiv.org/abs/2606.31693)：总分 20/25，相关但优先级低于 Top Picks。
- [Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark](https://arxiv.org/abs/2607.05443)：总分 24/25，相关但优先级低于 Top Picks。
- [From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier](https://arxiv.org/abs/2607.07779)：总分 23/25，相关但优先级低于 Top Picks。
- [LBR: Towards Mitigating Length Bias in Large Language Models for Recommendation](https://arxiv.org/abs/2607.04270)：总分 23/25，相关但优先级低于 Top Picks。

## 数据源失败或不确定性说明
- 未观察到阻断性 arXiv 官方源失败。
- Hugging Face Daily Papers 页面在本运行环境偶发网络不可达；本次以 arXiv 官方 export 为主。
- arXiv export 在 BJT 2026-07-11/12 没有新 submitted/updated 记录，本报告使用最新可见 release 批次作 continuity fallback。

## 附录：检索式 / 过滤 / 去重
- 类别：cs.AI, cs.CL, cs.CV, cs.LG, stat.ML, cs.SE, cs.IR；每类 submittedDate descending 前 100 条。
- 去重：seen_papers.json + 历史报告；history IDs 404；duplicates 57。
