FJAIResearch StudioResearch online

AI Paper Daily

AI 论文日报 · 2026-07-21

覆盖北京时间日期:2026-07-20、2026-07-21。聚焦 Agent、LLM reasoning/planning/tool use/memory、RAG、coding agents、evaluation/benchmark 与训练/推理基础设施。

执行摘要

candidate count64
new included count64
selected count15

Top 3:Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation;FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications;SEE: Structure-aware Exploring \& Exploiting for Long-horizon GUI Agent Trajectory Synthesis

Top Picks

#1 · Total 22/25

Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation

Siddharth Mishra-Sharma · 2026-07-20 · arXiv export API

一句话结论:Neural simulation-based inference enables parameter estimation for complex models, but typically requires the user to specify a simulator encoding a fixed model structure. We present a frame…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Neural simulation-based inference enables parameter estimation for complex models, but typically requires the user to specify a simulator encoding a fixed model structure. We present a framework for joint model selection and parameter estimation that combines…

实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 5Evidence 3Actionability 5

#2 · Total 20/25

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu, Atri Rudra, Beidi Chen · 2026-07-21 · arXiv export API

一句话结论:Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specif…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism.…

实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 4Evidence 2Actionability 5

#3 · Total 22/25

SEE: Structure-aware Exploring \& Exploiting for Long-horizon GUI Agent Trajectory Synthesis

Zhuohang Fan, Beichen Zhang, Yuanfa Li, Changqiao Wu, Wei Liu, Jian Luan · 2026-07-20 · arXiv export API

一句话结论:Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-h…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Graphical User Interface (GUI) agents powered by vision-language models hold promise for automating real-world mobile tasks. However, progress is limited by the lack of high-coverage, long-horizon interaction trajectories collected from element-rich and rapidl…

实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 4Evidence 4Actionability 5

#4 · Total 20/25

Planning with Transformers: Chain of Computation and Structured Context Windows

Ehsan Futuhi, Nathan R. Sturtevant · 2026-07-20 · arXiv export API

一句话结论:Large Language Models (LLMs) have had a remarkable impact across many areas of machine learning. However, recent studies have shown that they struggle to reliably solve planning problems. At…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Large Language Models (LLMs) have had a remarkable impact across many areas of machine learning. However, recent studies have shown that they struggle to reliably solve planning problems. At the same time, theoretical results have shown that transformers, the…

实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 3Substance 5Evidence 2Actionability 5

#5 · Total 17/25

Theoretical Foundations of $\max$@$k$ Reinforcement Learning

Riccardo Poiani, Martino Bernasconi, Andrea Celli · 2026-07-20 · arXiv export API

一句话结论:Reinforcement Learning is a cornerstone technique for modern large reasoning models. Usually, for difficult tasks such as code generation and theorem proving, the agent is evaluated by gener…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Reinforcement Learning is a cornerstone technique for modern large reasoning models. Usually, for difficult tasks such as code generation and theorem proving, the agent is evaluated by generating $K$ responses rather than sampling a single response, and perfor…

实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 4Novelty 4Substance 4Evidence 2Actionability 3

#6 · Total 21/25

Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication Loss

Kemal Devrim Kafadar, Eren Özaltun, Mahmud Efnan Şanlı, Feyza Orak, Emirhan Gazi, Kubilay Kağan Kömürcü · 2026-07-20 · arXiv export API

一句话结论:Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments. To maintain op…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Robust multi-agent coordination relies heavily on inter-agent communication, which is frequently disrupted by physical and environmental constraints in real-world deployments. To maintain operation during these intermittent communication failures, agents can e…

实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 4Substance 4Evidence 3Actionability 5

#7 · Total 22/25

EVOLVE: Efficient Learned Volume Compression with Variable-Rate Encoding on a Cross-Domain Database

Kaiyuan Tang, Maizhe Yang, Chaoli Wang · 2026-07-21 · arXiv export API

一句话结论:Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical. How…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Large-scale scientific simulations generate volumetric data at rates that far outpace advances in storage and network bandwidth, making effective lossy compression increasingly critical. However, conventional compressors often struggle to preserve fine structu…

实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 4Novelty 4Substance 5Evidence 4Actionability 5

#8 · Total 16/25

Rethinking the Suitability of Reinforcement Learning Algorithms Under Practical Transfer Constraints

Hany Hamed, Abhishek Naik, Colin Bellinger, A. Rupam Mahmood · 2026-07-20 · arXiv export API

一句话结论:Transfer-oriented reinforcement learning requires evaluating algorithms along dimensions that go beyond standard sample efficiency. We focus on two dimensions: practical efficiency, which as…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Transfer-oriented reinforcement learning requires evaluating algorithms along dimensions that go beyond standard sample efficiency. We focus on two dimensions: practical efficiency, which asks whether conclusions about algorithm suitability change under wall-c…

实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 4Novelty 4Substance 4Evidence 1Actionability 3

#9 · Total 19/25

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks

Damien Teney, Liangze Jiang, Hemanth Saratchandran, Simon Lucey · 2026-07-20 · arXiv export API

一句话结论:Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be key for p…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be key for pushing AI beyond merely scaling current designs. *Method.* We present…

实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 3Novelty 3Substance 5Evidence 4Actionability 4

#10 · Total 14/25

Patch Policy: Efficient Embodied Control via Dense Visual Representations

Gaoyue Zhou, Zichen Jeff Cui, Ada Langford, Bowen Tan, Yann LeCun, Lerrel Pinto · 2026-07-21 · arXiv export API

一句话结论:Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a sin…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Pretrained dense visual features from Vision Transformers (ViTs) are powerful yet have been underutilized in robot learning. Modern robot policies either compress each observation into a single global token, or rely on visual backbones trained from scratch, sa…

实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 3Novelty 3Substance 3Evidence 1Actionability 4

#11 · Total 18/25

Distributional Soft Bellman Operator under the Cramér Geometry

Keru Wang, Yixin Deng, Yao Lyu, Stephen Redmond, Shengbo Eben Li · 2026-07-20 · arXiv export API

一句话结论:Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evalua…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting…

实验/证据:Evidence 3/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 3Novelty 4Substance 5Evidence 3Actionability 3

#12 · Total 18/25

CORAL: Learning Amyloid Fibril Ligand Docking with Cooperative Binding Rewards

Yasheng Sun, Bohan Li, Youqi Tao, Jürgen Schmidhuber · 2026-07-20 · arXiv export API

一句话结论:A hallmark of neurodegenerative diseases such as Alzheimer's and Parkinson's is the aberrant aggregation of proteins into amyloid fibrils, and small molecules that selectively bind to these…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:A hallmark of neurodegenerative diseases such as Alzheimer's and Parkinson's is the aberrant aggregation of proteins into amyloid fibrils, and small molecules that selectively bind to these fibrils hold promise as diagnostics, imaging probes, and therapeutics.…

实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 3Novelty 3Substance 5Evidence 4Actionability 3

#13 · Total 16/25

FailureAtlas: A Taxonomy of Failure Modes in Multi-Provider LLM Serving Infrastructure

Vishal Pandey, Gopal Singh · 2026-07-20 · arXiv export API

一句话结论:Multi-provider LLM gateways reverse proxies that route, load-balance, and rate-limit requests across foundation-model APIs have become critical production infrastructure. Yet the failure mod…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Multi-provider LLM gateways reverse proxies that route, load-balance, and rate-limit requests across foundation-model APIs have become critical production infrastructure. Yet the failure modes specific to this architectural layer remain undocumented, scattered…

实验/证据:Evidence 2/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 3Novelty 3Substance 5Evidence 2Actionability 3

#14 · Total 17/25

Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy Optimization

Zijian Zhao, Sen Li · 2026-07-20 · arXiv export API

一句话结论:Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL). A central design question is how many neighboring…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Multi-agent policy optimization, exemplified by PPO-based methods, is a key branch of cooperative Multi-Agent Reinforcement Learning (MARL). A central design question is how many neighboring agents\footnote{In this paper, "neighbors" refer not only to physical…

实验/证据:Evidence 1/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 5Novelty 3Substance 3Evidence 1Actionability 5

#15 · Total 22/25

FlashPDE: A Drop-in Fused Triton Operator Library for Neural PDE Solvers

Peiyu Zang, Bosen Xie, Ruoxiang Xu, Yongqiang Cai · 2026-07-20 · arXiv export API

一句话结论:Physics-Informed Neural Networks (PINNs) solve PDEs by incorporating physical constraints into neural-network training, but large-scale problems are limited by automatic-differentiation memo…

为什么重要:贴近 Agent / LLM reasoning / coding / evaluation / personalization 研究线,适合快速转成复现实验或产品验证。

方法要点:Physics-Informed Neural Networks (PINNs) solve PDEs by incorporating physical constraints into neural-network training, but large-scale problems are limited by automatic-differentiation memory overhead and inefficient execution of grid-based PDE operators. We…

实验/证据:Evidence 4/5;需全文核验 benchmark、baseline、ablation 与代码可得性。

局限/风险:快筛基于官方元数据/摘要,结论强度以论文全文为准。

Lucian 下一步:抽取任务定义、指标与 baseline,加入 Auto Research 阅读/复现实验队列。

Relevance 4Novelty 4Substance 5Evidence 4Actionability 5

版本更新提醒

本次默认不重复收录历史已覆盖论文;未发现摘要层面足以单独列出的重大版本更新。

今日未纳入但可观察论文

数据源失败或不确定性说明

附录:检索式/过滤规则/去重状态摘要

Categories: cs.AI, cs.CL, cs.CV, cs.LG, stat.ML, cs.IR;按 published/updated 转北京时间过滤;去重读取 seen_papers.json 并扫描既有报告。