Agent 研究论文与资料索引¶
本页由 docs/research-manifest.json 自动生成;论文元数据只在统一 manifest
维护。背景、架构、公式、原文效果和本地实验请进入独立详情页。
已实现论文与资料¶
| 方向 | 方法 | 一作机构与日期 | 原作者代码 | 本地入口 |
|---|---|---|---|---|
| 编程 Agent | An Empirical Study of Harness Design for Coding Agents | University of Massachusetts Amherst,2026-09-17 | 未发现官方代码 | harness-design-study |
| 多 Agent | CERA-MoA: Co-Evolving Router and Agents for Mixture-of-Agents | IIIS, Tsinghua University,2026-09-16 | 未发现官方代码 | cera-moa |
| 轨迹精炼 | Dependency-Aware Trajectory Refinement for Efficient Multi-Turn Agent Fine-Tuning | ShanghaiTech University,2026-09-16 | 已开源 | dependency-refinement |
| 长期记忆 | Interactive Memory Learning for Long-Term Conversations | Harbin Institute of Technology, Shenzhen / Pengcheng Laboratory,2026-09-15 | 未发现官方代码 | interactive-memory |
| GUI Agent | Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents | Zhejiang University,2026-09-15 | 已开源 | evoskill-gui |
| 编程 Agent | RepoAtlas: Guiding Coding Agents via Evolving Multimodal Repository Views | Beihang University,2026-09-15 | 未发现官方代码 | repoatlas |
| 科学 Agent | ScienceBuddy: Recursive-in-Recursive Self-Improvement for Interactive Scientific Agents | Gen-Verse research collaboration,2026-09-15 | 已开源 | sciencebuddy |
| Agent 评测 | Verifiable Social Reasoning for LLM Assistants | Google Research / Hebrew University of Jerusalem / University of Cambridge,2026-09-15 | 未发现官方代码 | fuse-evaluator |
| Agent RL | HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning | Harbin Institute of Technology / Alibaba Cloud,2026-09-12 | 未发现官方代码 | harness-bandit |
| 技能进化 | COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization | The Chinese University of Hong Kong, Shenzhen,2026-09-10 | 已开源 | cobra-skills |
| Harness 自进化 | Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents | Chengdu Institute of Computer Applications, Chinese Academy of Sciences,2026-09-10 | 已开源 | ecdysis |
| 持久化优化规划 | MAPLE: Memory-Augmented Planning with Language and Evolution | Harbin Institute of Technology, Shenzhen,2026-09-10 | 已开源 | maple |
| Agent 记忆 | Grounding Agent Memory: Environment-Probing Curation for Enterprise Agents | Microsoft,2026-09-09 | 未发现官方代码 | grounded-memory |
| 搜索 Agent | SearchAtlas: Analyzing Agentic Search Strategies via Evidential Query Graphs | Duke University,2026-09-09 | 已开源 | searchatlas |
| Agent RL | T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks | National University of Singapore / Tencent,2026-09-09 | 未发现官方代码 | t1-terminal-rl |
| 技能检索 | When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents | Manulife,2026-09-09 | 已开源 | skill-retention |
| 环境反馈脚手架 | Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks | Fudan University,2026-09-08 | 已开源 | feedback-scaffold |
| 事件树记忆压缩 | MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging | Shanghai Jiao Tong University,2026-09-08 | 已开源 | memforest |
| 自进化程序图 | Procedural Graphs: Self-Evolving Execution Structures for LLM Agents | Google,2026-09-08 | 未发现官方代码 | procedural-graphs |
| Agentic recommendation memory | AtomRec: Evolving Atomic Memory for Agentic Recommendation | Xi'an Jiaotong-Liverpool University,2026-09-04 | 未发现官方代码 | atomrec |
| Hierarchical skill coevolution | CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution | Institute of Automation, Chinese Academy of Sciences,2026-09-04 | 已开源 | coskill |
| Structure-preserving verifier and reward | SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents | Institute of Science Tokyo,2026-09-03 | 未发现官方代码 | silr |
| Multi-harness RL audit | What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding Agents | New York University,2026-09-03 | 未发现官方代码 | multi-harness-rl |
| 论文到代码状态子规划 | DeepRepro: State-Aware Subplanning for Paper-to-Code Reproduction in Evolving Repositories | Institute of Computing Technology, Chinese Academy of Sciences,2026-08-27 | 已开源 | deeprepro |
| 红队技能进化 | RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution | City University of Hong Kong,2026-08-27 | 未发现官方代码 | redevoagent |
| Agent 技能预训练 | SPT: Skills as Pre-Training Data for Agentic Language Models | Beijing University of Posts and Telecommunications,2026-08-27 | 未发现官方代码 | spt |
| 软件轨迹筛选 | SWE-Prime: Fewer Trajectories, Better Performance | Sun Yat-sen University,2026-08-27 | 未发现官方代码 | swe-prime |
| 行为相关 Harness 验证 | Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification | Fudan University,2026-08-27 | 已开源 | harnesslens |
| Agent 数据质量 | What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents | Huawei Technologies Co., Ltd,2026-08-27 | 未发现官方代码 | ace-data |
| 可训练协同记忆 | When Memory Takes Gradients: Collaborative Vector Memory for Agentic Recommender Systems | Shenzhen Technology University,2026-08-27 | 未发现官方代码 | covemem |
| 视频研究自适应工具与反思 | AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research | Accio Team, Alibaba Group,2026-08-26 | 已开源 | adavdr |
| 反事实因果技能图 | CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval | Jilin University / Ant Group,2026-08-26 | 已开源 | caskg |
| Harness 即时生成与进化 | JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution | LV-NUS Lab,2026-08-26 | 已开源 | jit-agent |
| 在线进展成本路由 | ProgRouter: Online Progress-Guided Orchestration for Multi-Agent LLM Workflows under Quality-Cost Tradeoffs | Aston University,2026-08-26 | 未发现官方代码 | progrouter |
| 工作流 Prefix-State 调度 | TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving | University of Science and Technology of China,2026-08-26 | 未发现官方代码 | topas |
| ML 开发轨迹规划 | TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development | Carnegie Mellon University,2026-08-26 | 已开源 | traceml |
| 环境反馈提示训练 | AHEAD: Agentic Hints for Effective Agent Development | AWS AI Labs / Purdue University,2026-08-25 | 未发现官方代码 | ahead |
| 可验证技能锻造 | SkillForge: Automated Skill Discovery and Refinement for Tool-Using Agents | AMAP, Alibaba Group,2026-08-25 | 未发现官方代码 | skillforge |
| 多维验证工具自进化 | SMITH: Self-Improving Tool-Using Agents through Multi-Aspect Verification | Appier AI Research / National Taiwan University,2026-08-25 | 已开源 | smith |
| 异步单流 Agent RL | SPO++: Stabilizing Asynchronous Agentic Reinforcement Learning via Measure-Theoretic Token Correction | Renmin University of China,2026-08-25 | 未发现官方代码 | spo-plus-plus |
| 高斯 guidance Agent RL | Agent-G²: Gaussian Guidance for Agentic Reinforcement Learning | Baidu,2026-08-24 | 已开源 | agent-g2 |
| Harness 自动优化 | AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces | Microsoft / POSTECH,2026-08-24 | 未发现官方代码 | autosaddler |
| 动作级技能优化 | AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization | University of Science and Technology of China,2026-08-21 | 已开源 | auso |
| Agentic RL | SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning | Xiamen University,2026-08-20 | 未发现官方代码 | sapo |
| Agentic RL | RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training | Adelaide University,2026-08-19 | 未发现官方代码 | rtpo |
| Agentic RL | SPADE: Self-Play in Adaptive Synthetic Executable Environments | University of Washington,2026-08-19 | 已开源 | spade |
| Agentic RL | PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs | Xiamen University,2026-08-18 | 未发现官方代码 | planpo |
| Agentic RL | TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents | Authors did not disclose affiliation,2026-08-17 | 未发现官方代码 | trca |
| 记忆 | HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation | Institute of Automation, Chinese Academy of Sciences,2026-08-16 | 未发现官方代码 | hymem |
| 反思 | LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation | Authors did not disclose affiliation,2026-08-12 | 未发现官方代码 | loongreflect |
| Agentic RL / efficient long context | Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks | Capital One AI Foundations,2026-08-11 | 未发现官方代码 | sinkflex-rl |
| 自进化 | OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks | Tsinghua University,2026-08-10 | 已开源 | openloopevolve |
| 代码 Agent | Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution | Vanderbilt University,2026-08-07 | 未发现官方代码 | pmcoder |
| 递归 turn 信用 | AgentOPSD | Tsinghua University,2026-08-06 | 已开源 | agent-opsd |
| 代码检索 Agent | CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents | NetEase Guangzhou AI Lab,2026-08-06 | 未发现官方代码 | codegrep |
| 搜索 Agent RL | Contextual Information Policy Optimization for Search Agents | Beihang University,2026-08-06 | 未发现官方代码 | cipo |
| 环境 rehearsal / Agent RL | EnvACE | Shanghai Jiao Tong,2026-08-06 | 已开源 | envace |
| Harness 优化评测 | HarnessOpt-Bench: Evaluating LLMs at Harness Optimization | Scale AI,2026-08-06 | 未发现官方代码 | harnessopt-bench |
| 全局技能进化 | Learning Globally Reusable Skills for Coding Agents | Tianjin University,2026-08-06 | 未发现官方代码 | gse |
| 技能进化安全 | When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents | 论文未列机构,2026-08-06 | 未发现官方代码 | vag |
| Harness policy RL | EvoHarness-RL: Learning Self-Evolving Runtime Harness for Long-Horizon LLM Agents | University of Illinois Urbana–Champaign / Meta AI,2026-08-05 | 未发现官方代码 | evoharness-rl |
| 端到端 Agent 记忆 | MemoryCPT: An End-to-End Agent Memory Framework for Cost-Performance Trade-off | Hong Kong University of Science and Technology / Tencent LIGHTSPEED STUDIOS,2026-08-05 | 未发现官方代码 | memorycpt |
| 观测校准蒸馏 | OCSD | Nanjing University,2026-08-05 | 已开源 | ocsd |
| 环境派生中训练 | State2State: Environment-Derived Mid-Training for LLM Agents | Tsinghua University AIR / Alibaba Group,2026-08-05 | 已开源 | state2state |
| 工具规划 | ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning | Authors did not disclose affiliation,2026-08-04 | 未发现官方代码 | toollift |
| 可验证统一记忆 | VerMem | Sun Yat-sen University,2026-08-04 | 已开源 | vermem |
| 检索—记忆共进化 | CoEvo-Mem | 论文未列机构,2026-08-03 | 未发现官方代码 | coevo-mem |
| 搜索轨迹 hindsight | HindSearch: Trajectory-Level Hindsight Critique for Search-Augmented Reinforcement Learning | Santa Clara University,2026-08-03 | 已开源 | hindsearch |
| 工具规划 | HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents | University of New South Wales,2026-07-31 | 未发现官方代码 | hyperagent |
| Agent group credit | Group-Reflective Self-Distillation | Baidu Inc.,2026-07-30 | 已开源 | grsd |
| 多 Agent | MANTA: Multi-Agent Network Topology Adaptation for Self-Evolving Multi-Agent Systems | Cornell University,2026-07-30 | 未发现官方代码 | manta |
| CUA Reward 评测 | OSReward / OS-Shepherd | The University of Hong Kong,2026-07-30 | 已开源 | os-shepherd |
| Agent RL | TAPO | Peking University,2026-07-30 | 未发现官方代码 | tapo |
| 成本感知工具停止 | CAM-DF | Peking University,2026-07-29 | 未发现官方代码 | cam-df |
| 跨任务技能进化 | SkillRise | Zhejiang University,2026-07-29 | 已开源 | skillrise |
| Agentic RL / turn-level credit | CAST | University of Science and Technology of China,2026-07-28 | 已开源 | cast |
| Hierarchical skill memory | HiSkill | Beijing University of Posts and Telecommunications,2026-07-28 | 已开源 | hiskill |
| Continual agent memory | UniMem | CASIA,2026-07-28 | 未发现官方代码 | unimem |
| Agentic RL / hindsight skill | SEED | Tsinghua University,2026-07-16 | 已开源 | seed |
| Agentic OPD / rollout budgeting | TurnOPD | Academic author team,2026-07-07 | 未发现官方代码 | turn-opd |
| 研究自动化 | AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems | Kuaishou,2026-06-26 | 未发现官方代码 | agentx |
| 训练服务优化 | PROMPTS: Performance Optimization via Multi-Agent Planning for LLM Training and Serving | University of Maryland / Google,2026-05-18 | 未发现官方代码 | prompts |
| Step-aligned Agent RL | StepPO | University of Science and Technology of China,2026-04-20 | 未发现官方代码 | steppo |
| 策略—工具图共进化 | SEARL | Shanghai AI Laboratory,2026-04-09 | 未发现官方代码 | searl |
| 技能设计 | Memento-Skills | Memento Team,2026-03-19 | 已开源 | memento-skills |
| 主动记忆 | U-Mem | National University of Singapore,2026-02-25 | 已开源 | u-mem |
| 记忆技能 | MemSkill | Nanyang Technological University,2026-02-02 | 已开源 | memskill |
| 规划强化学习 | PEARL | 中国科学院信息工程研究所,2026-01-28 | 未发现官方代码 | pearl |
| 技能库强化学习 | SAGE | University of Wisconsin–Madison,2025-12-18 | 未发现官方代码 | sage |
| 零数据多 Agent | Agent0 | University of North Carolina at Chapel Hill,2025-11-20 | 未发现官方代码 | agent0 |
| Agentic RL 基础设施 | Agent-R1 | University of Science and Technology of China,2025-11-18 | 已开源 | agent-r1 |
| 过程记忆 | LEGOMem | Microsoft Research,2025-10-06 | 未发现官方代码 | legomem |
| 多轮用户 Agent RL | MUA-RL | Meituan,2025-08-26 | 已开源 | mua-rl |
| 工具学习 | ToolGrad: Efficient Tool-Use Dataset Generation with Textual Gradients | Google,2025-08-06 | 已开源 | toolgrad |
| Agent RL | Agent Lightning | Microsoft Research,2025-08-05 | 已开源 | agent-lightning |
| 工具记忆 | MemTool | PricewaterhouseCoopers Commercial Technology and Innovation Office,2025-07-29 | 未发现官方代码 | memtool |
| 网页 Agent RL | WebAgent-R1 | University of Virginia,2025-05-22 | 已开源 | webagent-r1 |
| Agent group credit | GiGPO | Nanyang Technological University,2025-05-16 | 未发现官方代码 | gigpo |
| 多轮 Agent RL | RAGEN | Northwestern,2025-04-24 | 已开源 | ragen |
| 工具强化学习 | ToolRL | University of Illinois Urbana-Champaign,2025-04-16 | 未发现官方代码 | toolrl |
| 推理中工具调用 | ReTool | ByteDance Seed,2025-04-15 | 未发现官方代码 | retool |
| 深度研究 RL | DeepResearcher | HKU,2025-04-04 | 已开源 | deepresearcher |
| 搜索 Agent RL | Search-R1 | University of Illinois Urbana-Champaign,2025-03-12 | 已开源 | search-r1 |
| 长时程 Agent RL | LOOP | Apple,2025-02-03 | 已开源 | loop |
| 通用软件 Agent | OpenHands | All-Hands-AI,2024-07-23 | 已开源 | openhands |
| 软件工程 ACI | SWE-agent | Princeton University,2024-05-06 | 已开源 | swe-agent |
| 通用 Agent 评测 | GAIA | Meta AI,2023-11-21 | 已开源 | gaia |
| 虚拟上下文 | MemGPT | University of California, Berkeley,2023-10-12 | 已开源 | memgpt |
| Agent 搜索 | LATS | University of Illinois Urbana-Champaign,2023-10-06 | 已开源 | lats |
| 多 Agent | AutoGen | Microsoft Research,2023-08-16 | 已开源 | autogen |
| 多 Agent 软件工程 | MetaGPT | DeepWisdom,2023-08-01 | 已开源 | metagpt |
| 解耦规划 | ReWOO | Microsoft Research,2023-05-29 | 已开源 | rewoo |
| 工具指令与评测 | ToolBench | Tsinghua University,2023-05-25 | 未发现官方代码 | toolbench |
| 终身学习 | Voyager | NVIDIA,2023-05-25 | 已开源 | voyager |
| 工具反馈 | CRITIC | Microsoft,2023-05-19 | 已开源 | critic |
| 推理搜索 | Tree of Thoughts | Princeton University,2023-05-17 | 已开源 | tree-of-thoughts |
| 记忆与反思 | Generative Agents | Stanford University,2023-04-07 | 已开源 | generative-agents |
| 多 Agent 协作 | CAMEL | King Abdullah University of Science and Technology,2023-03-31 | 已开源 | camel |
| 专家模型编排 | HuggingGPT | Zhejiang University,2023-03-30 | 已开源 | hugginggpt |
| 自我迭代 | Self-Refine | Carnegie Mellon University,2023-03-30 | 已开源 | self-refine |
| 自我反思 | Reflexion | Northeastern University,2023-03-20 | 已开源 | reflexion |
| 自动工具推理 | ART | University of Washington,2023-03-16 | 已开源 | art |
| 工具学习 | Toolformer | Meta AI Research,2023-02-09 | 未发现官方代码 | toolformer |
| 程序推理 | PAL | Carnegie Mellon University,2022-11-18 | 已开源 | pal |
| 推理与行动 | ReAct | Princeton University,2022-10-06 | 已开源 | react |
| 神经符号路由 | MRKL | AI21 Labs,2022-05-01 | 未发现官方代码 | mrkl |
| 具身规划 | SayCan | Robotics at Google,2022-04-04 | 已开源 | saycan |
| 浏览问答 | WebGPT | OpenAI,2021-12-17 | 已开源 | webgpt |
分类浏览: