AgentScout Logo Agent Scout

ArXiv cs.AI Weekly Papers — Week of June 4, 2026: Self-Evolving Agents and Multi-Agent Governance

31 papers collected this week with 25 agent-related papers (81%). Key trends: self-evolving agent frameworks surge (EvoDS, SkillPyramid, EvoDrive), LAP protocol fills agent-to-instrument gap, and domain benchmarks expose frontier model limitations.

AgentScout ·
#arxiv #ai-agents #papers #weekly-tracker #self-evolving-agents #multi-agent-systems
Analyzing Data Nodes...
SIG_CONF:CALCULATING
Verified Sources

Data Overview

  • Snapshot Week: 2026-05-28 to 2026-06-04
  • Tracker: ArXiv cs.AI Weekly Papers (view all snapshots: /tech/ai-agents/data/?tracker=arxiv-cs-ai-weekly)
  • Update Frequency: Weekly
  • Primary Sources: ArXiv cs.AI, ArXiv cs.CL

Key Facts

  • Who: 31 papers collected from ArXiv cs.AI and cs.CL categories
  • What: 25 agent-related papers (81%), including 12 multi-agent papers and 5 self-evolving agent frameworks
  • When: Week of May 28 - June 4, 2026
  • Impact: 3 new benchmarks, 1 new protocol (LAP), 7 papers with venue acceptance

🔺 Scout Intel: What Others Missed

Confidence: high | Novelty Score: 65/100

Three self-evolving agent papers (EvoDS, SkillPyramid, EvoDrive) appear in the same week, signaling a shift from static agent architectures toward autonomous skill acquisition. LAP protocol addresses a gap most coverage ignores: agent-to-instrument communication. While MCP handles model-to-tool and A2A handles agent-to-agent, LAP targets the physical instrument edge critical for autonomous scientific research. Hedge-Bench’s <16% frontier model performance on real hedge fund tasks exposes the gap between benchmark success and professional domain competence.

Key Implication: Agent frameworks are entering a consolidation phase where autonomous skill acquisition and standardized protocols replace manual prompt engineering. The 40% concentration on self-evolving systems suggests the field recognizes current limitations of static agent capabilities.

This Week’s Papers

#TitleArXiv IDTrendVenue/Improvement
1EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management2606.0384110KDD 2026, +28.9% over SOTA
2SkillPyramid: Hierarchical Skill Consolidation for Self-Evolving Agents2606.036929+38.0% reward, -27.7% steps
3LAP: Agent-to-Instrument Protocol for Autonomous Science2606.037559NEW protocol
4GAIATrace + Vidur-Agent: Multi-Model Agentic AI Systems Characterization2606.017258GAIATrace dataset, Vidur-Agent simulator
5Unified Context Evolution for LLM Agents2606.023048ALFWorld: 75.4% → 96.3%
6EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving2606.036788Self-improving LLM agents
7Hedge-Bench: Benchmarking Agents on Financial Reasoning Tasks2606.039187102 tasks, frontier <16%
8NovelAPIBench: Diagnosing Knowledge Gaps in LLM Tool Use2606.0365771.9K tasks, 5 domains
9Uncertainty-Aware Clarification with Information Gain2606.031357ICML 2026, +3.7% success rate
10Agentic CLEAR: Multi-Level Evaluation of LLM Agents2605.226087ACL

Self-Evolving Agent Frameworks

EvoDS (2606.03841) — Zherui Yang, Fan Liu, Yansong Ning, Hao Liu — KDD 2026

  • Focus: Autonomous data science with skill learning and adaptive context compression
  • Key Innovation: Self-evolving framework that acquires skills without manual intervention
  • Performance: +28.9% over SOTA on data science benchmarks

SkillPyramid (2606.03692) — Yuan Xiong et al.

  • Focus: Hierarchical skill consolidation for reusable experience
  • Key Innovation: Multi-level skill hierarchy enabling composition and reuse
  • Performance: +38.0% reward improvement, -27.7% steps on ALFWorld and WebShop

Unified Context Evolution (2606.02304) — Zixuan Zhu et al.

  • Focus: Gradient-free framework externalizing agent experience
  • Key Innovation: Typed Evolvable Context Units for memory management
  • Performance: ALFWorld 75.4% → 96.3%, WebShop 45.1% → 61.3%

EvoDrive (2606.03678) — Tong Nie et al.

  • Focus: Safety-critical autonomous driving scenario generation
  • Key Innovation: Pareto evolution via self-improving LLM agents
  • Domain: Autonomous driving

Multi-Agent Systems & Governance

LAP Protocol (2606.03755) — Linwu Zhu et al.

  • Type: Agent-to-Instrument Protocol
  • Gap Filled: Complements MCP (model-to-tool) and A2A (agent-to-agent)
  • Use Case: Autonomous scientific instruments

GAIATrace + Vidur-Agent (2606.01725) — Donghwan Kim et al.

  • Artifact: First token-level trace dataset for multi-model agentic systems
  • Tool: Vidur-Agent simulator for reproducible experiments
  • Benchmark: GAIA

Constraint State Governance (2605.10481) — Tianxiao Li

  • Focus: Safety in LLM multi-agent systems
  • Paradigm: Constraint drift prevention through state governance
  • Key Insight: Safe behavior must be maintained, not merely asserted

12 Angry AI Agents (2605.01986) — Ahmet Bahaddin Ersoz

  • Benchmark: Multi-agent decision-making using cinematic jury deliberation
  • Finding: 17/18 runs resulted in hung jury; anchoring is dominant failure mode
  • Insight: RLHF intensity determines deliberative flexibility

Benchmarks & Evaluation

BenchmarkDomainSizeKey Finding
Hedge-Bench (2606.03918)Financial reasoning102 tasksFrontier agents <16%
NovelAPIBench (2606.03657)Tool-use knowledge gaps1.9K tasks6 diagnostic categories
GAIATrace (2606.01725)Multi-agent tracesToken-levelFirst trace dataset
BigFinanceBench (2606.03829)Financial research workflows-Workflow-grounded

Protocols & Infrastructure

LAP (Agent-to-Instrument Protocol)

  • ArXiv: 2606.03755
  • Gap: Fills agent-to-instrument communication edge
  • Relation: Complements MCP (Anthropic) and A2A (Google)
  • Use Case: Autonomous scientific research

OpenAPI Documentation Agent-Ready

  • ArXiv: 2605.14312 — EASE 2026
  • Tool: Hermes multi-agent system
  • Result: Detected 2,450 smells in 600 endpoints
  • Purpose: MCP agent readiness

Continuum (KV Cache TTL)

  • ArXiv: 2511.02230
  • Focus: Multi-turn agent scheduling
  • Performance: 8x improvement in job completion time

Week-over-Week Summary

MetricThis WeekLast WeekChange
Total Papers315 (partial)+26
Agent-Related Papers255+20
Multi-Agent Papers121+11
Self-Evolving Agents50NEW
Avg Trend Score (Agent)6.47.2-0.8
Accepted Papers (venue)71+6

Notable Additions This Week:

  • EvoDS (KDD 2026) — first self-evolving data science agent with accepted venue
  • LAP protocol — new protocol category (agent-to-instrument)
  • Hedge-Bench — exposes frontier model gap in professional tasks
  • SkillPyramid — hierarchical skill consolidation framework

Papers from Last Week (Now Ranked Lower):

  • MUSE-Autoskill (2605.27366) — Trend: 8 → N/A
  • SIA (2605.27276) — Trend: 8 → N/A
  • FinHarness (2605.27333) — Trend: 7 → N/A
  • QUACK (2605.27068) — Trend: 7 → N/A
  • Alignment Tampering (2605.27355) — Trend: 6 → N/A
  1. Self-evolving agent frameworks surge: 3 major papers (EvoDS, SkillPyramid, EvoDrive) focus on autonomous skill acquisition, representing 40% of top-10 papers by trend score

  2. Multi-agent governance emerging: LAP protocol fills agent-to-instrument gap, Constraint State Governance addresses safety in LLM multi-agent systems

  3. Domain-specific benchmarks proliferate: Hedge-Bench (finance), NovelAPIBench (tool-use), BigFinanceBench reveal specialized evaluation needs

  4. Context management critical: Unified Context Evolution demonstrates 96.3% on ALFWorld through typed Evolvable Context Units

  5. Multi-agent characterization tools: GAIATrace + Vidur-Agent enable reproducible simulation of multi-model agentic systems

  6. RLHF alignment intensity key: 12 Angry AI Agents shows alignment level determines deliberative flexibility in multi-agent settings

Category Distribution

CategoryCountPercentage
cs.AI1858%
cs.CL413%
cs.MA413%
cs.SE26%
cs.DC13%
cs.OS13%
Other13%

Accepted Papers (with Venue)

PaperVenueArXiv ID
EvoDSKDD 20262606.03841
Uncertainty-Aware ClarificationICML 20262606.03135
Agentic CLEARACL2605.22608
Cattle TradeICLR 2026 Workshop2605.14537
OpenAPI DocumentationEASE 20262605.14312
LLM Agent SystemsIEEE AIIoT 20252505.16120
When to Re-PlanICML 2026 Workshop2606.03741

Previous Snapshots

This is the first snapshot for the ArXiv cs.AI Weekly Tracker. Future snapshots will be linked here.


Sources


Last updated: 2026-06-04 by AgentScout automated tracker. Collection duration: 180 seconds. Sources: 2/4 succeeded (ArXiv direct API rate-limited, HuggingFace 404).

ArXiv cs.AI Weekly Papers — Week of June 4, 2026: Self-Evolving Agents and Multi-Agent Governance

31 papers collected this week with 25 agent-related papers (81%). Key trends: self-evolving agent frameworks surge (EvoDS, SkillPyramid, EvoDrive), LAP protocol fills agent-to-instrument gap, and domain benchmarks expose frontier model limitations.

AgentScout ·
#arxiv #ai-agents #papers #weekly-tracker #self-evolving-agents #multi-agent-systems
Analyzing Data Nodes...
SIG_CONF:CALCULATING
Verified Sources

Data Overview

  • Snapshot Week: 2026-05-28 to 2026-06-04
  • Tracker: ArXiv cs.AI Weekly Papers (view all snapshots: /tech/ai-agents/data/?tracker=arxiv-cs-ai-weekly)
  • Update Frequency: Weekly
  • Primary Sources: ArXiv cs.AI, ArXiv cs.CL

Key Facts

  • Who: 31 papers collected from ArXiv cs.AI and cs.CL categories
  • What: 25 agent-related papers (81%), including 12 multi-agent papers and 5 self-evolving agent frameworks
  • When: Week of May 28 - June 4, 2026
  • Impact: 3 new benchmarks, 1 new protocol (LAP), 7 papers with venue acceptance

🔺 Scout Intel: What Others Missed

Confidence: high | Novelty Score: 65/100

Three self-evolving agent papers (EvoDS, SkillPyramid, EvoDrive) appear in the same week, signaling a shift from static agent architectures toward autonomous skill acquisition. LAP protocol addresses a gap most coverage ignores: agent-to-instrument communication. While MCP handles model-to-tool and A2A handles agent-to-agent, LAP targets the physical instrument edge critical for autonomous scientific research. Hedge-Bench’s <16% frontier model performance on real hedge fund tasks exposes the gap between benchmark success and professional domain competence.

Key Implication: Agent frameworks are entering a consolidation phase where autonomous skill acquisition and standardized protocols replace manual prompt engineering. The 40% concentration on self-evolving systems suggests the field recognizes current limitations of static agent capabilities.

This Week’s Papers

#TitleArXiv IDTrendVenue/Improvement
1EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management2606.0384110KDD 2026, +28.9% over SOTA
2SkillPyramid: Hierarchical Skill Consolidation for Self-Evolving Agents2606.036929+38.0% reward, -27.7% steps
3LAP: Agent-to-Instrument Protocol for Autonomous Science2606.037559NEW protocol
4GAIATrace + Vidur-Agent: Multi-Model Agentic AI Systems Characterization2606.017258GAIATrace dataset, Vidur-Agent simulator
5Unified Context Evolution for LLM Agents2606.023048ALFWorld: 75.4% → 96.3%
6EvoDrive: Pareto Evolution for Safety-Critical Autonomous Driving2606.036788Self-improving LLM agents
7Hedge-Bench: Benchmarking Agents on Financial Reasoning Tasks2606.039187102 tasks, frontier <16%
8NovelAPIBench: Diagnosing Knowledge Gaps in LLM Tool Use2606.0365771.9K tasks, 5 domains
9Uncertainty-Aware Clarification with Information Gain2606.031357ICML 2026, +3.7% success rate
10Agentic CLEAR: Multi-Level Evaluation of LLM Agents2605.226087ACL

Self-Evolving Agent Frameworks

EvoDS (2606.03841) — Zherui Yang, Fan Liu, Yansong Ning, Hao Liu — KDD 2026

  • Focus: Autonomous data science with skill learning and adaptive context compression
  • Key Innovation: Self-evolving framework that acquires skills without manual intervention
  • Performance: +28.9% over SOTA on data science benchmarks

SkillPyramid (2606.03692) — Yuan Xiong et al.

  • Focus: Hierarchical skill consolidation for reusable experience
  • Key Innovation: Multi-level skill hierarchy enabling composition and reuse
  • Performance: +38.0% reward improvement, -27.7% steps on ALFWorld and WebShop

Unified Context Evolution (2606.02304) — Zixuan Zhu et al.

  • Focus: Gradient-free framework externalizing agent experience
  • Key Innovation: Typed Evolvable Context Units for memory management
  • Performance: ALFWorld 75.4% → 96.3%, WebShop 45.1% → 61.3%

EvoDrive (2606.03678) — Tong Nie et al.

  • Focus: Safety-critical autonomous driving scenario generation
  • Key Innovation: Pareto evolution via self-improving LLM agents
  • Domain: Autonomous driving

Multi-Agent Systems & Governance

LAP Protocol (2606.03755) — Linwu Zhu et al.

  • Type: Agent-to-Instrument Protocol
  • Gap Filled: Complements MCP (model-to-tool) and A2A (agent-to-agent)
  • Use Case: Autonomous scientific instruments

GAIATrace + Vidur-Agent (2606.01725) — Donghwan Kim et al.

  • Artifact: First token-level trace dataset for multi-model agentic systems
  • Tool: Vidur-Agent simulator for reproducible experiments
  • Benchmark: GAIA

Constraint State Governance (2605.10481) — Tianxiao Li

  • Focus: Safety in LLM multi-agent systems
  • Paradigm: Constraint drift prevention through state governance
  • Key Insight: Safe behavior must be maintained, not merely asserted

12 Angry AI Agents (2605.01986) — Ahmet Bahaddin Ersoz

  • Benchmark: Multi-agent decision-making using cinematic jury deliberation
  • Finding: 17/18 runs resulted in hung jury; anchoring is dominant failure mode
  • Insight: RLHF intensity determines deliberative flexibility

Benchmarks & Evaluation

BenchmarkDomainSizeKey Finding
Hedge-Bench (2606.03918)Financial reasoning102 tasksFrontier agents <16%
NovelAPIBench (2606.03657)Tool-use knowledge gaps1.9K tasks6 diagnostic categories
GAIATrace (2606.01725)Multi-agent tracesToken-levelFirst trace dataset
BigFinanceBench (2606.03829)Financial research workflows-Workflow-grounded

Protocols & Infrastructure

LAP (Agent-to-Instrument Protocol)

  • ArXiv: 2606.03755
  • Gap: Fills agent-to-instrument communication edge
  • Relation: Complements MCP (Anthropic) and A2A (Google)
  • Use Case: Autonomous scientific research

OpenAPI Documentation Agent-Ready

  • ArXiv: 2605.14312 — EASE 2026
  • Tool: Hermes multi-agent system
  • Result: Detected 2,450 smells in 600 endpoints
  • Purpose: MCP agent readiness

Continuum (KV Cache TTL)

  • ArXiv: 2511.02230
  • Focus: Multi-turn agent scheduling
  • Performance: 8x improvement in job completion time

Week-over-Week Summary

MetricThis WeekLast WeekChange
Total Papers315 (partial)+26
Agent-Related Papers255+20
Multi-Agent Papers121+11
Self-Evolving Agents50NEW
Avg Trend Score (Agent)6.47.2-0.8
Accepted Papers (venue)71+6

Notable Additions This Week:

  • EvoDS (KDD 2026) — first self-evolving data science agent with accepted venue
  • LAP protocol — new protocol category (agent-to-instrument)
  • Hedge-Bench — exposes frontier model gap in professional tasks
  • SkillPyramid — hierarchical skill consolidation framework

Papers from Last Week (Now Ranked Lower):

  • MUSE-Autoskill (2605.27366) — Trend: 8 → N/A
  • SIA (2605.27276) — Trend: 8 → N/A
  • FinHarness (2605.27333) — Trend: 7 → N/A
  • QUACK (2605.27068) — Trend: 7 → N/A
  • Alignment Tampering (2605.27355) — Trend: 6 → N/A
  1. Self-evolving agent frameworks surge: 3 major papers (EvoDS, SkillPyramid, EvoDrive) focus on autonomous skill acquisition, representing 40% of top-10 papers by trend score

  2. Multi-agent governance emerging: LAP protocol fills agent-to-instrument gap, Constraint State Governance addresses safety in LLM multi-agent systems

  3. Domain-specific benchmarks proliferate: Hedge-Bench (finance), NovelAPIBench (tool-use), BigFinanceBench reveal specialized evaluation needs

  4. Context management critical: Unified Context Evolution demonstrates 96.3% on ALFWorld through typed Evolvable Context Units

  5. Multi-agent characterization tools: GAIATrace + Vidur-Agent enable reproducible simulation of multi-model agentic systems

  6. RLHF alignment intensity key: 12 Angry AI Agents shows alignment level determines deliberative flexibility in multi-agent settings

Category Distribution

CategoryCountPercentage
cs.AI1858%
cs.CL413%
cs.MA413%
cs.SE26%
cs.DC13%
cs.OS13%
Other13%

Accepted Papers (with Venue)

PaperVenueArXiv ID
EvoDSKDD 20262606.03841
Uncertainty-Aware ClarificationICML 20262606.03135
Agentic CLEARACL2605.22608
Cattle TradeICLR 2026 Workshop2605.14537
OpenAPI DocumentationEASE 20262605.14312
LLM Agent SystemsIEEE AIIoT 20252505.16120
When to Re-PlanICML 2026 Workshop2606.03741

Previous Snapshots

This is the first snapshot for the ArXiv cs.AI Weekly Tracker. Future snapshots will be linked here.


Sources


Last updated: 2026-06-04 by AgentScout automated tracker. Collection duration: 180 seconds. Sources: 2/4 succeeded (ArXiv direct API rate-limited, HuggingFace 404).

k29mey45s2cetoj6yvtu8████juvv0pknql7qugrfs9tusq6nnm6vcgul████brs7vjmv58b1q7yee57x4o27pgjv2bs░░░nsaoe48bv1ri89vqme0u57t3px5jsyh████0olqq7w1yj3i03ceyx3fuwkxbdjewb5m░░░uys2xffoahoijv94wreszlplrfqy4w58m░░░e85kkuujijbo5iniw2x8q6v21q8aq7yd░░░s7w6ocykm5s2x58fr7ovr29osmk9nhnuc░░░5zibuo8upv38se3centmuhq004bewujz░░░x91rdpr4bwfn60k39bxcj6h0w5k2jyq████xfn3uxidtq5xx18w4xzziaogffgtak░░░wuonr1yp51fvf81eaueqc9gta9o9fbyq████qt14wecsudbq5r5wljztbem7ooud01d8m████g39dzgbnaipnnoit0lickgk8pe1asm1go████cyd3rq8zalbgmv0i1ysufig96lg5yydqf████k1651fxzvys4n6z9go28n2jot708ful████cihfr2muxiia74z009iywr2xhztd9wgoc░░░a2pxvwvdeaj1a4qir9z9ogmywypmr3bkd░░░yy0gr3x5jksgwo5c0fynx7hkga7m4kj████q7y8zfy6yvfs6xfppuii95g65qf8mr8████tbl5ze9p8oc31f76m6bstmgtwwty4pri░░░729b0s48pf8twvc7n07i15mlli9moxiw░░░h28i9zau9665lc8jax8q54yg03a2wlbae████zgbxzcvw2ybtp2rm19bl16ana3k9zyi6░░░akaay21jps8oko2wxxq7zaccdgoypzura░░░eccyzdsxmzqqo0jj326ol6na5k16l9ux████7afm55wz8eq5rlosfzt2sc0i2etu3tmzaf░░░wz80j20bwahs90w8943920gq1vdaj8vwp████21nrkod6jdel9z6v3b3jpmkf1gfpp3sx░░░op4cxq0n3jcuz3ae3004tic8i9w8xsf████sty2qyvqh5rz9omvpsqgon6fcedsw7t1q████ztxm7cgsrdsldq1po9ptj9j5e0qq1zpl████ynvkp6ybm7rop38ktjrp2nz2ykygd6c████4wrq2qem9k9qetabwu3b0p8a5yrjulzfx░░░zh75u8dexae38rjf5knmni29af7wju8n3░░░wywmaubpofunpqr85ze29blcvqnikf4████v76dgegezz9yjids803s5z4v7hd41ms░░░c6k95gjwqbycnnzlaodk0fpwpcz03b16░░░ajpgo249v7p10kppesfkjrmsrw3oxytrn████8dciz5nazvlmz15pyzb7blagud7gvjlml████p68pazy19udbfa2rgedparunyretah░░░51phtrzb2pe2zajvuror6pm92bvcthl3f████ckmvoxi4ntjs9ole8noc8gsxzfzw7yol░░░6hm2le6lvqxa326pnhqsg7cob91msvt2w████7b8q4i79z9c034gyhj5tmce6fllyl7xhbx████01dt6d8qukrxf1ndn9zc5k9hbfr3yss34p████w9is39b6dojwkdo2cor1d8pp3x75q72zp░░░565oaidmb1qfdncqcw6tsua0ngl3s21r░░░j6j448rpieh8kblu5ljo47m875pfsqwqk████peovqwr7rat7bmee30e5fjgl0ggmuhk4q████3cctg5wi3kd