Skip to content
59

Comprehensive LLM Agent Research Collection

[Up-to-date] Large Language Model Agent: A Survey on Methodology, Applications and Challenges

2.9k stars119 forks350 entriesLast push Nov 7, 2025 (10 months ago)License none

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Overview

Read our survey paper here

Resource List >Agent Collaboration

Foam-Agent: Towards Automated Intelligent CFD Workflows

(2025) Arxiv

Why Do Multi-Agent LLM Systems Fail?

(2025) Arxiv

Linear formation control of multi-agent systems

(2025)

MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

(2025) Arxiv

A Survey of AI Agent Protocols

(2025) Arxiv

C^2: Scalable Auto-Feedback for LLM-based Chart Generation

(2025) *ACL

AgentRxiv: Towards Collaborative Autonomous Research

(2025) Arxiv

Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains

(2025) Arxiv

From Debate to Equilibrium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium

(2025) ICML

Chain of Agents: Large language models collaborating on long-context tasks

(2025) Blog

CS-Agent: LLM-based Community Search via Dual-agent Collaboration

(2025) Arxiv

MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use

(2025) Arxiv

CoMet: Metaphor-Driven Covert Communication for Multi-Agent Language Games

(2025) *ACL

Thought Communication in Multiagent Collaboration

(2025) Arxiv

Cache-to-Cache: Direct Semantic Communication Between Large Language Models

(2025) Arxiv

Adaptive Collaboration Strategy for LLMs in Medical Decision Making

(2024) NeurIPS

ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs

(2024) *ACL

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

(2024) ICLR

In 3 lists

Debating with More Persuasive LLMs Leads to More Truthful Answers

(2024) ICML

Roco: Dialectic multi-robot collaboration with large language models

(2024) Arxiv

AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning

(2024) *ACL

Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding

(2024) Arxiv

Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

(2024) *ACL

AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent Behaviors

(2024) ICLR

ChatDev: Communicative Agents for Software Development

(2024) *ACL

In 2 lists

ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

(2024) ICLR

A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration

(2024) COLM

AgentCoord: Visually Exploring Coordination Strategy for LLM-based Multi-Agent Collaboration

(2024) Arxiv

TradingAgents: Multi-Agents LLM Financial Trading Framework

(2024) Arxiv

AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

(2023) COLM

Improving Factuality and Reasoning in Language Models through Multiagent Debate

(2023) ICML

In 2 lists

Autonomous chemical research with large language models

(2023) Nature

In 3 lists

Resource List >Agent Construction

Planning with Multi-Constraints via Collaborative Language Agents

(2025) *ACL

Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

(2025) NeurIPS

SPeCtrum: A Grounded Framework for Multidimensional Identity Representation in LLM-Based Agent

(2025) Arxiv

Adaptive Thinking via Mode Policy Optimization for Social Language Agents

(2025)

On Architecture of LLM agents

(2025) Arxiv

Unified Mind Model: Reimagining Autonomous Agents in the LLM Era

(2025) Arxiv

ATLaS: Agent Tuning via Learning Critical Steps

(2025) Arxiv

Cognitive AI Memory: A Framework for More Human-like Memory in LLMs

(2025) Arxiv

Adaptive Graph Pruning: A Task-Adaptive Multi-Agent Collaboration Framework

(2025) Arxiv

Agents of Change: Self-Evolving LLM Agents for Strategic Planning

(2025) Arxiv

Reinforcing Large Language Model Reasoning through Multi-Agent Reflection

(2025) ICML

Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning

(2025) Arxiv

In 2 lists

BudgetThinker: Empowering Budget-aware LLM Reasoning with Control Tokens

(2025) Arxiv

A-MEM: Agentic Memory for LLM Agents

(2025) Arxiv

MemoCue: Empowering LLM-Based Agents for Human Memory Recall via Strategy-Guided Querying

(2025) Arxiv

Analyzing Information Sharing and Coordination in Multi-Agent Planning

(2025) Arxiv

AutoAgents: A Framework for Automatic Agent Generation

(2024) IJCAI

In 2 lists

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

(2024) ICLR

In 3 lists

Cognitive Architectures for Language Agents

(2024) TMLR

In 3 lists

Executable Code Actions Elicit Better LLM Agents

(2024) ICML

ChatDev: Communicative Agents for Software Development

(2024) *ACL

In 2 lists

Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents

(2024) CVPR/ICCV/ECCV

A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration

(2024) COLM

More Agents Is All You Need

(2024) TMLR

Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents

(2024) Arxiv

Empowering biomedical discovery with AI agents

(2024) Others

SMART-LLM: Smart Multi-Agent Robot Task Planning using Large Language Models

(2024) IROS

Perceive, Reflect, and Plan: Designing LLM Agent for Goal-Directed City Navigation without Instructions

(2024) Arxiv

Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning

(2024) Arxiv

PlanCritic: Formal Planning with Human Feedback

(2024) Arxiv

Enhancing Robot Task Planning: Integrating Environmental Information and Feedback Insights through Large Language Models

(2024) CCC

Devil's Advocate: Anticipatory Reflection for LLM Agents

(2024) Arxiv

Planning, Creation, Usage: Benchmarking LLMs for Comprehensive Tool Utilization in Real-World Complex Scenarios

(2024) *ACL

On the Structural Memory of LLM Agents

(2024) Arxiv

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society

(2023) NeurIPS

AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

(2023) COLM

AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

(2023) Arxiv

War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars

(2023) Arxiv

In 2 lists

Describe, Explain, Plan and Select: Interactive Planning with LLMs Enables Open-World Multi-Task Agents

(2023) NeurIPS

TPTU: Large Language Model-based AI Agents for Task Planning and Tool Usage

(2023) Arxiv

Resource List >Agent Evolution

Evolutionary optimization of model merging recipes

(2025) NMI

CREAM: Consistency Regularized Self-Rewarding Language Models

(2025) ICLR

KnowAgent: Knowledge-Augmented Planning for LLM-Based Agents

(2025) NAACL

STeCa: Step-level Trajectory Calibration for LLM Agent Learning

(2025) *ACL

SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

(2025) Arxiv

DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question Answering

(2025) Arxiv

Atom-Searcher: Enhancing Agentic Deep Research via Fine-Grained Atomic Thought Reward

(2025) Arxiv

PVPO: Pre-Estimated Value-Based Policy Optimization for Agentic Reasoning

(2025) Arxiv

SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents

(2025) Arxiv

LLM Collaboration With Multi-Agent Reinforcement Learning

(2025) Arxiv

VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models

(2025) Arxiv

In 2 lists

EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle

(2025) Arxiv

Self-Improving LLM Agents at Test-Time

(2025) Arxiv

CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

(2025) Arxiv

Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

(2024) Arxiv

Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization

(2024) ACL

Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement Learning

(2024) NeurIPS

A Survey on Self-Evolution of Large Language Models

(2024) Arxiv

LLM-Evolve: Evaluation for LLM’s Evolving Capability on Benchmarks

(2024) EMNLP

CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

(2024) ICLR

Iterative Translation Refinement with Large Language Models

(2024) EAMT

Agent Alignment in Evolving Social Norms

(2024) Arxiv

Mitigating the Alignment Tax of RLHF

(2024) EMNLP

Self-Rewarding Language Models

(2024) Arxiv

V-STaR: Training Verifiers for Self-Taught Reasoners

(2024) COLM

RLCD: Reinforcement learning from contrastive distillation for LM alignment

(2024) ICLR

LANGUAGE MODEL SELF-IMPROVEMENT BY REIN- FORCEMENT LEARNING CONTEMPLATION

(2024) ICLR

ProAgent: Building Proactive Cooperative Agents with Large Language Models

(2024) AAAI

Agent Planning with World Knowledge Model

(2024) NeurIPS

Refining Guideline Knowledge for Agent Planning Using Textgrad

(2024) ICKG

Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate

(2024) *ACL

LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error

(2024) *ACL

Richelieu: Self-Evolving LLM-Based Agents for AI Diplomacy

(2024) NeurIPS

Simulating Human-like Daily Activities with Desire-driven Autonomy

(2024) Arxiv

AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

(2023) NeurIPS

SELF-REFINE: Iterative Refinement with Self-Feedback

(2023) NeurIPS

Self-Evolution Learning for Discriminative Language Model Pretraining

(2023) EMNLP

Self-Evolved Diverse Data Sampling for Efficient Instruction Tuning

(2023) Arxiv

SELFEVOLVE: A Code Evolution Framework via Large Language Models

(2023) Arxiv

SELF-INSTRUCT: Aligning Language Models with Self-Generated Instructions

(2023) ACL

Large Language Models are Better Reasoners with Self-Verification

(2023) EMNLP

CODET: CODE GENERATION WITH GENERATED TESTS

(2023) ICLR

Evolving Diverse Red-team Language Models in Multi-round Multi-agent Games

(2023) Arxiv

Improving Factuality and Reasoning in Language Models through Multiagent Debate

(2023) ICML

In 2 lists

CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society

(2023) NeurIPS

STaR: Self-Taught Reasoner Bootstrapping Reasoning With Reasoning

(2022) NeurIPS

Resource List >Applications

An active inference strategy for prompting reliable responses from large language models in medical practice

(2025) npj Digital Medicine

An evaluation framework for clinical use of large language models in patient interaction tasks

(2025) Nature Medicine

Large Language Models lack essential metacognition for reliable medical reasoning

(2025) Nature Communications

Balancing autonomy and expertise in autonomous synthesis laboratories

(2025) Nature Computational Science

SimUSER: Simulating User Behavior with Large Language Models for Recommender System Evaluation

(2025) Arxiv

Swarm Autonomy: From Agent Functionalization to Machine Intelligence

(2025) Advanced Materials

ShowUI: One Vision-Language-Action Model for GUI Visual Agent

(2025) CVPR/ICCV/ECCV

Agent Laboratory: Using LLM Agents as Research Assistants

(2025) Arxiv

Towards Scientific Intelligence: A Survey of LLM-based Scientific Agents

(2025) Arxiv

In 2 lists

CitySim: Modeling Urban Behaviors and City Dynamics with Large-Scale LLM-Driven Agent Simulation

(2025) Arxiv

A Survey of AI for Materials Science: Foundation Models, LLM Agents, Datasets, and Tools

(2025) Arxiv

An Auditable Agent Platform For Automated Molecular Optimisation

(2025) Arxiv

PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation

(2025) Arxiv

Automated Clinical Problem Detection from SOAP Notes using a Collaborative Multi-Agent LLM Architecture

(2025) Arxiv

Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models

(2025) Arxiv

AlphaEvolve: A coding agent for scientific and algorithmic discovery

(2025) Arxiv

aiXiv: A Next-Generation Open Access Ecosystem for Scientific Discovery Generated by AI Scientists

(2025) Arxiv

GenoMAS: A Multi-Agent Framework for Scientific Discovery via Code-Driven Gene Expression Analysis

(2025) Arxiv

Motif: Intrinsic Motivation from Artificial Intelligence Feedback

(2024) ICLR

Baba Is AI: Break the Rules to Beat the Benchmark

(2024) ICML

Large language model-empowered agents for simulating macroeconomic activities

(2024) *ACL

CompeteAI: Understanding the Competition Dynamics in Large Language Model-based Agents

(2024) ICML

In 2 lists

Understanding the benefits and challenges of using large language model-based conversational agents for mental…

(2024) AMIA

Exploring Collaboration Mechanisms for LLM Agents

(2024) *ACL

Simulating Human Society with Large Language Model Agents: City, Social Media, and Economic System

(2024) WWW

Can large language models transform computational social science?

(2024) *ACL

AgentCF: Collaborative Learning with Autonomous Language Agents for Recommender Systems

(2024) SIGIR

On Generative Agents in Recommendation

(2024) SIGIR

In 2 lists

ChatDev: Communicative Agents for Software Development

(2024) *ACL

CRISPR-GPT: An LLM Agent for Automated Design of Gene-Editing Experiments

(2024) Arxiv

SciAgents: Automating Scientific Discovery Through Bioinspired Multi-Agent Intelligent Graph Reasoning

(2024) Advanced Materials

Medical large language models are susceptible to targeted misinformation attacks

(2024) npj Digital Medicine

CellAgent: An LLM-driven Multi-Agent Framework for Automated Single-cell Data Analysis

(2024) Arxiv

Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

(2023) NeurIPS

Language Models Meet World Models: Embodied Experiences Enhance Language Models

(2023) NeurIPS

ChessGPT: Bridging Policy Learning and Language Modeling

(2023) NeurIPS

Mindagent: Emergent gaming interaction

(2023) Arxiv

Exploring large language models for communication games: An empirical study on Werewolf

(2023) Arxiv

Language as reality: a co-creative storytelling game experience in 1001 nights using generative AI

(2023) AAAI

TradingGPT: Multi-Agent System with Layered Memory and Distinct Characters for Enhanced Financial Trading Performance

(2023) Arxiv

Using large language models to simulate multiple humans and replicate human subject studies

(2023) ICML

Generative Agents: Interactive Simulacra of Human Behavior

(2023) UIST

In 5 listsDetails

Self-collaboration Code Generation via ChatGPT

(2023) TOSEM

Language models can solve computer tasks

(2023) NeurIPS

ChemCrow: Augmenting large-language models with chemistry tools

(2023) Arxiv

In 3 lists

AlphaFlow: autonomous discovery and optimization of multi-step chemistry using a self-driven fluidic lab guided by…

(2023) Nature Communications

In 2 lists

Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

(2022) ICML

Stress-testing the resilience of the Austrian healthcare system using agent-based simulation

(2022) Nature Communications

Resource List >Datasets & Benchmarks

AgentHarm: Benchmarking Robustness of LLM Agents on Harmful Tasks

(2025) ICLR

AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction Simulator

(2025) *ACL

Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

(2025) *ACL

DCA-Bench: A Benchmark for Dataset Curation Agents

(2025) ICLR

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

(2025) Arxiv

MLE-Bench: Evaluating Machine Learning Agents on Machine Learning Engineering

(2025) ICLR

EgoLife: Towards Egocentric Life Assistant

(2025) Arxiv

DSBench: How Far Are Data Science Agents to Becoming Data Science Experts?

(2025) ICLR

Towards Internet-Scale Training For Agents

(2025) Arxiv

macOSWorld: An Interactive Benchmark for GUI Agents

(2025) Arxiv

Humanity's Last Exam

(2025) Arxiv

MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models

(2025) Arxiv, *ACL

IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis

(2025) Arxiv

SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks

(2025) Arxiv

MMSearch-Plus: A Simple Yet Challenging Benchmark for Multimodal Browsing Agents

(2025) Arxiv

MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents

(2025) *ACL

Establishing Best Practices for Building Rigorous Agentic Benchmarks

(2025) Arxiv

UserBench: An Interactive Gym Environment for User-Centric Agents

(2025) Arxiv

PillagerBench: A Competitive Multi-Agent Benchmark for Evaluating LLM-based Agents in Minecraft

(2025) Arxiv

UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI

(2025) CVPR/ICCV/ECCV

In 2 lists

Probe by Gaming: A Game-based Benchmark for Assessing Conceptual Knowledge in LLMs

(2025) Arxiv

NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents

(2025) Arxiv

LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries

(2025) Arxiv

Achilles Heel of Distributed Multi-Agent Systems

(2025) Arxiv

AgentBench: Evaluating LLMs as Agents

(2024) ICLR

AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents

(2024) *ACL

BENCHAGENTS: Automated Benchmark Creation with Agent Interaction

(2024) Arxiv

Benchmarking Data Science Agents

(2024) *ACL

Benchmarking Large Language Models as AI Research Agents

(2024) ICLR

Benchmarking Large Language Models for Multi-agent Systems: A Comparative Analysis of AutoGen, CrewAI, and TaskWeaver

(2024) Others

BLADE- Benchmarking Language Model Agents

(2024) *ACL

CRAB: Cross-platfrom agent benchmark for multi-modal embodied language model agents

(2024) NeurIPS

CToolEval: A Chinese Benchmark for LLM-Powered Agent Evaluation in Real-World API Interactions

(2024) *ACL

DA-Code: Agent Data Science Code Generation Benchmark for Large Language Models

(2024) *ACL

Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

(2025) NeurIPS

GTA: A Benchmark for General Tool Agents

(2024) NeurIPS

LaMPilot: An Open Benchmark Dataset for Autonomous Driving with Language Model Programs

(2024) CVPR/ICCV/ECCV

ML Research Benchmark

(2024) Arxiv

MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains

(2024) Arxiv

OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

(2024) CVPR/ICCV/ECCV

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

(2024) NeurIPS

Revisiting Benchmark and Assessment: An Agent-based Exploratory Dynamic Evaluation Framework for LLMs

(2024) Arxiv

Seal-Tools: Self-instruct Tool Learning Dataset for Agent Tuning and Detailed Benchmark

(2024) Others

Tapilot-Crossing: Benchmarking and Evolving LLMs Towards Interactive Data Analysis Agents

(2024) Arxiv

TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

(2024) Arxiv

Tur[k]ingBench: A Challenge Benchmark for Web Agents

(2024) Arxiv

Agent-FLAN: Designing Data and Methods of Effective Agent Tuning for Large Language Models

(2024) *ACL

AgentBank: Towards Generalized LLM Agents via Fine-Tuning on 50000+ Interaction Trajectories

(2024) *ACL

AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning

(2024) Arxiv

AgentTuning: Enabling Generalized Agent Abilities for LLMs

(2024) *ACL

Executable Code Actions Elicit Better LLM Agents

(2024) ICML

AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

(2024) *ACL

SheetAgent: Towards A Generalist Agent for Spreadsheet Reasoning and Manipulation via Large Language Models

(2024) WWW

GenoTEX: An LLM Agent Benchmark for Automated Gene Expression Data Analysis

(2024) Arxiv

FireAct: Toward Language Agent Fine-tuning

(2023) Arxiv

Resource List >Ethics

Medical large language models are vulnerable to data-poisoning attacks

(2025) Nature Medicine

Foundation Models and Fair Use

(2024) JMLR

Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model

(2023) JMLR

LLaMA: Open and Efficient Foundation Language Models

(2023)

Predictability and Surprise in Large Generative Models

(2022) FAccT

On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜

(2021) FAccT

In 2 lists

Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets

(2021) NeurIPS

GPT-3: Its Nature, Scope, Limits, and Consequences

(2020)

Energy and Policy Considerations for Modern Deep Learning Research

(2020) AAAI

Defending Against Neural Fake News

(2019) NeurIPS

Resource List >Security

RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage

(2025) Arxiv

Red-Teaming LLM Multi-Agent Systems via Communication Attacks

(2025) Arxiv

Unveiling Privacy Risks in LLM Agent Memory

(2025) Arxiv

AEIA-MN: Evaluating the Robustness of Multimodal LLM-Powered Mobile Agents Against Active Environmental Injection…

(2025) Arxiv

Firewalls to Secure Dynamic LLM Agentic Networks

(2025) Arxiv

AUTOHIJACKER: AUTOMATIC INDIRECT PROMPT INJECTION AGAINST BLACK-BOX LLM AGENTS

(2025) Arxiv

AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways

(2025) ACM Computing Survey

SAGA: A Security Architecture for Governing AI Agentic Systems

(2025) Arxiv

WebInject: Prompt Injection Attack to Web Agents

(2025) Arxiv

Web Fraud Attacks on LLM-driven Multi-Agent Systems

(2025) Arxiv

Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models

(2025) Arxiv

Beyond Data Privacy: New Privacy Risks for Large Language Models

(2025) Arxiv

Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents

(2025) Arxiv

PrivWeb: Unobtrusive and Content-aware Privacy Protection For Web Agents

(2025) Arxiv

DemonAgent: Dynamically Encrypted Multi-Backdoor Implantation Attack on LLM-based Agent

(2025) *ACL

CORBA: Contagious Recursive Blocking Attacks on Multi-Agent Systems Based on Large Language Models

(2025) Arxiv

G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-agent Systems

(2025) *ACL

AgentHarm: Benchmarking Robustness of LLM Agents on Harmful Tasks

(2025) ICLR

Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks

(2025) Arxiv

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

(2025) Arxiv

LLM-based Multi-Agent Systems: Techniques and Business Perspectives

(2024) Arxiv

BlockAgents: Towards Byzantine-Robust LLM-Based Multi-Agent Coordination via Blockchain

(2024) TURC

PROMPT INFECTION: LLM-TO-LLM PROMPT INJECTION WITHIN MULTI-AGENT SYSTEMS

(2024) Arxiv

AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

(2024) NeurIPS

AGENTPOISON: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases

(2024) NeurIPS

AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks

(2024) Arxiv

Imprompter- Tricking LLM Agents into Improper Tool Use

(2024) Arxiv

TARGETING THE CORE: A SIMPLE AND EFFECTIVE METHOD TO ATTACK RAG-BASED AGENTS VIA DIRECT LLM MANIPULATION

(2024) Arxiv

Prompt Injection as a Defense Against LLM-driven Cyberattacks

(2024) Arxiv

Evil Geniuses: Delving into the Safety of LLM-based Agents

(2024) Arxiv

AGENT SECURITY BENCH (ASB): FORMALIZING AND BENCHMARKING ATTACKS AND DEFENSES IN LLM-BASED AGENTS

(2024) Arxiv

AGENTHARM: A BENCHMARK FOR MEASURING HARMFULNESS OF LLM AGENTS

(2024) Arxiv

CLAS 2024: The Competition for LLM and Agent Safety

(2024) Arxiv

The Task Shield: Enforcing Task Alignment to Defend Against Indirect Prompt Injection in LLM Agents

(2024) Arxiv

WIPI: A New Web Threat for LLM-Driven Web Agents

(2024) Arxiv

Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

(2024) Arxiv

CORBA: Contagious Recursive Blocking Attacks on Multi-Agent Systems Based on Large Language Models

(2024) Arxiv

PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

(2024) ACL

Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In

(2024) Arxiv

AGENT-SAFETYBENCH: Evaluating the Safety of LLM Agents

(2024) Arxiv

INJECAGENT: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

(2024) Arxiv

PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

(2024) Arxiv

TrustAgent: Towards Safe and Trustworthy LLM-based Agents

(2024) *ACL

Watch Out for Your Agents! Investigating Backdoor Threats to LLM-Based Agents

(2024) NeurIPS

R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

(2024) *ACL

NetSafe: Exploring the Topological Safety of Multi-agent Networks

(2024) *ACL

A Trembling House of Cards? Mapping Adversarial Attacks against Language Agents

(2024) Arxiv

Resource List >Survey

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

(2025) Arxiv

Trust but Verify! A Survey on Verification Design for Test-time Scaling

(2025) Arxiv

A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers

(2025) Arxiv

In 2 lists

Evaluation and Benchmarking of LLM Agents: A Survey

(2025) Arxiv

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic…

(2025) Arxiv

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

(2025) Arxiv

Benchmark Evaluations, Applications, and Challenges of Large Vision Language Models: A Survey

(2025) Arxiv

The Future is Agentic: Definitions, Perspectives, and Open Challenges of Multi-Agent Recommender Systems

(2025) Arxiv

Commercial LLM Agents Are Already Vulnerable to Simple Yet Dangerous Attacks

(2025) Arxiv

Multi-Agent Collaboration Mechanisms: A Survey of LLMs

(2025) Arxiv

AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways

(2025) ACM Computing Survey

Large Model Based Agents: State-of-the-Art, Cooperation Paradigms, Security and Privacy, and Future Trends

(2024) Arxiv

Agent AI: Surveying the Horizons of Multimodal Interaction

(2024) Arxiv

Large Language Model based Multi-Agents: A Survey of Progress and Challenges

(2024) Arxiv

Large Multimodal Agents: A Survey

(2024) Arxiv

Understanding the planning of LLM agents: A survey

(2024) Arxiv

Computational Experiments Meet Large Language Model Based Agents: A Survey and Perspective

(2024) Arxiv

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security

(2024) Arxiv

The Landscape of Emerging AI Agent Architectures for Reasoning, Planning, and Tool Calling: A Survey

(2024) Arxiv

Exploring Large Language Model based Intelligent Agents: Definitions, Methods, and Prospects

(2024) Arxiv

Position Paper: Agent AI Towards a Holistic Intelligence

(2024) Arxiv

Large Language Model based Multi-Agents: A Survey of Progress and Challenges

(2024) IJCAI

LLM With Tools: A Survey

(2024) Arxiv

A Survey on the Memory Mechanism of Large Language Model based Agents

(2024) Arxiv

Large Language Model based Multi-Agents: A Survey of Progress and Challenges

(2024) Arxiv

A Survey on Large Language Model-Based Game Agents

(2024) Arxiv

Large Language Models and Games: A Survey and Roadmap

(2024) Arxiv

Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents

(2024) Arxiv

Security of AI Agents

(2024) Arxiv

PERSONAL LLM AGENTS: INSIGHTS AND SURVEY ABOUT THE CAPABILITY, EFFICIENCY AND SECURITY

(2024) Arxiv

The Emerged Security and Privacy of LLM Agent: A Survey with Case Studies

(2024) Arxiv

Inferring the Goals of Communicating Agents from Actions and Instructions

(2024) ICML Workshop

Recent advancements in LLM Red-Teaming: Techniques, Defenses, and Ethical Considerations

(2024) Arxiv

Deconstructing The Ethics of Large Language Models from Long-standing Issues to New-emerging Dilemmas: A…

(2024)

A survey on large language model based autonomous agents

(2023) FCS

The rise and potential of large language model based agents: a survey

(2023) SCIS

In 2 lists

Large Language Model Alignment: A Survey

(2023) Arxiv

Ethical and social risks of harm from Language Models

(2021) Arxiv

On the Opportunities and Risks of Foundation Models

(2021) Arxiv

Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims

(2020) Arxiv

In 2 lists

Actionable Auditing: Investigating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products

(2019) AIES

Resource List >Tools

ToolCoder: A Systematic Code-Empowered Tool Learning Framework for Large Language Models

(2025) Arxiv

VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use

(2025) Arxiv

Re-Invoke: Tool Invocation Rewriting for Zero-Shot Tool Retrieval

(2024) Arxiv

Chain of Tools: Large Language Model is an Automatic Multi-tool Learner

(2024) Arxiv

EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction

(2024) Arxiv

ToolGen: Unified Tool Retrieval and Calling via Generation

(2024) Arxiv

ToolNet: Connecting Large Language Models with Massive Tools via Tool Graph

(2024) Arxiv

ToolPlanner: A Tool Augmented LLM for Multi Granularity Instructions with Path Planning and Feedback

(2024) Arxiv

Making Language Models Better Tool Learners with Execution Feedback

(2024) *ACL

Leveraging Large Language Models to Improve REST API Testing

(2024) ICSE-NIER

LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error

(2024) *ACL

Skills-in-Context: Unlocking Compositionality in Large Language Models

(2024) *ACL

TaskMatrix.AI: Completing Tasks by Connecting Foundation Models with Millions of APIs

(2024) Others

Gorilla: Large Language Model Connected with Massive APIs

(2024) NeurIPS

LARGE LANGUAGE MODELS AS TOOL MAKERS

(2024) ICLR

Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents

(2023) Arxiv

Recommender AI Agent: Integrating Large Language Models for Interactive Recommendations

(2023) Arxiv

ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

(2023) Arxiv

TPTU-v2: Boosting Task Planning and Tool Usage of Large Language Model-based Agents in Real-world Systems

(2023) Arxiv

TPTU: Large Language Model-based AI Agents for Task Planning and Tool Usage

(2023) Arxiv

GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction

(2023) NeurIPS

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

(2023) *ACL

ChatCoT: Tool-Augmented Chain-of-Thought Reasoning on Chat-based Large Language Models

(2023) *ACL

ToolQA: A Dataset for LLM Question Answering with External Tools

(2023) NeurIPS

On the Tool Manipulation Capability of Open-source Large Language Models

(2023) Arxiv

RestGPT: Connecting Large Language Models with Real-World RESTful APIs

(2023) Arxiv

Toolformer: Language Models Can Teach Themselves to Use Tools

(2023) NeurIPS

WebCPM: Interactive Web Search for Chinese Long-form Question Answering

(2023) *ACL

ToolCoder: Teach Code Generation Models to use API search tools

(2023) Arxiv

ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases

(2023) Arxiv

ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings

(2023) NeurIPS

MultiTool-CoT: GPT-3 Can Use Multiple External Tools with Chain of Thought Prompting

(2023) *ACL

CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models

(2023) *ACL

GEAR: Augmenting Language Models with Generalizable and Efficient Tool Resolution

(2023) Arxiv

Dify

(2023)

In 14 listsDetails

LangChain

(2023)

In 20 listsDetails

WebGPT: Browser-assisted question-answering with human feedback

(2022) Arxiv

Task Bench: A Parameterized Benchmark for Evaluating Parallel Runtime Performance

(2020) SC

See category
94

Awesome OpenClaw Skills

VoltAgent/awesome-openclaw-skills

The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞

Fresh★ 53k830 entriesPushed today
92

Awesome DeepSeek Harness (DSH) Plugin

awesome-dsh-plugin/awesome-dsh-plugin

A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表

Fresh★ 17k1654 entriesPushed today
91

Awesome Guidelines

Kristories/awesome-guidelines

Programming style, best practices, and coding conventions.

Fresh★ 11k166 entriesPushed 2 days ago
90

Awesome

sindresorhus/awesome

😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]

Fresh★ 513k51 entriesPushed 28 days ago
90

Awesome Prompts

ai-boost/awesome-prompts

Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.

Fresh★ 9k288 entriesPushed yesterday
90

Awesome README

matiassingers/awesome-readme

A curated list of awesome READMEs

Fresh★ 22k143 entriesPushed yesterday