Comprehensive LLM Agent Research Collection
[Up-to-date] Large Language Model Agent: A Survey on Methodology, Applications and Challenges
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Overview
Resource List >Agent Collaboration
Why Do Multi-Agent LLM Systems Fail?
(2025) Arxiv
A Survey of AI Agent Protocols
(2025) Arxiv
Resource List >Agent Construction
On Architecture of LLM agents
(2025) Arxiv
ATLaS: Agent Tuning via Learning Critical Steps
(2025) Arxiv
Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
(2025) Arxiv
A-MEM: Agentic Memory for LLM Agents
(2025) Arxiv
Cognitive Architectures for Language Agents
(2024) TMLR
Editable Scene Simulation for Autonomous Driving via Collaborative LLM-Agents
(2024) CVPR/ICCV/ECCV
More Agents Is All You Need
(2024) TMLR
Empowering biomedical discovery with AI agents
(2024) Others
PlanCritic: Formal Planning with Human Feedback
(2024) Arxiv
On the Structural Memory of LLM Agents
(2024) Arxiv
War and Peace (WarAgent): Large Language Model-based Multi-Agent Simulation of World Wars
(2023) Arxiv
Resource List >Agent Evolution
VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
(2025) Arxiv
Self-Improving LLM Agents at Test-Time
(2025) Arxiv
Agent Alignment in Evolving Social Norms
(2024) Arxiv
Mitigating the Alignment Tax of RLHF
(2024) EMNLP
Self-Rewarding Language Models
(2024) Arxiv
Agent Planning with World Knowledge Model
(2024) NeurIPS
SELF-REFINE: Iterative Refinement with Self-Feedback
(2023) NeurIPS
CODET: CODE GENERATION WITH GENERATED TESTS
(2023) ICLR
Resource List >Applications
An active inference strategy for prompting reliable responses from large language models in medical practice
(2025) npj Digital Medicine
Large Language Models lack essential metacognition for reliable medical reasoning
(2025) Nature Communications
Balancing autonomy and expertise in autonomous synthesis laboratories
(2025) Nature Computational Science
Swarm Autonomy: From Agent Functionalization to Machine Intelligence
(2025) Advanced Materials
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
(2025) CVPR/ICCV/ECCV
On Generative Agents in Recommendation
(2024) SIGIR
Medical large language models are susceptible to targeted misinformation attacks
(2024) npj Digital Medicine
ChessGPT: Bridging Policy Learning and Language Modeling
(2023) NeurIPS
Mindagent: Emergent gaming interaction
(2023) Arxiv
Self-collaboration Code Generation via ChatGPT
(2023) TOSEM
Language models can solve computer tasks
(2023) NeurIPS
AlphaFlow: autonomous discovery and optimization of multi-step chemistry using a self-driven fluidic lab guided by…
(2023) Nature Communications
Stress-testing the resilience of the Austrian healthcare system using agent-based simulation
(2022) Nature Communications
Resource List >Datasets & Benchmarks
EgoLife: Towards Egocentric Life Assistant
(2025) Arxiv
Towards Internet-Scale Training For Agents
(2025) Arxiv
Humanity's Last Exam
(2025) Arxiv
MCPEval: Automatic MCP-based Deep Evaluation for AI Agent Models
(2025) Arxiv, *ACL
UnrealZoo: Enriching Photo-realistic Virtual Worlds for Embodied AI
(2025) CVPR/ICCV/ECCV
Achilles Heel of Distributed Multi-Agent Systems
(2025) Arxiv
AgentBench: Evaluating LLMs as Agents
(2024) ICLR
Benchmarking Data Science Agents
(2024) *ACL
BLADE- Benchmarking Language Model Agents
(2024) *ACL
GTA: A Benchmark for General Tool Agents
(2024) NeurIPS
ML Research Benchmark
(2024) Arxiv
FireAct: Toward Language Agent Fine-tuning
(2023) Arxiv
Resource List >Ethics
Medical large language models are vulnerable to data-poisoning attacks
(2025) Nature Medicine
Foundation Models and Fair Use
(2024) JMLR
Defending Against Neural Fake News
(2019) NeurIPS
Resource List >Security
Unveiling Privacy Risks in LLM Agent Memory
(2025) Arxiv
Firewalls to Secure Dynamic LLM Agentic Networks
(2025) Arxiv
AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
(2025) ACM Computing Survey
WebInject: Prompt Injection Attack to Web Agents
(2025) Arxiv
WIPI: A New Web Threat for LLM-Driven Web Agents
(2024) Arxiv
Resource List >Survey
AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
(2025) ACM Computing Survey
Large Multimodal Agents: A Survey
(2024) Arxiv
LLM With Tools: A Survey
(2024) Arxiv
Security of AI Agents
(2024) Arxiv
Inferring the Goals of Communicating Agents from Actions and Instructions
(2024) ICML Workshop
Large Language Model Alignment: A Survey
(2023) Arxiv
Resource List >Tools
Leveraging Large Language Models to Improve REST API Testing
(2024) ICSE-NIER
LARGE LANGUAGE MODELS AS TOOL MAKERS
(2024) ICLR
Related lists in Miscellaneous
See categoryAwesome OpenClaw Skills
VoltAgent/awesome-openclaw-skills
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
Awesome DeepSeek Harness (DSH) Plugin
awesome-dsh-plugin/awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
Awesome Guidelines
Kristories/awesome-guidelines
Programming style, best practices, and coding conventions.
Awesome
sindresorhus/awesome
😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]
Awesome Prompts
ai-boost/awesome-prompts
Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.
Awesome README
matiassingers/awesome-readme
A curated list of awesome READMEs