DSPy
Write LM pipelines declaratively, then compile — DSPy auto-optimizes prompts and few-shot demonstrations. The strongest engineering-first approach.
Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
Treats LLM feedback as "textual gradients" and backpropagates them to optimize prompts. Published in Nature.
Reflective Text Evolution — optimizes prompts, code, and agent configs. Claims +6–20 pts over GRPO on 6 tasks with fewer rollouts.
Evolutionary self-improvement for Hermes Agent — DSPy + GEPA (Genetic-Pareto Prompt Evolution) automatically evolves skills, tool descriptions, system prompts, and code via reflective search over execution traces (understands why things fail, not just that they failed); constraint gates (tests,…
Open-source infrastructure for continually self-improving agents — connects agent inference, feedback, learning, and versioned delivery; either train model weights (Slime + SGLang) or improve the agent harness itself (prompts, rules, skills) from interaction feedback, no local GPUs required;…
Reliability layer for self-hosted LLM tool-calling — guardrails (rescue parsing, retry nudges, response validation), optional workflow constraints (required_steps, prerequisites, terminal_tool), and built-in eval suite. MIT, 2.2k+ stars, Feb 2026
Hallucination gate for agents — the model proposes claims, deterministic tools check each against ground truth and return VERIFIED/REFUTED with evidence; only what survives counts as fact. Ships as MCP server + CLI, with reverify rollover for lossless context handoff across resets. Caught every…
Test-driven prompt engineering: regression tests, red teaming, model comparison, CI/CD integration. Acquired by OpenAI (Mar 2026) — remains open source.
Open eval framework and benchmark registry — standardizes LLM performance measurement.
Real-terminal agent benchmark (Stanford/Laude) — compile code, train models, set up servers in Docker-sandboxed environments; the de facto benchmark for agentic coding (2026).
LLM vulnerability scanner by NVIDIA — red teaming, prompt injection, jailbreak, and leakage detection.
Official OpenAI guide on designing agents to resist prompt injection — browser agents, defense principles (2026).
Bruce Schneier (Harvard/Lawfare): reframes prompt injection as a 7-stage malware kill chain; 21/36 documented attacks already traverse 4+ stages. Featured at Black Hat 2026.
7 packages (Python/Rust/TS/Go/.NET) — policy enforcement (<0.1ms), zero-trust agent identity (Ed25519 + SPIFFE), sandboxed execution; covers all OWASP Agentic Top 10; adapters for LangChain/CrewAI/ADK/OpenAI Agents SDK (Apr 2026)
Stress-test agents for goal drift and system-prompt violations across 6 value dimensions — multi-turn escalation, LLM-as-judge, interactive HTML reports; inspired by ICLR 2026 workshop paper (Apr 2026)
Autonomous red-teaming meta-harness for AI coding agents — recon → exploit → report against authorized targets, multi-agent offensive-security workflows, offline-model support; by elder-plinius (AGPL-3.0, 5.3k+ stars, July 2026)
Official OpenAI CLI and TypeScript SDK for finding, validating, and fixing security vulnerabilities — standard/deep scans, diff and working-tree targets, SARIF/CSV/JSON export, CI-native exit codes, pre-commit hooks (Apache-2.0, 8k+ stars, July 2026)
Security scanner for AI agent skills — detects vulnerabilities, malicious patterns, prompt injection, data exfiltration, and supply-chain risks in Claude Code, Codex, and MCP skills before installation (Apache-2.0, 14.7k+ stars, Mar 2026)
Unit testing for LLMs — G-Eval, hallucination, RAG faithfulness, agentic task metrics.
Open-source LLM engineering platform — tracing, evals, prompt management, A/B experiments.
Open-source AI observability & evaluation platform (Arize) — OpenTelemetry-native tracing for agents, LLM-as-judge evals, versioned datasets & experiments for prompt regression testing, prompt management with version control and replay, plus an MCP endpoint so Claude Code/Cursor can query traces…
Trace-native CI/CD for AI agents — grades every production trace as it lands (LLM-as-judge evaluators as trace-table columns), clusters failures into issues, freezes failing runs into hermetic replayable regression cases ($0 replay, no API keys), blocks the PR via CI gate, alerts via…
Most comprehensive — Cursor, Devin, Windsurf, Claude Code, v0, Lovable, Perplexity, Manus, Replit, Warp and 20+ more. Actively maintained.
20,000+ lines across 25+ tools (Claude Code, Cursor, Devin, Lovable, Manus, Windsurf, Kiro, v0, Codex, and more) — full tool definitions and internal agent logic; updated Mar 2026
Claude Code internal prompts — main system prompt, 18 tool descriptions, Plan/Explore/Task sub-agent prompts, 135+ version changelog
ChatGPT, Claude, Gemini system prompts and developer messages
Well-organized, includes tool call constraints and persona definitions
Focused on Claude system prompt analysis
Anthropic's systematic guide to managing the full context state—system prompts, tools, MCP, and message history—as a finite, curated resource. Reframes harness design as "what configuration of context produces the desired behavior?" rather than just prompt wording.
first-principles handbook on context design, orchestration, and optimization
curated papers, frameworks, and implementation guides
comprehensive, MIT-licensed collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems — context fundamentals/degradation/compression, memory systems, tool design, harness engineering, self-improvement loops; cited in academic research as…
Open-source codebase context layer for coding agents — builds a persistent, queryable code graph and pulls matching nodes into each prompt; 46% fewer tool calls, 42% token savings, 60% faster on a 162-run benchmark, 66% SWE-bench Verified (vs 54% cold); supports Claude Code, Cursor, Codex, Gemini…
"The ripgrep of AI context" (Red Hat): zero-dependency C++23 CLI + MCP server that lets coding agents find what they need without reading the whole repo — tree-sitter signatures at 74.7% fewer bytes than bodies, blast-radius queries, tests-to-run suggestions, and post-change verification with…
LangChain
CrewAI
Microsoft
OpenAI
Anthropic
Anthropic
Anthropic
Karpathy
Microsoft
OpenAI
ByteDance
OpenBMB / THUNLP / ModelBest / AI9Stars
Unicity
TinyHumans
HuggingFace
Astro
OSS
Vercel
ShawnPana
Alibaba/Qwen
DeusData
Tencent Cloud
rohitg00
Vercel
Portia Labs
Paperclip AI
Block
Moonshot AI
Yeachan Heo
UltraWorkers
Nous Research
Tracer Cloud
DeepSeek
TrueFoundry
CopilotKit
YC Software
Open-source infrastructure for continually self-improving agents — connects agent inference, feedback, learning, and versioned delivery; either train model weights (Slime + SGLang) or improve the agent harness itself (prompts, rules, skills) from interaction feedback, no local GPUs required;…
BitMiracle AI
Oomol
Official collection + spec (/spec/agent-skills-spec.md)
1000+ community skills, works across all major platforms
Vercel's official skills
Official docs & spec
Announcement post
When to use which
Official OpenAI post: "leveraging Codex in an agent-first world"
Component-by-component breakdown
TerminalBench 2.0 case study: 52.8% → 66.5%, same model
"The harness is the dataset. Competitive advantage is the trajectories it captures."
Architecture perspective
Sub-agents as context firewalls, practical patterns
Long-running agent design
Production harness: 4-tier routing, parallel worktrees, lifecycle hooks, 6 skills
LangChain's opinionated deep agent harness (used in TerminalBench)
Unified virtual filesystem for AI agents — mounts S3, GDrive, Slack, Gmail, Redis as one tree; agents use bash across every backend; Python/TypeScript SDKs, cache, snapshots (May 2026)
How Anthropic used parallel Claude sub-agents to build a C compiler — generator/evaluator harness patterns
Open-source loop/harness improvement skill — turns project and session evidence into prioritized improvements and verifiable next steps for Claude Code, Codex, Cursor, and other coding agents (July 2026)
Ryan Lopopolo's anthology, field guide, and agent context bundle for harness engineering — shaping context and tools so agents can recover intent, operate systems, respect authority, prove outcomes, and leave the next run better equipped (CC-BY-4.0, July 2026)
"Stop prompting. Design the loop." — practical patterns, starters & CLI (loop-audit, loop-init, loop-cost) for systems that discover work, hand it to agents, verify results, and persist state across Claude Code, Codex, Grok, and OpenCode; report-only week one, scores loops on a "Loop Ready" rubric…
The agent harness performance optimization system — skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond (MIT, 242k+ stars, Jan 2026)
NVIDIA's efficiency extension for the Pi coding agent — four opt-in harness mechanisms distilled from scaled auto-research loops (arXiv 2609.20519): Action Fusion (run follow-up validation in the same tool call), ObservationPack (stable handles with exact paged recall for repeated large outputs),…
"Let's think step by step" — zero-shot CoT milestone
Multi-path sampling + majority vote: GSM8K 57% → 74%
Reasoning + Acting interleaved — foundation of agent prompt design
LLM auto-generates and selects instructions — beats human prompts
Formalizes prompt engineering as expressivity problem — proves a fixed Transformer backbone can approximate any continuous function by varying only the prompt; decomposes switching into routing/arithmetic/composition
5W3H structured intent representation reduces cross-model output variance and avoids the dual-inflation bias of unstructured prompts; AI-expanded 5W3H matches manually crafted 5W3H across English, Japanese, and AI-assisted authoring
Textual gradient descent — source paper for many auto-optimization methods
Prompts as compilable programs — defines the engineering-first paradigm
Optimizes instructions and demonstrations across multi-stage LM programs
"Autograd for text" — LLM feedback as gradients, published in Nature
Reflective evolution outperforms GRPO by 6–20 pts with fewer rollouts
Treats prompts as structured objects; optimizes each semantic section independently with local textual gradients
Reframes prompt design as causal estimation — uses Double Machine Learning to isolate prompt effects
Memory-augmented APO that stores historical refinement insights and reuses them across iterations
Berkeley/Stanford (Stoica, Zou, Gonzalez): scales parallel prompt learning with up to 17x speedup over ACE/GEPA via parallel scans and dynamic batching; evaluated on AppWorld, Terminal-Bench, FiNER
Multi-agent prompt optimization framework that applies requirements engineering (elicitation, analysis, specification, validation) to generate production-ready system and user prompts for agent-based software development
Apple: embarrassingly simple self-distillation (SSD) — sample from model, fine-tune on raw unverified samples via cross-entropy; no reward model, no verifier, no RL; Qwen3-30B 42.4% → 55.3% pass@1 on LiveCodeBench v6; gains concentrate on hard problems; open source
NUS/CityUHK: closes the self-referential loop by treating the prompt agent's own system prompt as an optimization target alongside task-agent prompts; open-ended evolutionary search with an archive of stepping-stone candidates; two-stage pre-train/fine-tune pipeline generalizes to held-out tasks;…
≤5 words per reasoning step — 91% of CoT accuracy at 7.6% of the tokens; 76% latency reduction
IBM Research AI: replaces verbal CoT with short sequences of learned, reserved vocabulary tokens; up to 11.6× fewer reasoning tokens with comparable accuracy on math, instruction-following, and multi-hop reasoning
Longer CoT ≠ better reasoning — identifies "deep-thinking tokens" (high-revision tokens) as the true signal; enables cost-efficient test-time scaling
Detects overthinking/underthinking via confidence variance and applies steering vectors to redirect reasoning — ICLR 2026; works on DeepSeek-R1, QwQ, o3-class models
"Jagged" iterative reasoning — splits long reasoning into short segments with summaries, enabling unlimited depth without hitting context limits; ICLR 2026; +3–13% on MATH500/AIME24/GPQA
Google DeepMind: DeepSeek-R1/QwQ-32B superior reasoning emerges from simulating internal multi-agent dialogue — base models trained purely on reasoning accuracy spontaneously develop questioning, perspective-switching, and contradiction-resolving behaviors
For simple tasks, the model's final answer is already decodable from early-layer activations before CoT generates a single token — CoT produces genuine belief change only on hard problems; probe-guided early-exit reduces token generation by 80% on simple tasks
Diagnoses root cause of LLM agent long-horizon planning failures (stepwise reasoning induces greedy policy); FLARE (Future-aware Lookahead + Reward Estimation) lets LLaMA-8B surpass GPT-4o on planning benchmarks
Semi-formal reasoning using structured templates requiring explicit evidence — achieves 87% accuracy on code QA, 9 pp gain over standard agentic reasoning; enables interpretable code understanding for complex reasoning tasks
Contextual changes cause reasoning models to compress traces by up to 50%, reducing self-verification; simple problems unaffected but harder tasks suffer — critical finding for agent multi-turn reasoning
Challenges "SFT memorizes, RL generalizes" — reasoning SFT with long CoT does generalize cross-domain, conditional on optimization dynamics; discovers safety-reasoning tradeoff (reasoning improves but safety degrades); 152 HF likes
Identifies "template collapse" in agentic RL — models rely on fixed input-agnostic templates despite stable entropy; proposes mutual information (not entropy) as diagnostic for reasoning quality; Northwestern/Stanford/Microsoft; 49 HF likes
Google DeepMind: first systematic study of whether LLMs produce optimal plans (not just valid); reasoning-enhanced LLMs significantly outperform classical satisficing planners (LAMA) in complex multi-goal configurations
S³: inference-time procedure maintaining a population of partial denoising trajectories with verifier-based look-ahead and reward-tilted Gibbs distribution — first principled test-time scaling for discrete masked diffusion LMs
Side-by-Side (SxS) Interleaved Reasoning — makes disclosure timing a controllable decision in autoregressive generation; interleaves partial disclosures with continued private reasoning, releasing content only when supported by reasoning so far; improves accuracy–latency Pareto trade-offs on…
Google DeepMind: interactive workbench for open-ended mathematical research — ideation, literature search, computational exploration, theorem proving, theory building; manages uncertainty, tracks failed hypotheses, outputs native mathematical artifacts; scores 48% on FrontierMath Tier 4, a new…
Full overview of discrete / continuous / hybrid prompt optimization
Comprehensive survey unifying memory, skills, protocols, and harness engineering as four forms of "cognitive externalization" — traces progression from weights → context → harness using cognitive artifact theory; Shanghai Jiao Tong / UCL
Comprehensive survey treating context enrichment as a continuum — from in-context learning through RAG, GraphRAG, to CausalRAG; includes claim-audit framework and cross-paper evidence synthesis
Comprehensive survey of credit assignment methods for LLM RL (reasoning + agentic) — covers 47 papers from Jan 2024 to Apr 2026; traces shift from reasoning-focused to agentic/multi-agent CA methods
Comprehensive taxonomy of RAG security — poisoning, extraction, membership inference, jailbreaks, and privacy leakage attacks with corresponding defense strategies and future research directions
Graph-structured retrieval enabling multi-hop reasoning
Model decides when and how to retrieve
Agents embedded in RAG pipelines — dynamic, reasoning-driven retrieval beyond static pipelines
Hierarchical retrieval interfaces enabling agents to dynamically navigate multi-level knowledge structures
Meta AI: RAG for reasoning — decomposes trajectories into 32M reusable subquestion-subroutine pairs; retrieves procedural "how-to" knowledge within reasoning traces; +19.2% across math/science/coding
First Systematization of Knowledge for Agentic RAG — formalizes retrieval-generation loops as finite-horizon POMDPs; multi-dimensional taxonomy covering planning strategies, retrieval orchestration, memory paradigms, and tool coordination
RUC: file-based visual context management + progressive on-demand image loading — scales to 100-turn search horizons, SOTA on MM-BrowseComp and MMSearch-Plus
12 concrete reliability metrics across consistency, robustness, predictability, safety — capability gains ≠ reliability gains
Comprehensive survey: 3-layer framework (single-agent capabilities → self-evolving agents → multi-agent coordination); 202 Hugging Face likes
Decomposes web agent behavior into high-level planning, low-level grounding, and replanning — PDDL-structured plans outperform NL plans but grounding remains the dominant bottleneck; a single round of exploratory replanning substantially improves task success
End-to-end evaluation suite with 300 human-verified tasks across 9 categories — trajectory-aware grading over 2,159 rubric items; finds vanilla LLM judges miss 44% of safety violations and 13% of robustness failures
Benchmark built from 150 regulated prediction markets evaluated at 5 lifecycle checkpoints — models are most competitive early and on high-uncertainty markets; search improves pooled accuracy but degrades 12% of conditions
3D reliability surface R(k,ε,λ) unifying consistency, robustness, fault tolerance — chaos engineering for agents; ReAct outperforms Reflexion under stress; pass@1 overestimates reliability by 20–40%
Stanford: Python substrate that makes agent execution a first-class object — typed events, Git-like trace, deterministic fork/replay/intervene primitives; 5× faster fork than Docker, >95% prompt-cache reuse; CooperBench pair-coding success 28.8% → 54.7%, 58% lower wall-clock on TerminalBench-2
Tsinghua / Zhipu AI: argues the bottleneck in autonomous discovery is the environment, not the agent workflow — four environment-engineering dimensions (permissions, artifacts, budget, human-in-the-loop) enable off-the-shelf CLI agents to set SOTA on math, kernel engineering, and ML tasks at low…
UC Santa Cruz / MIT: six-state control-decision taxonomy and trajectory-failure vocabulary for separating outcome success from control-decision and trajectory quality; explicit label menus account for 14–40 pp of apparent agent capability
Defines loop engineering as a new layer above prompt, context, and harness engineering — loop spec anatomy (trigger, goal, five-level verification ladder, architecture, stopping rule, memory), design principles, and anti-patterns from a corpus of 50 real loops
Reconstructs prompt-dominant enterprise prototypes into code-owned, auditable harnesses — source-to-claim pipeline, seven validation dimensions (grounding, routing, trace, hygiene, recommendation language, runtime interfaces, latency), and the principle that "prompts are not guardrails"; validated…
HERA: 3-layer hierarchical framework that jointly evolves global orchestration strategies and local agent behaviors using experiential knowledge — role-aware prompt optimization drives targeted improvements for each agent's responsibilities
Brings credit assignment and policy gradient evolution from cooperative MARL into language space — enables LLM agents to autonomously evolve coordination strategies in dynamic environments
Reformulates topology selection as cooperative MARL — each agent selects communication actions that jointly induce round-wise communication graphs; improves coordination efficiency
LLM agents tend to cooperate in multi-round, non-zero-sum contexts rather than Nash equilibria — insights for designing cooperative multi-agent systems
Replaces free-text agent messages with explicit graph operations (traversal, subgraph fragments, updates) over a shared knowledge graph — 73% token reduction, 34% accuracy improvement, fully auditable reasoning chains
Topology selection (parallel/sequential/hierarchical/hybrid) matters more than model choice — AdaptOrch automatically picks the right topology per task; 12–23% improvement over static single-topology baselines across SWE-bench, GPQA, and RAG
Systematic academic analysis of MCP and A2A as complementary communication protocols; enterprise-grade multi-agent orchestration architecture covering governance, observability, and organizational adoption patterns
Progressive recursive delegation framework where search nodes pair local objectives with search modes, pass evidence upward, and recycle shared experience across sibling nodes; outperforms single-agent baselines on BrowseComp-Plus, WideSearch, DeepWideSearch, and GISA
Meta FAIR: task agent and meta agent unified in a single editable program — meta layer can modify itself (recursive self-improvement); validated on code, paper review, robotics, and olympiad math; 2.1k HF likes; open source (facebookresearch/HyperAgents)
Skill Generator iteratively refines agent skills while a Surrogate Verifier co-evolves to provide actionable feedback without ground-truth; surpasses human-written skills on SkillsBench in 5 rounds; works on Claude Code and Codex
Every agent interaction generates a next-state signal (user reply, tool output, GUI state) — OpenClaw-RL recovers all of them as live RL training sources via Hindsight-Guided On-Policy Distillation; one unified policy trains across conversation, terminal, SWE, and GUI tasks simultaneously (145 HF…
Continual meta-learning framework that jointly evolves a base LLM policy and a reusable skill library — skill-driven fast adaptation from failure trajectories + opportunistic gradient updates during idle periods; 21.4% → 40.6% accuracy on benchmarks (134 HF likes)
Framework enabling autonomous multi-agent evolution via persistent memory, asynchronous execution, and collaborative exploration — 3–10x higher improvement rates with fewer evaluations than evolutionary baselines; 251 HF likes
Cross-user trajectories continuously aggregated and refined by autonomous evolver into shared skill repository — collective skill evolution in multi-user agent ecosystems; 142 HF likes
Progressively withdraws skill documentation during training until agents operate zero-shot — +9.7% on ALFWorld, +6.6% on Search-QA with <0.5k tokens per step; 133 HF likes
Read-Write Reflective Learning over executable skill libraries — agents retrieve, execute, reflect, and rewrite their own skills without retraining the base model; evaluated on HLE and GAIA
120 adversarial scenarios across 5 high-privilege domains (SWE/finance/medical/legal/DevOps), 3 injection channels (skill files, email, web); 40–75% attack success rate; safety depends on model + framework stack, not model alone
DDIPE attack embeds malicious logic in skill documentation code examples; 1,070 adversarial skills across 15 MITRE ATT&CK categories; 11.6–33.5% bypass rate; responsible disclosure led to 4 confirmed vulnerabilities and 2 patches
First benchmark across 4 real functional domains (Web, Mobile, Embodied VLM/VLA) with 9 safety-risk categories; even the best agent completes <40% of tasks under full safety constraints
Two-week red-team study of live autonomous agents (email, Discord, shell, persistent memory) — documents 11 real attack categories including cross-agent unsafe practice propagation, identity spoofing, unauthorized resource consumption, and false task completion (32 HF likes)
Safety benchmark for browser/computer-use agents focused on long-horizon tasks where risk accumulates across many UI actions — useful for testing confirmation discipline, phishing resistance, and context drift
Introduces TVD framework and ISC-Bench — frontier models fail at 95.3% rate on dual-use professional tasks where capability and harm co-occur; advanced models are more vulnerable than earlier LLMs because their capabilities become liabilities
First unified survey spanning both LLM and VLM jailbreak — covers template, in-context, RL, and multimodal attack types; proposes 3-layer defense framework (perception / generation / parameter layers)
Dawn Song (UC Berkeley) et al. — first complete security survey for agentic AI systems (LLM + external tools/components); establishes threat model covering full attack surface and defense mechanisms; USENIX Security 2026
Greshake/Xiao/Suh et al. — security architecture paper arguing prompt injection must be handled at the system layer (permissioning, provenance, policy isolation), not by model alignment alone
Argues that prompt-based safety is architecturally insufficient for agents with execution capability; introduces Parallax, a plan-then-execute separation architecture with formal safety guarantees
Comprehensive threat model for world-model-equipped agents — adversarial attacks, goal misgeneralisation, deceptive alignment, automation bias; extends MITRE ATLAS and OWASP to world model stack
Demonstrates how attacks can autonomously propagate across interconnected LLM agents — worm-like self-spreading malware targeting agent ecosystems via MCP, tool chains, and shared memory
First systematic study of persistent memory poisoning — maps 4 write channels, 9 structural vulnerabilities, and 6 attack classes; introduces MPBench; shows current prompt-injection defenses are insufficient against cross-session memory manipulation
New category of indirect prompt injection in which malicious data is disguised as trusted data (metadata, tool outputs, context structures, identifiers), bypassing existing IPI defenses; demonstrates real-world attacks on web and coding agents including Claude Code, Codex, and Gemini CLI
Comprehensive review of medical reasoning methods + MR-Bench (real-world hospital data); reveals large gap between exam-level performance and authentic clinical decision-making
Truth-preserving patient simulation framework injecting controllable, clinically evidence-grounded noise — evaluates medical AI robustness under realistic imperfect patient data conditions
Minimal evidence extraction for medical AI explanations — identifies the smallest subset of input features sufficient for model decisions, improving interpretability without performance loss
Hierarchical fine-grained criteria modeling for medical LLM alignment — structured clinical evaluation rubrics with multi-level criteria decomposition for improved medical reasoning and safety
Exploratory study of LLM self-correction in medical QA — finds reflection can both correct and introduce errors; analyzes error correction dynamics across multiple reflection steps on MedQA, HeadQA, PubMedQA
MIT/Harvard: mixed-vendor multi-agent diagnosis outperforms single-vendor teams — complementary inductive biases surface correct diagnoses that homogeneous teams miss; SOTA on RareBench and DiagnosisArena
Focus agent architecture — autonomously consolidates history into a Knowledge block and prunes stale context; 22.7% token reduction on SWE-bench Lite, no accuracy loss
ACE treats contexts as evolving playbooks with Generator/Reflector/Curator roles and incremental delta updates; defeats brevity bias and context collapse; +10.6% on agent benchmarks, +8.6% on finance; Stanford/CMU/Salesforce
Defines context engineering as a standalone discipline for agentic AI; proposes a four-level maturity pyramid (Prompt Engineering → Context Engineering → Intent Engineering → Specification Engineering) and five context-quality criteria (relevance, sufficiency, isolation, economy, provenance)
First to unify LTM (add/update/delete) and STM (retrieve/summarize/filter) as tool-based actions via GRPO RL; 7B model achieves +49.59% over no-memory baseline across 5 benchmarks; ICLR 2026 MemAgents Workshop
End-to-end trainable sparse attention with linear complexity — scales to 100M tokens on 2×A800 GPUs with <9% degradation vs 16K baseline; Memory Interleaving enables multi-hop reasoning across scattered segments
Decomposes agent memory into 4 modules (extraction, management, storage, retrieval); systematic benchmark comparison of all methods; composite design from existing modules surpasses prior SOTA
Tsinghua / HKUST / SJTU: first data-management study of agent memory — 12 systems + 2 baselines across 5 workloads and 11 datasets; four-module framework (representation/storage, extraction, retrieval/routing, maintenance); finds no single architecture dominates and localized maintenance…
First benchmark focused on whether coding agents retrieve the right repository context before editing — measures relevance, latency, and downstream task success under realistic codebase navigation pressure
First large-scale empirical study of prompt compression trade-offs in production — 30K queries across multiple LLMs and 3 GPU classes; LLMLingua achieves up to 18% end-to-end speedup when prompt/ratio/hardware match; ECIR 2026; includes open-source profiler for latency break-even prediction
Memory mechanism that retrieves compressed reasoning "thoughts" rather than raw context — enables more efficient and reasoning-aware memory for long-horizon agents
Hierarchical graph-structured memory with role-aware modulation and temporal/confidence weighting; training-free, evaluated across multiple model scales
Context-ReAct paradigm with five atomic operations (Skip, Compress, Rollback, Snippet, Delete) for adaptive context management; proves expressive completeness of Compress; LongSeeker achieves 61.5% on BrowseComp and 62.5% on BrowseComp-ZH, substantially outperforming Tongyi DeepResearch and…
VISTA: typed, addressable context blocks + runtime proprioceptive dashboard (token usage, recency, access history, context pressure) + recoverable full-fidelity archive; training-free and model-agnostic; raises Gemini-3-Flash from 22.7% to 50.7% on LOCA-Bench, with gains on BrowseComp-Plus and GAIA
200-task benchmark across 12 constraint categories (resource, behavior, toolset, response) with step-level validation; no model exceeds 20% completion; models violate constraints in >50% of cases with limited self-correction
Comprehensive framework for understanding tool use in agentic systems — schema understanding, calling conventions, error handling, tool composition patterns
OpenTools: standardized tool schemas and lightweight wrappers for plug-and-play use across agent frameworks; intrinsic evaluation suite tracking correctness, robustness, regressions
Alibaba: addresses meta-cognitive deficit where agents blindly invoke tools — HDPO framework reduces unnecessary tool invocations from 98% to 2% while increasing reasoning accuracy; first paper on "when NOT to use tools"
Unified survey from single-tool call to multi-tool orchestration — covers reasoning-time planning, training/trajectory construction, safety, resource efficiency, open-environment completeness, and benchmark design (HIT & Harvard)
Evaluates whether agents can use actual Model Context Protocol servers rather than toy tool schemas — measures correctness, protocol handling, and real-world MCP interoperability
Lightweight signal-based taxonomy for sampling informative agent trajectories post-deployment — 82% informativeness vs 54% random; organizes signals across interaction, execution, and environment dimensions; 6.2k HF likes
Shifts evaluation from simple QA to multi-turn agentic assessment; newer benchmarks like SWE-bench Verified and Terminal-Bench test iterative agent behavior with execution feedback
Evaluates whether LLM agents maintain strategic coherence over long horizons — simulated startup over one-year horizon spanning hundreds of turns; tests consistent execution
Tests agent ability to handle user interruptions during mid-task execution — critical requirement for realistic deployment in dynamic environments
First CI-loop benchmark for long-term codebase maintainability — 100 tasks spanning 233 days and 71+ consecutive commits; shifts evaluation from static single-fix to dynamic long-horizon reasoning
565 real-world SE tasks measuring whether agent skills actually improve outcomes — 39/49 public skills give zero gain; average improvement only +1.2%; reveals fundamental gap in skill design
Benchmarks terminal-based coding agents on long-horizon programming tasks that require sustained planning, repo navigation, debugging, and recovery over many steps instead of single-fix patches
Evaluates whether agents can build complete software projects from requirements to implementation and validation, rather than solving isolated bug-fix tasks; targets end-to-end project delivery realism
Evaluates agents on compositional, real-world assistant tasks requiring planning, tool use, and recovery — closer to production deployment scenarios than static QA benchmarks
Realistic interactive benchmark for GUI agents in high-stakes professional workflows — 100 real-world e-commerce risk scenarios testing sequential decision-making under uncertainty
100 professional task scenarios across 10 industries and 65 domains — evaluates AI agents on realistic occupational workflows using language world models for environment simulation
Benchmarks multimodal agents on episodic scientific research workflows — literature search, figure extraction, cross-paper synthesis; built on smolagents with persistent memory and tool use
First forced-injection framework measuring how clarification value changes over the execution trajectory across goal/input/constraint/context dimensions; 6,000+ runs, 4 frontier models, 3 benchmarks; finds goal clarifications lose nearly all value after 10% execution, input clarifications retain…
ICML 2026: controlled comparisons show reasoning judges substantially improve accuracy on structured-verification tasks (math, coding) but yield limited or negative gains on simpler evaluations while costing significantly more compute; proposes RACER, a distributionally-robust routing policy that…
Modular benchmark with up to 20 application-oriented generation constraints per prompt; finds compliance degrades with constraint count and position (primacy/recency bias) — exposes multi-instruction conflict effects
Rubric-based RL with Token-Level Relevance Discriminator — solves credit assignment for instruction following by predicting which tokens satisfy specific constraints; fine-grained optimization
Discovers that schema key wording itself acts as an implicit instruction signal under constrained decoding — changing JSON key names alters model behavior even when semantic content is identical
Trivial lexical constraints (banning one punctuation mark) cause 14–48% response collapse in instruction-tuned LLMs — identified as planning failure via mechanistic analysis; base models show no collapse
Formalizes Compositional Behavioral Leakage (CBL) — prompt modules sharing a context window silently shift each other's behavior; introduces a three-channel perturbation protocol (volume / content / form) and detects Cohen's d = 0.63 content-channel interference in a deployed job-evaluation agent;…
NSHA: formulates hierarchical instruction resolution as constraint satisfaction, solved with SAT solver-guided inference-time reasoning — resolves conflicts between system prompts, user instructions, and tool outputs
Distribution-guided efficient fine-tuning for alignment — uses data distribution properties to guide selective parameter updates, improving alignment quality with reduced compute
Spatial reasoning as spatio-temporal evidence accumulation — VLM planner + hierarchical 2D/3D spatial tools + dual memory; training-free gains on open-source and closed-source VLMs; S-Agent-8B matches GPT-5.4 and Gemini 3 on spatial benchmarks
Overlays scene graphs onto input images at the pixel level to model object relationships — up to +11 percentage points on VQA and localization across 4 datasets, zero-shot
Inference-time framework exploiting MLLM attention patterns to identify relevant visual regions and text, then re-conditions generation on highlighted evidence — consistent VQA improvements, no training required
Systematic evaluation of agentic capability in multimodal LLMs — decomposes tasks into perception, reasoning, and action levels; reveals where agentic loops help vs. where they add overhead
First benchmark for Feynman diagram tasks — evaluates multistep diagrammatic reasoning requiring conservation laws, symmetry constraints, and graph topology; 2000+ tasks across Standard Model interactions
Benchmark for multimodal evidence retrieval and multi-hop reasoning over noisy web content — even strongest agent (Gemini-3.1-Pro) achieves only 40.1%; finds more search ≠ better performance
Converts inference-time zooming into training-time primitive — teaches MLLMs fine-grained perception in single forward pass; introduces ZoomBench (845 VQA across 6 perceptual dimensions); SOTA on fine-grained benchmarks
Unifies predictive imagination with reflective reasoning for driving foresight — action-derived trajectory guides next-frame generation, then reasons over the imagined frame to refine planning
Conversational framework for embodied AI development — batch simulation environment synthesis, automatic scene creation, controllable scene editing, and workflow execution via natural language
Open-source modular VLA framework — swappable backbone (VLM/world-model) and action heads, cross-embodiment learning, unified evaluation across LIBERO, SimplerEnv, RoboTwin, RoboCasa, BEHAVIOR-1K
Comprehensive survey of human-to-robot imitation learning — behavioral cloning, inverse reinforcement learning, adversarial imitation, and their combinations; includes taxonomy, benchmarks, and open challenges
100 detail-oriented embodied AI tasks spanning manipulation, navigation, and reasoning — evaluates fine-grained physical world understanding beyond coarse task completion
First unlearning method for VLA models — removes target behaviors while preserving general capabilities; introduces forget/retain/boundary splits and real-robot OXE benchmarks
Salesforce AI Research: complete tutorial for production voice agents — cascaded streaming pipeline (STT→LLM→TTS), ~750ms TTFA, function calling, full open-source codebase with 9 chapters
Tsinghua: long-term memory for realtime voice agents — "left brain" stores compressed facts (Mem0-level accuracy at ~300 tokens/query), "right brain" tracks emotional attribution across nodes; fully streaming architecture with speculative prefetch keeps added latency near zero; open model family +…
Data ingestion and RAG pipelines
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF to clean Markdown — Rust core with Node.js/Python bindings; agent/RAG document ingestion (Aug 2026)
Open-source HTML-to-video rendering framework built for agents — write HTML/CSS with seekable animations and render deterministic MP4s; agent skills, CLI, and hosted authoring workflows (Mar 2026)
Compress tool outputs, logs, files, and RAG chunks before they reach the LLM — 20% fewer tokens for coding agents, 60–95% fewer for JSON; ships as a library, proxy, and MCP server (Jan 2026)
Permanent memory for AI agents — append-only log + binary-tree summaries, 426-token prompt, plug-and-play with Claude Code/Codex/etc. via AGENTS.md/CLAUDE.md (July 2026)
Turn any codebase — plus docs, SQL schemas, configs, and PDFs — into a queryable knowledge graph. Ships as a /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI; local deterministic AST parsing, every edge explained, no vector store (Apr 2026)
Microsoft's LLM SDK — now merging with AutoGen into Microsoft Agent Framework (2026)
LLM gateway + observability + optimization
Official Pydantic agent runtime — typed tools, structured outputs, evals, production-ready (V1 stable)
Most widely used library for structured LLM outputs — typed extraction from any model, 3M+ monthly downloads
Realtime voice runtime for AI agents — keeps agents talking, working, and present while they think or use tools (no dead air during tool calls); pluggable STT/TTS and realtime providers, embeddable gateway, TUI/desktop apps, Agent Skills support, ACP-compatible; works with Claude Code, Codex,…
EleutherAI's unified LLM evaluation framework
Experiment tracking and LLMOps
Comprehensive prompt engineering reference (DAIR-AI)
Most comprehensive list of 2026 AI agents, frameworks & tools — 300+ resources, 20+ categories, updated monthly
Curated papers on LLM agents: methodology, applications, challenges — covers STRIDE, planning, tool use, memory, multi-agent (2026)
Papers and resources on agentic reasoning from foundational to multi-agent coordination — 3-layer framework (2026)
Curated papers on memory architectures for LLM agents — long-term, short-term, attention mechanisms (2026)
Curated 2025–2026 papers on agent engineering, memory, eval, and workflows
Claude-optimized prompts — XML tags, extended thinking, long-context patterns
Prompts for OpenAI Deep Research, Gemini Deep Research, Perplexity Labs
Curated papers on diffusion language models — LLaDA, Dream, MMaDA, consistency sampling, fast inference; 169 stars, actively maintained (2026)
Official production-ready prompts from Anthropic
The most complete open-source AI engineering curriculum — 523 lessons / 20 phases / ~342 hours; dedicated prompt engineering, agent engineering, MCP, and Agent Skills phases where every lesson ships a reusable artifact (prompt, skill, agent, MCP server); Python/TypeScript/Rust, MIT, 55k+ stars,…
22 Jupyter Notebook tutorials from basics to advanced — CoT, few-shot, templates, multi-language
152 installable Claude skills for automotive engineering — ISO 26262, ISO/SAE 21434, ISO 21448 SOTIF, AIAG-VDA, ASPICE, AUTOSAR; builder + reviewer pairs with xlsx deliverables
VoltAgent/awesome-openclaw-skills
The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
awesome-dsh-plugin/awesome-dsh-plugin
A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表
Kristories/awesome-guidelines
Programming style, best practices, and coding conventions.
sindresorhus/awesome
😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]
matiassingers/awesome-readme
A curated list of awesome READMEs
oz123/awesome-c
A curated list of awesome C frameworks, libraries, resources and other shiny things. Inspired by all the other awesome-... projects out there.