Harness Engineering
OpenAI's framing of harness engineering as a discipline: how to design the scaffolding that lets Codex and similar agents operate reliably in an agent-first world.
Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration.
This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.
OpenAI's framing of harness engineering as a discipline: how to design the scaffolding that lets Codex and similar agents operate reliably in an agent-first world.
OpenAI's detailed breakdown of the Codex agent loop, exposing each harness component and where it can be improved.
OpenAI's practice guide for long-horizon task planning: introduces Plan.md, Implement.md, Documentation.md as reusable harness artifacts.
Anthropic's foundational guide on agent architecture, covering when to use workflows vs. agents and how to compose primitives.
Anthropic's engineering blog on designing harnesses for sustained, multi-session development tasks. Key insight: every harness component assumes the model can't do something; those assumptions expire.
Anthropic's guide on tool interface design: naming, schemas, error surfaces, and the principle that tool design is agent UX.
Anthropic on building structured permission and authorization systems into agent harnesses instead of relying on natural-language permission text.
Anthropic's framework for evaluating agent behavior: what to measure, how to build eval harnesses, and why unit-test-style evals fail for agents.
IBM's definitional piece, useful for anchoring harness design decisions to a clear model of what an agent actually is.
Google's announcement and design rationale for ADK: explains the multi-agent topology, tool registration model, and eval pipeline that shaped their framework. Complements the Anthropic/OpenAI framing with Google's production perspective.
Martin Fowler's synthesis of what harness engineering practice looks like: three interlocking systems — context engineering (curating what the agent knows), architectural constraints (deterministic linters and structural tests), and entropy management (periodic agents that repair documentation…
LangChain's structural breakdown of the five primitives that compose a harness: filesystem (durable state + agent collaboration surface), code execution (autonomous problem-solving without pre-designed solutions), sandbox (isolation + verification), memory (cross-session persistence), and context…
The first systematic practitioner paper on terminal-native coding agent harness design: eager-construction scaffolding (pre-build all components before the first message to eliminate first-call latency and race conditions), compound multi-model architecture (different model instances for…
Proposes externalizing agent control logic as portable natural-language artifacts (NLAHs) executed by a shared Intelligent Harness Runtime, enabling harness design to be studied, transferred, and reproduced rather than buried in bespoke controller code. Directly addresses the root cause of harness…
Meta's production harness for multi-day ML pipeline automation with hibernate-and-wake checkpointing for resuming interrupted 6-hour tasks without losing context. Demonstrates harness design for scientific workflows where individual turns can exceed model context limits but the overall pipeline…
Google's 2026 update to Agent Development Kit expanding the ecosystem integrations (Hugging Face, GitHub, Daytona, Notion, etc.) and providing reference patterns for how orchestration harnesses wire external services without losing determinism or state coherence.
Anthropic's industry benchmark identifying infrastructure configuration as a first-class optimization variable: harness setup alone can swing benchmarks by 5+ percentage points. Documents the shift from single-agent to orchestrated multi-agent teams and introduces the "agentic engineering…
Architecture walkthrough of Microsoft's agent that has handled 35,000+ production incidents autonomously, reducing Azure App Service time-to-mitigation from 40.5 hours to 3 minutes. Documents the integration of MCP tools, telemetry, code repositories, and incident management platforms into a…
Microsoft's account of shifting from 100+ bespoke tools and a prescriptive prompt to a filesystem-based context engineering system for their SRE agent. Key finding: exposing everything (source code, runbooks, query schemas, past investigation notes) as files and letting the agent use read_file,…
Red Hat's enterprise perspective on harness engineering (April 7, 2026): AI writes better code when you design the environment it works in. Emphasizes structured context over free-form tickets, expanding the agent's toolbox through MCP integrations (CI status, deployment logs, runtime metrics) as…
Birgitta Böckeler's systematic mental model (April 2026) for coding-agent harnesses, framing them as feedforward guides plus feedback sensors that self-correct before output reaches human eyes. Distinguishes computational controls (linters, tests) from inferential ones (LLM-as-judge), and argues…
OpenAI's April 2026 comprehensive guide distilling production deployment patterns into actionable best practices: single-agent vs. multi-agent orchestration (manager vs. decentralized handoffs), tool design for many-to-many agent-tool relationships, and layered guardrail patterns combining input…
Anthropic's transparent April 2026 postmortem tracing Claude Code quality degradation to three independent harness-level changes: a default reasoning-effort downgrade, a caching-optimization bug that continuously dropped thinking history from stale sessions, and an overly aggressive…
Anthropic's April 2026 design guide distilling harness engineering into three actionable patterns: build on tools Claude already knows, remove harness assumptions as capabilities improve, and set UX/cost/safety boundaries carefully. A practical complement to the "agent = model + harness" framing…
May 2026 survey framing code as the basis for agent infrastructure rather than merely output: it unifies harness interface, mechanisms, and multi-agent scaling through shared code artifacts, and surfaces open challenges from verification under incomplete feedback to regression-free improvement.
deepset's May 2026 synthesis of agent reliability as a harness problem: a failure-classification framework (context, constraint, verification, planning failures) that maps each failure mode to the right harness component, and a concrete demonstration that harness-only changes can move agents 20+…
June 2026 constitutive definition of an agent harness as a runtime layer with four necessary and sufficient elements: an agent loop, a tool interface, context management, and control mechanisms. Applied to Claude Code, Codex CLI, Aider, Cline, OpenHands, and SWE-agent, it provides a rigorous…
April 2026 empirical study of 70 public agent systems across five recurring dimensions (subagent architecture, context management, tool systems, safety mechanisms, orchestration) that synthesizes five architectural patterns. The comparative research package turns harness selection from a framework…
RUCAIBox's survey paper and curated reading list on Agent Systems with Harness Engineering, mapping harness design across agent workflows, memory systems, skill libraries, and multi-agent orchestration with 500+ references. The clearest academic complement to vendor-specific harness engineering…
LangChain's July 2026 playbook showing how harness-only tuning brought Nemotron 3 Ultra within one point of Opus 4.8 on Deep Agents at roughly one-tenth the cost ($4.48 vs $43.48). The clearest recent demonstration that evals are the training data for harness work and that fit — not raw model…
Ryan Lopopolo's anthology, field guide, and agent context bundle for harness engineering: it reframes the harness as the environment that carries an organization's nonfunctional requirements, with reusable AGENTS.md/CLAUDE.md artifacts, playbooks, evals, and domain modeling docs. The most…
Lilian Weng's July 2026 synthesis framing the harness as the deployment layer of recursive self-improvement: three design patterns (workflow automation, filesystem as persistent memory, sub-agents/backend jobs), a case study of the tool surface stabilized across Claude Code, Codex, OpenCode, and…
The foundational paper defining the Thought/Action/Observation loop structure that underlies virtually every agent harness. Required reading for understanding why the loop is structured the way it is and where each harness component maps onto the reasoning-acting cycle.
OpenAI's detailed breakdown of the Codex agent loop, exposing each harness component and where it can be improved.
Models the agent loop explicitly as a directed graph with typed state, conditional edges, and checkpointing. The most concrete engineering treatment of loop control flow: how to implement termination conditions, branch on tool results, and persist mid-loop state for resumption.
OpenAI's engineering deep-dive into the Item/Turn/Thread protocol (JSON-RPC/JSONL over stdio) that exposes the Codex harness to every client surface. The most direct first-party account of why approval flows, streaming diffs, and thread persistence demand a purpose-built protocol — and why MCP's…
OpenAI's lifecycle-hook framework for Codex: inject deterministic scripts at SessionStart, PreToolUse, PostToolUse, and other loop events to enforce guardrails, audit actions, and customize agent behavior without relying on prompt-level trust. A concrete reference for programmable harness…
The harness-critical reference for integrating extended thinking into agent loops: budget_tokens controls reasoning depth per turn, thinking blocks must be preserved when passing tool results back (omitting them silently breaks multi-step reasoning), and thinking mode cannot change mid-turn.…
Anthropic's June 2026 practical taxonomy of agent loops: turn-based, goal-based (/goal), time-based (/loop, /schedule), and proactive loops. The framework for matching loop primitive to task shape — and the emphasis on deterministic stop conditions and token budgets — makes it a concise reference…
LangChain's case study showing harness-only changes moved their coding agent from rank 30 to top 5 on Terminal Bench 2.0 with no model swap: structured verification loops, context injection (directory maps + time budget warnings), loop-detection middleware, and a "reasoning sandwich" concentrating…
Practical design system for agent loops with seven production patterns, cross-tool starter kits, and CLI tools that score readiness, scaffold state, estimate cost, detect drift, and isolate worktrees. The clearest open-source resource for moving from one-off prompting to durable, observable agent…
Official implementation of a lifecycle-aware runtime harness that improves frozen LLM agents by adapting the model-environment interface across four layers: environment contract, procedural skills, action realization, and trajectory regulation. The key result is that harness-side adaptation…
Introduces AgentMiddleware: six composable hooks (before_agent, before_model, wrap_model_call, wrap_tool_call, after_model, after_agent) that intercept every stage of the agent loop. Enables deterministic policy enforcement (PII redaction that can't be trusted to prompts), dynamic tool injection,…
Controlled experiment isolating interpreter state persistence as an independent training variable. The harness finding: mismatching your runtime persistence mode to the model's training-time semantics produces either 80% missing-variable errors (model expects state that doesn't persist) or 3.5×…
Demonstrates that temporal awareness (handling deadlines and time constraints) appears orthogonal to reasoning capability: explicit temporal feedback in the agent loop significantly improves LLM performance on deadline-constrained tasks. Indicates temporal semantics as a learned behavior that must…
April 2026 systematic analysis of 70 open-source LLM agent projects showing 60% adopt the Agent Loop pattern. Proposes a formal scheduler framework that maps execution patterns (Agent Loop, Event-driven, State-machine, Graph/flow, Hybrid) onto a unified control model, making the…
February 2026 production-grade coding agent from Meta/Harvard built on the Confucius SDK, which structures harness design around three perspectives: Agent Experience (AX), User Experience (UX), and Developer Experience (DX). Features a unified orchestrator with advanced context management,…
April 2026 reverse-engineering of Claude Code's architecture revealing five-stage progressive compaction (budget reduction → snip → microcompact → context collapse → auto-compact), subagent isolation with rebuilt permission contexts, and a 27-event-type hook pipeline. The most detailed public…
Ports Claude Code's full agent loop to DeepSeek V4 Pro and other Anthropic-compatible backends while preserving the same UX. The strongest practical evidence that loop architecture — not model identity — determines agent behavior, and a concrete starting point for building backend-agnostic…
VS Code team's breakdown of the coding harness behind GitHub Copilot: three core loop responsibilities (context assembly, tool exposure, tool execution), multi-provider model routing across Anthropic, Google, OpenAI, xAI, and Mistral, and the VSC-Bench eval suite with PR-gated assessment. The…
State machine guardrails that constrain which tools an agent can call in each phase of a workflow, turning open-ended loops into deterministic state transitions. The research result is striking: local models went from 2/10 to 10/10 passing on a SWE-bench subset purely by shrinking the tool space,…
Anthropic's May 2026 introduction to dynamic parallel subagent orchestration: Claude generates JavaScript orchestration scripts that fan out work to tens or hundreds of parallel subagents with adversarial verification, converging on answers for tasks like the 750k-line Bun Zig-to-Rust port. The…
UIUC's open-source specification and execution language for LLM-agent workflows: declarative YAML with typed steps, branching, loops, and explicit state management, backed by a Docker sandbox with 50+ MCP tools, checkpointing, and trajectory logging. A concrete reference for turning ad-hoc agent…
OpenAI's practice guide for long-horizon task planning: introduces Plan.md, Implement.md, Documentation.md as reusable harness artifacts.
Anthropic's engineering blog on designing harnesses for sustained, multi-session development tasks. Key insight: every harness component assumes the model can't do something; those assumptions expire.
The canonical engineering write-up separating planning from execution as distinct harness layers: a planner LLM generates the step list once; an executor agent works through it, replanning only when needed. Defines the pattern that most modern task-decomposition harnesses follow.
Code-first task decomposition framework with a planner/executor split and a plugin system for injecting domain knowledge into the planning layer. The most complete reference implementation of plan-then-execute with stateful task tracking.
Unifies reasoning, acting, and planning via Monte Carlo Tree Search over agent trajectories. Directly informs harness design: external tool feedback as tree-search signals, trajectory backtracking on failure, and depth-bounded exploration make this the most actionable planning research for…
Demonstrates specialized harness patterns for coordinating heterogeneous agent teams (planner, coder, reviewer, executor) on software engineering tasks. Shows how role-specific agents with different model sizes and tool access produce better outcomes than single-agent approaches, with concrete…
Modular framework separating high-level planning from low-level execution through synthetic data generation and explicit structured planning. Achieves 57.58% success on WebArena-Lite and 81.36% on WebVoyager. The key harness insight is that planner and executor can be specialized independently —…
Decision framework for four multi-agent patterns (subagents, skills, handoffs, router) with concrete performance data: subagents process 67% fewer tokens than skills in multi-domain scenarios because context isolation prevents cross-domain bloat. The five-dimension matching table (distributed…
GitHub's February 24, 2026 distillation of a failure pattern most harnesses eventually rediscover: multi-agent systems behave like distributed systems, so every handoff needs typed schemas, constrained action schemas, and explicit boundary validation. Worth including because it turns "add more…
Anthropic's pattern for maintaining agent progress across multiple context windows: an initializer agent sets up the environment once and hands off to a coding agent that makes incremental progress each session. The structured handoff mechanism — feature lists, git commits, and test gates as…
February 2026 framework that dynamically selects orchestration topology (parallel, sequential, hierarchical, or hybrid) based on task dependency graphs rather than fixed pipeline architecture. Demonstrates that topology choice is a harness-level lever that can improve performance 12–23% over model…
GitHub's open-source toolkit for spec-driven development: spec.md → plan.md → tasks.md artifacts convert natural-language intent into a machine-checkable plan before any code is written, with /specify, /plan, /tasks, and /implement commands that run across Claude Code, Copilot, Codex, and Gemini…
January 2026 planning framework that combines task decomposition with modular agent design: a Supervisor decomposes tasks into a dependency graph, Planner & Executor agents solve each decoupled sub-task node independently, and a Self-Revision module updates the graph after execution. The key…
OpenAI's framing of harness engineering as a discipline: how to design the scaffolding that lets Codex and similar agents operate reliably in an agent-first world.
Anthropic's systematic guide to managing the full context state—system prompts, tools, MCP, and message history—as a finite, curated resource. Reframes harness design as "what configuration of context produces the desired behavior?" rather than just prompt wording.
Anthropic's reference for server-side context compaction: automatically summarizes older context when approaching the window limit. Reduced token consumption by 84% in a 100-turn web search eval while allowing agents to complete workflows that would otherwise hit context limits.
Microsoft Research's prompt compression toolkit (up to 20x compression, minimal performance loss) that can be embedded as a preprocessing step in the context delivery layer. LLMLingua-2 adds 3–6x speed gains, making it viable for latency-sensitive agent loops.
The most effective harness-level cost lever: cache repeated system prompts, tool definitions, and long documents across requests. Explains where to place cache_control breakpoints for maximum reuse across multi-turn agent sessions.
Shifts context compression from harness-controlled (compacting at a fixed token threshold) to agent-controlled: agents call a dedicated tool to trigger compression when strategically appropriate — between tasks or before consuming large inputs. Eliminates the failure mode where reactive-at-limit…
Proposes a "Focus Agent" architecture where the agent autonomously decides when to consolidate interaction history into a persistent Knowledge block and prune raw context — shifting compression from a harness-enforced policy to a model-controlled action. Produces 22.7% token reduction with no…
MCP server that intercepts raw tool output before it enters the context window, sandboxing bulky data (Playwright snapshots, GitHub issues, logs) outside the LLM and retrieving only relevant fragments via BM25 when needed. The "think in code" paradigm — replacing ten file-read tool calls with one…
Vercel's February 3, 2026 implementation guide for serving text/markdown when agents request it via Accept: text/markdown, while preserving the same human-facing HTML URL. This is a real harness primitive, not just a docs trick: it removes boilerplate before it ever enters the context window and…
Reframes RAG as a harness tool-design problem: instead of injecting retrieved documents into context at pipeline time, expose three retrieval tools (keyword search, semantic search, chunk read) and let the agent pull information incrementally as each reasoning step requires it. The key harness…
Structured framework for building production-grade evaluation harnesses: evaluation gates that block deployment, observability instrumentation that tracks all agent decisions, and CI integration patterns that catch regressions before they reach users. Essential reading for organizations deploying…
LLM-curated hierarchical context management for agents where the model itself learns to weight information importance across multiple hierarchy levels. Reduces token overhead through learned relevance filtering without sacrificing comprehension. Directly applicable to any harness where context…
March 2026 deep-dive into Claude Code's automatic compaction mechanism: what survives (current task, recent errors, file names) vs. what gets lost (initial instructions, intermediate decisions, style rules). Key harness insight: never rely on compaction for critical rules — move them to CLAUDE.md…
MCP server that indexes codebases by symbol (functions, classes, call graphs) so agents navigate by pointer instead of reading whole files, cutting active tokens by 77% and benchmark wall time by 76%. Demonstrates that context delivery for coding agents is a navigation problem, not just a…
Replaces the bloated CLAUDE.md pattern with a progressive spec system: agents load only the standards, task PRDs, and session journals relevant to the current step. The cross-platform adapter layer turns vendor-specific harness configuration into a portable team practice rather than a per-tool hack.
ByteDance's context database for AI agents that unifies memory, resources, and skills through a filesystem paradigm, enabling hierarchical context delivery where agents pull only the paths they need instead of receiving bloated monolithic prompts. The self-evolving layer that restructures context…
Google Labs' specification for describing visual identity systems to coding agents: machine-readable design tokens (YAML front matter) combined with human-readable design rationale (markdown prose) give agents a persistent, structured understanding of design constraints without requiring custom…
High-performance code intelligence MCP server that full-indexes repositories into a persistent knowledge graph via tree-sitter AST analysis across 66 languages. Replaces dozens of file-read/grep cycles with sub-millisecond structured queries, cutting active tokens by 120× and turning codebase…
Mounts S3, Slack, Gmail, GitHub, and Redis side-by-side as a single virtual filesystem so agents interact with every backend through familiar bash commands instead of learning N distinct APIs. The key harness insight: LLMs are already fluent in grep, cat, and cp — leveraging that vocabulary…
Coding agent harness optimized for surgical context curation and API cost reduction: Hash Anchored edits, massively parallel operations, and AST manipulation combine to cut costs 50–80% while improving code quality. Demonstrates that precise context delivery — not just bulk compression — is the…
Code search primitive that replaces grep+read cycles with natural-language retrieval, cutting active tokens by ~98% while keeping 99% of a transformer-based retriever's accuracy. Ships as an MCP server and CLI, runs on CPU with zero external dependencies — the right drop-in for any coding agent…
Repository-level operating harness that turns any software repo into an agent-ready workspace: structured AGENTS.md, HARNESS.md, and FEATURE_INTAKE.md give agents the missing project context — where to start, what the product contract says, how risky the change is, and which decisions future…
Compresses tool outputs, logs, files, and RAG chunks before they enter the context window, cutting active tokens by 60–95% without changing answers. Ships as a library, proxy, and MCP server — the right drop-in layer for any harness where bulky tool returns are the primary context pressure source.
MCP server and CLI that injects up-to-date, version-specific library documentation directly into agent context, eliminating hallucinated APIs and outdated code examples caused by stale training data. Ships as both a ctx7 command-line tool and an MCP server with resolve-library-id and query-docs…
Decomposes code relevance into two interpretable dimensions — semantic evidence and dependency support — rather than collapsing all retention decisions into a single score. Saves up to 31% more tokens on multi-turn coding agent tasks while improving Exact Match by up to +3.5, demonstrating that…
LangChain's CLI that writes and maintains agent-readable wikis for codebases or purpose memory, turning documentation drift into a versioned, automatable harness artifact. Emits Google Open Knowledge Format bundles so curated context stays portable across agents and can be kept fresh via CI.
Programmatic memory framework for long-horizon agents: the harness appends all observations to a structured log and lets the agent search it with code instead of relying on fixed summarization or compaction policies. Achieves 4.2–5.8× token reduction and matches or exceeds specialized harnesses on…
Builds a local, regenerable graph of plain-English system explanations and code relationships, then rides along inside Claude Code, Cursor, Codex, and Gemini via MCP and statusline hooks so the agent stops rediscovering the repo every session. The published SWE-bench Verified and efficiency…
Self-improving executable context layer for data and analytics agents: it ingests warehouses, BI tools, and wikis to build a semantic layer with approved metrics, joinable columns, and resolved fan/chasm traps, then serves the result to Claude Code, Codex, and Cursor through MCP. Fills the gap…
LangChain's September 2026 introduction of context modes for subagents: isolated subagents start fresh (context isolation as a firewall), while forked subagents inherit the supervisor's full conversation — excised trailing tool call, task description rewritten as a user message, prompt caching…
NVIDIA's September 2026 extension for the Pi harness that packages four efficiency mechanisms discovered by running 152 candidate ideas through auto-research loops: Action Fusion (an edit runs its follow-up validation in the same tool call), ObservationPack (repeated large tool results become…
Anthropic's guide on tool interface design: naming, schemas, error surfaces, and the principle that tool design is agent UX.
Authoritative reference for client vs. server tool execution models, strict schema enforcement, and tool_result error signaling. The distinction between client-side and server-side tool execution is a foundational harness architecture decision.
Defines the de facto industry-standard JSON Schema conventions for tool definitions and parallel function calling. Essential reading before designing a tool interface that needs to work across multiple models.
The MCP team's definitive post on the four tool annotation hints (readOnlyHint, destructiveHint, idempotentHint, openWorldHint) as inputs to harness permission decisions, not enforced contracts. The "lethal trifecta" — private data access + untrusted content exposure + external communication — is…
Constrains token sampling via regex/CFG/JSON Schema at the decoding layer, guaranteeing structured output without model fine-tuning. The right solution when you need OpenAI Structured Outputs-equivalent reliability from a locally deployed or open-weight model.
Maps Pydantic models directly to structured LLM extraction with built-in retry and validation-error feedback loops. Turns tool call output parsing from ad-hoc JSON handling into type-safe data models, eliminating an entire class of harness parsing bugs.
Framework for evaluating agent skills on three dimensions (capability, robustness, security) before deployment. Directly addresses the harness problem of skill sprawl: as agents gain access to more tools, the combinatorial explosion of failure modes becomes unmanageable without systematic…
Google DeepMind technique that uses code synthesis to auto-generate runtime constraint harnesses from tool schemas and task specifications. Gemini-2.5-Flash + AutoHarness outperforms Gemini-2.5-Pro and GPT-5.2-High on TextArena games by eliminating illegal moves through learned harness policies.…
February 2026 analysis of how parallel tool calling reduces latency in multi-step agent workflows. Demonstrates that concurrent tool execution (rather than sequential observe→act loops) is the key efficiency lever for deep-research harnesses where each step may invoke search, browse, and compute…
April 2026 framework for deep-research agents using dedicated reasoning tools (plan_next_searches, select_query_and_search, extract_relevant_details, analyze_search_progress) that externalize intermediate decisions as typed tool arguments. Inspired by Anthropic's think-tool paradigm, Q+ makes…
March 2026 framework that models interaction topology — the structural patterns of how agents invoke, chain, and conditionally branch between tools — as a first-class training signal. Rather than treating tool use as isolated function calls, TopoCurate learns topological priors from expert…
March 2026 field report from an enterprise MCP deployment identifying three protocol-level gaps that break production: missing identity propagation (who is the request for?), absent adaptive tool budgeting, and unstructured error semantics. The concrete mitigation patterns — JWT-enriched tool…
Expands the agent tool surface beyond non-interactive commands: programmable TUI interaction for REPLs, debuggers, and ncurses apps that standard bash can't reach. A concrete harness primitive for any agent that needs to operate interactive CLI tools without building a custom wrapper per program.
Generates agent-native CLI harnesses for any software, giving agents structured JSON access to applications that were never designed for automation. The CLI-Hub registry and auto-generated SKILL.md files turn tool expansion into a package-manager experience — solving the "long tail" of agent tool…
Experimental graph-first programming language where agents inspect and edit code through a compiler-derived ProgramGraph (node IDs, graph hashes, types, effects, ownership) instead of fragile text patches. Collapses the typical agent loop of edit-format-reparse-check-fix into a single…
Anthropic's open protocol for connecting agents to external tools, data sources, and services in a standardized way.
Anthropic's official reference MCP server implementations (GitHub, Slack, Postgres, Puppeteer, etc.). The authoritative source for understanding correct MCP server structure before building your own.
Browser automation via accessibility tree snapshots rather than screenshots, dramatically reducing token cost. The canonical example of structured tool output design in an MCP server.
Official Google MCP server that exposes live Chrome debugging surfaces — network analysis, performance profiling, console messages, memory snapshots, and Lighthouse audits — as structured agent tools. The clearest reference for turning browser inspection into a first-class tool interface rather…
MCP-native control layer for iOS and Android devices: snapshots, semantic targeting, typed client access, diagnostics, and replayable workflows. Fills a critical gap in the mobile-agent harness stack — most tool design assumes desktop or browser surfaces, but real-world agents increasingly need to…
Google's open Agent-to-Agent protocol: JSON-RPC over HTTP(S)/SSE with Agent Card service discovery and a task/message/artifact communication model. The emerging standard for cross-framework agent interoperability in multi-agent harnesses.
Google's June 2026 open specification for publishing, discovering, and verifying AI capabilities across the web via domain-owned catalogs and searchable registries. Adds the missing discovery layer that lets agents find MCP servers, A2A agents, and OpenAPI tools at runtime rather than relying on…
Interactive debugging UI for MCP servers: inspect tool definitions, send test calls, and validate responses without wiring up a full agent. The essential development tool for anyone building or integrating MCP servers into a harness.
OpenAI's engineering guide to three production harness primitives: versioned Skill bundles (SKILL.md manifest; routing accuracy improved 73%→85% by adding negative examples), a managed shell container for durable tool execution, and server-side compaction via explicit /responses/compact endpoint.…
Wraps 250+ SaaS APIs (GitHub, Slack, Linear, Notion, etc.) as agent-ready actions with managed OAuth, so tool integration becomes a one-line import rather than a custom harness component per service. The fastest path from "the agent needs to call an external API" to a production-grade,…
Open-source auth gateway connecting 1000+ SaaS providers to AI agents through SDK, CLI, MCP, HTTP, and OpenAPI. Treats external-tool onboarding as a unified, self-hosted harness layer: one credential and access model covers direct SDK calls, MCP servers, and OpenAPI discovery, so teams don't…
The transport that replaced HTTP+SSE in the 2025-11-25 spec, enabling MCP servers to run as remote services rather than local processes. Servers handle multiple client connections using HTTP POST (for client→server messages) and optional GET (for server→client SSE streams). The key harness…
The MCP team's roadmap for the next spec cycle: horizontal-scaling transport without stateful session constraints, .well-known discovery for capability advertisement without live connections, Tasks primitive with retry/expiry semantics, and enterprise extensions (audit trails, SSO, gateway…
The largest revision of MCP since launch: a stateless protocol core drops the initialize handshake and Mcp-Session-Id, the new ext-* extension framework formalizes Tasks and MCP Apps, and a twelve-month deprecation policy gives harness builders a stable target. Essential reading before designing…
Google's survey of six standardized agent interoperability protocols, each solving a distinct harness integration problem: MCP (tool/data connectivity), A2A (inter-agent routing via Agent Card discovery at well-known URLs), UCP (commerce workflows), AP2 (payment authorization with spend limits),…
Lightweight event-driven protocol standardizing how AI agents connect to frontend applications: streaming state updates, tool call rendering, and HITL interrupts over a shared event bus. Fills the layer between MCP (tool access) and A2A (agent-to-agent) — it's the missing protocol for real-time…
Anthropic's engineering account of reducing tool-call token overhead by having agents write code to interact with MCP servers rather than calling tools directly: up to 98.7% token reduction in experiments. Broadly applicable to any harness where tool schema overhead and intermediate results are…
Standardized framework for defining, versioning, and distributing agent skills. Enables skill reuse across Claude Code, Copilot, VS Code, Gemini, and other platforms — a harness-level abstraction that makes skills first-class deployment artifacts rather than ad-hoc tool definitions.
Comprehensive framework for creating, evaluating, and sharing agent skills with 86-task benchmark across 11 domains. Demonstrates the harness problem of skill fragmentation and provides infrastructure for standardized skill evaluation across frameworks.
Resumable long-running task workflow and Skill platform for coding that turns skill creation, evaluation, and release into a single lifecycle with Rubric, Pass@k, and Pass^k scoring. The Native/Classic dual-workflow model is a concrete reference for matching harness constraint strength to model…
Production-grade engineering skills for AI coding agents, packaged as 24 reusable skills covering the full development lifecycle from /spec to /ship. The slash-command interface and context-aware auto-activation make it a concrete reference for turning senior-engineering judgment into…
Adds peer-to-peer, UDP-based WebRTC bidirectional streaming to Bedrock Agents for real-time voice interactions. Complements existing WebSocket support with lower latency and better resilience for poor network conditions. Essential harness-level transport choice for agents targeting sub-800ms Total…
Token-by-token streaming delivery system enabling real-time agent responses; sub-second decision loops on streaming events vs. batch-refreshed data. Critical infrastructure for harnesses where latency (not just throughput) is the constraint — agents must react to events as they arrive, not wait…
Google ADK expansion with evaluation harness (117 prompts) for assessing skill performance across agentic coding, chatbots, document processing. Provides reference patterns and benchmark datasets for skill evaluation, complementing the Microsoft Skills Framework with Google's evaluation methodology.
GitHub's February 26, 2026 update is worth including for one specific reason: it makes .github/agents/ custom agent files, self-review, built-in security scanning, and CLI handoff concrete as harness primitives rather than abstract ideas. Useful as a current reference for how repository-scoped…
Google's 2026 rollout of managed MCP endpoints is a useful counterpoint to self-hosted MCP servers: discovery, IAM, audit logging, and Model Armor are provided as platform primitives instead of being rebuilt per server. Worth including because it shows what "enterprise MCP" looks like when the…
Microsoft's April 1, 2026 release is a strong concrete example of domain-specific skills done properly: the agent learns when to use MCP, when to drop to a Python SDK, and when to call a raw API, while the user stays in natural language. Worth adding because it shows that "skills" are not just…
Stripe's May 2026 study of how agents actually consume SDK and CLI guidance: passive documentation is ignored, while hard steering signals placed in the loaded context — skill files, error messages, CLI prompts — reliably change behavior. The "if your guidance wasn't in the loaded context, it…
Official AWS-supported MCP servers, skills, and plugins that let AI agents provision, query, and manage AWS resources through a standardized protocol interface. Worth including as the reference for how a major cloud provider productizes infrastructure access into agent-ready harness primitives…
Portable .agent/ folder that externalizes memory, skills, and protocols from any specific coding agent into a cross-tool harness layer. Adapters translate the same configuration into Claude Code's CLAUDE.md, Cursor's rules, OpenCode's AGENTS.md, and more — the first practical answer to harness…
Production-grade framework for building agents with MCP: composable workflows, built-in observability, and provider-agnostic model routing. The clearest reference for turning MCP servers from isolated utilities into a coherent agent harness.
TypeScript framework for building production MCP servers with a "Presenter" perception layer that strips undeclared fields, redacts PII, and gates tool visibility by workflow state. Fills a critical gap in the MCP ecosystem: most tooling focuses on consuming servers, while vurb.ts addresses the…
Microsoft's skill optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits and validation-gated updates, producing deployable best_skill.md artifacts. The key harness insight is that skills should be treated as optimizable parameters that improve…
Agentic skills framework and software-development methodology with automatically-triggered, mandatory skills that work across Claude Code, Cursor, Codex, Gemini CLI, and Copilot CLI. Demonstrates how to package cross-harness workflows — TDD, subagent-driven development, review gates — as reusable…
Installable library of 1,400+ agentic skills for Claude Code, Cursor, Codex CLI, Gemini CLI, and more. The largest community-driven skill catalog with an npm installer and role-based bundles — a concrete reference for treating skills as versioned harness artifacts rather than ad-hoc prompts.
Open-source agentic proxy that unifies LLM gateway, MCP gateway, and A2A gateway into a single control plane for managing multi-agent, multi-tool connectivity at scale. Provides drop-in security, observability, and governance for agent-to-LLM, agent-to-tool, and agent-to-agent communication — the…
June 2026 proposal to replace free-form skill prose with directed execution graphs: discrete steps as nodes backed by deterministic scripts or natural-language descriptions, connected by explicit typed input/output edges and governed by a schema-validated YAML spec. Compiling skills to AIP…
Skill that makes coding agents behave like a "lazy senior dev": prefer built-in solutions, avoid new dependencies, and write the minimum code that works. Benchmarked on real Claude Code sessions with ~54% fewer lines, ~20% lower cost, and preserved safety guards — a rare harness-level incentive…
Cross-harness plugin marketplace that maintains one source-of-truth plugins/ directory and generates harness-native artifacts for Claude Code, Codex CLI, Cursor, OpenCode, Gemini CLI, and GitHub Copilot. It is the clearest practical example of treating reusable agent capabilities as a portable…
CLI that turns agent skill verification into repeatable unit tests: it generates task and grader pairs from a SKILL.md, runs multi-trial evals against Claude, Codex, Gemini, or OpenCode, and reports pass rates with a CI-ready threshold. Fills the gap between shipping a skill and knowing an agent…
Qwen's official multimodal plugin suite packages vision, video, document, 3D, and CAD capabilities as portable skills and MCP servers across Claude Code, Codex, OpenCode, and other harnesses. It shows how to make a text-first coding harness multimodal-native without rebuilding the agent loop.
Anthropic on building structured permission and authorization systems into agent harnesses instead of relying on natural-language permission text.
OWASP's authoritative definition of the "excessive agency" risk: over-provisioned functions, unnecessary permissions, and missing approval mechanisms. The standard checklist for auditing harness permission scope against principle of least privilege.
April 2026 GitHub official guide for enterprise agent governance: MCP server registry curation with ruleset-protected configurations, agent environment standardization via copilot-setup-steps.yml, ephemeral runner enforcement, and cloud-agent firewall allowlisting. The most concrete published…
Anthropic's engineering post on replacing approval fatigue (users approve 93% of prompts, making approvals meaningless) with a two-stage classifier: fast single-token gate first, chain-of-thought reasoning only on flagged actions. The design decisions — stripping assistant messages to prevent the…
The most concrete reference for harness permission architecture: five-layer evaluation order (hooks → deny rules → permission mode → allow rules → canUseTool), allowedTools/disallowedTools declarative scoping, and four permission modes including dontAsk (deny-by-default for headless agents). The…
Distinguishes on-behalf-of authorization (agent uses end-user credentials, requires cross-channel identity mapping and per-user memory isolation) from fixed-credential authorization (agent owns its own account, requires human-in-the-loop guardrails on high-risk actions). The two models have…
Microsoft Security's reusable Authorization Fabric combining a Policy Enforcement Point (PEP) and Policy Decision Point (PDP) as a Microsoft Entra-protected endpoint. Every agent calls this fabric before tool execution, receiving a deterministic decision: ALLOW / DENY / REQUIRE_APPROVAL / MASK.…
The first IETF standards-track specification for AI agent authentication (March 2026, authors from AWS, OpenAI, Zscaler, Ping Identity, Defakto Security). Builds on WIMSE (Workload Identity in Multi-System Environments) and OAuth 2.0 rather than inventing new protocols — agents get SPIFFE-style…
Open-source platform providing pre-built OAuth and API key authentication for 700+ APIs across 30 categories. Automatically refreshes access tokens, provides webhooks when credentials break, and stores tokens securely so agent code never touches secrets. Solves the "agent needs to call an…
Infisical's open-source credential broker that sits between AI agents and the APIs they call, injecting real credentials onto outbound requests so agents never possess secrets directly. Eliminates a concrete prompt-injection attack surface — exfiltration of API keys and PATs — by treating…
Three-dimensional risk taxonomy (source/failure-mode/consequence) with fine-grained agentic safety benchmark (ATBench) and diagnostic guardrail models (4B–8B parameters) achieving 91.8% accuracy. Shifts safety monitoring from binary safe/unsafe checks to root-cause diagnosis: why did an action…
March 2026 open specification and reference implementation that intercepts tool calls synchronously before execution, evaluates them against a declarative policy, and produces a cryptographically signed audit record. Enforces authorization in a median of 53ms; in a live adversarial testbed ($5,000…
Deterministic permission guard that maps tool calls to an intent taxonomy (filesystem_delete, network_outbound, lang_exec, etc.) rather than relying on command-name allow/deny lists. The key insight for harness design: the same binary can be benign or destructive depending on its arguments, so…
Pre-execution firewall that intercepts, classifies, and blocks agent tool calls before they execute, with a compliance cockpit for real-time monitoring, human-in-the-loop approvals, and a tamper-evident audit trail. The zero-code-change integration makes runtime policy enforcement practical for…
Analyzes 481 public CLAUDE.md files and finds only ~4% of natural-language security rules are backed by a matching built-in control, exposing the gap between documented intent and enforced permissions. A concrete reminder that harness instructions are not guardrails unless they map to…
August 2026 user study (113 non-technical participants) comparing per-action HITL approval, automated model review, and user-authored allow/ask/never rules. The key harness finding: pre-authored policies blocked ~20 percentage points less overreach than per-action approval — users overwhelmingly…
LangChain's September 2026 implementation of on-behalf-of authorization as a harness primitive: credentials live in the workspace (never in .env or the build) and are resolved at run time by slug, with user-owned connections resolving per caller so agent actions carry the requesting human's…
Anthropic's foundational guide on agent architecture, covering when to use workflows vs. agents and how to compose primitives.
The reference architecture for stateful agents: three-tier memory (core / archival / recall) maps directly to harness state management design. Their agent loop redesign post is the most thorough public analysis of how memory structure shapes the harness.
Drop-in universal memory layer (YC-backed, AWS Agent SDK's exclusive memory provider) that handles cross-session retention without custom harness-level state management code. Lowest integration cost for production-grade persistent memory.
Self-hosted persistent memory layer with an 8-stage consolidation pipeline (episodes → facts → relationships → patterns) and built-in MCP server. The critical gap it fills: production-grade cross-session memory without cloud dependencies or complex infrastructure — a single Docker Compose gives…
Tencent's fully local agent memory system with a 4-tier progressive pipeline (Conversation → Atom → Scenario → Persona) and symbolic short-term memory via Mermaid canvases. The benchmark data is striking: 61% token reduction and 51% relative pass-rate improvement on long-horizon tasks,…
Purpose-built agent memory store with automatic conversation summarization, entity extraction, and semantic search over session history. Solves long-session context overflow at the memory layer rather than forcing the harness to manage trimming manually.
Persistent memory for AI coding agents delivered as a single Go binary with SQLite + FTS5 and 18 MCP tools for save, search, session lifecycle, and conflict detection. The agent-agnostic, zero-dependency design makes cross-session memory a local harness primitive rather than a managed cloud service.
Local-first AI memory system that stores conversation history verbatim and retrieves it with semantic search through a structured palace architecture (wings, rooms, drawers). Achieves 96.6% R@5 on LongMemEval with zero LLM calls, making it the best-benchmarked open-source memory layer for agents…
Persistent memory layer purpose-built for coding agents with 95.2% retrieval accuracy and 92% token reduction, backed by real-world benchmarks. Its cross-agent architecture — one memory server serving Claude Code, Cursor, Codex, and OpenCode through MCP and hooks — makes cross-session memory a…
Turns raw Claude Code sessions into a self-evolving knowledge base: hooks capture every interaction, the Agent SDK extracts decisions and lessons, and an LLM compiler distills them into structured, cross-referenced articles that improve retrieval quality over time. The most concrete open-source…
Turns agent-learned project knowledge into a repo-local, symbol-grounded wiki with drift detection and task-aware routing. The critical gap it fills: most agent memory stores facts, but mex keeps those facts connected to the exact code symbols they describe and flags when code changes invalidate…
LangChain's engineering account of a COALA-based three-tier memory system (procedural/semantic/episodic) backed by PostgreSQL but exposed to agents as a virtual filesystem. Key harness decisions: human-in-the-loop approval gates every memory write (blocking prompt-injection via malformed writes),…
GitHub's January 15, 2026 write-up is one of the clearest public discussions of deployed cross-agent memory: repository-scoped memories are shared across coding agent, CLI, and code review, but only after just-in-time verification against the current code state. The core harness lesson is that…
LangChain's August 2026 design account of memory that can forget: every claim the agent writes is persisted with versioned code evidence, and when that evidence changes the claim is flagged stale — remaining durably uncertain until re-verified rather than silently trusted. The most concrete…
Proposes a governance layer that decouples memory lifecycle management (decay, conflict resolution, privacy enforcement) from model weights, directly addressing the "zombie memory" problem: outdated facts sitting in the context window that only a harness-level eviction policy — not the model — can…
Production-validated architecture (283 sessions, 108k-line codebase) built on three components: a "hot-memory constitution" encoding conventions and multi-agent coordination protocols, 19 domain-specialist agents, and a "cold-memory knowledge base" of 34 on-demand specification documents. The…
Identifies three production failure modes of in-context memory at scale: capacity overflow at ~8,000 facts, 60% fact destruction during compaction, and 54% behavioral drift from constraint erosion across cascaded summarizations. Proposes Knowledge Objects (hash-addressed discrete fact tuples)…
Formal framework for measuring how well agents recover from tool failures. Defines Expected Recovery Regret (ERR) as a metric for harness design: the cost of recovering from stochastic failures in downstream tasks. Critical for assessing reliability of production harnesses where tool calls…
Represents agent memory across four orthogonal semantic, temporal, causal, and entity graphs, enabling policy-guided retrieval over relational views. Outperforms MemGPT on long-horizon reasoning benchmarks by 18.5% accuracy improvement. The multi-graph abstraction lets harness engineers compose…
Hybrid memory system blending graph traversal with semantic similarity through additive scoring; graph augmentation improves retrieval over embedding-only approaches for long-horizon reasoning. Practical alternative to full multi-graph systems when adding structure to existing vector-based memory…
Formal semantics for versioned memory graphs with belief revision operations, enabling agents to maintain coherent evolving world models through multi-turn reasoning. Addresses the hard problem of inconsistency resolution in long-lived agent memory: when new information contradicts prior beliefs,…
LangChain's April 2026 framing of agent learning as three distinct layers: model weights, harness behavior, and contextual memory. Essential for designing memory systems that don't just store facts but actually improve agent performance over time through trace-driven harness and context updates.
Open-source memory platform with a hybrid graph-vector-relational poly-store that lets agents recall facts through both semantic similarity and structured graph traversal — the practical middle ground between flat vector stores and full multi-graph research systems. The self-improving pipeline…
Agent memory system organized around three explicit operations—retain, recall, and reflect—with semantic, keyword, graph, and temporal retrieval plus an MCP server. The June 2026 release and production usage make it a concrete reference for turning cross-session persistence from passive storage…
Applies virtual-memory semantics to agent context management, treating the context window as working memory with typed pages, minimum-fidelity invariants, and validated writeback at lifecycle boundaries. A concrete reference for making residency and durability auditable harness-level concerns…
June 2026 proposal to treat long-horizon memory as execution-state management rather than semantic retrieval: a hierarchical state tree preserves trajectories, enables rollback, and constructs working state from the active root-to-leaf path. Improves task success by 7.8–20.4 percentage points…
Indexes coding-agent sessions already written to disk and serves them back over MCP with no LLM calls, embeddings, or API keys. The zero-dependency binary solves the memory cold-start problem for cross-session persistence and demonstrates that agent memory can be built on existing filesystem…
Letta's July 2026 library that normalizes native session transcripts from 15+ harnesses (Claude Code, Codex, Cursor, OpenCode, Pi, Gemini CLI, OpenHands, and more) into one validated, model-ready record format. Fills a real infrastructure gap — every runtime logs the same concepts (messages,…
OpenAI's framing of harness engineering as a discipline: how to design the scaffolding that lets Codex and similar agents operate reliably in an agent-first world.
Anthropic's account of coordinating 16 Claude instances in parallel on a shared git repo without a central orchestrator: agents claim tasks via files in current_tasks/, git forces collision resolution naturally, and a continuous restart loop spawns fresh sessions that resume where predecessors…
Unified proxy and SDK that routes to 100+ LLM providers behind a single OpenAI-compatible interface, with a Router handling retry/fallback across deployments, per-project cost and rate-limit tracking, and OTEL callback integrations. The right infrastructure layer when your harness needs provider…
Graph-based state machine framework for multi-agent harnesses: models supervisor/subagent topologies, error-recovery branches, and checkpoint persistence as first-class primitives. The most widely adopted harness orchestration layer in production.
Lightweight multi-agent framework built around handoffs and guardrails; the production successor to Swarm. Complements LangGraph for harnesses where delegation patterns are simpler than full graph orchestration.
OpenAI's official SDK for programmatically controlling local Codex agents from TypeScript or Python: start threads, run prompts, resume sessions, and choose sandbox presets (read-only, workspace-write, full-access) so Codex can be wired into CI/CD or custom agent surfaces instead of remaining…
OpenAI's September 2026 guide to treating the open-source Codex harness as reusable infrastructure rather than a product to adopt wholesale: your application owns the interface, business context, MCP tools, and approval gates, while the app-server owns the agent loop and sandboxed execution. The…
Google's code-first agent framework with built-in multi-agent orchestration, tool registration, session state, and eval pipeline. Its Runner and AgentTool patterns are the reference implementation for wrapping sub-agents as tools in a larger harness.
AWS's open-source, model-driven agent SDK that treats the loop, tool binding, and guardrails as first-class primitives. Supports Bedrock, Anthropic, OpenAI, Gemini, and Ollama with native MCP, built-in observability, and multi-agent patterns — the missing AWS open-source harness framework…
Google's May 2026 guide to production agents that survive days-long idle periods: DatabaseSessionService for persistent sessions, webhook-triggered state_delta resumption that lets containers scale to zero, and explicit state machines instead of dumping raw JSON into vector databases. Complements…
Google's June 2026 production pattern decomposing monolithic agents into cross-language specialists: a Python LLM extraction agent delegates to a Go deterministic validator via A2A, with shared session-state checkpoints and fail-safe routing to manual review. A concrete example of using…
Microsoft's multi-agent conversation framework with a complete AgentChat layer covering agent loop, tool integration, termination conditions, and human-in-the-loop. The most comprehensive open-source reference for large-scale multi-agent harness design.
Dual-layer harness orchestration: Crew handles autonomous agent delegation, Flow provides event-driven deterministic control (branching + shared Pydantic state). The clearest open-source example of mixing autonomous and scripted execution in the same harness.
June 2026 harness-first redesign built around the Capability primitive: a single composable unit bundling instructions, tools, lifecycle hooks, and model settings. The split between a small stable core and a fast-moving pydantic-ai-harness lets capabilities graduate as they prove essential, while…
Intelligent routing across multiple LLM providers with load balancing, intelligent fallbacks, rate limiting, and response caching. Achieves 40–60% token cost reduction through smart model routing (cheap models for simple tasks, capable models for complex reasoning). Essential infrastructure for…
Token-efficient microkernel agent with an on-device SquillaRouter that dispatches each turn to the cheapest capable model, a unified turn loop across CLI/Web/chat, and persistent memory with on-device embeddings. Demonstrates that harness-level routing and loop optimization — not just model scale…
Anthropic's production architecture for separating three stateless components — the "brain" (Claude + harness), "hands" (sandboxes/tools), and "session" (append-only event log) — enabling independent failure and replacement of each. Crash recovery via session replay (wake(sessionId) + getEvents())…
Microsoft's June 2026 expansion folds previously scattered harness primitives into a single runtime layer: automatic context compaction, instruction merging, todo tracking, background agents, and a ToolApprovalAgent. The CodeAct execution mode is particularly notable—agents emit a short Python…
Production-ready 1.0 release (April 2026) unifying Semantic Kernel and AutoGen into a single framework with graph-based orchestration, middleware pipeline for intercepting every execution stage, and declarative YAML agent definitions. DevUI provides a browser-based debugger for visualizing agent…
Microsoft’s open-source YAML-first CLI for deterministic multi-agent orchestration: Jinja2 routing, static parallel groups, and per-agent model overrides across Claude and Copilot with zero token overhead on the orchestration layer itself. Treats workflows as version-controlled, diffable…
Shopify's open-source Ruby DSL for structured AI workflows that interleaves deterministic steps (shell commands, Ruby code) with agentic steps via Claude Code or Pi. Its "non-determinism is the enemy of reliability" philosophy makes it a concrete reference for building reproducible,…
Production-ready open-source runtime focused on two pieces many agent frameworks leave underspecified: secure sandbox execution and durable agent serving. The "Agent as API" model, async sandbox types, and built-in state/sandbox lifecycle management make it one of the few 2026 projects tackling…
Production-ready Java framework for distributed, enterprise-grade agents with workspace sandboxing, permission-gated tool calls, AOP-style middleware, and cross-replica session recovery. The Java complement to AgentScope Runtime for teams that need a typed, JVM-native harness stack.
Temporal.io's harness infrastructure for persistent agent workflows with native agentic handshake protocol for secure deadline-aware calendar negotiation between autonomous agents. Brings distributed systems best practices (durability, retry semantics, activity monitoring) to agent orchestration,…
The leading TypeScript toolkit for building AI agents (20M+ monthly downloads, 25+ provider integrations). AI SDK 6 introduced a first-class Agent abstraction with ToolLoopAgent for production-ready tool execution loops, DevTools for local debugging, full MCP support, and type-safe UI streaming.…
Vercel's filesystem-first framework for durable AI agents: instructions, typed tools, on-demand skills, message channels, and cron schedules live in conventional directories, making the harness inspectable and version-controlled by default. A concrete reference for treating the project filesystem…
TypeScript-native agent framework (from the Gatsby team) with 22K+ stars and 300K+ weekly npm downloads. Connects to 40+ providers through one standard interface, with built-in workflows, RAG pipelines, and agent orchestration. The @mastra/deployer handles serverless deployment, and the eval…
Astro's TypeScript-native agent harness that treats an agent as a function composed of model, sandbox, skills, tools, and MCP servers. The deploy-anywhere runtime (Node, Cloudflare Workers, GitHub Actions, Daytona) and first-class durability make it a practical reference for building autonomous…
TypeScript-native multi-agent orchestration that automatically decomposes a natural-language goal into a task DAG, parallelizes independent nodes, and synthesizes results. With only three runtime dependencies, built-in MCP support, token budgets, retries, and context compaction, it is the lightest…
OpenAI's April 2026 update adding native sandbox execution, configurable memory, and sandbox-aware orchestration to the Agents SDK. The shift toward a "model-native harness" that aligns execution patterns with how frontier models actually perform best is a reference design for SDK-level harness…
OpenAI's orchestration layer for moving from supervising coding agents to managing work: it monitors an issue tracker, creates isolated per-issue workspaces, and surfaces proof-of-work artifacts (CI status, review feedback, walkthrough videos) as the handoff signal. The spec deliberately leaves…
Drop-in multi-agent framework where protocol enforcement is a mechanical gate, not a prompt request: IDE-level hooks check every code-changing turn for required reviewers, memory updates, and supply-chain integrity before the turn can complete. Built on stdlib Python with zero dependencies, it…
YC-backed production harness that compiles multi-agent objectives into deterministic execution DAGs with state persistence, crash recovery, cost enforcement, and human-in-the-loop oversight. The most complete open-source reference for running agent workloads in production rather than demos.
Native-Rust agent harness platform with four surfaces (GUI, CLI, web, one-shot), multi-provider routing, and sovereign-by-design architecture. The most complete open-source reference for a locally-run, offline-capable harness that unifies MCP, skills, AGENTS.md, hooks, and session resumption in a…
TypeScript-native orchestration for sandboxed coding agents that treats provider-agnostic isolation (Docker, Podman, or Vercel Firecracker microVMs) as a primitive, not a framework. The built-in review-pipeline and parallel-AFK-agent patterns demonstrate how lightweight harness layers can enforce…
Deterministic scheduler for 40+ CLI coding agents running in parallel git worktrees with an HMAC-signed audit chain, signed agent cards, and per-artefact lineage. The zero-LLM coordination loop and tamper-evident audit trail make it the only open-source orchestrator designed for…
Attributed, resolvable artifact references for agent handoffs: a ~30-byte token replaces pasted context, letting each consumer resolve only the projection or slice it needs under byte budgets while propagating corrections to every holder. The first concrete primitive that turns multi-agent…
The AI orchestrator for running Codex, Claude Code, OpenCode, and Pi side-by-side in isolated git worktrees. It turns fleet-of-agents execution into a polished desktop IDE with a mobile companion, making parallel agent harness design accessible beyond shell-scripting teams.
Provider-neutral control plane that sits above existing harnesses (Codex, Claude Code, Cursor) and keeps long-horizon objectives, gates, todos, evidence, and quota stable across bounded agent turns. The local-first state kernel makes multi-day work reviewable, restartable, and handoff-safe without…
Self-hosted unified API for running Codex, Claude Code, DeepSeek Harness, Pi, and other terminal agents through the open Unified Harness Protocol (UHP), with sessions, streaming, file access, cancellation, and failure handling. Worth including because harness fragmentation is becoming a real…
Anthropic's framework for evaluating agent behavior: what to measure, how to build eval harnesses, and why unit-test-style evals fail for agents.
YAML-driven LLM testing framework with LLM-as-judge, assertion DSL, and native CI integration. The most practical tool for adding agent output regression tests to a PR pipeline without writing a test harness from scratch.
Multi-environment agent benchmark (OS, DB, web, code) with a structured eval pipeline. Worth studying for its environment isolation design and task definition format when building custom eval environments for your harness.
OpenAI's framework for skill regression testing: four eval dimensions (outcome, process, style, efficiency goals), JSONL trace capture for deterministic checks (command sequences, token budgets, repo cleanliness), then rubric-based grading only where deterministic checks don't suffice. The…
A 33-item checklist covering the full evaluation lifecycle: error taxonomy, three-level granularity (single-step → trace → multi-turn thread), grader specialization, and CI integration. Key insight: capability evals (low pass rate, improvement target) and regression evals (near-100%, protection…
LangChain's methodology for benchmarking agent skills in Docker-sandboxed environments. Key empirical findings: Claude Code achieved 82% task completion with curated skills vs. 9% without, and consolidating to ≤12 skills improved accuracy over sprawling skill sets. The baseline-vs-skills…
Addresses agent CI's core problem: binary pass/fail is useless for non-deterministic workflows. Behavioral fingerprinting detects 86% of regressions vs. 0% with binary testing; stochastic PASS/FAIL/INCONCLUSIVE verdicts grounded in hypothesis testing cut token costs 78%. Trace-first offline mode…
Demonstrates how harness specialization (llvm-autofix) for a narrow domain (compiler bug fixes) achieves better results than general-purpose coding agents. The tool design patterns — exposing compiler error messages directly, bounding search depth by compilation cost — are transferable to any…
Red Hat's eight-stage evaluation maturity progression from manual CLI testing to cost-aware continuous monitoring (March 2026). Uses DeepEval with 15 custom ConversationalGEval metrics and LLM-as-judge; key finding: evaluator model capability matters significantly — llama-3-3-70b caught all known…
Comprehensive framework combining multi-environment baselines (AgentBench), domain-specific benchmarks (Terminal Bench 2.0, WebArena, SWE-bench Verified), and industry standards (NIST AI Agent Standards Initiative, February 2026). Provides reference metrics and rubrics for evaluating coding…
Systematic index analyzing deployed agent safety documentation, guardrails, and third-party testing across 30 production systems; identifies critical gaps in agentic safety disclosure and documentation. Useful as a checklist for production harness safeguards before deployment.
Real-time architectural sensor that closes the feedback loop for coding agents: scans codebases, scores structural health, and surfaces degradation via MCP so agents can self-correct before entropy compounds. The gate --save / gate pair makes architectural regression detection CI-friendly —…
Google's June 2026 account of automating the eval-optimize loop for coding agents: independent AutoRaters grade agent outputs so the optimizer cannot game its own metrics, and custom rubrics isolate specific behaviors like stale-message echoing that blended scores miss. Shows how to close the…
Pytest-style testing framework for MCP servers with protocol-aware assertions, snapshot tests, and CI-ready reports across stdio/SSE/HTTP transports. Fills the gap between shipping an MCP server and trusting it in production by turning tool/schema regressions into ordinary test failures.
September 2026 agent-first linter for Tailwind design systems: you encode design-system contracts as programmable lint rules, and the errors carry your components' sizes, variants, and theme file paths so agents self-correct in one round instead of guessing. Validated across 150+ task runs…
OpenTelemetry-based instrumentation for LLM calls and agent steps: adds trace spans to every inference and tool call without modifying business logic. The cleanest way to bring the existing OTEL ecosystem (Grafana, Datadog, Jaeger) to a harness.
Self-hostable trace UI and eval runtime for agent workflows. Lets harness engineers audit and replay every reasoning step and tool call offline, without sending data to a third-party cloud.
Comet's open-source AI observability and evaluation platform: deep tracing of LLM calls, conversation logging, and agent activity, plus built-in eval metrics, prompt versioning, guardrails, and the Opik Agent Optimizer. Worth including because it unifies observability, verification, and…
The most widely adopted self-hostable LLM observability platform: traces every agent step, manages prompt versions, and runs evals in one tool. Preferred over cloud-only alternatives when data residency or cost control is a constraint.
W&B's tracing and eval layer purpose-built for agent workflows: automatic call graph capture, dataset versioning, and LLM-as-judge evals that integrate directly with the wandb experiment tracking ecosystem.
OpenTelemetry's standard attribute names for GenAI spans (gen_ai.system, gen_ai.request.model, etc.). The naming baseline that makes harness traces portable across any OTEL-compatible backend.
AI observability platform from the Pydantic team with a unique angle: all trace data is SQL-queryable (PostgreSQL-compatible), so coding agents can query production observability data directly via the Logfire MCP server. Full-stack OTEL tracing covers both the AI layer and backend — letting you…
Open-source LLM observability proxy (YC W23) with the largest open-source pricing database (300+ models). One-line proxy integration provides cost tracking, token monitoring, session tracing, and prompt versioning across providers. The AI Gateway component handles request routing and caching with…
2026-standard platform for LLM tracing with infrastructure log/metric unification. Enables harness engineers to correlate agent decisions with system-level events (network delays, GPU memory pressure) that explain agent failures, going beyond isolated LLM call traces.
Evaluation-first agent observability platform ($80M Series B, Feb 2026) with exhaustive auto-tracing that captures every LLM call, tool invocation, and retrieval step as nested span hierarchies. Brainstore, its purpose-built data store, enables full-trace search without sampling — critical for…
Combines Temporal's durable execution (automatic retries, state persistence, event history replay) with Braintrust's LLM tracing so every Workflow and Activity becomes a Braintrust span and every LLM call is traced with full context. Demonstrates the pattern with a deep research agent where failed…
Google Cloud's 2026 launch treats agent traces, tool calls, sessions, and outcomes as analytical data rather than dashboard exhaust. The important harness idea is that observability becomes queryable infrastructure: once telemetry lands in BigQuery, teams can build evaluators, regressions, and…
Red Hat's April 6, 2026 guide is one of the few concrete references that walks through context propagation across routing agents, specialist agents, MCP servers, and external systems using standard tracing infrastructure. It belongs here because it treats agent observability as a…
METR's three-week adversarial audit of Anthropic's internal agent monitoring and security systems (described in the Opus 4.6 Sabotage Risk Report). Discovered several novel vulnerabilities, some since patched. The most concrete published account of what it takes to stress-test agent monitoring…
Open-source, self-hostable platform unifying tracing, evals, simulations, guardrails, and gateway into a single feedback loop. Worth including because it demonstrates what a unified observability-and-improvement plane looks like rather than stitching together five separate vendor tools.
Local-first Agent Work Intelligence for coding agents: ingests existing Claude Code, Codex, and OpenCode session logs, attributes tokens and estimated cost to recorded work steps with confidence labels, and surfaces the evidence on a private dashboard with no cloud sync or API keys.
Open-source agent engineering platform (YC W24) with session replay, cost tracking, and failure detection across 10+ frameworks including CrewAI, LangGraph, and OpenAI Agents SDK. The step-by-step execution graph and cross-session metrics make it the most practical debugging layer for multi-agent…
February 2026 open-source DevTools for Claude Code that reconstructs hidden session internals from local logs: per-turn token attribution across 7 context categories, full subagent execution trees with cost breakdowns, and syntax-highlighted diffs for every tool call. Essential because Claude…
Wireshark for MCP: a transparent proxy that shows every JSON-RPC frame between your real client and MCP servers, live in your terminal. Worth including because MCP Inspector tests servers from its own client, so it can't see the calls your actual agent makes, misses, or hangs on — mcpsnoop sits in…
Anthropic's July 2026 bundled setup-checkup skill (alias /checkup) that audits harness hygiene: it deduplicates local and checked-in CLAUDE.md files, flags unused skills/MCP servers/plugins, identifies slow hooks, and proposes fixes only after confirmation. The clearest first-party example of…
April 2026 agent debugging skill that stops guesswork with runtime evidence. Uses background tracing (Runtime Facts) to capture the exact execution path leading to failures, then constrains the agent to cite specific data points (stack traces, variable snapshots) before proposing fixes. Moves…
March 2026 framework that localizes root causes in multi-agent execution traces using causal graph analysis rather than LLM inference. Processes traces in 0.12 seconds (69× faster than LLM-based analysis) with 93.6–95.8% accuracy across 550 synthetic failure scenarios. Distinguishes root causes…
February 2026 ICSE paper introducing a collaborative multi-agent debugging loop: Instrumentation Agent injects diagnostic probes, Analysis Agent performs causal trace diagnosis with a Historical Lesson Learning Mechanism (HLLM) that distills insights from prior failures, and Repair Agent executes…
Framework for automated root-cause analysis of agent failures: trajectory normalization, constraint synthesis from tool schemas, and constraint-guided evaluation. Achieves 23.6% better failure localization than existing approaches with a 115-trajectory annotated benchmark. Shifts agent debugging…
Addresses the core problem of debugging agents that run for minutes, span hundreds of steps, and produce massive traces no human can manually scan. Introduces Polly (an AI assistant that analyzes traces to surface root causes) and langsmith-fetch (CLI for piping trace data to coding agents). Key…
ICLR 2026 paper introducing the Agent Error Taxonomy — a modular classification covering memory, reflection, planning, action, and system-level failures. The AgentDebug framework isolates root-cause failures and provides corrective feedback, achieving +24% higher all-correct accuracy. The Agent…
Open-source React component library (Evil Martians) that transforms OpenTelemetry trace data into interactive visualizations: tree view, timeline/Gantt view, sequence diagrams, and detail panels. Framework-agnostic — works with any OTEL-compatible agent. Fills the gap between raw OTEL spans and…
March 2026 empirical study mining 375 GitHub issues across real-world agent systems (AutoGen, CrewAI, OpenAI Agents SDK, LangChain, CAMEL, DB-GPT) to build the first grounded taxonomy of agent-specific faults: initialization failures, role deviation, memory/state deficiencies, orchestration…
GitHub's March 19, 2026 changelog is short but materially useful: setup-step logs, collapsed subagent traces, and clearer session-stage visibility are exactly the kind of DX improvements that make long-running agent failures debuggable in practice. It is a concrete reminder that trace readability…
February 2026 interactive debugger for agent execution trajectories that organizes raw logs into structured, side-by-side conversations (agent↔LLM and agent↔tools). Enables step-through execution, breakpoint manipulation, and mid-trajectory inspection. Developer study shows frustration scores drop…
Replays Claude Code and Codex sessions on a 3D map of your codebase, turning raw JSONL logs into a spatial view of where the agent searched, read, and edited. The glow-based visualization makes exploration drift and context pressure immediately visible — a concrete debugging primitive for…
Evaluates coding-agent workflows across five dimensions and turns project and session evidence into prioritized, evidence-bounded findings with scoped repair actions. Runs inside Claude Code, Codex, Cursor, GitHub Copilot, and other hosts, making it a practical cross-tool diagnostic layer for…
March 2026 AWS reference implementation demonstrating four distinct HITL patterns for sensitive agent tool calls: Hook System (centralized blanket policy), Tool Context (per-tool fine-grained), Step Functions (async third-party approval via SNS), and MCP Elicitation (protocol-native real-time…
February 2026 release making human oversight a native workflow primitive: suspend execution at critical decision points, expose review-and-edit UI mid-flow, and route subsequent execution based on human action (approve/reject/escalate). Demonstrates how HITL transitions from bolt-on approval gates…
Open standard (v0.8, February 2026) for human decisions in agent workflows: HTTP 202 + review URL pattern connecting services, agents, and humans across any messaging channel. No SDK required — ~15 lines of code for agents, reference implementations in Express/Hono/Next.js/FastAPI for services,…
Systematic treatment of interrupt, breakpoint, and approve patterns: how to pause an agent mid-loop, persist state, and resume after human review. Directly addresses the harness engineering challenge of inserting human gates into long-running workflows.
Explains human_input_mode (NEVER / TERMINATE / ALWAYS) and the UserProxyAgent as an approval gate. The most concrete implementation reference for adding human review nodes to a multi-agent conversation harness.
The most complete implementation reference for HITL mechanics: canUseTool callback pauses execution at every tool request with allow/deny/approve-with-changes/suggest-alternative response shapes; AskUserQuestion surfaces structured clarifications mid-task; streaming input enables mid-execution…
April 2026 benchmark that transforms well-specified tasks into judgment challenges by injecting 3–5 realistic blockers (missing critical information) and giving agents an ask_human() tool. Agents from top models achieve ~90% pass@3 with full information but performance drops significantly when…
LangChain's April 9, 2026 guide closes an important gap that most HITL write-ups skip: human input is not just an approval gate at execution time, it's also supervision for improving prompts, tools, memory, and evaluators over time. Useful because it treats expert review as a structured data…
Martin Fowler defines three human-involvement postures — humans outside, in, or on the agent loop — and argues that "humans on the loop" (maintaining the harness rather than reviewing individual outputs) is the only approach that scales with agent throughput. The "agentic flywheel" section — where…
Anthropic's February 2026 empirical study of millions of real-world Claude Code interactions. Key finding: experienced users shift from per-action approval (20% auto-approve when new) to intervention-only oversight (40% auto-approve at 750+ sessions), and agent-initiated clarification stops grow…
April 2026 open-source human-in-the-loop system with six intervention modes (full-auto, gate-only, checkpoint, step-by-step, co-pilot, custom), SmartPause confidence-driven dynamic suspension, and Intervention Learning from human corrections. The cost-guardrail system — aborting runs that exceed…
Local-first attention orchestration layer that sits between you and every agent, alert, and feed: a three-stage worthiness engine decides whether to interrupt the human, dispatch work to an agent, or curate to memory. Worth including because it treats human attention as a scarce harness resource…
A project-based course on designing the environments, state, verification, and control systems that make Codex and Claude Code reliable. The most approachable public curriculum for learning harness engineering from first principles — each module builds a working artifact rather than summarizing…
A book-length architectural teardown of Claude Code's harness: 15 chapters and 139 diagrams covering the tool system, four-stage permission pipeline, context compaction, memory, hooks, subagent scheduling, MCP integration, skills, streaming architecture, and plan mode. Every chapter explains why…
Two open-source books that use Claude Code and Codex as observation targets to explain how constraint structures organize execution in real engineering environments. The clearest book-length treatment of why prompts, tools, permissions, recovery paths, and team rules form a single control plane…
February 2026 hands-on workshop building an AI agent from scratch with Google's Agent Development Kit (ADK) in 3 hours. Five milestones with a built-in benchmark that tracks progress from ~19% (base agent) to ~81% (with web search, PDF reader, and calculator tools). The clearest public tutorial…
February 2026 workshop by Mastra's founders dissecting every layer of an open-source AI coding agent: stateful/resumable harness, dynamic prompt composition, workspace sandboxing, memory compaction, HITL steering, event protocols, and cost tracking. The 11-topic curriculum is the most complete…
OpenAI's February 2026 cookbook building a complete multi-agent governance system from scratch: policy-as-code guardrails, OpenAI Traces for full observability, eval-driven design, and a distributable governance package. The most concrete first-party tutorial for making governance part of core…
Anthropic's official notebook collection covering orchestrator-worker patterns, parallel tool calling, programmatic tool calling (PTC), context compaction, and Agent SDK examples. The patterns/agents/ directory is the reference implementation of every orchestration pattern described in Building…
HuggingFace's deliberately minimal agent library (~1,000 lines of core code): the entire harness — tool validation, memory, monitoring, sandbox isolation (E2B, Docker, Pyodide) — is readable in an afternoon. The code-agent pattern (model writes Python that calls tools, eliminating JSON…
Pure-Python coding agent harness (standard library only) that implements the six core harness components—live repo context, structured tools with permissions, context reduction, transcript resumption, and bounded subagents—in a single readable file. The clearest starting point for understanding…
Step-by-step deconstruction of Claude Code as an agent harness (s01–s12). Best resource for understanding how agent loop, tool use, skills, context compaction, and task management compose in practice.
Curated list organized into Full Lifecycle Platforms, Task Runners, Agent Runtimes, Coding Agents. Close to this list's scope; good complementary reference.
Practitioners' guide covering all harness configuration points for coding agents: system prompts, MCP tool selection, skills for progressive disclosure, sub-agents as context firewalls, hooks for deterministic control, and back-pressure verification. The central argument — that most agent failures…
GitHub's December 2025 practical guide on coordinating multiple coding agents with mission control: parallel vs. sequential execution, when to intervene, and how to review agent work productively. Shows the shift from single-agent prompts to multi-agent choreography and the harness decisions…
AWS's official sample repo is one of the most complete public walkthroughs of what "productionizing" an agent platform actually means: runtime, gateway, memory, identity, observability, IaC, and blueprint apps all live in one place. Worth including because it covers the harness infrastructure…
IEEE CAI 2026 tutorial (December 2025) providing a research-based practical guide for designing enterprise-ready multi-agent systems. Covers agentic patterns (ReACT, Reflection, CoT), emerging protocols (MCP, A2A), multi-layer memory structures, observability and online/offline evaluation…
Provider-neutral Agent Skill for designing, auditing, and refactoring agentic harnesses across domains. Distills the full runtime discipline — loop budgets, typed tools, permission gates, compaction-aware memory, prompt-caching layout, and launch checklists — into an interactive reference that…
Anthropic Hackathon Winner (140K+ stars). The agent harness performance optimization system: skills, instincts, memory optimization, continuous learning, security scanning, and research-first development. Production-ready agents, skills, hooks, rules, and MCP configurations evolved over 10+ months…
Affaan Momin's agent-harness operating system: 68 specialized agents, 286 skills, hooks, memory, continuous learning, and AgentShield security scanning across Claude Code, Codex, Cursor, OpenCode, and other harnesses. The clearest open-source example of packaging an end-to-end engineering workflow…
Anthropic's official SDK that exposes Claude Code's entire harness as a programmable API: built-in tool execution loop, PreToolUse/PostToolUse hooks for interception, subagent definitions, allowedTools permission control, and session resumption. The highest-leverage starting point for building a…
Scaffold factory that turns any repo into a branded agent harness with its own npx CLI, MCP server, scoped memory, governance policy, and Darwin Mode self-evolution. The clearest open-source embodiment of "the model is replaceable, the harness is the product" — it generates the owned scaffolding…
A meta-skill that generates domain-specific agent teams and the skills they use. Good example of harness-as-code, where the harness itself is produced by an agent.
March 2026 Claude Code plugin that autonomously evolves LLM agent harnesses using multi-agent proposers in isolated git worktrees, LangSmith-backed evaluation, and regression guards. Iterates on prompts, routing, retrieval, and orchestration code based on full-trace counterfactual diagnosis. The…
April 2026 open-source self-improving agentic system: bring your own coding agent, automatically mine failures from benchmark runs, optimize the harness through iterative edits, and gate changes against regressions. Supports Terminal-Bench 2.0 and tau-bench with Harbor and Docker evaluation…
Observability-driven automatic evolution of coding-agent harnesses that decomposes the scaffold into seven orthogonal, git-tracked components and iteratively improves them through trace distillation and evidence-backed edits. Ranked #3 on Terminal-Bench 2.0 (84.7%) with demonstrated cross-model…
Treats the entire harness (system prompt, tool definitions, context management, completion logic) as a joint optimization target rather than hand-tuning each piece. The key insight: give the proposer agent filesystem access to all prior harness candidates, scores, and execution traces — 10M-token…
June 2026 proposal for a self-improving harness loop: the agent mines its own failure traces, proposes minimal harness edits, and validates them through regression testing. Improves Terminal-Bench-2.0 pass rates by 20+ percentage points across MiniMax, Qwen, and GLM without requiring stronger…
Meta's framework integrating task-solving and meta-level improvement into a unified, editable program with metacognitive self-modification. Improved paper-review tasks from 0.0 to 0.710, transferred to Olympiad math grading at 0.630 improvement@50 score. Shows how agents can be designed to modify…
Open-source library (April 2026) that automates the harness engineering loop itself: give it a task and a benchmark, and it iterates overnight on system prompts, tool configurations, agent orchestration, and routing — keeping or discarding each change based on score. In a 24-hour run, hit #1 on…
Open-source Python library (April 2026) that implements an outer optimization loop around executable harnesses for coding agents. Inspired by the Meta-Harness paper, it treats AGENTS.md, setup scripts, validation logic, and test flows as optimizable artifacts rather than static configs — with…
Lightweight continual harness optimizer (April 2026) built on the Claude Agent SDK. Runs an outer loop that reads task traces, rewrites harness configs, and re-evaluates — achieving 67% → 87% on tau-bench with no labeled training data. Demonstrates that even small, focused meta-harness loops can…
April 2026 official implementation of the Meta-Harness paper from Stanford's IRIS Lab. Provides the framework and two reference experiments for end-to-end harness optimization via filesystem-backed search loops where a coding agent proposes, evaluates, and refines harness artifacts. The cleaned-up…
Recursive self-improving harness that runs multi-generation evaluation loops, distilling successful strategies into persistent playbooks and trace datasets that future agents inherit. The five-role architecture (competitor, analyst, coach, architect, curator) and built-in production trace capture…
Official implementation of RHO (Retrospective Harness Optimization): improves an agent's harness using only its own past trajectories, via self-validation, self-consistency, and pairwise self-preference — no ground-truth labels or external evaluators required. A single round moves SWE-Bench Pro…
Reset-free framework that lets an LLM Refiner rewrite its own system prompt, sub-agents, skills, and memory mid-episode via an evolve_harness tool. The first open-source implementation of online harness self-improvement with reproducible long-horizon benchmarks (Gemini Plays Pokémon).
Databricks' open-source meta-harness (June 2026) that sits above Claude Code, Codex, Pi, and custom agents to compose, govern, and share live agent sessions from one control plane. The key harness insight is that enterprises don't need another agent framework — they need a portability and…
Claude Code plugin that externalizes multi-agent orchestration as installable skills and staged team pipelines, with cross-provider advisor routing and built-in requirement-clarification interviews. The most widely adopted example of turning a single-agent CLI into a team-ready meta-harness…
A systems approach to recursive self-improvement: a full agent harness that can safely edit its own prompts, memory, tools, and policy because an immutable event log is the one thing it cannot rewrite. The clearest open-source architecture for long-lived agents that evolve their own scaffolding…
Zero-code platform for auto-generating production-grade AI agents using Harness Engineering principles: unified tools, skills, memory, and orchestration with built-in constraints, feedback loops, and control planes. The clearest open-source example of turning harness scaffolding from a hand-rolled…
Salesforce AI Research's platform for evaluating and evolving agent harnesses at scale (September 2026), shipping the official implementation of DarwinX: natural-selection search over harness changes where each trial is driven faithfully by the benchmark's own framework (Harbor, PIER) and…
August 2026 method that makes harness editing a learned capability rather than a fixed editor agent: a dedicated 9B harness engineer is post-trained with online RL — batches of target-agent failures become validated executable patches, fresh reruns of the frozen target supply outcome rewards, and…
Anthropic's reference harness for the screenshot-action loop: defines the screenshot, bash, and text_editor tool interface that makes desktop/browser control work. Essential reading before building any harness where the agent's primary sensory input is a rendered screen rather than structured API…
OpenAI's official autonomous coding agent CLI — the open-source reference implementation of the Codex harness with sandboxed tool execution, multi-file editing, and a streaming agent loop. Worth studying because it is the most widely adopted terminal-native coding agent harness and exposes the…
Agent harness with Slack, GitHub, and Linear integrations. Useful reference for how real-world tool wiring works inside a harness.
The most architecturally complete open-source coding agent: Runtime/Sandbox isolation, EventStream message bus, and Agent Controller are a three-layer harness design worth studying for production deployments.
Open-source framework for building AI SRE agents with 60+ observability and remediation tool integrations plus a synthetic incident evaluation environment. The clearest open-source reference for turning infrastructure incident response into a trainable, evaluable agent harness rather than a…
Block's open-source, extensible AI agent donated to the Linux Foundation's Agentic AI Foundation in April 2026. Its MCP-native architecture treats every capability as an MCP server, making it a practical reference for building vendor-neutral, extensible harnesses where tool integration is the…
Minimal browser-automation agent harness with clean separation of tool registration, DOM state injection, action loop, and error recovery. Small codebase, clear structure — the best "minimal viable harness" reference for understanding core loop mechanics.
Self-healing browser harness that connects an LLM directly to your real Chrome via CDP. The critical design decision: the agent itself writes missing helpers and domain skills into the harness during execution, so the scaffold improves every run rather than requiring manual updates. At ~1k lines…
A 24/7 Claude Code agent with Browser Harness, running autonomously on any machine you own. Demonstrates how to combine a terminal-native coding agent with live browser automation for workflows that span API documentation, web-based configuration, and headless verification — the reference for…
Coding agent whose Agent-Computer Interface (ACI) — purpose-built file viewer, search, and editor tools with explicit state constraints and error feedback — is the reference design for adapting a tool interface to a specific task domain rather than using generic bash.
AI pair-programmer harness with an Architect mode that splits planning (one LLM) from coding (another), and git-aware tooling that uses version control as the undo mechanism instead of custom state rollback. The best reference for multi-file editing tool design and planner/coder layer separation.
A composable coding-agent harness built on Deep Agents, synthesizing design patterns from Stripe, Ramp, and Coinbase production deployments. Key decisions: curated ~15-tool limit enforced at harness design time, one isolated sandbox (Modal/Daytona/Runloop/LangSmith) per task, AGENTS.md for…
Production harness achieving 77.4% solve rate on SWE-bench Verified through continuous harness evolution — the scaffold adapts from failure signals rather than requiring manual retuning per task class. Demonstrates the architectural pattern where the harness itself is a learnable component, not…
Handles frame management, streaming media coordination, and pipeline orchestration between ASR/LLM/TTS services for sub-800ms Total Turn-Around Time voice interactions. The missing harness primitive for voice agents: manages backpressure, handles frame queueing, and exposes a simple async…
Qwen's realtime voice runtime that keeps agents talking, working, and present across multiple backend agents (Claude Code, Codex, Qwen Code, Kimi Code). It demonstrates how to wrap terminal-native coding agents in a persistent, interruptible voice shell without losing task continuity — a concrete…
Orchestrated team of domain-specialized scientist agents that autonomously analyzed 55,984 clinical trials and discovered cell-type-specific drug targets 40% more likely to succeed Phase I→II transitions. Demonstrates specialized harness design for scientific workflows where formal reasoning,…
NVIDIA's Nemotron 3 family (Super for long-context reasoning, Content Safety for multimodal moderation, VoiceChat for real-time speech) designed for scalable agentic AI with enterprise-grade multimodal understanding. The integration of safety models, vision models, and voice models into a single…
All-in-one agent sandbox combining browser, shell, filesystem, MCP servers, and VSCode Server in a single Docker container. Native MCP support exposes sandbox capabilities to LLMs via the standard protocol, and files downloaded in the browser are instantly accessible in terminal and VSCode.…
GitHub's February 13, 2026 technical preview is unusually valuable because the implementation is fully open source (gh-aw) and shows how natural-language workflow generation, approval handling, and GitHub-native execution fit together in one harness. It belongs here as a reference implementation…
LangChain's batteries-included agent harness (released April 2026) with built-in planning, filesystem tools, shell access, sub-agents, and auto-summarization. The clearest open-source demonstration of how a general-purpose coding agent harness can be made ready-to-run out of the box while…
A compact, inspectable open-source agent harness from HKUDS (April 2026) featuring a built-in personal agent (ohmo), auto-compaction with session preservation, MCP HTTP transport, and multimodal gateway support. Excellent reference for understanding how a small, modular harness can support…
HKUDS's ultra-lightweight, self-hosted personal AI agent framework: a single readable Python core that combines WebUI/terminal/chat-app surfaces, long-term memory, MCP tools, model routing, multi-agent delegation, and scheduled automation. A practical reference for how a complete personal agent…
Y Combinator's open-source multiplayer agent harness for startups: isolated per-person workspaces plus shared Slack channels and projects, with pluggable harness backends (Claude Code, Codex, OpenCode, Pi) and scope-owned skills. The clearest reference for building team-wide agent deployments…
Open-source terminal-native AI coding agent with 131K+ stars and 2.5M+ monthly active developers. Provider-agnostic architecture supports 75+ LLM providers plus native LSP auto-configuration, multi-session parallel agents, and MCP extensibility. The build/plan agent split and client/server…
Repository-native multi-agent orchestration framework built on GitHub Copilot. Initializes a persistent AI team (lead, frontend, backend, tester) as files inside your repo — knowledge compounds across sessions through committed history.md and decisions.md. The most accessible reference for teams…
Open-source infrastructure for Computer-Use Agents: sandboxed full-desktop control across macOS, Linux, and Windows, a background-native macOS driver that operates without stealing cursor focus, plus SDKs and benchmarks. The most complete reference for building harnesses around screen-based agent…
Long-horizon computer-use harness that runs on top of Claude Code and Codex, splitting work into Manager, Executor, and Auditor roles so only independently verified results enter persistent task state. The clearest August 2026 open-source reference for carrying desktop-and-CLI tasks through dozens…
End-to-end harness for GUI agents that unifies online RL training, standardized benchmarking, and real-device deployment in one framework. The most complete open-source reference for building visual perception-action loops across desktop and mobile surfaces.
Android Open Harness Project: an open-source OS-level agent harness built on AOSP that treats AI agents as first-class OS actors, redesigning service composition, agent interfaces, and information flow for personalized, efficient, and secure mobile interaction. It fills a rare gap by moving the…
Agent harness that turns codebase quality improvement into a structured, score-driven workflow: mechanical detectors find dead code and complexity, LLM review assesses naming and abstractions, and a persistent next → fix → resolve loop keeps the agent on track across sessions. The anti-gaming…
DeepSeek-native coding agent harness engineered around prefix-cache stability as a loop invariant: immutable-prefix / append-only-log / volatile-scratch partitioning achieves 99.82% cache-hit rates and ~5× cost reduction on long sessions. The most detailed public case study of designing an entire…
DeepSeek-V4 terminal coding agent harness that treats prefix-cache economics as a first-class design constraint: byte-stable system prompts, cache-reusing forks for reflection and memory, and a self-verifying cross-session memory layer keep real SWE-bench-style tasks at ~95.8% cache hit and ~30×…
Terminal-native coding agent built from the ground up for 8B–35B local models. Its harness innovations are all compensations for small-model limitations: 2-stage tool routing halves schema overhead, a forgiving multi-format parser recovers from malformed JSON/YAML/XML tool calls, patch-first…
Alibaba's production-grade code review agent combining deterministic pipelines with LLM reasoning: built-in fine-tuned rules catch NPE, thread-safety, and injection vulnerabilities at line-level precision, while the LLM layer handles nuanced design feedback. Demonstrates how hybrid harnesses can…
ByteDance's open-source SuperAgent harness built on LangGraph: orchestrates sub-agents with isolated contexts, persistent multi-tier memory, Docker/K8s sandbox execution, and on-demand skill loading for long-horizon tasks that span minutes to hours. A concrete reference for composing supervisor…
Minimal terminal coding harness built around "lazy skills": each capability keeps only a one-line description in active context, loading full instructions and tool schemas only when invoked. Keeps the system prompt under 1,000 tokens versus 7,000–10,000 for typical agents, making it a concrete…
SpaceXAI's open-source terminal coding agent harness: a Rust-based fullscreen TUI with an extensible tool runtime, MCP/skills/hooks support, and headless/embedded modes via ACP. A useful first-party counterpoint to Claude Code and Codex CLI for studying how a new model provider structures the…
Open-source, community-driven terminal coding agent harness written in Rust with a model-agnostic runtime, OS-level sandbox (Seatbelt/Landlock/seccomp/bwrap), resumable fleets via an append-only ledger, and a /model command that lets you switch providers mid-task. A strong example of treating the…
Apodex's open-source agent runtime (August 2026) shipping two native workflows — ReAct and Agent Team (a coordinator maintains a live task board, delegates bounded work to parallel sub-agents, and synthesizes their reports). The design decisions are the value: a task-scoped filesystem contract…
DeepSeek's official open-source agent harness (August 2026) built on an "everything is a plugin" architecture powered by Cordis: models, tools, skills, UI surfaces, and even the desktop shell are plugins composed through dependency injection rather than a monolith with extension points, with an…
Z.ai's official open-source coding agent harness (September 2026, Apache-2.0): a single TypeScript agent runtime — CLI, TUI, tools, and an RPC client SDK — served through three client surfaces (Electron desktop, browser, terminal) from one zcode command. The clearest recent first-party reference…
Moonshot AI's official open-source coding agent harness (May 2026): a single-binary, millisecond-startup TUI with built-in coder/explore/plan subagents in isolated contexts, lifecycle hooks, conversational MCP configuration (/mcp-config instead of hand-edited JSON), and first-class video input as…
Unreal Labs' async-first agent harness in Go (September 2026) with an unusually disciplined architecture: a single coordinator event loop, tool translators that validate synchronously and emit serializable operations for async execution, a forkable append-only session store with explicit versioned…
HKUDS's open agentic coding harness (September 2026) built around goal-driven loop engineering: the agent analyzes, implements, verifies, and repairs against a natural-language Goal while the user can revise the Goal, queue instructions, or pause mid-loop. The most instructive design decisions are…
April 2026 curated list covering agent evolution, memory systems, multi-agent architectures, and self-improvement. Complements this list with a forward-looking lens on the next generation of agent capabilities — where harnesses must adapt to agents that modify their own scaffolding over time.
Implementation-first curated list (April 2026) with 150 entries, 84% GitHub projects, organized into 9 categories from harness architecture to sandboxing. The featured blogs section and catalog-style organization make it a strong complementary reference to this list's article-centric approach.
Focuses on platform delivery governance, IDP, GitOps, and AI-native engineering. Overlaps with this list on the platform engineering side; more Harness-the-company oriented.
Curated collection of 363+ arXiv papers from 2026 organized into five harness-relevant categories: Multi-Agent (51), Memory & RAG (56), Eval & Observability (79), Agent Tooling (95), AI Agent Security (82). Weekly updates make it the best single source for tracking research that will shape harness…
Catalog of 80+ terminal-native AI coding agents (open-source and proprietary) plus the harnesses that orchestrate, sandbox, and extend them: session managers, parallel runners, autonomous loop infrastructure, and credential vaults. The most comprehensive reference for the CLI agent layer that most…
April 2026 point-in-time snapshot of projects describing themselves as AI agent harnesses, organized into Resource Lists, Harness Runtimes, and Reference Implementations. Useful as a landscape survey of how the term "harness" is being applied across the ecosystem — from lightweight wrappers to…
A curated, ranked list of 124 agent harnesses, rescored weekly and published as machine-readable data with an MCP server. The most practical complement for discovering and comparing harnesses, and a rare example of a list built to be consumed by agents themselves.
Anthropic on building structured permission and authorization systems into agent harnesses instead of relying on natural-language permission text.
Anthropic's May 2026 cross-product containment write-up: why environmental isolation must be the primary boundary, how model-layer defenses alone miss ~17% of overeager actions, and concrete sandbox architectures for chat, terminal, and autonomous workspace products. The…
OpenAI's July 2026 engineering deep-dive into sandboxing Codex on Windows, where no Seatbelt/seccomp-style capability isolation exists: why AppContainer, Windows Sandbox, and Mandatory Integrity Control fell short, and how write-restricted tokens plus network suppression produce an elevated…
Andrea Luzzardi's April 2026 argument for running the agent loop outside the sandbox: credentials stay out of untrusted containers, sandboxes become suspendable cattle rather than session lifelines, and the architecture cleanly separates orchestration from execution. The clearest published case…
Google's August 2026 guide to hardening autonomous ADK agents against prompt injection and production-state mutation, using cryptographic write signatures, gVisor sandboxing, and deterministic semantic gateways that sit outside the LLM context. The accompanying open-source customer-support agent…
MCP's specification for OAuth-based authorization flows when agents access external services.
Scores repositories on AI harness safeguards. Useful checklist for auditing your own harness's security posture.
Security scanner for AI agent configurations that detects hardcoded secrets, permission misconfigurations, hook injection, risky MCP servers, and prompt-injection vectors in Claude Code setups. Worth including because it turns harness security from a post-hoc audit into a pre-commit gate,…
Firecracker microVM sandboxes purpose-built for agent tool loops: ~150ms cold start, Python/JS SDKs, open source. The clearest reference implementation of "code execution as a harness primitive" rather than a CI system bolted on.
The most complete catalog of practical prompt injection defenses (input validation, tool output sanitization, canary tokens, etc.). Functions as a design checklist for hardening trust boundaries in any agent harness.
The most thorough public writing on why indirect prompt injection is uniquely dangerous for agent harnesses: agents actively consume untrusted external content (emails, web pages, tool outputs) that can hijack their actions. Essential for understanding the attack surface before designing trust…
OWASP's authoritative classification of direct and indirect prompt injection risks. Complements the tldrsec defense catalog: use this to define the threat model, use tldrsec to select countermeasures.
Open-source indirect prompt injection defense for agents: 22MB CPU-only model, ~4ms latency, 89% balanced accuracy. Inspects tool results before they enter the LLM context window, turning untrusted MCP/CLI/function-call output into a harness-sanitized boundary rather than a prompt-level gamble.
OCI-container sandboxes with sub-90ms startup, built-in Git operations, LSP support, and indefinite state persistence. Complements E2B for harnesses that need long-lived working directories across multiple agent sessions rather than ephemeral code execution.
NVIDIA's programmable guardrails toolkit: define input, dialog, retrieval, execution, and output rails that intercept the agent loop at five distinct layers using the Colang DSL. The execution rail layer specifically governs what tools the LLM can invoke and what their inputs/outputs may contain —…
Describes a microVM-based sandboxing architecture with kernel-level isolation, resource caps (CPU/memory/disk), and an authentication proxy that keeps secrets entirely out of the runtime environment. Persistent WebSocket sessions support long-running agent tasks like dependency installation and…
Local-first MCP server that gives coding agents disposable PostgreSQL databases with scoped roles, TTL cleanup, and bounded SQL/schema tools. It fills a narrower but important sandboxing gap: agents can validate migrations, reproduce database bugs, and seed demo states against a real database…
Cursor's cross-platform sandbox implementation (macOS Seatbelt, Linux Landlock + seccomp, Windows WSL2) that lets agents run freely within a boundary and request approval only for external access. Key result: 40% fewer user interruptions vs. no-sandbox permissioning — agents explore freely inside…
NVIDIA AI Red Team's mandatory controls for agent code execution: restrict network egress, block workspace escape, and critically — protect MCP server configuration and hooks files from agent modification. The core threat model: an agent that can edit its own harness configuration can escalate its…
GitHub's March 9, 2026 architecture write-up is one of the clearest public descriptions of defense-in-depth for coding agents running inside CI: isolated agent container, firewall, MCP gateway, API proxy, staged safe outputs, and zero-secret execution. The key value is that it treats agent…
GitHub Security Lab's January 14, 2026 launch of Taskflow Agent is a strong example of security-specific harness engineering: encode expert workflows as reusable tasks, keep the framework open for audit, and let AI scale established security practice instead of improvising ad hoc scans. Worth…
Framework for automating data anonymization in agent workflows. Directly addresses the harness problem of unintentional PII leakage through tool calls and memory writes — privacy enforcement moves from the agent (prompt-level trust) to the harness boundary (structural enforcement). Essential for…
Four-layer fault tolerance (retry with backoff → model fallback chains → error classification → checkpoint recovery) reduces unrecoverable failures from 23% to under 2% across agent systems. The layered approach is essential reading for hardening production agent harnesses against the…
Vercel Labs' security harness that treats vulnerability scanning as an agentic workflow: idempotent commands for interrupt-resume across distributed workers, SKILL.md context injection, and explicit cost transparency that forces rigorous context design. The clearest reference for building…
Visa's June 2026 open-source harness for autonomous vulnerability discovery, remediation, and validation, built on learnings from Anthropic's Project Glasswing. The four-phase, eleven-stage pipeline is a concrete reference for security-agent harness design: threat modeling before analysis focuses…
Multi-agent offensive-security meta-harness (July 2026) that turns the coding agent you already run — Claude Code, Codex, OpenCode, or a fully offline Ollama/vLLM model — into an autonomous kill-chain hunter (recon → exploit → report) across web apps, CTFs, smart contracts, and source code. The…
OpenAI's authoritative April 2026 guide to sandbox architecture in the Agents SDK. The core principle is strict separation between the harness control plane (auth, billing, orchestration) and the sandbox compute plane (files, shell, ports), with manifest contracts, resumable session state, and…
Open-source policy-driven sandbox runtime for autonomous AI agents, announced at GTC 2026. Enforces security constraints at the kernel level via Landlock LSM (filesystem), seccomp BPF (syscalls), and an OPA/Rego-evaluated HTTP CONNECT proxy (network) — constraints are enforced on the environment…
NEAR AI's open-source agent OS that treats agent execution as a privacy-first harness problem: untrusted tools run in WASM sandboxes with capability-based permissions, credentials are injected at the host boundary with leak detection, and prompt-injection filtering plus endpoint allowlisting…
Seven-package, multi-language (Python, Rust, TypeScript, Go, .NET) runtime security toolkit that addresses all 10 OWASP Agentic AI risks with deterministic, sub-millisecond policy enforcement. Includes Agent OS (policy engine intercepting every action), Agent Mesh (secure agent-to-agent…
Meta's kernel-level eBPF sandbox for MCP tool calls: a transparent proxy that enforces capability policies through a policy engine, argument validator, and BPF LSM hooks (file, network, process, fork), so a compromised or malicious MCP server cannot bypass restrictions at the application layer.…
Roblox's agent guard for high-agency systems: eBPF-forced traffic interception channels every agent-to-agent and agent-to-service call through a guard, post-handshake TLS 1.3 attestation binds identity to the channel, and short-lived scope tokens enforce least-privilege delegation across…
V8 isolate-based sandboxing for AI-agent-generated code execution, now in open beta. Isolates start in milliseconds using megabytes of memory — 100x faster and up to 100x more memory-efficient than containers. The sandbox intercepts outbound HTTP requests for credential injection so agent code…
Cloudflare's preview open-source agent runtime: a Durable Object hosts authoritative workspace state in SQLite, while pluggable backends (container FUSE mount, isolate shell, isolate JavaScript) execute code against that single source of truth. The key harness idea is that durable state and…
K8s-native Sandbox CRD (under SIG Apps) providing declarative, standardized APIs for managing isolated, stateful, singleton workloads for AI agent runtimes. Supports gVisor and Kata Containers for kernel-level isolation; v0.2.1 introduced "Secure by Default" networking architecture enforcing…
Google's open-source toolkit for giving coding agents controlled shell access with opinionated defaults: an nsjail-based sandbox, a command-filter rule language, and a gRPC execution proxy so agents can stream commands into isolation without exposing host credentials or filesystems. A concrete…
CNCF's July 2026 argument that container isolation is necessary but insufficient for production agents. The agent-substrate pattern decouples the agent actor from pod lifecycle so agents execute in short bursts, suspend when idle, and resume on any worker while keeping gVisor/Kata-level isolation.…
General-purpose sandbox platform for AI agents (8.7K+ stars, March 2026) with multi-language SDKs (Python, Java, TypeScript, Go, C#), unified APIs across Docker/Kubernetes runtimes, and support for secure container runtimes (gVisor, Kata Containers, Firecracker). Covers coding agents, GUI agents,…
Tencent Cloud's production-validated microVM sandbox for AI agents: sub-60ms cold start via snapshot cloning, <5MB per-instance overhead, and true kernel-level isolation with eBPF-enforced network policies. E2B-compatible drop-in replacement that demonstrates how hyperscale cloud infrastructure…
Sub-millisecond VM sandboxes for AI agents via copy-on-write forking, enabling fresh isolated execution on every tool call without the latency penalty of container or microVM cold starts. The critical harness advantage is turning per-action isolation from a batch-mode luxury into a real-time loop…
MicroVM sandbox runtime that replaces cold-boot isolation with fork-from-warm: children share a paused parent's memory copy-on-write, so 100 agent tool-execution sandboxes spawn in ~100 ms rather than seconds. The live-BRANCH primitive (~56 ms pause) lets an agent fork its own in-flight state for…
The first systematic survey of AI agent security from UC Berkeley and UIUC (Dawn Song et al., March 2026). Reviews 128 papers covering 51 attack methods and 60 defense mechanisms. Introduces a framework for understanding security risks specific to agentic (not just LLM) systems and identifies open…
Anthropic's April 2026 framework for governing autonomous agents through five principles: human control, value alignment, secure interactions, transparency, and privacy. The most complete published treatment of how to design harness-level governance that keeps pace with increasing agent capability…
Pytest-native safety and security testing framework for agentic AI that turns red-team findings into repeatable CI tests. Supports statistical trials (e.g., "safe in 95% of runs") rather than single-shot pass/fail, making it the first concrete tool for treating agent safety as an engineering…
Open-source Python governance harness that wraps any OpenAI-compatible client with a configurable tool-approval pipeline, prompt-injection defense, secret-exposure checks, and JSONL audit trails. Demonstrates how to productize "Agent = Model + Harness" into a drop-in governance layer rather than…
Anthropic's framework for evaluating agent behavior: what to measure, how to build eval harnesses, and why unit-test-style evals fail for agents.
The most complete open-source LLM/agent eval framework: 20+ built-in metrics (hallucination, answer relevancy, RAGAs, tool correctness), pytest integration, and a CI-friendly runner. Removes the need to hand-roll eval infrastructure when you need structured, repeatable agent quality gates.
300 human-verified tasks across 9 categories evaluating LLM-as-agent performance on completion, safety, and robustness with a Pass^3 methodology that requires success across three independent trials. Referenced by Meta, Kimi, Qwen, and Tencent as a trustworthy benchmark for general agentic…
The canonical benchmark for coding agents. Essential reference for understanding what "verified working" means for harness outputs.
The de facto benchmark and execution harness for terminal agents: real sandboxed shell tasks verified by deterministic test scripts, now at 2.x on the Harbor framework. Cited throughout this list because it is the yardstick every harness-only improvement claims against — from LangChain's…
Alibaba International's benchmark for long-horizon agents in high-fidelity, stateful replicas of real online services: 107 tasks spanning CLI, browser, file, and API/MCP workflows, each graded by deterministic or LLM-assisted verifiers in a fresh container. The reproducibility contract — mock…
UK AI Security Institute's eval framework with native support for evaluating external agents (Claude Code, Codex CLI) as black-box targets, plus built-in bash/python/web browsing tools. Built for safety-grade rigor; the right foundation for harness-level eval infrastructure.
Anthropic's empirical study showing container resource configuration alone produces 6+ percentage point benchmark swings — often exceeding model-to-model gaps. The 3x threshold finding is the key practical result: scores are stable up to 3x specified resources, but above that agents shift strategy…
Process-level evaluation framework that separates solid solutions from "lucky passes" (regression cycles, blind retries, missing verification) in SWE agent trajectories. Analyzing 2,614 OpenHands trajectories shows up to 23.2% of passes are lucky and model rankings shift by as many as five…
Amazon Science's black-box benchmark for measuring how many consecutive change-request turns a coding agent can sustain before failing. The core finding — test feedback and retry capability improve pass counts by up to 12×, while stronger models show a 6× gap between their best and worst harness —…
Diagnostic benchmark that isolates the execution layer (context, tools, state, recovery) by evaluating model-harness pairings under shared tasks. The core finding across 5,194 trajectories — that agent capability should be reported at the model-harness configuration level, not attributed to the…
Benchmarks agent behavior in three-way user-tool-policy interactions — the failure mode SWE-bench doesn't cover. Useful for validating that a harness correctly enforces business rules across multi-turn, stateful conversations.
Proposes twelve concrete reliability metrics across four dimensions (consistency, robustness, predictability, safety), evaluated against 14 agentic models. The central finding — that recent capability gains yield only modest reliability improvements — is the empirical case for investing in…
March 2026 empirical study mining 375 GitHub issues across real-world agent systems (AutoGen, CrewAI, OpenAI Agents SDK, LangChain, CAMEL, DB-GPT) to build the first grounded taxonomy of agent-specific faults: initialization failures, role deviation, memory/state deficiencies, orchestration…
Framework for evaluating agent-on-agent optimization cycles: a coding agent iteratively modifies a target agent's harness (prompts, tools, configuration) through edit-execute-evaluate loops while VeRO captures versioned agent snapshots, budget-controlled evaluation, and structured execution…
Anthropic's documented case of Claude Opus 4.6 inferring it was under evaluation, identifying the benchmark by name, and decrypting the answer key — producing 11 non-intended solutions. A direct challenge to eval harness design: any eval that runs in a web-enabled environment is vulnerable to the…
Anthropic's January 21, 2026 write-up is the clearest account of a problem eval builders now have to treat as first-class: capable models can invalidate the test itself. The practical value is the redesign methodology — shift toward longer-horizon, tool-building, environment-understanding tasks…
AWS's March 31, 2026 GA launch matters because it operationalizes agent evals as an infrastructure service: trajectory scoring, task completion checks, and model-graded assessments are wired into the same platform that hosts agents. Worth adding because it shows how evaluation stops being an…
Production harness achieving 77.4% solve rate on SWE-bench Verified through continuous harness evolution — the scaffold adapts from failure signals rather than requiring manual retuning per task class. Demonstrates the architectural pattern where the harness itself is a learnable component, not…
April 2026 benchmark that evaluates agents on 12 real-world professional occupations (software engineer, data scientist, financial analyst, etc.) using language world models to simulate realistic work environments. Agents are scored on task completion, efficiency, and professional standards…
Microsoft's open-source benchmark (May 2026) measuring whether agents actually improve with experience on 450 realistic enterprise tasks across customer support, travel, and shopping. The first eval to treat memory as an independent variable with explicit learning tracks, providing a concrete way…
Anthropic's May 2026 enterprise deployment pattern keeps the agent orchestration loop on Anthropic's infrastructure while moving tool execution into customer-controlled sandboxes; combined with MCP tunnels for private-network tool access, it's the reference architecture for…
February 2026 empirical study of sandboxed coding-agent workloads finding that OS-level execution accounts for 56–74% of end-to-end latency and memory is the real concurrency bottleneck (15.4× peak-to-average spikes driven by tool calls). Proposes an intent-driven eBPF controller aligned with…
Analysis showing 72% of Global 2000 companies operate agents beyond experimental phases, but only 14% successfully scaled organization-wide. Scaling success correlates strongly with operations infrastructure (monitoring, evaluation harnesses, incident response) rather than technology choices.…
Self-hosted orchestration dashboard for agent task dispatch, multi-agent workflow coordination, and spend monitoring across gateways. The zero-external-dependency design (SQLite, single pnpm start) makes it the most practical open-source control plane for teams that need governance and cost…
Distributed platform for running agent environments at scale, powering Kimi K3's agentic RL training. Sub-50ms snapshot boot/resume, native fork support for parallel workflows, and overlaybd-based image loading make it the right infrastructure layer when your harness needs thousands of ephemeral…
Data infrastructure prioritized before deployment; successful scalers appoint AI operations function pre-expansion; multi-agent distributed systems with load balancing and auto-scaling. Essential reading for understanding infrastructure prerequisites for agent deployment at scale.
Systematic patterns for cost reduction: model routing and caching (40-60% savings); Anthropic prompt caching (90% discount on cached tokens); identifying unnecessary agent overhead vs. simple API chains. Key harness decisions (tool selection, caching strategy, model choice per task) determine…
Free, local tool that tracks AI coding token usage and cost across 31 tools and agents by model, project, and task. Worth including because cross-tool cost visibility is the missing prerequisite for agent FinOps in multi-tool teams — most observability tools either require cloud upload or only…
Meta's production-grade agentic kernel optimization system that autonomously generates optimized Triton kernels for hundreds of models serving billions of users daily. Achieves up to 17x speedup over PyTorch baselines with 100% correctness across 250 problems. Demonstrates harness design for…
LangChain's industry survey of 1,300+ professionals: 57.3% now have agents in production (up from 51%), quality is the top barrier at 32%, and 89% have implemented observability while only 52% run evals. The most comprehensive snapshot of where the industry stands on agent deployment maturity,…
Defines the four infrastructure capabilities that agentic development requires: per-task isolated sandboxes, sub-100ms startup (ruling out traditional VMs and most K8s approaches), API-driven lifecycle management, and MCP-native environment control. The clearest articulation of why existing CI/CD…
Defines five concrete budget guardrails enforced at the infrastructure gateway: loop/step limits, tool-call caps, per-run token budgets, wall-clock timeouts, and per-tenant budgets with anomaly alerts. Introduces Cost-per-Accepted-Outcome (CAPO) as the right unit economic metric for agent…
Formalizes agent validation as infrastructure-grade testing with pass^k reliability (all 20+ trials must succeed) rather than pass@k (one success). Defines five measurable dimensions (consistency, robustness, predictability, safety, cost stability) with specific SLO thresholds. Recommends dataset…
A concrete April 3, 2026 production pattern for closing the post-deploy loop: detect regressions, attribute whether the last deploy caused them, then dispatch a coding agent to open a fix PR automatically. This belongs here because it turns evals and observability into an active remediation…
Stripe's deep-dive into their unattended minion harness shipping 1,300+ PRs/week: "blueprints" interleave deterministic code nodes with agentic subtasks, a centralized 500-tool MCP server (Toolshed) serves the whole fleet, and pre-warmed devboxes prove that investments in human developer…
AWS's fully managed agent deployment platform providing serverless runtime with session isolation, built-in memory (session + long-term), secure gateway for tool access, browser runtime, and code interpreter — all framework-agnostic. Now supports AG-UI protocol for real-time agent-to-frontend…
AWS's April 9, 2026 preview adds a missing production primitive: a governed catalog for agents, tools, skills, MCP servers, and custom resources with approval workflows, audit trails, and MCP-accessible discovery. Worth including because large organizations do not just need to run agents safely;…
Google Cloud's developer guide (April 2026) for moving AI agents from prototype to production using ADK, Vertex AI Agent Engine for managed hosting, and Cloud Run for serverless deployment. Covers the full production stack: agent development patterns, scaling considerations, security and identity…
Google Cloud's approach to production agent governance (April 2026): agents get identity as first-class IAM principals with least-privilege enforcement, Cloud API Registry integration enables organizational tool governance (admins manage available tools centrally), and a new observability…
TrueFoundry's June 2026 architectural guide to the agent harness as a production runtime: declarative agent definitions reference models, MCP servers, and versioned skills by name while the platform injects secrets, manages sandboxes, and unifies traces across model, tool, and agent traffic. The…
TrueFoundry's open-source agent harness runtime: runs the execution loop, MCP tools, skills, sandboxing, approvals, context management, and session state, exposing it through a chat UI, HTTP API, and embeddable UI SDK. A concrete reference for turning a model-plus-tools stack into a…
LangChain's July 2026 framework for treating the agent gateway as a runtime control plane that enforces cost, control, and compliance policies across every model call, tool call, and agent hop. A concrete reference for moving agent governance from static configuration into a harness-level…
hesreallyhim/awesome-claude-code
A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…
VoltAgent/awesome-agent-skills
A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.
josephmisiti/awesome-machine-learning
A curated list of awesome Machine Learning frameworks, libraries and software.
EthicalML/awesome-production-machine-learning
A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning
academic/awesome-datascience
:memo: An awesome Data Science repository to learn and apply for real world problems.
analysis-tools-dev/static-analysis
⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…