Skip to content
86

Awesome Autoresearch

A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.

2.5k stars193 forks105 entriesLast push Sep 21, 2026 (9 days ago)License Other

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

General-purpose descendants

kayba-ai/recursive-improve

Recursive self-improvement framework where agents capture execution traces, analyze failure patterns, and apply targeted fixes with keep-or-revert evaluation.

vukrosic/auto-research

Docs-only control plane for an open autonomous AI research lab — file-based operating model for human direction and agent execution.

uditgoenka/autoresearch

Claude Code skill that generalizes autoresearch into a reusable loop for software, docs, security, shipping, debugging, and other measurable goals.

leo-lilinxiao/codex-autoresearch

Codex-native autoresearch skill with resume support, lessons across runs, optional parallel experiments, and mode-specific workflows.

junjunjunbong/research-loop

Autoresearch-style Agent Skill for Codex and Claude Code with a deterministic runner, plan-hash approval, isolated Git worktrees, authoritative metric evaluation, and an append-only experiment ledger.

xieyulai/steer

Governed experiment framework where coding agents edit training code and run rounds while the task, scorer, and evidence stay fixed.

SeeleAI/Thoth

Dashboard-first Claude Code and Codex runtime for autoresearch, with durable runs, locked work items, visible ledgers, and reviewable verdicts.

supratikpm/gemini-autoresearch

Gemini CLI skill that generalises autoresearch to any measurable goal. Gemini-native: uses Google Search grounding as a live verification source inside the loop, true headless overnight mode via --yolo --prompt, and 1M token context. Also works in Antigravity IDE via .agents/skills/.

davebcn87/pi-autoresearch

pi extension plus dashboard for persistent experiment loops, live metrics, confidence tracking, and resumable autoresearch sessions.

drivelineresearch/autoresearch-claude-code

Claude Code plugin/skill port of pi-autoresearch, with a clean experiment-loop workflow and a concrete biomechanics case study.

greyhaven-ai/autocontext

Closed-loop control plane for repeated agent improvement, with evaluation, persistent knowledge, staged validation, and optional distillation into cheaper local runtimes.

In 2 lists

Necmttn/ax

Local retro loop for AI coding agents: captures session traces, turns repeated friction into proposals, and tracks accepted fixes as experiments.

jmilinovich/goal-md

Generalizes autoresearch into a GOAL.md pattern for repos where the agent must first construct a measurable fitness function before it can optimize.

james-s-tayler/lazy-developer

Claude Code skill that orchestrates autoresearch across a prioritized sequence of optimization goals (coverage, test speed, build speed, complexity, LOC, performance) using GOAL.md as the engine. Supports standalone and Ralph Mode multi-instance execution.

mutable-state-inc/autoresearch-at-home

Collaborative fork of upstream autoresearch that adds experiment claiming, shared best-config syncing, hypothesis exchange, and swarm-style coordination across many single-GPU agents.

zkarimi22/autoresearch-anything

Generalizes autoresearch to any measurable metric — system prompts, API performance, landing pages, test suites, config tuning, SQL queries. "If you can measure it, you can optimize it."

Entrpi/autoresearch-everywhere

Cross-platform expansion that auto-detects hardware config and starts the loop. The "glue and generalization" half of autoresearch.

ShengranHu/ADAS

Automated Design of Agentic Systems — ICLR 2025. Meta-agents that invent novel agent architectures by programming them in code.

In 2 lists

MaximeRobeyns/self_improving_coding_agent

SICA: Self-Improving Coding Agent that edits its own codebase. ICLR 2025 Workshop paper demonstrating scaffold-level self-improvement on coding benchmarks.

peterskoett/self-improving-agent

Alternative self-improving agent architecture with reflection and meta-learning cycles.

metauto-ai/HGM

Huxley-Gödel Machine for coding agents — applies self-improvement to SWE-bench performance via meta-level optimization.

gepa-ai/gepa

GEPA (Genetic-Pareto) — ICLR 2026 Oral. Reflective prompt evolution that outperforms RL (GRPO) on benchmarks. Optimizes any textual parameters against any metric using natural language reflection.

In 3 lists

sentient-agi/EvoSkill

Automated skill discovery for coding agents: evolves reusable skills and prompts from failed trajectories against benchmarks, with support for Claude Code, Codex CLI, OpenCode, OpenHands, and Goose.

MrTsepa/autoevolve

GEPA-inspired autoresearch for self-play: mutate code strategies, evaluate head-to-head, rate with Elo/Bradley-Terry, branch from the Pareto front. Agent reads match traces to target mutations. Works as a Claude Code skill.

HKUDS/ClawTeam

Agent swarm intelligence for autoresearch — spawns parallel GPU research directions, distributes work across agents, aggregates results.

In 2 lists

Orchestra-Research/AI-Research-SKILLs

Comprehensive skill library including autoresearch orchestration with two-loop architecture (inner optimization + outer synthesis).

In 2 lists

WecoAI/aideml

AIDE: Tree-search ML engineering agent that autonomously improves model performance via iterative code generation and evaluation.

In 4 listsDetails

weco.ai

Weco: Cloud platform for AIDE with observability, experiment tracking, and managed runs — brings the autoresearch loop into production.

In 2 lists

Research-agent systems

aiming-lab/AutoResearchClaw

End-to-end research pipeline that turns a topic into literature review, experiments, analysis, peer review, and paper drafts; broader than autoresearch, but clearly in the same lineage.

In 5 listsDetails

OpenLAIR/dr-claw

Open-source research workspace with sequential idea-to-paper pipelines and integrated autoresearch tool packs.

In 2 lists

OpenRaiser/NanoResearch

End-to-end autonomous research engine that plans experiments, generates code, runs jobs locally or on SLURM, analyzes real results, and writes papers grounded in those outputs.

In 3 lists

kaust-ark/ARK

ARK (Automatic Research Kit): idea + venue → paper pipeline orchestrating 6 agents — proposal analysis, literature search, Slurm experiments, LaTeX drafting, iterative peer review. Controlled via CLI, web dashboard, or Telegram.

wanshuiyin/Auto-claude-code-research-in-sleep

Markdown-first research workflows for Claude Code and other agents, centered on autonomous literature review, experiments, paper iteration, and cross-model critique.

In 7 listsDetails

skyllwt/AutoSci

Wiki-centric full-lifecycle research platform built on Claude Code, realizing Karpathy's LLM-Wiki vision. 20+ skills cover the full loop: ingest → ideate → novelty check → experiment design / run / eval → paper writing. Research state lives in a structured knowledge wiki with an interactive graph.

Sibyl-Research-Team/AutoResearch-SibylSystem

Fully autonomous AI scientist built on Claude Code, with explicit AutoResearch lineage, multi-agent research iteration, GPU experiment execution, and a self-evolving outer loop.

wjc2830/Easy-AutoResearch-for-DeepLearning

Claude Code skill that runs an autoresearch-style, human-gated deep-learning loop across six roles, versioned experiments, and evidence-checked completion.

eimenhmdt/autoresearcher

Early open-source package for automating scientific workflows, currently centered on literature-review generation with an ambition toward broader autonomous research.

hyperspaceai/agi

Distributed, peer-to-peer research network where autonomous agents run experiments, gossip findings, maintain CRDT leaderboards, and archive results to GitHub across multiple research domains.

Human-Agent-Society/CORAL

CORAL: Autonomous multi-agent evolution for open-ended discovery (arXiv:2604.01658). Long-running agents with shared persistent memory, asynchronous execution, and heartbeat-based interventions; SOTA on 10 math/algorithmic/systems tasks.

In 2 lists

SakanaAI/AI-Scientist

The AI Scientist: First comprehensive system for fully automatic scientific discovery. From idea generation to paper writing with minimal human supervision.

In 5 listsDetails

SakanaAI/AI-Scientist-v2

Workshop-level automated scientific discovery via agentic tree search. Removes template dependency from v1, generalizes across research domains.

In 2 lists

AweAI-Team/AiScientist

AiScientist: long-horizon ML research lab with hierarchical orchestration and File-as-Bus coordination — workspace files act as the durable system of record. Drives autonomous paper-reproduction (PaperBench) and competition-style MLE-Bench iteration loops under fixed compute/time budgets. (arXiv…

HKUDS/AI-Researcher

NeurIPS 2025 paper. Full end-to-end research automation: hypothesis → experiments → manuscript → peer review. Production version at novix.science.

In 2 lists

openags/Auto-Research

OpenAGS: Orchestrates a team of AI agents across the full research lifecycle — lit review, hypothesis generation, experiments, manuscript writing, and peer review.

SamuelSchmidgall/AgentLaboratory

End-to-end autonomous research workflow: idea → literature review → experiments → report. Supports both autonomous and co-pilot modes.

In 2 lists

AgentRxiv

Collaborative autonomous research framework where agent laboratories share a preprint server to build on each other's work iteratively.

JinheonBaek/ResearchAgent

Iterative research idea generation over scientific literature with LLMs. Multi-agent review and feedback loops.

du-nlp-lab/MLR-Copilot

Autonomous ML research framework — generates ideas, implements experiments, analyzes results.

MASWorks/ML-Agent

Reinforcing LLM agents for autonomous ML engineering. Learns from trial and error to improve model performance.

In 2 lists

PouriaRouzrokh/LatteReview

Low-code Python package for automated systematic literature reviews via AI-powered agents.

LitLLM/LitLLM

AI-powered literature review assistant using RAG for accurate, well-structured related-work sections in academic writing.

Agent Laboratory

Three-phase research pipeline: Literature Review → Experimentation → Report Writing, with specialized agents for each phase.

In 2 lists

happyhappy-jun/writing-driven-autoresearch

Autoresearch-style harness that keeps a submittable paper from the first minute and drives every experiment from the claims in that draft, looping modify → measure → verify → revise. 1st place at the Ralphthon@ICML 2026 autonomous-research hackathon.

AutoResearch-Factory/Agon

End-to-end research orchestrator built on one cornerstone principle, Prompt Economy (reusable loops, not one-off prompts), plus five supporting rules; runs scientist/coder/auditor loops across 10+ disciplines, same reusable-loop lineage as autoresearch but scaled to full research programs.

In 2 lists

Platform ports and hardware forks

gianfrancopiana/openclaw-autoresearch

OpenClaw port of pi-autoresearch; autonomous experiment loop for any optimization target with statistical confidence scoring.

miolini/autoresearch-macos

Widely adopted macOS fork that adapts upstream autoresearch for Apple Silicon / MPS while preserving the original loop shape.

trevin-creator/autoresearch-mlx

MLX-native Apple Silicon port that keeps the upstream fixed-budget val_bpb loop while removing the PyTorch/CUDA dependency entirely.

jsegov/autoresearch-win-rtx

Windows-native RTX fork focused on consumer NVIDIA GPUs, with explicit VRAM floors and a practical desktop setup path.

iii-hq/n-autoresearch

Multi-GPU autoresearch infrastructure with structured experiment tracking, adaptive search strategy, crash recovery, and queryable orchestration around the classic train.py loop.

lucasgelfond/autoresearch-webgpu

Browser/WebGPU port that lets agents generate training code, run experiments in-browser, and feed results back into the loop without a Python setup.

tonitangpotato/autoresearch-engram

Fork with persistent cognitive memory — frequency-weighted retrieval of cross-session knowledge for improved experiment continuity.

Colab/Kaggle T4 port

Adapts autoresearch for free T4 GPUs (Google Colab / Kaggle) with zero cost and zero local setup. Key changes: Flash Attention 3 → PyTorch SDPA, removes H100-only kernel dependency.

In 5 listsDetails

ArmanJR-Lab/autoautoresearch

Jetson AGX Orin port with a director — a Go binary that acts as a "creative director" injecting novelty (arxiv papers + DeepSeek Reasoner) into the loop to escape local minima. Includes multi-experiment comparison (baseline vs director-guided) with detailed stall analysis.

Domain-specific adaptations

mattprusak/autoresearch-genealogy

Applies the autoresearch pattern to genealogy, using structured prompts, archive guides, source checks, and vault workflows to iteratively expand and verify family-history research.

ArchishmanSengupta/autovoiceevals

Uses adversarial callers plus keep-or-revert prompt edits to harden voice AI agents across Vapi, Smallest AI, and ElevenLabs.

chrisworsey55/atlas-gic

Applies the autoresearch keep-or-revert loop to trading agents, optimizing prompts and portfolio orchestration against rolling Sharpe ratio instead of model loss.

In 2 lists

RightNow-AI/autokernel

Applies the autoresearch loop to GPU kernel optimization: profile bottlenecks, edit one kernel, benchmark, keep or revert, repeat.

ElliotXie/autozyme

Multi-agent framework that applies the autoresearch keep-or-revert loop to CPU-side scientific software: profile a target function, generate one optimization candidate, benchmark for speed while preserving the original outputs, keep or revert, repeat.

In 3 lists

Agent-Analytics/autoresearch-growth

Applies autoresearch to landing-page positioning and A/B test candidates, using analytics snapshots and measured experiment results to seed subsequent rounds.

Rkcr7/autoresearch-sudoku

Enhanced autoresearch workflow where an AI agent iteratively rewrites and benchmarks a Rust sudoku solver, ultimately beating leading human-built solvers on hard benchmark sets.

jeongph/autospec

Reads natural-language business rules and autonomously builds a Spring Boot service with tests via the keep-or-revert loop. Evaluates with Gradle build + JUnit XML. 119-line skeleton to 950 lines in 5 cycles.

vlasenkoalexey/tpu_performance_autoresearch_wiki

Applies the autoresearch keep-or-revert loop to TPU model performance (MFU / tokens-per-sec) on v6e hardware: profiles each run through an XProf MCP server, makes one model-code change per experiment, and keeps or reverts against measured MFU. Pairs the loop with a Karpathy-style LLM wiki for…

Evaluation & benchmarks

snap-stanford/MLAgentBench

Benchmark suite for evaluating AI agents on ML experimentation tasks. 13 tasks from CIFAR-10 to BabyLM.

OpenAI/mle-bench

OpenAI's benchmark for measuring how well AI agents perform at ML engineering.

In 3 lists

chchenhui/mlrbench

MLR-Bench: Evaluating AI agents on open-ended ML research. 201 tasks from NeurIPS/ICLR/ICML workshops.

gersteinlab/ML-Bench

Evaluates LLMs and agents for ML tasks on repository-level code.

THUDM/AgentBench

Comprehensive benchmark for LLM-as-Agent evaluation across 8 distinct environments. ICLR 2024.

In 7 listsDetails

Notable use cases and writeups

Shopify Liquid optimization

Tobi Lütke shared an autoresearch-style optimization run on Shopify's Liquid engine, with public traces showing major parse/render speedups and allocation reductions. (tweet)

In 2 lists

Driveline baseball biomechanics

Public autoresearch-style experiment loop for pitch-velocity prediction from biomechanics data, with large reported gains in model quality.

Tennis XGBoost prediction + reward hacking writeup

Nick Oak documents an autoresearch-inspired loop for tennis match prediction, including where the optimization setup went wrong. (repo · gamed branch)

Vesuvius Challenge ink detection swarm

Multi-agent experimental loop applied to ancient-scroll ink detection, with a strong writeup on cross-scroll generalization improvements.

Earth system model optimization

Hybrid workflow where an LLM proposes equation structures and a search process tunes parameters, showing how the autoresearch pattern extends into scientific modeling. (tweet)

GPUMode eigh kernel competition retrospective

Eliana documents a (mostly) autonomous system that took 3rd in GPUMode's eigh kernel competition: 1,293 experiments over two weeks, an 8.56x speedup over torch, and a detailed account of caught reward hacks and failure modes. (tweet)

The Agentic Researcher

Paper: "A Practical Guide to AI-Assisted Research in Mathematics and Machine Learning." Cites autoresearch as the canonical example of automated ML experiment pipelines.

Scaling Autoresearch to GPU Clusters

SkyPilot blog on running autoresearch on H100/H200 clusters with cloud orchestration.

Self-Improving Coding Agents

Addy Osmani's practical guide to setting up self-improving agent loops with Claude Code.

autoresearch@home: Distributed AI Research

SETI@home model applied to autoresearch — contribute GPU time to collective model optimization.

Claude Code + AutoResearch for Self-Improving Skills

MindStudio guide to building self-improving AI skills using Claude Code with autoresearch patterns.

100 ML Experiments Overnight

Particula technical breakdown with domain-agnostic fork applications.

PM's Guide to Autoresearch

Product manager's guide covering setup, community forks, and real-world applications.

Autoresearch 101 Builder's Playbook

Substack deep-dive on applying autoresearch patterns to prompts, agents, and workflows with concrete examples.

Kingy AI Technical Breakdown

Detailed technical walkthrough of the autoresearch loop architecture, mutation operators, and fitness function design.

Fortune Feature

Business and industry context on why autoresearch matters for the future of autonomous AI agents.

What's Missing in Autonomous Research?

Codes 56 autonomous research systems along seven axes and argues the real gap isn't generating research artifacts (most systems can) but catching weak ones before they ship.

ai-agents-2030/awesome-deep-research-agent

Curated list of deep research agent papers and systems.

YoungDubbyDu/LLM-Agent-Optimization

Papers on LLM agent optimization methods.

VoltAgent/awesome-ai-agent-papers

Curated AI agent papers from 2026 — agent engineering, memory, evaluation, workflows, and autonomous systems.

In 3 listsDetails

masamasa59/ai-agent-papers

AI agent research papers updated biweekly via automated arxiv search with curated selection.

tmgthb/Autonomous-Agents

Autonomous agents research papers, updated daily.

HKUST-KnowComp/Awesome-LLM-Scientific-Discovery

EMNLP 2025 survey on LLMs in scientific discovery.

In 2 lists

openags/Awesome-AI-Scientist-Papers

Collection of AI Scientist / Robot Scientist papers.

In 2 lists

agenticscience.github.io

Survey: "From AI for Science to Agentic Science: A Survey on Autonomous Scientific Discovery."

dspy.ai/GEPA

DSPy integration of GEPA reflective prompt optimizer for compound AI systems.

OpenAI Cookbook: Self-Evolving Agents

Cookbook for autonomous agent retraining using GEPA-style reflective evolution.

WecoAI/awesome-autoresearch

Curated list of AutoResearch use cases with verifiable traces and progress charts, organized by domain (LLM training, GPU kernels, voice agents, trading, etc.).

See category
94

Table of Contents

hesreallyhim/awesome-claude-code

A hand-picked collection of the finest of resources for the most awesome of agents, Claude Code, the undisputed champion of coding companions, from the unstoppable team…

Fresh★ 55k202 entriesPushed today
94

Awesome Agent Skills

VoltAgent/awesome-agent-skills

A curated collection of 1000+ agent skills from official dev teams and the community, compatible with Claude Code, Codex, Gemini CLI, Cursor, and more.

Fresh★ 35k839 entriesPushed today
93

Awesome Machine Learning

josephmisiti/awesome-machine-learning

A curated list of awesome Machine Learning frameworks, libraries and software.

Fresh★ 74k1188 entriesPushed 7 days ago
92

Awesome Production Machine Learning

EthicalML/awesome-production-machine-learning

A curated list of awesome open source libraries to deploy, monitor, version and scale your machine learning

Fresh★ 21k519 entriesPushed 3 days ago
92

AWESOME DATA SCIENCE

academic/awesome-datascience

:memo: An awesome Data Science repository to learn and apply for real world problems.

Fresh★ 30k881 entriesPushed today
91

Static Analysis

analysis-tools-dev/static-analysis

⚙️ A curated list of static analysis (SAST) tools and linters for all programming languages, config files, build tools, and more. The focus is on tools which improve…

Fresh★ 15k528 entriesPushed 8 days ago