Skip to content
69

Awesome LLM Reasoning

From Chain-of-Thought prompting to OpenAI o1 and DeepSeek-R1 🍓

3.7k stars214 forks155 entriesLast push Apr 20, 2026 (5 months ago)License MIT

This page lists names, links and short descriptions. The original list on GitHub is the source and belongs to its authors.

Survey >2025

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

[code]

Recent Advances in Large Language Model Benchmarks Against Data Contamination: From Static to Dynamic Evaluation.

[code]

Survey >2024

Attention Heads of Large Language Models: A Survey.

[code]

Internal Consistency and Self-Feedback in Large Language Models: A Survey.

[code]

Puzzle Solving using Reasoning of Large Language Models: A Survey.

[code]

Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Survey >2022

Towards Reasoning in Large Language Models: A Survey.

[code]

Reasoning with Language Model Prompting: A Survey.

[code]

Analysis >2025

New Trends for Modern Machine Translation with Large Reasoning Models.

Analysis >2024

Are Your LLMs Capable of Stable Reasoning?

[code]

From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond.

To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning.

Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

A Peek into Token Bias: Large Language Models Are Not Yet Genuine Reasoners.

[code]

Iteration Head: A Mechanistic Study of Chain-of-Thought

Do Large Language Models Latently Perform Multi-Hop Reasoning?

Premise Order Matters in Reasoning with Large Language Models.

The Impact of Reasoning Step Length on Large Language Models.

Large Language Models Cannot Self-Correct Reasoning Yet.

At Which Training Stage Does Code Data Help LLM Reasoning?

Analysis >2023

Measuring Faithfulness in Chain-of-Thought Reasoning.

Faith and Fate: Limits of Transformers on Compositionality.

In 2 lists

Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting.

[code]

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity.

Large Language Models Can Be Easily Distracted by Irrelevant Context.

by Google Research et al., 2023 - Adding the instruction "Feel free to ignore irrelevant information given in the questions." consistently improves robustness to irrelevant context.

In 3 lists

On Second Thought, Let's Not Think Step by Step! Bias and Toxicity in Zero-Shot Reasoning.

Bias and Toxicity in Zero-Shot Reasoning.

In 2 lists

Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters.

[code]

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

[code]

Analysis >2022

Emergent Abilities of Large Language Models.

[blog]

In 2 lists

Can language models learn from explanations in context?

Analysis >2025

JudgeLRM: Large Reasoning Models as a Judge.

Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination.

[code]

CRANE: Reasoning with constrained LLM generation.

Sketch-of-Thought: Efficient LLM Reasoning with Adaptive Cognitive-Inspired Sketching.

[code]

Self-rewarding correction for mathematical reasoning.

Competitive Programming with Large Reasoning Models.

s1: Simple test-time scaling.

(2025) - small open reasoning recipe via budget-forcing.

In 2 lists

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

[project]

In 2 lists

Towards System 2 Reasoning in LLMs: Learning How to Think With Meta Chain-of-Thought.

Analysis >2024

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

[code]

PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models

[code] Simin Chen, XiaoNing Feng, Xiaohong Han, Cong Liu, Wei Yang FSE'24

DRT-o1: Optimized Deep Reasoning Translation via Long Chain-of-Thought.

[code]

MALT: Improving Reasoning with Multi-Agent LLM Training.

SmartAgent: Chain-of-User-Thought for Embodied Personalized Agent in Cyber World.

Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions.

[code] [model]

In 2 lists

Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning.

Deliberate Reasoning for LLMs as Structure-aware Planning with Accurate World Model.

[code]

Interpretable Contrastive Monte Carlo Tree Search Reasoning.

Training Language Models to Self-Correct via Reinforcement Learning.

OpenAI o1.

(2024-2025) - test-time-compute reasoning systems.

In 2 lists

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents.

DotaMath: Decomposition of Thought with Code Assistance and Self-correction for Mathematical Reasoning.

[code]

LLM-ARC: Enhancing LLMs with an Automated Reasoning Critic.

Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning.

Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models.

[code]

Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing.

Self-playing Adversarial Language Game Enhances LLM Reasoning.

Evaluating Mathematical Reasoning Beyond Accuracy.

Advancing LLM Reasoning Generalists with Preference Trees.

LLM3: Large Language Model-based Task and Motion Planning with Motion Failure Reasoning.

[code]

Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking.

GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements.

Chain-of-Thought Reasoning Without Prompting.

V-STaR: Training Verifiers for Self-Taught Reasoners.

InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Self-Discover: Large Language Models Self-Compose Reasoning Structures.

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

K-Level Reasoning with Large Language Models.

Efficient Tool Use with Chain-of-Abstraction Reasoning.

Teaching Language Models to Self-Improve through Interactive Demonstrations.

Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic.

[code]

Chain-of-Verification Reduces Hallucination in Large Language Models.

Skeleton-of-Thought: Large Language Models Can Do Parallel Decoding.

Question Decomposition Improves the Faithfulness of Model-Generated Reasoning.

[code]

Let's Verify Step by Step.

process-supervised reward models for reasoning.

In 2 lists

REFINER: Reasoning Feedback on Intermediate Representations.

[project] [code]

Active Prompting with Chain-of-Thought for Large Language Models.

[code]

Language Models as Inductive Reasoners.

Analysis >2023

Boosting LLM Reasoning: Push the Limits of Few-shot Learning with Reinforced In-Context Pruning.

Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning.

[code]

Recursion of Thought: A Divide and Conquer Approach to Multi-Context Reasoning with Language Models.

[code] [poster]

Reasoning with Language Model is Planning with World Model.

Reasoning Implicit Sentiment with Chain-of-Thought Prompting.

[code]

Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

[code]

In 4 listsDetails

SatLM: Satisfiability-Aided Language Models Using Declarative Prompting.

[code]

ART: Automatic multi-step reasoning and tool-use for large language models.

Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data.

[code]

Synthetic Prompting: Generating Chain-of-Thought Demonstrations for Large Language Models.

Generating Chain-of-Thought Demonstrations for Large Language Models.

In 2 lists

Faithful Chain-of-Thought Reasoning.

Rethinking with Retrieval: Faithful Large Language Model Inference.

by University of Pennsylvania et al., 2022 - They shows the potential of enhancing LLMs by retrieving relevant external knowledge based on decomposed reasoning steps obtained through chain-of-thought (CoT) prompting. I predict we're going to see many of these types of retrieval-enhanced LLMs in…

In 2 lists

LAMBADA: Backward Chaining for Automated Reasoning in Natural Language.

Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions.

[code]

Large Language Models are Reasoners with Self-Verification.

[code]

Can Retriever-Augmented Language Models Reason? The Blame Game Between the Retriever and the Language Model.

[code]

Complementary Explanations for Effective In-Context Learning.

Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

[code]

Unsupervised Explanation Generation via Correct Instantiations.

PAL: Program-aided Language Models.

[project] [code]

Solving Math Word Problems via Cooperative Reasoning induced Language Models.

[code]

Large Language Models Can Self-Improve.

Mind's Eye: Grounded language model reasoning through simulation.

Automatic Chain of Thought Prompting in Large Language Models.

[code]

Language Models are Multilingual Chain-of-Thought Reasoners.

Ask Me Anything: A simple strategy for prompting language models.

[code]

Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning.

[project] [code]

Making Large Language Models Better Reasoners with Step-Aware Verifier.

In 2 lists

Least-to-most prompting enables complex reasoning in large language models.

Self-consistency improves chain of thought reasoning in language models.

Multi-path sampling + majority vote: GSM8K 57% → 74%

In 5 listsDetails

Analysis >2022

Retrieval Augmentation for Commonsense Reasoning: A Unified Approach.

[code]

Language Models of Code are Few-Shot Commonsense Learners.

[code]

Solving Quantitative Reasoning Problems with Language Models.

[blog]

In 2 lists

Large Language Models Still Can't Plan.

[code]

Large Language Models are Zero-Shot Reasoners.

"Let's think step by step" — zero-shot CoT milestone

In 4 listsDetails

Iteratively Prompt Pre-trained Language Models for Chain of Thought.

[code]

Chain of Thought Prompting Elicits Reasoning in Large Language Models.

[blog]

In 5 listsDetails

Analysis >2025

Introducing Visual Perception Token into Multimodal Large Language Model.

[code] [model] [dataset]

LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

[project] [code] [model]

Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

[project][code][dataset] Wenqi Zhang, Mengna Wang, Gangao Liu, Xu Huixin, Yiwei Jiang, Yongliang Shen, Guiyang Hou, Zhe Zheng, Hang Zhang, Xin Li, Weiming Lu, Peng Li, Yueting Zhuang. Preprint'25

In 2 lists

Analysis >2024

Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

[code] [model]

LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

code model

In 2 lists

Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models.

[project] [code]

In 2 lists

Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs.

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

[project]

Chain-of-Table: Evolving Tables in the Reasoning Chain for Table Understanding.

Link-Context Learning for Multimodal LLMs.

[code]

Analysis >2023

Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models.

G-LLaVA: Solving Geometric Problems with Multi-Modal Large Language Model.

Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models.

[project] [code]

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

[project] [code] [demo]

ViperGPT: Visual Inference via Python Execution for Reasoning.

[project] [code]

Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

[code]

In 3 lists

Multimodal Chain-of-Thought Reasoning in Language Models.

[code]

In 3 lists

Visual Programming: Compositional Visual Reasoning without Training.

[project] [code]

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

[project] [code]

Analysis >2025

Learning to Reason from Feedback at Test-Time.

[code]

S²R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning.

[code]

rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

[code]

In 2 lists

Analysis >2024

MathScale: Scaling Instruction Tuning for Mathematical Reasoning.

Analysis >2023

Learning Deductive Reasoning from Synthetic Corpus based on Formal Logic.

[code]

Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step.

[code]

Specializing Smaller Language Models towards Multi-Step Reasoning.

Large Language Models Are Reasoning Teachers.

[code]

Teaching Small Language Models to Reason.

They finetune a student model on the chain of thought (CoT) outputs generated by a larger teacher model. For example, the accuracy of T5 XXL on GSM8K improves from 8.11% to 21.99% when finetuned on PaLM-540B generated chains of thought.

In 2 lists

Distilling Multi-Step Reasoning Capabilities of Large Language Models into Smaller Models via Semantic Decompositions.

Analysis >2022

Scaling Instruction-Finetuned Language Models.

by Google - They find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks. Flan-PaLM 540B achieves SoTA performance on several benchmarks. They…

In 2 lists

Other Useful Resources

Agent Shadow Brain

Self-evolving AI coding intelligence with infinite memory (TurboQuant), genetic algorithm self-evolution, predictive bug detection, PageRank knowledge graphs, swarm intelligence, and adversarial defense.

In 2 lists

LLM Reasoners

A library for advanced large language model reasoning.

Chain-of-Thought Hub

Benchmarking LLM reasoning performance with chain-of-thought prompting.

In 2 lists

Omni Skills Forge

50,000+ curated AI agent skills for Claude Code, Cursor, Copilot, Windsurf, Cline. Visual dashboard, one-click install, skill doctor, auto-update.

In 2 lists

ThoughtSource

Central and open resource for data and tools related to chain-of-thought reasoning in large language models.

In 2 lists

AgentChain

Chain together LLMs for reasoning & orchestrate multiple large models for accomplishing complex tasks.

In 2 lists

google/Cascades

Python library which enables complex compositions of language models such as scratchpads, chain of thought, tool use, selection-inference, and more.

LogiTorch

PyTorch-based library for logical reasoning on natural language.

salesforce/LAVIS

One-stop Library for Language-Vision Intelligence.

In 4 listsDetails

facebookresearch/RAM

A framework to study AI models in Reasoning, Alignment, and use of Memory (RAM).

See category
94

Awesome OpenClaw Skills

VoltAgent/awesome-openclaw-skills

The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞

Fresh★ 53k830 entriesPushed today
92

Awesome DeepSeek Harness (DSH) Plugin

awesome-dsh-plugin/awesome-dsh-plugin

A curated list of plugins for DeepSeek Harness (dsh) · DeepSeek Harness 插件精选列表

Fresh★ 17k1654 entriesPushed today
91

Awesome Guidelines

Kristories/awesome-guidelines

Programming style, best practices, and coding conventions.

Fresh★ 11k166 entriesPushed 2 days ago
90

Awesome

sindresorhus/awesome

😎 Awesome lists about all kinds of interesting topics [NOTE: Pull requests are temporarily disabled until I have a chance to catch up with the existing ones]

Fresh★ 513k51 entriesPushed 28 days ago
90

Awesome Prompts

ai-boost/awesome-prompts

Curated list of chatgpt prompts from the top-rated GPTs in the GPTs Store. Prompt Engineering, prompt attack & prompt protect. Advanced Prompt Engineering papers.

Fresh★ 9k288 entriesPushed yesterday
90

Awesome README

matiassingers/awesome-readme

A curated list of awesome READMEs

Fresh★ 22k143 entriesPushed yesterday