Skip to content

DeepSoftwareAnalytics/Awesome-Issue-Resolution

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

19 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

✨ Awesome Issue Resolution

Advances and Frontiers of LLM-based Issue Resolution in Software Engineering A Comprehensive Survey

GitHub Stars Forks Awesome Paper arXiv Tables Contributors Papers Count

📖 Documentation Website | 📄 Full Paper | 📋 Tables & Resources

🎙️ Interactive Exploration:

NotebookLM Discord Issues

Awesome Issue Resolution

📖 Abstract

Based on a systematic review of 176 papers and online resources, this survey establishes a holistic theoretical framework for Issue Resolution in software engineering. We examine how Large Language Models (LLMs) are transforming the automation of GitHub issue resolution. Beyond the theoretical analysis, we have curated a comprehensive collection of datasets and model training resources, which are continuously synchronized with our GitHub repository and project documentation website.

🔍 Explore This Survey:


📚 Complete Paper List

Total: 176 works across 14 categories

📊 Evaluation Datasets

Benchmarks for evaluating issue resolution systems

  • SWE-bench Lite: SWE-bench: Can Language Models Resolve Real-world Github Issues? (2024)
  • SWE-bench Verified: Introducing SWE-bench Verified | OpenAI (2024)
  • SWE-bench-java: SWE-bench-java: A GitHub Issue Resolving Benchmark for Java (2024) arXiv
  • Visual SWE-bench: CodeV: Issue Resolving with Visual Data (2025) ACL DOI
  • SWE-Lancer: SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? (2025)
  • FEA-Bench: FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature Implementation (2025) ACL
  • Multi-SWE-bench: Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving (2025) OpenReview
  • SWE-PolyBench: SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents (2025) arXiv
  • SWE-bench Multilingual: SWE-smith: Scaling Data for Software Engineering Agents (2025) OpenReview
  • SwingArena: SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving (2025) arXiv
  • SWE-bench Multimodal: SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? (2025) OpenReview
  • OmniGIRL: Omnigirl: A multilingual and multimodal benchmark for github issue resolution (2025)
  • SWE-bench-Live: SWE-bench Goes Live! (2025) OpenReview
  • SWE-Factory: SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks (2025) arXiv
  • SWE-MERA: SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks (2025) arXiv
  • SWE-Perf: SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories? (2025) OpenReview
  • SWE-Bench Pro: SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? (2025) arXiv
  • SWE-InfraBench: SWE-InfraBench: Evaluating Language Models on Cloud Infrastructure Code (2025) OpenReview
  • SWE-Sharp-Bench: SWE-Sharp-Bench: A Reproducible Benchmark for C# Software Engineering Tasks (2025) arXiv
  • SWE-fficiency: SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads? (2025) arXiv
  • SWE-Compass: SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models (2025) arXiv
  • SWE-EVO: SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios (2025) arXiv

🎯 Training Datasets

Datasets for training issue resolution systems

  • SWE-bench-extra: SWE-bench: Can Language Models Resolve Real-world Github Issues? (2024)
  • Multi-SWE-RL: Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving (2025) OpenReview
  • R2E-Gym: R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents (2025) OpenReview
  • SWE-Synth: SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs (2025) arXiv
  • LocAgent: OrcaLoca: An LLM Agent Framework for Software Issue Localization (2025) OpenReview
  • SWE-Smith: SWE-smith: Scaling Data for Software Engineering Agents (2025) OpenReview
  • SWE-Fixer: SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution
  • SWELoc: SweRank: Software Issue Localization with Code Ranking (2025) arXiv
  • SWE-Gym: Training Software Engineering Agents and Verifiers with SWE-Gym
  • SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner (2025) OpenReview
  • SWE-Factory: SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks (2025) arXiv
  • Skywork-SWE: Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs (2025) arXiv
  • RepoForge: RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale (2025) arXiv
  • SWE-Mirror: SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories (2025) arXiv
  • SWE-Lego: SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving (2026) arXiv

🤖 Single-Agent Systems

Individual autonomous agents for issue resolution

  • SWE-agent: Swe-agent: Agent-computer interfaces enable automated software engineering (2024)
  • Aider (2026) Website
  • Devin: SWE-bench technical report (2025) Website
  • PatchPilot: PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification (2025) OpenReview
  • LCLM: Putting It All into Context: Simplifying Agents with LCLMs (2025) arXiv
  • DGM: Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents (2025) arXiv
  • Trae Agent: Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling (2025) arXiv
  • Live-SWE-agent: SE-Agent: Self-Evolution Trajectory Optimization in Multi-Step Reasoning with LLM-Based Agents (2025) OpenReview
  • Lita: Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs (2025) arXiv
  • TOM-SWE: TOM-SWE: User Mental Modeling For Software Engineering Agents (2025) arXiv
  • Confucius Code Agent: Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases (2025) arXiv

👥 Multi-Agent Systems

Collaborative multi-agent frameworks

  • MAGIS: MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution (2024) OpenReview
  • AutoCodeRover: AutoCodeRover: Autonomous Program Improvement (2024) DOI
  • CodeR: CodeR: Issue Resolving with Multi-Agent and Task Graphs (2024) arXiv
  • OpenHands: OpenHands: An Open Platform for AI Software Developers as Generalist Agents (2025) OpenReview
  • AgentScope: SWE-Bench - AgentScope (2025) Website
  • OrcaLora: OrcaLoca: An LLM Agent Framework for Software Issue Localization (2025) OpenReview
  • DEI: Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents (2025) OpenReview
  • MarsCode Agent: MarsCode Agent: AI-native Automated Bug Fixing (2024) arXiv
  • Lingxi: Lingxi/docs/Lingxi Technical Report 2505.pdf at master · lingxi-agent/Lingxi (2026) GitHub
  • Devlo: Achieving SOTA on SWE-bench (2026) Website
  • Refact.ai Agent: AI Coding Agent for Software Development - Refact.ai (2025) Website
  • HyperAgent: HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale (2024) arXiv
  • SWE-Search: SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement (2025) OpenReview
  • CodeCoR: CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation (2025) arXiv
  • Agent KB: Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving (2025) arXiv
  • SWE-Debate: SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution (2026)
  • SWE-Exp: SWE-Exp: Experience-Driven Software Issue Resolution (2025) arXiv
  • Meta-RAG: Meta-RAG on Large Codebases Using Code Summarization (2025) arXiv

🔄 Workflow-Based Methods

Structured pipeline approaches

  • Agentless: Demystifying LLM-Based Software Engineering Agents (2025) Website
  • Conversational Pipeline: Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench (2024) arXiv
  • SynFix: SynFix: Dependency-Aware Program Repair via RelationGraph Analysis (2025) ACL DOI
  • CodeV: CodeV: Issue Resolving with Visual Data (2025) ACL DOI
  • GUIRepair: Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing (2025)

🛠️ Tool-Augmented Methods

Methods leveraging external tools

  • MAGIS: MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution (2024) OpenReview
  • AutoCodeRover: AutoCodeRover: Autonomous Program Improvement (2024) DOI
  • SWE-agent: Swe-agent: Agent-computer interfaces enable automated software engineering (2024)
  • Alibaba LingmaAgent: Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration (2025) DOI
  • OpenHands: OpenHands: An Open Platform for AI Software Developers as Generalist Agents (2025) OpenReview
  • SpecRover: SpecRover: Code Intent Extraction via LLMs (2025)
  • MarsCode Agent: MarsCode Agent: AI-native Automated Bug Fixing (2024) arXiv
  • RepoGraph: RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
  • SuperCoder2.0: SuperCoder2.0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer (2024) arXiv
  • EvoCoder: LLMs as Continuous Learners: Improving the Reproduction of Defective Code in Software Issues (2024) arXiv
  • AEGIS: AEGIS: An Agent-based Framework for General Bug Reproduction from Issue Descriptions (2025)
  • CoRNStack: CoRNStack: High-Quality Contrastive Data for Better Code Retrieval and Reranking (2025) OpenReview
  • OrcaLoca: OrcaLoca: An LLM Agent Framework for Software Issue Localization (2025) OpenReview
  • DARS: DARS: Dynamic Action Re-Sampling to Enhance Coding Agent Performance by Adaptive Tree Traversal (2025) arXiv
  • Otter: Otter: Generating Tests from Issues to Validate SWE Patches (2025) OpenReview
  • Quadropic Insiders: Quadropic Insiders : Syntheo Tops Swelite Feb (2025) Website
  • Issue2Test: Issue2Test: Generating Reproducing Test Cases from Issue Reports (2025) arXiv
  • KGCompass: Enhancing repository-level software repair via repository-aware knowledge graphs (2025) arXiv
  • CoSIL: Issue Localization via LLM-Driven Iterative Code Graph Searching (2025)
  • InfantAgent-Next: InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction (2025) arXiv
  • Co-PatcheR: Co-PatcheR: Collaborative Software Patching with Component(s)-specific Small Reasoning Models (2025) arXiv
  • SWERank: SweRank: Software Issue Localization with Code Ranking (2025) arXiv
  • Nemotron-CORTEXA: Nemotron-CORTEXA: Enhancing LLM Agents for Software Engineering Tasks via Improved Localization and Solution Diversity (2025) OpenReview
  • LCLM: Putting It All into Context: Simplifying Agents with LCLMs (2025) arXiv
  • SACL: SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization (2025) arXiv
  • SWE-Debate: SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution (2026)
  • OpenHands-Versa: Coding Agents with Multimodal Browsing are Generalist Problem Solvers
  • SemAgent: SemAgent: A Semantics Aware Program Repair Agent (2025) arXiv
  • Repeton: Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles (2025) arXiv
  • cAST: cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree (2025) arXiv
  • Prometheus: Prometheus: Unified Knowledge Graphs for Issue Resolution in Multilingual Codebases (2025) arXiv
  • Git Context Controller: Git Context Controller: Manage the Context of LLM-based Agents like Git (2025) arXiv
  • Trae Agent: Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling (2025) arXiv
  • BugPilot: BugPilot: Complex Bug Generation for Efficient Learning of SWE Skills (2025) arXiv
  • TestPrune: When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution (2025) arXiv
  • Meta-RAG: Meta-RAG on Large Codebases Using Code Summarization (2025) arXiv
  • InfCode: InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution (2025) arXiv
  • GraphLocator: GraphLocator: Graph-guided Causal Reasoning for Issue Localization (2025) arXiv

🧠 Memory-Enhanced Methods

Systems with memory mechanisms

  • Infant Agent: Infant Agent: A Tool-Integrated, Logic-Driven Agent with Cost-Effective API Usage (2024) arXiv
  • EvoCoder: LLMs as Continuous Learners: Improving the Reproduction of Defective Code in Software Issues (2024) arXiv
  • Learn-by-interact: Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments (2025) OpenReview
  • DGM: Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents (2025) arXiv
  • ExpeRepair: EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair (2025) arXiv
  • Agent KB: Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving (2025) arXiv
  • SWE-Exp: SWE-Exp: Experience-Driven Software Issue Resolution (2025) arXiv
  • RepoMem: Improving Code Localization with Repository Memory (2025) arXiv
  • AgentDiet: Improving the Efficiency of LLM Agent Systems through Trajectory Reduction (2025) arXiv
  • ReasoningBank: ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory (2025) arXiv
  • MemGovern: MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences (2026) arXiv

📚 Supervised Fine-Tuning (SFT)

Models trained via supervised fine-tuning

  • Lingma SWE-GPT: SWE-GPT: A Process-Centric Language Model for Automated Software Improvement (2025)
  • ReSAT: Repository Structure-Aware Training Makes SLMs Better Issue Resolver (2024) arXiv
  • Scaling data collection: Scaling Data Collection for Training SWE Agents (2024) arXiv
  • CodeXEmbed: CodeXEmbed: A Generalist Embedding Model Family for Multilingual and Multi-task Code Retrieval (2025) OpenReview
  • SWE-Gym: Training Software Engineering Agents and Verifiers with SWE-Gym
  • Thinking Longer: Thinking Longer, Not Larger: Enhancing Software Engineering Agents via Scaling Test-Time Compute (2025) arXiv
  • Search for training: Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents (2025) arXiv
  • Co-PatcheR: Co-PatcheR: Collaborative Software Patching with Component(s)-specific Small Reasoning Models (2025) arXiv
  • MCTS-Refined CoT: MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue Resolution (2025) arXiv
  • SWE-Swiss: SWE-Swiss: A Multi-Task Fine-Tuning and RL Recipe for High-Performance Issue Resolution (2025) arXiv
  • Devstral: Devstral: Fine-tuning Language Models for Coding Agent Applications (2025) arXiv
  • Kimi-Dev: Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents (2025) arXiv
  • SWE-Compressor: Context as a Tool: Context Management for Long-Horizon SWE-Agents (2025) arXiv
  • SWE-Lego: SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving (2026) arXiv
  • Agentic Rubrics: Agentic Rubrics as Contextual Verifiers for SWE Agents (2026) arXiv

🎮 Reinforcement Learning (RL)

Models trained via reinforcement learning

  • SWE-RL: SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution (2025) OpenReview
  • SoRFT: SoRFT: Issue Resolving with Subtask-oriented Reinforced Fine-Tuning (2025) ACL
  • SEAlign: SEAlign: Alignment Training for Software Engineering Agent (2026)
  • SWE-Dev1: SWE-Dev: Evaluating and Training Autonomous Feature-Driven Software Development (2025) arXiv
  • Satori-SWE: Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering (2025) arXiv
  • Agent-RLVR: Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards (2025) arXiv
  • DeepSWE: DeepSWE: Training a State-of-the-Art Coding Agent from Scratch by Scaling RL (2025) arXiv
  • SWE-Dev2: SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling (2025) arXiv
  • Tool-integrated RL: Tool-integrated Reinforcement Learning for Repo Deep Search (2025) arXiv
  • SWE-Swiss: SWE-Swiss: A Multi-Task Fine-Tuning and RL Recipe for High-Performance Issue Resolution (2025) arXiv
  • SeamlessFlow: SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling (2025) arXiv
  • DAPO: Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning (2025) arXiv
  • CoreThink: CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs (2025) arXiv
  • CWM: CWM: An Open-Weights LLM for Research on Code Generation with World Models (2025) arXiv
  • EntroPO: Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization (2025) arXiv
  • Kimi-Dev: Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents (2025) arXiv
  • FoldGRPO: Scaling Long-Horizon LLM Agent via Context-Folding (2025) arXiv
  • GRPO-based Method: A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning (2025) OpenReview
  • TSP: Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair (2025) ACL DOI
  • Self-play SWE-RL: Toward Training Superintelligent Software Agents through Self-Play SWE-RL (2025) arXiv
  • SWE-Playground: Training Versatile Coding Agents in Synthetic Environments (2025) arXiv
  • Supervised RL: Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning (2025) arXiv
  • OSCA: Scaling LLM Inference Efficiently with Optimized Sample Compute Allocation (2025) ACL DOI
  • SWE-RM: SWE-RM: Execution-free Feedback For Software Engineering Agents (2025) arXiv
  • One Tool Is Enough: One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents (2025) arXiv
  • Let It Flow: Let It Flow: Agentic Crafting on Rock and Roll, Building the ROME Model within an Open Agentic Learning Ecosystem (2025) arXiv
  • KAT-Coder: KAT-Coder Technical Report (2025) arXiv
  • Seed1.5-Thinking: Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning (2025) arXiv
  • Deepseek V3.2: DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models (2025) arXiv
  • Kimi-K2-Instruct: Kimi K2: Open Agentic Intelligence (2025) arXiv
  • GLM-4.6: gpt-oss-120b & gpt-oss-20b model card (2025) arXiv
  • Qwen3-Coder: Qwen3 Technical Report (2025) arXiv
  • GLM-4.6: Glm-4.5: Agentic, reasoning, and coding (arc) foundation models (2025) arXiv
  • Minimax M2: MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention (2025) arXiv
  • LongCat-Flash-Think: Introducing LongCat-Flash-Thinking: A Technical Report (2025) arXiv
  • MiMo-V2-Flash: MiMo-V2-Flash Technical Report (2026) arXiv

⚡ Inference-Time Scaling

Methods for scaling at inference time

  • SWE-Search: SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement (2025) OpenReview
  • ReasoningBank: CodeMonkeys: Scaling Test-Time Compute for Software Engineering (2025) arXiv
  • SWE-PRM: When Agents go Astray: Course-Correcting SWE Agents with PRMs (2025) arXiv
  • SIADAFIX: SIADAFIX: issue description response for adaptive program repair (2025) arXiv

📥 Data Collection Methods

Techniques for collecting training data

  • SWE-rebench: SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents (2025) OpenReview
  • RepoLaunch: SWE-bench Goes Live! (2025) OpenReview
  • SWE-Factory: SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks (2025) arXiv
  • SWE-MERA: SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks (2025) arXiv
  • RepoForge: RepoForge: Training a SOTA Fast-thinking SWE Agent with an End-to-End Data Curation Pipeline Synergizing SFT and RL at Scale (2025) arXiv
  • Multi-Docker-Eval: Multi-Docker-Eval: A `Shovel of the Gold Rush' Benchmark on Automatic Environment Building for Software Engineering (2025) arXiv

🔬 Data Synthesis Methods

Approaches for synthetic data generation

  • Learn-by-interact: Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic Environments (2025) OpenReview
  • R2E-Gym: R2E-Gym: Procedural Environment Generation and Hybrid Verifiers for Scaling Open-Weights SWE Agents (2025) OpenReview
  • SWE-Synth: SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs (2025) arXiv
  • SWE-smith: SWE-smith: Scaling Data for Software Engineering Agents (2025) OpenReview
  • SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner (2025) OpenReview
  • SWE-Mirror: SWE-Mirror: Scaling Issue-Resolving Datasets by Mirroring Issues Across Repositories (2025) arXiv

📈 Data Analysis

Analysis of datasets and benchmarks

  • SWE-bench Verified: Introducing SWE-bench Verified | OpenAI (2024)
  • Patch Correctness: Are "Solved Issues" in SWE-bench Really Solved Correctly? An Empirical Study (2025) arXiv
  • UTBoost: UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench (2025) arXiv
  • Trustworthiness: Is Your Automated Software Engineer Trustworthy? (2025) arXiv
  • Rigorous agentic benchmarks: Establishing Best Practices for Building Rigorous Agentic Benchmarks (2025) arXiv
  • The SWE-Bench Illusion: The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason (2025) arXiv
  • Revisiting SWE-Bench: Revisiting SWE-Bench: On the Importance of Data Quality for LLM-Based Code Models (2025) DOI
  • SPICE: SPICE: An Automated SWE-Bench Labeling Pipeline for Issue Clarity, Test Coverage, and Effort Estimation (2025)
  • Data contamination: Does SWE-Bench-Verified Test Agent Ability or Model Memory? (2025) arXiv

🔍 Methods Analysis

Comparative analysis of different methods

  • Context Retrieval: On The Importance of Reasoning for Context Retrieval in Repository-Level Code Editing (2024) arXiv
  • Evaluating software development agents: Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios (2025) DOI
  • Overthinking: The Danger of Overthinking: Examining the Reasoning-Action Dilemma in Agentic Tasks (2025) arXiv
  • Beyond final code: Beyond Final Code: A Process-Oriented Error Analysis of Software Development Agents in Real-World GitHub Scenarios (2025) arXiv
  • GSO: GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents (2025) arXiv
  • Dissecting the SWE-Bench Leaderboards: Dissecting the SWE-Bench Leaderboards: Profiling Submitters and Architectures of LLM- and Agent-Based Repair Systems (2025) arXiv
  • Security analysis: How Safe Are AI-Generated Patches? A Large-scale Study on Security Risks in LLM and Agentic Automated Program Repair on SWE-bench (2025) arXiv
  • Failures analysis: An Empirical Study on Failures in Automated Issue Solving (2025) arXiv
  • SeaView: SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow (2025) arXiv
  • SWEnergy: SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs (2026)
  • Strong-Weak Model Collaboration: An Empirical Study on Strong-Weak Model Collaboration for Repo-level Code Generation (2025) arXiv
  • Agents in the Wild (2025) Website

📋 Statistical Tables

Comprehensive tables and statistics about issue resolution datasets, methods, and benchmarks.

Evaluation & Training Datasets

A comprehensive survey and statistical overview of issue resolution datasets. We categorize these datasets based on programming language, modality support, source repositories, data scale (Amount), and the availability of reproducible execution environments.

Dataset Language Multimodal Repos Amount Environment Link
Single-PL Datasets
SWE-Fixer Python 856 115,406 GitHub HuggingFace HuggingFace
SWE-smith Python 128 50k GitHub HuggingFace
SWE-Lego Python 3,251 32,119 GitHub HuggingFace
SWE-rebench Python 3,468 21,336 GitHub HuggingFace
SWE-bench-train Python 37 19k GitHub HuggingFace
SWE-Flow Python 74 18,081 GitHub
Skywork-SWE Python 2,531 10,169 -
R2E-Gym Python 10 8,135 GitHub HuggingFace
RepoForge Python - 7.3k -
SWE-bench-extra Python 2k 6.38k HuggingFace
SWE-Gym Python 11 2,438 GitHub HuggingFace
SWE-bench Python 12 2,294 GitHub HuggingFace
SWE-bench-java Java 19 1,797 GitHub HuggingFace
FEA-bench Python 83 1,401 GitHub HuggingFace
SWE-bench-Live Python 164 1,565 GitHub HuggingFace
Loc-Bench Python - 560 GitHub HuggingFace
SWE-bench Verified Python - 500 GitHub HuggingFace
SWE-bench Lite Python 12 300 GitHub HuggingFace
SWE-MERA Python 200 300 GitHub HuggingFace
SWE-Bench-CL Python 8 273 GitHub
SWE-Sharp-Bench C# 17 150 GitHub HuggingFace
SWE-Perf Python 12 140 GitHub HuggingFace
Visual SWE-bench Python 11 133 GitHub HuggingFace
SWE-EVO Python 7 48 GitHub
Multi-PL Datasets
SWE-Mirror Python, Rust, Go 40 60k -
Multi-SWE-bench Java, JS, TS, Go, Rust, C, C++ 76 4,723 GitHub HuggingFace
Swing-Bench Python, Go, C++, Rust 400 2300 -
SWE-PolyBench Python, Java, JS, TS 21 2,110 GitHub HuggingFace HuggingFace
SWE-Compass Python, JS, TS, Java, C, C++, Go, Rust, Kotlin, C# - 2,000 GitHub HuggingFace
SWE-Bench Pro Python, Go, TS 41 1,865 GitHub HuggingFace
SWE-bench++ Python, Go, TS, JS, Ruby, PHP, Java, Rust, C++, C#, C 3,971 1,782 GitHub HuggingFace
SWE-Lancer JS, TS - 1,488 GitHub
OmniGIRL Python, TS, Java, JS 15 959 GitHub HuggingFace
SWE-bench Multimodal JS, TS, HTML, CSS 17 619 GitHub HuggingFace
SWE-fficiency Python, Cython 9 498 GitHub
SWE-Factory Python, Java, JS, TS 12 430 GitHub HuggingFace
SWE-bench-Live-MultiLang & Windows Python, JS, TS, C, C++, C#, Java, Go, Rust 238 418 GitHub HuggingFace HuggingFace
SWE-bench Multilingual C, C++, Go, Java, JS, TS, Rust, Python, Ruby, PHP 42 300 GitHub HuggingFace
SWE-InfraBench Python, TS - 100 -

Training Trajectory Datasets

A survey of trajectory datasets used for agent training or analysis. We list the programming language, number of source repositories, and total trajectories for each dataset.

Dataset Language Repos Amount Link
SWE-Fixer Python 856 69,752 GitHub HuggingFace
SWE-rebench Python 1,823 67,074 HuggingFace
R2E-Gym Python 10 3,321 GitHub HuggingFace
SWE-Synth Python 11 3,018 GitHub HuggingFace
SWE-Factory Python 10 2,809 GitHub HuggingFace
SWE-Gym Python 11 491 GitHub HuggingFace
SWE-Lego Python 3251 14.6k GitHub

SFT-based Methods

Overview of SFT-based methods for issue resolution. This table categorizes models by their base architecture and training scaffold (Sorted by Performance).

Model Name Base Model Size Arch. Training Scaffold Res.(%) Code Data Model
SWE-rebench-openhands-Qwen3-235B-A22B Qwen3-235B-A22B 235B-A22B MoE OpenHands 59.9 - HuggingFace HuggingFace
SWE-Lego-Qwen3-32B Qwen3-32B 32B Dense OpenHands 57.6 GitHub HuggingFace HuggingFace
SWE-rebench-openhands-Qwen3-30B-A3B Qwen3-30B-A3B 30B-A3B MoE OpenHands 49.7 - HuggingFace HuggingFace
Devstral Mistral Small 3 22B Dense OpenHands 46.8 - Website HuggingFace
Co-PatcheR Qwen2.5-Coder-14B 3×14B Dense PatchPilot-mini 46.0 GitHub - HuggingFace
SWE-Swiss-32B Qwen2.5-32B-Instruct 32B Dense Agentless 45.0 GitHub HuggingFace HuggingFace
SWE-Lego-Qwen3-8B Qwen3-8B 8B Dense OpenHands 44.4 GitHub HuggingFace HuggingFace
Lingma SWE-GPT Qwen2.5-72B-Instruct 72B Dense SWESynInfer 30.2 GitHub - -
SWE-Gym-Qwen-32B Qwen2.5-Coder-32B 32B Dense OpenHands, MoatlessTools 20.6 GitHub - HuggingFace
Lingma SWE-GPT Qwen2.5-Coder-7B 7B Dense SWESynInfer 18.2 GitHub - -
SWE-Gym-Qwen-14B Qwen2.5-Coder-14B 14B Dense OpenHands, MoatlessTools 16.4 GitHub - HuggingFace
SWE-Gym-Qwen-7B Qwen2.5-Coder-7B 7B Dense OpenHands, MoatlessTools 10.6 GitHub - HuggingFace

RL-based Methods

A comprehensive overview of specialized models for issue resolution, categorized by parameter size. The table details each model's base architecture, the training scaffold used for rollout, the type of reward signal employed (Outcome vs. Process), and their performance results (Res. %) on issue resolution benchmarks.

Model Name Base Model Size Arch. Train. Scaffold Reward Res.(%) Code Data Model
560B Models (MoE)
LongCat-Flash-Think LongCatFlash-Base 560B-A27B MoE R2E-Gym Outcome 60.4 GitHub - HuggingFace
72B Models
Kimi-Dev Qwen 2.5-72B-Base 72B Dense BugFixer + TestWriter Outcome 60.4 GitHub - HuggingFace
SWE-RL Llama-3.3-70B-Instruct 70B Dense Agentless-mini Outcome 41.0 GitHub - -
Multi-turn RL(Nebius) Qwen2.5-72B-Instruct 72B Dense SWE-agent Outcome 39.0 - - -
Agent-RLVR-RM-72B Qwen2.5-Coder-72B 72B Dense Localization + Repair Outcome 27.8 - - -
Agent-RLVR-72B Qwen2.5-Coder-72B 72B Dense Localization + Repair Outcome 22.4 - - -
32B Models
OpenHands Critic Qwen2.5-Coder-32B 32B Dense SWE-Gym - 66.4 GitHub - HuggingFace
KAT-Dev-32B Qwen3-32B 32B Dense - - 62.4 - - HuggingFace
SWE-Swiss-32B Qwen2.5-32B-Instruct 32B Dense - Outcome 60.2 GitHub HuggingFace HuggingFace
FoldAgent Seed-OSS-36B-Instruct 36B Dense FoldAgent Process 58.0 GitHub Website -
SeamlessFlow-32B Qwen3-32B 32B Dense SWE-agent Outcome 45.8 GitHub - -
DeepSWE Qwen3-32B 32B Dense R2E-Gym Outcome 42.2 GitHub HuggingFace HuggingFace
SA-SWE-32B - 32B Dense SkyRL-Agent - 39.4 - - -
OpenHands LM v0.1 Qwen2.5-Coder-32B 32B Dense SWE-Gym - 37.2 GitHub - HuggingFace
SWE-Dev-32B Qwen2.5-Coder-32B 32B Dense OpenHands Outcome 36.6 GitHub - HuggingFace
Satori-SWE Qwen2.5-Coder-32B 32B Dense Retriever + Code editor Outcome 35.8 GitHub HuggingFace HuggingFace
SoRFT-32B Qwen2.5-Coder-32B 32B Dense Agentless Outcome 30.8 - - -
Agent-RLVR-32B Qwen2.5-Coder-32B 32B Dense Localization + Repair Outcome 21.6 - - -
14B Models
Agent-RLVR-14B Qwen2.5-Coder-14B 14B Dense Localization + Repair Outcome 18.0 - - -
SEAlign-14B Qwen2.5-Coder-14B 14B Dense OpenHands Process 17.7 - - -
7-8B Models
SeamlessFlow-8B Qwen3-8B 8B Dense SWE-agent Outcome 27.4 GitHub - -
SWE-Dev-7B Qwen2.5-Coder-7B 7B Dense OpenHands Outcome 23.4 GitHub - HuggingFace
SoRFT-7B Qwen2.5-Coder-7B 7B Dense Agentless Outcome 21.4 - - -
SWE-Dev-8B Llama-3.1-8B 8B Dense OpenHands Outcome 18.0 GitHub - HuggingFace
SEAlign-7B Qwen2.5-Coder-7B 7B Dense OpenHands Process 15.0 - - -
SWE-Dev-9B GLM-4-9B 9B Dense OpenHands Outcome 13.6 GitHub - HuggingFace

General Foundation Models

Overview of general foundation models evaluated on issue resolution. The table details the specific inference scaffolds (e.g., OpenHands, Agentless) employed during the evaluation process to achieve the reported results.

Model Name Size Arch. Inf. Scaffold Reward Res.(%) Code Model
MiMo-V2-Flash 309B-A15B MoE Agentless Outcome 73.4 GitHub HuggingFace
KAT-Coder - - Claude Code Outcome 73.4 - Website
Deepseek V3.2 671B-A37B MoE Claude Code, RooCode - 73.1 GitHub HuggingFace
Kimi-K2-Instruct 1T MoE Agentless Outcome 71.6 - HuggingFace
Qwen3-Coder 480B-A35B MoE OpenHands Outcome 69.6 GitHub HuggingFace
GLM-4.6 355B-A32B MoE OpenHands Outcome 68.0 - HuggingFace
gpt-oss-120b 116.8B-A5.1B MoE Internal tool Outcome 62.0 GitHub HuggingFace
Minimax M2 230B-10B MoE R2E-Gym Outcome 61.0 GitHub HuggingFace
gpt-oss-20b 20.9B-A3.6B MoE Internal tool Outcome 60.0 GitHub HuggingFace
GLM-4.5-Air 106B-A12B MoE OpenHands Outcome 57.6 - -
Minimax M1-80k 456B-A45.9B MoE Agentless Outcome 56.0 GitHub Website
Minimax M1-40k 456B-A45.9B MoE Agentless Outcome 55.6 GitHub Website
Seed1.5-Thinking 200B-A20B MoE - Outcome 47.0 GitHub -
Llama 4 Maverick 400B-A17B MoE mini-SWE-agent Outcome 21.0 GitHub HuggingFace
Llama 4 Scout 109B-17B MoE mini-SWE-agent Outcome 9.1 GitHub HuggingFace


🚀 Quick Start

Windows:

run.bat

Linux/Mac:

chmod +x run.sh
./run.sh

Options:

  • [1] Add Paper - Interactive paper entry with duplicate check
  • [2] Add Table - Update statistical tables
  • [3] Batch Import - Import papers from CSV template
  • [4] Sync & Build - Render website and sync README

Manual operations:

# Local preview
mkdocs serve

# Deploy (or push to GitHub for auto-deploy via Actions)
mkdocs gh-deploy


🤝 Contributing

We welcome contributions! To add new papers or tables:

  1. Fork this repository
  2. Run run.bat (Windows) or run.sh (Linux/Mac)
  3. Or manually edit YAML/CSV files in data/ directory
  4. Submit a PR with your changes

🌟 Related Work

Code Generation

The application of LLMs in the programming domain has witnessed explosive growth. Early research focused primarily on function-level code generation, with benchmarks such as HumanEval serving as standard metrics. However, generic benchmarks often fail to capture the nuances of real-world development. To bridge this gap, recent initiatives have attempted to extend evaluation tasks to align more closely with realistic software development scenarios, revealing the limitations of general models in specialized domains. Concurrently, methods are also evolving to capture these broader contexts. While foundational approaches primarily relied on SFT or standard retrieval-augmented generation, RL-based methods emerged as a pivotal direction for handling complex coding tasks.

Related:

Automated Software Generation

The primary goal of this task is to autonomously construct complete and executable software systems starting from high-level natural language requirements. Unlike code completion, it necessitates covering the Software Development Life Cycle (SDLC), including requirement analysis, system design, coding, and testing. To address the complexity and potential logic inconsistencies in this process, state-of-the-art frameworks leverage multi-agent collaboration, simulating human development teams to decompose complex tasks into streamlined and verifiable workflows.

Related:

Automated Software Maintenance

Issue resolution is intrinsically linked to the broader domain of automated software maintenance. Methodologies established in this field are frequently encapsulated as callable tools to augment the capabilities of LLMs in software development tasks.

Related:

Automated Environment Setup

Recent initiatives focus on automating the configuration of runtime environments for entire repositories. This capability develops in parallel with data construction for issue resolution.

Related:

Related Surveys

Existing surveys primarily focus on code generation or other tasks within the software engineering domain. This paper bridges this gap by offering the first systematic survey dedicated to the entire spectrum of issue resolution, ranging from non-agent approaches to the latest agentic advancements.

Related:


📄 Citation

If you use this project or related survey in your research or system, please cite the following:

Li, Caihua, Guo, Lianghong, Wang, Yanlin, et al. (2026). Advances, Frontiers, and Future of Issue Resolution in Software Engineering: A Comprehensive Survey. TechRxiv. DOI: 10.36227/techrxiv.176779734.47868328/v2

BibTeX:

@article{li2026advances,
  title={Advances, Frontiers, and Future of Issue Resolution in Software Engineering: A Comprehensive Survey},
  author={Li, Caihua and Guo, Lianghong and Wang, Yanlin and Guo, Daya and Tao, Wei and Shan, Zhenyu and Liu, Mingwei and Chen, Jiachi and Song, Haoyu and Tang, Duyu and Zhang, Hongyu and Zheng, Zibin},
  journal={TechRxiv},
  year={2026},
  page={1375056},
  dor={10.36227/techrxiv.176779734.47868328/v2},
  publisher={IEEE}
}

Once published on arXiv or at a conference, please replace the entry with the official citation information (authors, DOI/arXiv ID, conference name, etc.).

🙏 Acknowledgements

We would like to express our sincere gratitude to:

  • The authors of cited papers who provided valuable feedback on how their work is presented in this survey, greatly improving its accuracy and comprehensiveness.

  • All contributors who have helped improve this project through issues, pull requests, and discussions.

  • The open-source community for developing the amazing tools and frameworks that made this project possible.

Special Thanks

  • @chao-peng (Dr. Chao Peng), ByteDance Software Engineering Lab, for providing valuable suggestions on the Challenges and Opportunities section of our survey.

  • @EuniAI/awesome-code-agents for providing an excellent reference on managing survey papers through documentation systems and inspiring our project structure.


📬 Contact

If you have any questions or suggestions, please contact us through:


📜 License

This project is licensed under the MIT License - see the LICENSE file for details.


⭐ Star this repository if you find it helpful!

Made with ❤️ by the DeepSoftwareAnalytics team

Documentation | Paper | Tables | About | Cite

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

No releases published

Packages

No packages published