Chain-of-Thought Prompt Optimization via Adversarial Learning
preprint
OA: closed
Abstract
Chain-of-Thought (CoT) prompting has shown strong reasoning capabilities in solving complex problems. To further enhance accuracy, various methods have been proposed to improve both the prompt design process and the quality of the prompts themselves. However, whether manually crafted or automatically generated, prompts often lack effective strategies for optimization and evaluation. To address this limitation, we introduce the Adversarial Chain-of-Thought (adv-CoT) framework, inspired by adversarial learning. Specifically, adv-CoT comprises a generator and a discriminator, both powered by large language models (LLMs). The framework iteratively optimizes an initial CoT prompt by the following processes: the generator produces responses based on the input and the CoT prompt, while the discriminator distinguishes between generated outputs and ground-truth answers. This adversarial process continuously refines the prompt, forming a minimax game that jointly enhances both the generator and the discriminator. adv-CoT enables automatic prompt optimization and output evaluation through the use of LLMs. Extensive experiments on 12 datasets demonstrate the effectiveness of our approach, consistently improving accuracy across commonsense, factual, symbolic, and arithmetic tasks.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00