Evaluating the impact of prompting styles on LLM accuracy for AIME math questions

(1) American High School

https://doi.org/10.59720/25-137
Cover photo for Evaluating the impact of prompting styles on LLM accuracy for AIME math questions

Researchers are beginning to use LLMs to solve math problems but their ability to perform multi-step reasoning remains unclear. In this study, we investigated whether various prompting styles could enhance the problem-solving abilities of Large Language Models (LLMs) for difficult math problems such as those from the American Invitational Mathematics Exam (AIME). We hypothesized that, of the five prompting styles we tested, providing detailed AIME solutions (previous AIME problems with step-by-step solutions would lead to the highest LLM accuracy on new AIME problems. To test this hypothesis, we used three LLMs (ChatGPT Free, ChatGPT Plus, and Gemini) as tools to examine the effects of five prompting styles: a baseline with the problem alone, a study guide, past AIME problems with the answer alone, past AIME problems with detailed solutions, and Olympiad problems with detailed solutions. Our dataset included four AIME exams (2024 AIME I, 2024 AIME II, 2025 AIME I, and 2025 AIME II), totaling 60 problems. Although performance varied, particularly for Gemini, the detailed solution style did not significantly improve accuracy over the baseline or other prompting styles. This suggests that, compared to other prompting styles, detailed AIME solutions do not significantly improve LLM performance on multi-step math problems. Our work provides one approach for testing LLM performance in math and shows the necessity of deeper approaches beyond prompting to improve LLM performance.

Download Full Article as PDF