Interpretable Adaptive Sampling for LLM Test-Time Scaling

2026-08-04Artificial Intelligence

Artificial Intelligence
AI summary

The authors propose a smart way to decide how many times a language model should try to answer a question based on how hard the question seems and how confident the model is. Instead of using the same number of tries for every question, their method uses fewer tries for easy questions and more for hard ones. This approach helps save computing power and is easier to understand. They tested it on tasks like question-answering and math problems and found it works better than some common methods while using less effort.

test-time scalinglarge language modelsadaptive samplingprompt complexitymodel confidencefuzzy controllerquestion answeringmathematical reasoninginference efficiencycompute budget
Authors
Mobina Kashaniyan, Ali Jannesari
Abstract
Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that spend the same compute on easy and difficult prompts. These fixed budgets are also difficult to inspect because they do not explain why a given prompt receives a particular number of samples. We propose adaptive} test-time scaling with a lightweight fuzzy controller that maps interpretable signals, including estimated prompt complexity and model confidence, to a per-query sampling budget. The controller assigns fewer samples to easier or more confident prompts and more samples to harder or less certain prompts, making inference-time compute inspectable rather than fixed or opaque. We evaluate under a fair-alignment protocol with matched decoding settings and controlled answer selection, and compare against best-of-$N$, compute-aware scaling, and self-certainty-based baselines on question-answering and mathematical reasoning tasks. Across models and datasets, adaptive fuzzy control improves over several standard baselines and remains close to a selector-matched full-budget control while reducing the average number of samples. These findings suggest that interpretable adaptive sampling is a practical direction for more efficient test-time reasoning in large language models.