Efficient Reasoning at the Edge: A Systematic Review of Small Reasoning Models, Test-Time Scaling, Quantization, and Energy-Aware AI Inference
DOI:
https://doi.org/10.51483/IJAIML.6.11s.2026.1583-1592Keywords:
Edge artificial intelligence; small reasoning models; test-time scaling; quantization; energy-aware inference; edge–cloud computing; hardware-aware optimization.Abstract
Edge devices are increasingly required to have reasoning capabilities and deployed under extreme constraints of memory, latency, energy, network and privacy. In this systematic review, we give an overview of the current research on small reasoning models, test-time scaling, model quantization, hardware-aware optimization and energy-aware edge–cloud inference. Literature was identified by using structured multi-database search and pre-defined eligibility criteria, quality-assessment and data-extraction criteria. The papers analyzed here show that it is possible to obtain valuable reasoning abilities even with much smaller architectures for compact transformers, knowledge distillation, parameter-efficient adaptation, and carefully selected training data. The difficult tasks' performance is further improved by adaptive test-time computation in those more reasoning tokens, more candidate paths and/or verification are provided for the difficult tasks as needed. Post training and quantization-aware methods help to lower bandwidth and memory usage, but could introduce errors to the arithmetic operations and may impact the calibration for multi-step reasoning/inference without any hardware-specific optimizations. Apart from the energy efficient deployment, coordinated voltage and frequency scaling, early exit, workload scheduling, model selection and selective cloud offloading are also important. The compromise that is required between accuracy, speed, memory usage and energy usage makes it impossible to achieve the best performance in all of the above for any one hardware configuration across all application classes in mobile, wearable, robotic, automotive and Internet-of-Things applications. To move forward, there is a need for standardized hardware-level benchmarking, powerful compressed reasoning, adaptation without compromising privacy, explainable decision mechanisms, special accelerators, and self-adaptive systems that can dynamically balance reasoning quality to latency, energy, thermal limits and carbon intensity during actual deployment.





