Efficient Reasoning at the Edge: A Systematic Review of Small Reasoning Models, Test-Time Scaling, Quantization, and Energy-Aware AI Inference

Authors

  • Dr. G. Jagan Naik
  • Dr. Burla Srinivas
  • Dr. A. Ramesh
  • Y Amrutha
  • Dr. Mohammad Shahbaz Khan
  • Vankudoth Ramesh

DOI:

https://doi.org/10.51483/IJAIML.6.11s.2026.1583-1592

Keywords:

Edge artificial intelligence; small reasoning models; test-time scaling; quantization; energy-aware inference; edge–cloud computing; hardware-aware optimization.

Abstract

Edge devices are increasingly required to have reasoning capabilities and deployed under extreme constraints of memory, latency, energy, network and privacy. In this systematic review, we give an overview of the current research on small reasoning models, test-time scaling, model quantization, hardware-aware optimization and energy-aware edge–cloud inference. Literature was identified by using structured multi-database search and pre-defined eligibility criteria, quality-assessment and data-extraction criteria. The papers analyzed here show that it is possible to obtain valuable reasoning abilities even with much smaller architectures for compact transformers, knowledge distillation, parameter-efficient adaptation, and carefully selected training data. The difficult tasks' performance is further improved by adaptive test-time computation in those more reasoning tokens, more candidate paths and/or verification are provided for the difficult tasks as needed. Post training and quantization-aware methods help to lower bandwidth and memory usage, but could introduce errors to the arithmetic operations and may impact the calibration for multi-step reasoning/inference without any hardware-specific optimizations. Apart from the energy efficient deployment, coordinated voltage and frequency scaling, early exit, workload scheduling, model selection and selective cloud offloading are also important. The compromise that is required between accuracy, speed, memory usage and energy usage makes it impossible to achieve the best performance in all of the above for any one hardware configuration across all application classes in mobile, wearable, robotic, automotive and Internet-of-Things applications. To move forward, there is a need for standardized hardware-level benchmarking, powerful compressed reasoning, adaptation without compromising privacy, explainable decision mechanisms, special accelerators, and self-adaptive systems that can dynamically balance reasoning quality to latency, energy, thermal limits and carbon intensity during actual deployment.

Downloads

Published

2026-09-22

How to Cite

Naik, D. G. J., Srinivas, D. B., Ramesh, D. A., Amrutha, Y., Khan, D. M. S., & Ramesh, V. (2026). Efficient Reasoning at the Edge: A Systematic Review of Small Reasoning Models, Test-Time Scaling, Quantization, and Energy-Aware AI Inference. International Journal of Artificial Intelligence and Machine Learning, 6(11s), 1583–1592. https://doi.org/10.51483/IJAIML.6.11s.2026.1583-1592