Submitted:
13 June 2026
Posted:
15 June 2026
You are already at the latest version
Abstract
Tiny Machine Learning (TinyML) has emerged as a significant advancement in embedded Artificial Intelligence (AI), enabling machine learning inference directly on resource-constrained microcontrollers and ultra-low-power edge devices. By integrating lightweight machine learning models with embedded systems, TinyML facilitates real-time, energy-efficient, and privacy-preserving intelligence at the edge of Internet of Things (IoT) ecosystems. This chapter presents a comprehensive introduction to TinyML, examining its evolution from conventional cloud-centric AI and Edge AI architectures toward distributed embedded intelligence. The chapter discusses the fundamental architecture of TinyML systems, key model optimization and deployment techniques, including quantization, pruning, and model compression, as well as hardware-aware design strategies for efficient on-device inference. Furthermore, major application domains such as healthcare, consumer electronics, industrial automation, agriculture, and environmental monitoring are explored to demonstrate the practical relevance of TinyML across diverse sectors. The chapter also evaluates the principal advantages and limitations of TinyML and outlines a practical development workflow encompassing hardware selection, software frameworks, data acquisition, model training, optimization, and deployment. Overall, TinyML represents a critical enabling technology for scalable, low-power, and autonomous intelligent systems, supporting the next generation of edge computing and IoT applications.
Keywords:
TinyML
; Edge AI
; Internet of Things (IoT)
; embedded systems
; low-power machine learning
; microcontrollers
I. Introduction
Artificial Intelligence (AI) and Machine Learning (ML) have become fundamental technologies driving innovation across a wide range of domains, including healthcare, industrial automation, smart agriculture, consumer electronics, and wearable systems. Traditionally, most AI applications have relied on cloud computing infrastructures, where data generated by end devices are transmitted to powerful remote servers for processing, analysis, and decision-making. These cloud-based platforms leverage high-performance computing resources such as Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs) to train and execute complex machine learning and deep learning models [1]. While cloud AI provides significant computational power and high predictive accuracy, it also introduces several challenges, including communication latency, continuous internet dependency, high bandwidth consumption, increased energy usage, and concerns regarding data privacy and security. As reported by Ray [2], the continuous transmission of large volumes of sensor data from edge devices to remote cloud servers can significantly reduce the efficiency of latency-sensitive applications such as healthcare monitoring, autonomous systems, and industrial control.
The rapid expansion of the Internet of Things (IoT) has further intensified these challenges. Billions of connected sensors and embedded devices continuously generate data at the network edge, often operating under strict constraints in terms of memory, computational capability, and energy consumption. In many real-world scenarios, transmitting all collected data to centralized cloud infrastructures is neither efficient nor practical [3]. To address these limitations, the concept of Edge AI emerged, enabling data processing and machine learning inference to be performed closer to the data source. By relocating computation from the cloud to edge devices and gateways, Edge AI reduces latency, minimizes network traffic, and improves system responsiveness. However, many Edge AI solutions still depend on relatively powerful processors, embedded GPUs, or specialized accelerators, making them unsuitable for highly constrained low-power devices. Capogrosso et al. [4] highlighted that numerous edge computing platforms require substantial hardware resources and energy consumption, limiting their deployment in ultra-low-power embedded environments.
The need for intelligent processing on resource-constrained devices led to the emergence of Tiny Machine Learning (TinyML), a rapidly growing field that enables machine learning inference directly on microcontrollers and low-power embedded systems. TinyML integrates machine learning algorithms, embedded computing, and energy-efficient hardware to bring intelligence to devices operating within milliwatt-level power budgets. According to Soro [3], TinyML represents the convergence of embedded systems and machine learning, allowing intelligent decision-making to occur locally on the device without continuous reliance on cloud connectivity. By performing inference directly at the sensing layer, TinyML enables faster response times, lower communication costs, enhanced privacy protection, and significantly reduced energy consumption.
Unlike conventional AI systems, TinyML devices typically operate with only a few kilobytes of RAM and limited flash memory. Consequently, traditional deep learning models cannot be deployed directly on such platforms without substantial optimization. To address this challenge, TinyML employs a variety of model optimization techniques, including quantization, pruning, and model compression. Quantization reduces numerical precision to decrease memory requirements and computational complexity, while pruning removes redundant network parameters to create compact and efficient models. These techniques enable machine learning models to operate within the stringent resource limitations of embedded hardware while maintaining acceptable inference accuracy [4,5].
The growth of TinyML has been accelerated by significant advancements in both hardware platforms and software ecosystems. Modern microcontroller-based development boards such as the Arduino Nano 33 BLE Sense, ESP32, STM32, and Raspberry Pi Pico provide affordable and energy-efficient platforms for deploying TinyML applications. In parallel, software frameworks such as TensorFlow Lite for Microcontrollers (TFLM), Edge Impulse, and ARM-NN have simplified the development, optimization, and deployment of machine learning models on embedded devices. These tools allow developers to train models on high-performance computing systems and subsequently deploy optimized versions onto resource-constrained hardware. Ray [2] emphasized that such frameworks play a crucial role in bridging the gap between machine learning development environments and embedded deployment platforms.
TinyML is increasingly finding applications across numerous domains. In healthcare, TinyML-powered wearable devices enable real-time monitoring of physiological parameters such as heart rate, body temperature, and blood oxygen saturation. In agriculture, intelligent sensor nodes can monitor soil moisture, crop health, and environmental conditions while operating autonomously for extended periods. Industrial systems utilize TinyML for predictive maintenance, fault detection, and anomaly identification, whereas smart home devices employ TinyML for speech recognition, keyword spotting, gesture recognition, and intelligent automation. As noted by Schizas et al. [6], TinyML enables embedded devices to perform intelligent tasks locally, reducing dependence on cloud infrastructures while improving privacy, reliability, and operational efficiency.
Despite its considerable advantages, TinyML continues to face several challenges. Severe memory and computational constraints, hardware heterogeneity, deployment complexity, and the difficulty of maintaining high model accuracy under resource limitations remain active research issues. Furthermore, the absence of universally accepted benchmarking methodologies and standardized deployment frameworks continues to hinder large-scale adoption. Nevertheless, ongoing advancements in lightweight neural network architectures, model optimization techniques, and embedded hardware technologies are rapidly expanding the capabilities of TinyML systems. By enabling intelligent, low-power, real-time, and privacy-preserving inference directly on embedded devices, TinyML represents a major step toward the realization of distributed, autonomous, and scalable edge intelligence for next-generation IoT ecosystems.
II. Evolution of Tinyml from Edge AI
Artificial Intelligence (AI) systems were originally developed around centralized cloud computing infrastructures equipped with high-performance computational resources, including Graphics Processing Units (GPUs), Tensor Processing Units (TPUs), and large-scale data centers. These infrastructures enabled the execution of computationally intensive Machine Learning (ML) and Deep Learning (DL) algorithms capable of processing vast amounts of data with high predictive accuracy. In traditional Internet of Things (IoT) architectures, sensors and embedded devices primarily functioned as data collection units, continuously transmitting sensory information to remote cloud servers where data analysis, model execution, and decision-making processes were performed. This cloud-centric approach offered significant computational capability, centralized resource management, and scalability for large-scale intelligent applications.
Despite these advantages, the rapid growth of IoT ecosystems exposed several limitations of cloud-based intelligence. As billions of connected devices began generating continuous streams of data, the need to transmit information to remote servers introduced communication latency, increased bandwidth consumption, higher energy requirements, and dependence on reliable network connectivity. These limitations became particularly problematic in real-time applications where immediate response and continuous operation are critical [2]. Consequently, cloud-centric AI architectures proved insufficient for many latency-sensitive domains, including healthcare monitoring, industrial automation, autonomous systems, and intelligent surveillance.
The shortcomings of cloud-based processing became even more apparent in applications requiring near real-time decision-making. For example, autonomous vehicles, industrial control systems, and wearable healthcare devices often cannot tolerate delays associated with transmitting data to distant cloud servers and waiting for inference results. Furthermore, the continuous transmission of sensitive information raises significant concerns regarding data privacy, cybersecurity, and system reliability. Centralized processing architectures may increase exposure to network vulnerabilities and unauthorized access, particularly in applications involving personal health information and critical industrial data [5]. These challenges motivated researchers and engineers to explore distributed intelligence architectures capable of relocating computational tasks closer to the source of data generation.
This shift led to the emergence of Edge Artificial Intelligence (Edge AI), a paradigm in which data processing and machine learning inference are performed closer to the network edge using embedded processors, edge gateways, and local computing nodes. By executing inference within the operational environment rather than relying exclusively on remote cloud infrastructures, Edge AI significantly reduces communication latency, minimizes bandwidth requirements, improves system reliability, and enables faster response generation. In addition, Edge AI allows systems to maintain partial functionality even under limited or intermittent network connectivity, thereby enhancing operational resilience in distributed IoT environments [6].
Although Edge AI represented a significant advancement over cloud-centric architectures, many Edge AI solutions continued to rely on relatively powerful hardware platforms such as embedded GPUs, AI accelerators, multicore processors, and high-memory edge computing devices. While these platforms provide substantial computational capability, they often consume considerable power and require hardware resources that exceed the capabilities of low-cost embedded systems [4]. Consequently, many Edge AI approaches remained unsuitable for ultra-low-power devices operating under strict memory, computational, and energy constraints.
The increasing demand for intelligent processing directly on resource-constrained embedded devices ultimately led to the emergence of Tiny Machine Learning (TinyML). TinyML extends the principles of Edge AI by enabling machine learning inference directly on microcontrollers and ultra-low-power embedded systems operating at the extreme edge of the network. Rather than performing inference on edge servers or gateways, TinyML embeds intelligence directly within sensing devices themselves, allowing data to be processed where it is generated. This approach significantly reduces communication overhead, lowers energy consumption, enhances privacy preservation, and enables real-time decision-making without continuous cloud connectivity [3].
From a broader perspective, the evolution from Cloud AI to Edge AI and ultimately TinyML represents a progressive migration of intelligence from centralized infrastructures toward deeply embedded autonomous systems. Cloud AI centralized computation within large-scale data centers, Edge AI distributed intelligence across localized computing environments, and TinyML embedded intelligence directly into sensing devices. This evolution has transformed IoT systems from passive data acquisition networks into intelligent, autonomous, energy-efficient, and privacy-aware ecosystems capable of making decisions at the point of data generation.
Figure 1.
Architectures of Cloud AI to Edge AI and TinyML.

The increasing demand for intelligent processing within ultra-low-power Internet of Things (IoT) environments ultimately led to the emergence of Tiny Machine Learning (TinyML). TinyML extends the principles of Edge AI by enabling machine learning inference directly on microcontroller-based systems operating at the extreme edge of the network. Unlike conventional Edge AI architectures that often rely on gateways, edge servers, or dedicated processing units, TinyML embeds intelligence directly within sensing devices themselves. This approach enables devices to analyze data locally, generate predictions in real time, and respond autonomously without continuous communication with external computing infrastructures. As a result, TinyML significantly reduces communication overhead, lowers latency, improves privacy, and enables energy-efficient operation within highly constrained environments [3].
A distinguishing characteristic of TinyML is its ability to execute machine learning models on devices with extremely limited computational resources. Typical TinyML platforms operate with only a few kilobytes of RAM, limited flash storage, and strict milliwatt-level power budgets. Despite these constraints, TinyML systems are capable of performing intelligent tasks such as classification, anomaly detection, speech recognition, and sensor data analysis in real time. This capability has enabled the integration of artificial intelligence into a wide range of embedded applications where conventional Edge AI solutions may be impractical due to power, cost, or hardware limitations [3,6].
The evolution from Cloud AI to TinyML also reflects a progressive shift toward greater computational efficiency and hardware-aware intelligence. Traditional machine learning and deep learning models were primarily designed to maximize predictive performance and computational scalability, often without considering the strict memory, processing, and energy constraints of embedded systems. Consequently, deploying such models on resource-constrained hardware requires extensive optimization and adaptation. To address these challenges, TinyML employs lightweight inference techniques such as quantization, pruning, model compression, knowledge distillation, and memory-efficient neural network architectures. These optimization strategies reduce model size, computational complexity, and energy consumption while maintaining acceptable levels of predictive accuracy [4,5].
In recent years, TinyML research has increasingly focused on hardware-software co-design approaches, in which machine learning models and embedded hardware are optimized simultaneously to maximize overall system efficiency. Rather than treating hardware and software as independent components, this approach considers the capabilities and limitations of the target platform during model design and optimization. Such integration enables more efficient utilization of memory, processing resources, and energy budgets, thereby improving the practicality of deploying machine learning algorithms on highly constrained embedded devices [4].
Figure 2.
Paramitarized diffenre of CloudML, EdgeML and TinyML.

From a conceptual perspective, the progression from Cloud AI to Edge AI and ultimately TinyML represents the gradual migration of intelligence from centralized infrastructures toward deeply embedded autonomous systems. Cloud AI centralized computation within large-scale data centers, Edge AI distributed intelligence across localized computing environments, and TinyML embedded intelligence directly into sensing devices operating at the point of data generation. This evolution has transformed IoT systems from passive data collection networks into intelligent, autonomous, real-time, and energy-efficient ecosystems capable of making decisions locally while minimizing dependence on external computational resources. Consequently, TinyML is increasingly recognized as a foundational technology for the next generation of scalable, privacy-preserving, and autonomous intelligent systems.
III. TinyML Architecture
Tiny Machine Learning (TinyML) is a machine learning paradigm that enables intelligent data processing and inference directly on resource-constrained embedded devices such as microcontrollers, sensors, and low-power edge nodes. Unlike traditional AI systems that rely on cloud infrastructures for computation, TinyML performs inference locally on the device, reducing latency, bandwidth usage, and energy consumption while improving privacy and real-time responsiveness [2,3].
The architecture of a TinyML system follows a layered approach in which data are collected from sensors, processed locally, analyzed using optimized machine learning models, and transformed into actionable decisions. The major layers of the TinyML architecture are described below.
A. Sensor Layer
The Sensor Layer serves as the entry point of the TinyML system. It consists of sensing devices such as microphones, cameras, accelerometers, gyroscopes, and environmental sensors that collect data from the surrounding environment and convert it into digital signals for further processing [6]. The quality of sensor data directly affects the accuracy and reliability of the deployed TinyML model.
B. Data Acquisition Layer
The Data Acquisition Layer collects and organizes raw sensor data before processing. It manages tasks such as sensor sampling, buffering, and data transfer while ensuring efficient operation under strict energy and memory constraints. Efficient data acquisition is essential for maintaining system performance and low power consumption.
C. Data Preprocessing Layer
Raw sensor data often contain noise and redundant information. Therefore, preprocessing techniques such as filtering, normalization, and feature extraction are applied to prepare the data for inference. Effective preprocessing reduces computational complexity and improves model efficiency while preserving important information required for accurate predictions [2,5].
D. TinyML Model Layer
The TinyML Model Layer represents the intelligence core of the architecture. Machine learning models such as ANNs, CNNs, and RNNs are typically trained on powerful computing systems and then optimized using techniques such as quantization, pruning, and model compression before deployment [3,4]. These optimizations reduce memory and computational requirements while maintaining acceptable accuracy.
E. Inference Engine Layer
The Inference Engine Layer executes the optimized model on the target hardware and generates predictions from incoming data. Frameworks such as TensorFlow Lite for Microcontrollers (TFLM), CMSIS-NN, and ARM-NN provide lightweight runtime environments for efficient model execution on embedded systems [3,4]. This enables real-time inference without continuous cloud connectivity.
F. Decision and Actuation Layer
The Decision and Actuation Layer converts model predictions into practical actions. Depending on the application, the system may trigger an alarm, activate an actuator, classify an event, recognize a gesture, or detect anomalies. This layer enables autonomous system operation and real-time response to environmental changes [5].
G. Cloud Communication Layer (Optional)
Although TinyML primarily operates on-device, cloud communication may be used for remote monitoring, data storage, or system updates. Instead of transmitting raw sensor data, only essential information such as predictions, alerts, or summarized results is typically sent to the cloud, reducing bandwidth usage while preserving privacy and security [5,6].
Figure 3.
Illustration of TinyML Architecture.

The TinyML workflow begins with sensors collecting data from the environment. The acquired data are then preprocessed and passed to an optimized machine learning model deployed on a microcontroller. The inference engine executes the model and generates predictions in real time. Finally, the decision layer converts these predictions into actions, while optional cloud communication may be used for monitoring and management. This layered architecture enables intelligent, low-power, and real-time machine learning directly on embedded devices [2,3].
IV. Model Optimization and Deployment Pipeline
TinyML enables machine learning models to run efficiently on resource-constrained devices such as microcontrollers, sensors, and edge nodes. However, these devices have limited memory, storage, processing power, and energy availability. Therefore, machine learning models must be carefully optimized before deployment. TinyML achieves this through a series of optimization techniques that reduce model size and computational requirements while maintaining acceptable prediction accuracy [2,3,4,5].
A. Model Training and Baseline Model Development
The TinyML workflow begins with training a machine learning model on powerful computing platforms such as cloud servers, workstations, or GPUs. At this stage, the primary goal is to achieve high prediction accuracy without considering hardware limitations. Commonly used models include Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Long Short-Term Memory (LSTM) networks.
However, these models are often too large for direct deployment on microcontrollers. Therefore, additional optimization steps are required to transform them into lightweight TinyML models.
Example: A CNN model for image classification may occupy 20 MB of memory, which exceeds the storage capacity of most microcontrollers.
B. Model Compression
Model compression reduces the size of a trained model by removing redundant information and unnecessary parameters. This process decreases storage requirements and computational complexity while preserving most of the model’s predictive performance [5]. The objective of model compression is to preserve model accuracy while reducing the number of parameters, memory requirements, and computational operations.
Example:
Original Model Size = 10 MB
Compressed Model Size = 1 MB
Benefits: Reduced storage requirements, Lower memory consumption, Faster model execution, Improved deployment feasibility
C. Quantization
Quantization further reduces model size by converting high-precision parameters into lower-precision representations. While traditional neural networks commonly use 32-bit floating-point (FP32) values, TinyML models often use 8-bit integer (INT8) representations [2,3]. This significantly reduces memory usage and improves inference speed with only a minor impact on accuracy.
Example:
Before Quantization: Weight = 0.874532 (FP32),
After Quantization: Weight = 0.87 (INT8 Representation)
Benefits: Up to 75% reduction in memory usage, Faster arithmetic computations, Lower energy consumption, Improved inference efficiency
D. Pruning
Pruning is a structural optimization technique that removes unnecessary neurons, connections, or parameters from a neural network. Research has shown that many network weights contribute minimally to the final prediction and can therefore be removed without significantly affecting model performance. It pruning creates sparse neural networks by eliminating low-importance parameters. Further emphasize that pruning reduces both computational complexity and memory utilization, making models more suitable for TinyML deployment [2,4].
Example: A TinyML model contains 1,000 parameters. During pruning, 200 low-importance parameters are removed because they contribute very little to the prediction. The optimized model now contains only 800 parameters, reducing memory usage and computational cost while preserving most of its accuracy.
Benefits: Reduced parameter count, Lower computational complexity smaller memory footprint, Faster execution speed
E. Knowledge Distillation
Knowledge Distillation is a model transfer technique in which a large and highly accurate model, referred to as the Teacher Model, transfers its learned knowledge to a smaller Student Model. The student model learns to mimic the predictions of the teacher while maintaining a significantly smaller architecture.
Knowledge distillation enables lightweight TinyML models to achieve performance levels comparable to larger deep learning models while requiring substantially fewer resources [3].
Example: Teacher Model → Knowledge Transfer → Student Model
Benefits: High predictive accuracy, Reduced model complexity, Efficient deployment on embedded devices
F. Efficient Neural Network Architectures
TinyML also relies on lightweight neural network architectures specifically designed for embedded systems. Examples include MobileNet, SqueezeNet, TinyCNN, and EfficientNet-Lite [6]. These architectures use fewer parameters and computational operations while maintaining competitive accuracy.
Example:
Traditional CNN → Millions of Parameters
MobileNet → Thousands of Parameters
Benefits: Reduced computational requirements, Faster inference execution, Improved energy efficiency, better suitability for microcontrollers
G. Hardware-Aware Optimization
After the model architecture has been optimized, hardware-aware optimization is performed. This stage adapts the model according to the characteristics of the target hardware platform, including processor architecture, memory capacity, energy budget, and available accelerators.
This process as hardware-software co-design, where model design and hardware capabilities are jointly optimized to maximize efficiency. Hardware-aware optimization ensures that the model fully utilizes the available computational resources while minimizing power consumption [4].
Example:
Model A → Optimized for ARM Cortex-M4
Model B → Optimized for Edge TPU
Benefits: Improved hardware utilization, Lower power consumption, Enhanced execution efficiency, better platform compatibility
H. On-Device Inference
The final stage of the TinyML pipeline is on-device inference. After all optimization procedures have been completed, the model is deployed to the target embedded device using lightweight frameworks such as TensorFlow Lite for Microcontrollers (TFLM), ARM-NN, CMSIS-NN, or Edge Impulse. At this stage, sensor data are processed locally, predictions are generated in real time, and actions are performed without continuous cloud connectivity. The on-device inference is the defining characteristic of TinyML, enabling ultra-low-latency, privacy-preserving, and energy-efficient artificial intelligence at the edge [3,6].
Example: Sensor Data → TinyML Model → Prediction → Action
Benefits: Real-time decision making, Reduced network bandwidth usage, Enhanced privacy and security, Offline operation capability
Figure 4.
Illustration of TinyML Core deployment pipeline.

The TinyML optimization pipeline begins with model training on high-performance computing systems. The trained model is then compressed, quantized, pruned, and optimized for the target hardware platform. Finally, the optimized model is deployed on a microcontroller where it performs local inference and generates real-time decisions. This process enables intelligent, low-power, and efficient machine learning directly on embedded devices [2,3,4,5,6].
V. Applications of TinyML
TinyML has emerged as a transformative technology that enables machine learning inference directly on resource-constrained embedded devices such as microcontrollers, sensors, and edge nodes. By performing data processing locally, TinyML reduces latency, minimizes communication overhead, lowers energy consumption, and enhances data privacy. These advantages make TinyML particularly suitable for Internet of Things (IoT) applications, where devices often operate under limited computational resources and intermittent network connectivity [2,3,6].
The growing availability of low-power hardware platforms and lightweight machine learning frameworks has accelerated the adoption of TinyML across a wide range of application domains. Unlike traditional cloud-based AI systems that depend heavily on remote servers for computation, TinyML enables intelligent decision-making directly on the device, allowing real-time responses and autonomous operation in resource-constrained environments [4,5].
As a result, TinyML is increasingly being integrated into various sectors, including healthcare, consumer electronics, industrial automation, agriculture, and environmental monitoring. The following sections discuss how TinyML is applied in these domains and highlight its role in enabling efficient, low-power, and intelligent embedded systems.
Figure 5.
TinyML Application Domains.

A. Healthcare Application
In healthcare systems, TinyML is widely used for wearable health monitoring, respiratory analysis, activity recognition, and patient diagnostics. Embedded TinyML models can continuously monitor physiological signals such as heart rate, body temperature, motion patterns, and oxygen saturation directly on wearable devices without continuously transmitting sensitive data to external cloud servers. This localized inference improves privacy, reduces latency, and enables real-time anomaly detection. According to Elhanashi et al., TinyML-based wearable systems are increasingly being used for health monitoring and activity recognition applications in resource-constrained healthcare environments.
B. Consumer Application
In consumer electronics and smart home systems, TinyML enables intelligent voice recognition, keyword spotting, gesture recognition, occupancy detection, and personalized automation directly on embedded devices. Smart speakers, home automation systems, and portable consumer devices increasingly use TinyML to provide always-on sensing capability while maintaining ultra-low-power operation. By processing voice and sensor data locally, TinyML significantly reduces cloud dependency and improves user privacy.
C. Industrial Application
Industrial IoT and predictive maintenance represent another major application domain of TinyML. Modern industries require continuous monitoring of machinery, motors, pumps, and manufacturing systems to detect operational abnormalities and reduce maintenance costs. TinyML-based industrial sensors can locally analyze vibration signals, acoustic patterns, and thermal conditions to detect anomalies directly at the machine level. This enables predictive maintenance systems to identify faults before critical failures occur while minimizing communication overhead and cloud processing requirements. Elhanashi et al. emphasized that TinyML significantly enhances industrial automation through localized anomaly detection and predictive maintenance capabilities.
D. Agricultural Application
TinyML is also transforming agriculture and environmental monitoring systems by enabling intelligent sensing in remote and resource-constrained environments. Smart agriculture systems use TinyML for crop monitoring, precision irrigation, livestock tracking, and soil condition analysis. Similarly, environmental monitoring systems use TinyML for air quality monitoring, weather analysis, forest fire detection, and wildlife monitoring. Since TinyML systems consume very low power, they are highly suitable for long-term autonomous deployments in remote outdoor environments where continuous cloud connectivity may not always be available.
Table 1.
Comparative Table as per Domain.
| Application Domain | Cloud AI Approach | Edge AI Approach | TinyML Approach |
|---|---|---|---|
| Healthcare Monitoring | Physiological data continuously transmitted to cloud servers for processing and diagnosis | Local edge gateway performs preliminary health analytics before cloud synchronization | Wearable microcontroller directly performs real-time health inference and anomaly detection |
| Consumer Electronics | Voice and sensor data sent to cloud assistants for processing | Smart home hub locally processes user commands and environmental data | Embedded TinyML model performs keyword spotting and local voice recognition directly on-device |
| Industrial IoT | Industrial sensor data continuously uploaded to centralized cloud platforms for fault analysis | Edge servers perform localized predictive maintenance and anomaly analytics | Embedded industrial sensors perform direct vibration and fault inference at the sensing layer |
The evolution from Cloud AI to Edge AI and ultimately TinyML has fundamentally changed how intelligent IoT applications are designed and deployed. Earlier cloud-centric systems relied entirely on remote servers for data processing, while Edge AI introduced localized gateway-based analytics. TinyML further advances this evolution by enabling complete inference directly on embedded microcontrollers and sensing devices themselves.
The widespread adoption of TinyML demonstrates the transition of embedded systems from passive sensing platforms to intelligent autonomous systems capable of localized real-time decision-making. By integrating lightweight Machine Learning models directly into ultra-low-power embedded hardware, TinyML enables faster response, improved privacy, reduced communication overhead, and greater operational autonomy across modern IoT ecosystems. As highlighted by Elhanashi et al., TinyML is expected to play a critical role in the future development of intelligent edge computing systems, enabling scalable, efficient, and distributed AI-driven infrastructures for next-generation applications.
VI. Advantages and Limitations of TinyML
TinyML has emerged as an important advancement in embedded Artificial Intelligence by enabling lightweight Machine Learning models to execute directly on low-power microcontrollers and resource-constrained embedded devices. Unlike conventional cloud-based AI systems that rely heavily on remote servers for computation and inference, TinyML performs localized processing directly at the sensing layer. This enables real-time intelligence, reduced communication overhead, improved privacy, and energy-efficient operation within IoT and embedded environments. Due to these capabilities, TinyML is increasingly being adopted across healthcare, industrial automation, consumer electronics, smart agriculture, and environmental monitoring systems. However, despite its growing importance, TinyML also faces several limitations associated with memory constraints, computational capability, and deployment complexity. According to Tsoukas et al., TinyML systems must continuously balance inference accuracy, energy efficiency, and hardware limitations to achieve reliable embedded intelligence.
A. Advantages of TinyML
One of the major advantages of TinyML is its ability to perform real-time inference directly on embedded devices without depending on continuous cloud communication. Since data processing occurs locally, TinyML significantly reduces latency and enables faster response generation in applications such as healthcare monitoring, industrial automation, and smart wearables. Ray highlighted that localized embedded inference improves operational responsiveness while minimizing network dependency in IoT systems.
TinyML also enables ultra-low-power operation because lightweight models are optimized for constrained embedded hardware operating under milliwatt-level power budgets. This makes TinyML highly suitable for battery-powered wearable devices, remote sensors, and long-term IoT deployments. Furthermore, TinyML improves privacy preservation by processing sensitive sensor data locally instead of continuously transmitting information to cloud servers. Offline functionality is another important advantage, as TinyML systems can continue operating even in environments with unstable or unavailable internet connectivity.
In addition, TinyML supports scalable distributed intelligence by enabling intelligent processing directly across multiple low-cost embedded sensing devices. According to Schizas et al., TinyML significantly improves energy efficiency and reduces communication overhead in large-scale IoT ecosystems.
B. Limitations of TinyML
Despite its advantages, TinyML faces several technical limitations primarily caused by severe hardware constraints. Most microcontrollers contain limited RAM, flash memory, and processing capability, restricting the size and complexity of deployable machine learning models. Large deep learning architectures therefore require aggressive optimization techniques such as quantization, pruning, and model compression before deployment.
Limited computational capability is another major challenge because embedded devices generally operate at low clock frequencies and lack dedicated AI acceleration hardware. As a result, complex inference tasks such as high-resolution image analysis and large-scale neural processing may exceed the capability of many TinyML platforms. Capogrosso et al. observed that maintaining inference accuracy while reducing computational complexity remains one of the most challenging aspects of TinyML system design.
TinyML systems also face deployment and portability challenges due to hardware heterogeneity across embedded platforms. Differences in processor architecture, memory organization, and runtime environments often require hardware-specific optimization and deployment strategies. Additionally, environmental variability and sensor noise may reduce model generalization performance in real-world applications.
Table 2.
Advantages vs Disadvantages.
| Advantages | Disadvantages |
|---|---|
| Real-time on-device inference | Limited memory and storage |
| Ultra-low-power operation | Restricted computational capability |
| Reduced latency and bandwidth usage | Reduced model complexity |
| Enhanced privacy preservation | Complex optimization requirements |
| Offline functionality | Hardware heterogeneity |
| Scalable distributed intelligence | Deployment challenges |
Overall, TinyML represents a significant advancement in embedded AI by enabling localized, energy-efficient, and real-time intelligent processing directly on constrained devices. Although several technical challenges remain, ongoing improvements in lightweight neural architectures, model optimization techniques, and embedded hardware platforms continue to expand the capabilities and practical adoption of TinyML systems across modern IoT ecosystems.
VII. How to Get Started with TinyML
Before beginning TinyML development, it is important to understand how TinyML differs from traditional machine learning systems. Conventional AI models are typically designed for cloud environments with abundant computational resources, large memory capacity, and continuous power availability. In contrast, TinyML targets resource-constrained embedded devices such as microcontrollers, where memory, processing power, and energy resources are limited. Consequently, TinyML models must be lightweight, efficient, and carefully optimized for embedded execution. This enables intelligent inference directly on devices, providing benefits such as low latency, reduced power consumption, enhanced privacy, and offline operation [1,2,3].
A. Selecting TinyML Hardware Platforms
The first step in TinyML development is selecting a suitable hardware platform. Most TinyML applications are built on microcontroller-based development boards equipped with sensors and communication interfaces. Popular platforms include Arduino Nano 33 BLE Sense, ESP32, Raspberry Pi Pico, STM32, and Seeed Studio Wio Terminal [1,8].
When selecting hardware, factors such as memory capacity, processor performance, power consumption, sensor availability, and communication support should be considered. Development boards with built-in sensors are often preferred for prototyping because they simplify data collection and experimentation.
B. Setting Up the TinyML Software Environment
After selecting the hardware, the next step is preparing the software environment. TinyML development combines machine learning frameworks with embedded programming tools. Python is commonly used for data preparation, feature extraction, model training, and optimization.
For deployment, TensorFlow Lite for Microcontrollers (TFLM) is one of the most widely adopted frameworks because it enables machine learning models to execute efficiently on microcontrollers [1,7]. Development environments such as Arduino IDE and PlatformIO are commonly used for firmware compilation and deployment. Platforms such as Edge Impulse further simplify the workflow by integrating data acquisition, model training, optimization, and deployment within a unified environment.
C. Data Acquisition and Pre-Processing
Data collection is a critical stage in TinyML development because model performance depends heavily on the quality of the training data. Unlike many traditional machine learning applications that rely on publicly available datasets, TinyML systems often require data to be collected directly from the target sensors.
The collected data may include audio signals, motion patterns, vibration measurements, temperature readings, or environmental parameters. Since raw sensor data often contain noise and redundant information, pre-processing techniques such as filtering, normalization, and feature extraction are applied before model training. These operations improve model accuracy while reducing computational complexity and memory requirements [2,5].
D. Model Training and Optimization
After pre-processing, the extracted features are used to train a machine learning model. Depending on the application, lightweight neural networks such as compact CNNs, dense neural networks, or sequence-based models may be employed.
Once the desired accuracy is achieved, optimization techniques such as quantization, pruning, and model compression are applied to reduce memory usage and computational overhead [4,5]. Quantization is particularly important because it converts floating-point parameters into lower-precision formats such as INT8, significantly reducing model size and improving inference speed on embedded devices [4].
E. Deployment and Real-Time Inference
The final stage involves deploying the optimized model onto the target microcontroller. The model is converted into an embedded-compatible format and integrated with the device firmware using tools such as TensorFlow Lite for Microcontrollers, Arduino IDE, or PlatformIO [1,7].
After deployment, the device continuously collects sensor data, performs local inference, and generates predictions in real time. Depending on the application, the output may trigger an alert, activate an actuator, display information, or transmit a notification. To ensure reliable operation, developers should evaluate prediction accuracy, inference latency, memory usage, and power consumption under real-world conditions [2,7].
VIII. Conclusions and Future Scope of TinyML
Tiny Machine Learning (TinyML) has emerged as a significant advancement in embedded Artificial Intelligence by enabling machine learning inference directly on resource-constrained microcontrollers and edge devices. By shifting intelligence from centralized cloud infrastructures to the device level, TinyML addresses critical challenges associated with latency, bandwidth consumption, energy usage, privacy, and continuous internet dependency. Through optimization techniques such as model compression, quantization, pruning, and hardware-aware deployment, TinyML enables efficient execution of machine learning models within highly constrained environments.
This chapter discussed the evolution of TinyML from Cloud AI and Edge AI, its layered architecture, core optimization techniques, major application domains, advantages, limitations, and practical deployment workflow. The ability of TinyML to perform real-time, low-power, and privacy-preserving inference makes it a key enabling technology for modern IoT systems, including healthcare monitoring, industrial automation, smart agriculture, consumer electronics, and environmental sensing.
Despite existing challenges related to memory limitations, computational constraints, and deployment complexity, the future of TinyML remains highly promising. Advances in embedded hardware, AI accelerators, lightweight neural architectures, Federated Learning, Neural Architecture Search (NAS), and Tiny Transformers are expected to further enhance TinyML capabilities. As IoT ecosystems continue to expand, TinyML will play an increasingly important role in enabling intelligent, autonomous, energy-efficient, and scalable edge computing systems. Ultimately, TinyML is expected to become a foundational technology for next-generation smart devices and pervasive AI applications.
References
- Warden, P.; Situnayake, D. TinyML: Machine Learning with TensorFlow Lite on Arduino and Ultra-Low-Power Microcontrollers; O’Reilly Media: Sebastopol, CA, USA, 2019. [Google Scholar]
- Ray, S. A Review on TinyML: State-of-the-Art and Prospects. J. King Saud. Univ.-Comput. Inf. Sci. 2022, vol. 34(no. 4), 1595–1623. [Google Scholar] [CrossRef]
- Soro, M. TinyML for Ubiquitous Edge AI. MITRE Technical Report, 2020. [Google Scholar]
- Capogrosso, D.; Khan, M. T.; Sheikholeslami, A. A Machine Learning-Oriented Survey on Tiny Machine Learning. IEEE Access 2023, vol. 11, 11179–11220. [Google Scholar]
- Tsoukas, K.; Boulis, A.; Tsolis, D. A Review on the Emerging Technology of TinyML. Electronics 2024, vol. 13(no. 2), 1–28. [Google Scholar] [CrossRef]
- Schizas; Tsiftsis, T.; Voznak, M. TinyML for Ultra-Low Power AI and Large Scale IoT Deployments: A Review. Sensors 2022, vol. 22(no. 16), 1–35. [Google Scholar]
- Kesavan, E.; Chelladurai, S. Practical TinyML: Hands-on Model Training and Deployment. J. Mark. Soc. Res. 2025, vol. 2(no. 7), 219–230. [Google Scholar]
- Kesavan; Chelladurai, H. “Arduino Machine Learning Tutorial: Introduction to TinyML with Wio Terminal,” Hardware.ai, YouTube Video, 2023. Available. YouTube Tutorial. (accessed on 16 May 2026).
- Warden, P. TinyML. Computer 2019, vol. 52(no. 5), 38–42. [Google Scholar]
- Tsoukas, V.; Boumpa, A.; Gkogkidis, E.; Kakarountas, A. A Review on the Emerging Technology of TinyML. Electronics 2024, vol. 13(no. 2), 1–28. [Google Scholar] [CrossRef]
- Banbury; Reddi, V. J.; Warden, P.; Jeffries, N.; Kiraly, C.; Montino, P.; Loh, J. M.; Stoneman, C.; Plunkett, M.; Dragone, A. N.; Rus, D. Benchmarking TinyML Systems: Challenges and Direction. arXiv 2020, arXiv:2003.04821. [Google Scholar]
- 12] A. Banbury et al., MLPerf Tiny Benchmark. Proc. Mach. Learn. Syst. (MLSys) 2021, vol. 3, 502–524.
- Lin, M.; Chen, R.; Wang, S.; Li, Y.; Wang, Z. MCUNet: Tiny Deep Learning on IoT Devices. Adv. Neural Inf. Process. Syst. (NeurIPS) 2020, vol. 33, 11711–11722. [Google Scholar]
- Lin, J.; Chen, W.; Lin, Y.; Cohn, C.; Han, S. MCUNetV2: Memory-Efficient Patch-Based Inference for Tiny Deep Learning. Adv. Neural Inf. Process. Syst. (NeurIPS) 2021, vol. 34, 31172–31184. [Google Scholar]
- Panda, P.; Leong, P. H. W.; Roy, K. Recent Advances in TinyML: Models, Applications and Challenges. IEEE Des. Test. 2023, vol. 40(no. 2), 30–40. [Google Scholar]
- Kwon, H.; Banbury, A.; Loh, J. M.; Stoneman, C.; Reddi, V. J. Understanding Reuse, Performance, and Hardware Cost of DNN Dataflows: A Comprehensive Analysis for TinyML. IEEE Micro 2021, vol. 41(no. 3), 55–63. [Google Scholar]
- Han, D.; Lee, J.; Park, S.; Choi, K. TinyML Applications and Challenges in Internet of Things: A Survey. IEEE Internet Things J. 2023, vol. 10(no. 8), 6542–6563. [Google Scholar]
- Krishnan; Kumar, R. S. TinyML in Smart Healthcare and Wearable Systems: Opportunities and Challenges. Sensors 2024, vol. 24(no. 3), 1–24. [Google Scholar] [CrossRef]
- Bhattacharya, S.; Lane, N. D. From Smart to Tiny: Machine Learning for Resource-Constrained Embedded Systems. Commun. ACM 2022, vol. 65(no. 1), 50–58. [Google Scholar]
- Blouw, P.; Chao, X.; Hsu, C.; Lowe-Power, T.; Krishnamurthy, R. Benchmarking Keyword Spotting Efficiency on Neuromorphic Hardware for TinyML Applications. IEEE Embed. Syst. Lett. 2021, vol. 13(no. 4), 173–176. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.