Skip to content

Digital marketing

Edge AI – artificial intelligence closer to the user

Read the articleQuestions and answers

Article cover: Edge AI – artificial intelligence closer to the user
Edge AI is an approach in which machine learning models run directly on the endpoint device, rather than “somewhere in the cloud”. As a result, decisions are made on site, in milliseconds, which matters greatly in systems that respond to real-time events. For many people, the question “does it work without the internet?” also remains important — in Edge AI, inference can take place offline, and the network is usually needed mainly to update models. Processing on the device limits the removal of sensitive data, such as images, audio or biometrics, from the place where it is created. This, in turn, makes it easier to minimise data and supports GDPR compliance, and with streams (e.g. video) it can also reduce transfer and storage costs. In the next part, you will see how Edge AI differs from cloud and fog computing and when this approach makes the most sense.

what is Edge AI and why is it important?

Edge AI means running machine learning models directly on the endpoint device (e.g. a phone, camera or sensor), rather than in the cloud. The most important consequence is that inference can work offline, and internet connectivity is sometimes needed only to update the model or synchronise results. In practice, this translates into a shorter response time, because the decision does not have to “travel” to the data centre and back. This approach works particularly well in streaming workloads (frame by frame or in time windows), where predictable, consistent latency matters.

Edge AI matters because it limits sending sensitive data beyond the device, which reduces the risk of leakage and supports a privacy-by-design approach. Instead of transmitting raw video or audio, the device can send only metadata, such as the number of people or the event alert itself. In addition, in places with poor coverage (e.g. warehouses, mines, vehicles, critical infrastructure), local inference ensures continuity of operation, and the results can be synchronised later. In many cases, on-site processing is also more cost-effective, because with continuous streams (e.g. 1080p video), transfer and data storage quickly begin to dominate costs.

Technology and privacy What is Edge AI and why is it important?
  1. 01Model on the endpoint deviceMachine learning locally (e.g. smartphone, camera).
  2. 02Offline inferenceOperation without the cloud, independent of the network.
  3. 03Shorter response timeImmediate decisions, predictable latency.
  4. 04Greater privacyLimits data transmission, privacy-by-design.

In short: local, faster and more secure data analysis without relying on the cloud.

how does Edge AI differ from cloud and fog computing?

Edge AI differs from the cloud in that processing and decision-making take place locally on the device, rather than in a remote data centre. In a cloud model, data has to be sent to the server, which increases latency as well as transfer and storage costs. In Edge AI, a decision can be made in milliseconds, which can be critical in applications requiring a fast response. For example, the difference between 10–30 ms locally and 200–800 ms in the cloud can affect effectiveness in scenarios such as emergency braking in robotics or fall detection in care.

Fog computing is an intermediate layer between the edge and the cloud, e.g. an IoT gateway in a plant that aggregates streams and runs models larger than those on a single sensor, while still operating closer to the data source than the cloud. An edge-first architecture is often used: decisions are made locally, while aggregates, metrics and selected samples used for training go to the cloud. When the aim is to reduce the number of errors without constantly transmitting data, hybrid approaches are used, in which a lightweight “detector” runs on the device, and a heavier “verifier” is launched in the cloud only for uncertain cases. Edge is not always “automatically” faster — it usually wins in inference, but only provided that the model is well optimised and matched to the hardware.

the role of privacy and data security in Edge AI

Privacy and data security in Edge AI are crucial, because processing takes place where the data is created, which limits its exposure beyond the device. In many deployments, the principle of data minimisation works best: raw audio/video is analysed locally, and only results, statistics or events are sent onwards. For example, a camera can pass on metadata (e.g. number of people, alert) instead of the raw image, and retain the footage only as a short incident clip. This approach supports privacy-by-design and makes it easier to meet GDPR requirements, because it limits the transfer of personal data.

Security in Edge AI is built through encryption, integrity controls and identity management for devices across the fleet. Data at rest should be encrypted (e.g. AES-256), and communication secured with TLS 1.2/1.3 with mutual authentication, e.g. in an MQTT over TLS scheme with client certificates. Secure Boot and cryptographic signatures for firmware and models reduce the risk of software tampering, and where possible TPM/TEE-type mechanisms are also used. In addition, component isolation (e.g. containers and least privilege), key rotation and a unique device identity are used to make it harder to impersonate edge nodes.

Auditability and resilience to risks characteristic of ML models operating in the field are also important. Logs of model decisions (time, version, confidence) and the update path help explain incidents without the need to store full input data, unless this is necessary. Models may be exposed to adversarial attacks or data poisoning, so it is worth monitoring unusual input distributions (drift) and limiting automatic training on unverified data. Since edge devices are often physically located “in the field”, hardware attacks are also realistic (e.g. removal of the storage medium, debug ports), so measures such as tamper-evident enclosures, disabling debug and disk encryption are used, among others.

Technology and security The role of privacy and data security in Edge AI
  1. 01Local processingLimits data exposure
  2. 02Data minimisationOnly results and metadata
  3. 03Privacy-by-designEases GDPR compliance
  4. 04Advanced security measuresEncryption and integrity

Local processing in Edge AI supports privacy by limiting data transfer and minimising exposure.

optimisation of AI models for edge devices

Optimising AI models for edge devices comes down to adapting them to the constraints of memory, computing power and the requirements of stable latency on specific hardware. It usually starts with compression and choosing the architecture so that the model “fits” within the MB/ms budget while maintaining the expected throughput (e.g. FPS). In practice, both lightweight mobile networks (MobileNetV3, EfficientNet-Lite, ShuffleNet) and trimmed-down detection variants (e.g. YOLO nano/tiny) are used. At the same time, the runtime and deployment path are selected (e.g. TensorRT, OpenVINO, TFLite, ONNX Runtime), most often without rewriting the model, but with the need to export it and check operator compatibility.

  • Quantisation (e.g. INT8) to reduce model size and speed up inference, usually with only a small drop in quality after calibration or when using quantization-aware training.
  • Pruning (removing weights/channels) to limit the number of computations and memory usage; in CNNs, channel pruning can reduce computations by 20–50%, provided the runtime/hardware actually makes use of it.
  • Distillation (teacher → student), which makes it possible to preserve the quality of a small model better than training “from scratch”.
  • Compilation and tuning for the runtime (e.g. TensorRT/OpenVINO/TFLite) and checking operator compatibility after export.
  • Multi-task (multi-head) models instead of several separate models, to save memory and time in the pipeline.

End-to-end profiling delivers the biggest gains, because the bottleneck often lies not in the inference itself, but in preprocess and postprocess. If the accelerator is not “speeding things up”, a common cause is postprocessing on the CPU (e.g. NMS in detection), so it is worth measuring resize/normalise, inference and the final stage separately. In vision systems, moving preprocessing to the ISP/GPU and using zero-copy between the camera and the accelerator brings significant benefits, and can reduce latency by several to a dozen milliseconds at 30 FPS. After each modification, regression tests for quality (accuracy/F1/mAP) and latency stability should be carried out on the target device, because what works offline may behave differently in a real stream.

Applications of Edge AI in different industries

Edge AI is used in many industries where fast response, working on data streams and limiting the sending of raw data outside the device matter. In intelligent video surveillance, cameras can detect people, vehicles or intrusions and generate alerts without continuously streaming to the cloud. To reduce false alarms (e.g. from shadows), in practice object detection is used instead of simple motion detection, and zones and confidence thresholds are set. In retail (traffic analysis), edge devices count customers, measure queue length and detect empty shelves, passing on metrics instead of storing images.

In consumer devices, Edge AI strengthens audio functions and privacy protection, because some computations can remain on the device itself. Wake word in voice assistants often works locally, and only after activation does a short audio fragment go to the cloud or get analysed on the spot. In smart home and home automation, locally running models and rules enable control of heating, lighting and security without worrying about network latency, and the internet is mainly needed for remote viewing and updates. In education, Edge AI on school tablets can support handwriting recognition, translations or offline exercises, with progress synchronised only after reconnection.

In industry and medicine, Edge AI is used for rapid event detection and signal analysis as close to the source as possible. In predictive maintenance, models running on a gateway or MCU analyse vibrations and acoustics, catching deviations from normal operation, often in the form of anomaly detection trained on “healthy” data. In health devices (watches, bands), Edge AI supports ECG analysis, saturation monitoring and fall detection, and can pass only the event and a signal summary to further systems. In automotive (ADAS and in-cabin functions), on-device processing helps detect fatigue and distraction, where minimal latency and reliability are crucial, and cabin data is often not sent.

Technology and applications Applications of Edge AI across different industries
  1. 01Intelligent monitoringObject detection, rapid response
  2. 02Traffic analysis in retailCustomer counting, without streaming
  3. 03Consumer devicesPrivacy protection, local processing

Main benefits: Speed, reduced data transfer and device-level privacy protection.

How to choose hardware and accelerators for Edge AI?

Hardware and accelerators for Edge AI are selected for a specific model and the application requirements, such as latency, throughput, power consumption and thermal constraints. CPU is a universal solution, but usually slower for matrix computations, so for vision tasks and stable operation at high FPS, GPU or NPU/TPU is more often used. In smartphones, typical accelerators are Qualcomm Hexagon (DSP/NPU) used by Android NNAPI, designed for low power consumption in always-on scenarios (e.g. background speech). In practice, the “CPU or accelerator” decision should be based on whether you can maintain constant latency on the target device, not just on a test result on a computer.

Among the most common edge platforms are NVIDIA Jetson, Google Coral, Intel solutions and microcontrollers for TinyML. NVIDIA Jetson (Nano, Xavier, Orin) is popular in robotics and vision thanks to support for CUDA and TensorRT for optimisation, and Jetson Orin Nano can run real-time detection models at a power draw of around a few to a dozen watts (depending on the configuration). Google Coral with Edge TPU accelerates TFLite inference, especially INT8, but requires compatibility with the Edge TPU compiler and usually quantisation and specific operators. Intel (e.g. NUC-class mini PCs) with OpenVINO allows you to accelerate inference on CPU/GPU/iGPU and VPU, often without changing the application code, with a significant improvement in inference time for classification and detection.

  • Model size (MB) and device memory – at the edge, typical memory limits are around 256 MB–8 GB, so the model often needs to be smaller or compressed.
  • Latency budget (ms) and throughput (FPS) – e.g. for a 30 FPS camera, you usually need an NPU/GPU, because CPU may not maintain constant latency during detection.
  • Power consumption (W) and thermals – TDP limits, throttling and operation in a fanless enclosure can reduce FPS over time, so cooling and power limits are planned in advance.
  • Data type and scenario – for 1 kHz audio classification, an MCU (TinyML) may be enough, whereas for 720p image segmentation you usually need a Jetson, an NPU in the camera or a high-performance iGPU.

In many deployments, it is also worth considering specialised devices such as smart cameras with a SoC, which have a built-in ISP and AI accelerator. This architecture makes it possible to perform image processing and inference in a single chip, so the system mainly sends events and results, without adding a separate computer. However, compatibility limitations must be taken into account: not every model will run on every accelerator (e.g. Edge TPU requires compatibility and usually INT8). Ultimately, platform selection should cover not only the question of “will it work”, but also long-term operational stability and field conditions (temperature, power supply, connectivity).

Managing and deploying Edge AI in a production environment

Managing and deploying Edge AI in production means treating models like software, with versioning, rollout control and monitoring of performance across a fleet of devices. In practice, you need a model registry and a mapping of model versions to specific devices and deployment groups so you have clarity on “what is running where”. Over-the-air updates should be atomic and allow rollback, because an interrupted update can disable the device. An A/B strategy (two partitions) with automatic rollback works well if the health check does not pass in around 60–120 seconds.

Safe field deployments are carried out through gradual rollouts and post-update quality control. First, you run the new model on 1–5% of devices (canary), measure the metrics and only then expand the deployment to limit the risk of a “catastrophic” update. At the same time, data drift is monitored (e.g. lighting changes, seasonality, new products), because it can reduce model performance in the real environment. On the operational side, decision logs (time, model version, confidence) and stability metrics are key, not just a one-off test result.

Maintaining Edge AI also means collecting data sensibly so that models can be improved and resilience to connectivity outages can be preserved. Instead of sending the entire stream, diagnostic samples are chosen: low-confidence cases, new data clusters, user errors, usually around 0.1–1% of frames, and event clips. With unstable connectivity, buffering and store-and-forward are used (e.g. MQTT queues with a local broker and file rotation) to avoid filling up the device memory. In practice, operational tools such as AWS IoT Greengrass, Azure IoT Edge, balena.io, KubeEdge and Eclipse Mosquitto/MQTT are used, and field actions are made easier by feature flags (e.g. raising the confidence threshold without deploying new firmware).

what are the challenges and limitations of Edge AI?

The challenges and limitations of Edge AI stem mainly from the fact that edge devices have limited memory, computing power and thermal budget, so not every model can be run “as in the cloud”. A typical memory range at the edge is around 256 MB–8 GB, which often forces smaller architectures or compression. On top of that come TDP limits and throttling: during longer operation, especially in a fanless enclosure, a drop in FPS is often the result of overheating and clock speed reduction. Edge AI usually provides faster inference, but only when the model is well optimised and matched to the hardware.

Compatibility can also be a limitation, as can the real performance of the entire pipeline, not just the network itself. Not every model will run on every accelerator (e.g. Edge TPU requires compatibility with the compiler and usually quantisation plus specific operators), so already at the design stage you need to check export and runtime compatibility. In practice, bottlenecks are often found in preprocessing and postprocessing rather than in inference itself, so end-to-end profiling is essential (separately: preprocess, inference, postprocess). If the question comes up “will it be as accurate as in the cloud?” — it can be slightly less accurate, but well-chosen lightweight architectures and quantisation often maintain quality while delivering a major speed gain.

Reliability, security and maintenance costs at fleet scale can also be challenging. Devices operating in the field should have degradation modes (e.g. NPU failure → a simpler model on CPU or rules), decision thresholds, and watchdogs and health checks, so that decisions remain operationally safe. Models may be exposed to adversarial attacks or data poisoning, especially when they learn from field data, and physical attacks on hardware are also realistic (e.g. access to storage media or debug ports), which requires a well thought-out security plan. In practice, TCO includes not only hardware, but also energy, connectivity, service and the model lifecycle — and, for example, a camera with an NPU may be 20–40% more expensive, while at the same time reducing the costs of video transfer and storage over a year.

FAQ

Frequently asked questions

How does Edge AI work without the internet?

Models run locally on the end device, so inference can take place offline. The internet is mainly needed for model updates or result synchronisation.

How is Edge AI different from the cloud?

In Edge AI, processing and decision-making happen on the device, rather than in a remote data centre. As a result, latency is lower because data does not need to be sent to and from the cloud.

Why is Edge AI important for data privacy?

Because it limits moving sensitive data, such as images, audio or biometrics, away from where it was created. Instead of raw data, only metadata, results or alerts can be sent.

When does Edge AI make the most sense?

It works best where fast response, consistent latency and working with data streams matter. It is also useful in places with weak coverage, where local inference ensures business continuity.

Which industries most often use Edge AI?

The article mentions, among others, video monitoring, retail, smart home, education, industry, medicine and automotive. In these areas, Edge AI is used for rapid event detection, signal analysis and reducing data transfer.

How is hardware selected for Edge AI?

The choice depends on latency, bandwidth, power consumption, memory and the type of data. CPU is versatile, but for vision tasks GPU, NPU, TPU, Jetson, Coral or Intel solutions are more often chosen.

Contents