Publicado

2026-08-12

Modelo de Inteligencia Artificial en el borde basado en Ollama para inferencia local en dispositivos de Internet de las cosas

Ollama-Based edge Artificial Intelligence model for local inference on Internet of things devices

DOI:

https://doi.org/10.15446/dyna.v93n242.125408

Palabras clave:

Internet de las cosas, dispositivo embebido, inteligencia artificial, modelo extenso de lenguaje, computación en el borde (es)
Internet of things, embedded device, artificial intelligence, large language model, edge computing (en)

Descargas

Autores/as

El Internet de las Cosas (IoT) depende cada vez más de la computación en la nube para el procesamiento de datos y la toma de decisiones, lo que introduce riesgos de latencia, conectividad y privacidad en dominios sensibles como la vida asistida y los hogares inteligentes. Este artículo presenta la Arquitectura de Borde Integrada con Ollama (OLLiE), que permite el razonamiento local en dispositivos IoT embebidos mediante modelos de lenguaje de gran tamaño (LLM) cuantizados. OLLiE fue implementada en un escenario de hogar inteligente asistido utilizando Raspberry Pi 5, ThingsBoard, Redis, Ollama y modelos LLM locales y basados en la nube de las familias Gemma, Phi y Qwen. La evaluación consideró métricas funcionales, incluyendo precision, recall y F1-score, y métricas no funcionales, como latencia, uso de RAM, carga de CPU, temperatura y consumo energético estimado. Los resultados revelan un compromiso entre la velocidad de inferencia, la privacidad y la sostenibilidad computacional. Los modelos en la nube alcanzaron menor latencia (2.5–3.1 s), mientras que los modelos locales preservaron el procesamiento en el borde con mayor latencia (76–115 s). Entre ellos, phi-3.5-mini (Q2_K) obtuvo la mejor calidad de inferencia local.

The Internet of Things (IoT) increasingly relies on cloud computing for data processing and decision-making, introducing latency, connectivity, and privacy risks in sensitive domains such as assisted living and smart homes. This paper presents the Ollama-Integrated Edge Architecture (OLLiE), which enables local reasoning on embedded IoT devices using quantized large language models (LLMs). OLLiE was implemented in an assisted smart home scenario using Raspberry Pi 5, ThingsBoard, Redis, Ollama, and local and cloud-based LLMs from the Gemma, Phi, and Qwen families. The evaluation considered functional metrics, including precision, recall, and F1-score, and non-functional metrics, such as latency, RAM usage, CPU load, temperature, and estimated energy consumption. Results reveal a trade-off between inference speed, privacy,
and computational sustainability. Cloud models achieved lower latency (2.5–3.1 s), whereas local models preserved edge processing with higher latency (76–115 s). Among them, phi-3.5-mini (Q2_K) achieved the best local inference quality.

Referencias

[1] Mouha, R.A., Internet of Things (IoT). Journal of Data Analysis and Information Processing, 9, pp. 77–101, 2021. DOI: https://doi.org/10.4236/jdaip.2021.92006.

[2] Hajjaji, Y., Boulila, W., Farah, I.R., Romdhani, I., and Hussain, A., Big data and IoT-based applications in smart environments: a systematic review. Comput. Sci. Rev., 39, art. 100318, 2021. DOI: https://doi.org/10.1016/j.cosrev.2020.100318.

[3] Oliveira, F., Costa, D.G., Assis, F., and Silva, I., Internet of intelligent things: a convergence of embedded systems edge computing and machine learning. Internet of Things, 26, art. 101153, 2024. DOI: https://doi.org/10.1016/j.iot.2024.101153.

[4] Arjunan, G., Optimizing edge ai for real-time data processing in iot devices: challenges and solutions. International Journal of Scientific Research and Management (IJSRM), 11(06), pp. 944–953, 2023. DOI: https://doi.org/10.18535/ijsrm/v11i06.ec2.

[5] Wakili, A., and Bakkali, S., Privacy-preserving security of IoT networks: a comparative analysis of methods and applications. Cyber Security and Applications, 3, art. 100084, 2025. DOI: https://doi.org/10.1016/j.csa.2025.100084.

[6] Kalita, A., Large Language Models (LLMs) for semantic communication in edge-based IoT networks. Arxiv, art. 20970, 2024. DOI: https://doi.org/10.48550/arXiv.2407.20970.

[7] Alwahedi, F., Aldhaheri, A., Ferrag, M.A., Battah, A., and Tihanyi, N., Machine learning techniques for IoT security: current research and future vision with generative AI and large language models. Internet of Things and Cyber-Physical Systems, 4, pp. 167–185, 2024. DOI: https://doi.org/10.1016/j.iotcps.2023.12.003.

[8] McIntosh, F., Murina, S., Chen, L., Vargas, H.A., and Becker, A.S., Keeping private patient data off the cloud: a comparison of local LLMs for anonymizing radiology reports. European Journal of Radiology Artificial Intelligence, 2, art. 100020, 2025. DOI: https://doi.org/10.1016/j.ejrai.2025.100020.

[9] Liu, F., Kang, Z., and Han, X., Optimizing RAG Techniques for automotive industry pdf chatbots: a case study with locally deployed ollama models. In: AIIIP ’24: Proceedings of the 2024 3rd International Conference on Artificial Intelligence and Intelligent Information

Processing, New York, pp. 152–159, DOI: https://doi.org/10.1145/3707292.3707358.

[10] Husom, E.J., Goknil, A., Astekin, M., Shar, L.K., Kåsen, A., Sen, S., Mithassel, B.A., and Soylu, A., Sustainable LLM inference for edge ai: evaluating quantized llms for energy efficiency, output accuracy, and inference latency, acm trans. Internet Things, 6(4), pp. 1–35, 2025. DOI: https://doi.org/10.1145/3767742.

[11] Kok, I., Demirci, O., and Ozdemir, S., When IoT Meet LLMs: applications and challenges. In: 2024 IEEE International Conference on Big Data (BigData), Los Alamitos, CA, IEEE Computer Society, pp. 7075–7084, art. 5187, 2024. DOI: https://doi.org/10.1109/BigData62323.2024.10825187.

[12] Alhafnawi, M., Abu-Ein, A., Bany-Salameh, H., Jararweh, Y., and Al- Hazaimeh, O., Intelligent dynamic bandwidth allocation for real-time IoT in fog-based optical networks. Simulation Modelling Practice and Theory, 142, art. 103126, 2025. DOI: https://doi.org/10.1016/j.simpat.2025.103126.

[13] Shen, Y., Shao, J., Zhang, X., Lin, Z., Pan, H., Li, D., Zhang, J., and Letaief, K.B., large language models empowered autonomous edge AI for connected intelligence. IEEE Communications Magazine, 62(10), art. 140, 2024. DOI: https://doi.org/10.1109/MCOM.001.2300550.

[14] Cao, K., Liu, Y., Meng, G., and Sun, Q., An overview on edge computing research. IEEE, 8, pp. 85714–85728, 2020. DOI: https://doi.org/10.1109/ACCESS.2020.2991734.

[15] Hua, H., Li, Y., Wang, T., Dong, N., Li, W., and Cao, J., Edge computing with artificial intelligence: a machine learning perspective. ACM Computing Surveys, 55(9), pp. 1–35, 2023. DOI: https://doi.org/10.1145/3555802.

[16] Hartmann, M., Hashmi, U.S., and Imran, A., Edge computing in smart health care systems: review, challenges, and research directions. Transactions on Emerging Telecommunications Technologies, 33(3), art. 3710, 2022. DOI: https://doi.org/10.1002/ett.3710.

[17] Naveed, H., Khan, A.U., Qiu, S., Saqib, M., Anwar, S., Usman, M., Akhtar, N., Barnes, N., and Mian, A., A comprehensive overview of large language models. ACM Trans. Intell. Syst. Technol., 16(5), pp. 1-72, 2025. DOI: https://doi.org/10.1145/3744746.

[18] An, T., Zhou, Y., Zou, H., and Yang, J., IoT-LLM: a framework for enhancing large language model reasoning from real-world sensor

data, arXiv, art. 2429, 2025. DOI: https://doi.org/10.48550/arXiv.2410.02429.

[19] Chen, X., Wu, W., Li, L., and Ji, F., LLM-Empowered IoT for 6G Networks: architecture, challenges, and solutions. IEEE Internet of Things Magazine, 8(6), pp. 34–41, 2025. DOI: https://doi.org/10.1109/MIOT.2025.3582641.

[20] ThingsBoard, ThingsBoard—Open-source IoT (Internet of Things) Platform, [online]. 2026. [date of reference November 24th of 2025]. Available at: https://thingsboard.io/.

[21] Hosny, K.M., Magdi, A., Salah, A., El-Komy, O., and Lashin, N.A., Internet of things applications using Raspberry-Pi: a survey. International Journal of Electrical and Computer Engineering, 13(1), pp. 902–910, 2023. DOI: https://doi.org/10.11591/ijece.v13i1.pp902-910.

[22] Raspberry, Pi-Ltd, Raspberry Pi AI HAT+, [online]. 2024. [date of reference December 14th 2025]. Available at: https://www.farnell.com/datasheets/4423344.pdf

[23] Hiltgen, D., Get up and running with OpenAI gpt-oss, DeepSeek-R1, Gemma 3 and other models, [online]. 2026. [date of reference November 25th 2025]. Available at: https://github.com/ollama/ollama

[24] Redis, Redis.io, [online]. 2025. [date of reference November 24th 2025]. Available at: https://redis.io/

[25] The Pandas Development Team, Pandas-dev/pandas: Pandas, [online]. 2020. [date of reference December 14th 2025]. Available at: https://doi.org/10.5281/zenodo.3509134.

[26] Rodola, G., Cross-platform lib for process and system monitoring in Python, [online]. 2025. [date of reference December 1st 2025]. Available at: https://github.com/giampaolo/psutil.

[27] Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, É., Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 12, pp. 2825–2830, 2011.

[28] Hunter, J.D., Matplotlib: A 2D graphics environment. Computing in Science and Engineering, 9(3), pp. 90–95, 2007. DOI: https://doi.org/10.1109/MCSE.2007.55.

[29] Fradkin, A., Demand for LLMs: descriptive evidence on substitution, Market Expansion, and Multihoming, arXiv, art. 15440, 2025. DOI: https://doi.org/10.48550/arXiv.2504.15440.

Dimensions

PlumX

Visitas a la página del resumen del artículo

34

Descargas

Los datos de descarga aún no están disponibles.

Cómo citar

[1]
“Modelo de Inteligencia Artificial en el borde basado en Ollama para inferencia local en dispositivos de Internet de las cosas”, DYNA, vol. 93, no. 242, pp. 128–138, Aug. 2026, doi: 10.15446/dyna.v93n242.125408.