Use this link to cite:
https://hdl.handle.net/2183/49607 Análisis y evaluación de estrategias de aprendizaje federado
Loading...
Identifiers
Publication date
Authors
Cárdenas Corral, Antón
Advisors
Other responsabilities
Universidade da Coruña. Facultade de Informática
Journal Title
Bibliographic citation
Type of academic work
Academic degree
Abstract
[Resumen] Este Trabajo de Fin de Grado aborda el estudio experimental del aprendizaje federado, un paradigma de entrenamiento distribuido en el que varios nodos entrenan modelos de forma local y comparten únicamente actualizaciones del modelo con un servidor central, evitando así la transferencia directa de los datos. Partiendo de una implementación de referencia del algoritmo Federated Averaging (FedAvg), el trabajo analiza el efecto de la heterogeneidad y el desbalanceo de los datos entre nodos, evalúa una representación del aprendizaje federado basada en deltas de pesos (incluyendo distintas precisiones numéricas), y estudia diversas estrategias para reducir el coste de comunicación (esparsificación mediante codificación por coordenadas, por longitud de racha y codificación entrópica de Huffman). Además, para preservar la privacidad de las actualizaciones, analiza el uso de cifrado homomórfico (esquemas Cheon-Kim-Kim-Song (CKKS) y Brakerski/Fan-Vercauteren (BFV)). Todos los experimentos se evalúan sobre el conjunto de datos CIFAR-10 utilizando una arquitectura de red neuronal convolucional común, comparando en todos los casos el rendimiento, el coste de comunicación y el tiempo de ejecución frente a un entrenamiento centralizado de referencia.
Los resultados muestran que FedAvg aproxima de forma notable el rendimiento centralizado en condiciones favorables, pero se degrada de forma acusada frente a distribuciones de datos no homogéneas; que la representación basada en deltas reproduce fielmente el comportamiento de FedAvg, y que reducir su precisión numérica a float16 o int8 apenas afecta a la accuracy mientras reduce sustancialmente el ancho de banda necesario para la comunicación; que la esparsificación por capa combinada con cuantización a int8 ofrece el mejor compromiso entre precisión y ahorro de comunicación; y que el cifrado homomórfico preserva la precisión del modelo sin pérdida alguna, a costa de un incremento muy significativo del volumen de datos transmitido, siendo el esquema BFV sistemáticamente más eficiente en comunicación que CKKS para la configuración evaluada.
[Abstract] This Bachelor’s Thesis presents an experimental study of federated learning, a distributed training paradigm in which multiple nodes train models locally and share only model updates with a central server, thereby avoiding the direct transfer of data. Starting from a reference implementation of the FedAvg algorithm, the work analyzes the effect of data heterogeneity and imbalance across nodes, evaluates a representation of federated learning based on weight deltas (including several numeric precisions) and studies various strategies for reducing communication cost (sparsification via coordinate-list, run-length encoding and Huffman entropy coding). In addition, and for preserving the privacy of the updates, the use of homomorphic encryption (CKKS and BFV schemes) has been analyzed. All experiments are evaluated on the CIFAR-10 dataset using a common convolutional neural network architecture, comparing in every case performance, communication cost, and execution time against a centralized training baseline. The results show that FedAvg closely approaches centralized performance under favorable conditions, but degrades sharply under non homogeneous data distributions; that the delta-based representation faithfully reproduces FedAvg’s behavior, and that reducing its numeric precision to float16 or int8 barely affects accuracy while substantially reducing bandwidth communication; that per-layer sparsification combined with int8 quantization offers the best trade-off between accuracy and communication savings; and that homomor phic encryption preserves model accuracy without any loss, at the cost of a very significant increase in the volume of transmitted data, with the BFV scheme being consistently more communication-efficient than CKKS under the evaluated configuration.
[Abstract] This Bachelor’s Thesis presents an experimental study of federated learning, a distributed training paradigm in which multiple nodes train models locally and share only model updates with a central server, thereby avoiding the direct transfer of data. Starting from a reference implementation of the FedAvg algorithm, the work analyzes the effect of data heterogeneity and imbalance across nodes, evaluates a representation of federated learning based on weight deltas (including several numeric precisions) and studies various strategies for reducing communication cost (sparsification via coordinate-list, run-length encoding and Huffman entropy coding). In addition, and for preserving the privacy of the updates, the use of homomorphic encryption (CKKS and BFV schemes) has been analyzed. All experiments are evaluated on the CIFAR-10 dataset using a common convolutional neural network architecture, comparing in every case performance, communication cost, and execution time against a centralized training baseline. The results show that FedAvg closely approaches centralized performance under favorable conditions, but degrades sharply under non homogeneous data distributions; that the delta-based representation faithfully reproduces FedAvg’s behavior, and that reducing its numeric precision to float16 or int8 barely affects accuracy while substantially reducing bandwidth communication; that per-layer sparsification combined with int8 quantization offers the best trade-off between accuracy and communication savings; and that homomor phic encryption preserves model accuracy without any loss, at the cost of a very significant increase in the volume of transmitted data, with the BFV scheme being consistently more communication-efficient than CKKS under the evaluated configuration.
Description
Keywords
Aprendizaje federado Aprendizaje distribuido Compresión de gradientes Esparsificación Cuantización Codificación de Huffman Cifrado homomórfico Coordinate List (COO) Run-Length Encoding (RLE) CKKS Privacidad de datos Redes neuronales convolucionales BFV Federated learning Distributed learning Gradient compression Sparsification Quantization Huffman coding; Homomorphic encryption COO RLE Data privacy Convolutional neural network
Editor version
Rights
Todos os dereitos reservados. Prohíbese a reprodución, transformación, distribución e comunicación pública da obra por terceiros. Permítese a súa visualización e a descarga dunha copia privada para o uso persoal.








