Use this link to cite:
https://hdl.handle.net/2183/49399 Visión activa basada en aprendizaje por refuerzo profundo para la mejora de clasificadores de imágenes preentrenados
Loading...
Identifiers
Publication date
Authors
Martínez Martínez, Alejandro
Other responsabilities
Universidade da Coruña. Facultade de Informática
Journal Title
Bibliographic citation
Type of academic work
Academic degree
Abstract
[Resumen] Los clasificadores de imágenes preentrenados procesan de forma determinista la totalidad de los píxeles que se les presentan, sin distinguir entre regiones informativas y zonas irrelevan tes u ocluidas. Este trabajo investiga si un agente de Aprendizaje por Refuerzo (RL) puede dotar a estos clasificadores de una capacidad de atención dinámica, manipulando una venta na de observación móvil para maximizar la confianza del clasificador en la clase correcta, sin modificar en ningún momento sus pesos.
La experimentación se desarrolla en dos fases de complejidad creciente. En la primera, sobre Imagenette con oclusión artificial mediante Random Erasing, un agente PPO aprende a esquivar las regiones dañadas de la imagen, mejorando la precisión del clasificador estático en +2,63 puntos porcentuales y generalizando zero-shot a un clasificador no visto durante el entrenamiento. En la segunda, se escala el problema a ImageNet completo sin oclusión artifi cial, donde la ausencia de una señal de mejora clara expone una cadena de comportamientos patológicos —convergencia prematura, Reward Hacking mediante oscilaciones de encuadre, ceguera estructural a clases competidoras— que se diagnostican y corrigen de forma sistemá tica a lo largo de más de veinticinco experimentos.
Los resultados confirman que el paradigma de visión activa es viable cuando existe una señal causal clara entre la acción del agente y la recompensa, y permiten delimitar con pre cisión las condiciones bajo las cuales deja de aportar valor: cuando el clasificador ya está bien calibrado sobre el encuadre original, el margen de mejora por reencuadre es sistemática mente pequeño y ruidoso. Este resultado negativo, documentado y diagnosticado con rigor, constituye una contribución sobre las limitaciones reales del paradigma y sienta las bases de líneas de trabajo futuro orientadas a superarlas mediante memoria recurrente, extractores de características preentrenados y señales de recompensa semánticamente más ricas
[Abstract] Pretrained image classifiers process every pixel they are given in a deterministic manner, without distinguishing informative regions from irrelevant or occluded ones. This work in vestigates whether a Reinforcement Learning (RL) agent can endow these classifiers with a dynamic attention capability, manipulating a movable observation window to maximise the classifier’s confidence in the correct class, without ever modifying its weights. The experimentation is carried out in two phases of increasing complexity. In the first, on Imagenette with artificial occlusion via Random Erasing, a PPO agent learns to evade damaged regions of the image, improving the static classifier’s accuracy by +2.63 percentage points and generalising zero-shot to a classifier unseen during training. In the second, the problem is scaled to the full ImageNet dataset without artificial occlusion, where the absence of a clear improvement signal exposes a chain of pathological behaviours —premature convergence, Reward Hacking through oscillating reframing, structural blindness to competing classes—which are systematically diagnosed and corrected across more than twenty-five experiments. The results confirm that the active vision paradigm is viable when a clear causal signal ex ists between the agent’s action and the reward, and they allow the conditions under which it ceases to add value to be precisely delimited: when the classifier is already well calibrated on the original frame, the margin for improvement through reframing is systematically small and noisy. This negative result, rigorously documented and diagnosed, constitutes a methodolog ical contribution on the real limitations of the paradigm and lays the groundwork for future work aimed at overcoming them through recurrent memory, pretrained feature extractors, and semantically richer reward signals
[Abstract] Pretrained image classifiers process every pixel they are given in a deterministic manner, without distinguishing informative regions from irrelevant or occluded ones. This work in vestigates whether a Reinforcement Learning (RL) agent can endow these classifiers with a dynamic attention capability, manipulating a movable observation window to maximise the classifier’s confidence in the correct class, without ever modifying its weights. The experimentation is carried out in two phases of increasing complexity. In the first, on Imagenette with artificial occlusion via Random Erasing, a PPO agent learns to evade damaged regions of the image, improving the static classifier’s accuracy by +2.63 percentage points and generalising zero-shot to a classifier unseen during training. In the second, the problem is scaled to the full ImageNet dataset without artificial occlusion, where the absence of a clear improvement signal exposes a chain of pathological behaviours —premature convergence, Reward Hacking through oscillating reframing, structural blindness to competing classes—which are systematically diagnosed and corrected across more than twenty-five experiments. The results confirm that the active vision paradigm is viable when a clear causal signal ex ists between the agent’s action and the reward, and they allow the conditions under which it ceases to add value to be precisely delimited: when the classifier is already well calibrated on the original frame, the margin for improvement through reframing is systematically small and noisy. This negative result, rigorously documented and diagnosed, constitutes a methodolog ical contribution on the real limitations of the paradigm and lays the groundwork for future work aimed at overcoming them through recurrent memory, pretrained feature extractors, and semantically richer reward signals
Description
Keywords
Visión Activa Aprendizaje por Refuerzo Clasificación de Imágenes Reward Hacking Visión por Computador Aprendizaje por Transferencia Proximal Policy Optimization POMDP ImageNet Memoria Recurrente Active Vision Reinforcement Learning Image Classification Reward Hacking Computer Vision Transfer Learning Recurrent Memory
Editor version
Rights
Attribution-NonCommercial-ShareAlike 4.0 International







