An extended evaluation of the impact of different modules in ST-VQA systems - La Rochelle Université Accéder directement au contenu
Chapitre D'ouvrage Année : 2020

An extended evaluation of the impact of different modules in ST-VQA systems

Résumé

Scene Text VQA has been recently proposed as a new challenging task in the context of multimodal content description. The aim is to teach traditional VQA models to read text contained in natural images by performing a semantic analysis between the visual content and the textual information contained in associated questions to give the correct answer. In this work, we present results obtained after evaluating the relevance of different modules in the proposed frameworks using several experimental setups and baselines, as well as to expose some of the main drawbacks and difficulties when facing this problem. We makes use of a strong VQA architecture and explore key model components such as suitable embeddings for each modality, relevance of the dimension of the answer space, calculation of scores and appropriate selection of the number of spaces in the copy module, and the gain in improvement when additional data is sent to the system. We make emphasis and present alternative solutions to the out-of-vocabulary (OOV) problem which is one of the critical points when solving this task. For the experimental phase, we make use of the TextVQA database, which is one of the main databases targeting this problem.
Fichier principal
Vignette du fichier
ICPRAI_camera_ready.pdf (1.51 Mo) Télécharger le fichier
Origine : Fichiers produits par l'(les) auteur(s)

Dates et versions

hal-03026917 , version 1 (26-11-2020)

Identifiants

Citer

Viviana Beltrán, Mickaël Coustaty, Nicholas Journet, Juan Caicedo, Antoine Doucet. An extended evaluation of the impact of different modules in ST-VQA systems. ICPRAI: International Conference on Pattern Recognition and Artificial Intelligence, pp.562-574, 2020, ⟨10.1007/978-3-030-59830-3_49⟩. ⟨hal-03026917⟩
34 Consultations
115 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More