Résumés
Résumé
Si historiquement, les bibliothèques numériques patrimoniales furent d’abord alimentées par des images, elles profitèrent rapidement de la technologie OCR pour indexer les collections imprimées afin d’améliorer le service de recherche d’information offert aux utilisateurs. Mais l’accès aux ressources iconographiques n’a pas connu les mêmes progrès et ces dernières demeurent dans l’ombre : indexation manuelle lacunaire, hétérogène et impossible à généraliser ; silos par genre documentaire ; recherche dans le contenu des images encore peu opérationnelle sur les collections patrimoniales. Aujourd’hui, il serait pourtant possible de mieux valoriser ces ressources en exploitant les énormes volumes d’OCR produits durant les deux dernières décennies (tant comme descripteur textuel que pour l’identification automatique des illustrations des imprimés), en profitant de la maturité des techniques d’intelligence artificielle (en particulier l’apprentissage profond ou deep learning), pour mettre ainsi en valeur ces gravures, dessins, photographies, cartes, etc., pour leur valeur propre, mais aussi comme point d’entrée dans les collections, en favorisant découverte et rebond.
Cet article décrit une approche ETL (extract-transform-load) appliquée aux images d’une bibliothèque numérique à vocation encyclopédique : identifier et extraire l’iconographie partout où elle se trouve (dans les collections d’images, mais aussi dans les imprimés) ; transformer, harmoniser et enrichir ses métadonnées descriptives grâce à l’IA ; intégrer ces données dans une application web dédiée à la recherche iconographique. Cette approche est qualifiée de pragmatique à double titre, puisqu’il s’agit de valoriser des ressources numériques existantes tout en mettant à profit les acquis de l’IA.
Abstract
If historically, heritage digital libraries were initially made up of images, they rapidly benefited from the optical character recognition (OCR) technology to index print collections and improve reference services for users. However, access to iconographic resources has not experienced the same progression, remaining somewhat difficult to access. Manual indexation is not very efficient, it is varied and impossible to apply uniformly. Searching the content of an image is not as effective with heritage collections. Today, it is possible to improve the use of these resources by exploiting large volumes of OCR produced over the past two decades (both the textual descriptors as well as the automatic identification of the illustrations in the printed documents) and to take advantage of proven artificial intelligence techniques, especially deep learning. In doing so, it will showcase engravings, drawings, photographs, maps, etc. as such but also the point of entry to the collections by improving discovery and connections.
This article describes an ETL (extract-transform-load) approach as it applies to the images in a digital library with an encyclopedic vocation. There are three components: 1) identify and extract the iconography wherever it is found, either in images or in the printed documents, 2) transform, harmonise and enrich the descriptive metadata with the help of artificial intelligence, and 3) incorporate this data into a web application dedicated to iconographic research. This is a two-pronged approach because it highlights existing digital resources and takes advantage of the benefits of artificial intelligence.
Parties annexes
Bibliographie
- Bermès, E. (2017, août). Text, Data and Link-Mining in Digital Libraries : Looking for the Heritage Gold. Communication présentée à la conférence IFLA Satellite Meeting, Digital Humanities – Opportunities and Risks : Connecting Libraries and Research, Berlin, Allemagne. Repéré à www.ifla.org/files/assets/academic-and-research-libraries/conferences/emmanuelle_bermes_keynote.pdfGoogle Scholar Rechercher cette référence bibliographique sur Google Scholar
- Bibliothèque nationale de France (BnF). (2017). Enquête auprès des usagers de la bibliothèque numérique Gallica. Repéré à www.bnf.fr/documents/mettre_en_ligne_patrimoine_enquete.pdfGoogle Scholar Rechercher cette référence bibliographique sur Google Scholar
- Breiteneder, C. et Eidenberger, H. (2000, février). Content-Based Image Retrieval in Digital Libraries. Communication présentée à la Conférence internationale de Kyoto sur les bibliothèques numériques, Japon. doi.org/10.1109/DLRP.2000.942186Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Chiron, G., Doucet, A., Coustaty, M., Visani, M. et Moreux, J.-P. (2017, juin). Impact of OCR Errors on the Use of Digital Libraries. Communication présentée à la 17e conférence commune ACM/IEEE sur les bibliothèques numériques, Toronto, Ontario. doi.ieeecomputersociety.org/10.1109/JCDL.2017.799158210.1109/JCDL.2017.7991582 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Coustaty, M., Pareti, R., Vincent, N. et Ogier, J.-M. (2011). Towards Historical Document Indexing : Extraction of Drop Cap Letters. International Journal on Document Analysis and Recognition, 14(3), 243-254. Repéré à hal.archives-ouvertes.fr/hal-00916007/document10.1007/s10032-011-0152-x Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Datta, R., Joshi, D., Li, J. et Wang, J. (2008). Image Retrieval : Ideas, Influences, and Trends of the New Age. ACM Computing surveys, 40(2), [5]. Repéré à infolab.stanford.edu/~wangz/project/imsearch/review/JOUR/datta.pdf10.1145/1348246.1348248 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K. et Fei-Fei, L. (2009, juin). ImageNet : A Large-Scale Hierarchical Image Database. Communication présentée à la conférence « IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2009 », Miami, Floride. doi.org/10.1109/CVPRW.2009.520684810.1109/CVPR.2009.5206848 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Douze, M., Szlam, A., Hariharan, B. et Jégou, H. (2018, juin). Low-Shot Learning with Large-Scale Diffusion. Communication présentée à la conférence « IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2018 », Salt Lake City, Utah. Repéré à openaccess.thecvf.com/content_cvpr_2018/papers/Douze_Low-Shot_Learning_With_CVPR_2018_paper.pdf10.1109/CVPR.2018.00353 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Feaster, P. (2016, 31 octobre). Time-Based Image Averaging [Billet de blogue]. Repéré à griffonagedotcom.wordpress.com/2016/10/31/time-based-image-averaging.Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Freire, N., Robson, G., Howard, J. B., Manguinhas, H. et Isaac, A. (2017). Metadata Aggregation : Assessing the Application of IIIF and Sitemaps Within Cultural Heritage. Dans J. Kamps, G. Tsakonas, Y. Manolopoulos, L. Illiadis et I. Karydis (dir.), Research and Advanced Technology for Digital Libraries. TPDL 2017. doi.org/10.1007/978-3-319-67008-9_1810.1007/978-3-319-67008-9_18 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Ganascia, J.-G. (2017). Le mythe de la Singularité. Faut-il craindre l’intelligence artificielle ? Paris, France : Le Seuil.Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Ginosar, S., Rakelly, K., Sachs, S., Yin, B., Lee, C., Krähenbühl, P. et Efros, A. A. (2015). A Century of Portraits : A Visual Historical Record of American High School Yearbooks. IEEE Transactions on Computational Imaging, 3(3), 421-431.10.1109/TCI.2017.2699865 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Gordea, S. et Haskiya, D. (2017). Europeana DSI 2 – Access to Digital Resources of European Heritage. MS6.1 : Advanced Image Discovery Development Plan. Repéré à https://pro.europeana.eu/files/Europeana_Professional/Projects/Project_list/Europeana_DSI-2/Milestones/ms6.3-advanced-image-discovery-development-plan.pdfGoogle Scholar Rechercher cette référence bibliographique sur Google Scholar
- Gunthert, A. (2017, 10 juin). Le « visual turn » n’a pas eu lieu [Billet de blogue]. Repéré à imagesociale.fr/4603Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Karpathy, A. et Fei-Fei, L. (2017). Deep Visual-Semantic Alignments for Generating Image Descriptions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(4), 664-676. doi.org/10.1109/TPAMI.2016.259833910.1109/TPAMI.2016.2598339 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Lai, H. P., Visani, M., Boucher, A. et Ogier, J.-M. (2014). A New Interactive Semi-Supervised Clustering Model for Large Image Database Indexing. Pattern Recognition Letters, 37, 94-106. doi.org/10.1016/j.patrec.2013.06.01410.1016/j.patrec.2013.06.014 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Langlais, P.-C. (2017). Identifier les rubriques de presse ancienne avec du topic modeling. Repéré à numapresse.hypotheses.orgGoogle Scholar Rechercher cette référence bibliographique sur Google Scholar
- Moiraghi, E. (2018). Explorer des corpus d’images. L’IA au service du patrimoine. Repéré à bnf.hypotheses.org/2809Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Moreux, J.-P. (2016). Innovative Approaches of Historical Newspapers : Data Mining, Data Visualization, Semantic Enrichment. Facilitating Access for Various Profiles of Users. Repéré à http://library.ifla.org/2076/1/S21-2016-moreux-en.pdfGoogle Scholar Rechercher cette référence bibliographique sur Google Scholar
- Nottamkandath, A., Oosterman, J., Ceolin, D. et Fokkink, W. (2014). Automated Evaluation of Crowdsourced Annotations in the Cultural Heritage Domain. URSW’14 Proceedings of the 10th International Workshop on Uncertainty Reasoning for the Semantic Web, 1259, 25-36. Repéré à http://ceur-ws.org/Vol-1259/ursw2014_submission_5.pdfGoogle Scholar Rechercher cette référence bibliographique sur Google Scholar
- Pan, S. J. et Yang, Q. (2010). A Survey on Transfer Learning. IEEE Transactions on Knowledge and Data Engineering, 22(10), 1345-1359. doi.org/10.1109/TKDE.2009.19110.1109/TKDE.2009.191 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Picard, D., Gosselin, P.-H. et Gaspard, M.-C. (2015). Challenges in Content-Based Image Indexing of Cultural Heritage Collections. IEEE Signal Processing Magazine, 32(4), 95-102. Repéré à hal.archives-ouvertes.fr/hal-01164409/document10.1109/MSP.2015.2409557 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J. et Wojna, Z. (2016, juin). Rethinking the Inception Architecture for Computer Vision. Communication présentée à la conférence « 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) », Nevada, États-Unis. doi.org/10.1109/CVPR.2016.30810.1109/CVPR.2016.308 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Underwood, T. (2012, 7 avril). Topic Modeling Made Just Simple Enough [Billet de blogue]. Repéré à tedunderwood.com/2012/04/07/topic-modeling-made-just-simple-enoughGoogle Scholar Rechercher cette référence bibliographique sur Google Scholar
- Velcin, J., Soulages, J.-C., Kurpiel, S., Dias, L., Del Vecchio, M. et Aubrun, F. (2017). Fouille de textes pour une analyse comparée de l’information diffusée par les médias en ligne : une étude sur trois éditions du Huffington Post. Repéré à hal.archives-ouvertes.fr/hal-01571265/documentGoogle Scholar Rechercher cette référence bibliographique sur Google Scholar
- Viana, M., Nguyen, Q.-B., Smith, J. et Gabrani, M. (2017, novembre). Multimodal Classification of Document Embedded Images. Communication présentée à la conférence « 12th IAPR International Workshop, GREC 2017 », Kyoto, Japon. http://doi.org/10.1007/978-3-030-02284-6_410.1007/978-3-030-02284-6_4 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Wan, G. et Liu, Z. (2008). Content-Based Information Retrieval and Digital Libraries. Information Technology and Librairies, 27(1), 41-47. doi.org/10.6017/ital.v27i1.326210.6017/ital.v27i1.3262 Google Scholar Rechercher cette référence bibliographique sur Google Scholar
- Wang, K., Yin, Q., Wang, W., Wu, S. et Wang, L. (2016). A Comprehensive Survey on Cross-Modal Retrieval. Repéré à arxiv.org/pdf/1607.06215.pdfGoogle Scholar Rechercher cette référence bibliographique sur Google Scholar
- Welinder, P., Branson, S., Belongie, S. et Perona, P. (2010). The Multidimensional Wisdom of Crowds. NIPS’10 Proceedings of the 23rd International Conference on Neural Information Processing Systems, 2, 2424-2432.Google Scholar Rechercher cette référence bibliographique sur Google Scholar
