Sweden’s National Library training AI to parse centuries of data

other

Summary

For the past 500 years, the National Library of Sweden has collected virtually every word published in Swedish, from priceless medieval manuscripts to present-day pizza menus. Thanks to a centuries-old law that requires a copy of everything published in Swedish to be submitted to the library — also known as Kungliga biblioteket, or KB — its collections span from the obvious to the obscure: books, newspapers, radio and TV broadcasts, internet content, Ph.D. dissertations, postcards, menus, and video games. “Between the exponential growth of digital data and ongoing work digitising physical collections that date back hundreds of years, we’ll never be finished adding to our collections.” Soon after KBLab was established in 2019, Börjeson saw the potential for training transformer language models on the library’s vast archives. “It’s one of the big advantages of NVIDIA for us, as a small lab that doesn’t have 50 engineers available to optimise AI training for every project.” In addition to transformer models that understand Swedish text, KBLab has an AI tool that transcribes sound to text, enabling the library to transcribe its vast collection of radio broadcasts so that researchers can search the audio records for specific content. “When you search the library’s databases for a specific term, we should be able to return results that include text, audio and video.” KBLab has partnered with researchers at the University of Gothenburg, who are developing downstream apps using the lab’s models to conduct linguistic research — including a project supporting the Swedish Academy’s work to modernise its data-driven techniques for creating Swedish dictionaries.

Classifications

industries
No industries detected
applications
AI & Machine learning

AskAI Classifications

Labels
No AI classifications detected

Linked Companies