Skip to Main Content (Press Enter)

Logo UNIOR
  • ×
  • Home
  • Corsi
  • Insegnamenti
  • Persone
  • Strutture

UNIFIND
Logo UNIOR

|

UNIFIND

unior.it
  • ×
  • Home
  • Corsi
  • Insegnamenti
  • Persone
  • Strutture

Safeguarding Minors: A Study on Automatic Age-based Classification of Online Multimedia Content Using Text Transcripts

Tesi di Dottorato
Data di Pubblicazione:
2025
Abstract:
When we talk about online multimedia content, we refer to those elements that can enrich the user experience through the use of audio and visual material, such as videos, podcasts, live, and webinars. It often happens that videos, with content or language not
suitable for children, are seen by minors who can be attracted and influenced by what they see.
Automatic content classification would make the identification of content dangers quick
and applicable to a large amount of data. Furthermore, considering that all online multimedia content contains textual components, a linguistic approach is particularly useful.
The aim of this thesis is to explore the topic of automatic classification of texts related
to multimedia content (such as subtitles and transcriptions) to assess the appropriateness of such content for different age groups, with a focus on protecting minors from exposure to unsuitable material. This study investigates machine learning-based approaches, comparing various classifiers, as well as semantic and lexical methods. Furthermore, resources have been developed for both Italian and English languages to support this classification process.
Chapter 1 outlines the background and current state of multimedia content classification
research, providing an overview of what has already been done in the field of computational linguistics on this problem and identifying the elements that still need to be developed.
In Chapter 2, the methodology for carrying out the work is illustrated. Two corpora have
been created: one for Italian and one for English. The initial idea was to focus exclusively
on TED Talks, a particular type of content that covers a wide range of topics presented by
experts in various fields. These talks are generally of high quality and well-curated, making them ideal for linguistic and content analysis. In addition, TED Talks are available in many languages, facilitating the creation of multilingual corpora and comparative analysis in different languages. However, after some experiments, it was decided to add other texts to the corpora, specifically for children, extracted from YouTube Kids and other children’s channels, and specifically for adults, as they deal with themes related to sex, drugs, violence, and fear, or are characterized by vulgar language.
In order to conduct supervised classification, the corpus has been annotated following
the AGCOM (Authority for Communications Guarantees) Guidelines on the Classification of Audiovisual Works Intended for the Web and Video Games (AGCOM 2017).
High-specialized lexicons have been used to extract the linguistic indicators of the main
features used to rate multimedia online content. The dictionaries created are the dictionary of Violence, Drugs, and Bad language.
In addition, a study of emotions has been conducted, focusing particularly on emotions
that indicate negative sentiment such as Anger, Disgust, and Fear. The emotions have been analyzed thanks to the NRC Emotion Intensity Lexicon (Saif M Mohammad 2017): a list of words that have been assigned a value ranging from 0 to 1 in relation to the eight main emotions theorized by Plutchik (1984). The semantic information contained in the texts played a fundamental role in the classification.
All texts considered for the analysis were represented in a network based on their semantic proximity. Specifically, a distributional semantic matrix was created to extract
similarity values between words and to perform a semantic expansion of the more significant words in the transcripts. Once the texts were vectorized, the similarity values between text vectors were calculated using the Cosine Similarity algorithm. After comparing all texts and generating a large network of text-per-text edges, this
Tipologia CRIS:
5.13 Tesi di dottorato
Elenco autori:
Paone, Antonietta
Link alla scheda completa:
https://unora.unior.it/handle/11574/244940
Link al Full Text:
https://unora.unior.it//retrieve/handle/11574/244940/240837/Safeguarding%20Minors:%20A%20Study%20on%20Automatic%20Age-based%20Classification%20of%20Online%20Multimedia%20Content%20Using%20Text%20Transcripts.pdf
https://unora.unior.it//retrieve/handle/11574/244940/240931/Giudizio.pdf
  • Utilizzo dei cookie

Realizzato con VIVO | Designed by Cineca | 26.9.2.0