Publication Date:
2026
abstract:
This dataset contains a corpus of 329 annotated conversations between Italian university students (EFL learners) and AI-based chatbots (ChatGPT and Pi.AI), collected between May and December 2024 within the PRIN 2022 project *UNITE – Universally Inclusive Technologies to Practice English*.
The corpus includes interactions from three institutions (University of Bologna, University of Macerata, and University of Naples “L’Orientale”) and consists of learner-driven tasks such as small talk and role-play activities.
All conversations are annotated using a custom semantic tagset (DIS-TAG). DIS-TAG identifies lexical and discourse features related to mobility, sensory perception, and instructional language, enabling the analysis of normative discourse patterns in chatbot responses.
The corpus includes interactions from three institutions (University of Bologna, University of Macerata, and University of Naples “L’Orientale”) and consists of learner-driven tasks such as small talk and role-play activities.
All conversations are annotated using a custom semantic tagset (DIS-TAG). DIS-TAG identifies lexical and discourse features related to mobility, sensory perception, and instructional language, enabling the analysis of normative discourse patterns in chatbot responses.
Iris type:
5.10 Banca dati
Keywords:
annotated corpus, DISTAG, semantic annotation ai, chatgpt, Pi.AI,
List of contributors: