Vollständiger Abstract
Worum geht es in dieser Arbeit?
Suitable speech data is crucial for the development of Mispronunciation Detection and Diagnosis (MDD) systems, which rely on examples of correct pronunciations and mispronunciations. This paper describes the creation of a novel dataset for British English MDD systems, with a focus on developing an accent-specific approach using deep learning methods. The dataset consists of over 7000 utterances from 29 speakers who were identified as using the ‘Modern Received Pronunciation’ (MRP) accent, and includes orthographic transcripts and automatically generated phone sequences. The MRP dataset aims to contribute to the development of both single-accent MDD systems and broader British English MDD. To observe the impact of single-accent training data, two identical CNN–LSTM models with Connectionist Temporal Classification (CTC) loss were trained to perform MDD. One model was trained on TIMIT, the other on our MRP data. When evaluated on L2 Arctic speech, which is annotated from an American English perspective, the TIMIT-trained model would normally be expected to perform better due to accent alignment. However, the MRP-trained model produced a lower Phoneme Error Rate (PER) and a higher rate of Correct Diagnosis. These exploratory results show that MRP can support MDD modelling, even when evaluated against an American English pronunciation target. The MRP dataset therefore offers an immediate resource for developing RP-focused pronunciation tools, a template for constructing comparable corpora across other UK accents, and an openly accessible dataset for future MDD and ASR research.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Daniel Tweddle, Eugenio Donati, Julie Wall
- Quelle
- Electronics
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2079-9292
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Daniel Tweddle, Eugenio Donati, Julie Wall (2026). The MRP Corpus: A Speech Dataset of Modern Received Pronunciation. Electronics. https://doi.org/10.3390/electronics15173873
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1