EUVIMEDEuropean Health Evidence
Uhr 10/10Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

The MRP Corpus: A Speech Dataset of Modern Received Pronunciation

Daniel Tweddle, Eugenio Donati, Julie Wall

Electronics · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Suitable speech data is crucial for the development of Mispronunciation Detection and Diagnosis (MDD) systems, which rely on examples of correct pronunciations and mispronunciations. This paper describes the creation of a novel dataset for British English MDD systems, with a focus on developing an accent-specific approach using deep learning methods. The dataset consists of over 7000 utterances from 29 speakers who were identified as using the ‘Modern Received Pronunciation’ (MRP) accent, and includes orthographic transcripts and automatically generated phone sequences. The MRP dataset aims to contribute to the development of both single-accent MDD systems and broader British English MDD. To observe the impact of single-accent training data, two identical CNN–LSTM models with Connectionist Temporal Classification (CTC) loss were trained to perform MDD. One model was trained on TIMIT, the other on our MRP data. When evaluated on L2 Arctic speech, which is annotated from an American English perspective, the TIMIT-trained model would normally be expected to perform better due to accent alignment. However, the MRP-trained model produced a lower Phoneme Error Rate (PER) and a higher rate of Correct Diagnosis. These exploratory results show that MRP can support MDD modelling, even when evaluated against an American English pronunciation target. The MRP dataset therefore offers an immediate resource for developing RP-focused pronunciation tools, a template for constructing comparable corpora across other UK accents, and an openly accessible dataset for future MDD and ASR research.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Daniel Tweddle, Eugenio Donati, Julie Wall
Quelle
Electronics
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2079-9292
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Daniel Tweddle, Eugenio Donati, Julie Wall (2026). The MRP Corpus: A Speech Dataset of Modern Received Pronunciation. Electronics. https://doi.org/10.3390/electronics15173873
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1