The corpus is available in Kielipankki - the Language Bank of Finland (lat.csc.fi, http://urn.fi/urn:nbn:fi:lb-1001100133; download at http://urn.fi/urn:nbn:fi:lb-2016021201).
Regarding the LAT version of the corpus (http://urn.fi/urn:nbn:fi:lb-1001100133): at the moment it contains only one subcorpus (FBC-1). However the entire corpus is being converted for the LAT system and also the FBC-2 subcorpus will be published in LAT as soon as the conversion process is completed.
Apply for access rights: https://lbr.csc.fi
The Finnish Broadcast Corpus is divided into two main parts: FBC-1 and FBC-2.
The Finnish Broadcast Corpus 1, FBC-1 contains 65 radio and tv recordings broadcast by YLE – the Finnish Broadcasting Company during the year 2003. Parts of the audio and video material have been annotated either manually or automatically in various levels: e.g., utterance (orthographic transcript), word, phone. FBC-1 was compiled under an initiative called Integrated Resources for Speech Technology and Spoken Language Research in Finland, funded by the Academy of Finland. It is CSC’s first multimodal corpus.
Details of the size of FBC-2 are being updated.
The material in the FBC-1 represents four categories:
* Radio monologues
- broadcast telegraph news (24 × 3 minutes, Nov. 2003)
- broadcast lectures of the week (8 × 14 minutes).
* Radio dialogues
- unfinished recordings of the Moninaisuusfoorumi event (5 × 1h).
* TV monologues
- broadcast main news read by Arvi Lind ja Eeva Polttila (15 × 30 minutes, September - November 2003), including the very last news telecast by Arvi Lind on October 15, 2003
* TV dialogues
- broadcast Aamu-TV programs (13 × ca. 12 minutes, 2003).
* WAV audio format
* HQ_Pure audio format (44,1–48 KHz) (supported by the Puh-Editor, which is now obsolete)
* HQ_Pure audio format (16 KHz) (supported by the Puh-Editor, which is now obsolete)
* MPEG2 video
The purpose of the resource use must be outlined in a research plan.