Artigo Revisado por pares

CALBC SILVER STANDARD CORPUS

2010; Imperial College Press; Volume: 08; Issue: 01 Linguagem: Inglês

10.1142/s0219720010004562

ISSN

1757-6334

Autores

Dietrich Rebholz‐Schuhmann, Antonio Jimeno Yepes, Erik M. van Mulligen, Kang Ning, Jan A. Kors, David Milward, Peter Corbett, Ekaterina Buyko, Elena Beißwanger, Udo Hahn,

Tópico(s)

Bioinformatics and Genomic Networks

Resumo

The CALBC initiative aims to provide a large-scale biomedical text corpus that contains semantic annotations for named entities of different kinds. The generation of this corpus requires that the annotations from different automatic annotation systems be harmonized. In the first phase, the annotation systems from five participants (EMBL-EBI, EMC Rotterdam, NLM, JULIE Lab Jena, and Linguamatics) were gathered. All annotations were delivered in a common annotation format that included concept identifiers in the boundary assignments and that enabled comparison and alignment of the results. During the harmonization phase, the results produced from those different systems were integrated in a single harmonized corpus ("silver standard" corpus) by applying a voting scheme. We give an overview of the processed data and the principles of harmonization — formal boundary reconciliation and semantic matching of named entities. Finally, all submissions of the participants were evaluated against that silver standard corpus. We found that species and disease annotations are better standardized amongst the partners than the annotations of genes and proteins. The raw corpus is now available for additional named entity annotations. Parts of it will be made available later on for a public challenge. We expect that we can improve corpus building activities both in terms of the numbers of named entity classes being covered, as well as the size of the corpus in terms of annotated documents.

Referência(s)