Cross-species gene normalization by species inference

Chih Hsuan Wei, Hung Yu Kao

研究成果: Article同行評審

45 引文 斯高帕斯(Scopus)

摘要

Background: To access and utilize the rich information contained in the biomedical literature, the ability to recognize and normalize gene mentions referenced in the literature is crucial. In this paper, we focus on improvements to the accuracy of gene normalization in cases where species information is not provided. Gene names are often ambiguous, in that they can refer to the genes of many species. Therefore, gene normalization is a difficult challenge.Methods: We define " gene normalization" as a series of tasks involving several issues, including gene name recognition, species assignation and species-specific gene normalization. We propose an integrated method, GenNorm, consisting of three modules to handle the issues of this task. Every issue can affect overall performance, though the most important is species assignation. Clearly, correct identification of the species can decrease the ambiguity of orthologous genes.Results: In experiments, the proposed model attained the top-1 threshold average precision (TAP-k) scores of 0.3297 (k=5), 0.3538 (k=10), and 0.3535 (k=20) when tested against 50 articles that had been selected for their difficulty and the most divergent results from pooled team submissions. In the silver-standard-507 evaluation, our TAP-k scores are 0.4591 for k=5, 10, and 20 and were ranked 2nd, 2nd, and 3rd respectively.Availability: A web service and input, output formats of GenNorm are available at http://ikmbio.csie.ncku.edu.tw/GN/.

原文English
文章編號S5
期刊BMC Bioinformatics
12
發行號SUPPL. 8
DOIs
出版狀態Published - 2011 十月 3

All Science Journal Classification (ASJC) codes

  • 結構生物學
  • 生物化學
  • 分子生物學
  • 電腦科學應用
  • 應用數學

指紋

深入研究「Cross-species gene normalization by species inference」主題。共同形成了獨特的指紋。

引用此