...
首页> 外文期刊>BMC Medical Informatics and Decision Making >Developing a cardiovascular disease risk factor annotated corpus of Chinese electronic medical records
【24h】

Developing a cardiovascular disease risk factor annotated corpus of Chinese electronic medical records

机译:建立中国电子病历中标注心血管疾病危险因素的语料库

获取原文
           

摘要

Background Cardiovascular disease (CVD) has become the leading cause of death in China, and most of the cases can be prevented by controlling risk factors. The goal of this study was to build a corpus of CVD risk factor annotations based on Chinese electronic medical records (CEMRs). This corpus is intended to be used to develop a risk factor information extraction system that, in turn, can be applied as a foundation for the further study of the progress of risk factors and CVD. Results We designed a light annotation task to capture CVD risk factors with indicators, temporal attributes and assertions that were explicitly or implicitly displayed in the records. The task included: 1) preparing data; 2) creating guidelines for capturing annotations (these were created with the help of clinicians); 3) proposing an annotation method including building the guidelines draft, training the annotators and updating the guidelines, and corpus construction. Meanwhile, we proposed some creative annotation guidelines: (1) the under-threshold medical examination values were annotated for our purpose of studying the progress of risk factors and CVD; (2) possible and negative risk factors were concerned for the same reason, and we created assertions for annotations; (3) we added four temporal attributes to CVD risk factors in CEMRs for constructing long term variations. Then, a risk factor annotated corpus based on de-identified discharge summaries and progress notes from 600 patients was developed. Built with the help of clinicians, this corpus has an inter-annotator agreement (IAA) F1-measure of 0.968, indicating a high reliability. Conclusion To the best of our knowledge, this is the first annotated corpus concerning CVD risk factors in CEMRs and the guidelines for capturing CVD risk factor annotations from CEMRs were proposed. The obtained document-level annotations can be applied in future studies to monitor risk factors and CVD over the long term.
机译:背景技术心血管疾病(CVD)已成为中国的主要死亡原因,通过控制危险因素可以预防大多数情况。这项研究的目的是基于中国电子病历(CEMR)建立CVD危险因素注释的语料库。该语料库旨在用于开发风险因素信息提取系统,进而可以用作进一步研究风险因素和CVD的基础。结果我们设计了一个简短的注释任务,以捕获具有明确或隐含显示在记录中的指标,时间属性和主张的CVD危险因素。任务包括:1)准备数据; 2)创建捕获注释的准则(这些准则是在临床医生的帮助下创建的); 3)提出注释方法,包括构建指南草案,培训注释者和更新指南以及语料库构建。同时,我们提出了一些创造性的注释准则:(1)为了研究危险因素和CVD的进展,对阈值以下的医学检查值进行了注释; (2)由于相同的原因,可能的和负面的风险因素也受到关注,因此我们为注释创建了断言; (3)我们为CEMR中的CVD危险因素添加了四个时间属性,以构建长期变化。然后,开发了基于不确定的出院摘要和600名患者的病历注释的风险因素注释语料库。该语料库在临床医生的帮助下构建,其注释者间协议(IAA)F 1 度量为0.968,表明其可靠性很高。结论据我们所知,这是关于CEMR中CVD危险因素的第一个注释语料,并提出了从CEMR中捕获CVD危险因素注释的指南。获得的文档级注释可以在将来的研究中应用,以长期监控风险因素和CVD。

著录项

相似文献

  • 外文文献
  • 中文文献
  • 专利
获取原文

客服邮箱:kefu@zhangqiaokeyan.com

京公网安备:11010802029741号 ICP备案号:京ICP备15016152号-6 六维联合信息科技 (北京) 有限公司©版权所有
  • 客服微信

  • 服务号