用于自动结构法律文件的语料库

论文标题

用于自动结构法律文件的语料库

Corpus for Automatic Structuring of Legal Documents

论文作者

Kalamkar, Prathamesh, Tiwari, Aman, Agarwal, Astha, Karn, Saurabh, Gupta, Smita, Raghavan, Vivek, Modi, Ashutosh

论文摘要

在人口稠密的国家中，悬而未决的法律案件呈指数增长。需要开发处理和组织法律文件的技术。在本文中，我们介绍了一个新的语料库来构建法律文件。特别是，我们以英语介绍了一系列法律判断文件，这些文件被分为局部和连贯的部分。这些部分中的每一个都带有来自预定义角色列表的标签。我们开发了基线模型，以根据注释语料库自动预测法律文档中的修辞角色。此外，我们展示了修辞角色在提高总结和法律判断预测任务的绩效方面的应用。我们发布了语料库和基线模型代码以及纸张。

In populous countries, pending legal cases have been growing exponentially. There is a need for developing techniques for processing and organizing legal documents. In this paper, we introduce a new corpus for structuring legal documents. In particular, we introduce a corpus of legal judgment documents in English that are segmented into topical and coherent parts. Each of these parts is annotated with a label coming from a list of pre-defined Rhetorical Roles. We develop baseline models for automatically predicting rhetorical roles in a legal document based on the annotated corpus. Further, we show the application of rhetorical roles to improve performance on the tasks of summarization and legal judgment prediction. We release the corpus and baseline model code along with the paper.

下载PDF全文

下载文献需遵守相关版权规定

论文标题