A two-headed BERT for the legal domain, doing Named Entity Recognition and Relation Extraction in one pass over HUDOC court decisions. To our knowledge a first, with 96.25% NER accuracy and 95.96% RE accuracy.
Automatically Identifying Legal Concepts in Court Decision Texts
Abstract
While there are many repositories available online for legal documents, its utilization, that is, identifying and extracting relevant information from this huge stack of unstructured information, remains a challenge in the legal domain. It becomes essential that this challenge is addressed to capitalize on the available legal resources. Thus, our objective in this project is to identify and extract the entities and relations, such that important legal concepts are captured with the help of Natural Language Processing. Since deep learning approaches have been used, an annotated dataset needed to be created using raw text from the HUDOC database. This has been performed for both entities and relations, by combining manual, rule-based, and semi-supervised approaches. The task of identifying entities and extracting relations has been achieved via two steps, Named Entity Recognition and Relation Extraction. In this project, pre-existing models on the legal domain have been created as benchmark models. After that, a BERT model with two output heads has been implemented to perform both NER and RE, which is a first in the legal domain (according to our findings). The Two-Headed BERT Model performance was significantly higher than the benchmark models performance, with an accuracy of 96.25% and 95.96% for Named Entity Recognition and Relation Extraction respectively.