Publication

PathologyBERT - Pre-trained Vs. A New Transformer Language Model for Pathology Domain.

Downloadable Content

Persistent URL
Last modified
  • 06/25/2025
Type of Material
Authors
    Thiago Santos, Emory UniversityAmara Tariq, Mayo Clinic, Phoenix, Arizona, USA.Susmita Das, Indian Institute of Technology (IIT), Centre of Excellence in Artificial Intelligence, Kharagpur, West Bengal, India.Kavyasree Vayalpati, Arizona State University, School of Computing and Augmented Intelligence, Tempe, Arizona, USA.Geoffrey Smith, Emory UniversityHari Trivedi, Emory UniversityImon Banerjee, Emory University
Language
  • English
Date
  • 2022
Publisher
  • AMIA Annual Symposium Proceesings Archive
Publication Version
Copyright Statement
  • ©2022 AMIA - All rights reserved.
Title of Journal or Parent Work
Volume
  • 2022
Start Page
  • 962
End Page
  • 971
Abstract
  • Pathology text mining is a challenging task given the reporting variability and constant new findings in cancer sub-type definitions. However, successful text mining of a large pathology database can play a critical role to advance 'big data' cancer research like similarity-based treatment selection, case identification, prognostication, surveillance, clinical trial screening, risk stratification, and many others. While there is a growing interest in developing language models for more specific clinical domains, no pathology-specific language space exist to support the rapid data-mining development in pathology space. In literature, a few approaches fine-tuned general transformer models on specialized corpora while maintaining the original tokenizer, but in fields requiring specialized terminology, these models often fail to perform adequately. We propose PathologyBERT - a pre-trained masked language model which was trained on 347,173 histopathology specimen reports and publicly released in the Huggingface1 repository2. Our comprehensive experiments demonstrate that pre-training of transformer model on pathology corpora yields performance improvements on Natural Language Understanding (NLU) and Breast Cancer Diagnose Classification when compared to nonspecific language models.
Keywords
Research Categories
  • Computer Science

Tools

Relations

In Collection:

Items