Libraries, University of Nebraska-Lincoln

Copyright, Fair Use, Scholarly Communication, etc.

Peer Review Analyze: A Novel Benchmark Resource for Computational Analysis of Peer Reviews

Tirthankar Ghosal, Charles University
Sandeep Kumar, Indian Institute of Technology Patna
Prabhat Kumar Bharti, Indian Institute of Technology Patna
Asif Ekbal, Indian Institute of Technology Patna

Document Type

Article

Date of this Version

1-27-2022

Citation

PLoS ONE (January 27, 2022) 17(1): e0259238

https://doi.org/10.1371/journal.pone.0259238

Dataset and codes are available at https://www.iitp.ac.in/~ai-nlp-ml/resources.html#Peer-Review-Anazlyze or https://github.com/Tirthankar-Ghosal/Peer-Review-Analyze-1.0

Comments

License: Creative Commons Attribution 4.0 (CC BY 4.0)

Abstract

Peer Review is at the heart of scholarly communications and the cornerstone of scientific publishing. However, academia often criticizes the peer review system as non-transparent, biased, arbitrary, a flawed process at the heart of science, leading to researchers arguing with its reliability and quality. These problems could also be due to the lack of studies with the peer-review texts for various proprietary and confidentiality clauses. Peer review texts could serve as a rich source of Natural Language Processing (NLP) research on understanding the scholarly communication landscape, and thereby build systems towards mitigating those pertinent problems. In this work, we present a first of its kind multi-layered dataset of 1199 open peer review texts manually annotated at the sentence level (*17k sentences) across the four layers, viz. Paper Section Correspondence, Paper Aspect Category, Review Functionality, and Review Significance. Given a text written by the reviewer, we annotate: to which sections (e.g., Methodology, Experiments, etc.), what aspects (e.g., Originality/Novelty, Empirical/Theoretical Soundness, etc.) of the paper does the review text correspond to, what is the role played by the review text (e.g., appreciation, criticism, summary, etc.), and the importance of the review statement (major, minor, general) within the review. We also annotate the sentiment of the reviewer (positive, negative, neutral) for the first two layers to judge the reviewer’s perspective on the different sections and aspects of the paper. We further introduce four novel tasks with this dataset, which could serve as an indicator of the exhaustiveness of a peer review and can be a step towards the automatic judgment of review quality. We also present baseline experiments and results for the different tasks for further investigations. We believe our dataset would provide a benchmark experimental testbed for automated systems to leverage on current NLP state-of-the-art techniques to address different issues with peer review quality, thereby ushering increased transparency and trust on the holy grail of scientific research validation.

Download

Included in

Intellectual Property Law Commons, Scholarly Communication Commons, Scholarly Publishing Commons

COinS

Libraries, University of Nebraska-Lincoln

Copyright, Fair Use, Scholarly Communication, etc.

Peer Review Analyze: A Novel Benchmark Resource for Computational Analysis of Peer Reviews

Document Type

Date of this Version

Citation

Comments

Abstract

Included in

Search

Browse

Author Corner

Links

Libraries, University of Nebraska-Lincoln

Copyright, Fair Use, Scholarly Communication, etc.

Peer Review Analyze: A Novel Benchmark Resource for Computational Analysis of Peer Reviews

Authors

Document Type

Date of this Version

Citation

Comments

Abstract

Included in

Share

Search

Browse

Author Corner

Links