Wharton School researchers recently found that authors’ rankings of their own artificial intelligence papers may help identify influential work overlooked by traditional peer review.
The study, published on Aug. 24 in Nature Computational Science, analyzed authors’ comparative rankings of their own papers among those with multiple submissions to the 2023 International Conference on Machine Learning. The team embarked upon this study to determine an additional measure for identifying high-impact AI research given that the field has grown at a rate much larger than the number of experienced peer reviewers in the same areas.
Authors were forced to be objective because they could not give all of their work a perfect score, and the researchers found that higher-ranked papers received more academic citations than lower-ranked papers.
The team included statistics and data science professors Weijie Su and Bingxin Zhao, and fourth-year Applied Mathematics and Computational Science doctoral student Buxin Su.
“We use self-rankings to identify papers that need additional attention and guide Area Chairs toward those papers,” Su wrote in a statement.
In the experiment, authors with multiple submissions were asked to rank their papers by perceived scientific quality shortly after the submission deadline and before reviews were released. Reviewers then evaluated the papers without seeing the rankings.
The study collected rankings from 1,342 authors covering 2,592 submissions. After matching the responses with citation data, the researchers analyzed a final sample of 797 authors and 1,527 unique papers.
Papers ranked highest by their authors received an average of twice as many citations over the following 16 months as those ranked lowest. The pattern held among both accepted and rejected papers.
RELATED:
New Penn Med AI tool aims to improve clinical decision-making
Penn researchers find nearly half of observational social science studies overstate causality
Of the 22 papers in the sample that received more than 150 citations, 17 — or about 77% — were ranked first by at least one of their authors. The researchers also found that author rankings predicted future citation counts more accurately than peer-review scores.
The study used citations as a measure of scientific impact, though citation counts do not necessarily reflect a paper’s quality. Factors including the popularity of a research topic, the timing of a paper’s public release, and an author’s professional visibility can affect how often it is cited.
The researchers proposed using differences between author rankings and reviewer scores to identify papers that might require closer examination, but they did not propose using self-rankings to directly determine whether a paper should be accepted.
Under the system, author rankings are converted into scores that can be compared with reviewers’ initial evaluations. The 10% of papers with the largest evaluation differences are flagged for an area chair, who recommends whether it should be accepted.
Area chairs can only see the size of the discrepancy, not whether the author-based score was higher or lower than the review score. They can then examine the paper and its reviews more closely or recruit additional reviewers.
According to Su, the system is designed to reduce incentives for authors to manipulate their rankings. Placing a weaker paper first could draw additional scrutiny to the submission rather than improve its likelihood of acceptance.
ICML incorporated the approach into its 2026 review process and conducted a randomized experiment to measure whether displaying the discrepancy categories affected how area chairs handled papers.
Among flagged papers, area chairs who could see the categories wrote approximately 71% more comment text per paper than those who could not, according to Su. They also communicated more frequently with individual reviewers, and the team observed more subsequent discussion between reviewers and authors.
The experiment measured participation in the review process but did not establish whether the system led to better acceptance decisions.
“The broader implication is that author self-ranking can help conferences direct limited reviewing attention toward cases that deserve closer scrutiny,” Su wrote.






