SkillAgentSearch skills...

paper-preflight

Pre-submission integrity gate for LaTeX papers: every reference verified against real scholarly records. No LLM guessing.

Install / Use

claude mcp add amos689 -- npx -y github:amos689/paper-preflight

If the server publishes to npm under a different name, use that package instead — check the repo README.

About this skill
🔌

MCP Server

Model Context Protocol server

Quality Score

79/100

Supported Platforms

Claude Code
Claude Desktop

Our assessment of paper-preflight

paper-preflight scores 79/100 on our quality scale, 797th of 963 AI & Machine Learning skills we index.

Its MCP Server is 28 KB long, well organised into 21 sections with 18 code examples: a thorough specification that gives an agent plenty to work with.

It has 43 GitHub stars, so there is little community track record yet; judge it on its content.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
7/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated yesterday, so paper-preflight is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

paper-preflight compared with similar skills

All 4 of these similar skills score higher than paper-preflight; compare them before choosing.

SkillScoreStarsUpdatedFormat
paper-preflight (this skill)by amos68979431d agoMCP Server
claude-memby thedotmack10099.1k1d agoCLAUDE.md
Agent-Reachby Panniantong10095.3k2d agoCLAUDE.md
Understand-Anythingby Egonex-AI10085.8k1d agoCLAUDE.md
headroomby headroomlabs-ai10074.9ktodayCLAUDE.md

Frequently asked questions

How do I install paper-preflight?
Run claude mcp add amos689 -- npx -y github:amos689/paper-preflight. The install tabs above show the steps for each supported agent.
Which AI agents does paper-preflight work with?
It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
Is paper-preflight safe to use?
It is MIT-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is paper-preflight still maintained?
The repository was last updated yesterday, so paper-preflight is actively maintained.
<!-- mcp-name: io.github.amos689/paper-preflight --> <div align="center">

paper-preflight

Check every reference of a LaTeX paper against real scholarly records before you submit.<br> No LLM guessing, no false accusations.

CI PyPI version MIT license paper-preflight MCP server on Glama

Python 3.11 to 3.14 Input: LaTeX, BibTeX and PDF Checked against six scholarly databases No LLM in the verdicts Read-only MCP tools

Tested on Windows, Linux and macOS English and Simplified Chinese

English · 简体中文 · Quick start · User guide · Releases · Feedback

</div>

paper-preflight checking the demo paper: errors for an undefined citation key, a DOI that belongs to another paper, a reference no source knows, a retracted paper and a duplicate entry key; warnings for a published preprint, two entries for the same work, a wrong year and a LaTeX-escaped DOI

Language models invent references, and copy-pasted BibTeX carries wrong years, wrong authors and dead DOIs. paper-preflight reads your .tex and .bib files and asks Crossref, dblp, arXiv, DataCite, PubMed and OpenAlex (and Semantic Scholar, if you have a key; Open Library for books without a DOI; GitHub, PyPI, CRAN, Hugging Face and OpenML for software, models and datasets) about every cited work:

  • Does it exist?
  • Does it match what you wrote?
  • Has it been retracted?
  • Has the preprint you cite been published since?

When it cannot tell, it says so instead of guessing.

Status: 0.x releases, on the way to a stable 1.0. False positives are the bugs we most want to hear about: please open an issue.

The repository's demo paper cites eleven works, several of them wrong on purpose. A real run, against the live sources:

$ paper-preflight check examples/demo-paper
paper-preflight 0.8.0 · main.tex · 12 entries, 12 cited keys

error   CIT001 main.tex:31
    Citation key 'nonexistent2023' is not defined in any bibliography file (1 use(s)).
error   REF001 refs.bib:43
    The doi of 'devlin2019bert' (10.1109/cvpr.2016.90) resolves to a different work in Crossref: "Deep Residual Learning for Image Recognition" (He et al., 2016).
error   REF003 refs.bib:66
    'lindqvist2024quantum' was not found in Crossref, dblp and Semantic Scholar, and every source responded. Check that the work exists and that its title is correct.
error   REF004 refs.bib:73
    'wakefield1998ileal' has been retracted (reported by Crossref, OpenAlex). Cite it only if the text discusses the retraction.
error   CIT002 refs.bib:128
    Entry key 'kingma2015adam' is already defined at line 47; BibTeX ignores this one.
warning REF015 refs.bib:31
    'he2015residual' cites a preprint that has been published in CVPR (2016), DOI 10.1109/cvpr.2016.90. Cite the published version and keep the eprint field.
warning CIT004 refs.bib:37
    Entries 'devlin2019bert' and 'he2016deep' look like the same work (same DOI).
warning REF013 refs.bib:51
    'kingma2015adam' gives the year 2016, but dblp records 2014, 2015.
warning REF017 refs.bib:112
    The doi of 'tacl2019example' contains LaTeX escapes: '10.1162/tacl\_a\_00276'. Write it as: 10.1162/tacl_a_00276
info    REF005 refs.bib:73
    'wakefield1998ileal' has a published correction (reported by Crossref).
info    REF090 refs.bib:95
    'zhou2016ml' could not be verified: non-Latin titles are not supported yet; grey literature without an identifier (book, report, software, web page).
info    CIT003 refs.bib:116
    Entry 'lecun1998gradient' is never cited.

References: 7 verified · 1 metadata mismatch · 1 identifier conflict · 1 not found · 1 cannot determine
5 error(s) · 4 warning(s) · 3 info

Each finding is backed by a record (or by every source answering "no"). The correct NeurIPS paper is verified through dblp even though Crossref only holds fake copies of it, the book without an identifier is confirmed by Open Library, and the Chinese one, whose script is not supported yet, is reported as "cannot determine" instead of "not found".

What it catches

| Rule | Finding | |---|---| | REF001 | The DOI or arXiv ID points to a different paper | | REF002 | The DOI or arXiv ID does not exist | | REF003 | The work was not found in any source, and every source answered | | REF004 · REF005 | The work was retracted, or has an expression of concern or a correction | | REF010–REF014 | Authors, title, year or venue differ from the real record | | REF015 | A cited preprint has been formally published | | REF016 | The registry has a DOI the entry lacks (offered as a safe fix) | | REF017 | An identifier is written so that links break (10.1162/tacl\_a\_00276, …v1) | | REF018 | A cited arXiv preprint was withdrawn by its authors | | REF019 · REF021 | A linked repository, package or dataset does not exist; a linked web page is gone, with no archived copy | | REF020 | A preprint is cited with a venue no catalogue of journals and conferences has | | CIT001–CIT008 | Undefined, duplicate, unused or near-duplicate citation keys; broken .bib syntax | | REF090 | Cannot determine, always with the reason (source unavailable, grey literature, …) |

Every rule has a page: what it checks, when it can be wrong, and what to do. paper-preflight explain REF003 prints it in the terminal.

How accurate is it?

Four measurements, all against the live sources: hallucinations found in published papers, the bibliographies of real papers, a head-to-head with published tools, and a public benchmark.

On hallucinations that got past peer review

GPTZero published 151 hallucinated references it found in NeurIPS 2025 papers and ICLR 2026 submissions, each confirmed by its staff. Pasted as plain text, as the papers printed them:

| References | Flagged | Cannot determine | Verified | |---|---|---|---| | 151 | 142 (94%) | 9 | 0 |

  • None of them is verified. The 9 left undecided are web pages and blog posts, a workshop no source indexes, real titles given with invented authors where several works share the title, and two references too garbled to search. Each is listed with its reason in evals/results/gptzero.md.
  • GPTZero's own tool found these, so they are the hallucinations a search can find; recall on every kind of hallucination is lower (see HALLMARK below).

On real papers

Weeks of arXiv papers (20 each, cs, stat, q-bio, quant-ph and astro-ph), chosen mechanically, each list fixed before the changes it measures, run once, with every warning and error reviewed by hand:

| Papers first submitted | Version | References | Flags | Real problems | False positives | Unclear | False positives per 100 references | |---|---|---|---|---|---|---|---| | 2026-08-19..25 | 0.5.0 | 962 | 53 | 41 | 10 | 2 | 1.0 | | 2026-08-26..09-01 | 0.5.1 | 780 | 101 | 81 | 19 | 1 | 2.4 | | 2026-09-02..08 | 0.5.2 | 975 | 48 | 31 | 13 | 4 | 1.3 | | 2026-09-09..15 | 0.5.3 | 1,067 | 98 | 83 | 14 | 1 | 1.3 | | 2026-09-16..22 | 0.6.0 candidate | 868 | 89 | 76 | 13 | 0 | 1.5 | | 2026-09-23..29 | 0.6.0 | 786 | 48 | 42 | 6 | 0 | 0.8 | | 2026-09-30..10-06 | 0.7.0 | 1,014 | 89 | 79 | 7 | 3 | 0.7 | | 2026-06-24..30, an earlier week | 0.8.0 | 1,127 | 143 | 132 | 10 | 1 | 0.9 |

  • 0.8.0: one false alarm every two papers (56 references on average), against 132 real problems: 93 cited preprints since published, 17 identifiers written as links, wrong authors or given names in 10 entries (in four, most given names are invented), 6 wrong years, 2 misquoted titles, a wrong venue, a DOI that does not exist and an arXiv ID of another paper. Six false alarms are registries' own writing ('Ueber' for 'Über', a title that lost its '3/4'), three are author fields written by hand ('and 324 others'), one an ITU-R report no source indexes. This week was sampled from before the first one, never used until then, so that 0.8.0 need not wait for next week's papers; checked months after it was written, it may read a little better than a fresh week, which follows as a control.
  • 0.7.0: about one false alarm every three papers, against 79 real problems. The false alarms are three articles a registry dates only by their online appearance, a book against its online edition, two symbols a registry writes its own way, and a paper's code offered as its published version; all are fixed in 0.8.0.
  • The two weeks before (0.6.0): 1.5, then 0.8. 8 of the first week's 13 false alarms were registries' own errors (a short author list, misspelt names, an HTML entity, a typo, a wrong year); 5 are fixed in 0.6.0.
  • One week was over our target of 1.5 (2.4); 11 of its 19 false alarms are fixed in 0.5.2.
  • In the first week, one false alarm every two papers (48 references on average), against 41 real problems: 16 errors in the entries (wrong years, given names and titles, a missing first author, one paper's title with another's authors, identifiers written so that links break) and 25 cited preprints that have since been published.
  • The false alarms are mostly registry records with errors of their own (author lists that stop short, a garbled symbol, an English given name) and collaboration names in author lists ("MAGPI Team"); two come from a title cited without its last words.
  • Seven earlier batches of 20 papers were used to find false positives, each first measured as it came out (0.1.0: 4.5 per 100 references; 0.1.1: 2.3; 0.1.2 before its last fixes: 3.0; 0.1.2: 1.7; 0.2.1: 1.9; 0.3.0: 1.9; 0.4.0: 1.2). Details in evals/README.md.

Next to other tools

Badalova & Mayr (2026) checked 104 references by hand and published what five tools flagged. On the same references, with their labels:

| Tool | Precision [95% CI] | Recall | False flags per 100 correct references | |---|---|---|---| | CheckIfExist | 47.7% [36.0%, 59.6%] | 93.9% | 47.9 | | HalluCiteChecker | 47.4% [32.5%, 62.7%] | 54.5% | 28.2 | | Hallucinator | 50.9% [38.3%, 63.4%] | 87.9% | 39.4 | | HalRef | 31.2% [21.9%, 42.2%] | 72.7% | 74.6 | | RefChecker | 47.1% [35.7%, 58.8%] | 97.0% | 50.7 | | paper-preflight | 72.5% [57.2%, 83.9%] | 87.9% | 15.5 |

The sample is small, so the intervals are wide. Some flags count as false here because the study labels a reference correct when the work exists: five of paper-preflight's flags on such references point at real errors (a wrong author, a broken DOI). Two causes of false flags found in this data were fixed, and four names the study's CSV garbled were restored, before the run above; the first run measured 62.8%. See evals/results/badalova-mayr.md.

On a benchma

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars43
CategoryAI
Updated1d ago
Forks7

Languages

Python

Trust signals

97/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 info