paper-preflight
Pre-submission integrity gate for LaTeX papers: every reference verified against real scholarly records. No LLM guessing.
Install / Use
claude mcp add amos689 -- npx -y github:amos689/paper-preflightIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AI & Machine LearningSupported Platforms
Our assessment of paper-preflight
paper-preflight scores 79/100 on our quality scale, 797th of 963 AI & Machine Learning skills we index.
Its MCP Server is 28 KB long, well organised into 21 sections with 18 code examples: a thorough specification that gives an agent plenty to work with.
It has 43 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated yesterday, so paper-preflight is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 97/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
paper-preflight compared with similar skills
All 4 of these similar skills score higher than paper-preflight; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| paper-preflight (this skill)by amos689 | 79 | 43 | 1d ago | MCP Server |
| claude-memby thedotmack | 100 | 99.1k | 1d ago | CLAUDE.md |
| Agent-Reachby Panniantong | 100 | 95.3k | 2d ago | CLAUDE.md |
| Understand-Anythingby Egonex-AI | 100 | 85.8k | 1d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.9k | today | CLAUDE.md |
Frequently asked questions
- How do I install paper-preflight?
- Run
claude mcp add amos689 -- npx -y github:amos689/paper-preflight. The install tabs above show the steps for each supported agent. - Which AI agents does paper-preflight work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is paper-preflight safe to use?
- It is MIT-licensed and scores 97/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is paper-preflight still maintained?
- The repository was last updated yesterday, so paper-preflight is actively maintained.
Skill content
View source on GitHubCheck every reference of a LaTeX paper against real scholarly records before you submit.<br> No LLM guessing, no false accusations.
English · 简体中文 · Quick start · User guide · Releases · Feedback
</div>
Language models invent references, and copy-pasted BibTeX carries wrong years, wrong authors
and dead DOIs. paper-preflight reads your .tex and .bib files and asks Crossref, dblp,
arXiv, DataCite, PubMed and OpenAlex (and Semantic Scholar, if you have a key; Open Library for
books without a DOI; GitHub, PyPI, CRAN, Hugging Face and OpenML for software, models and
datasets) about every cited work:
- Does it exist?
- Does it match what you wrote?
- Has it been retracted?
- Has the preprint you cite been published since?
When it cannot tell, it says so instead of guessing.
Status: 0.x releases, on the way to a stable 1.0. False positives are the bugs we most want to hear about: please open an issue.
The repository's demo paper cites eleven works, several of them wrong on purpose. A real run, against the live sources:
$ paper-preflight check examples/demo-paper
paper-preflight 0.8.0 · main.tex · 12 entries, 12 cited keys
error CIT001 main.tex:31
Citation key 'nonexistent2023' is not defined in any bibliography file (1 use(s)).
error REF001 refs.bib:43
The doi of 'devlin2019bert' (10.1109/cvpr.2016.90) resolves to a different work in Crossref: "Deep Residual Learning for Image Recognition" (He et al., 2016).
error REF003 refs.bib:66
'lindqvist2024quantum' was not found in Crossref, dblp and Semantic Scholar, and every source responded. Check that the work exists and that its title is correct.
error REF004 refs.bib:73
'wakefield1998ileal' has been retracted (reported by Crossref, OpenAlex). Cite it only if the text discusses the retraction.
error CIT002 refs.bib:128
Entry key 'kingma2015adam' is already defined at line 47; BibTeX ignores this one.
warning REF015 refs.bib:31
'he2015residual' cites a preprint that has been published in CVPR (2016), DOI 10.1109/cvpr.2016.90. Cite the published version and keep the eprint field.
warning CIT004 refs.bib:37
Entries 'devlin2019bert' and 'he2016deep' look like the same work (same DOI).
warning REF013 refs.bib:51
'kingma2015adam' gives the year 2016, but dblp records 2014, 2015.
warning REF017 refs.bib:112
The doi of 'tacl2019example' contains LaTeX escapes: '10.1162/tacl\_a\_00276'. Write it as: 10.1162/tacl_a_00276
info REF005 refs.bib:73
'wakefield1998ileal' has a published correction (reported by Crossref).
info REF090 refs.bib:95
'zhou2016ml' could not be verified: non-Latin titles are not supported yet; grey literature without an identifier (book, report, software, web page).
info CIT003 refs.bib:116
Entry 'lecun1998gradient' is never cited.
References: 7 verified · 1 metadata mismatch · 1 identifier conflict · 1 not found · 1 cannot determine
5 error(s) · 4 warning(s) · 3 info
Each finding is backed by a record (or by every source answering "no"). The correct NeurIPS paper is verified through dblp even though Crossref only holds fake copies of it, the book without an identifier is confirmed by Open Library, and the Chinese one, whose script is not supported yet, is reported as "cannot determine" instead of "not found".
What it catches
| Rule | Finding |
|---|---|
| REF001 | The DOI or arXiv ID points to a different paper |
| REF002 | The DOI or arXiv ID does not exist |
| REF003 | The work was not found in any source, and every source answered |
| REF004 · REF005 | The work was retracted, or has an expression of concern or a correction |
| REF010–REF014 | Authors, title, year or venue differ from the real record |
| REF015 | A cited preprint has been formally published |
| REF016 | The registry has a DOI the entry lacks (offered as a safe fix) |
| REF017 | An identifier is written so that links break (10.1162/tacl\_a\_00276, …v1) |
| REF018 | A cited arXiv preprint was withdrawn by its authors |
| REF019 · REF021 | A linked repository, package or dataset does not exist; a linked web page is gone, with no archived copy |
| REF020 | A preprint is cited with a venue no catalogue of journals and conferences has |
| CIT001–CIT008 | Undefined, duplicate, unused or near-duplicate citation keys; broken .bib syntax |
| REF090 | Cannot determine, always with the reason (source unavailable, grey literature, …) |
Every rule has a page: what it checks, when it can be wrong, and what to
do. paper-preflight explain REF003 prints it in the terminal.
How accurate is it?
Four measurements, all against the live sources: hallucinations found in published papers, the bibliographies of real papers, a head-to-head with published tools, and a public benchmark.
On hallucinations that got past peer review
GPTZero published 151 hallucinated references it found in NeurIPS 2025 papers and ICLR 2026 submissions, each confirmed by its staff. Pasted as plain text, as the papers printed them:
| References | Flagged | Cannot determine | Verified | |---|---|---|---| | 151 | 142 (94%) | 9 | 0 |
- None of them is verified. The 9 left undecided are web pages and blog posts, a workshop
no source indexes, real titles given with invented authors where several works share the
title, and two references too garbled to search. Each is listed with its reason in
evals/results/gptzero.md. - GPTZero's own tool found these, so they are the hallucinations a search can find; recall on every kind of hallucination is lower (see HALLMARK below).
On real papers
Weeks of arXiv papers (20 each, cs, stat, q-bio, quant-ph and astro-ph), chosen mechanically, each list fixed before the changes it measures, run once, with every warning and error reviewed by hand:
| Papers first submitted | Version | References | Flags | Real problems | False positives | Unclear | False positives per 100 references | |---|---|---|---|---|---|---|---| | 2026-08-19..25 | 0.5.0 | 962 | 53 | 41 | 10 | 2 | 1.0 | | 2026-08-26..09-01 | 0.5.1 | 780 | 101 | 81 | 19 | 1 | 2.4 | | 2026-09-02..08 | 0.5.2 | 975 | 48 | 31 | 13 | 4 | 1.3 | | 2026-09-09..15 | 0.5.3 | 1,067 | 98 | 83 | 14 | 1 | 1.3 | | 2026-09-16..22 | 0.6.0 candidate | 868 | 89 | 76 | 13 | 0 | 1.5 | | 2026-09-23..29 | 0.6.0 | 786 | 48 | 42 | 6 | 0 | 0.8 | | 2026-09-30..10-06 | 0.7.0 | 1,014 | 89 | 79 | 7 | 3 | 0.7 | | 2026-06-24..30, an earlier week | 0.8.0 | 1,127 | 143 | 132 | 10 | 1 | 0.9 |
- 0.8.0: one false alarm every two papers (56 references on average), against 132 real problems: 93 cited preprints since published, 17 identifiers written as links, wrong authors or given names in 10 entries (in four, most given names are invented), 6 wrong years, 2 misquoted titles, a wrong venue, a DOI that does not exist and an arXiv ID of another paper. Six false alarms are registries' own writing ('Ueber' for 'Über', a title that lost its '3/4'), three are author fields written by hand ('and 324 others'), one an ITU-R report no source indexes. This week was sampled from before the first one, never used until then, so that 0.8.0 need not wait for next week's papers; checked months after it was written, it may read a little better than a fresh week, which follows as a control.
- 0.7.0: about one false alarm every three papers, against 79 real problems. The false alarms are three articles a registry dates only by their online appearance, a book against its online edition, two symbols a registry writes its own way, and a paper's code offered as its published version; all are fixed in 0.8.0.
- The two weeks before (0.6.0): 1.5, then 0.8. 8 of the first week's 13 false alarms were registries' own errors (a short author list, misspelt names, an HTML entity, a typo, a wrong year); 5 are fixed in 0.6.0.
- One week was over our target of 1.5 (2.4); 11 of its 19 false alarms are fixed in 0.5.2.
- In the first week, one false alarm every two papers (48 references on average), against 41 real problems: 16 errors in the entries (wrong years, given names and titles, a missing first author, one paper's title with another's authors, identifiers written so that links break) and 25 cited preprints that have since been published.
- The false alarms are mostly registry records with errors of their own (author lists that stop short, a garbled symbol, an English given name) and collaboration names in author lists ("MAGPI Team"); two come from a title cited without its last words.
- Seven earlier batches of 20 papers were used to find false positives, each first measured
as it came out (0.1.0: 4.5 per 100 references; 0.1.1: 2.3; 0.1.2 before its last fixes: 3.0;
0.1.2: 1.7; 0.2.1: 1.9; 0.3.0: 1.9; 0.4.0: 1.2). Details in
evals/README.md.
Next to other tools
Badalova & Mayr (2026) checked 104 references by hand and published what five tools flagged. On the same references, with their labels:
| Tool | Precision [95% CI] | Recall | False flags per 100 correct references | |---|---|---|---| | CheckIfExist | 47.7% [36.0%, 59.6%] | 93.9% | 47.9 | | HalluCiteChecker | 47.4% [32.5%, 62.7%] | 54.5% | 28.2 | | Hallucinator | 50.9% [38.3%, 63.4%] | 87.9% | 39.4 | | HalRef | 31.2% [21.9%, 42.2%] | 72.7% | 74.6 | | RefChecker | 47.1% [35.7%, 58.8%] | 97.0% | 50.7 | | paper-preflight | 72.5% [57.2%, 83.9%] | 87.9% | 15.5 |
The sample is small, so the intervals are wide. Some flags count as false here because the study
labels a reference correct when the work exists: five of paper-preflight's flags on such
references point at real errors (a wrong author, a broken DOI). Two causes of false flags found
in this data were fixed, and four names the study's CSV garbled were restored, before the run
above; the first run measured 62.8%. See
evals/results/badalova-mayr.md.
On a benchma
Truncated for display — read the full file on GitHub.
Related Skills
claude-mem
99.1kPersistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Agent-Reach
95.3kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Understand-Anything
85.8kGraphs that teach > graphs that impress. Turn any code into an interactive knowledge graph you can explore, search, and ask questions about. Works with Claude Code, Codex, Cursor, Copilot, Gemini CLI, and more.
headroom
74.9kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
