content-refinement-agent
Step 5 of the PaperOrchestra pipeline (arXiv:2604.05018). Iteratively refine drafts/paper.tex by simulating peer review and applying targeted revisions, with strict accept/revert halt rules, deterministic 0-100 decision bands (Accept/Minor/Major/Reject) that drive a target-met early stop, and a Devi…
Install / Use
npx skills add Ar9av/PaperOrchestra --skill content-refinement-agentInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Our assessment of content-refinement-agent
content-refinement-agent scores 92/100 on our quality scale, 956th of 2,855 Automation skills we index (top 34%).
Its SKILL.md is 19 KB long, well organised into 20 sections with 12 code examples: a thorough specification that gives an agent plenty to work with.
It has 664 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 12 days ago, so content-refinement-agent is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-04. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
content-refinement-agent compared with similar skills
All 4 of these similar skills score higher than content-refinement-agent; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| content-refinement-agent (this skill)by Ar9av | 92 | 664 | 12d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 90.1k | 18d ago | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.6k | 1d ago | MCP Server |
| rufloby ruvnet | 100 | 73.8k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 11d ago | SKILL.md |
Frequently asked questions
- How do I install content-refinement-agent?
- Run
npx skills add Ar9av/PaperOrchestra --skill content-refinement-agent. The install tabs above show the steps for each supported agent. - Which AI agents does content-refinement-agent work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is content-refinement-agent safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is content-refinement-agent still maintained?
- The repository was last updated 12 days ago, so content-refinement-agent is actively maintained.
Skill content
View source on GitHubname: content-refinement-agent description: Step 5 of the PaperOrchestra pipeline (arXiv:2604.05018). Iteratively refine drafts/paper.tex by simulating peer review and applying targeted revisions, with strict accept/revert halt rules, deterministic 0-100 decision bands (Accept/Minor/Major/Reject) that drive a target-met early stop, and a Devil's Advocate concession-threshold guard that blocks acceptance on unresolved critical findings. Maintains a worklog and snapshots each iteration so revert is real, not symbolic. TRIGGER when the orchestrator delegates Step 5 or when the user asks to "refine the draft", "iterate on the paper", or "run peer review on this paper". data_access_level: verified_only
Content Refinement Agent (Step 5)
Faithful implementation of the Content Refinement Agent from PaperOrchestra (Song et al., 2026, arXiv:2604.05018, §4 Step 5, App. F.1 pp. 49–51).
Cost: ~5–7 LLM calls (App. B), typically ~3 refinement iterations, each consisting of one reviewer call and one revision call.
The paper highlights this step as one of the largest contributors to overall quality: refinement alone accounts for +19% (CVPR) and +22% (ICLR) absolute acceptance-rate improvement (Fig. 4). Get this step right.
Inputs
workspace/drafts/paper.tex— output of Step 4workspace/inputs/conference_guidelines.mdworkspace/inputs/experimental_log.md— used as ground truth for the hallucination checkworkspace/citation_pool.json/workspace/refs.bib— the allowed bibliography
Outputs
workspace/refinement/iter1/,iter2/,iter3/— per-iteration snapshots containingpaper.tex,paper.pdf,review.json,score.jsonworkspace/refinement/worklog.json— append-only history of decisionsworkspace/final/paper.texandworkspace/final/paper.pdf— copy of the best accepted snapshot
The refinement loop
prev_score = score(paper.tex) # baseline from initial draft
snapshot iter0/
for iter in 1..ITER_CAP (default 3):
1. simulate_review(paper.tex) → review.json
(uses `references/reviewer-rubric.md` rubric)
2. apply_revision(paper.tex, review.json) → new_paper.tex
(uses verbatim Refinement Agent prompt at `references/prompt.md`)
3. snapshot iter<N>/ with new_paper.tex, review.json
latexmk -pdf new_paper.tex → iter<N>/paper.pdf
4. score(new_paper.tex) → curr_score
5. decide via score_delta.py:
- if curr.overall > prev.overall: ACCEPT
- elif curr.overall == prev.overall and net_subaxis ≥0: ACCEPT
- else: REVERT
6. apply_worklog.py to append the decision
7. if REVERT or no actionable weaknesses or iter == ITER_CAP: HALT
paper.tex ← new_paper.tex (only on ACCEPT)
prev_score ← curr_score
cp <best iter>/paper.tex → workspace/final/paper.tex
The "best" snapshot at HALT is the one with the highest accepted overall score. On a REVERT halt, the best is the iteration immediately before the revert.
Step-by-step
0. Pre-refinement integrity gate
Before snapshotting or scoring the initial draft, run two gates in order:
Gate A — AI failure modes (load references/ai-failure-modes.md, runs once):
Load references/ai-failure-modes.md (which points to skills/shared/ai_failure_modes.md).
Run all 7 checks against the draft and the inputs. This gate runs once only,
at the start of iteration 1.
- CONFIRMED failure → write HALT entry to worklog.json, report to user, stop.
- SUSPECTED failure → add WARNING comment to paper.tex, log in worklog.json, continue.
- No failures → proceed.
Gate B — Claim-evidence provenance (runs once, WARN gate):
python skills/paper-orchestra/scripts/claim_evidence_gate.py \
--paper workspace/drafts/paper.tex \
--log workspace/inputs/experimental_log.md \
--out workspace/claim_evidence_report.json \
--out-md workspace/claim_evidence_map.md
The gate sorts every number in the draft into supported (the value is in
experimental_log.md), attributed (the sentence carries a citation or a
prior-work cue), or needs evidence (neither). Only the third category is a
finding. See references/claim-evidence-map.md.
Exit 0 → PASS, proceed normally.
Exit 1 → WARN: unsupported numeric claims found. Log in worklog.json as:
{gate: "claim_evidence", status: "WARN", unsupported_count: N, report: "workspace/claim_evidence_report.json"}
Pass the needs evidence rows of workspace/claim_evidence_map.md to the
revision agent in Step 3 as an additional instruction: "The following values
appear in the paper but cannot be corroborated in experimental_log.md and
carry no citation — restate them from logged values, attribute them, or remove
the claim. Do not weaken the sentence into vagueness to make the number
defensible." A row that survives two iterations should be deleted rather than
reworded again.
Do NOT halt on Gate B warnings; the revision agent will address them.
Gate C — Reverse outline (runs once, advisory):
python skills/content-refinement-agent/scripts/reverse_outline.py \
--paper workspace/drafts/paper.tex \
--out workspace/reverse_outline.md \
--json workspace/reverse_outline.json
Strips the draft to one line per paragraph — its topic sentence — and flags
paragraphs with no topic sentence, two messages, or a citation dump. Read
workspace/reverse_outline.md before the first reviewer call and pass it into
the reviewer call as the input for the Logical Flow axis. Structural
findings enter the revision agenda as reorder / merge / split / cut
instructions; sentence-level rewriting cannot fix a sequencing problem, and
iterations spent polishing a misordered section still count against the budget.
See references/reverse-outline.md.
Gate D — Read research brief (every run, no exit code):
If workspace/research_brief.md exists, read it before all reviewer calls.
Pass the "Sections where evidence was thin" list from §4 as additional
context to the Devil's Advocate reviewer. This surfaces the highest-risk
sections for CRITICAL scrutiny.
0b. Snapshot the initial draft
python skills/content-refinement-agent/scripts/snapshot.py \
--src workspace/drafts/paper.tex \
--dst workspace/refinement/iter0/
This creates iter0/paper.tex. Then compile to iter0/paper.pdf:
cd workspace/refinement/iter0/ && latexmk -pdf -interaction=nonstopmode paper.tex
Score it (see Step 1 below) → iter0/score.json.
1. Simulate peer review
For each iteration N starting from 1:
Writing quality pre-check (start of every iteration): Load
references/writing-quality-check.md and run the 5-category checklist
(Categories A–E) against the current draft. Note violations and add them to
the revision agenda.
Update critique memory before the reviewer call (iter N ≥ 2 only — skip for iter 1):
python skills/content-refinement-agent/scripts/update_critique_memory.py \
--worklog workspace/refinement/worklog.json \
--review workspace/refinement/iter<N-1>/review.json \
--iter <N> \
--out workspace/refinement/critique_memory.json
This produces critique_memory.json with focus_on (persistent unresolved
issues) and do_not_reflag (already-resolved issues). Inject both lists into
the reviewer system prompt verbatim:
CRITIQUE MEMORY — you must honour this before reviewing:
FOCUS ON (flagged in prior iterations, not yet resolved — prioritise these):
<critique_memory.focus_on items, one per line>
DO NOT RE-FLAG (already addressed in prior iterations):
<critique_memory.do_not_reflag items, one per line>
This prevents the reviewer from re-discovering already-fixed issues and from missing genuinely stuck problems.
Regenerate the reverse outline for the current draft (reverse_outline.py,
Gate C) and include workspace/reverse_outline.md in the reviewer's user
message. The reviewer scores Logical Flow against the topic-sentence sequence
rather than against its impression of the prose, which is what makes that axis
move for structural reasons instead of stylistic ones.
Load references/reviewer-rubric.md as the system prompt for the simulated
reviewer call. The reviewer reads iter<N-1>/paper.pdf (or paper.tex if
your host LLM lacks PDF input) and produces a JSON of strengths,
weaknesses, questions, and per-axis scores.
The rubric is structured to mimic AgentReview (Jin et al., 2024) — the paper's chosen evaluator. We ship a faithful rubric in the references directory; the host agent's LLM does the actual reviewing.
Devil's Advocate reviewer: One simulated reviewer must be designated the DA
following references/da-reviewer.md. The DA challenges core claims from first
principles (causal overclaiming, ablation coverage, baseline fairness,
generalization claims, novelty inflation) rather than surface polish. If the DA
issues a CRITICAL finding that remains unaddressed after all reviewers weigh in,
that finding blocks the "refinement accepted" decision regardless of rubric scores.
Log DA CRITICAL findings in worklog.json: {da_critical: true, finding: "..."}.
Record the DA's per-round findings and concession decisions in
workspace/refinement/da_concessions.json (schema in references/da-reviewer.md)
and enforce the concession-threshold protocol deterministically — this stops the
simulated DA from sycophantically caving:
python skills/content-refinement-agent/scripts/concession_guard.py \
--log workspace/refinement/da_concessions.json \
--out workspace/refinement/iter<N>/da_guard.json
# exit 0 = clear; exit 1 = standing CRITICAL → force REVERT this iteration;
# exit 2 = a concession was rejected (caving/consecutive) → DA must restate;
# exit 3 = schema error.
The guard rejects any concession made at rebuttal_score < 4 or in a round
immediately following another concession, and restores the affected finding to
"standing". A standing CRITICAL (exit 1) overrides an ACCEPT into a REVERT.
Save to workspace/refinement/iter<N>/review.json.
2. Score the draft
The reviewer call produces both qualitative feedback and a per-axis score:
{
"axis_scores": {
"scientific_depth": {"score": 65, "justification": "..."},
"technical_execution": {"score": 70, "justification": "..."},
"logical_flow": {"score": 60, "justification": "..."},
"writing_clarity": {"score": 55, "justification": "..."},
"evidence_presentation":{"score": 72, "justification": "..."},
"academic_style": {"score": 68, "justification": "..."}
},
"overall_score": 64.5,
"decision_band": "Major Revision",
"strengths": [...],
"weaknesses": [...],
"questions": [...]
}
Save to iter<N>/score.json. (Combined with review.json if your host
emits one document; the schemas overlap.)
decision_band is derived deterministically from overall_score — Accept
(≥80) / Minor Revision (65–79) / Major Revision (50–64) / Reject (<50). Fill it
in with python skills/content-refinement-agent/scripts/decision_band.py --score-json iter<N>/score.json rather than by hand, so it can never disagree
with the number. The bands drive the target-met halt in Step 5.
3. Apply revision
Load the verbatim Content Refinement Agent prompt at references/prompt.md.
Prepend the Anti-Leakage Prompt. Inputs:
paper.tex— current draftpaper.pdf— compiled PDF (multimodal context if available)conference_guidelines.mdexperimental_log.md— ground truth for numeric claimsworklog.json— history of previous changescitation_pool.json— the allowed bibliographyreviewer_feedback— the JSON from Step 1
The prompt instructs the model to address weaknesses, integrate question answers, and emit two output blocks:
- A worklog JSON
{addressed_weaknesses[], integrated_answers[], actions_taken[]} - The full revi
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
90.1kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Scrapling
85.6k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
ruflo
73.8k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
