SkillAgentSearch skills...

siege

Work on one hard mathematics problem in rounds: each round a judge writes a few self-contained questions, independent workers answer them, and the judge keeps a ledger of what is proved, refuted and open, until the judge concludes or the rounds run out, and then a proof.md says plainly what is and i…

Install / Use

npx skills add anthropics/claude-plugins-official --skill siege

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

92/100

Supported Platforms

Universal

Tags

Our assessment of siege

siege scores 92/100 on our quality scale, 879th of 4,583 Development & Engineering skills we index (top 20%).

Its SKILL.md is 73 KB long, well organised into 11 sections and no code examples: long enough that it reads more like full documentation than a focused instruction file, which agents can find harder to follow.

With 37,497 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
21/30
Structure
13/20
Description
15/15
Adoption
19/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated today, so siege is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

siege compared with similar skills

All 4 of these similar skills score higher than siege; compare them before choosing.

SkillScoreStarsUpdatedFormat
siege (this skill)by anthropics9237.5ktodaySKILL.md
ai-job-searchby MadsLorentzen10045.2k2d agoCLAUDE.md
claude-howtoby luongnv8910041.8k7d agoCLAUDE.md
algorithmic-artby anthropics100177.9k15d agoSKILL.md
pptxby anthropics100177.9k15d agoSKILL.md

Frequently asked questions

How do I install siege?
Run npx skills add anthropics/claude-plugins-official --skill siege. The install tabs above show the steps for each supported agent.
Which AI agents does siege work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is siege safe to use?
It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is siege still maintained?
The repository was last updated today, so siege is actively maintained.

name: siege description: "Work on one hard mathematics problem in rounds: each round a judge writes a few self-contained questions, independent workers answer them, and the judge keeps a ledger of what is proved, refuted and open, until the judge concludes or the rounds run out, and then a proof.md says plainly what is and is not proved. A run takes hours and dozens of worker runs. Usage: /math-proof:siege [NAME=value settings] <the problem, stated in full, or the path of a file holding it>." argument-hint: "[NAME=value ...] [DIR=run-directory] <problem statement | problem-file>" disable-model-invocation: true disallowed-tools: WebSearch, WebFetch, AskUserQuestion allowed-tools: Read, Write, Edit, Glob, Grep, Agent, Bash(python3 ${CLAUDE_SKILL_DIR}/scripts/ledger.py *), Bash(python3 ${CLAUDE_PLUGIN_ROOT}/skills/siege/scripts/ledger.py *), Bash(python ${CLAUDE_SKILL_DIR}/scripts/ledger.py *), Bash(python ${CLAUDE_PLUGIN_ROOT}/skills/siege/scripts/ledger.py *), Bash(mkdir *), Bash(cp *), Bash(mv *), Bash(cat *), Bash(test *), Bash(ls *), Bash(wc *), Bash(cmp *), Bash(grep *), Bash(printf *), Bash(echo *)

math-proof: siege

The approach

The skill works on one problem in at most MAX_ROUNDS rounds. Each round a judge writes a few self-contained questions, at most WAVE of them before round ESC_ROUND. The steps below call them queries. Each question goes to a fresh worker that sees only that question and the problem statement. The judge reads the answers, rewrites the running summary, and adds to a ledger of claims marked PROVED, REFUTED or OPEN. One claim is the goal, set in round 1; the judge may set a new goal at a later round, and in particular may raise it to a more significant statement once it is proved (the proved goal then stays in the ledger as the fallback result, to which the judge returns if two waves bring the raised goal no nearer; when the goal is raised you tell the user and copy the proved result to DIR/result-so-far.md). When an answer holds a complete written proof or disproof of the goal, two more workers check that exact text line by line. Once both pass, the judge concludes or raises the goal. Before round MIN_ROUNDS, that is the only way to conclude. From round ESC_ROUND on, if no complete proof or disproof is in hand, a round may hold up to WAVE_DEEP questions for higher-effort workers. All but at most two of them aim at the one statement still missing from the proof. After the rounds end, the judge writes a self-contained proof.md, a worker checks it, and the judge finalizes it. Unless the goal the run ends on was both proved and checked twice during the rounds, there are also extra revision passes and a last wave of full proof attempts before finalizing. proof.md has a Status section that says plainly what is and is not proved. This section is only a summary: you do none of the mathematics yourself, and you follow the steps below exactly, launching a fresh sub-agent for every judge step and every worker.

You are the ORCHESTRATOR of this protocol. You do not do the mathematics yourself and you do not judge it: the judging is done by fresh math-proof-judge subagents, one per step, and the reasoning by fresh math-proof-worker or math-proof-worker-deep subagents, one per query (which of the two is fixed by the round number — see Rules). Your job is to run the steps below exactly, keep the files in order, act on the one mechanical verdict the bookkeeping script prints, and never skip, merge or reorder steps because the problem looks easy or hard. Do not read the workers' answer files yourself (use test -s to see whether one exists; the bookkeeping script, not you, decides whether it is finished); do not summarize mathematics in your own words anywhere a judge will read it — pass files, not paraphrases. Work unattended to the end: there is nobody to answer questions.

Arguments. The invoking message reads: $ARGUMENTS It gives the problem and, optionally, settings. Read it this way. Tokens of the form NAME=value at its start, where NAME is a word of two or more capital letters and underscores and value is a whole number (or, for DIR, a path, quoted if it contains spaces), are settings (the seven below, or DIR, the run directory; any other such NAME is an error, see Settings); remove them. If what remains is a single line that, taken as a whole — surrounding whitespace and one pair of enclosing quotation marks removed, backslash-escaped spaces read as spaces — is the path of an existing file (it may contain spaces; check with Read or Glob, not the shell), that file is the problem file; if no such file exists and what remains can only be a file path — a single line ending in .md, .txt or .tex, or a single token (no spaces once the quotes are removed) containing "/" or "" — tell the user in one sentence that no file exists at the absolute path you looked for (give it) and that the problem can instead be given in full as text after the command, and stop; otherwise everything that remains, to the end of the message, IS the problem statement, verbatim — mathematics, line breaks and all (MAX_ROUNDS=8 is a setting; "n=3", "N=pq", "AB=AC" and "f(x)=…" are mathematics). The run directory DIR defaults to ./math-proof-run under the current directory; use DIR's absolute path everywhere below. If the message holds neither a readable problem file nor any problem text, say so in one or two sentences — with the usage, /math-proof:siege [NAME=value …] <problem statement, or the path of a file holding it>, and that a stopped run is resumed by giving its original line again in the same directory — and stop. Settings (use these unless the invoking message overrides them by name): ESC_ROUND = 4 (the escalation round: from round ESC_ROUND on, if the loop is still running, the wave cap rises, every worker is a math-proof-worker-deep, and the plan brief carries its escalation clauses); WAVE = 4 queries per round at most in rounds before ESC_ROUND and WAVE_DEEP = 10 queries per round at most from round ESC_ROUND on; MAX_ROUNDS = 14; MIN_ROUNDS = 4 (the judge may conclude freely from round MIN_ROUNDS on, earlier only with an audited chain in the ledger — the script decides); REFINE_STEPS = 2; MAX_COMMIT = 9 (both apply to the FULL proof tail; a run whose final goal was certified during the rounds gets the SHORT tail — see "The proof tail"). The invoking message may override any of these seven by name, with tokens of the form NAME=value (for example MAX_ROUNDS=8 WAVE_DEEP=6) placed before the problem, at the start of the invoking message: use the values as given, and treat such a token there (NAME a word of two or more capitals and underscores, value a whole number) whose NAME is neither one of the seven nor DIR as an error — tell the user in one sentence that NAME is not a setting of /math-proof:siege, that the accepted names are DIR, ESC_ROUND, WAVE, WAVE_DEEP, MAX_ROUNDS, MIN_ROUNDS, REFINE_STEPS and MAX_COMMIT, and that if the token is part of the problem itself the problem can be given as a file path instead — and stop before creating anything. (A leading token whose left-hand side is a single letter or not all capitals, or whose right-hand side is not a whole number, and any "=" further inside the problem statement, is mathematics, not a setting.) Wherever a setting is named below, it means the value in force. The bookkeeping script is scripts/ledger.py in this skill's own folder: python3 ${CLAUDE_SKILL_DIR}/scripts/ledger.py (called SCRIPT below). If that placeholder was not filled in — the path before /scripts does not exist — use ${CLAUDE_PLUGIN_ROOT}/skills/siege/scripts/ledger.py, and if that one is unfilled too, Glob for **/skills/siege/scripts/ledger.py under ~/.claude and use its absolute path (if several match, the newest); in every case quote the script path in the command if it contains spaces. Before Setup, run SCRIPT once with no further arguments: it should print a line beginning "usage:". If instead Python runs but says it cannot open the script file, the path is wrong, not Python: resolve it again by the fallbacks above and run once more, and if no ledger.py can be found tell the user the plugin's files are not where expected (reinstall math-proof from /plugin) and stop. If instead the shell says python3 cannot be found, or anything else comes back that is not the usage line (on Windows, a reply that Python "was not found" and can be installed from the Store is this case), use python in place of python3 in SCRIPT from then on and run it once more; if that fails too, tell the user in one or two sentences that /math-proof:siege could not run its bookkeeping script — quote the command and the shell's reply — that it needs Python 3.7 or later on the PATH as python3 or python, and that the same line given again will work once that is fixed; then stop.

Setup

  1. If DIR/state.md already exists, this is a resume: check that the existing DIR/problem.md is the same problem you were given (compare the text, ignoring differences in whitespace and line endings; for a file, cmp) — if it differs, say in one sentence that DIR holds a run on a different problem and that DIR=<another directory> selects a fresh one, and stop; if it is the same and state.md records "phase: finished", that run is complete: say where DIR/proof.md is (and DIR/result-so-far.md, if it exists), that DIR=<another directory> starts a fresh run, and stop; otherwise Read DIR/problem.md in full, read state.md and resume from the phase it records instead of starting over, never redoing a step whose output files exist and never rewriting DIR/problem.md (if the invoking message gives settings that differ from those recorded in state.md, include one line saying the recorded ones govern this run in the same message as your next tool call — a notice, not a stop). If the phase recorded is a wave, start with SCRIPT answers on that wave's query stems (see the wave rule under Rules) and launch workers only for the queries the script does not report answered, each partial one with its {EARLIER} paragraph; that launch counts as the wave's first, so its one re-run still follows. A query already answered whose index line is missing gets the line "{Q} | answered | (finished before this session resumed)"; a query that already has an index line gets no second one — its new status is appended to that line as " | re-run: status | abstract" instead. Otherwise create DIR and DIR/judge/ (mkdir -p) and put the problem at DIR/problem.md: if it came as a file, copy that file there byte for byte with cp; if it came as text in the invoking message, Write exactly that text (nothing added, removed or reworded) to DIR/problem.md. Then, as your very next action, Read DIR/problem.md in full — you paste its text into every judge brief and every worker prompt (see Rules).
  2. Write DIR/state.md (protocol math-proof siege, the seven settings with the values in force — marking any the invoking message overrode —, "phase: round 1 plan, attempt 1"). On a resume, the values recorded in state.md govern.

The round loop — for r = 1, 2, …, MAX_ROUNDS

(a) Plan. Launch ONE math-proof-judge whose prompt is the PLAN BRIEF below with {r}, DIR and the round-dependent slots filled in (and, on a retry, the correction paragraph the script gave you added at the top). The slots: {CAP} is this round's cap on the number of queries — the Setting WAVE while r < ESC_ROUND, the Setting WAVE_DEEP when r ≥ ESC_ROUND — and the same number goes into the check command of (b); {MAX_ROUNDS} is the MAX_ROUNDS setting, written as a number; {HORIZON}, {CONCLUDE_RULE}, {ESC_SUMMARY} and {ESC_WAVE} are the texts given after the brief for this r (several are empty before round ESC_ROUND: an empty slot inserts nothing, not even a space or a blank line). This is attempt 1 of round r. (b) Check. First, if DIR/judge/screen_r{r-1}.md exists and round r−

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars37.5k
CategoryDevelopment
Updated15h ago
Forks4.2k

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions