SkillAgentSearch skills...

self-managed-context

This skill should be used when a model gets read-write control over its own live context window instead of a harness-scheduled compaction policy: the context exposed as an editable file the model rewrites with code tools, model-driven eviction and in-place updates, the harness invariants that keep s…

Install / Use

npx skills add muratcankoylan/Agent-Skills-for-Context-Engineering --skill self-managed-context

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

96/100

Supported Platforms

Universal

Tags

Our assessment of self-managed-context

self-managed-context scores 96/100 on our quality scale, 217th of 4,570 Development & Engineering skills we index (top 5%).

Its SKILL.md is 24 KB long, well organised into 26 sections with 3 code examples: a thorough specification that gives an agent plenty to work with.

With 18,026 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
18/20
Description
15/15
Adoption
18/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 27 days ago, so self-managed-context is actively maintained.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

self-managed-context compared with similar skills

All 4 of these similar skills score higher than self-managed-context; compare them before choosing.

SkillScoreStarsUpdatedFormat
self-managed-context (this skill)by muratcankoylan9618.0k27d agoSKILL.md
ai-job-searchby MadsLorentzen10045.3k2d agoCLAUDE.md
claude-howtoby luongnv8910041.8k8d agoCLAUDE.md
algorithmic-artby anthropics100177.9k15d agoSKILL.md
pptxby anthropics100177.9k15d agoSKILL.md

Frequently asked questions

How do I install self-managed-context?
Run npx skills add muratcankoylan/Agent-Skills-for-Context-Engineering --skill self-managed-context. The install tabs above show the steps for each supported agent.
Which AI agents does self-managed-context work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is self-managed-context safe to use?
It is MIT-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is self-managed-context still maintained?
The repository was last updated 27 days ago, so self-managed-context is actively maintained.

name: self-managed-context description: "This skill should be used when a model gets read-write control over its own live context window instead of a harness-scheduled compaction policy: the context exposed as an editable file the model rewrites with code tools, model-driven eviction and in-place updates, the harness invariants that keep self-editing safe (pinned prefix, edit gate, edit receipts, budget readouts, rollback on overflow), the prefix-cache cost of mid-context edits, and steering or training the model's own context-editing strategy. Route note content and fixed-threshold summarization to context-compression, cache-stable layout under harness control to context-optimization, file offloading to filesystem-context, and loop governance to self-improvement-loops."

Self-Managed Context

This skill covers agents that manage their own context window: the model, not the harness, decides what stays in the live context, what is compacted, what is evicted, and what is rewritten in place. The reference implementation is Context Language Models (CLMs), which mirror the editable part of the conversation into a file that the model edits with ordinary code tools, then re-parse that file into the next prompt. Applied zero-shot to existing models on a shared agent backbone, this matched or beat harness-scheduled and action-based context management on accuracy and compute across deep-research, terminal-coding, and multi-hour repository-optimization tasks (claim-self-managed-context-zero-shot-results, claim-self-managed-context-long-horizon-results).

The controlling trade-off: model control buys adaptivity (verbatim retention, surgical in-place updates, eviction on demand) and creates three problems the harness must absorb. Edits break prefix caching, so a badly placed edit can cost more compute than the tokens it frees. Models do not know how full their context is. And a model-writable context is a persistence channel for whatever the model writes into it, including instructions. Most of this skill is the harness contract that makes the first property worth the other three.

When to Activate

Activate this skill when:

  • Giving a model write access to its own live context instead of compacting on a fixed schedule
  • Designing a context-as-a-file protocol: turn headers, parse-back, pinned prefix, edit receipts
  • An agent must keep exact values, update state in place, or evict offloaded material, and fixed summarization loses or duplicates it
  • Pricing mid-context edits under prefix caching, or evaluating suffix cache reuse
  • Budget readouts, tiered nudges, or rollback for an agent that times its own compaction
  • Steering, evolving, or training a model's context-editing strategy
  • Diagnosing a self-editing agent that never edits, wipes useful state, livelocks after overflow, or carries injected instructions in its own notes

Do not activate this skill for adjacent work owned by other skills:

  • Replacement-note content and fixed-threshold summarization: context-compression.
  • Cache-stable layout and observation masking under harness control: context-optimization.
  • File offloading while the live context stays append-only: filesystem-context.
  • Latent KV transfer between agents: latent-briefing.
  • Cross-session stores: memory-systems.
  • Loop governance and acceptance gates: self-improvement-loops. This skill supplies the context-editing signal; that skill governs the loop.

Core Concepts

The Control Spectrum

Context-management designs differ in who decides the transition from one context to the next. Append-only agents extend the context each step; a self-managing agent produces the whole next context as a function of the current one.

| Level | Who decides | Examples | Characteristic failure | | --- | --- | --- | --- | | Harness-scheduled | Harness, at a threshold or every turn | Threshold summarization, per-turn state rewrite | Wrong timing; summaries lose or invent verbatim state | | Action-based | Model chooses when; harness defines what | Self-compaction tools, offload-and-retrieve tools, context folding | Strategy bounded by the action set; offload without eviction | | Model-controlled | Model chooses when and what, with general tools | Context as an editable file | Edit cache cost, budget blindness, persisting self-written instructions |

A diagnostic built to isolate context management from reasoning (verbatim retention, in-place board updates, offload-and-evict) found no fixed strategy perfect even on simple synthetic tasks (claim-self-managed-context-contextbench-pilot). Each failure maps to a missing capability: summaries cannot hold exact values, append-only designs re-emit full state for every small update, and tool-based offloading cannot remove the original from the window.

Context as an Editable File

The implementation needs no new model capability:

  1. Split the context into a pinned prefix (system prompt and task) and an editable region (every later turn).
  2. Before each command, render the editable region to a file, each turn preceded by a header such as [[CTX_TURN 7 role=tool]].
  3. The model edits the file with the same shell and code tools it uses for the task, locating text with code (header regexes, unique anchors) instead of retyping it.
  4. After the command, re-read the file, parse it tolerantly back into messages, and replace the live context.

The load-bearing choice is general tools, not context tools. There is no summarize action and no eviction API; the model can delete, rewrite, merge, or annotate. Observed behaviors went beyond any predefined action set: a model-invented notes role, a self-defined compaction helper reused across a run, and an in-context subagent scoreboard maintained through in-place edits at small context size (claim-self-managed-context-emergent-behaviors).

The Harness Owns the Invariants

Unrestricted edits are safe only because the safety contract lives outside the editable region.

| Invariant | Mechanism | Failure prevented | | --- | --- | --- | | Pinned prefix | System and task messages are never rendered into the file; parse-back re-pins them from originals | The model rewriting its own instructions or task | | Role folding | Any parsed role except assistant becomes user; stray text becomes a user note | Model-minted system authority | | Edit gate | fit: growth allowed only if the result stays under budget; shrink: any growth is rejected | Edits that duplicate instead of replace | | Receipts | One line whenever a command changes or targets the file: edit applied, applied but grew, rejected with the rule, failed, or matched nothing | Silent no-op edits the model believes succeeded | | Free edit turns | A turn whose edit is applied, that prints nothing, and that exits 0 costs no task step | Housekeeping suppressed by step budgets | | Structural legality | Parse-back drops emptied turns and merges adjacent same-role turns; rollback never orphans a tool call from its result | Malformed message lists rejected by the API |

Edit Position Sets the Cost

Prefix caching reuses computation only up to the first changed token. An edit forces everything after it to be re-processed, so its cost scales with the text that follows it, not with the size of the edit. In an illustrative turn, an edit at the start of the context cost several times an append-only turn (claim-self-managed-context-edit-cost). Three operating rules follow, and the reference system prompt states all three:

  • Batch. One large compaction beats many small edits; each edit pays for its own tail.
  • Mind the tail. Do not compact a small early region under a long, still-useful tail. Wait and compact head and tail together, unless the limit is near.
  • Be generous in the replacement. The tail is re-read regardless, so a detailed summary is nearly free relative to the re-prefill it already triggered.

Measure cost as prefix-reuse compute (decode, prefill, and re-prefill across the trajectory), not as context length. A shorter context with frequent early edits can cost more than a longer append-only one.

Budget Awareness Is Not Native

Models estimate their own context length poorly at long lengths, often emitting the same bucketed values regardless of actual length; token-count hints near the estimation point fix much of the error (claim-self-managed-context-length-awareness). A self-managing agent therefore needs a deterministic readout supplied by the harness:

  • Count tokens locally with a fixed tokenizer, calibrated to the server's reported prompt tokens with a clamped ratio so a gateway anomaly cannot corrupt the budget.
  • Show the current size on every tool result.
  • Escalate nudges by tier. Low fill: informational only, with no how-to, because compaction instructions this early trigger premature wholesale deletion. Mid fill: finish the unit in flight, then tidy once. Near the limit: compact settled spans without wiping them.
  • Size the urgent band from the largest recent tool output, not a fixed percentage, so one large observation cannot jump past the warning straight into overflow.
  • On overflow, roll back the newest turns to free just enough room for one compaction, name the commands whose output blew the budget, and deepen the rollback only after repeated ignored rollbacks (claim-self-managed-context-harness-defaults).

Strategy Is Steerable, Evolvable, and Trainable

Moving context management into model behavior makes it learnable at three levels of cost:

| Lever | Mechanism | Evidence | | --- | --- | --- | | Instruction | One sentence sets a compaction threshold, semantic boundaries, or backup-before-edit | claim-self-managed-context-steering | | Skill evolution | A proposer rewrites the context-management skill from contrastive rollouts; a paired-SE gate on a held-out split accepts or rejects | claim-self-managed-context-skill-evolution | | RL | Stepwise GRPO with a success-gated efficiency advantage applied only to context-edit tokens | claim-self-managed-context-rl |

The common rule across the two optimization levers: task success is the primary signal and cost only re-ranks successful attempts. Rewarding edit frequency or removed volume invites reward hacking through unnecessary edits that discard information and break prefix reuse.

Detailed Topics

Suffix Cache Reuse

When the serving stack is under your control, Suffix Cache Reuse (SCR) relocates cached states for spans that survive an edit instead of re-prefilling them: keys are re-rotated to new positions, and hybrid linear-attention layers fork their recurrent state. Surviving spans keep slightly stale states that encode the old prefix; the reference implementation caps relocation at a small number of spans per edit to bound that approximation. SCR matched standard serving accuracy at a fraction of its compute (claim-self-managed-context-scr-savings). Most of its savings came from chat templates that strip prior reasoning blocks, which forces re-prefill of everything after the first stripped block even for agents that never edit their context. Check whether your serving path strips reasoning before attributing re-prefill cost to model edits.

Capability Dependence

Self-management delegates a decision, so gains scale with the model's ability to make it. A smaller model edited less often, skipped editing entirely on many tasks, and ran near the limit, while the larger model kept substantial headroom (claim-self-managed-context-capability-gap). Untrained small models can trail a harness-scheduled baseline until RL closes the gap (claim-self-managed-context-rl). For weak models, a hybrid works: keep a harness-scheduled fallback beneath the self-managed layer, or steer with explicit thresholds.

Editable Context as a Persistence Channel

Anything the model writes into its context is read by every later turn as context. A frontier lab reported a model writing jailbreak-style instructions into its own compactio

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars18.0k
CategoryDevelopment
Updated27d ago
Forks1.5k

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions