SkillAgentSearch skills...

oswright

MCP server for Windows desktop automation. Incremental perception instead of a screenshot per step: measured 16.9x fewer tokens per action than the alternatives, with the benchmarks to check it.

Install / Use

claude mcp add Ask-812 -- npx -y github:Ask-812/oswright

If the server publishes to npm under a different name, use that package instead — check the repo README.

About this skill
🔌

MCP Server

Model Context Protocol server

Quality Score

83/100

Category

Automation

Supported Platforms

Claude Code
Claude Desktop

Our assessment of oswright

oswright scores 83/100 on our quality scale, 2218th of 2,891 Automation skills we index.

Its MCP Server is 30 KB long, well organised into 26 sections with 17 code examples: a thorough specification that gives an agent plenty to work with.

It has 3 GitHub stars, so there is little community track record yet; judge it on its content.

Substance
30/30
Structure
20/20
Description
15/15
Adoption
3/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 23 days ago, so oswright is actively maintained.
  • Our last check on 2026-09-18 found the source still online.
  • It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

Safety scan

No issues found

Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.

AI review by kimi-k2.7-code on 2026-10-02. Automated pattern scan on 2026-10-02. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.

oswright compared with similar skills

All 4 of these similar skills score higher than oswright; compare them before choosing.

SkillScoreStarsUpdatedFormat
oswright (this skill)by Ask-81283323d agoMCP Server
Agent-Reachby Panniantong10095.5k2d agoCLAUDE.md
headroomby headroomlabs-ai10074.9ktodayCLAUDE.md
CowAgentby zhayujie10047.3ktodayCLAUDE.md
Scraplingby D4Vinci10086.7ktodayMCP Server

Frequently asked questions

How do I install oswright?
Run claude mcp add Ask-812 -- npx -y github:Ask-812/oswright. The install tabs above show the steps for each supported agent.
Which AI agents does oswright work with?
It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
Is oswright safe to use?
Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is oswright still maintained?
The repository was last updated 23 days ago, so oswright is actively maintained.

OSWright

PyPI Tests Python License

Desktop automation for AI agents, without paying for a screenshot every step.

mcp-name: io.github.Ask-812/oswright

An MCP server that lets an LLM drive real desktop applications — the desktop equivalent of Playwright MCP. It keeps a model of the screen between actions and re-reads only the parts that changed, so the same work costs an order of magnitude fewer tokens.

OSWright transcribing an invoice into an expense form

Eight fields read off an invoice and typed into an expense form, verified by the application itself. Same task, same result, 7.4× less context than returning a screenshot after every action. Every number on screen is measured during the run — regenerate the whole thing with python benchmarks/record_demo.py.

Why this exists

Most GUI agents re-perceive the entire screen on every step: screenshot, OCR, hand the model an image, repeat. Measured on a live desktop, the median observation changes 0.012% of the screen's pixels. Re-reading everything does far more work than the change warrants, and charges ~2,800 image tokens whether anything happened or not.

OSWright asks the compositor what changed, rescans only that, and answers element lookups from the cheapest source that can. The claims below are measured on this machine and reproducible from benchmarks/ — including the ones that did not come out in its favour.

Key Features

  • Cross-platform. Windows (Win32 API), Linux (pynput/X11), macOS (pynput/Quartz).
  • Accessibility tree. Find elements deterministically by role and name via Windows UI Automation — 100% accurate, instant, no model needed.
  • Fast OCR. Windows OCR (built-in, instant) with EasyOCR fallback for Linux/macOS. Results are cached automatically.
  • Lightweight on Windows. No PyTorch download — Windows uses the built-in OCR engine, so a full install is a few MB rather than a few GB.
  • Image matching. Locates elements by template image via OpenCV.
  • Window management. List, focus, minimize, close, and screenshot specific windows.
  • Screenshot diffing. Detect when the screen changes with wait_for_change.
  • Clipboard access. Read and write system clipboard for data transfer.
  • App launcher. Launch applications and wait for them to load.
  • Auto-snapshot. Every action returns a screenshot so the agent always sees current state.
  • 43 MCP tools. Screen, OCR, UIA, mouse, keyboard, windows, clipboard, and compound actions.
  • Incremental perception. Rescans only the parts of the screen that changed, and can return what changed instead of a full screenshot — ~21× fewer tokens per step.
  • Screen memory. Recognises screens it has read before and reuses them, verified by pixels — 89× cheaper than reading again.
  • Speculative perception. Learns what actions do and confirms the expected result instead of re-reading — 19–23× cheaper, with a surprise report when the interface does something unexpected.
  • Adaptive waiting. Waits for the screen to actually settle rather than sleeping a fixed 300 ms — 11.9 s saved over a 50-step task.
  • Resolution cascade. Element lookups stop at the cheapest method that works; repeat lookups cost ~0.05 ms.
  • DPI-correct. Coordinates are physical pixels everywhere, so clicks land correctly on scaled displays.
  • Test suite. 254 automated tests; the desktop-driving ones skip themselves when no display is available.

Requirements

  • Python 3.10 or newer
  • VS Code, Cursor, Windsurf, Claude Desktop, or any other MCP client

Getting started

First, install the OSWright MCP server with your client.

Standard config works in most tools:

{
  "mcpServers": {
    "oswright": {
      "command": "uvx",
      "args": ["oswright"]
    }
  }
}

Note: If you don't have uvx, you can use pip install oswright and then set "command": "oswright" directly.

<details> <summary>Claude Desktop</summary>

Follow the MCP install guide, use the standard config above.

</details> <details> <summary>Claude Code</summary>
claude mcp add oswright uvx oswright
</details> <details> <summary>VS Code</summary>

Add to your user or workspace settings.json under mcp.servers:

{
  "mcp": {
    "servers": {
      "oswright": {
        "command": "uvx",
        "args": ["oswright"]
      }
    }
  }
}

Or use the VS Code CLI:

code --add-mcp '{"name":"oswright","command":"uvx","args":["oswright"]}'
</details> <details> <summary>Cursor</summary>

Go to Cursor Settings -> MCP -> Add new MCP Server. Name it oswright, use command type with the command uvx oswright.

</details> <details> <summary>Windsurf</summary>

Follow Windsurf MCP documentation. Use the standard config above.

</details> <details> <summary>Cline</summary>

Add to your cline_mcp_settings.json:

{
  "mcpServers": {
    "oswright": {
      "type": "stdio",
      "command": "uvx",
      "args": ["oswright"],
      "disabled": false
    }
  }
}
</details> <details> <summary>Goose</summary>

Go to Advanced settings -> Extensions -> Add custom extension. Name it oswright, use type STDIO, and set the command to uvx oswright.

</details> <details> <summary>Using pip instead of uvx</summary>

If you prefer a standard pip install:

pip install oswright

Then use this config:

{
  "mcpServers": {
    "oswright": {
      "command": "oswright"
    }
  }
}

Or run directly:

python -m oswright
</details>

Incremental perception

Most GUI agents re-perceive the entire screen on every step: full screenshot, full OCR, then hand the model a fresh image. Measured on a live desktop, the median observation changes 0.012% of pixels — so a full rescan does roughly 240× more work than the change warrants, and the screenshot it returns costs ~2,800 image tokens whether anything happened or not.

OSWright keeps a model of the screen between observations and rescans only the regions that actually moved.

observe()  ->  {"changed": true,
                "added":   [{"text": "Saved", "x": 812, "y": 447}],
                "removed": ["Unsaved changes"],
                "screen_fraction_scanned": 0.015}

Measured on this machine over a 14-step agent loop:

| | v0.4.0 (full OCR + screenshot) | incremental | |---|---|---| | Median latency per step | 212 ms | 33 ms | | Tokens per observation | ~2,764 | ~49 | | Tokens over 14 steps | 38,696 | 1,025 | | Screen re-read | 100% | 16% |

The busier the screen, the larger the gap: full OCR scales with how much text is on screen, whereas the incremental path scales with how much changed. The same comparison measures 6.5× on a quiet desktop and 14.3× with a dense web page open. Re-measure with benchmarks/ rather than trusting these.

Cost is a proxy, though, and a cheaper perception path that quietly degraded accuracy would be worse than none. So it is checked against task completion: scripted tasks driving the real tool surface across four applications, graded against each application's own state — UI Automation for Calculator and Explorer, the window title for Chrome and VS Code — never against OCR.

| configuration | Calculator | File Explorer | Chrome | tokens | |---|---|---|---|---| | v0.4-style (full screenshot) | 9/9 | 3/3 | 3/3 | 118,858 | | delta only | 9/9 | 3/3 | 3/3 | 5,252 | | delta + memory | 9/9 | 3/3 | 3/3 | 5,099 | | delta + memory + prediction | 9/9 | 3/3 | 3/3 | 7,981 |

Accuracy is identical across every configuration while token cost falls 23×. Run it with python benchmarks/bench_tasks.py.

Why both pixels and accessibility

The design bets that neither perception path wins everywhere. Turning each half off measures that rather than asserting it:

| configuration | Calculator | File Explorer | Chrome | |---|---|---|---| | full cascade | 9/9 | 3/3 | 3/3 | | accessibility only | 9/9 | 0/3 | 0/3 | | pixels only | 6/9 | 3/3 | 3/3 |

Accessibility-only — the posture most Windows GUI agents take — is perfect on XAML and blind on a Win32 list view and on web content. Probed against VS Code it sees 18 elements, the entire IDE being a single node named Chrome Legacy Window, while OCR reads 94 including every filename.

Pixels-only fails Calculator's buttons, because the button a human reads as 7 is named Seven, and Windows OCR returns no digits from Calculator at all.

The cascade is the only configuration that passes everywhere.

The resolution cascade

find_element and click_element stop at the first method that can answer, so cost tracks how novel the request is rather than how large the screen is:

| Rung | Method | Typical cost | |---|---|---| | 0 | Already in the screen model | ~0.05 ms | | 1 | Rescan only what changed | ~70 ms | | 2 | Accessibility tree (knows a Button is a button) | ~40 ms | | 3 | App's own text buffer via UIA TextPattern — exact characters | ~400 ms | | 4 | Full-screen OCR | ~250 ms |

Looking up text the model already knows is ~5,000× cheaper than the v0.4.0 path (0.05 ms versus 244 ms). The response reports which rung answered, so you can see what a task is actually costing.

Rung 3 is worth understanding: UIA's TextRange.FindText searches the application's own text buffer and returns exact bounding rectangles. It is immune to font, DPI, antialiasing and OCR error. It sits below the pixel rungs only because scanning a window's controls for it costs a few hundred milliseconds of cross-process COM — it is the accurate rung, not the fast one.

Note on ordering. These rungs are ordered by measurement, not by theory. The common advice is to make the accessibility tree primary, but on real applications it is not always cheaper: walking Chrome's tree took 537 ms here, slower than a full-screen OCR pass, and VS Code exposed only 18 elements to it. Neither pixels nor accessibility wins everywhere, which is why this is a cascade rather than a choice.

Asking the compositor instead of looking

On Windows, the desktop compositor already knows which pixels changed and exposes them through DXGI Desktop Duplication. Asking it costs 0.14 ms and transfers no pixels, against tens of milliseconds to capture a frame and discover it was identical — so an idle observation skips the capture entirely.

When something has changed, the compositor is left holding that frame, so its pixels are read directly from the GPU rather than grabbed a second time through a different API — 1.5–2.3× faster than mss in measurements here.

The compositor is used only as a fast negative for change detection. When it reports a change, the dirty regions still come from hashing the captured frame: the two are measured over slightly different intervals, so compositor rectangles can under-report relative to the pixels actually captured, and an under-reported region is text that never gets re-read. It degrades silently to tile hashing and normal capture wherever Desktop Duplication is unavailable.

Enable delta observations for action tools with --observation-mode delta. The default remains screenshot for compatibility with existing clients.

Reproduce all of this yourself: see benchmarks/. The reasoning behind each decision, including the dead

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars3
CategoryAutomation
Updated23d ago
Forks1

Languages

Python

Trust signals

92/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

1 low