oswright
MCP server for Windows desktop automation. Incremental perception instead of a screenshot per step: measured 16.9x fewer tokens per action than the alternatives, with the benchmarks to check it.
Install / Use
claude mcp add Ask-812 -- npx -y github:Ask-812/oswrightIf the server publishes to npm under a different name, use that package instead — check the repo README.
MCP Server
Model Context Protocol server
Quality Score
Category
AutomationSupported Platforms
Tags
Our assessment of oswright
oswright scores 83/100 on our quality scale, 2218th of 2,891 Automation skills we index.
Its MCP Server is 30 KB long, well organised into 26 sections with 17 code examples: a thorough specification that gives an agent plenty to work with.
It has 3 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated 23 days ago, so oswright is actively maintained.
- Our last check on 2026-09-18 found the source still online.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful.
AI review by kimi-k2.7-code on 2026-10-02. Automated pattern scan on 2026-10-02. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
oswright compared with similar skills
All 4 of these similar skills score higher than oswright; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| oswright (this skill)by Ask-812 | 83 | 3 | 23d ago | MCP Server |
| Agent-Reachby Panniantong | 100 | 95.5k | 2d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.9k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.3k | today | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 86.7k | today | MCP Server |
Frequently asked questions
- How do I install oswright?
- Run
claude mcp add Ask-812 -- npx -y github:Ask-812/oswright. The install tabs above show the steps for each supported agent. - Which AI agents does oswright work with?
- It is written for Claude Code and Claude Desktop, as a MCP Server file. Other agents that read the same format can often use it too.
- Is oswright safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. An AI review of the same text found nothing harmful. It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is oswright still maintained?
- The repository was last updated 23 days ago, so oswright is actively maintained.
Skill content
View source on GitHubOSWright
Desktop automation for AI agents, without paying for a screenshot every step.
mcp-name: io.github.Ask-812/oswright
An MCP server that lets an LLM drive real desktop applications — the desktop equivalent of Playwright MCP. It keeps a model of the screen between actions and re-reads only the parts that changed, so the same work costs an order of magnitude fewer tokens.

Eight fields read off an invoice and typed into an expense form, verified by the
application itself. Same task, same result, 7.4× less context than returning
a screenshot after every action. Every number on screen is measured during the
run — regenerate the whole thing with python benchmarks/record_demo.py.
Why this exists
Most GUI agents re-perceive the entire screen on every step: screenshot, OCR, hand the model an image, repeat. Measured on a live desktop, the median observation changes 0.012% of the screen's pixels. Re-reading everything does far more work than the change warrants, and charges ~2,800 image tokens whether anything happened or not.
OSWright asks the compositor what changed, rescans only that, and answers
element lookups from the cheapest source that can. The claims below are measured
on this machine and reproducible from benchmarks/ — including
the ones that did not come out in its favour.
Key Features
- Cross-platform. Windows (Win32 API), Linux (pynput/X11), macOS (pynput/Quartz).
- Accessibility tree. Find elements deterministically by role and name via Windows UI Automation — 100% accurate, instant, no model needed.
- Fast OCR. Windows OCR (built-in, instant) with EasyOCR fallback for Linux/macOS. Results are cached automatically.
- Lightweight on Windows. No PyTorch download — Windows uses the built-in OCR engine, so a full install is a few MB rather than a few GB.
- Image matching. Locates elements by template image via OpenCV.
- Window management. List, focus, minimize, close, and screenshot specific windows.
- Screenshot diffing. Detect when the screen changes with
wait_for_change. - Clipboard access. Read and write system clipboard for data transfer.
- App launcher. Launch applications and wait for them to load.
- Auto-snapshot. Every action returns a screenshot so the agent always sees current state.
- 43 MCP tools. Screen, OCR, UIA, mouse, keyboard, windows, clipboard, and compound actions.
- Incremental perception. Rescans only the parts of the screen that changed, and can return what changed instead of a full screenshot — ~21× fewer tokens per step.
- Screen memory. Recognises screens it has read before and reuses them, verified by pixels — 89× cheaper than reading again.
- Speculative perception. Learns what actions do and confirms the expected result instead of re-reading — 19–23× cheaper, with a
surprisereport when the interface does something unexpected. - Adaptive waiting. Waits for the screen to actually settle rather than sleeping a fixed 300 ms — 11.9 s saved over a 50-step task.
- Resolution cascade. Element lookups stop at the cheapest method that works; repeat lookups cost ~0.05 ms.
- DPI-correct. Coordinates are physical pixels everywhere, so clicks land correctly on scaled displays.
- Test suite. 254 automated tests; the desktop-driving ones skip themselves when no display is available.
Requirements
- Python 3.10 or newer
- VS Code, Cursor, Windsurf, Claude Desktop, or any other MCP client
Getting started
First, install the OSWright MCP server with your client.
Standard config works in most tools:
{
"mcpServers": {
"oswright": {
"command": "uvx",
"args": ["oswright"]
}
}
}
<details> <summary>Claude Desktop</summary>Note: If you don't have
uvx, you can usepip install oswrightand then set"command": "oswright"directly.
Follow the MCP install guide, use the standard config above.
</details> <details> <summary>Claude Code</summary>claude mcp add oswright uvx oswright
</details>
<details>
<summary>VS Code</summary>
Add to your user or workspace settings.json under mcp.servers:
{
"mcp": {
"servers": {
"oswright": {
"command": "uvx",
"args": ["oswright"]
}
}
}
}
Or use the VS Code CLI:
code --add-mcp '{"name":"oswright","command":"uvx","args":["oswright"]}'
</details>
<details>
<summary>Cursor</summary>
Go to Cursor Settings -> MCP -> Add new MCP Server. Name it oswright, use command type with the command uvx oswright.
Follow Windsurf MCP documentation. Use the standard config above.
</details> <details> <summary>Cline</summary>Add to your cline_mcp_settings.json:
{
"mcpServers": {
"oswright": {
"type": "stdio",
"command": "uvx",
"args": ["oswright"],
"disabled": false
}
}
}
</details>
<details>
<summary>Goose</summary>
Go to Advanced settings -> Extensions -> Add custom extension. Name it oswright, use type STDIO, and set the command to uvx oswright.
If you prefer a standard pip install:
pip install oswright
Then use this config:
{
"mcpServers": {
"oswright": {
"command": "oswright"
}
}
}
Or run directly:
python -m oswright
</details>
Incremental perception
Most GUI agents re-perceive the entire screen on every step: full screenshot, full OCR, then hand the model a fresh image. Measured on a live desktop, the median observation changes 0.012% of pixels — so a full rescan does roughly 240× more work than the change warrants, and the screenshot it returns costs ~2,800 image tokens whether anything happened or not.
OSWright keeps a model of the screen between observations and rescans only the regions that actually moved.
observe() -> {"changed": true,
"added": [{"text": "Saved", "x": 812, "y": 447}],
"removed": ["Unsaved changes"],
"screen_fraction_scanned": 0.015}
Measured on this machine over a 14-step agent loop:
| | v0.4.0 (full OCR + screenshot) | incremental | |---|---|---| | Median latency per step | 212 ms | 33 ms | | Tokens per observation | ~2,764 | ~49 | | Tokens over 14 steps | 38,696 | 1,025 | | Screen re-read | 100% | 16% |
The busier the screen, the larger the gap: full OCR scales with how much text is
on screen, whereas the incremental path scales with how much changed. The same
comparison measures 6.5× on a quiet desktop and 14.3× with a dense web page
open. Re-measure with benchmarks/ rather than trusting these.
Cost is a proxy, though, and a cheaper perception path that quietly degraded accuracy would be worse than none. So it is checked against task completion: scripted tasks driving the real tool surface across four applications, graded against each application's own state — UI Automation for Calculator and Explorer, the window title for Chrome and VS Code — never against OCR.
| configuration | Calculator | File Explorer | Chrome | tokens | |---|---|---|---|---| | v0.4-style (full screenshot) | 9/9 | 3/3 | 3/3 | 118,858 | | delta only | 9/9 | 3/3 | 3/3 | 5,252 | | delta + memory | 9/9 | 3/3 | 3/3 | 5,099 | | delta + memory + prediction | 9/9 | 3/3 | 3/3 | 7,981 |
Accuracy is identical across every configuration while token cost falls 23×.
Run it with python benchmarks/bench_tasks.py.
Why both pixels and accessibility
The design bets that neither perception path wins everywhere. Turning each half off measures that rather than asserting it:
| configuration | Calculator | File Explorer | Chrome | |---|---|---|---| | full cascade | 9/9 | 3/3 | 3/3 | | accessibility only | 9/9 | 0/3 | 0/3 | | pixels only | 6/9 | 3/3 | 3/3 |
Accessibility-only — the posture most Windows GUI agents take — is perfect on
XAML and blind on a Win32 list view and on web content. Probed against VS Code
it sees 18 elements, the entire IDE being a single node named Chrome Legacy Window, while OCR reads 94 including every filename.
Pixels-only fails Calculator's buttons, because the button a human reads as 7
is named Seven, and Windows OCR returns no digits from Calculator at all.
The cascade is the only configuration that passes everywhere.
The resolution cascade
find_element and click_element stop at the first method that can answer,
so cost tracks how novel the request is rather than how large the screen is:
| Rung | Method | Typical cost | |---|---|---| | 0 | Already in the screen model | ~0.05 ms | | 1 | Rescan only what changed | ~70 ms | | 2 | Accessibility tree (knows a Button is a button) | ~40 ms | | 3 | App's own text buffer via UIA TextPattern — exact characters | ~400 ms | | 4 | Full-screen OCR | ~250 ms |
Looking up text the model already knows is ~5,000× cheaper than the v0.4.0 path (0.05 ms versus 244 ms). The response reports which rung answered, so you can see what a task is actually costing.
Rung 3 is worth understanding: UIA's TextRange.FindText searches the
application's own text buffer and returns exact bounding rectangles. It is
immune to font, DPI, antialiasing and OCR error. It sits below the pixel rungs
only because scanning a window's controls for it costs a few hundred
milliseconds of cross-process COM — it is the accurate rung, not the fast one.
Note on ordering. These rungs are ordered by measurement, not by theory. The common advice is to make the accessibility tree primary, but on real applications it is not always cheaper: walking Chrome's tree took 537 ms here, slower than a full-screen OCR pass, and VS Code exposed only 18 elements to it. Neither pixels nor accessibility wins everywhere, which is why this is a cascade rather than a choice.
Asking the compositor instead of looking
On Windows, the desktop compositor already knows which pixels changed and exposes them through DXGI Desktop Duplication. Asking it costs 0.14 ms and transfers no pixels, against tens of milliseconds to capture a frame and discover it was identical — so an idle observation skips the capture entirely.
When something has changed, the compositor is left holding that frame, so its
pixels are read directly from the GPU rather than grabbed a second time through
a different API — 1.5–2.3× faster than mss in measurements here.
The compositor is used only as a fast negative for change detection. When it reports a change, the dirty regions still come from hashing the captured frame: the two are measured over slightly different intervals, so compositor rectangles can under-report relative to the pixels actually captured, and an under-reported region is text that never gets re-read. It degrades silently to tile hashing and normal capture wherever Desktop Duplication is unavailable.
Enable delta observations for action tools with --observation-mode delta.
The default remains screenshot for compatibility with existing clients.
Reproduce all of this yourself: see benchmarks/.
The reasoning behind each decision, including the dead
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
95.5kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.9kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
CowAgent
47.3kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Scrapling
86.7k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
