computer-use
Use Cua Driver through MCP to inspect and operate the user's native desktop apps on Windows, macOS, or Linux. Trigger when a task requires a desktop application, native file dialog, OS window, signed-in browser UI, screenshot-grounded interaction, or a result that must be verified in the application…
Install / Use
npx skills add xuzhougeng/wisp-science --skill computer-useInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Development & EngineeringSupported Platforms
Our assessment of computer-use
computer-use scores 83/100 on our quality scale, 2763rd of 4,657 Development & Engineering skills we index.
Its SKILL.md is 5.4 KB long, split into 6 sections with 1 code example: a solid amount of guidance for an agent.
With 1,167 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 8 days ago, so computer-use is actively maintained.
- It is released under AGPL-3.0, a copyleft license: you can use it, but modified versions you distribute must carry the same license.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-02. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
computer-use compared with similar skills
All 4 of these similar skills score higher than computer-use; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| computer-use (this skill)by xuzhougeng | 83 | 1.2k | 8d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 88.6k | 17d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.3k | today | CLAUDE.md |
| rufloby ruvnet | 100 | 73.7k | today | CLAUDE.md |
| CowAgentby zhayujie | 100 | 47.2k | today | CLAUDE.md |
Frequently asked questions
- How do I install computer-use?
- Run
npx skills add xuzhougeng/wisp-science --skill computer-use. The install tabs above show the steps for each supported agent. - Which AI agents does computer-use work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is computer-use safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It is AGPL-3.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is computer-use still maintained?
- The repository was last updated 8 days ago, so computer-use is actively maintained.
Skill content
View source on GitHubname: computer-use description: "Use Cua Driver through MCP to inspect and operate the user's native desktop apps on Windows, macOS, or Linux. Trigger when a task requires a desktop application, native file dialog, OS window, signed-in browser UI, screenshot-grounded interaction, or a result that must be verified in the application. Keep browser page work in browser-use when the page bridge is sufficient." fold_cue: "instead_of=blind_pixel_clicking use=list_windows_then=get_window_state — keep an exact window target and refresh state after navigation or UI changes"
Computer Use — operate native desktop apps with Cua Driver
Wisp uses the installed Cua Driver as an MCP server. Cua Driver owns the platform-specific desktop integration; this Skill owns the agent workflow and the safety boundary. It can operate native apps, native file dialogs, and browser windows that are visible to the host desktop.
Before the first action
Use this Skill only when the Cua Driver tools are advertised in the current conversation. If they are absent, do not invent tool names or fall back to blind shell input. Tell the user to install Cua Driver, add a stdio MCP connection with:
command: cua-driver
args: mcp
Then ask them to reconnect the MCP service. The driver must run in the interactive user session. An SSH or service-session process cannot see the user's desktop. On macOS, Accessibility and Screen Recording permission must be granted to the Cua Driver app identity. On Windows and Linux, report any interactive-session or display-server refusal as a capability boundary.
If the available Cua Driver tool exposes a health, doctor, or permission
status call, use it first. Otherwise call list_apps or list_windows as a
read-only connection check. A process starting successfully is not evidence
that the desktop is controllable.
Tool selection
Use the exact tool schemas advertised by the connected Cua Driver server. The common names are:
list_appsandlaunch_appfor application discovery and startup;list_windowsfor exact process/window identity;get_window_statefor the accessibility tree plus a window screenshot;get_desktop_statefor the primary desktop screenshot and desktop identity;click,type_text,press_key,hotkey,scroll, anddragfor input;- window or session cleanup tools when the driver advertises them.
Do not guess a selector, process ID, window ID, element index, or coordinate.
Read the current state first. For input, prefer a window target with an exact
pid and window_id; use the returned accessibility element_index when the
control exposes a semantic action. Use window-local pixel coordinates only
when the element is not actionable semantically. Use a desktop target only
for deliberate foreground screen actions.
Observe → act → verify
For every meaningful action:
- Discover the app and select one exact window. If several candidates match, stop and resolve the ambiguity instead of choosing by title alone.
- Call
get_window_stateand keep the resulting window identity and fresh element references together. Treat element indexes as stale after a page navigation, dialog transition, window recreation, or material UI change. - Perform one bounded action. Prefer background delivery when the target and platform support it. Request foreground delivery only for that action when the application requires focus and interrupting the user's desktop is acceptable.
- Read the same target again and verify the application state or external artifact. A successful input dispatch is not proof that the application handled it.
- If the result is stale, ambiguous, refused, or unverifiable, follow the returned refusal code and re-observe. Do not retry the same blind action.
For a native save or export, verify the actual path and file existence with a
filesystem tool after the application reports completion. For a visual canvas,
verify the screenshot and, where possible, an application-owned state or
exported artifact. For a browser page, use browser-use page tools when they
provide the needed operation; use Cua Driver for browser chrome, native
dialogs, or a page surface that the browser bridge cannot access.
Safety boundaries
- Ask for confirmation before sending, posting, purchasing, deleting, submitting, or otherwise committing an irreversible external action.
- Never type passwords, API keys, payment data, or one-time codes. Have the user enter them in the visible application and continue after confirmation.
- Do not use desktop control to solve CAPTCHA or bypass human verification.
- Do not treat
effect: confirmedas a universal success signal; inspect its evidence and verify the application-owned result. - Keep one foreground input sequence serialized. Do not drive two windows with concurrent keyboard or pointer actions.
- If a target disappears, permissions change, or the driver returns a structured refusal, report the concrete reason and stop or re-observe as the refusal instructs.
First smoke task
For a new installation, use a reversible task such as opening Calculator,
entering 6 × 7, and reading back 42. For this project’s acceptance task,
open Inkscape, make one small edit, export through the native dialog, and
verify the resulting SVG exists at the requested path. Record the platform,
driver version, delivery mode, and whether verification was semantic, visual,
or filesystem-based.
Related Skills
Agent-Reach
88.6kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.3kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
ruflo
73.7k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
CowAgent
47.2kOpen-source personal AI assistant & Agent Harness. Plans tasks, runs tools and skills, self-evolves with memory and knowledge. Multi-agent, multi-model, multi-channel. Lightweight, extensible, one-line install.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
