dsh-browser-application
Browser extension for the DeepSeek Harness desktop app. It lets the model see the page you are on as text with numbered controls, and act on it: click, type, scroll, navigate, manage tabs. It asks first, keeps passwords in the page, and connects only to 127.0.0.1.
Install / Use
npx skills add youbaiyun/dsh-browser-applicationInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
AutomationSupported Platforms
Tags
Our assessment of dsh-browser-application
dsh-browser-application scores 74/100 on our quality scale, 2612th of 2,870 Automation skills we index.
Its SKILL.md is 10 KB long, well organised into 10 sections with 6 code examples: a thorough specification that gives an agent plenty to work with.
It has 2 GitHub stars, so there is little community track record yet; judge it on its content.
Maintenance, license and trust
- The repository was last updated yesterday, so dsh-browser-application is actively maintained.
- It is released under the MIT license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 92/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
dsh-browser-application compared with similar skills
All 4 of these similar skills score higher than dsh-browser-application; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| dsh-browser-application (this skill)by youbaiyun | 74 | 2 | 1d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 91.8k | 20d ago | CLAUDE.md |
| Scraplingby D4Vinci | 100 | 85.8k | 1d ago | MCP Server |
| rufloby ruvnet | 100 | 73.9k | today | MCP Server |
| algorithmic-artby anthropics | 100 | 177.9k | 13d ago | SKILL.md |
Frequently asked questions
- How do I install dsh-browser-application?
- Run
npx skills add youbaiyun/dsh-browser-application. The install tabs above show the steps for each supported agent. - Which AI agents does dsh-browser-application work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is dsh-browser-application safe to use?
- It is MIT-licensed and scores 92/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is dsh-browser-application still maintained?
- The repository was last updated yesterday, so dsh-browser-application is actively maintained.
Skill content
View source on GitHubdsh-browser
让 dsh 桌面端以纯文本读取并操作一个浏览器标签页的 MV3 扩展 + 桥接插件。这是 Lum1104/dsh-browser 的衍生重写版,按「兼容性高、去臃肿化优先」重构,兼顾稳定性与高效。
结构(四个包同一个版本号 0.37)
packages/protocol 零依赖的线上协议(帧校验、能力/授权分离)
packages/bridge dsh 桥接插件(WebSocket 服务端 + browser_* 工具,跑在 dsh 内)
extension Chrome/Firefox MV3 扩展本体
与原版的关键差异
| 维度 | 原版 | 本版 |
|---|---|---|
| 版本 | 根 0.2.1 / 扩展 0.3.1 / 桥 0.0.7 各自漂移 | 统一 0.37(要发布到 npm 的桥接包是 0.37.0 —— npm 的 semver 要求三段,Chrome 的 manifest 接受两段 ✓)(桌面端插件列表里显示的就是桥接的这个版本) |
| Node | ^22.19 \|\| >=24 | >=20 |
| TypeScript | 扩展 5.6 与桥 6.0 分裂 | 单一 ^5.7 |
| 依赖 | 根包 35 个 @deepseek-ai/*(RC)、node_modules 600MB | 根包 0 个 dsh 依赖;扩展零运行时依赖;桥把 dsh 全声明为 peerDependencies(运行时 ctx.get() 探测宿主,不内嵌) |
| 桥接死依赖 | React peer + 3 个 web 端 dsh-client-ui-*/locale peer + dsh.client 注入块 | 全部移除(扩展面板自带,不往 dsh 的 web 客户端注入 UI) |
| 桥接协议 | 自带 protocol.ts(12.7 KB)与扩展各存一份 | 共享 @dsh-browser/protocol,由打包器内联进两份产物 |
| 桥接冗余 | 另有 web 端 client.js(7.2 KB) | 删除 |
| 注入 | manifest 全局 content_scripts 注入每个站点的每个 iframe | 按需注入:仅在受控标签页触发 |
| 浏览器下限 | Chrome 116 / Firefox 140 | Chrome 114 / Firefox 128 |
| 构建 | 3 个 vite 配置 + shared + build.mjs(5 文件) | 1 个 build.mjs(3 次隔离构建) |
| 测试运行器 | vitest + jsdom | 单用 Node 内置 node --test(类型擦除直跑) |
核心机制
- 文本快照 + 动作执行:不截图、不做图像识别(协议层
textOnly: true)。页面渲染成结构化文本:标题/URL/正文(readability-lite)+ 编号交互清单(含 ARIA role 控件)+ 表单字段(含masked/checked/required),支持delta差分与region局部快照。 - 安全不变量:敏感字段(
type=password、autocomplete=credit-card|cc-*、id/name/aria-label 命中password|passwd|credit|card|cvv|cvc|secret|pwd)的值一律掩码成••••;可访问名永不使用输入框的当前值(仅 submit/button/reset 类的 value 作名),并有单测钉死。 - 工具面(15):
browser_snapshot/click/type/press/scroll/navigate/open_tab/list_tabs/follow_tab/close_tab/back/forward/reload/get_text/wait。完整参数面:delta/region/replace/amount/selector/ms/active/tabId/index/text/key/direction/url,以及 6 个帧局部工具上的frame。 - 帧路由:快照组合主帧与所有可访问 iframe,子帧标注为
[frame N] <origin>;后台记住N → frameId,后续带frame的工具路由到同一帧;每帧按 80/20 分配快照预算。 - 超出上游的能力(在上游实现之上补的能力,不是删减):
browser_type能填<select>(按 option value → 可见标签 → 1-based 序号依次匹配,失败时列出可用选项)并用 true/false 设置 checkbox/radio;browser_wait支持等待条件(selector/text,未出现则以timeout错误码失败),不再只是固定延时。 - 导航后自动快照:content 脚本在新文档就绪时宣告(等价原版
DSH_CONTENT_READY),后台在导航类动作(navigate/back/forward/reload)后握手等待新文档并自动快照,推送到面板——不必再要求模型多跑一轮。 - 图片清单(不接视觉也有用):移植的管线原本完全丢掉裸
<img>——图片不在交互选择器里,innerText也不带alt。现在快照末尾新增Images:节:按文档顺序、按最近标题分段(序号与实际分段一致),每条给at=#N(可点击的交互编号)、pos、kind、size、alt、near(紧邻文本锚)与识别结果desc;作者已给文本的图片标state=text,不花识别调用。定位仍由 DOM 给出,模型永远不需要坐标。格式为单行定长文法、自由文本只在引号内,因此页面无法伪造分节。 覆盖的图片表面:<img>(含currentSrc)、<canvas>、<video poster>(封面是普通图片 URL,走同一条取图路径)、内联<svg>(无字节可取,因此命名为state=text、未命名的标state=unreadable reason=inline-svg,绝不会交给取图路径)、以及 CSS 背景图——内联background-image精确采集,类声明的走有界 computed-style 扫描(先按"无文字、尺寸够大、内部没有已被采集的 img/canvas/video/svg"筛掉绝大多数元素,再解析样式,扫描上限 400)。 渲染在后台:内容脚本只做采集并带外返回ImageEntry[],纯字符串渲染在shared/image-manifest.ts,由后台在把结果交给模型前追加——因为识别是异步的,后台才是持有识别结果与缓存的一侧;iframe 各帧各自追加,所以img N与at=#N都是帧局部的。预算按比例切分(IMAGE_BUDGET_SHARE = 0.25):有图时内容脚本按 75% 渲染正文、后台按 25% 渲染该节,两侧读同一常量因而相加不超协商预算;图多时先让图片节自己缩,正文不被挤掉。 - 取图与识别缓存(后台):
background/image-fetch.ts在 service worker 里取图——扩展上下文持有 host permissions,能直接读内容脚本永远读不到的跨域图片字节(内容脚本画布会被污染);带credentials:'include'以覆盖登录后的媒体,解码后用OffscreenCanvas缩到长边 1024 并转 WebP。失败被分类而不是抛出(network/http-error/not-an-image/decode-failed/too-large),因为模型必须被告知"这里有图但读不到",而不是看不到任何东西。background/image-cache.ts按图片身份(URL)缓存结果与失败(成功 6 小时、失败 5 分钟、LRU 300 条),会话存储做 MV3 worker 被回收后的兜底。background/image-pipeline.ts是识别器的接缝:直连云端与经桌面桥接只差这一个被注入的函数,取图/缩放/缓存/渲染完全相同。 - 桌面中转识别(两条互补的取图路径):协议新增
image.call/image.result一对帧,hello.ok的 policy 增加imageRecognition。扩展先自己取图(只有它带登录态),拿到字节就随帧发下去;取不到才把 URL 交给桌面,由桌面用不受 CSP/host permission/企业策略约束的网络栈去取——这是桌面中转真正的价值,也是为什么帧里带的是source而不是固定的那一种。桌面侧vision.ts调 V4.1 Flash(OpenAI 兼容/chat/completions,图片走 data URL,thinking关闭,usage回读以便确认开关真的生效),image-relay.ts负责取字节或按 URL 兜底并分类失败。凭据留在桌面:visionApiKey为空时hello.ok报imageRecognition: false,扩展于是走自己的路,不会把帧发给一个不会应答的桌面。 - 外接 API 直连(识别这一步对用户完全不可见):
background/direct-recognizer.ts让扩展自己调外接 API,配置从chrome.storage.local读(visionEndpoint/visionApiKey/visionModel/visionThinking/visionTimeoutMs),不新增任何界面——面板、桌面端、web 端都不显示任何识别相关的 UI。顺序是先直连、再桌面中转:未配置(no-direct-vision)才落到下一条,而真实失败(429 等)当场结算,不会为同一张图付两次钱。shared/vision-settings.ts是独立模块,因此改这块不触碰 fork 自己的settings.ts。关键保证:识别这一步不建会话、不写文件、不产生对话条目——它只是后台里一次普通外发调用;缓存用chrome.storage.session(内存,浏览器关闭即清)。两条路径共用protocol/src/vision-contract.ts的提示词与解析器,所以「谁调模型」不改变问了什么、也不改变什么算合格回答。 - 面板 markdown:模型回复用
marked渲染、DOMPurify净化;流式期间保持纯文本(避免逐帧重解析),定稿后渲染;回复中的链接在新标签打开,不跳出面板。 - 审批确认:状态变更动作默认先问用户(
unrestrictedBrowserAccess开关、受信源、单次允许)。 - 按需注入:无 manifest
content_scripts,脚本只在受控标签页经chrome.scripting注入,导航后重新注入。 - 标签页亲和:工具绑定单一受控标签页;
ask模式下切换标签页会阻塞并询问 keep/follow。 - 会话桥接:面板提示词经
rpc帧发往桌面 dsh,assistant 文本流式回传。 - 桥接插件:
/ext/bridge上的 token 认证 WebSocket,hello握手协商 caps/policy,browser_*工具以tool.call帧下发给扩展,特权网关方法对非回环远端一律拒绝。
看图(图像识别)
界面里只有一个开关:面板 → 设置 → 看图(关 / 低 / 标准 / 增强)。默认关,图片不出本机。
调用去哪是部署配置,不是设置项 —— 面板刻意不提供文本框(见 extension/tests/control-page.spec.ts 的 offers no control… 那条断言)。
桌面中转(推荐):在桌面端 profile 的 cordis.patch.yml 里给桥接加一个 key,地址与模型已有默认值(https://api.deepseek.com/v1 / deepseek-v4.1-flash),思考强制关闭:
- id: bridge-browser
disabled: false
config:
visionApiKey: <key>
visionModel: <可读图的模型 id>
浏览器直连:没有界面入口,按部署写入扩展存储即可(非回环地址还需在 manifest 的 connect-src 放行):
chrome.storage.local.set({ dshSettings: { visionEndpoint: 'https://…/v1', visionModel: '…', visionApiKey: '…' } })
tools/vision-stub.mjs 与 tools/vision-proxy.mjs 是开发工具:前者量本地开销,后者让扩展在不改 manifest 的前提下走云端。日常用法不需要它们。
命令
pnpm install
pnpm -r run typecheck # 全仓类型检查(协议 + 桥接 + 扩展)
pnpm -r run test # 全部单测
pnpm --filter dsh-browser-extension run build # Chrome → extension/dist
pnpm --filter dsh-browser-extension run build:firefox # Firefox → extension/dist-firefox
pnpm --filter dsh-browser-application run build # 桥接插件 → packages/bridge/lib
端到端评测(benchmark/)
把"高效"从形容词变成数字的仪器:同一个模型、同一份 profile、同一组任务、同一台机器,对比两种浏览器执行后端——playwright(runner 内置的基线插件,暴露与产品一致的 browser_* 工具契约)与 extension(本仓库真实的桥接 + 已构建的扩展)。
六个任务覆盖浏览器 agent 的基本形态(读 / 单步 / 表单 / 搜索 / 多步 / 动态加载),按 seed 变化以防过拟合,跑在 benchmark/site/ 的夹具站点上。每个任务记四个数:成功率(120 秒超时算失败)、完成耗时 p50/p90/mean(从模型收到任务到 DSH turn 结束)、平均工具调用次数、平均 prompt token。
必须成对跑的理由:绝对耗时被模型支配(生成约占一半),只有对照才能把后端自己的贡献分离出来。
pnpm --dir benchmark install-browser # 下载与 lockfile 中 playwright-core 匹配的 Chrome for Testing
node benchmark/run.mjs --dry-run --smoke # 不调用模型的基础设施检查
node benchmark/run.mjs --smoke # 调一次真实模型(消耗额度)
node benchmark/run.mjs # 正式:6 任务 × 5 seed × 2 后端 = 60 次
node benchmark/run.mjs --tasks order_lookup,contact_form --seeds 1-5 --trials 2
三个前置条件,缺一个它就跑不起来:
- 一个能接受
--profile web --patch … -- --no-open --port N的 dsh 命令行 ✗ —— 上游仓库把@deepseek-ai/dsh当依赖,那里pnpm exec dsh直接可用;这份重写是独立工作区,没有这个依赖(桥接是装进桌面端 profile 的),所以要自己指:BENCHMARK_DSH_COMMAND="node D:/path/to/dsh.mjs"。 - Windows 上用
pnpm.cmd—— 全局安装同时提供pnpm/pnpm.cmd/pnpm.ps1,而child_process.spawn只能执行前两者,且 Node 18+ 还要求.cmd必须经由 shell。脚本已处理这两点,并把实际使用的命令写进失败信息里。 - Chrome for Testing 或 Chromium —— 新版 Google Chrome Stable 会忽略
--load-extension,不能作为扩展后端的自动评测浏览器。
benchmark/tests/ 是评测工具自身的单测(node --test tests/*.test.mjs,不需要模型与浏览器)。
状态
四项能力已全部补齐,另加图片清单节、后台取图/缓存、外接 API 直连与桌面中转识别。下表列出各套测试,具体条数以 pnpm -r run test 的输出为准(此前这里写死的数字已经漂移过两次,所以不再写死)。
| 能力 | 实现 | 验证 |
|---|---|---|
| 内容管线(与原版等价 + 扩展) | extract / privacy / ids / snapshot / actions 全量移植:可访问名优先级、敏感字段掩码、forms 数组、ARIA role、稳定 id、delta、region;另外补了 <select>/checkbox 填写、browser_wait 等待条件、Images: 图片清单节 | jsdom 测试(含安全不变量断言) |
| 帧路由 | [frame N] 编号 + N → frameId 路由表 + 每帧 80/20 预算 | 5 项单测(frames.spec.ts) |
| 导航后自动快照 | content 就绪宣告(DSH_CONTENT_READY 等价)+ 后台握手 + 导航后自动快照推面板 | 随 e2e 与双构建通过 |
| 面板 markdown | marked + DOMPurify,流式纯文本 / 定稿渲染,链接拦截 | 双构建通过 |
| 桥接插件测试套件 | 原版 spec 全量迁移(vitest) | 14 个 spec 文件,162 项 |
| 端到端 | Playwright 加载已构建扩展 + 真实 BridgeServer:零配置发现 → caps 协商 → session.create/prompt | 1 项,真跑通(含在上面 14 个文件里) |
测试运行器:协议用 Node 内置 node --test;扩展与桥接用 vitest(桥接套件自原版原样迁移,不重写断言)。
另有 1 项跳过:locales.spec.ts 里「与商店上架文案一致」的检查——store-assets/store-listing.md 不存在时跳过;补上那份文案后它会开始跑(它同时管着三语描述 132 字符的上限)。
e2e 需要一个仍遵守 --load-extension 的浏览器——Playwright 自带 Chromium 可以,系统 Chrome/Edge 137+ 默认忽略该开关;缺浏览器或 extension/dist 时自动跳过。可用 PLAYWRIGHT_CHROMIUM_PATH 指定。
分层原则(为什么这些代码在它现在的位置)
- 纯逻辑与接线分开:校验进来方向的数据、生成给用户看的审批文案、编码字节、解析帧号与组合跨帧快照——这些是纯的,因此放在可单测的模块里(
shared/payloads.ts、background/action-summary.ts、background/base64.ts、background/frame-routing.ts)。接线留在background/index.ts。 - 组合根也可测:
tests/chrome-stub.ts提供够用的chromeAPI 桩,使background/index.ts能在测试里被加载,于是"每个事件都有监听器""content-ready 会被答复""非面板的 port 被忽略"这些接线本身有了断言。 - 类型与守卫同源:
APPROVAL_DECISIONS/TAB_AFFINITY_DECISIONS/TAB_SWITCH_MODES是运行时数组,类型与守卫都从它派生——给面板加一个按钮却忘了教后台接受它,是"按钮点了没反应"这类 bug 的常见来源。 - 成本保证可验证:
visionThinking: off无法从请求体证明生效(服务商可以静默忽略),所以启动时发一次 1×1 图的自检并回读usage;只有真的看到 reasoning token 才告警,看不到就保持安静(不误报)。 - 未使用的导入/参数是错误:
tsconfig.base.json开了noUnusedLocals+noUnusedParameters,这类腐化不能再静默堆积。 - 跨包不共享小工具:
isRecord这类三行类型守卫在协议包与桥接包各自保留一份,因为把线协议包变成通用工具包不值得。
extension/tools/preview-images.ts 是开发工具(不参与构建与发布):它打印一份真实结构页面上的 Images: 节,可以在不接模型的前提下审阅格式。
待办(目标外):background/index.ts 可继续拆分;桥接的 composition.spec.ts(4 用例,真实 Loader)与 session-purge.spec.ts(12 用例)未迁移,需要再补约 15 个 dsh 包作为 devDependency。
未迁移带走的验证(不是"没做",是"没人看着"):组装给模型的提示必须保持纯 ASCII(网页因此无法伪造面板来源标记)——上游就没有断言过它,这条性质现在只写在 packages/bridge/src/index.ts 的注释里。要恢复它,把那段字面量提成一个具名常量再断言其 charCodeAt 全在 ASCII 范围内即可,是纯提取、不改文本。
许可
MIT。本工程是 Lum1104/dsh-browser 的派生重写(上游引擎的归属与文件范围见上一版套件里的 COPYRIGHT-who-owns-what.md),因此 LICENSE 保留了上游的版权声明,并加上本工程自己的那一行:
Copyright (c) 2026 Yuxiang Lin
Copyright (c) 2026 youbaiyun
MIT 要求分发时随附版权声明与许可文本,所以这个文件必须跟着源码与扩展包一起走,不能删。
Related Skills
Agent-Reach
91.8kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Scrapling
85.8k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
ruflo
73.9k🌊 The original agent harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, federation, vector RAG integration, and native Claude Code / Codex / Hermes and many more Integrated
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
