16 skills found
firerpa / lamdaAndroid Full-Stack Device Control Platform: WebRTC/H.264 remote desktop, UI/OCR/image-matching automation, one-click MITM, built-in Frida, proxy/VPN/frp/P2P networking, MCP/Agent, 160+ APIs, designed for multi-device clusters and engineered deployments.
giancarloerra / SocratiCodeEnterprise-grade (40m+ LOC) codebase intelligence, zero-setup, local & private Plugin/Skill/Extension or MCP: hybrid semantic search, polyglot dependency graphs, symbol-level impact analysis & call-flow, interactive HTML viewer, cross-project & branch-aware search, DB/API/infra knowledge.
SylphxAI / pdf-reader-mcpGive your AI agent eyes for PDFs — structured text, tables, OCR, visual evidence, and page-level citations via MCP. Native Rust, local-first.
AI272 / speakerSpeaker is a Codex skill project for academic presentations: read real.pptx, combine text extraction, PPTX structure parsing, page-by-page rendering, OCR, and visual review to generate page-by-page speaker notes, and write a clean version of the lecture into the PowerPoint comment area.
AmrDab / clawdcursorclawdcursor compiles whatever's on screen into one UI map — accessibility tree and OCR fused into stable, addressable elements, with a screenshot only when needed — then drives apps through reusable scripts, verifying every action and routing it through a single safety gate.
guimatheus92 / mcp-video-analyzerMCP server that turns any video — YouTube, Instagram, TikTok, Loom, X, Vimeo, direct URLs, local files — into transcripts, key frames, OCR text, and metadata for AI agents.
kitlau86 / agent-vision-mcpAn MCP server that gives non-vision LLMs the ability to "see" images. Plug in any OpenAI-compatible vision API (Gemini, Qwen-VL, OpenAI, or self-hosted) and your text-only model — like DeepSeek in Claude Code — can analyze screenshots, OCR text, read charts, and more.
MikkoParkkola / nabToken-lean web microfetch for LLM agents: any URL → clean markdown via CLI, MCP server, and Claude Code plugin. Real browser-cookie auth, passkeys, anti-bot reach, on-by-default prompt-injection defense, plus on-device multimodal ASR/OCR. A single Rust binary — not a browser.
HarperZ9 / gatherResearch intake that reaches the hard places: web, video, papers, scanned PDFs, browser, OCR, and audio into structured research packets. DOM extraction and change tracking built in; provenance rides along on every item.
opendatalab / omnidocbench-eval-helperHelp users deploy, validate, run, and parse OmniDocBench evaluations. Use this skill whenever the user mentions OmniDocBench, document parsing/OCR benchmark scoring, MinerU or other model evaluation on OmniDocBench, CDM formula metrics, end2end/md2md configs, Docker/conda deployment, remote SSH/H-cl…
BytePioneer-AI / wecom-app-ops企业微信自建应用(wecom-app)运维与使用技能包。用于:定位并回发图片/语音/文件;使用 saved: 稳定路径做 OCR/MCP/二次处理;规范 target(wecom-app:user:xxx / user:xxx / 裸ID / @accountId);排查发送失败;配置入站媒体保留策略与语音转码策略。
hassan-alnator / vibe-eyesAutomated visual testing framework for Claude and MCP-compatible clients. Generate unit tests, run E2E browser automation, perform visual regression testing, and extract text with OCR - all through natural language commands.
AuroraKON6 / backendBoss auto-apply tool - Playwright + Spring Boot + AI/OCR
AuroraKON6 / frontendBoss auto-apply tool - Playwright + Spring Boot + AI/OCR
Asenay-Tech / git-autonomyAI bookkeeping engine that automates 98–100% of financial tasks—uploads, OCR, categorization, fraud checks, forecasting, and reporting. Designed for companies in logistics, retail, restaurants, warehouses, and enterprise workflows.
tygwan / craft-syncfoliocraft에서 생성/배포된 프로젝트 사이트 정보를 메인 포트폴리오(`dev/portfolio`)의 데이터 파일과 동기화합니다. ```