deepstream-sop
Use this skill when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection (GEBD) plus VLM classification.
Install / Use
npx skills add NVIDIA/skills --skill deepstream-sopInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
OperationsSupported Platforms
Our assessment of deepstream-sop
deepstream-sop scores 90/100 on our quality scale, 185th of 487 Operations skills we index (top 38%).
Its SKILL.md is 18 KB long, split into 6 sections with 1 code example: a thorough specification that gives an agent plenty to work with.
With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.
Maintenance, license and trust
- The repository was last updated 5 days ago, so deepstream-sop is actively maintained.
- It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
- Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
deepstream-sop compared with similar skills
All 4 of these similar skills score higher than deepstream-sop; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| deepstream-sop (this skill)by NVIDIA | 90 | 3.4k | 5d ago | SKILL.md |
| Agent-Reachby Panniantong | 100 | 86.0k | 13d ago | CLAUDE.md |
| headroomby headroomlabs-ai | 100 | 74.0k | today | CLAUDE.md |
| crawl4aiby unclecode | 100 | 84.4k | 3d ago | MCP Server |
| Scraplingby D4Vinci | 100 | 84.4k | today | MCP Server |
Frequently asked questions
- How do I install deepstream-sop?
- Run
npx skills add NVIDIA/skills --skill deepstream-sop. The install tabs above show the steps for each supported agent. - Which AI agents does deepstream-sop work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is deepstream-sop safe to use?
- It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is deepstream-sop still maintained?
- The repository was last updated 5 days ago, so deepstream-sop is actively maintained.
Skill content
View source on GitHubname: "deepstream-sop" description: > Use this skill when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection (GEBD) plus VLM classification. Trigger even if the user does not name it: verify operator step sequence, detect missing or out-of-order SOP steps, score factory/work-cell video for procedure compliance, run VLM-based SOP checking on industrial cameras, or call /v1/chat/completions with a file, RTSP, or Basler camera. Also trigger for its internals: SOPVideoProcessor, DeepStream GEBD model (e.g. DDM) via Triton CAPI, nvds_custom_postprocess, Cosmos Reason 1/2 vLLM, SSE streaming, Kafka NvProto/JSON output, Basler/Pylon camera + emulation, Docker compose, chunk-level latency. Do NOT trigger for generic DeepStream pipelines, object detection/tracking, NIM imports, or video summarization. owner: "windy@nvidia.com" service: "deepstream-sop" version: "1.0.0" license: "CC-BY-4.0 AND Apache-2.0" reviewed: "2026-04-08" metadata: author: "Wind Yuan windy@nvidia.com" tags: - deepstream - sop - vlm - triton - gpu languages: - python frameworks: - deepstream - triton - fastapi domain: video-analytics
DeepStream SOP Inference Microservice Skill
This skill guides AI coding assistants in building, extending, and debugging the NVIDIA DeepStream SOP (Standard Operating Procedure) Inference Microservice — a GPU-accelerated pipeline for temporal action detection and VLM-based SOP compliance monitoring on industrial video feeds.
Reference repository: https://github.com/NVIDIA/sop-monitoring-blueprints/tree/main/microservices/sop-inference-bp
Local reference code: sop-inference-bp/ directory (from a local clone of the repository)
Models
Model-agnostic at both inference stages — swap via env var (and Triton dir for GEBD).
| Stage | Role | Model class | Default | Swap via |
|------|------|-------------|---------|----------|
| Stage 1 (CV) | Per-frame boundary scoring → chunk segmentation | Generic Event Boundary Detection (GEBD) | DDM (MCG-NJU/DDM) via Triton Python backend | Replace triton_model_repo/<model>/ + DDM_MODEL_PATH (§ 5) |
| Stage 3 (VLM) | Per-chunk action classification | Vision-language model via vLLM | Cosmos Reason 1 7B (Reason 2 also supported) | Set VLLM_MODEL_PATH to a different HF ID or local path |
"GEBD" = swappable Stage-1 slot; "DDM" = the default architecture (terms used interchangeably).
Chunking is selectable per request (§ 2): default ddm-net uses GEBD; uniform produces fixed-length chunks and bypasses Stage-1 GEBD (§ 3, § 6). DDM temporal window is configurable via FRAMES_PER_SIDE / SEQUENCE_BATCH (§ 4, § 5), with optional TensorRT (§ 5).
Architecture Overview
Runs in a Docker container (nvds-action-sop) alongside a Kafka container. Full diagram: references/sop_architecture.svg.
Data flow through the 4-stage SOPVideoProcessor pipeline (per-request):
Input Sources Docker Container: nvds-action-sop
───────────── ──────────────────────────────────────────────────
Video Files ──┐ FastAPI Server (port 8300)
RTSP Streams ─┤── base64/ ├─ /v1/chat/completions → SOPProcessManager
Basler Camera ┘ file/rtsp/ │
camera │ ModelInitializer: VLM first, then DDM dummy pipeline
│ 4 Thread Pools: cv(32), clip(32), vlm(64), vlm_req(64)
│
▼ SOPVideoProcessor (per-request)
┌────────────────────────────────────────────────┐
│ Stage 1: DeepStream Pipeline (GPU) │
│ Source → nvstreammux → tee1 │
│ ├─[inference] queue1 → nvdspreprocess │
│ │ → nvinferserver (Triton CAPI + DDM) │
│ │ → InferOutputTensorParser → score_queue │
│ ├─[frames] queue3 → nvvideoconvert │
│ │ → capsfilter → appsink │
│ │ → DecodedFrameRetriever → frame_queue │
│ └─[RTSP out] queue → convert → H.264 enc │ (optional, § 18)
│ → rtppay → udpsink → RTSPServer (§ 18) │ opt-in only
│ │ boundary scores │
│ ▼ │
│ Stage 2: Clip Post-Process │
│ Boundary detection → chunk segmentation │
│ │ video frames + timestamps │
│ ▼ │
│ Stage 3: VLM Inference │
│ Embedded vLLM (Cosmos Reason 1/2) │
│ Frame sampling at VLM_FPS → classification │
│ │ action labels │
│ ▼ │
│ Stage 4: SOP Checker │
│ Sequence validation → missing/misordered │
│ │ chunk results │
│ ▼ │
│ final_queue │
└────────────────────────────────────────────────┘
│
Output ▼
────── ┌─────────────────┐
SSE Stream (chat.completion.chunk) │ Kafka Messages │
Non-streaming (chat.completion) │ (JSON/Protobuf) │
Prometheus metrics (/v1/metrics) └────────┬────────┘
▼
Docker Container: kafka
(apache/kafka:3.7.0)
Section Index
Each section is a standalone file in references/ — load only what your task needs.
| § | File | Responsibility |
|---|------|---------------|
| 1 | skill_01_fastapi_endpoints.md | FastAPI endpoints, server init, Prometheus metrics |
| 2 | skill_02_pydantic_schemas.md | Request/response Pydantic models (api_types.py) |
| 3 | skill_03_deepstream_pipeline.md | DeepStream pyservicemaker pipeline, tensor parser, dummy pipeline |
| 4 | skill_04_config_templates.md | nvdspreprocess / nvinferserver config templates + rendering |
| 5 | skill_05_triton_ddm_model.md | Triton model repo, config.pbtxt, model.py, ddm_net.py |
| 5b | skill_05b_custom_postprocess.md | C++ postprocess plugin, Makefile, IOptions API |
| 6 | skill_06_sop_process_manager.md | SOPProcessManager, SOPVideoProcessor, VLLMInference, Kafka |
| 6b | skill_06b_sop_checker.md | SOP sequence and checker compliance: MissingNumberDetector, SopCheckerCache, SopCheckerRequest/Response |
| 7 | skill_07_sse_streaming.md | SSE generator, stream response formatting, dummy test mode |
| 8 | skill_08_basler_camera.md | Basler camera support, Pylon SDK, emulation, formats |
| 9 | skill_09_docker_build_deploy.md | Docker build, deploy, .env configuration |
| 10 | skill_10_test_suite.md | Test suite coverage, assertions, running tests |
| 11 | skill_11_env_variables.md | All environment variables reference |
| 12 | skill_12_evaluation_workflow.md | End-to-end eval workflow: static checks, build, launch, tests, API/camera/Kafka checks, report |
| 13 | skill_13_verification_curl.md | Verification steps and curl examples |
| 14 | skill_14_implementation_checklist.md | Implementation checklist: file copy list, generated files, Docker prereqs, verification |
| 15 | skill_15_latency_measurement.md | TTFC and C2C latency measurement for file input via SSE streaming |
| 16 | skill_16_message_schema.md | Kafka message schema selection (JSON default vs NvProtoSchema) and extending messages with custom data |
| 17 | skill_17_camera_latency_measurement.md | Camera / live-stream chunk_e2e latency measurement using internal pipeline timestamps |
| 18 | skill_18_rtsp_streaming_output.md | OPT-IN RTSP streaming output: tee1-tap re-stream, RTSPStreamingServer, SW_ENCODER toggle. Generate only when user explicitly requests RTSP |
For end-to-end evaluation, read § 12 first; load build/test/curl/latency/camera/Kafka as needed.
§ 18 is opt-in — generate only when the user explicitly requests RTSP output; otherwise skip § 18 and the RTSP_* rules below.
Key Files Map
The full source-to-target file mapping lives in
skill_14_implementation_checklist.md:
- Files copied verbatim from
references/(non-trivial algorithms — cycle detection, qwen_vl_utils preprocessing, DeepStreamIOptionsAPI, protobuf sources) with the rationale per file. - Files copied as adaptable templates (Dockerfile, compose.yaml, Triton
config and
model.py,ddm_net.py, Pylon emulation config, etc.). - Files generated from skill sections — each annotated with the Critical Rules below that the generation must follow exactly.
- Docker build prerequisites and post-build verification checklist.
Config files (nvds_preprocess_template.txt, nvds_inference_template.txt,
vlm_prompts.txt) are used as-is from configs/.
When skill_06b is loaded, read configs/actions.json from the project root and run the
§ 6b-G generation workflow to produce nvds_action_detector/missing_number_detector.py.
If configs/actions.json is absent or invalid, fall back to copying the reference file.
Critical Rules
Each rule's full detail lives in the linked
skill_NN_*.mdreference file.
| Tag | Rule summary | Details in |
|-----|---|---|
| MANAGER_INIT_IN_MAIN | SOPProcessManager init in main() before uvicorn.run() — not inside lifespan() | skill_01_fastapi_endpoints.md |
| NAMED_KWARGS | create_video_processor() uses named kwargs; camera args as separate kwargs | skill_06_sop_process_manager.md |
| LIVE_REQUIRES_STREAM_TRUE | stream: true required for live inputs (RTSP / camera) | skill_08_basler_camera.md |
| VLM_DISABLED_DISABLES_SOP_CHECKER | `DISABLE_VLM_INFEREN
Truncated for display — read the full file on GitHub.
Related Skills
Agent-Reach
86.0kGive your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
headroom
74.0kCompress tool outputs, logs, files, and RAG chunks before they reach the LLM. 20% fewer tokens for coding agents, 60-95% fewer tokens for JSON, same answers. Library, proxy, MCP server.
crawl4ai
84.4kOpen-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Scrapling
84.4k🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
