SkillAgentSearch skills...

deepstream-sop

Use this skill when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection (GEBD) plus VLM classification.

Install / Use

npx skills add NVIDIA/skills --skill deepstream-sop

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

90/100

Category

Operations

Supported Platforms

Universal

Our assessment of deepstream-sop

deepstream-sop scores 90/100 on our quality scale, 185th of 487 Operations skills we index (top 38%).

Its SKILL.md is 18 KB long, split into 6 sections with 1 code example: a thorough specification that gives an agent plenty to work with.

With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
15/20
Description
15/15
Adoption
15/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 5 days ago, so deepstream-sop is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

deepstream-sop compared with similar skills

All 4 of these similar skills score higher than deepstream-sop; compare them before choosing.

SkillScoreStarsUpdatedFormat
deepstream-sop (this skill)by NVIDIA903.4k5d agoSKILL.md
Agent-Reachby Panniantong10086.0k13d agoCLAUDE.md
headroomby headroomlabs-ai10074.0ktodayCLAUDE.md
crawl4aiby unclecode10084.4k3d agoMCP Server
Scraplingby D4Vinci10084.4ktodayMCP Server

Frequently asked questions

How do I install deepstream-sop?
Run npx skills add NVIDIA/skills --skill deepstream-sop. The install tabs above show the steps for each supported agent.
Which AI agents does deepstream-sop work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is deepstream-sop safe to use?
It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is deepstream-sop still maintained?
The repository was last updated 5 days ago, so deepstream-sop is actively maintained.

name: "deepstream-sop" description: > Use this skill when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection (GEBD) plus VLM classification. Trigger even if the user does not name it: verify operator step sequence, detect missing or out-of-order SOP steps, score factory/work-cell video for procedure compliance, run VLM-based SOP checking on industrial cameras, or call /v1/chat/completions with a file, RTSP, or Basler camera. Also trigger for its internals: SOPVideoProcessor, DeepStream GEBD model (e.g. DDM) via Triton CAPI, nvds_custom_postprocess, Cosmos Reason 1/2 vLLM, SSE streaming, Kafka NvProto/JSON output, Basler/Pylon camera + emulation, Docker compose, chunk-level latency. Do NOT trigger for generic DeepStream pipelines, object detection/tracking, NIM imports, or video summarization. owner: "windy@nvidia.com" service: "deepstream-sop" version: "1.0.0" license: "CC-BY-4.0 AND Apache-2.0" reviewed: "2026-04-08" metadata: author: "Wind Yuan windy@nvidia.com" tags: - deepstream - sop - vlm - triton - gpu languages: - python frameworks: - deepstream - triton - fastapi domain: video-analytics

DeepStream SOP Inference Microservice Skill

This skill guides AI coding assistants in building, extending, and debugging the NVIDIA DeepStream SOP (Standard Operating Procedure) Inference Microservice — a GPU-accelerated pipeline for temporal action detection and VLM-based SOP compliance monitoring on industrial video feeds.

Reference repository: https://github.com/NVIDIA/sop-monitoring-blueprints/tree/main/microservices/sop-inference-bp Local reference code: sop-inference-bp/ directory (from a local clone of the repository)


Models

Model-agnostic at both inference stages — swap via env var (and Triton dir for GEBD).

| Stage | Role | Model class | Default | Swap via | |------|------|-------------|---------|----------| | Stage 1 (CV) | Per-frame boundary scoring → chunk segmentation | Generic Event Boundary Detection (GEBD) | DDM (MCG-NJU/DDM) via Triton Python backend | Replace triton_model_repo/<model>/ + DDM_MODEL_PATH (§ 5) | | Stage 3 (VLM) | Per-chunk action classification | Vision-language model via vLLM | Cosmos Reason 1 7B (Reason 2 also supported) | Set VLLM_MODEL_PATH to a different HF ID or local path |

"GEBD" = swappable Stage-1 slot; "DDM" = the default architecture (terms used interchangeably).

Chunking is selectable per request (§ 2): default ddm-net uses GEBD; uniform produces fixed-length chunks and bypasses Stage-1 GEBD (§ 3, § 6). DDM temporal window is configurable via FRAMES_PER_SIDE / SEQUENCE_BATCH (§ 4, § 5), with optional TensorRT (§ 5).


Architecture Overview

Runs in a Docker container (nvds-action-sop) alongside a Kafka container. Full diagram: references/sop_architecture.svg.

Data flow through the 4-stage SOPVideoProcessor pipeline (per-request):

Input Sources                    Docker Container: nvds-action-sop
─────────────                    ──────────────────────────────────────────────────
Video Files ──┐                  FastAPI Server (port 8300)
RTSP Streams ─┤── base64/       ├─ /v1/chat/completions → SOPProcessManager
Basler Camera ┘   file/rtsp/       │
                  camera           │ ModelInitializer: VLM first, then DDM dummy pipeline
                                   │ 4 Thread Pools: cv(32), clip(32), vlm(64), vlm_req(64)
                                   │
                                   ▼ SOPVideoProcessor (per-request)
                                   ┌────────────────────────────────────────────────┐
                                   │ Stage 1: DeepStream Pipeline (GPU)             │
                                   │   Source → nvstreammux → tee1                  │
                                   │    ├─[inference] queue1 → nvdspreprocess       │
                                   │    │  → nvinferserver (Triton CAPI + DDM)      │
                                   │    │  → InferOutputTensorParser → score_queue  │
                                   │    ├─[frames]  queue3 → nvvideoconvert         │
                                   │    │  → capsfilter → appsink                   │
                                   │    │  → DecodedFrameRetriever → frame_queue    │
                                   │    └─[RTSP out] queue → convert → H.264 enc    │  (optional, § 18)
                                   │       → rtppay → udpsink → RTSPServer (§ 18)   │  opt-in only
                                   │              │ boundary scores                 │
                                   │              ▼                                 │
                                   │ Stage 2: Clip Post-Process                     │
                                   │   Boundary detection → chunk segmentation      │
                                   │              │ video frames + timestamps        │
                                   │              ▼                                 │
                                   │ Stage 3: VLM Inference                         │
                                   │   Embedded vLLM (Cosmos Reason 1/2)            │
                                   │   Frame sampling at VLM_FPS → classification   │
                                   │              │ action labels                    │
                                   │              ▼                                 │
                                   │ Stage 4: SOP Checker                           │
                                   │   Sequence validation → missing/misordered     │
                                   │              │ chunk results                    │
                                   │              ▼                                 │
                                   │         final_queue                            │
                                   └────────────────────────────────────────────────┘
                                          │
Output                                    ▼
──────                             ┌─────────────────┐
SSE Stream (chat.completion.chunk) │ Kafka Messages   │
Non-streaming (chat.completion)    │ (JSON/Protobuf)  │
Prometheus metrics (/v1/metrics)   └────────┬────────┘
                                            ▼
                                   Docker Container: kafka
                                   (apache/kafka:3.7.0)

Section Index

Each section is a standalone file in references/ — load only what your task needs.

| § | File | Responsibility | |---|------|---------------| | 1 | skill_01_fastapi_endpoints.md | FastAPI endpoints, server init, Prometheus metrics | | 2 | skill_02_pydantic_schemas.md | Request/response Pydantic models (api_types.py) | | 3 | skill_03_deepstream_pipeline.md | DeepStream pyservicemaker pipeline, tensor parser, dummy pipeline | | 4 | skill_04_config_templates.md | nvdspreprocess / nvinferserver config templates + rendering | | 5 | skill_05_triton_ddm_model.md | Triton model repo, config.pbtxt, model.py, ddm_net.py | | 5b | skill_05b_custom_postprocess.md | C++ postprocess plugin, Makefile, IOptions API | | 6 | skill_06_sop_process_manager.md | SOPProcessManager, SOPVideoProcessor, VLLMInference, Kafka | | 6b | skill_06b_sop_checker.md | SOP sequence and checker compliance: MissingNumberDetector, SopCheckerCache, SopCheckerRequest/Response | | 7 | skill_07_sse_streaming.md | SSE generator, stream response formatting, dummy test mode | | 8 | skill_08_basler_camera.md | Basler camera support, Pylon SDK, emulation, formats | | 9 | skill_09_docker_build_deploy.md | Docker build, deploy, .env configuration | | 10 | skill_10_test_suite.md | Test suite coverage, assertions, running tests | | 11 | skill_11_env_variables.md | All environment variables reference | | 12 | skill_12_evaluation_workflow.md | End-to-end eval workflow: static checks, build, launch, tests, API/camera/Kafka checks, report | | 13 | skill_13_verification_curl.md | Verification steps and curl examples | | 14 | skill_14_implementation_checklist.md | Implementation checklist: file copy list, generated files, Docker prereqs, verification | | 15 | skill_15_latency_measurement.md | TTFC and C2C latency measurement for file input via SSE streaming | | 16 | skill_16_message_schema.md | Kafka message schema selection (JSON default vs NvProtoSchema) and extending messages with custom data | | 17 | skill_17_camera_latency_measurement.md | Camera / live-stream chunk_e2e latency measurement using internal pipeline timestamps | | 18 | skill_18_rtsp_streaming_output.md | OPT-IN RTSP streaming output: tee1-tap re-stream, RTSPStreamingServer, SW_ENCODER toggle. Generate only when user explicitly requests RTSP |

For end-to-end evaluation, read § 12 first; load build/test/curl/latency/camera/Kafka as needed.

§ 18 is opt-in — generate only when the user explicitly requests RTSP output; otherwise skip § 18 and the RTSP_* rules below.


Key Files Map

The full source-to-target file mapping lives in skill_14_implementation_checklist.md:

  • Files copied verbatim from references/ (non-trivial algorithms — cycle detection, qwen_vl_utils preprocessing, DeepStream IOptions API, protobuf sources) with the rationale per file.
  • Files copied as adaptable templates (Dockerfile, compose.yaml, Triton config and model.py, ddm_net.py, Pylon emulation config, etc.).
  • Files generated from skill sections — each annotated with the Critical Rules below that the generation must follow exactly.
  • Docker build prerequisites and post-build verification checklist.

Config files (nvds_preprocess_template.txt, nvds_inference_template.txt, vlm_prompts.txt) are used as-is from configs/.

When skill_06b is loaded, read configs/actions.json from the project root and run the § 6b-G generation workflow to produce nvds_action_detector/missing_number_detector.py. If configs/actions.json is absent or invalid, fall back to copying the reference file.


Critical Rules

Each rule's full detail lives in the linked skill_NN_*.md reference file.

| Tag | Rule summary | Details in | |-----|---|---| | MANAGER_INIT_IN_MAIN | SOPProcessManager init in main() before uvicorn.run() — not inside lifespan() | skill_01_fastapi_endpoints.md | | NAMED_KWARGS | create_video_processor() uses named kwargs; camera args as separate kwargs | skill_06_sop_process_manager.md | | LIVE_REQUIRES_STREAM_TRUE | stream: true required for live inputs (RTSP / camera) | skill_08_basler_camera.md | | VLM_DISABLED_DISABLES_SOP_CHECKER | `DISABLE_VLM_INFEREN

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars3.4k
CategoryOperations
Updated5d ago
Forks412

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions