SkillAgentSearch skills...

digital-health-clinical-asr-eval

Stage 3 of Clinical ASR Flywheel. Score a NeMo manifest, produce the five-section KER leaderboard (by-ipa_source diagnostic). Not for ASR auth (/riva-asr).

Install / Use

npx skills add NVIDIA/skills --skill digital-health-clinical-asr-eval

Installs into whichever agent you are using.

About this skill
📄

SKILL.md

Installable skill definition

Quality Score

92/100

Supported Platforms

Universal

Our assessment of digital-health-clinical-asr-eval

digital-health-clinical-asr-eval scores 92/100 on our quality scale, 8th of 40 Healthcare & Life Sciences skills we index (top 20%).

Its SKILL.md is 18 KB long, well organised into 18 sections with 1 code example: a thorough specification that gives an agent plenty to work with.

With 3,421 GitHub stars, it is one of the more widely adopted skills in the catalogue.

Substance
30/30
Structure
17/20
Description
15/15
Adoption
15/20
Freshness
15/15

Maintenance, license and trust

  • The repository was last updated 5 days ago, so digital-health-clinical-asr-eval is actively maintained.
  • It is released under the Apache-2.0 license, a permissive license that allows use, modification and commercial use with attribution.
  • Its trust signals score 100/100, with no cautions. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.

digital-health-clinical-asr-eval compared with similar skills

All 4 of these similar skills score higher than digital-health-clinical-asr-eval; compare them before choosing.

SkillScoreStarsUpdatedFormat
digital-health-clinical-asr-eval (this skill)by NVIDIA923.4k5d agoSKILL.md
Agent-Reachby Panniantong10086.0k13d agoCLAUDE.md
headroomby headroomlabs-ai10074.0ktodayCLAUDE.md
crawl4aiby unclecode10084.4k3d agoMCP Server
Scraplingby D4Vinci10084.4ktodayMCP Server

Frequently asked questions

How do I install digital-health-clinical-asr-eval?
Run npx skills add NVIDIA/skills --skill digital-health-clinical-asr-eval. The install tabs above show the steps for each supported agent.
Which AI agents does digital-health-clinical-asr-eval work with?
It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
Is digital-health-clinical-asr-eval safe to use?
It is Apache-2.0-licensed and scores 100/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
Is digital-health-clinical-asr-eval still maintained?
The repository was last updated 5 days ago, so digital-health-clinical-asr-eval is actively maintained.

name: "digital-health-clinical-asr-eval" description: "Stage 3 of Clinical ASR Flywheel. Score a NeMo manifest, produce the five-section KER leaderboard (by-ipa_source diagnostic). Not for ASR auth (/riva-asr)." version: "1.1.0" author: "Ben Randoing brandoing@nvidia.com" tags:

  • clinical-asr
  • eval
  • ker
  • leaderboard
  • flywheel tools:
  • Read
  • Write
  • Bash
  • Skill license: Apache-2.0 compatibility: "NVIDIA_API_KEY (required) for hosted ASR NIMs via NVCF. A NeMo-format manifest produced by /digital-health-clinical-asr-build (or an externally-provided manifest carrying the clinical-extension fields). All ASR call shapes and WER/CER/KER/SER scoring recipes are inlined — no sibling agent skill required." metadata: author: "Ben Randoing brandoing@nvidia.com" tags:
    • clinical-asr
    • flywheel
    • eval
    • ker
    • leaderboard team: healthcare-tme domain: ai-ml stage: 3 previous_skill: digital-health-clinical-asr-build next_skill: digital-health-clinical-asr-finetune

<!-- SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. SPDX-License-Identifier: Apache-2.0 -->

Clinical ASR Flywheel — Stage 3 (Eval)

⚠ Agent: read the Critical Workflow Rules section below before answering. This SKILL.md is self-contained — evals/, references/, and assets/ are pointers, not load-bearing. Answer methodology questions from this file directly; only invoke tools when the user explicitly asks to execute against a real manifest.

You are the score-and-route stage. The user arrives with a NeMo-format manifest.jsonl (either from /digital-health-clinical-asr-build or carried in from elsewhere). You transcribe it via the chosen ASR NIM, score four metrics, produce a five-section leaderboard, and read the decision tree to decide whether the user should advance to /digital-health-clinical-asr-finetune, loop back to /digital-health-clinical-asr-build, or stop and harden the eval.

This skill does not generate audio. If the manifest is missing or empty, send the user back to /digital-health-clinical-asr-build.

Audio leaves your environment — disclose this to the user before any clip is sent

This stage transmits each manifest row's WAV file plus its reference text to an external NVIDIA service. Surface this before invoking the first ASR call:

| Service | What gets sent | When | |---|---|---| | NVIDIA NVCF Parakeet/Nemotron ASR (grpc.nvcf.nvidia.com) | Every audio clip referenced by the manifest (raw PCM bytes), plus the reference transcript and the clinical-extension metadata for scoring | Step 3b, one call per manifest row |

The clips should be synthetic audio generated by Stage 2 (Magpie TTS over a user-curated term list) — not real patient audio. Do not pass real ASR recordings, real patient encounters, or any PHI through this skill. Scoring then runs locally (pure-Python WER/CER/KER/SER, or jiwer if installed). The scoring step itself does not transmit anything; only the ASR step does.

Critical workflow rules (apply on every activation)

For methodology questions (leaderboard structure, KER definition, decision tree), answer from this file. Don't invoke tools, call other skills, or run scripts unless the user explicitly asks to execute against a real manifest. Surface these facts in any response:

  1. Off-ramp first. If the user is asking about something outside scoring, route and stop without running any workflow:
    • ASR model-catalog selection / comparison / alternative NIMs → /riva-asr
    • ASR auth (API keys, bearer tokens, function IDs) → /riva-asr
    • ASR gRPC protocol, streaming, batching, chunking, retries → /riva-asr
    • NIM deploy / riva-build / riva-deploy → /riva-asr-custom
    • NGC / Docker / NVIDIA Container Toolkit → /riva-nim-setup
    • No manifest yet → /digital-health-clinical-asr-build
    • Wants to fine-tune now with a known KER → /digital-health-clinical-asr-finetune
  2. Default ASR NIM is nvidia/parakeet-tdt-0.6b-v2 (NVCF function-id d3fe9151-442b-4204-a70d-5fcc597fd610, offline gRPC). Env-var overrides: ASR_MODEL_NAME (leaderboard display name), ASR_NVCF_FUNCTION_ID (swap to a different hosted NIM — e.g. Whisper Large v3 b702f636-… while the Parakeet backend is faulting, or a fine-tuned NIM), ASR_ENDPOINT (self-hosted gRPC; takes precedence). Echo the chosen NIM and the resolved function-id back before spending API credits.
  3. ASR transcription is inlined in Step 3b (NVCF gRPC + riva.client.ASRService.offline_recognize, same auth pattern as Stage 1). For deeper protocol/auth questions, alternative NIM catalogs, or self-hosted Riva NIM configuration, defer to /riva-asr.
  4. KER is the headline. Per-row check: the flagged term words must appear in order, contiguous, adjacent in the normalized hypothesis. cefazolin → cefa zolin is a miss. Aggregate WER hides clinically dangerous failures; both are reported, KER is the gate.
  5. The by-ipa_source split is the most informative single number in the leaderboard. The merriam-webster vs magpie_g2p delta proves the SSML override pipeline is doing real work. Read it aloud to the user.
  6. Special-case routing. merriam-webster rows good, magpie_g2p rows bad → pronunciation-coverage gap, not a model gap. Route back to /digital-health-clinical-asr-build Step 2d. Do NOT recommend /digital-health-clinical-asr-finetune as a first response.
  7. Five-section leaderboard order. Headline (WER/CER/KER/SER) → KER by entity_category → KER by ipa_source → KER by noise_level → Per-term KER worst-first. The by-ipa_source section is mandatory; it is the proof the SSML pipeline works.

Purpose

Score a clinical-ASR manifest, produce a five-section KER leaderboard, and route the user via the post-eval decision tree. Methodology details (metric definitions, normalization, leaderboard order, special-case routing) live in Critical Workflow Rules above and Instructions below.

When to use this skill

Activate on user phrases like:

  • "Score my ASR manifest"
  • "What's the KER on Parakeet TDT v2?"
  • "Run the eval on cycle-N"
  • "Compare two ASR models on the clinical benchmark"
  • "Generate the leaderboard"
  • "I have a manifest.jsonl, how do I score it?"
  • "Why is KER 0.4 when WER is 0.07?"
  • "Should we fine-tune?" (this is the eval-side question — the post-eval decision tree lives in this skill)

Literal-keyword non-activation check — if the user's message contains any of authenticate, API key, bearer, function ID, gRPC, streaming, chunking, batching, transcription retry, riva-build, riva-deploy, NIM deploy, NGC, Docker, Container Toolkit, or asks "which ASR model is best" / "compare models" / "vendor differences" — do NOT activate the scoring workflow. Apply Critical Workflow Rule #1 above to route to the right sibling skill and stop. This applies even if the user mentions "KER" or "eval" alongside the keyword.

Prerequisites

  • A NeMo-format manifest with the clinical extension fields (term, entity_category, ipa_source, voice_id, noise_level, context_type). The schema is documented in the build skill's references/manifest-schema.md.
  • NVIDIA_API_KEY exported (Stage 1 prerequisite still applies).
  • nvidia-riva-client + soundfile installed (Stage 1 prerequisite). For self-hosted Riva NIM details, see /riva-asr Option B.
  • Audio files actually present on disk — run the audio-existence pre-flight from the manifest-schema reference before spending API credits.

Instructions

3a. Pick the ASR NIM

Default: nvidia/parakeet-tdt-0.6b-v2 via NVCF gRPC (offline), function-id d3fe9151-442b-4204-a70d-5fcc597fd610. NVIDIA's current English ASR recommendation — fastest/cheapest in the catalog, and supported in NeMo's stock SFT recipe so the Stage 3 baseline and a Stage 4 fine-tune ride the same model family.

Three runtime env-var override knobs (ASR_MODEL_NAME for leaderboard display, ASR_NVCF_FUNCTION_ID to swap to a different hosted NIM, ASR_ENDPOINT for self-hosted gRPC) plus the full alternate-NIM catalog (Parakeet TDT 1.1B, Parakeet CTC 1.1B, Whisper Large v3, Nemotron streaming) with function IDs and call-shape notes: references/offline-asr-recipe.md.

Echo the chosen NIM, the resolved function-id, and any env-var overrides to the user before spending API credits. A 200-row manifest on hosted Parakeet TDT v2 is cheap; an accidental run against the wrong model on a 1,000-row manifest is not.

3b. Transcribe

For each row in manifest.jsonl, transcribe audio_filepath and write per_sample.json (one JSON object per row, JSONL or a JSON array — caller's choice):

{
  "audio_filepath": "...",
  "ref": "<row.text>",
  "hyp": "<asr output>",
  "term": "<row.term>",
  "entity_category": "<row.entity_category>",
  "ipa_source": "<row.ipa_source>",
  "voice_id": "<row.voice_id>",
  "noise_level": "<row.noise_level>",
  "context_type": "<row.context_type>"
}

Recipe (full Python in references/offline-asr-recipe.md): transcribe_manifest(api_key, manifest_path, out_path, language_code="en-US") opens an offline gRPC stream to NVCF (or to ASR_ENDPOINT if set for self-hosted Riva), calls riva.client.ASRService.offline_recognize per row — sentences in a clinical manifest are ≤ 30 s so no streaming/batching needed — and writes the JSONL above. Same auth_for shape as the Stage 1 setup smoke test. The agent harness passes api_key explicitly; the recipe reads the three env-var overrides (ASR_NVCF_FUNCTION_ID, ASR_MODEL_NAME, ASR_ENDPOINT) at the top so auditors see the knobs in one place.

Whisper fallback (when Parakeet's NVCF backend faults with CUDA illegal-memory-access from Triton) and self-hosted Riva NIM (ASR_ENDPOINT=localhost:50051) env-var patterns: see references/offline-asr-recipe.md (§Whisper fallback, §Self-hosted Riva NIM).

Resilience knobs deferred to the user. If NVCF returns RESOURCE_EXHAUSTED mid-batch, the loop raises on that row; re-run from the failing row. Streaming/batching/retry-with-backoff are out of scope — see /riva-asr.

3c. Score four metrics

For every row, compute:

| Metric | What it measures | Why we keep it | |---|---|---| | WER | Word error rate (Levenshtein on tokens, after normalization) | Industry standard; blunt instrument for clinical | | CER | Character error rate | Catches near-misses on long compound names | | KER ★ | Keyword error rate — did the flagged term appear in the hypothesis (normalized, contiguous match)? | Headline clinical signal | | SER | Sentence error rate (1 if any wrong, 0 if perfect) | Sanity bound; what the doctor experiences |

Normalization (apply to both ref and hyp before all four metrics):

  1. Lowercase.
  2. NFKD-normalize (smart quotes → ASCII, etc.).
  3. Strip punctuation except hyphen.
  4. Collapse whitespace runs to a single space.

Inline scoring recipes — normalize / edit_distance / wer / cer / ker / ser (pure-Python, no jiwer dependency): see references/scoring-recipes.md. Aggregate across rows by taking mean(per-row score) for each metric.

Strict KER — term words must appear in order, adjacent in the normalized hypothesis. This is conservative: cefazolin → cefa zolin counts as a miss. That's the right call clinically — a downstream pharmacy lookup will fail on the misspelled token.

KER does not punish surrounding errors. A row where the term is correct and the rest of the sentence is garbage still scores KER=0; the WER on that row will surface the broader problem separately.

3d. Breakdowns + leaderboard

Write a five-section markdown leaderboard, in this order:

  1. Headline — overall WER, CER, KER, SER for the chosen model.
  2. KER by entity_category — drug vs procedure vs anatomy vs ... This i

Truncated for display — read the full file on GitHub.

Related Skills

View on GitHub
GitHub Stars3.4k
CategoryHealthcare
Updated5d ago
Forks412

Languages

Python

Trust signals

100/100

From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.

No cautions