scientific-critical-thinking
Evaluating scientific evidence and claims. Covers study design hierarchy (RCT to expert opinion), effect sizes (OR, RR, NNT, Cohen's d), confounding, p-value vs clinical significance, GRADE quality assessment, reproducibility, and bias types (selection, information, confounding, reporting)
Install / Use
npx skills add jaechang-hits/SciAgent-Skills --skill scientific-critical-thinkingInstalls into whichever agent you are using.
SKILL.md
Installable skill definition
Quality Score
Category
Healthcare & Life SciencesSupported Platforms
Our assessment of scientific-critical-thinking
scientific-critical-thinking scores 89/100 on our quality scale, 38th of 48 Healthcare & Life Sciences skills we index.
Its SKILL.md is 18 KB long, well organised into 14 sections with 2 code examples: a thorough specification that gives an agent plenty to work with.
It has 367 GitHub stars, a meaningful sign that others use it.
Maintenance, license and trust
- The repository was last updated 37 days ago, so scientific-critical-thinking is actively maintained.
- No license is declared. By default that means all rights are reserved: you can read it, but reusing or redistributing it is not clearly permitted. Ask the author before building on it commercially.
- Its trust signals score 88/100, with 1 caution from licensing, adoption, age or documentation. These come from repository metadata, not a code audit — read the skill file before letting an agent act on it.
Safety scan
No issues foundOur scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands.
Automated pattern scan on 2026-10-05. It catches known dangerous patterns, not every risk — read a skill before letting an agent act on it.
scientific-critical-thinking compared with similar skills
All 4 of these similar skills score higher than scientific-critical-thinking; compare them before choosing.
| Skill | Score | Stars | Updated | Format |
|---|---|---|---|---|
| scientific-critical-thinking (this skill)by jaechang-hits | 89 | 367 | 37d ago | SKILL.md |
| algorithmic-artby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| pptxby anthropics | 100 | 177.9k | 12d ago | SKILL.md |
| designby nextlevelbuilder | 100 | 130.2k | 13d ago | SKILL.md |
| ui-ux-pro-maxby nextlevelbuilder | 100 | 130.2k | 13d ago | SKILL.md |
Frequently asked questions
- How do I install scientific-critical-thinking?
- Run
npx skills add jaechang-hits/SciAgent-Skills --skill scientific-critical-thinking. The install tabs above show the steps for each supported agent. - Which AI agents does scientific-critical-thinking work with?
- It is written for Universal, as a SKILL.md file. Other agents that read the same format can often use it too.
- Is scientific-critical-thinking safe to use?
- Our scan of the whole file found no instruction hijacking, hidden characters, credential access, data exfiltration or destructive commands. It declares no license and scores 88/100 on trust signals. Skills are instructions an agent will follow, so read the file before installing it and do not approve commands you do not understand.
- Is scientific-critical-thinking still maintained?
- The repository was last updated 37 days ago, so scientific-critical-thinking is actively maintained.
Skill content
View source on GitHubname: "scientific-critical-thinking" description: "Evaluating scientific evidence and claims. Covers study design hierarchy (RCT to expert opinion), effect sizes (OR, RR, NNT, Cohen's d), confounding, p-value vs clinical significance, GRADE quality assessment, reproducibility, and bias types (selection, information, confounding, reporting). Use when reading a paper or assessing claims." license: "CC-BY-4.0"
Scientific Critical Thinking: Evaluating Evidence and Claims
Overview
Scientific critical thinking is the disciplined application of logical and methodological standards to evaluate whether a study's design, analysis, and interpretation support its conclusions. It is the skill that separates a researcher who synthesizes evidence from one who accumulates it. This guide covers the hierarchy of evidence, the mechanics of common biases, effect size interpretation, the p-value controversy, GRADE evidence grading, and common logical fallacies in the interpretation of scientific literature.
Key Concepts
1. Study Design Hierarchy
Study designs vary in their ability to support causal inference. The hierarchy below applies to questions about the effect of an intervention or exposure on an outcome:
Systematic reviews and meta-analyses of RCTs (highest causal certainty)
↓
Randomized Controlled Trials (RCTs)
↓
Non-randomized controlled trials / cluster-randomized trials
↓
Prospective cohort studies (follow exposure → outcome forward in time)
↓
Retrospective cohort studies
↓
Case-control studies (compare exposed vs. unexposed given outcome)
↓
Cross-sectional studies (measure exposure and outcome simultaneously)
↓
Case series and case reports
↓
Expert opinion, mechanistic reasoning, animal models (lowest causal certainty)
Important exceptions: For questions about rare outcomes, case-control designs are often more efficient than cohort studies. For questions about diagnostic accuracy, randomized designs are usually inappropriate — cross-sectional or cohort designs with verified reference standards are preferred. For harm questions, RCTs are often infeasible (ethical constraints), making large cohort studies the best available evidence.
2. Effect Measures and Their Interpretation
Effect measures quantify the relationship between an exposure/intervention and an outcome. Confusing them is a leading source of misinterpretation.
| Measure | Formula | Use case | Key interpretation | |---------|---------|----------|-------------------| | Risk Ratio (RR) | Risk in exposed / Risk in unexposed | Cohort studies, RCTs | RR = 2.0: exposed group has twice the risk | | Odds Ratio (OR) | Odds in exposed / Odds in unexposed | Case-control studies, logistic regression | Approximates RR when outcome is rare (<10%); overestimates RR for common outcomes | | Hazard Ratio (HR) | Instantaneous event rate ratio | Survival analysis (Cox regression) | HR = 0.7: 30% lower hazard of event per time unit in treated group | | Number Needed to Treat (NNT) | 1 / Absolute Risk Reduction | Clinical decision-making | NNT = 20: treat 20 patients to prevent 1 event | | Absolute Risk Reduction (ARR) | Risk_control − Risk_treated | Clinical impact | ARR = 2%: intervention reduces absolute event rate by 2 percentage points | | Cohen's d | (μ₁ − μ₂) / σ_pooled | Continuous outcomes, psychology | d = 0.2 small; 0.5 medium; 0.8 large | | Pearson r | Correlation coefficient | Association, not causal | r = 0.1 small; 0.3 medium; 0.5 large (Cohen 1988) |
Common error: Reporting only the relative risk reduction (e.g., "50% reduction in risk") without the absolute risk reduction. A treatment that reduces risk from 2% to 1% has a 50% relative reduction but only 1% absolute reduction (NNT = 100). The relative measure appears more impressive but the absolute measure is clinically relevant.
3. Bias Types
Bias is systematic deviation of results or inferences from the truth. Unlike random error (reduced by larger samples), bias is directional and not correctable by increasing sample size.
Selection bias: Systematic difference in characteristics between those selected and not selected for study.
- Examples: Healthy worker effect (workers healthier than general population), loss to follow-up bias (sicker patients drop out), volunteer bias
- Detection: Compare baseline characteristics of included vs. excluded participants; assess attrition patterns
Information bias: Systematic error in measuring exposure or outcome.
- Recall bias: Cases remember exposure better than controls (especially case-control studies for rare diseases)
- Observer bias: Assessors aware of exposure status rate outcomes differently
- Detection: Look for blinding of outcome assessors; validated measurement instruments; objective vs. self-reported outcomes
Confounding: A variable associated with both the exposure and the outcome, creating a spurious or masked association.
- Positive confounding: Confounder inflates the observed association
- Negative confounding: Confounder masks a true association
- Detection: Compare crude and adjusted effect estimates; large change (>10%) indicates confounding
- Control methods: Randomization (RCTs), multivariable adjustment, propensity score methods, restriction, matching
Reporting bias: Selective reporting of outcomes or results based on their statistical significance or direction.
- Publication bias: Positive results are published; null results are not
- Outcome reporting bias: Predefined outcomes not reported if non-significant; unplanned analyses reported if significant
- Detection: Compare registered protocol vs. published outcomes; funnel plot asymmetry for meta-analyses
4. The P-value and Statistical vs. Clinical Significance
A p-value is the probability of observing results at least as extreme as those obtained, under the null hypothesis. It is NOT the probability that the null hypothesis is true, nor the probability that the finding is a false positive.
Correct interpretation: p < 0.05 means that, if the null hypothesis were true, fewer than 5% of equally designed studies would produce results this extreme or more. It does NOT indicate the effect is large, clinically meaningful, or replicable.
Statistical significance ≠ clinical significance: With large enough samples, even trivially small effects become statistically significant. A blood pressure drug that reduces systolic BP by 0.8 mmHg (95% CI 0.3–1.3, p = 0.001) is statistically significant but clinically irrelevant.
Confidence intervals are more informative than p-values: A 95% CI of [0.5 kg, 35 kg weight loss] and a 95% CI of [0.5 kg, 1.5 kg weight loss] can both have p < 0.05, but the clinical implications are vastly different. Always focus on CI width and range, not just whether it excludes the null.
5. GRADE Evidence Certainty
GRADE classifies the certainty of evidence across four levels:
| GRADE level | Meaning | Typical starting point | |-------------|---------|----------------------| | High | Further research very unlikely to change confidence | Consistent RCTs, large effect, no bias | | Moderate | Further research likely to have important impact | RCTs with limitations, or strong consistent observational | | Low | Further research very likely to have important impact | Observational studies, or RCTs with serious limitations | | Very low | Any estimate is very uncertain | Case series, expert opinion, very inconsistent results |
GRADE certainty can be downgraded for: risk of bias, inconsistency (heterogeneity), indirectness (different population/outcome), imprecision (wide CI), and publication bias. It can be upgraded for: large effect (OR > 5), dose-response relationship, or all plausible confounders would reduce the effect.
Decision Framework
How should I evaluate this study?
│
├── Step 1: What question is being answered?
│ ├── Intervention effectiveness → Need RCT or high-quality cohort
│ ├── Diagnostic accuracy → Need cross-sectional vs reference standard
│ ├── Prognosis → Need prospective cohort
│ └── Harm / rare exposure → Case-control or large cohort acceptable
│
├── Step 2: Is the study design appropriate?
│ ├── Design matches question → Proceed
│ └── Mismatch → Major limitation (flag)
│
├── Step 3: What are the key threats to validity?
│ ├── Selection bias → Who was included/excluded? Loss to follow-up?
│ ├── Information bias → Blinding? Validated instruments?
│ └── Confounding → What was adjusted for? Residual confounders?
│
├── Step 4: Are the effect estimates clinically meaningful?
│ ├── Effect size large enough to matter clinically?
│ ├── CI narrow enough to be informative?
│ └── Absolute vs relative risk reported?
│
└── Step 5: How certain is the evidence overall? (GRADE)
├── High certainty → Confident conclusion
├── Moderate certainty → Likely true; note limitations
├── Low certainty → Uncertain; more research needed
└── Very low certainty → Cannot draw conclusions
| Claim type | Appropriate response | Red flags requiring skepticism | |------------|---------------------|-------------------------------| | "Drug X reduces mortality by 50%" | Ask: 50% relative or absolute? What was baseline risk? | Only relative risk reported; no CI provided | | "Observational study shows cause" | Downgrade to "association"; list plausible confounders | Authors use "causes" without adjustment | | "Significant p-value proves effect" | Check effect size and CI; assess clinical relevance | p = 0.04 with N = 50,000 and tiny effect | | "Single RCT is definitive" | Check for replication; assess risk of bias | Funded by manufacturer; no blinding | | "Preprint shows breakthrough" | Await peer review; check for reproducibility | No data/code sharing; sensational press release | | "N-of-1 case report demonstrates treatment" | Note limited generalizability; no control | Used to support policy without cohort evidence |
Best Practices
-
Read the Methods before the Results: The Discussion section is written by the authors to support their conclusions. The Methods section is where you independently assess whether the data can support those conclusions. Specifically: what were the pre-specified primary outcomes? Do the reported outcomes match those in the Methods or the registered protocol?
-
Always seek the pre-registration record: ClinicalTrials.gov, PROSPERO, and OSF registrations contain the original protocol. Comparing the pre-specified primary outcome to what was reported in the abstract is the single most efficient check for outcome reporting bias. A change in primary outcome without explanation is a major red flag.
-
Distinguish statistical and clinical significance explicitly: For every effect estimate, ask: if this effect is real, would it matter to a patient or a biological system? A genomic variant with OR = 1.05 (p = 1e-12) in a GWAS of 500,000 people is a genuine association but contributes negligibly to disease risk prediction.
-
Identify the funding source and conflicts of interest: Industry-funded trials are not automatically invalid, but industry funding is associated with more favorable outcomes for the sponsor's product. Assess whether conflicts are disclosed, whether the funder had access to data or participated in analysis, and whether an independent statistician reviewed the data.
-
Check for multiple testing without correction: When a paper tests 20 outcomes, 1 will be statistically significant at p < 0.05 by chance alone. Look for corrections (Bonferroni, Benjamini-Hochberg FDR) in genomics, proteomics, and other high-throughput studies. Absence of correction in a paper reporting 50 comparisons invalidates the significance claims.
-
Require absolute risk data before accepting clinical conclusions: For any binary outcome, request (or calculate) the absolute r
Truncated for display — read the full file on GitHub.
Related Skills
algorithmic-art
177.9kCreating algorithmic art using p5.js with seeded randomness and interactive parameter exploration. Use this when users request creating art using code, generative art, algorithmic art, flow fields, or particle systems.
pptx
177.9kUse this skill any time a .pptx or .potx file is involved in any way — as input, output, or both. This includes: creating slide decks, pitch decks, or presentations; reading, parsing, or extracting text from any .pptx or .potx file (even if the extracted content will be used elsewhere, like in an em…
design
130.2kComprehensive design skill: brand identity, design tokens, UI styling, logo generation (55 styles, Gemini, Atlas Cloud, or MuAPI AI), corporate identity program (50 deliverables, CIP mockups), HTML presentations (Chart.js), banner design (22 styles, social/ads/web/print), icon design (15 styles, SVG…
ui-ux-pro-max
130.2kUI/UX design intelligence for web, mobile, and desktop. This skill should be used when designing, building, reviewing, or fixing interfaces, including pages, components, design systems, accessibility, interaction, responsive layout, typography, color, charts, and stack-specific UI implementation.
Languages
Trust signals
From repository metadata: license, adoption, age and documentation. Not a code audit — see the Safety scan above for what the skill file itself contains.
