目錄
Step 1 qPCR / RT-qPCR
- 2025NEWGuidelineMIQE 2.0: Revision of the Minimum Information for Publication of Quantitative Real-Time PCR Experiments Guidelines.Clin Chem 71(6):634. 10.1093/clinchem/hvaf043
- 2009GuidelineThe MIQE guidelines: minimum information for publication of quantitative real-time PCR experiments.Clin Chem. 10.1373/clinchem.2008.112797
- 2001Analysis of relative gene expression data using real-time quantitative PCR and the 2^−ΔΔCT method.Methods. 10.1006/meth.2001.1262
- 2002geNorm: accurate normalization by geometric averaging of multiple internal control genes.Genome Biol. 10.1186/gb-2002-3-7-research0034
- 2020GuidelineThe Digital MIQE Guidelines Update (dMIQE 2020).Clin Chem 66(8):1012. 10.1093/clinchem/hvaa125
- 2025NEWMIQE 2.0 and the Urgent Need to Rethink qPCR Standards.Int J Mol Sci 26(11):4975. 10.3390/ijms26114975
Step 2 Western Blot 定量
- 2008The use of total protein stains as loading controls.J Neurosci Methods. 10.1016/j.jneumeth.2008.05.003
- 2010Reversible Ponceau staining as a loading control alternative to actin in Western blots.Anal Biochem. 10.1016/j.ab.2010.02.036
- 2013Stain-Free technology as a normalization tool in Western blot analysis.Anal Biochem. 10.1016/j.ab.2012.10.010
- 2013A defined methodology for reliable quantification of Western blot data.Mol Biotechnol. 10.1007/s12033-013-9672-6
- 2020Antibody validation for Western blot: by the user, for the user.J Biol Chem / Anal Biochem. 10.1016/j.ab.2020.113608
- 2025NEWSuperior normalization using total protein for western blot analysis of human adipocytes.PLOS ONE. 10.1371/journal.pone.0328136
Step 3 ELISA & Standard Curves
- 2007Appropriate calibration curve fitting in ligand binding assays.AAPS J. 10.1208/aapsj0902029
- 2005The five-parameter logistic: a characterization and comparison with the four-parameter logistic.Anal Biochem. 10.1016/j.ab.2005.04.035
- 2008Limit of Blank, Limit of Detection and Limit of Quantitation.Clin Biochem Rev 29 Suppl 1:S49. PMID: 18852857
- 2023Scaling of an antibody validation procedure.eLife 12:RP91645. 10.7554/eLife.91645
- 2024NEWYCharOS protocol for antibody validation.Nat Protoc. 10.1038/s41596-024-01108-6
- 2018GuidelineBioanalytical Method Validation Guidance for Industry.FDA.gov
Step 4 Flow Cytometry
- 2008GuidelineMIFlowCyt: the minimum information about a flow cytometry experiment.Cytometry A. 10.1002/cyto.a.20623
- 2021GuidelineGuidelines for the use of flow cytometry and cell sorting in immunological studies (3rd ed.).Eur J Immunol. 10.1002/eji.202170126
- 2020GuidelineMIFlowCyt-EV: a framework for standardized reporting of EV flow cytometry experiments.J Extracell Vesicles. 10.1080/20013078.2020.1713526
- 2001Spectral compensation for flow cytometry.Cytometry. 10.1002/cyto.1163
- 2025NEWSpectral Flow Cytometry: The Current State and Future.Cells. PMC12193525
Step 5 Image Quantification
- 2012Fiji: an open-source platform for biological-image analysis.Nat Methods. 10.1038/nmeth.2019
- 2017QuPath: open source software for digital pathology image analysis.Sci Rep. 10.1038/s41598-017-17204-5
- 2021CellProfiler 4: improvements in speed, utility and usability.BMC Bioinformatics. 10.1186/s12859-021-04344-9
- 2021Cellpose: a generalist algorithm for cellular segmentation.Nat Methods 18:100. 10.1038/s41592-020-01018-x
- 2025NEWCellpose3: one-click image restoration for improved cellular segmentation.Nat Methods 22:592. 10.1038/s41592-025-02595-5
- 1993Measurement of co-localization of objects in dual-colour confocal images.J Microsc. 10.1111/j.1365-2818.1993.tb03313.x
- QUAREP-LiMiInitiativeQuality Assessment and Reproducibility for Instruments & Images in Light Microscopy.quarep.org
Step 6 EDA
- 1977BookExploratory Data Analysis.Addison-Wesley. ISBN 978-0201076165
- 1965An analysis of variance test for normality (complete samples).Biometrika 52:591. 10.1093/biomet/52.3-4.591
- 2011Power comparisons of Shapiro-Wilk, Kolmogorov-Smirnov, Lilliefors and Anderson-Darling.J Stat Modeling and Analytics 2(1):21.
- 2013Detecting outliers: Do not use SD around the mean, use MAD around the median.J Exp Soc Psychol. 10.1016/j.jesp.2013.03.013
- 2025NEWTesting for normality in regression models: mistakes abound (but may not matter).R Soc Open Sci. 10.1098/rsos.241904
Step 7 Hypothesis Testing
- 1947The generalization of "Student's" problem when several different population variances are involved.Biometrika. 10.1093/biomet/34.1-2.28
- 2016StatementThe ASA's Statement on p-Values: Context, Process, and Purpose.Am Stat 70(2):129. 10.1080/00031305.2016.1154108
- 2019Moving to a World Beyond "p < 0.05".Am Stat 73(sup1):1. 10.1080/00031305.2019.1583913
- 2017Why psychologists should by default use Welch's t-test instead of Student's t-test.Int Rev Soc Psychol. 10.5334/irsp.82
- 1993BookAn Introduction to the Bootstrap.Chapman & Hall. ISBN 978-0412042317
Step 8 Multiple Testing Correction
- 1995Controlling the false discovery rate: a practical and powerful approach to multiple testing.JRSS B 57(1):289. 10.1111/j.2517-6161.1995.tb02031.x
- 1979A simple sequentially rejective multiple test procedure.Scand J Stat 6(2):65.
- 2003Statistical significance for genomewide studies.PNAS 100(16):9440. 10.1073/pnas.1530509100
- 1955A multiple comparison procedure for comparing several treatments with a control.JASA 50:1096. 10.1080/01621459.1955.10501294
- 2024NEW2dGBH: Two-dimensional group Benjamini-Hochberg procedure.Bioinformatics 40(2):btae035. 10.1093/bioinformatics/btae035
Step 9 Regression & Dose-Response
- 2015Dose-Response Analysis Using R.PLOS ONE 10(12):e0146021. 10.1371/journal.pone.0146021
- 2004BookFitting Models to Biological Data Using Linear and Nonlinear Regression.Oxford UP. ISBN 978-0195171792
- 2011Guidelines for accurate EC50/IC50 estimation.Pharm Stat 10(2):128. 10.1002/pst.426
- 2012BookFundamentals of Enzyme Kinetics, 4th ed.Wiley-Blackwell.
- 2011The original Michaelis constant: Translation of the 1913 Michaelis-Menten paper.Biochemistry 50:8264. 10.1021/bi201284u
Step 10 Effect Size & Power
- 1988BookStatistical Power Analysis for the Behavioral Sciences, 2nd ed.Routledge. ISBN 978-0805802832
- 1981Distribution theory for Glass's estimator of effect size.J Educ Stat 6(2):107. 10.3102/10769986006002107
- 2007G*Power 3: A flexible statistical power analysis program.Behav Res Methods 39:175. 10.3758/BF03193146
- 2001The abuse of power: the pervasive fallacy of power calculations for data analysis.Am Stat 55:19. 10.1198/000313001300339897
- 2024NEWPower to Detect What? Considerations for Planning and Evaluating Sample Size.Pers Soc Psychol Rev. 10.1177/10888683241228328
Step 11 Omics Quantification
- 2012Measurement of mRNA abundance using RNA-seq data: RPKM measure is inconsistent among samples.Theory Biosci. 10.1007/s12064-012-0162-3
- 2014Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2.Genome Biol. 10.1186/s13059-014-0550-8
- 2010edgeR: a Bioconductor package for differential expression analysis.Bioinformatics. 10.1093/bioinformatics/btp616
- 2014voom: precision weights unlock linear model analysis tools for RNA-seq read counts.Genome Biol. 10.1186/gb-2014-15-2-r29
- 2018Heavy-tailed prior distributions for sequence count data (apeglm).Bioinformatics. 10.1093/bioinformatics/bty895
- 2014Accurate proteome-wide label-free quantification by delayed normalization and MaxLFQ.Mol Cell Proteomics. 10.1074/mcp.M113.031591
Step 12 Visualization & Reporting
- 2016Bookggplot2: Elegant Graphics for Data Analysis.Springer. 10.1007/978-3-319-24277-4
- 2021seaborn: statistical data visualization.J Open Source Softw 6(60):3021. 10.21105/joss.03021
- 2015Beyond Bar and Line Graphs: Time for a New Data Presentation Paradigm.PLOS Biol. 10.1371/journal.pbio.1002128
- 2020SuperPlots: Communicating reproducibility and variability in cell biology.J Cell Biol. 10.1083/jcb.202001064
- 2014UpSet: Visualization of Intersecting Sets.IEEE Trans Vis Comput Graph. 10.1109/TVCG.2014.2346248
- 2019BookFundamentals of Data Visualization.O'Reilly. free online
Step 13 Reproducibility & Reporting
- 2020GuidelineThe ARRIVE guidelines 2.0: Updated guidelines for reporting animal research.PLOS Biol. 10.1371/journal.pbio.3000410
- 2010GuidelineCONSORT 2010 Statement: updated guidelines for reporting parallel group randomised trials.BMJ. 10.1136/bmj.c332
- 2016The FAIR Guiding Principles for scientific data management and stewardship.Sci Data. 10.1038/sdata.2016.18
- 2012Raise standards for preclinical cancer research.Nature 483:531. 10.1038/483531a
- 20161,500 scientists lift the lid on reproducibility.Nature 533:452. 10.1038/533452a
- —PortalEnhancing the QUAlity and Transparency Of health Research.equator-network.org
Step 14 AI / ML for Quantification
- 2024NEWGuidelineTRIPOD+AI statement: updated guidance for reporting clinical prediction models using regression or machine learning methods.BMJ. 10.1136/bmj-2023-078378
- 2025NEWGuidelineThe STARD-AI reporting guideline for diagnostic accuracy studies using AI.Nat Med. 10.1038/s41591-025-03953-8
- 2020GuidelineCLAIM: Checklist for Artificial Intelligence in Medical Imaging.Radiol Artif Intell. 10.1148/ryai.2020200029
- 2017A Unified Approach to Interpreting Model Predictions (SHAP).NeurIPS. arXiv:1705.07874
- 2015The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets.PLOS ONE. 10.1371/journal.pone.0118432
- 2019Calibration: the Achilles heel of predictive analytics.BMC Med. 10.1186/s12916-019-1466-7
📌 教學註記與細節
以下整理本頁引用脈絡中容易被忽略、但會直接影響資料判讀與報告品質的補充說明。內容皆為延伸閱讀,所引文獻已收錄於上方各 Step;目的在於把零散注記集中於一處,方便日後查閱。
Below is a consolidated set of notes that often slip past readers but materially affect data interpretation and reporting quality. These are extended discussion items — the underlying references already appear in the Step sections above, and gathering them here is meant to make later look-up easier.
E1 · MIQE 2.0 vs MIQE 2009 (Bustin 2025, Clin Chem)
MIQE 2009 (Bustin et al., Clin Chem 2009) 奠定了 qPCR 報告的最低資訊清單;MIQE 2.0 (Bustin et al., Clin Chem 71(6):634, 2025) 為十六年來首次正式修訂。新增重點:(1) RT 步驟單獨報告,priming 策略、RT 酵素、RT 效率皆需揭露;(2) 內部校正基因 (reference gene) 必須以 geNorm/NormFinder 等方法驗證穩定性,「使用 GAPDH 或 β-actin」已不再被接受為正當理由;(3) dPCR (digital PCR) 另立 dMIQE 2020 (Huggett, Clin Chem 66(8):1012)。Mehle 2025 (Int J Mol Sci 26(11):4975) 進一步指出許多期刊仍允許 2009 版的舊報告格式,造成 reproducibility gap。建議:作者於投稿時即填寫 MIQE 2.0 checklist 並附 supplementary。
MIQE 2009 (Bustin et al., Clin Chem 2009) established the minimum-information checklist for qPCR; MIQE 2.0 (Bustin et al., Clin Chem 71(6):634, 2025) is the first major revision in 16 years. Key additions: (1) the RT step is reported separately — priming strategy, RT enzyme and RT efficiency must all be disclosed; (2) reference genes must be validated for stability with geNorm/NormFinder — "we used GAPDH or β-actin" is no longer an acceptable justification; (3) digital PCR has its own dMIQE 2020 (Huggett, Clin Chem 66(8):1012). Mehle 2025 (Int J Mol Sci 26(11):4975) notes that many journals still accept legacy 2009-style reports, perpetuating a reproducibility gap. Recommendation: fill the MIQE 2.0 checklist at submission and attach it as supplementary material.
E2 · Western Blot — Total Protein Normalization (Hagstrom 2025, PLOS ONE)
單一 housekeeping band (β-actin、GAPDH、α-tubulin) 作為 loading control 已被多篇文獻證實在病理樣本、藥物處理、發育階段中本身會變動,且其線性動態範圍 (LDR) 通常 < 2 logs。Total Protein Normalization (TPN) — 以 Stain-Free、Ponceau S 或 REVERT/SYPRO Ruby 染整條膜 — 提供 ~3 logs LDR,且不依賴任何個別蛋白的表達穩定性。Hagstrom et al. 2025 (PLOS ONE) 直接比較 GAPDH/β-actin/vinculin 與 TPN 於人類脂肪細胞,發現 TPN 顯著降低樣本間 CV 並改變部分 target 的顯著性判讀。Aldridge 2008、Romero-Calvo 2010、Gürtler 2013、Taylor 2013 為 TPN 的奠基方法學。實作要點:(a) 將 sample loading 控制在 LDR 內;(b) 報告 TPN 影像 (整條 lane 訊號) 為 figure supplement;(c) 抗體驗證 (Pillai-Kastoori 2020) 仍不可省。
A single housekeeping band (β-actin, GAPDH, α-tubulin) as loading control has been shown across multiple studies to vary in disease samples, drug treatments and developmental stages, and its linear dynamic range (LDR) is typically < 2 logs. Total Protein Normalization (TPN) — staining the entire membrane with Stain-Free, Ponceau S or REVERT/SYPRO Ruby — gives ~3 logs LDR and does not depend on the expression stability of any single protein. Hagstrom et al. 2025 (PLOS ONE) directly compared GAPDH/β-actin/vinculin vs TPN in human adipocytes and found TPN significantly reduced between-sample CV and changed the significance call for several targets. Aldridge 2008, Romero-Calvo 2010, Gürtler 2013 and Taylor 2013 are the foundational TPN methodology papers. Practical points: (a) keep sample loading within the LDR; (b) report the TPN image (whole-lane signal) as a figure supplement; (c) antibody validation (Pillai-Kastoori 2020) is still required.
E3 · ELISA — 5-PL vs 4-PL (Findlay & Dillard 2007, AAPS J)
4-PL (four-parameter logistic) 是 ligand-binding assays 的標準 calibration curve,但其前提是曲線「對稱」(symmetric around the inflection point)。對許多 immunoassays — 尤其 sandwich ELISA 在 high-dose hook、low-end matrix effect 等情況 — 曲線實際上是 asymmetric,此時 5-PL (added asymmetry parameter g) 可顯著改善 back-calculation 準確度。Gottschalk & Dunn 2005 (Anal Biochem) 提供 5-PL 的特徵化,Findlay & Dillard 2007 (AAPS J) 提供模型選擇準則:以 AIC + 標準曲線 weighted residuals plot 判斷是否需要 5-PL。Armbruster & Pry 2008 提供 LOB/LOD/LOQ 的標準定義。FDA Bioanalytical Method Validation Guidance 2018 為法規層面權威依據。錯誤示警:直接用 linear regression 跨越 4 logs 動態範圍是常見錯誤,殘差呈非隨機 U 形時即需改用 4-PL/5-PL。
4-PL (four-parameter logistic) is the standard calibration curve for ligand-binding assays, but it assumes the curve is symmetric around the inflection point. For many immunoassays — especially sandwich ELISA with high-dose hook or low-end matrix effects — the curve is asymmetric, and 5-PL (an added asymmetry parameter g) significantly improves back-calculation accuracy. Gottschalk & Dunn 2005 (Anal Biochem) characterize 5-PL; Findlay & Dillard 2007 (AAPS J) give model-selection criteria — use AIC plus a weighted-residuals plot of the standard curve to decide whether 5-PL is needed. Armbruster & Pry 2008 provide standard definitions of LOB/LOD/LOQ. The FDA Bioanalytical Method Validation Guidance 2018 is the regulatory authority. Common pitfall: forcing a linear regression across 4 logs of dynamic range is a frequent mistake — when residuals show a non-random U-shape, switch to 4-PL/5-PL.
E4 · Flow Cytometry — Compensation vs Spectral Unmixing
Conventional flow cytometry 使用 compensation matrix (Roederer 2001, Cytometry) 處理 fluorochrome 之間的 spectral spillover:以 single-stained controls 估計 spillover coefficients,從每個 detector 線性減去其他 channel 的貢獻。此方法假設 spectrum 為 delta 函數 — 只看一個 peak detector。Spectral cytometry (Bonilla 2025, Cells PMC12193525) 則記錄整個 emission spectrum (32-64 detectors),以最小平方法或非負矩陣分解進行 unmixing。優勢:(1) 處理高度重疊的 fluorochromes (e.g., BB515 + FITC);(2) 區分 autofluorescence 為一個獨立 "fluorochrome";(3) 大幅擴增 panel 維度 (40+ 顏色)。注意:spectral unmixing 並非「免 compensation 的魔法」— 仍需 single-stained controls + autofluorescence control,且 unmixing error 會放大為 negative values。Cossarizza 2021 (Eur J Immunol, 3rd ed) 為實作 standard。MIFlowCyt (Lee 2008) + MIFlowCyt-EV (Welsh 2020) 為 reporting guideline。
Conventional flow cytometry uses a compensation matrix (Roederer 2001, Cytometry) to handle spectral spillover between fluorochromes: spillover coefficients are estimated from single-stained controls and subtracted linearly from each detector. This assumes spectra are delta functions — only the peak detector is used. Spectral cytometry (Bonilla 2025, Cells PMC12193525) records the entire emission spectrum across 32-64 detectors and performs unmixing via least-squares or non-negative matrix factorization. Advantages: (1) handles highly overlapping fluorochromes (e.g., BB515 + FITC); (2) treats autofluorescence as a separate "fluorochrome"; (3) dramatically expands panel dimensionality (40+ colors). Caveat: spectral unmixing is not "compensation-free magic" — single-stained + autofluorescence controls are still required, and unmixing error is amplified into negative values. Cossarizza 2021 (Eur J Immunol, 3rd ed) is the implementation standard; MIFlowCyt (Lee 2008) + MIFlowCyt-EV (Welsh 2020) provide the reporting guidelines.
E5 · Tukey 1977 — EDA Fences vs MAD-based Outlier Detection
Tukey 的 boxplot fences (Q1 − 1.5·IQR, Q3 + 1.5·IQR; Tukey 1977 EDA) 是視覺化 outlier 的標準工具,但其數學依據是「在常態分布下,約 0.7% 的資料會落於 fence 外」。當資料 (a) 嚴重偏態 (e.g., qPCR Ct、cytokine concentration)、(b) 樣本數 < 30、(c) 真實存在 mixture model 時,Tukey fences 會誤判過多 outliers 或漏抓 outliers。Leys et al. 2013 (J Exp Soc Psychol) 推薦改用 MAD (median absolute deviation) 為基準:outlier 條件為 |x − median| > k · MAD (k=2.5 為 conservative, k=3 為 very conservative)。MAD 對 50% 污染仍 robust,遠優於 mean ± n·SD (breakdown point = 0)。實作:R `mad()` 預設帶 consistency constant 1.4826;Python `scipy.stats.median_abs_deviation(scale='normal')`。Knief & Forstmeier 2025 (R Soc Open Sci) 指出回歸殘差的常態性測試本身就受 outlier 干擾,建議改看 QQ-plot + Cook's distance 而非 Shapiro-Wilk on residuals。
Tukey's boxplot fences (Q1 − 1.5·IQR, Q3 + 1.5·IQR; Tukey 1977 EDA) are the standard visualization tool for outliers, but their basis is that under a normal distribution about 0.7% of data falls outside the fence. When data are (a) heavily skewed (e.g., qPCR Ct, cytokine concentrations), (b) sample size < 30, or (c) a true mixture model, Tukey fences over- or under-call outliers. Leys et al. 2013 (J Exp Soc Psychol) recommend MAD (median absolute deviation) instead: an outlier is |x − median| > k · MAD (k=2.5 conservative, k=3 very conservative). MAD remains robust under 50% contamination, far better than mean ± n·SD (breakdown point = 0). Implementation: R `mad()` includes a consistency constant of 1.4826 by default; Python uses `scipy.stats.median_abs_deviation(scale='normal')`. Knief & Forstmeier 2025 (R Soc Open Sci) further note that normality tests on regression residuals are themselves perturbed by outliers — prefer QQ plots + Cook's distance over Shapiro-Wilk on residuals.
E6 · ASA 2016/2019 — Beyond "p < 0.05"
Wasserstein & Lazar 2016 (Am Stat 70(2):129) 為 American Statistical Association 史上首份對 p-value 的官方 statement,列出六大原則:(1) p-value 衡量資料與某模型的不相容程度,不衡量該假說為真的機率;(2) p-value 不衡量效應大小或結果的重要性;(3) 不應僅以 p < 0.05 作為科學結論的依據;(4) 完整 reporting 需含 effect size、CI、prior 資訊;(5) p-value 不衡量 hypothesis 的證據力;(6) 在 NHST 之外存在多種互補方法 (Bayesian、likelihood、equivalence testing)。2019 年的後續 (Am Stat 73(sup1):1) 主張 "moving to a world beyond p < 0.05",呼籲 (i) 廢除 "statistically significant" 這個二分標籤、(ii) 整合 ATOM 框架 (Accept uncertainty, be Thoughtful/Open/Modest)。Greenland 2016 (Eur J Epidemiol) 補充 25 個常見的 p-value 誤解。實作:在論文中報告 effect size + 95% CI + exact p-value (不要寫 "p < 0.05" 或 "n.s.");用 figure 呈現 raw data + group-level estimate (Weissgerber 2015)。
Wasserstein & Lazar 2016 (Am Stat 70(2):129) is the American Statistical Association's first-ever official statement on p-values. Six principles: (1) a p-value measures incompatibility between data and a model, not the probability that the hypothesis is true; (2) p-values do not measure effect size or importance; (3) scientific conclusions should not be based on whether p < 0.05; (4) complete reporting requires effect size, CI and prior information; (5) p-values do not measure the strength of evidence; (6) Bayesian, likelihood and equivalence-testing alternatives complement NHST. The 2019 follow-up (Am Stat 73(sup1):1) calls for "moving to a world beyond p < 0.05" — (i) abandon the "statistically significant" dichotomy; (ii) adopt the ATOM framework (Accept uncertainty, be Thoughtful/Open/Modest). Greenland 2016 (Eur J Epidemiol) catalogs 25 common p-value misinterpretations. Practice: report effect size + 95% CI + exact p-value (avoid "p < 0.05" or "n.s."); show raw data + group-level estimates in figures (Weissgerber 2015).
E7 · BH-FDR vs Bonferroni — When to Use Which
Bonferroni 校正 (α/m) 控制 FWER (family-wise error rate),即「至少一個 false positive」的機率 ≤ α。其代價:當 m 很大 (genomics 中 m = 10⁴ ~ 10⁶) 時 power 趨近於 0。Benjamini-Hochberg 1995 (JRSS B 57(1):289) 改為控制 FDR (false discovery rate, E[V/R]),即「被宣告為 discovery 的結果中,預期 false positive 比例」≤ q。FDR 在大尺度檢定 (microarray、RNA-seq、ChIP-seq、GWAS) 中成為標準。決策樹:(a) m < 20 且每個假設都關鍵 (e.g., 多重 endpoint clinical trial) → Bonferroni or Holm;(b) m > 100 且可容忍少量 false positives → BH-FDR;(c) 已知 p-value 並非 independent (e.g., LD 中的 SNPs) → BY (Benjamini-Yekutieli) or permutation-based。Storey & Tibshirani 2003 (PNAS) 提出 q-value (local FDR) 取代 BH 的 step-up,可估計每筆檢定的 FDR。Liu 2024 (Bioinformatics 40(2):btae035) 提出 2dGBH,針對 grouped/structured hypotheses (e.g., gene × tissue) 提供 power-adaptive FDR 控制。注意:BH 假設 p-value 在 H₀ 下為 uniform — 對於 discrete tests (Fisher exact) 需特殊處理。
Bonferroni (α/m) controls the family-wise error rate (FWER) — the probability of "at least one false positive" ≤ α. Cost: when m is large (10⁴–10⁶ in genomics), power approaches 0. Benjamini-Hochberg 1995 (JRSS B 57(1):289) instead controls the false discovery rate (FDR, E[V/R]) — the expected proportion of false positives among declared discoveries ≤ q. FDR is the standard for large-scale testing (microarray, RNA-seq, ChIP-seq, GWAS). Decision tree: (a) m < 20 with every hypothesis critical (e.g., multi-endpoint clinical trial) → Bonferroni or Holm; (b) m > 100 with some false positives tolerable → BH-FDR; (c) p-values known to be dependent (e.g., LD-linked SNPs) → BY (Benjamini-Yekutieli) or permutation-based. Storey & Tibshirani 2003 (PNAS) introduced the q-value (local FDR) as a per-test FDR estimate replacing BH step-up. Liu 2024 (Bioinformatics 40(2):btae035) proposes 2dGBH, a power-adaptive FDR procedure for grouped/structured hypotheses (e.g., gene × tissue). Note: BH assumes p-values are uniform under H₀ — discrete tests (Fisher exact) need special handling.
E8 · Post-hoc Power Myth (Hoenig & Heisey 2001, Am Stat)
「我的 t-test p = 0.08 (非顯著),所以 post-hoc 計算 power = 0.45,可見研究 underpowered」— 這是統計學上最普遍但最錯誤的推論。Hoenig & Heisey 2001 (Am Stat 55:19) 證明:post-hoc power 是 observed effect size 的單調函數,與 p-value 是「一個硬幣的兩面」(p-value 高 ⇔ observed power 低)。重複計算 observed power 不會提供新資訊,更不能用來解釋為何沒有達到顯著。正確做法:(1) 在實驗 *之前* 以 minimum effect size of interest (SESOI) 進行 a priori power analysis (Cohen 1988, Faul 2007 G*Power);(2) 實驗 *之後* 報告 effect size + 95% CI,CI 寬度即「實驗的 informativeness」;(3) 若 CI 包含 SESOI 與 null,結論為 "inconclusive",不是 "no effect";(4) 若需 formal equivalence test,用 TOST (two one-sided tests) 而非 post-hoc power。Giner-Sorolla 2024 (Pers Soc Psychol Rev) 提供 SESOI 的選擇框架。Goodman & Berlin 1994 為此議題的奠基論文。
"My t-test p = 0.08 (non-significant), so post-hoc power = 0.45, therefore my study is underpowered" — this is the most common but most incorrect inference in statistics. Hoenig & Heisey 2001 (Am Stat 55:19) prove that post-hoc power is a monotonic function of the observed effect size and is essentially "the other side of the p-value" (high p ⇔ low observed power). Recomputing observed power gives no new information and cannot explain why significance was not reached. Correct approach: (1) before the experiment, run a-priori power analysis with a minimum effect size of interest (SESOI; Cohen 1988, Faul 2007 G*Power); (2) after the experiment, report effect size + 95% CI — the CI width is the "informativeness" of the study; (3) if the CI contains both SESOI and null, conclude "inconclusive", not "no effect"; (4) for formal equivalence testing, use TOST (two one-sided tests), not post-hoc power. Giner-Sorolla 2024 (Pers Soc Psychol Rev) provides a framework for choosing SESOI. Goodman & Berlin 1994 is the foundational paper.
E9 · Cohen's d vs Hedges' g — Small-sample Bias
Cohen's d = (M₁ − M₂) / s_pooled 是兩組效應量的標準度量。其估計在小樣本 (n < 20 per group) 下會 *系統性高估* 真實效應量,bias 約為 1 + 3/(4·df − 1)。Hedges 1981 (J Educ Stat 6(2):107) 提出 small-sample correction factor J(df) = 1 − 3/(4·df − 1),定義 Hedges' g = J(df) · d。在 df → ∞ 時 g → d;在 df = 10 時 g ≈ 0.923 · d。Cohen 1988 (Statistical Power Analysis 2nd ed) 提供經典 thresholds:d = 0.2 small, 0.5 medium, 0.8 large — 但這只是 "in the absence of other information" 的 default,每個學科應建立自己的 SESOI。對於 within-subject design (paired),應使用 dz 或 drm;對於 ANOVA,應改用 η² / η_p² / ω²。R: `effsize::cohen.d(hedges.correction=TRUE)`;Python: `pingouin.compute_effsize(eftype='hedges')`。Lakens 2013 (Front Psychol) 為實作 review。
Cohen's d = (M₁ − M₂) / s_pooled is the standard effect size for two-group comparisons. Its estimator systematically *overestimates* the true effect in small samples (n < 20 per group), with bias ≈ 1 + 3/(4·df − 1). Hedges 1981 (J Educ Stat 6(2):107) proposed a small-sample correction J(df) = 1 − 3/(4·df − 1), defining Hedges' g = J(df) · d. As df → ∞, g → d; at df = 10, g ≈ 0.923 · d. Cohen 1988 (Statistical Power Analysis 2nd ed) gives the classic thresholds d = 0.2 small / 0.5 medium / 0.8 large — but these are defaults "in the absence of other information" and each field should establish its own SESOI. For within-subject (paired) designs, use dz or drm; for ANOVA, η² / η_p² / ω². R: `effsize::cohen.d(hedges.correction=TRUE)`; Python: `pingouin.compute_effsize(eftype='hedges')`. Lakens 2013 (Front Psychol) is the implementation review.
E10 · TPM Is Not for DE (Wagner et al. 2012)
TPM (transcripts per million) 與 RPKM/FPKM 都對「樣本內」的 sequencing depth 進行了標準化,但 *並未* 對「樣本間」的 RNA composition 差異做校正。Wagner et al. 2012 (Theory Biosci) 指出 RPKM 在跨樣本比較時 inconsistent — 因為 RPKM 的分母 (library size in millions) 隨著 *highly expressed gene 的存在與否* 而改變。同樣的,TPM 雖然滿足「各 gene TPM 加總 = 10⁶」的 invariant,仍受 composition bias 影響:當少數高表現 gene (e.g., globin、rRNA contamination) 變動時,所有 gene 的 TPM 都被「擠壓」。正確做法:DE analysis 使用 raw counts + DESeq2 median-of-ratios (Love 2014, Genome Biol)、edgeR TMM (Robinson 2010, Bioinformatics) 或 limma-voom (Law 2014, Genome Biol),這些方法都採 size factor 估計來校正 composition bias。MaxLFQ (Cox 2014, MCP) 為 label-free proteomics 的對應概念。TPM/FPKM 適合:(a) 視覺化單一 gene 跨樣本的 expression;(b) 不同 gene 之間的相對表達。不適合:(c) 不同 sample 之間的差異表達檢定。apeglm (Zhu 2018) 為 log2FC shrinkage 標準。
TPM (transcripts per million) and RPKM/FPKM normalize within-sample sequencing depth but *do not* correct between-sample RNA composition differences. Wagner et al. 2012 (Theory Biosci) showed RPKM is inconsistent across samples because the denominator (library size in millions) changes with the *presence or absence of highly expressed genes*. Likewise, even though TPM satisfies the invariant "sum of all TPMs = 10⁶", it is still affected by composition bias: when a few highly-expressed genes (e.g., globin, rRNA contamination) shift, every gene's TPM is "squeezed". Correct approach: do DE analysis on raw counts with DESeq2 median-of-ratios (Love 2014, Genome Biol), edgeR TMM (Robinson 2010, Bioinformatics) or limma-voom (Law 2014, Genome Biol) — all estimate size factors to correct composition bias. MaxLFQ (Cox 2014, MCP) is the equivalent for label-free proteomics. TPM/FPKM is suitable for: (a) visualizing a single gene across samples; (b) relative expression between genes. Not suitable for: (c) DE testing between samples. apeglm (Zhu 2018) is the standard for log2FC shrinkage.
E11 · Weissgerber 2015 — Bar Plot Critique
Weissgerber et al. 2015 (PLOS Biol) 分析了三本頂級生理學期刊上百篇論文中的 bar plot,指出當 n 較小時 (常見於 wet lab,n = 3-10),同一個 mean ± SEM 的 bar 可能來自完全不同的資料分布 — uniform、bimodal、含 outliers 都會產生相似的 summary。這導致讀者無法判斷:(a) 真實的 sample size、(b) 個別資料點分布、(c) 變異來源。建議替代圖:(i) dot plot / strip plot 直接顯示每個資料點;(ii) box plot 顯示 median + IQR;(iii) violin plot 顯示 density;(iv) Lord et al. 2020 (J Cell Biol) 提出 SuperPlot — 對 cell biology 的 nested data (每個 biological replicate 內含多個 technical replicates),以兩層 dot plot 同時呈現「每細胞」與「每樣本」的 variability,並以後者進行統計檢定 (避免 pseudoreplication)。實作:ggplot2 `geom_jitter() + stat_summary()`;seaborn `stripplot()` / `swarmplot()` + `pointplot()`。Lex et al. 2014 (UpSet, TVCG) 為 multi-set Venn 的替代。color: viridis / Okabe-Ito 為 colorblind-safe palettes (Wilke 2019)。
Weissgerber et al. 2015 (PLOS Biol) analyzed bar plots in hundreds of papers across top physiology journals and showed that, at small n (common in wet-lab work, n = 3-10), the same mean ± SEM bar can come from very different distributions — uniform, bimodal or outlier-laden samples all produce similar summaries. Readers therefore cannot judge: (a) the actual sample size, (b) the distribution of individual data points, or (c) the source of variation. Recommended alternatives: (i) dot plot / strip plot to show every data point; (ii) box plot showing median + IQR; (iii) violin plot showing density; (iv) Lord et al. 2020 (J Cell Biol) introduced SuperPlot — for nested cell-biology data (each biological replicate containing many technical replicates), a two-layer dot plot simultaneously displays per-cell and per-sample variability, with statistics run on the latter (avoiding pseudoreplication). Implementation: ggplot2 `geom_jitter() + stat_summary()`; seaborn `stripplot()` / `swarmplot()` + `pointplot()`. Lex et al. 2014 (UpSet, TVCG) is the multi-set Venn alternative. Color: viridis / Okabe-Ito are colorblind-safe palettes (Wilke 2019).
E12 · MIQE / ARRIVE / CONSORT — Scope Distinction
三大 reporting guideline 各有領域,不可混用:(1) MIQE (Bustin 2009/2025, Clin Chem) — qPCR / RT-qPCR 實驗的最低資訊清單;dMIQE 2020 為 digital PCR 對應版本。(2) ARRIVE 2.0 (Percie du Sert 2020, PLOS Biol) — *動物實驗* 報告規範 (10 項 essential + recommended);強調 randomization、blinding、sample size justification 與 attrition。(3) CONSORT 2010 (Schulz 2010, BMJ) — *人類隨機對照試驗* 報告規範;含 trial flow diagram、primary/secondary outcomes、ITT vs per-protocol。其他相關:(a) STARD (診斷準確性研究)、(b) STROBE (觀察性流行病學)、(c) PRISMA (systematic review/meta-analysis)、(d) CARE (case reports)、(e) TRIPOD+AI (Collins 2024) 與 STARD-AI (Salim 2025) 為 AI/ML 模型的延伸版本。EQUATOR Network (equator-network.org) 為所有 health research reporting guidelines 的官方 portal。實作:選對 guideline + 完整 checklist 於 supplementary 提供;許多期刊已將其列為 mandatory submission requirement。
The three major reporting guidelines have distinct scopes and must not be conflated: (1) MIQE (Bustin 2009/2025, Clin Chem) — minimum information for qPCR / RT-qPCR; dMIQE 2020 is the digital PCR analog. (2) ARRIVE 2.0 (Percie du Sert 2020, PLOS Biol) — *animal-experiment* reporting (10 essential + recommended items); emphasizes randomization, blinding, sample-size justification and attrition. (3) CONSORT 2010 (Schulz 2010, BMJ) — *human randomized controlled trials*; includes trial flow diagram, primary/secondary outcomes, ITT vs per-protocol. Related: (a) STARD (diagnostic accuracy), (b) STROBE (observational epidemiology), (c) PRISMA (systematic review / meta-analysis), (d) CARE (case reports), (e) TRIPOD+AI (Collins 2024) and STARD-AI (Salim 2025) are AI/ML extensions. EQUATOR Network (equator-network.org) is the official portal for all health-research reporting guidelines. Practice: choose the right guideline, provide a complete checklist as supplementary material — many journals now make this mandatory at submission.
E13 · TRIPOD+AI 2024 Supersedes TRIPOD 2015
TRIPOD 2015 (Collins, Reitsma, Altman, Moons; BMJ 2015) 為 *clinical prediction model* (regression-based) 的報告規範,含 22 個 checklist items 涵蓋 model development 與 external validation。TRIPOD+AI (Collins et al. 2024, BMJ; 10.1136/bmj-2023-078378) 為其 AI/ML 延伸版本,從 22 擴增為 27 items,新增重點:(1) AI/ML 演算法的明確 specification (architecture、hyperparameters、software version);(2) data leakage 預防 (training/validation/test 嚴格分割,避免 patient-level leakage);(3) interpretability/explainability (SHAP、LIME、attention maps);(4) fairness assessment (subgroup performance、algorithmic bias);(5) calibration 必須報告 (不僅 discrimination/AUC; Van Calster 2019)。STARD-AI (Salim 2025, Nat Med) 為 *diagnostic accuracy studies* 用 AI 的對應版本;CLAIM (Mongan 2020, Radiol AI) 為 medical imaging 的 checklist。實作建議:論文投稿前完成 TRIPOD+AI checklist + 公開 code/model weights (若可);Lundberg 2017 (NeurIPS, SHAP) + Saito 2015 (PLOS ONE, PR vs ROC) 為 imbalanced classification 的標準參考。注意 — TRIPOD+AI 並非廢除 TRIPOD 2015,而是補充:純粹 regression-based prediction model 仍可用 TRIPOD 2015,但 ML 元素 (random forest、XGBoost、neural net) 必須改用 TRIPOD+AI。
TRIPOD 2015 (Collins, Reitsma, Altman, Moons; BMJ 2015) is the reporting guideline for *clinical prediction models* (regression-based), with 22 checklist items covering model development and external validation. TRIPOD+AI (Collins et al. 2024, BMJ; 10.1136/bmj-2023-078378) is its AI/ML extension, expanded from 22 to 27 items. Key additions: (1) explicit specification of AI/ML algorithms (architecture, hyperparameters, software version); (2) data-leakage prevention (strict train/validation/test split, no patient-level leakage); (3) interpretability/explainability (SHAP, LIME, attention maps); (4) fairness assessment (subgroup performance, algorithmic bias); (5) calibration must be reported (not just discrimination/AUC; Van Calster 2019). STARD-AI (Salim 2025, Nat Med) is the AI extension for *diagnostic accuracy studies*; CLAIM (Mongan 2020, Radiol AI) is the medical-imaging checklist. Practice: complete the TRIPOD+AI checklist before submission and release code / model weights when possible; Lundberg 2017 (NeurIPS, SHAP) + Saito 2015 (PLOS ONE, PR vs ROC) are standard references for imbalanced classification. Note — TRIPOD+AI does not retire TRIPOD 2015; it supplements it. Pure regression-based prediction models can still use TRIPOD 2015, but ML elements (random forest, XGBoost, neural net) require TRIPOD+AI.