Vollständiger Abstract
Worum geht es in dieser Arbeit?
ABSTRACT Public proteomics repositories have expanded rapidly, but large‐scale reuse still depends on whether deposited studies are computationally reusable. We provide an evidence‐based snapshot of this gap by evaluating 500 non‐redundant human ProteomeXchange datasets released in 2025. Experimental design annotation was classified as metadata‐based when a deposited table (e.g., SDRF, spreadsheet, or text table) explicitly linked samples, files, channels, or quantitative columns to biological groups; sample‐based when groups could only be inferred from sample, column or channel names; and no information when reliable group assignment was not possible. Sample‐based datasets represented the largest category, comprising 195 deposits (39.0%), followed by datasets lacking reliable design information (170; 34.0%) and metadata‐based datasets (135; 27.0%). Among metadata‐based deposits, 81 (60.0%) included additional contextual covariates. Processed quantitative outputs were fully available in 245 datasets (49.0%), partially available in 168 (33.6%), and absent in 87 (17.4%). Combining processed quantitative outputs with available biological design information, 189 datasets (37.8%) supported direct matrix‐level reuse. Under a stricter scenario requiring metadata‐based design, contextual covariates and available processed quantitative outputs, only 64 datasets (12.8%) fulfilled all criteria. These findings suggest that scalable reuse requires more consistent machine‐actionable sample‐to‐file mapping, interpretable design annotation, processing information and processed outputs, while preserving raw data for harmonized reanalysis.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Giuseppe G. F. Leite, Victor Corasolla Carregari, Alexandre Keiji Tashima, Reinaldo Salomão
- Quelle
- PROTEOMICS
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1615-9853, 1615-9861
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Giuseppe G. F. Leite, Victor Corasolla Carregari, Alexandre Keiji Tashima, Reinaldo Salomão (2026). How Reusable Are Public Proteomics Datasets in Practice?. PROTEOMICS. https://doi.org/10.1002/pmic.70175
Kontext