Introduction
Shared instrumentation and core facilities are foundational components of the modern biomedical research enterprise as they centralize advanced technologies, specialized expertise, and high-value instrumentation. These facilities enable investigators across disciplines to access capabilities that would otherwise be beyond the reach of individual laboratories. As hubs within the institutional research ecosystem, core facilities foster collaboration and team science by connecting researchers, technologies, and data across laboratories and departments. They also play a critical educational role by introducing investigators and trainees to emerging technologies and providing training in experimental design, data generation, and analysis. By adopting standardized protocols and best practices, core facilities contribute significantly to research rigor, reproducibility, and transparency.1–7
The increasing scale and complexity of data generated by modern biomedical technologies further elevate the importance of shared core resource facilities in shaping institutional data management practices.6 Because cores generate a substantial portion of institutional research data across platforms, such as genomics, proteomics, and imaging, they are uniquely positioned to embed best practices for data stewardship directly into experimental workflows. Core facilities can ensure clear data provenance and readiness for downstream sharing and reuse by implementing standardized operating procedures, structured metadata capture, and consistent documentation of analytical parameters.
Importantly, shared resource facilities represent a strategic leverage point for advancing FAIR data practices across institutions. Operating as centralized environments that serve diverse research communities, cores can implement consistent workflows for metadata capture, data quality control, and documentation across multiple projects. Through their consultant role in experimental design and user training, core staff can introduce responsible data management practices at the earliest stages of research. By embedding FAIR-aligned data practices8 into routine service operations, core facilities can help operationalize emerging policy expectations — such as the National Institutes of Health (NIH) Data Management and Sharing Policy9 — while fostering a broader cultural shift toward more transparent, interoperable, and reusable research data.10–14 Three important steps in achieving this goal include: (1) connecting clinical information and patient data to biosample-derived data generated at core facilities, (2) encouraging the use of standardized formats for datatypes and minimum metadata, and (3) institutional support and infrastructure that enables these efforts.
Connecting Shared Core Resource-Generated Data with Clinical Information
A critical first step in operationalizing FAIR data practices in shared resource facilities is ensuring that experimental data generated from patient-derived biospecimens are appropriately linked to relevant clinical information.15 Many technologies housed in core facilities — including genomics, proteomics, imaging, and other high-throughput platforms — analyze samples derived from tissues, blood, or other biospecimens obtained from patients during clinical care or clinical studies. While the resulting molecular or imaging datasets are often deposited in repositories or shared through publications, the absence of accompanying clinical context substantially limits their interpretability and downstream reuse. To enable meaningful secondary analysis and integration across studies, experimental data must be connected to patient-derived clinical metadata describing the individuals from whom the biospecimens were obtained.16,17
Recent workshops convened by the National Cancer Institute (NCI) Office of Data Sharing identified a core set of clinical features that are broadly required to support translational and clinical cancer research using human-derived data.18 These features span several major categories, including demographics, disease characteristics, molecular and diagnostic profiles, treatment exposures, and clinical outcomes (Table 1). Establishing systematic mechanisms for connecting experimental outputs generated in core facilities to structured clinical metadata is a key operational challenge. One practical approach may be to establish metadata pipelines that connect biospecimen records and clinical data sources to experimental datasets, using laboratory information management systems (LIMS).
Standardizing Data Formats and Metadata to Enable Interoperability
In addition to connecting experimental data with clinical metadata, a critical requirement for making core-generated datasets FAIR is to adopt standardized data formats. A recent survey of shared instrumentation cores at cancer centers by the NCI Office of Data Sharing demonstrated that data produced by shared resource facilities, including proteomics, metabolomics, genomics, and imaging data, are typically generated using vendor-specific instruments and proprietary software (Figure 1). This data are then mainly shared in heterogeneous formats, making it difficult to integrate datasets across platforms, institutions, and studies. As a result, even when datasets are publicly shared, the lack of standardized formats can significantly limit their discoverability, interoperability, and reuse.
Equally important is the structured capture of metadata during the experimental workflow. Core facilities routinely collect instrument settings, assay conditions, and analytical parameters as part of standard operating procedures. Extending these workflows to include structured fields for sample preparation details, and study-level metadata ensures that datasets can be accurately contextualized and linked to laboratory data maintained elsewhere in institutional data systems such as electronic laboratory notebooks (ELNs). These practices also support provenance tracking, which is critical for reproducibility and downstream data integration.19
Adopting community-driven standards is essential for addressing these challenges. In several technology domains commonly supported by shared resource facilities, well-established standards have already been developed. For example, the Human Proteome Organization Proteomics Standards Initiative (HUPO-PSI) has introduced formats such as mzML,20 mzIdentML,21 and mzTab22 to standardize mass spectrometry data and other associated metadata.23,24 Similarly, the metabolomics community has promoted standardized reporting and data exchange frameworks through initiatives such as the Metabolomics Standards Initiative (MSI)25,26 and repositories including MetaboLights. In imaging, the Open Microscopy Environment (OME) has developed interoperable file formats such as OME-TIFF and OME-NGFF,27 along with metadata guidelines such as REMBI (Recommended Metadata for Biological Images)28to facilitate consistent data annotation and sharing.
Despite the availability of these frameworks, adopting them across laboratories and instrument vendors remains uneven.
Shared resource facilities can play a pivotal role in advancing data interoperability by embedding standards-compliant workflows into routine service pipelines. For example, core facilities can implement automated export pipelines that convert proprietary instrument outputs into community-supported formats, incorporate standardized metadata templates into sample submission workflows, and provide guidance to investigators on selecting appropriate repositories and data standards for their datasets. By integrating these practices directly into the data generation process, core facilities can ensure that datasets produced across institutional research programs are structured in ways that enable seamless discovery, integration, and reuse.
Building an Institutional Infrastructure to Support FAIR Data Generation in Shared Resource Facilities
Operationalizing FAIR data practices within shared resource facilities requires more than just adopting standards and capturing appropriate metadata; it also depends on the availability of institutional infrastructure that supports consistent data management across the research lifecycle. Because core facilities generate large volumes of highly structured experimental data across multiple technologies, they are uniquely positioned to serve as integration points between laboratory workflows and institutional data ecosystems. Establishing the appropriate technical and organizational infrastructure (such as LIMS systems) that are integrated with ELNs and biobanks can enable these facilities to efficiently capture, manage, and disseminate research data in ways that align with FAIR principles and funder expectations.17,29,30
Furthermore, partnerships with institutional libraries, research informatics groups, and information technology departments31 can support the development of data repositories, metadata catalogs, and data discovery tools that improve the findability and accessibility of core-generated datasets. Institutional data commons32 initiatives and federated data platforms33 also provide mechanisms for harmonizing datasets across multiple research programs, which enable investigators to identify and reuse data generated across different laboratories and technologies.
Finally, core facility staff routinely provide technical training to investigators and trainees on experimental methods and instrument operation; these interactions also represent valuable opportunities to introduce best practices for data management, metadata documentation, and data sharing. By integrating data stewardship into routine research workflows and training activities, shared resource facilities can help cultivate a research culture in which FAIR data practices are viewed as a standard component of high-quality scientific work.
Together, these institutional infrastructure elements — data management systems, coordinated governance structures, and workforce training — create the foundation needed for shared resource facilities to generate datasets that are well-documented, interoperable, and ready for responsible sharing. Investments in these capabilities will be essential for enabling core facilities to fulfill their role as operational engines for FAIR data generation.
Conclusion
In conclusion, shared resource facilities — such as genomics, proteomics, imaging, and other biophysical cores — play a vital role in operationalizing federal data sharing policy. These facilities generate large-scale datasets whose downstream value depends on consistent linkage to a well-defined clinical context and other metadata. Furthermore, standardization of data across core facilities is critical for the data to be used together. Aligning shared resource practices with NIH data sharing expectations can improve compliance, enhance interoperability, and increase the long-term impact of publicly funded cancer research.
Financial Support/Conflict of Interest
This manuscript is the result of funding in whole or in part by the National Institutes of Health (NIH). It is subject to the NIH Public Access Policy. Through acceptance of this federal funding, NIH has been given a right to make this manuscript publicly available in PubMed Central upon the official date of publication as defined by NIH. The contributions of the NIH author(s) are considered works of the United States government. The findings and conclusions presented in this paper are those of the author(s) and do not necessarily reflect the views of the NIH or the U.S. Department of Health and Human Services. The authors have no conflicts of interest to disclose.
