Metadata, context, and provenance are important to collect for data that are ready to be published as well as pre-publication or controlled access data as this contextual information is essential for future reuse including future use by the originating research team. The Genesis Mission has created data cards, model cards, agent cards, and tool cards (collectively referred to as xCards) to capture and format the context needed to make digital artifacts machine actionable and reusable. xCards can be created iteratively and in phases as part of the data lifecycle.
The Genesis Mission xCards are intended to be documentation for humans and machines about scientific assets that support discovery, access, interoperability, reusability, governed use, and AI use. Creating an xCard enables the described resource to be discoverable, helping agentic systems find the right resources for research and provides for workflow creation.
Collect the Required Metadata and Provenance
Collect the required metadata and provenance by consulting the requirements of the targeted repositories and Genesis Mission xCards. Align with any required ontologies or controlled vocabularies and formats.
Genesis Mission Data Card template: The published technical report provides a description of the data card including supporting ontologies and controlled vocabularies, links to versioned data card templates, corresponding schema, and supporting artifacts including a validation plug-in.
McSpadden, D., Kuchar, O., Neher, C., Walker, V., Bez, J., & Biven, L. (2026). Data Cards for Standardized Metadata Across DOE-Aligned Data Initiatives: Toward Transparent, Interoperable, and Governed Dataset Documentation. DOI: https://doi.org/10.2172/3377514
McSpadden, D., Kuchar, O. A., Walker, V., Neher, C., Bez, J. L., & Biven, L. (2026, July 1). Genesis data card schema, template and supporting tools [Dataset]. Energy Data eXchange (EDX), Jefferson Lab. https://doi.org/10.18141/3376151.
Tool card template documentation for AI-readiness tools: https://github.com/AI-ModCon/BaseData_Toolcards
Genesis Mission Model and Agent Card Templates
Automate Metadata, Context, and Provenance Capture
Some of the information for the xCards can be captured from related resources to enable xCard creation at scale. Capture and collect metadata, context and provenance automatically, when possible. Minimum requirements are described in the data card technical report. Some provenance and metadata may be captured as a distinct digital object and related to the primary data through unique identifiers and formal relationships.
Build data cards as part of the data lifecycle: Initiate a data card for datasets at the time of creation or at the time a unique identifier is assigned. Treat the data card as a living file that can be completed, augmented, and enhanced in parallel or as part of the data lifecycle. For example, as provenance artifacts are created and collected, they can be incorporated or linked through the data card. To help automate this process in agentic pipelines, a SKILL is being provided to introspect the available data and available documentation, automatically filling out the data card template. The latest version of the SKILL can be found at: https://github.com/AI-ModCon/BaseData_Skills/tree/main/skills/datacard-generator
Contribute to a FAIR, Machine-Actionable Web of Science
Elements in the FAIR and machine-actionable web of science are uniquely identified, richly and unambiguously defined, and have well defined relationships with other digital objects in the web. This is achieved through the appropriate use of digital object identifiers PIDs, DOIs, ARKs etc, collection of metadata and provenance and governance information, and linkages among these through citations and cross-referencing.
To maximize the value and long-term utility of data, it is essential to apply the FAIR Findable, Accessible, Interoperable, and Reusable principles throughout the entire project lifecycle. This practice should not be limited to final data products intended for publication but should also encompass raw or working data and internal, non-public data resources. Adhering to FAIR principles at all stages improves data quality, facilitates reproducibility, and ensures that data is well-managed and prepared for future use or sharing.
Assign Unique Identifiers Throughout the Data Lifecycle
Identifiers allow for unambiguous citation and referencing to data, code, models, and other research artifacts. Assigning unique identifiers early in the data life cycle, for example at data creation, even before data are published. Identifiers can be used to link digital objects, for example, an AI model and the data used to train it.
ARK identifier service: On behalf of the Genesis Mission, the Oak Ridge Leadership Computing Facility OLCF is developing OLCF Atlas, a researcher-driven identifier service to track research artifacts during the pre-publication, non-public phase of the scientific lifecycle. Atlas is not intended to replace DOIs or other PIDs; instead, it fills an earlier lifecycle gap: helping researchers identify, describe, relate, and track scientific products before they are ready for formal publication. Atlas is not yet available to RFA teams. More information will be shared soon.
OSTI Data ID Service: OSTI's data ID service assigns and registers Digital Object Identifiers DOIs for publicly released scientific data products. This service can be used when data are ready to be shared or published as open or controlled access. Mapping between DOIs and, for example, ARK IDs used in the pre-publication phase should be maintained for seamless referencing. https://www.osti.gov/pids/doi-services/doe-data-id-service
Repository ID: Many repositories assign a persistent identifier to data when the data are submitted or as part of the curation and publication process. In this case, an OSTI DOI is not needed. The repository DOI and citation can be used when announcing the data as scientific and technical information to OSTI.
Leverage Best Practices and Tools Related to FAIR Data
Genesis Mission catalog of AI ready data assessment tools is a centralized repository containing both a standardized template and completed "tool cards" for AI-readiness tools. These structured documents cover tools across the entire data-to-model lifecycle; from raw data ingestion and preprocessing to validation, profiling, governance, and production monitoring. Each tool card provides a comprehensive overview of a tool's capabilities, detailing its core function, supported data formats, deployment methods, and the dataset readiness levels it helps achieve. Designed to be readable by both humans and AI agents, these cards make it easy to seamlessly filter, search, and match the right tools to specific tasks.
Tools and Guidelines to support the raw to AI ready data pipeline:
DSAgt is an agentic toolkit for building AI-ready data preparation pipelines. DSAgt connects an MCP-compatible AI coding agent to code registration, a semantic knowledge base, skills discovery and creation, execution provenance, and observability infrastructure. It wraps these capabilities around a user's existing agentic CLI or VS Code extension Claude Code, Goose, Codex, etc. - https://github.com/AI-ModCon/dsagt
A technical overview of how to go from raw to curated data in three stages using a lakehouse approach: https://escholarship.org/uc/item/7049c4mp
Submit a support ticket for any of the following:
Connect your RFA team with services like data management planning, supercharging your scientific workflows with AI best practices, and cross-cutting AI capabilities.
Have an issue, suggestion, or addition, to this repository.