July Spotlight: How Federal Data and Artificial Intelligence Could Accelerate Childhood Cancer Research
By: Diya Sriramagiri, Founder of Project 46
Every child diagnosed with cancer generates valuable information, from tumor genetics and medical images to treatment responses and long-term outcomes. However, this information is often stored across separate hospitals, research institutions, and databases. When researchers cannot easily combine these records, discovering patterns in rare childhood cancers becomes much more difficult. The National Cancer Institute’s Childhood Cancer Data Initiative (CCDI) is working to change this. Launched in 2020, CCDI is a federally supported initiative designed to collect, organize, and share childhood cancer data. Its goal is to help researchers better understand cancer biology, improve diagnoses, develop more precise treatments, and learn from the experiences of children, adolescents, and young adults with cancer, regardless of where they receive care. Childhood cancers are much rarer than adult cancers. Although this means fewer children are diagnosed, it also means researchers have fewer cases available to study for each cancer subtype. Data from a single hospital may not be enough to identify meaningful genetic patterns or determine why certain patients respond differently to treatment.CCDI addresses this problem through its Data Ecosystem, a connected network of platforms that allows researchers to locate and analyze clinical, genomic, imaging, and biospecimen data. Resources within the ecosystem include the Childhood Cancer Clinical Data Commons, National Childhood Cancer Registry, Childhood Cancer Data Catalog, Molecular Targets Platform, and CCDI cBioPortal Cancer Data Explorer.Instead of requiring researchers to search through disconnected databases, CCDI’s Data Federation Resource allows information stored across multiple sources to be queried as though it were part of one unified system. Standardizing and connecting these datasets can make research more efficient and allow scientists to study larger, more representative groups of patients.
Artificial intelligence is most useful when it has access to large quantities of high-quality, well-organized data. CCDI is helping create this foundation by generating, standardizing, and sharing pediatric cancer datasets that can be used to train and evaluate machine-learning models. AI does not independently “cure” childhood cancer or automatically produce a new therapy. Instead, it can help researchers process complex information more quickly, generate hypotheses, and identify promising targets that can then be investigated through laboratory studies and clinical trials. One of CCDI’s major programs is the Molecular Characterization Initiative (MCI), which provides eligible children, adolescents, and young adults with advanced molecular testing at no cost. The testing examines DNA and RNA from tumor and blood samples to help researchers and care teams better understand the changes driving an individual patient’s cancer. As of July 2026, the initiative had enrolled more than 9,000 participants. Eligibility currently includes certain newly diagnosed patients age 25 or younger with central nervous system tumors, soft-tissue sarcomas, high-risk neuroblastoma, rare tumors, or metastatic Ewing sarcoma who receive care through qualifying institutions. Data produced through this program can support individual diagnoses while also contributing to a broader research resource. When combined with treatment and outcome information, these molecular profiles may help researchers identify groups of patients who share the same cancer-driving alteration—even when their cancers were originally given different names.CCDI also maintains a Molecular Targets Platform, which connects information about childhood cancers with molecules involved in cancer growth, associated diseases, and drugs. These connections can help researchers prioritize possible therapeutic targets and investigate whether an existing treatment might benefit a different group of pediatric patients.
CCDI has been supported through a federal investment of $50 million per year, proposed for ten years beginning in 2020. Continued support matters because creating a functional national data infrastructure requires more than simply collecting information. Data must be securely stored, standardized, updated, connected across institutions, and made usable for researchers while protecting patient privacy. Federal support can also help ensure that discoveries are not limited to children treated at a small number of major research centers. One of CCDI’s foundational goals is to gather data from children, adolescents, and young adults regardless of where they receive care. Reaching that goal could help researchers understand whether geography, race, ethnicity, insurance status, or access to specialized hospitals influences diagnosis, treatment, and survival. However, data inclusion alone will not eliminate disparities. Families still need access to molecular testing, clinical trials, transportation, specialists, and the treatments that research ultimately produces. Data-driven discoveries must therefore be paired with policies that make new advances accessible to children in underserved and rural communities.
The Childhood Cancer Data Initiative demonstrates how federal investment, nationwide collaboration, and responsible uses of artificial intelligence can strengthen pediatric cancer research. By connecting information that was previously scattered across institutions, CCDI can help researchers investigate rare cancers, recognize hidden patterns, identify potential treatment targets, and design more precise studies. Every child’s experience with cancer contains information that could help the next patient. CCDI is building the infrastructure needed to learn from those experiences, but its long-term success will depend on sustained funding, responsible data sharing, patient privacy, and a commitment to ensuring that future discoveries reach every child who needs them.