AI data under the microscope: Accelerating and securing the AI data supply chain for the health and biopharmaceutical sectors
Table of Contents
- Executive summary
- Introduction
- Health and biopharma data in the AI supply chain
- Use cases for health and biopharma AI-related data
- The fractured global regulatory landscape
- Current tensions and future issues
- Conclusions and recommendations
- Author and acknowledgements
Executive summary
Health and biopharmaceutical companies occupy an important frontier in artificial intelligence (AI) research and development and the data that goes into it. Organizations in these sectors can, if behaving irresponsibly or harmfully, collect, analyze, use, and share data in ways that undermine patient privacy, create cybersecurity risks to individuals and society, threaten US national security, and inflict harm on people receiving care, particularly individuals in already marginalized or disadvantaged populations. At the same time, health and biopharmaceutical companies have tremendous opportunities to leverage data in line with established patient consent frameworks, apply data to developing novel AI technologies that advance research and development, and drive innovations that improve patient care, treat disease, and increase health equity—sometimes, without the tremendous costs that come with general-purpose large language model development. Given the importance of data and data-related components across the health and biopharmaceutical AI value chain, this report presents a framework for the data in the AI supply chain, maps it against the sectors in question, and offers five real-world use cases that illustrate varied opportunities, risks, and relationship dynamics. It then surveys fourteen major regulatory regimes around the world (AI-specific and non-AI-specific) that impact the data components of the AI supply chain for the health and biopharmaceutical sectors, describes current tensions and future issues, and ends with three high-level recommendations for how governments can navigate this frontier safely, securely, and responsibly for societal benefit.
Introduction
Health and biopharmaceutical (biopharma) industries continue to generate and collect significant amounts of data around the world and, increasingly, to build and deploy artificial intelligence (AI) technologies. Indeed, startups and research organizations in the health and biopharma sectors are using emerging AI models to predict the highly complex structures of proteins, improve early-stage diagnosis of diseases, bolster genetic sequencing and analysis capabilities, and potentially accelerate scientific discovery. These activities can further important functions—with much clearer societal benefit than many other use cases, such as general-purpose large language model chatbots in workplaces—including health condition diagnosis, disease research, and vaccine development. At the same time, the use of data as well as the development and deployment of AI raise critical questions about transparency, auditability, privacy, cybersecurity, inequity, national security, and more. Further, the globalized nature of the AI supply chain writ large as well as the globalized nature of the health and biopharma industries raise complicated geopolitical questions in an increasingly tense world, around cross-border access to citizen data (including personal health data), ideas of sovereignty, and national security.
As the companies and organizations in the health and biopharma sectors leverage AI technologies, they interact with many data components across the AI supply chain: training data, testing data, AI models, model architectures, model weights, application programming interfaces (APIs), and software development kits (SDKs). Health and biopharma companies do so subject to a varied, shifting patchwork of global laws and regulations, from the United States to the European Union (EU) to China to India, some specific to AI and others more broadly.
Altogether, the breadth and complexity of these components and the sometimes compatible, sometimes clashing nature of regulatory regimes raise many urgent policy questions for the future of health and biopharma AI-related activities. Poorly formed and implemented policies could expose patient data with unforeseen consequences, lead to the compromise of AI models, yield inadequately tested diagnostic or predictive AI systems, disproportionately and unnecessarily stall important health research, and much more. Thoughtful and carefully implemented policies, on the other hand, could accelerate the responsible, adequately protected use of health and biopharma data components in ways that genuinely advance patient outcomes—while often using far less energy, far less compute, and far more consent-permissioned data than many generic large language models currently depend upon. These patient outcomes could include earlier and more accurate prediction of disease, improved disease treatments, and more equitable treatment access and quality within and between populations.
This report unpacks the health and biopharma policy challenges across the data components of the AI supply chain, with the aim of identifying major tensions, flagging open questions, and offering potential next steps for future-proofed policymaking. The author begins by introducing a framework from a prior Atlantic Council publication to review the core data components in the AI supply chain, including how they overlap and interact. Then, this framework is used to lay out the spectrum of health and biopharma examples of those data components in the AI supply chain, from a number of health and biopharma organizations around the world working on different or overlapping problems. After those examples of health and biopharma AI-related data components, mapped against the supply chain framework, the report details five exemplative case studies of health and biopharma use cases. These range from cancer prediction to the enhancement of genetic sequencing using large-scale, global datasets to cross-border activities from a company that raises US national security questions.
Next, the report describes the fractured global regulatory landscape for the health and biopharma data components of the AI supply chain through the lens of fourteen regulatory regimes, including in the United States, the European Union, China, and India. The subsequent section details the current tensions in those regulatory regimes and lists a number of pressing, open questions for policymakers around the world vis-à-vis data privacy, data security, AI innovation, and the like as applied to the health and biopharma data spheres. Finally, the report concludes with three overarching, high-level recommendations for policymakers to identify current tensions, close current legal and regulatory gaps, and chart better paths forward.
- Governments should use existing best-practice privacy and cybersecurity principles when pursuing regulations on data collection, data use, cross-border data transfers, and AI model deployment that impact health and biopharma, including the healthcare profession’s patient consent ethical framework and standards bodies’ specifications for encryption, data minimization, and protection against reidentification.
- The United States, European Union, India, and China should issue clarifications about how the health and biopharma data components of the AI supply chain sit within their non-AI-specific and AI-specific laws, drawing on the tensions and open questions identified in this report.
- Governments should clarify criteria for any situation where they will treat health or biopharma AI-related data activities differently than AI-related data activities in other sectors.
Enabling the future of responsible innovation in the health and biopharma sectors, with the goal of delivering better and more equitable health outcomes for all, depends in part on policymakers and other stakeholders navigating these complexities. The time to do so is now.
Health and biopharma data in the AI supply chain
A complex supply chain that includes organizations, people, activities, information, and resources underpins AI technologies that enable AI research, development, deployment, and more.1This section is pulled and adapted from the earlier Atlantic Council report establishing this framework: Justin Sherman, Securing the Data in the AI Supply Chain, Atlantic Council, September 2025, https://www.atlanticcouncil.org/in-depth-research-reports/issue-brief/securing-data-in-the-ai-supply-chain/. The AI supply chain includes human talent, compute, institutional and individual stakeholders, and, data, the last of which is the focus of this report.
Data components fit into the AI supply chain because researchers, developers, deployers, users, maintainers, governors, securers, and attackers of AI systems depend upon and access different kinds of data that are transmitted, stored, and analyzed in different ways in order to make it all happen. They also fit into the AI supply chain because a wide range of entities around the world—from individuals who publish self-labeled datasets to corporations that analyze AI model outputs—provide, access, and use the underlying data as well. Conceptualizing this data as part of a supply chain echoes the concept of AI as a value chain (referring to the business activities that deliver value to customers), though focused specifically on data components.2See: Beatriz Botero Arcila, AI Liability Along the Value Chain (San Francisco: Mozilla Foundation, April 2025), https://blog.mozilla.org/netpolicy/files/2025/03/AI-Liability-Along-the-Value-Chain_Beatriz-Arcila.pdf; Max von Thun and Daniel A. Hanley, Stopping Big Tech from Becoming Big AI (San Francisco: Mozilla Foundation, October 2024), https://blog.mozilla.org/wp-content/blogs.dir/278/files/2024/10/Stopping-Big-Tech-from-Becoming-Big-AI.pdf; SPEAR Invest, “Diving Deep into the AI Value Chain,” NASDAQ, December 18, 2023, https://www.nasdaq.com/articles/diving-deep-into-the-ai-value-chain. See also: “The Value Chain,” Harvard Business School: Institute for Strategy and Competitiveness, accessed June 17, 2026, https://www.isc.hbs.edu/strategy/business-strategy/Pages/the-value-chain.aspx.
As detailed in a previous Atlantic Council report, the data in the AI supply chain covers many data types, sources, and formats—all of which need safeguards to enable competition, boost public trust, and protect against the leaks, exploitation, and other risks delineated above. The data in the AfI supply chain includes the data describing an AI model’s properties and behavior, as well as the data associated with building and using a model. It also includes the AI models themselves and the different systems that facilitate the movement of data into and out of models.3While recognizing the necessity of evaluating AI in relation to the social, political, and economic systems that researchers, companies, and others operate within and use to build AI technologies—such as exploitative labor systems and the environmental system—this report focuses, for scope- and length-limitation purposes, on a typology of the digital and data elements themselves of relevance for AI R&D. For essential reading on other systems that generate data, move data into AI systems, and much more, see: Tamara Kneese, Climate Justice and Labor Rights: Part I: AI Supply Chains and Workflows (New York: AI Now Institute, August 2023), https://ainowinstitute.org/general/climate-justice-and-labor-rights-part-i-ai-supply-chains-and-workflows; Kashmir Hill, Your Face Belongs to Us: A Tale of AI, a Secretive Startup, and the End of Privacy (New York: Penguin Random House, 2023); Billy Perrigo, “Exclusive: OpenAI Used Kenyan Workers on Less Than $2 Per Hour to Make ChatGPT Less Toxic,” TIME, January 18, 2023, https://time.com/6247678/openai-chatgpt-kenya-workers/; Adrienne Williams, Milagros Miceli, and Timnit Gebru, “The Exploited Labor Behind Artificial Intelligence,” Noema Magazine, Berggruen Institute, October 13, 2022, https://www.noemamag.com/the-exploited-labor-behind-artificial-intelligence/. This report conceptualizes seven parts of the data and core data systems in the AI supply chain, which are laid out in the visualized framework and tables below:
- Training data
- Testing data
- Models (themselves)
- Model architectures
- Model weights
- Application Programming Interfaces (APIs)
- Software Development Kits (SDKs)
This framework draws, at a high level, inspiration from a November 2024 paper by Qiang Hu et al on the large language model (LLM) supply chain. The paper envisioned a framework for understanding the components and processes that go into LLMs.4Qiang Hu et al., “Large Language Model Supply Chain: Open Problems from the Security Perspective,” arXiv, November 3, 2024, https://arxiv.org/abs/2411.01604. See: LLM supply chain map on page 2. Nonetheless, this paper differs, focusing on the data components themselves rather than the activities to produce them (like dataset processing); delineating data components by their properties and functional differences (such as distinguishing between training data and testing data); and looking at the data supply chain for AI technologies broadly (instead of just LLMs). This framework also differs in that it focuses specifically on the need for security.
Notably, the first five of these seven components are data per se, or the models themselves. The last two, however—APIs and SDKs—are neither data nor models themselves; instead, they are code and software systems that enable data to pass into, extract from, and otherwise collect around AI models. For example, business users of an LLM may use an API to submit questions to the chatbot in batches; mobile consumers using an AI image recognition application may, whether they know it or not, depend on an SDK to take their snapshot of a bird and submit to a cloud-hosted AI model, which then returns back through the SDK’s code the species of the bird in question. The concept of data components of the AI supply chain does not list out every possible software system that could interact with AI data components, but it includes APIs and SDKs because of their prevalence and their security relevance in delivering AI data to and from cloud systems and mobile devices. After all, many of the major AI commercial companies in the United States, India, China, the European Union, and elsewhere offer access to APIs to use their models (including submitting queries to chatbots and uploading images to recognition models).
This supply chain is highly relevant to the health and biopharma industries. Companies in the healthcare and biopharma sectors are developing, deploying, procuring, using, testing, and maintaining AI systems. They must contend with enabling the value of the data in their AI supply chains while doing so securely. The costs of failing to innovate securely and responsibly are serious. Disruptions to data utility in this context can undermine activities that serve a societal benefit, such as disease research and vaccine development. In this context, compromises of data security can expose trade secrets (such as proprietary training datasets or testing methodologies), violate consumer or patient privacy (such as through the theft or leaking of genetic data), bring regulatory scrutiny (such as through violating data protection requirements), and detract from health research and other R&D goals (such as disrupting projects). More broadly, these compromises can undermine trust in the use of AI-related datasets and models specifically tailored for disease research and public health benefit, potentially forestalling important future health innovations.
Enabling the value of health and biopharma data in the AI supply chain, while securing it, first requires mapping such data in the AI supply chain itself. Table 1 lists each of the seven data components of the AI supply chain, defines them, and provides examples of health or biopharma companies and other entities involved in that component or its sourcing. Again, this focuses primarily on traditional AI models (e.g., excludes agents), and does not include other elements of the AI supply chain (e.g., human talent, compute). In addition, it does not consider the outputs and exhaust from AI model usage, including metadata. In doing so, this table lays the foundation for the discussion of use cases for health and biopharma AI-related data in the next section.
Many health and biopharma data components of the AI supply chain span multiple categories of components at once. Google DeepMind’s AlphaFold model, which helps to predict the structure of proteins, is entangled with multiple other component areas (including the model itself), the training data used to teach the model, and the testing data used to test the model. The U-Net architecture is a model architecture, but its use in many health- and biopharma-related AI systems—such as in recent work on medical image segmentation and multicellular biological systems5See, e.g., Nima Hassanpour and Abouzar Ghavami, “Deep Learning-based Bio-Medical Image Segmentation Using UNet Architecture and Transfer Learning,” arXiv, May 24, 2023, https://arxiv.org/abs/2305.14841; Tien Comlekoglu et al., “Surrogate Modeling of Cellular-Potts Agent-Based Models as a Segmentation Task Using the U-Net Neural Network Architecture,” PLOS Computational Biology (November 3, 2025), https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1013626. —makes it much more present throughout several aspects of the data in the AI supply chain. Health tech company Tempus advertises real-world, multimodal clinical, molecular, and imaging data for therapy development and trials,6“Real-World Multimodal Data for Therapy Development,” Tempus, accessed June 17, 2026, https://www.tempus.com/solutions/real-world-data/. at the same time as it uses AI models to build structured patient cohorts from unstructured patient records and data.7“AI in Healthcare in 2026 and Beyond,” Tempus, accessed June 17, 2026, https://www.tempus.com/content/article/ai-in-healthcare/.
Plenty of private-sector companies also maintain their own internal AI-related data components. Based on conversations with industry, while companies will ultimately release many of the products and services that stem from R&D using the AI-related data—such as a future vaccine—it varies how much of the underlying AI-related data components are publicly shared or disclosed to a third party, based on regulatory requirements, clinical trial obligations, and business decisions, among other factors. For example, the fact that the table above primarily contains examples of datasets themselves (e.g., training datasets, testing datasets) from public research consortiums, universities, and governments rather than private-sector biopharma companies may reflect the fact that many in the latter category build and maintain their own internal AI-related training and testing datasets that they cannot (e.g., for regulatory reasons) or will not (i.e., by choice) publish. This underscores, among other points, the limits of conceptually applying ideas about open-source large language models to the varied field of more focused health or biopharma AI models and related data components.
“Open” access to health and biopharma AI-related data is one thing, while the ability for any organization or individual to effectively make use of that data for the purposes of training, testing, or improving an AI model is another. It requires access to computing resources and data storage, medical knowledge, business or clinical research mechanisms through which to deploy and ethically test the model in practice, and so on. Similarly, having access on the public internet to health or biopharma AI-related data components, from model weights to APIs to training data, does not mean everyone can equally compromise the data component’s security or the privacy of the people in the underlying data either—whether with malintent or with ethical intentions to help bolster subsequent protections. As other researchers and practitioners have explored,8See, e.g., David Gray Widder, Meredith Whittaker, and Sarah Myers West, “Why ‘Open’ AI Systems Are Actually Closed, and Why This Matters,” Nature 635 (2024): 827-33, https://www.nature.com/articles/s41586-024-08141-1; Matt Davies and Jai Vipra, “Mapping Global Approaches to Public Compute,” Ada Lovelace Institute, November 4, 2024, https://www.adalovelaceinstitute.org/policy-briefing/global-public-compute/; Sara Ann Brackett, Securing Cloud Infrastructure for AI, Atlantic Council, March 31, 2026, https://www.atlanticcouncil.org/in-depth-research-reports/issue-brief/securing-cloud-infrastructure-ai/; Noah Martin and Fahad Dogar, “Divided at the Edge— Measuring Performance and the Digital Divide of Cloud Edge Data Centers,” Proceedings of the ACM on Networking 1, no. 16 (November 2023), https://dl.acm.org/doi/10.1145/3629138. there are many elements of the AI supply chain and the computing stack for AI that impact how an organization or individual could leverage any of these data components.
Other ways to conceptually parse the health and biopharma data components of the AI supply chain include:
- Use restrictions: Across this supply chain, AI-related health and biopharma data includes data that is open, free, and publicly available for download and unrestricted use. It includes data that is private, corporate, and internal-only—in other words, proprietary. It also includes data that is somewhere in the middle: licensed by a government agency to a particular set of companies for a defined set of purposes; sold by a data broker to a limited pool of customers for their own internal use; or, perhaps, shared under a research agreement between a university, a start-up, and a public interest research group. Similarly, health and biopharma organizations can make AI model weights or AI models themselves free for public use, free for noncommercial use, restricted to certain users, closed-source (for internal use only), and so forth. They can do the same with APIs and SDKs.
- Consent type: Health and biopharma organizations may functionally get or may be legally required to obtain different types of individual consent (or not) for AI-related data collection, use, and transfer. As discussed further in the regulatory section below, these constraints or lack thereof will depend on the jurisdiction, type of data, use case for the data in question, among other factors. Some AI-related datasets may require written consent from a patient or other clear, explicit consent-acquisition steps. Other AI-related datasets may require none, with legal systems permitting health and biopharma companies to generate synthetic data, not tied to any real person, using statistical and machine learning-driven methods—or to buy health and genetic data on the commercial market. Secondary uses of biospecimens raise their own questions about data consent, because patients may agree to have their biospecimen collected for a particular purpose at a particular point in time, but where the organization may keep that biospecimen for future use is presently unknown (e.g., it might be determined through later scientific developments).
- Functional use: Companies and research organizations may employ AI data components for various purposes, ranging from early-stage R&D to regulatory approval.
- Population scope: Health and biopharma data components in the AI supply chain may relate to different populations, including by geography and by demographic. Geographically, distinct AI-related data components in the health and biopharma sectors might differently encompass the world, a few regions, one specific region, one country, and so forth. Demographically, different AI-related data components for those same sectors might differently encompass various age groups, races, ethnicities, nationalities, sexes, genders, and so on. The potential for population-scope variation mostly refers to dataset components of the AI supply chain, such as one medical image training dataset covering adults in the Middle East and another covering only teenagers in Brazil, because the data itself can be geographically or demographically bounded. But it is possible for the likes of an AI model or an SDK to be geographically bounded, too, if its use or export is restricted under laws (discussed more later) or if it has been developed specific to a region or population (and thus not trained on other data or built for use elsewhere, even if it remains technically available to other regions or groups).
- Degree of reidentifiability: Components differ in their degree of reidentifiability—that is, how quickly data that is supposedly disconnected from an individual’s identity can be tied back to that identity. Decades of computer science and statistics literature have shown that many data types that are allegedly “deidentified” or “anonymized” can be quickly linked back to specific people, with variations in ease of reidentifiability based on the specifics of the data type, the dataset, other available data, the specific kinds of masking applied to the dataset, and so forth.9See, e.g., Latanya Sweeney, “Risks to Patient Privacy: A Re-Identification of Patients in Maine and Vermont Statewide Hospital Data,” Harvard Kennedy School, October 2018, https://www.hks.harvard.edu/publications/risks-patient-privacy-re-identification-patients-maine-and-vermont-statewide-hospital. Nonetheless, there are some important differences among health and biopharma data types. Genetic and genomic data sit at the high end of risk—uniquely identifying, unchanging, and implicating biological relatives who may have never consented to data collection and analysis. Aggregated datasets, or those to which differential privacy or other masking techniques have been applied, may sit lower on the risk scale. However, these dataset states are not fixed, and computer science and statistical advances, the growth of available third-party data, and other factors can increase the reidentification risk to these datasets over time.
- Provenance and auditability: Components differ in whether their origin and handling can be traced and independently verified. A training set built from documented, consented clinical sources with a recorded chain of custody is a different governance object than one aggregated from data brokers, scraped repositories, or undisclosed partners—even when the two are interchangeable for training. The same holds for models and weights, some released with datasheets, evaluations, and license terms and others arriving as opaque artifacts. Because downstream developers inherit whatever provenance gaps exist upstream, this axis determines whether an organization can actually attest to how a component was built—a prerequisite for both compliance and security assurance.
- Revocability: Components differ in whether they can be corrected, withdrawn, or deleted once in the supply chain. A database row can be expunged, while its influence on an already trained model’s weights generally cannot (at least not cleanly). The inability to “revoke” certain data components in the AI context creates tension with rights such as the EU’s General Data Protection Regulation (GDPR) erasure and with later-revoked consent, and it compounds as data propagates into testing sets, fine-tuning sets, and synthetic generators. Such a tension surfaces where a downstream obligation (e.g., delete this patient’s data) cannot be met by an upstream component that has already learned from it.
- Retention and lifecycle: At any large health or biopharma organization, data necessary for patient health outcomes, internal recordkeeping, or ongoing research could be retained indefinitely (barring legal or regulatory requirements that indicate otherwise). Conversely, internal retention policies from a data security and cybersecurity perspective could require a company to delete data it is not using and has held for some period of years to minimize the risk of it getting breached (and the associated liability and other costs of such an incident). Companies may add to publicly accessible data components as there is more data available, such as by uploading more public training or testing data to an existing dataset, while they may choose to overwrite other components, rendering the prior versions inaccessible—as AI model vendors typically do when they update their models to new versions. These issues of retention and data lifecycle management can vary by data type, organization, and use case, with significant implications for individual privacy, cybersecurity, data provenance, and the ability to responsibly innovate.
Use cases for health and biopharma AI-related data
While the above section breaks down the health and biopharma data components in the AI supply chain, this section details a sample of five use cases for AI-related data in those same sectors. It showcases some of the many reasons why AI-related health and biopharma data can create opportunities for societal benefit when handled appropriately, but can create risks to security and privacy when designed or handled poorly. The use cases also illustrate how health and biopharma sector entities can use AI-related data with varied use restrictions, consent types, and geographic and demographic population scopes. In that vein, they underscore the complications of attempting to neatly bucket health and biopharma AI-related data components—such as publicly accessible versus not publicly accessible, or private-sector versus not-for-profit—when there can be much overlap and interconnection among the various buckets of data component accessibility and organization structure.
These use cases span:
- Isomorphic Labs: A British company leveraging publicly available, US government-funded protein data to enhance predictive AI functions for drug design;
- AI Medical Service (AIM): A Japanese company aggregating image training data from more than one hundred globally dispersed research institutes to develop a cancer-diagnosing AI model, with a focus on Japan initially;
- Basecamp Research: A network of companies and research organizations that feeds into the UK-based company’s AI-powering, proprietary database that it has used, among others, for frontier bio models that could enable AI-programmable therapeutics;
- Tempus: An American AI medicine and genomic testing data company building proprietary datasets across multiple types and clinical areas, and then offering the data as a service; and
- BGI Genomics: A Chinese genomics company that operates globally, conducting AI-related data activity that might otherwise be innocuous or standard risk, but due to its headquarters and reported work with the Chinese government there are potential risks to US national security.
Use case example #1: Predictive drug design
The UK company Isomorphic Labs has built what it calls the Isomorphic Labs Drug Design Engine, a computational drug-design model that builds on AlphaFold 3, Google DeepMind’s protein structure-predicting model that was released in 2024.10“The Isomorphic Labs Drug Design Engine Unlocks a New Frontier Beyond AlphaFold,” Isomorphic Labs, February 10, 2026, https://www.isomorphiclabs.com/articles/the-isomorphic-labs-drug-design-engine-unlocks-a-new-frontier; “AlphaFold 3 Predicts the Structure and Interactions of All of Life’s Molecules,” Isomorphic Labs, May 8, 2024, https://www.isomorphiclabs.com/articles/alphafold-3-predicts-the-structure-and-interactions-of-all-of-lifes-molecules. Proteins have complex 3D structures, and the original AlphaFold showcased in 2020 the ability to predict these structures in minutes, rather than several years.11“AlphaFold,” Google DeepMind, accessed August 11, 2026, https://deepmind.google/science/alphafold/. Doing so with speed and accuracy supports advancement in fields such as drug discovery and disease understanding.12Marios G. Krokidis et al., “AlphaFold 3: An Overview of Applications and Performance Insights,” International Journal of Molecular Sciences 26, no. 8 (April 2025): 3671, https://pmc.ncbi.nlm.nih.gov/articles/PMC12027460/.The new model more than doubles the accuracy of AlphaFold 3 on a “challenging protein-ligand generalization benchmark.”13“Accurate Predictions of Novel Biomolecular Interactions with IsoDDE,” Isomorphic Labs, February 10, 2026, https://storage.googleapis.com/isomorphiclabs-website-public-artifacts/isodde_technical_report.pdf, 1. An application programming interface (API) makes AlphaFold protein structure predictions available for free to the public.14“AlphaFold Protein Structure Database,” EBI, accessed August 13, 2026, https://alphafold.ebi.ac.uk.
AlphaFold was trained on structures from the Protein Data Bank.15Ibid., 16. The Protein Data Bank is funded by the US National Science Foundation, Department of Energy, National Cancer Institute, National Institute of Allergy and Infectious Diseases, and National Institute of General Medical Sciences and consists of 3D structure data for large biological molecules (protein, DNA, and RNA) essential for biology, health, energy, and biotech R&D.16“About RSCB PDB: A Living Digital Data Resource That Enables Scientific Breakthroughs Across The Biological Sciences,” RSCB, accessed June 17, 2026, https://www.rcsb.org/pages/about-us/index. Visitors to the Protein Data Bank website can use its built-in tools, such as an advanced search query builder, to filter the data by determination methodology, scientific name of the source organism, taxonomy, input data, polymer entity type, and many more variables.17“Advanced Search Query Builder,” RCSB, accessed June 18, 2026, https://www.rcsb.org/search/advanced.
Isomorphic Labs’ goals for the model include predicting protein-binding affinity to rank and optimize potential molecules across diverse chemical series during drug design programs.18“The Isomorphic Labs Drug Design Engine.” “Our dedicated drug design teams,” the company says, “are using these capabilities every day across our programs—to understand unseen structures, identify uncharacterized pockets, and create novel chemical matter in the pursuit of new medicines for patients.”19“The Isomorphic Labs Drug Design Engine.” This use case captures how private-sector health or biopharma companies can leverage publicly accessible protein data and already released AI models to create newly robust AI models tailored for predictive functions and future drug designs, fusing tailored data with public and private capabilities.
Use case example #2: Cancer diagnostic support
Japanese company AI Medical Service, or AIM, built an AI model called “gastroAI-model G” to aid in gastric cancer diagnostics.20“AI Medical Service Inc. Announces Regulatory Approval of Gastric AI-based Endoscopic Diagnostic Imaging Support System,” AI Medical Service, December 26, 2023, https://en.ai-ms.com/news/product/20231226. The model processes images gathered from an endoscopy system’s video processor, scans the candidate lesions to detect whether a gastric lesion is a candidate for biopsy, and, as needed, alerts the physician of its findings and provides diagnostic assistance by overlaying a rectangle on the image.21“AI Medical Service Inc. Announces.” Since April 2021, AIM has collaborated with the National University Hospital of Singapore on this effort.22“AIM Applies for Approval to Manufacture and Market the World’s First Gastric Cancer AI,” AI Medical Service, August 31, 2021, https://en.ai-ms.com/news/product/20210831. AIM trained the model on lesion image data capturing early gastric cancer from over more than one hundred collaborative research institutes, which it describes as “world-class medical institutions.”23“AIM Applies for Approval.”
This use case illustrates how health and biopharma organizations can today leverage image-based health data for diagnostic purposes. It also represents an example of a private-sector company requiring approval from a government before debuting an AI model that makes diagnostic assessments. In this case, AIM required approval from the Japanese government because Japan regulates software functioning for a diagnostic or therapeutic purpose differently than software that does not make diagnoses or inform therapy treatments.24Yurika Inoue, “Japan’s Evolving AI and Digital Health Regulations: Legal Developments and Outlook,” International Bar Association, December 3, 2025, https://www.ibanet.org/japan-ai-digital-health-regulations. This case also shows how cross-border collaborations and cross-border sharing of AI training data can enable a company to offer a relevant health or biopharma AI model in one country, and only one country to start, but with the subsequent goal of global expansion.
Use case example #3: Cross-border, AI-powering, proprietary database
The UK-based AI company Basecamp Research runs an initiative called the Trillion Gene Atlas. Launched in March 2026 with Anthropic and US biotech companies Ultima Genomics and PacBio, the database currently includes more than ten billion novel genes from over one million new species compared to current public data.25“Our Data,” Basecamp Research, accessed June 18, 2026, https://basecamp-research.com/bcr-data/. Its partners directly help with data processing for the effort. For example, Ultima Genomics uses its next-generation sequencing systems to deliver “high-throughput, whole-genome, and multi-omics sequencing” at scale and lower cost for the Trillion Atlas initiative.26“Our Data.” It also draws on a network of scientific collaborators across thirteen countries to establish “a scalable evolutionary genomics pipeline purpose-built for AI training.”27Basecamp Research, “Basecamp Research Launches Trillion Gene Atlas to Scale AI-Designed Therapeutics,” March 18, 2026, https://basecamp-research.com/wp-content/uploads/2026/03/BCR-TGA.pdf, 2.
A January 2026 Basecamp Research paper described foundation models developed with this data, which the paper says can enable AI-programmable therapeutics through predictive and generative genomic and protein benchmarks.28Geraldene Munsamy et al., “Designing AI-Programmable Therapeutics with the EDEN Family of Foundation Models,” Basecamp Research, January 2026, https://basecamp-research.com/wp-content/uploads/2026/01/BCR_Designing-programmable-therapeutics-with-the-EDEN-family-of-foundation-models.pdf, 1-2. It was coauthored with experts from NVIDIA, the University of Pennsylvania, Johns Hopkins University, the University of Oxford, the Centre for Genomic Regulation in Barcelona, and other companies and research centers.29Munsamy et al., “Designing AI-programmable therapeutics,” 1. The model, called EDEN, was built on a Llama3-style architecture and trained on up to 9.7 trillion nucleotide tokens from the company’s database.30Munsamy et al., “Designing AI-programmable therapeutics,” 9. For the collection of samples that went into the database used for EDEN’s training, the authors acknowledged contributions from numerous organizations around the world, including the British Antarctic Survey, Heritage Malta, CENIBiot in Costa Rica, Balaton Limnological Research Institute in Hungary, and AJESH International in Cameroon.31Munsamy et al., “Designing AI-programmable therapeutics,” 36.
This use case illustrates that a private-sector company in the health or biopharma sector can leverage global networks of both corporate and nonprofit organizations (e.g., university) to build a dataset that is ultimately proprietary. Organizations across sectors can collaborate on AI-related data components and initiatives in ways where neither datasets nor other data components of the AI supply chain, such as AI models themselves, necessarily cleanly fall into public- or private-sector buckets.
Use case example #4: Proprietary but licensable, AI-enabling data
American AI medicine and genomic testing data company Tempus advertises several real-world datasets. It offers more than 8.5 million “deidentified research records” for scientific discovery, around four hundred thousand records with whole transcriptomic profiles, more than two million records with imaging data, more than 1.5 million records where matched clinical data is linked to genomic information, and a database the company says is one hundred times larger than the public Cancer Genome Atlas database.32“Tempus Real-World Data,” Tempus, accessed June 18, 2026, https://www.tempus.com/solutions/real-world-data/#overview. The company advertises two mechanisms by which other organizations can purchase access to Tempus data:
- Data collaborations: “Work with experienced Tempus scientists and biostatisticians to perform a scoped analysis based on your research objectives and criteria. Receive end-to-end support designing and executing retrospective data analyses.”
- Advanced analytics: “Leverage Tempus Lens to conduct preliminary analyses and identify cohorts of interest, then migrate your records to Workspaces to generate additional insights in our cloud-based coding environment. Layer on AI agents for faster data curation and abstraction.”33“Tempus Real-World Data.”
Tempus adds that it has data pipelines with over 4,500 US hospitals, “enabling us to ingest, structure, and harmonize multimodal data at an unprecedented scale.” Its use cases for this kind of data-meets-AI span oncology, cardiology, neurology, radiology, pathology, and other areas.34Ryan Fukushima, “Advancing the Frontier of AI in Healthcare,” Tempus, October 9, 2025, https://www.tempus.com/content/article/advancing-the-frontier-of-ai-in-healthcare/. The company says that 95 percent of the top twenty pharma oncology companies (based on their publicly available 2024 segment revenue) and over 250 biopharma companies use its services.35“Tempus Real-World Data.”
This example illustrates that companies can build proprietary datasets that are themselves data components of the AI supply chain (e.g., potential training data) and can also be used to develop other components of the AI supply chain (e.g., testing data, AI models). It also illustrates that even though the databases themselves are proprietary—they are ostensibly unique to a company and privately held—the company can still make them available to others, for a price. Organizations building proprietary, AI-related health and biopharma datasets and making them available as a service could additionally represent the data as “deidentified,” potentially intending to speak to consent requirements or privacy laws in various jurisdictions.
Use case example #5: Cross-border national security risk
This fifth use case is quite different, intended to showcase the security risks of particular activities within the health or biopharma sectors.
BGI Genomics, a Chinese company, leverages health and genetic data to train AI models and produces other data components in the AI supply chain (such as AI models for health and biopharma functions). Its Genos model, for instance, is a ten-billion-parameter open-source genomic foundation model that is trained on 636 genomes from diverse populations36.“BGI Group CEO Dr. Yin Ye: AI Is Redefining the Efficiency and Boundaries of Genetic Testing,” BGI, March 17, 2026, https://en.genomics.cn/en-news-7477.html. BGI has also debuted an AI platform that ingests genomic, epigenetic, microbiome, imaging, and lifestyle data to “generate actionable insights for disease prevention and long-term health planning” shortly after it launched a clinical genomics laboratory in Riyadh.37“BGI Genomics Taps AI to Power Healthcare Transformation in Saudi Vision,” BGI, September 21, 2025, https://www.bgi.com/global/news/BGI%20Genomics%20Taps%20AI%20to%20Power%20Healthcare%20Transformation%20in%20Saudi%20Vision. “As [genetic] sequencing becomes more accessible, laboratories and clinicians face an explosion of data, while the capacity to interpret results with consistency and clinical relevance has not scaled at the same pace,” the company has written.38“AI Is Redefining the Efficiency and Boundaries.” And, as the company paraphrased its CEO saying, “the next phase of progress depends on moving from generating data to understanding it—quickly, reproducibly, and in a way that can support real clinical and public health workflows.”39“AI Is Redefining the Efficiency and Boundaries.”
There is a US national security risk dimension to these activities. The US National Counterintelligence and Security Center (NCSC) has warned that while the bioeconomy can yield many benefits to the food supply, healthcare, environmental quality, and other important areas, actors such as foreign governments can use genomic data and technology to “identify genetic vulnerabilities in a population” and surveil people.40Protecting Critical and Emerging U.S. Technologies from Foreign Threats, National Counterintelligence and Security Center, October 2021, https://permanent.fdlp.gov/gpo173109/FINALNCSCEmergingTechnologiesFactsheet10222021.pdf, 6. In particular, the agency has stated that Chinese firms are collecting genetic data globally in an effort to develop the largest bio-database in the world.41Julian E. Barnes, “U.S. Warns of Efforts by China to Collect Genetic Data,” New York Times, October 22, 2021, https://www.nytimes.com/2021/10/22/us/politics/china-genetic-data-collection.html. For example, the US Department of Defense updated its formal list of entities identified as Chinese military companies (the so-called 1260H list) in October 2022 to add BGI Genomics.42US Department of Defense. Entities Identified as Chinese Military Companies Operating in the United States in Accordance with Section 1260H of the William M. (“Mac”) Thornberry National Defense Authorization Act for Fiscal Year 2021 (Public Law 116-283). Arlington: Department of Defense, October 2022. https://media.defense.gov/2022/Oct/05/2003091659/-1/-1/0/1260h companies.pdf, 1. BGI Group is affiliated, per a June 2026 update to the list, with China’s People’s Liberation Army (PLA) and “is a military-civil fusion contributor to the Chinese defense industrial base.”43US Department of Defense. Entities Identified as Chinese Military Companies Operating in the United States in Accordance with Section 1260H of the William M. (Mac) Thornberry National Defense Authorization Act for Fiscal Year 2021 (Public Law 116-283, Section 1260H, as amended) (Arlington: Department of Defense, June 2026. https://media.defense.gov/2026/Jun/08/2003945537/-1/-1/1/entities-identified-as-chinese-military-companies-operatiung-in-the-united-states-in-accordance-with-section-1260h.pdf, 3. A bipartisan group of members of Congress have worried that BGI’s access to the US market could provide it with genetic data that it can use to then undermine US national security.44Ken Dilanian, “Congress Wants to Ban China’s Largest Genomics Firm from Doing Business in the US. Here’s Why,” NBC News, January 25, 2024, https://www.nbcnews.com/politics/national-security/congress-wants-ban-china-genomics-firm-bgi-from-us-rcna135698. For example, there are concerns that such data could be used to fuel the development of bioweapons or in the service of next-generation capabilities whose details and subsequent risks have yet to be realized.
This case is relevant for several reasons. BGI appears to be conducting activity that many companies and other organizations around the world conduct: collecting genetic data, using genomic insights to predict health outcomes, building AI models using health and genomic data, and more. It does so globally, as do many other firms, as evinced by its recent AI and AI-related data work in Saudi Arabia (among other places). But BGI is also a company headquartered in China, which lacks strong rule of law, meaning that Chinese security agencies can easily compel BGI to hand over data of interest.45Beyond many press stories to this effect, the US government has publicized this risk with respect to health and genetic data in particular; US National Counterintelligence and Security Center. China’s Collection of Genomic and Other Healthcare Data from America: Risks to Privacy and U.S. Economic and National Security, National Counterintelligence and Security Center, February 2021, https://www.dni.gov/files/NCSC/documents/SafeguardingOurFuture/NCSC_China_Genomics_Fact_Sheet_2021revision20210203.pdf. Further, BGI reportedly cooperates with China’s PLA on Chinese defense and national security objectives. These phrases ostensibly refer to activities related to biosecurity or bioweaponry. Combined, this raises the risk that what might otherwise be innocuous or standard-risk activity—a company collecting AI-related health and genomic data, transferring it across borders, potentially exposing it to breaches, and the like—becomes a risk to US national security due to the Chinese government acquiring and using American or other genetic data for Chinese security ends.
The fractured global regulatory landscape
Countries around the world are debating how to apply existing regulations to AI-related data and AI technologies themselves or pass new laws to handle their transparency, innovation, fairness, security, privacy, and other implications. Other countries have already made decisions about how to apply existing regulations or have already passed new laws for similar purposes.
Some of these laws and regulations are AI-specific, including the EU AI Act and China’s rules for public-facing algorithmic and generative AI services. China’s AI-specific rules operate alongside broader personal information protection, data security, cybersecurity, health-sector, and human genomics resources requirements. A single health-related AI activity may therefore trigger several of China’s regimes at once.46Kenton Thibaut, Balancing Openness and Control: Cross-Border Health Data and AI Governance in China, Atlantic Council, June 23, 2026, https://www.atlanticcouncil.org/in-depth-research-reports/report/balancing-openness-and-control-cross-border-health-data-and-ai-governance-in-china/. In other cases, the laws and regulations in question are more general but can apply to AI data components and AI use cases. For example, the US Health Insurance Portability and Accountability Act (HIPAA), the US Federal Trade Commission’s (FTC) unfair and deceptive business acts and practices (Section 5 of the FTC Act) authority,47See, e.g., US Federal Trade Commission, “FTC Announces Crackdown on Deceptive AI Claims and Schemes,” September 25, 2024, https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes. and the Food and Drug Administration’s (FDA) regulation of medical devices writ large can all regulate AI-related data components in the health and biopharma sectors, even though they are not focused solely on AI or are not focused on it at all.48See, e.g., “Artificial Intelligence-Enabled Medical Devices,” US Food and Drug Administration, updated June 16, 2026, https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devices. And sometimes, countries and blocs are using laws in both buckets. Most prominently, the EU’s GDPR applies to AI-related data as one part of a broader tapestry of data protection requirements, while the European Union has also passed and continues to amend its AI Act to specifically regulate AI systems themselves.
Researchers, companies, and even government agencies continue to develop privacy-enhancing technologies to operate amid these regulations. Federated learning is a decentralized approach to machine learning where individual devices keep personal data on the device and collectively work to train a centralized model.49See, e.g., “What is Federated Learning?,” IBM, accessed August 24, 2022, https://research.ibm.com/blog/what-is-federated-learning. Differential privacy is a mathematical definition of privacy, whereby adding a person’s data into a dataset or taking it out of that dataset does not change the overall dataset’s behavior—ensuring that individual-level information about participants is not leaked.50See, e.g., “Differential Privacy,” Harvard University Privacy Tools Project, accessed August 13, 2022, https://privacytools.seas.harvard.edu/differential-privacy. Synthetic data generation aims to produce datasets that mirror real-world data so they are useful for tasks like healthcare model training, but without containing real personal data per se. And homomorphic encryption, to give one more example, allows organizations to perform computation on encrypted data.51See, e.g., “What is Homomorphic Encryption?,” IBM, accessed August 13, 2026, https://www.ibm.com/think/topics/homomorphic-encryption. Not all of these measures would necessarily bring a company into compliance with all the regulations in question, but they are technological developments for data privacy, obfuscation, and security that continue to evolve in tandem with complex regulations.
For their part, health and biopharma sector organizations collecting AI-related data and developing, deploying, procuring, using, testing, and maintaining AI systems must contend with this fractured global regulatory landscape. Some requirements are similar across regimes. For example, India’s Digital Personal Data Protection Act, like some other national-level privacy regimes, is loosely modeled after the EU GDPR. Health and biopharma companies operating under the EU GDPR may already be providing rights to consumers, such as the right to access their data, that also ensure compliance with the Indian law. In other cases, countries’ requirements for health and biopharma AI-related data components are significantly different or even contradictory. The United States and China, for instance, have both implemented restrictions on the outbound transfer on genetic and genomic data for national security reasons, which are intended to limit or outright prohibit the transfer of such data from one to the other.
Table 2 below shows fourteen exemplative regulatory regimes around the world that govern or substantially implicate health and biopharma data components of the AI supply chain, from genetic training data to medical image testing data to biotech AI models. The table is not meant to be comprehensive and certainly does not cover every law and regulation in question, let alone all the nuances, relevant case law, potential legal gaps, and so forth. However, it does capture the variety of laws and regulations that impact health and biopharma data components of the AI supply chain. In doing so, it shows how different countries approach these laws and regulations in different ways (such as regulating AI as its own specific issue area versus applying non-AI-specific laws and regulations to AI technologies) and how those decisions create a global patchwork.
The table specifies, for each of the fourteen exemplative regimes, the regime name; whether the regime is AI-specific or broad (but encompassing AI-related data components); whether the regime applies to a region, to a country, or to multiple countries in a cross-border fashion; a summary of the regime; the data components of the AI supply chain that are clearly covered by the regime; and a short description of the regime’s implications for the health and biotech sectors.
China, the European Union, and the United States have regulatory regimes that stack on top of one another. The Chinese case provides an illustrative example. A health dataset may qualify as sensitive personal information under China’s Personal Information Protection Law (PIPL) and as important data under its Data Security Law (DSL) and Cybersecurity Law (CSL) framework, while genomic content may also trigger the HGR and hospital-held information may face sector-specific storage or cybersecurity rules. Compliance with one pathway does not displace the others. A clinical dataset excluded from HGR information, for example, may still require PIPL consent, a personal information protection impact assessment, and a Cyberspace Administration of China (CAC) transfer mechanism.
China’s health-sector rules add a more targeted and granular regulatory layer. The 2014 Population Health Information Measures prohibit covered population health information from being stored on overseas servers. The 2018 Health and Medical Big Data Measures and 2022 healthcare cybersecurity rules (and the 2026 update) impose additional storage, lifecycle, access-control, and security duties.52https://www.nhc.gov.cn/mohwsbwstjxxzx/s8553/201809/742308fad8ed47fb85b9675ee4b6c6eb.html; see “China’s Hospitals Get Their Own Data Rulebook: Reading the 2026 Healthcare Data Security & PI Measures,” Data Compliance China, June 4, 2026, https://datacompliancechina.com/posts/china-healthcare-data-rulebook-2026/ The coverage of these measures depends on the institution, dataset, and processing context.
China has also eased some lower-risk transfer procedures. The 2024 CAC provisions raised numerical thresholds, expanded exemptions, extended the validity of successful security assessments, and authorized free-trade-zone negative lists.53Cùjìn Hé Guīfàn Shùjù Kuà Jìng Liúdòng Guīdìng (促进和规范数据跨境流动规定) Guójiā hùliánwǎng xìnxī bàngōngshì lìng dì 16 hào (国家互联网信息办公室令 第16号) [Regulations on Promoting and Regulating Cross-Border Data Flow, Order from the State Internet Information Office No. 16] (promulgated and effective March 22, 2024), Cyberspace Administration of China, https://www.cac.gov.cn/2024-03/22/c_1712776611775634.htm The resulting system reflects managed openness: lower-risk transfers may proceed through defined channels, while important data, data held by any organization legally classified as a critical information infrastructure operator (CIIO), large volumes of sensitive personal information, and covered HGR, remain subject to closer state review.
The terms in the laws and regulations themselves are also complex. Any organization complying with them will contend with both straightforward definitions and those that are opaque or vague, leaving much more up for interpretation. This, in turn, contributes to variations in regulatory enforcement and the need for health and biopharma companies to understand their regulators’ current and future enforcement postures. In the United States, for example, the FTC regulates unfair or deceptive business acts or practices under its Section 5 authority. The Commission defines “deceptive” practices as practices involving a material representation, omission, or practice likely to mislead a reasonable consumer.54“FTC Policy Statement on Deception,” US Federal Trade Commission, October 1983, https://www.ftc.gov/system/files/documents/public_statements/410531/831014deceptionstmt.pdf, 1. This definition is quite straightforward: companies wishing to avoid a deceptive practices claim should tell the truth to consumers.
By contrast, the Commission has long defined “unfair” practices as those that cause a substantial injury, which is not outweighed by any offsetting consumer or competitive benefits that the practice also yields, and which consumers could not reasonably have avoided.55“FTC Policy Statement on Unfairness,” US Federal Trade Commission, December 17, 1980, https://www.ftc.gov/legal-library/browse/ftc-policy-statement-unfairness. On its face, the definition is easy to follow. However, the FTC has fluctuated widely in recent years in its approach to enforcing “unfair” data or AI practices under Section 5 of the FTC Act. It has recently argued, breaking with some previous Commission postures, that certain uses of particularly sensitive personal data are in and of themselves harmful enough to make them unfair—such as companies selling identifiable data about individuals’ 24-7 movements, whether or not it is clear that the recipients of the data subsequently took any identifiable action with that data.56“FTC to Ban Kochava and Subsidiary from Selling Sensitive Location Data to Settle Charges They Sold Location Data Linked to Millions of Mobile Devices,” US Federal Trade Commission, May 4, 2026, https://www.ftc.gov/news-events/news/press-releases/2026/05/ftc-ban-kochava-subsidiary-selling-sensitive-location-data-settle-charges-they-sold-location-data. There are plenty of reasons why such disclosures in and of themselves cause harm.57For whatever it is worth, I agree with the FTC’s argument in the cited Kochava case. However, debate persists as to whether a category of behavior (such as selling certain kinds of data) should be considered “unfair” underscores the variation in how future FTC enforcers could interpret this definition, and thus scrutinize health or biopharma company handling of AI-related data components, compared to the more obvious definition of a “deceptive” practice.
Most of these laws put controls on how health or biopharma companies and organizations could collect, analyze, transfer, and use any AI-related data components, such as training data composed of people’s genetic information. Most of the above-listed regimes do not specify what happens when a company transfers an AI model out of the country that was trained on data that the country’s laws and regulations do cover. For example, India’s Digital Personal Data Protection Act does not clearly describe this issue. The US Data Security Program does not clearly describe this issue, either—perhaps an area for a future advisory opinion from the Justice Department. The European Union presents a thornier situation, given the overlap of the GDPR) with the AI Act, and the related complexities of the EU-US Data Privacy Framework.
EU data privacy and security regulations do not clearly specify whether biopharma companies that train AI models on EU data can or cannot transfer the models themselves out of the European Union, without implicating the GDPR or other requirements. For example, in December 2024 the European Data Protection Board (EDPB) issued Opinion 28/2024 on the processing of personal data in the context of AI models. The EDPB said that “claims of an AI model’s anonymity should be assessed by competent [Supervisory Authorities] on a case-by-case basis, since the EDPB considers that AI models trained with personal data cannot, in all cases, be considered anonymous.”58“Opinion of the Board (Art. 64): Opinion 28/2024 on Certain Data Protection Aspects Related to the Processing of Personal Data in the Context of AI Models,” European Data Protection Board, December 2024, https://www.edpb.europa.eu/system/files/2024-12/edpb_opinion_202428_ai-models_en.pdf, 2. Essentially, it depends on the specific case whether or not an American, Japanese, or other organization could train an AI model on EU citizen data and transfer the resulting model out of the European Union without triggering GDPR obligations on the AI model itself. Doing so creates several complications.
On the one hand, there are plenty of reasons to take a case-by-case regulatory approach to organizations transferring AI models trained on EU citizen data out of the European Union. A health organization could use two similar health datasets to train two different AI models, making one a significant source of privacy concern (thus subject to some outbound-transfer restriction under GDPR) and the other not (thus unrestricted) due to variations in the model architectures. To give another example, two biopharma companies could train similar-functionality AI models on datasets with different levels of underlying data identifiability, likewise making the risks of subsequent training data leakage or model reverse engineering outside of the European Union much different.
However, making the rules vague and so case-specific creates uncertainty for health and biopharma companies seeking to plan for how to continue conducting clinical trials, vaccine development, and other such functions in an era of AI usage. This can lead to variability among each EU member state’s enforcement, meaning that a health or biopharma organization that collects EU citizen data from multiple member states under a single GDPR framework could subsequently be subject to varied post-collection enforcement by country on the resulting model’s transfer. Such enforcement variation can create a risk to privacy insofar as member states whose privacy enforcement bodies are particularly Big Tech-friendly could be incentivized to lower the floor of outbound transfer protections in cases where protections should be robust. It also leaves open-ended the question of what level of protection on the underlying data and the model itself is sufficient to mitigate risks of data leakage or reverse-engineering—because the risk cannot be eliminated entirely.
China provides one partial clarification relevant to this issue. CAC guidance treats data stored in China as exported when a person or organization accesses, queries, retrieves, downloads, or calls that data from outside China.59“Q&A on Data Cross-Border Security Management Policies (October 2025)”, Office of the Central Cyberspace Affairs Commission, October 31, 2025, https://www.cac.gov.cn/2025-10/31/c_1763633376984070.htm The location of the access or calling activity controls this determination. However, whether that rule reaches overseas use of an AI model trained on protected Chinese data remains unresolved—the answer would likely depend on whether the model, its outputs, or the access process provides or reveals covered personal information or important data.
Most major data privacy and security regimes, as well as emerging AI regulatory regimes, have some explicit coverage of data or AI models relevant to health and biopharma. Within these regulatory frameworks, countries commonly define personal health data as any data describing someone’s past, present, and future conditions, treatments, and so on. Governments, even in countries with otherwise conflicting perspectives on tech regulation, frequently treat genetic data and genomic data as among the most sensitive, if not the most sensitive, data types. And many AI-specific and non-AI-specific measures cover as a sensitive category the use of AI models for health- and biopharma-related purposes, whether for public agency monitoring of disease outbreaks or private determination of diagnoses or treatments. The above-cited and other laws and regulations have varied terms, frameworks, restrictions, exemptions, and incentives for health and biopharma AI-related data components and use cases. But these treatments of health and biopharma data and AI technologies as important and sensitive are relatively consistent.
Lastly, as mentioned, some of these regimes represent clashing perspectives. EU politicians are increasingly discussing how to create regulations that hive off the European technology and data sphere from the United States and create a relatively autonomous European technology stack. Meanwhile, the current US federal government perspective, overall, is to remove and curtail enforcement of technology- and data-related regulations, with the goal of accelerating AI research and development—even if rolling back regulations creates substantial privacy, cybersecurity, and other risks to Americans, companies, and the country. The United States has implemented cross-border data transfer regulations for national security perspectives, especially focused on curtailing transfers of genetic and genomic data to Chinese organizations and significantly limiting transfers of personal health data to them.
China applies its own controls to outbound health and genetic data, while recent CAC rules and free-trade-zone pilots create more predictable channels for some lower-risk transfers. The Chinese government’s approach combines selective facilitation with retained state discretion over important data, CIIO-held data, large-scale sensitive personal information, and human genetic resources. These regimes are in conflict with one another. Opposing views of what data should flow where feed into the overall patchwork of laws and regulations for health- and biopharma-related AI data and data supply chain components.
Current tensions and future issues
Policymakers, companies, nonprofit research organizations, and other stakeholders charting the future of health and biopharma, data components of the AI supply chain, and responsible innovation should use the areas covered and gaps left by the legal and regulatory patchwork to identify paths forward. Both the benefits and the frictions that these laws and regulations create are instructive for identifying areas for improvement and harmonization.
This section starts by describing some of the reasons why the health and biopharma sectors’ interactions with training data, testing data, models (themselves), model architectures, model weights, APIs, and SDKs are differentiated compared to other sectors. These are:
- Healthcare is not advertising. Health and biopharma functions related to the data components of the AI supply chain—compiling training datasets; transferring protein folding or genetic-analysis AI models across borders; using AI models to generate synthetic health or genetic data—can carry significant risks but, on the whole, still serve fundamentally important societal functions around healthcare, disease research, vaccine development, and the like.
- Some health and biopharma activities, either by science or regulation (or both), require or benefit from access to data from multiple countries to function effectively. Data from different populations across countries can ensure that treatments indeed work across populations, improve health equity, significantly drive down the subsequent costs to societies of treatments that only work for a subset of the population, and help better tackle global diseases and pandemics. Some regulatory structures can also facilitate or encourage this multi-country, multi-population data collection for maximally effective, ethical clinical trials.
- Health and biopharma activity related to AI data components is often guided by quite robust ethical frameworks that cover patients, health decisions, and patient data. Informed patient consent, for example, is a bedrock principle of healthcare that is quite relevant for laws, regulations, and practices related to the data components of the AI supply chain, such as how data could be used in AI model construction or subsequent AI-driven diagnostics or analysis.
After detailing what makes the health and biopharma sectors unique in this area—compared to, say, activities carried out by social media companies or general-purpose LLM chatbot vendors—this section highlights some of the open legal and regulatory questions for policymakers and other stakeholders. These span whether or how laws and regulations can or should reference existing data best practices for health and biopharma entities, such as from requirements for clinical trial data; if health and biopharma activities should be treated differently under the AI-related and AI data-related regimes; how organizations should approach the export of AI models trained on covered health or genetic data, as well as the provision of those models via API to entities outside the jurisdiction; how synthetic data-generating AI models fit into the regulatory picture; and, among others, whether national security restrictions on cross-border data transfers should or should not have exemptions for healthcare and biopharma activity.
Listing these questions is designed to serve two purposes. These questions connect the prior section’s discussion of the global regulatory patchwork to this section’s discussion of what is unique about the health and biopharma sectors. And these questions also tee up issues for the subsequent, concluding discussion and recommendations.
What is unique to health and biopharma versus other sectors?
Many of the tensions that the legal and regulatory patchwork for the AI supply chain’s data components pose are not unique to the health and biopharma sectors. Countless global organizations seek to transfer data across borders. Many companies face regulatory uncertainty when building new technologies or deploying them with impacts on society or specific populations. But there are several reasons why these health and biopharma activities are different from other use cases, whether in transportation, finance, or advertising. Spelling them out facilitates a more nuanced conversation about future regulations.
The first, overarching point is that healthcare is not advertising. Disease research, clinical trial development for vaccines, and other activities are fundamentally different from the running of advertisements for retail goods, entertainment events, technology services, and the like. Certainly, these health- and biopharma-related activities (and sectors) have their fair share of problems, including critical issues around equity and privacy. This statement is not intended to downplay those issues. Nonetheless, a private-sector company collecting individuals’ health conditions so it can target them with an ad on a social media platform is different than a private-sector company collecting individuals’ health conditions so it can map patterns in surgical needs, improve early-stage diagnostics with AI models, or monitor for signs of potential disease outbreaks. The latter activities serve different purposes with far greater potential for societal benefit.
Again, the nature of health and biopharma work can make the AI-related and AI data-related risks of errors, biases, data leaks, and the like much more serious than in other contexts; a major error in a system’s diagnostic finding has the potential to hurt someone more than a poorly targeted social media ad.60Again, there are nuances, but this is meant as a general statement. But the fact remains that regulations recklessly accelerating,61This includes policymakers significantly weakening regulations, eliminating them entirely, or choosing not to implement new ones to address novel risks. responsibly guiding, or curtailing62This includes regulations that create enough friction or uncertainty to significantly and harmfully curtail important research and other activities. health research, clinical trials, vaccine development, and disease surveillance will have more immediate impacts on human life, human wellness, and societal safety than those regulations impacting other sectors that intersect with the data in the AI supply chain.
Second, some health and biopharma activities, either by science or regulation (or both), require or benefit from access to data collected from multiple countries to function effectively. On the scientific front, having more data from multiple countries can improve health equity and ensure that proposed treatments are effective for different groups of people. Groups historically underrepresented in or excluded from clinical research can have “distinct disease presentations or health circumstances that affect how they will respond to an investigational drug or therapy,” meaning that comprehensive, including globalized, datasets can help improve the generalizability of findings.63Kirsten Bibbins-Domingo, Alex Helman, eds., “Why Diverse Representation in Clinical Research Matters and the Current State of Representation within the Clinical Research Ecosystem,” in Improving Representation in Clinical Trials and Research: Building Research Equity for Women and Underrepresented Groups (Washington, DC: National Academies Press, 2022), https://doi.org/10.17226/26479. Beyond the moral implications of not covering these populations in health and biopharma work, failing to do so can also cost US society alone hundreds of billions of dollars, as medical problems that could have been avoided early on through inclusive clinical trials show up later and at much greater cost.64111Bibbins-Domingo, Helman, eds., “Why Diverse Representation in Clinical Research Matters”. In this context, companies, nonprofit research organizations, and others gaining access to data from different countries is therefore not (or is not merely) a “surveillance” function, such as an advertising company or chatbot LLM vendor wishing to profile even more people through training data. Representative data can help to ensure that therapies work across populations,65See, e.g., Annette S. Gross et al., “Clinical Trial Diversity: An Opportunity for Improved Insight into the Determinants of Variability in Drug Response,” British Journal of Clinical Pharmacology 88, no. 6 (February 2022): 2700-2717, https://pmc.ncbi.nlm.nih.gov/articles/PMC9306578/. which especially important in situations such as global pandemics. For example, the pharmaceutical companies that developed a COVID-19 vaccine leveraged clinical trial participants and data from around the world to understand the virus and develop a mitigation.66See, e.g., “Pfizer and BioNTech Conclude Phase 3 Study of COVID-19 Vaccine Candidate, Meeting All Primary Efficacy Endpoints,” Pfizer, November 18, 2020, https://www.pfizer.com/news/press-release/press-release-detail/pfizer-and-biontech-conclude-phase-3-study-covid-19-vaccine.
Regulations can simultaneously lead health and biopharma companies to pursue multi-country data collection and analysis. The FDA issued draft guidance in September 2024, for instance, encouraging sponsors to use, where appropriate, clinical data from outside the United States for oncology clinical trials—to ensure that the data applies to the diverse US population and accounts for the prevalence, presentation, causes, or severity of a disease across countries or regions.67“FDA Issues Draft Guidance on Conducting Multiregional Clinical Trials in Oncology,” US Food and Drug Administration, September 16, 2024, https://www.fda.gov/news-events/press-announcements/fda-issues-draft-guidance-conducting-multiregional-clinical-trials-oncology. For discussion of how the FDA’s posture on this may effectively change or have changed since then, see, e.g., Jacqueline R. Berman, “Key Considerations for Foreign Clinical Trials When Looking Abroad for Product Development,” Morgan Lewis, April 1, 2025, https://www.morganlewis.com/pubs/2025/04/key-considerations-for-foreign-clinical-trials-when-looking-abroad-for-product-development. European regulators have developed a unified system through which clinical trial sponsors can submit applications to run a clinical trial in several European countries, rather than the prior process of submitting separate applications to national authorities and ethics committees in each country.68“Clinical Trials Regulation,” European Medicines Agency, accessed June 21, 2026, https://www.ema.europa.eu/en/human-regulatory-overview/research-development/clinical-trials-human-medicines/clinical-trials-regulation.
Third, health and biopharma activity related to AI data components is often guided by quite robust ethical frameworks that cover patients, health decisions, and patient data. This is the case even if regulatory requirements and coverage can vary country to country. Informed consent, for one, is a cornerstone of medicine in many countries around the world, centering the right of patients to make informed and voluntary treatment decisions.69See, e.g., Parth Shah et al., “Informed Consent,” National Institutes of Health, November 24, 2024, https://www.ncbi.nlm.nih.gov/books/NBK430827/. Under this ethical principle, providers should fully inform patients about the nature of procedures or interventions, the potential risks and benefits, and the available, alternative treatments.70Shah et al., “Informed Constent.. These are notions of consent that far surpass that imposed in practice in many privacy laws and regulations, such as the United States’ broken notice-and-choice approach to consumer disclosure through lengthy, hard-to-read, and take-it-or-leave-it privacy policies.
For clinical research in particular, the 1947 Nuremberg Code, the 1979 Belmont Report, the 1991 US Common Rule, the 2000 Declaration of Helsinki (updated last year), and the 2002 Council for International Organizations of Medical Sciences (CIOMS) guidelines (updated in 2016) are used around the world to guide its ethical conduct.71“Ethics in Clinical Research,” US National Institutes of Health, accessed June 21, 2026, https://www.cc.nih.gov/recruit/ethics. Together, and along with other sources, professionals reference seven principles that should guide clinical research activities: social and clinical value, scientific validity, fair subject selection, favorable risk-benefit ratio, independent review, informed consent, and respect for potential and enrolled subjects.72“Ethics in Clinical Research.” The World Health Organization, to give another example, emphasizes centering public health needs and equitable health advancements; involving patients, the public, and communities to align research with public needs and sustain trust; ensuring randomized controlled trials are ethical, efficient, informative, risk-based, and proportionate; and addressing local health research needs and expanding cross-border collaborations, among others.73“Guidance for Best Practices for Clinical Trials,” World Health Organization, accessed June 21, 2026, https://www.who.int/our-work/science-division/research-for-health/implementation-of-the-resolution-on-clinical-trials/guidance-for-best-practices-for-clinical-trials. Although not surveyed comprehensively here, extensive literature exists on these subjects.
This is not to say that health and biopharma activity related to AI data components should not be regulated—far from it. But to underscore that unlike some other sectors, where AI-applicable ethical frameworks are still nascent, many relevant concepts already exist for the health and biopharma spaces. Some guidelines may fall short, and some organizations may not follow them. Nonetheless, having many ingrained codes of ethics adds a layer of both thinking about responsibility and AI-useful frameworks to the discussion. Policymakers building AI-related or AI data-related regulations, companies determining what uses of AI-related data components are ethical, and stakeholders interacting with the health and biopharma AI ecosystems can all reference these concepts to understand tradeoffs, identify opportunities and problems, and build solutions.
Additionally, existing heath and biopharma regulations could apply to AI use cases, serving as another layer of governance. There are instances in which companies or other organizations use AI to develop a medicine or vaccine, but where the medicine or vaccine still goes through its own regulatory approval. Use of AI data components or AI models can still introduce new risks—older health or biopharma regulations may not cover every scenario or harm—but an underlying layer of governance can be developed over it.
Open legal and regulatory questions
The current legal and regulatory landscape leaves questions unanswered for health and biopharma organizations engaged in activities related to the data in the AI supply chain. These open questions, which policymakers and other stakeholders should seek to address in the coming months, include:
- How do data privacy, cross-border data transfer, AI model, and other regulations point to or draw lessons from health best practices, such as clinical trial data requirements?
- Should laws and regulations treat health and genetic data used in an AI context for health research, drug development, and similar public-interest functions differently than the same data used by organizations for purposes such as online advertising?
- Can organizations train an AI model on health or genetic data from one country (or bloc) and then transfer the model beyond that jurisdiction? Are there requirements they must satisfy before doing so? If the organizations need to obtain a case-by-case determination from a regulator, is there any standard that the regulator uses to make such a determination?
- What if the health or biopharma organization trains an AI model on health or genetic data from that jurisdiction and does not transfer the model from the jurisdiction per se—for instance, hosting it on a server in that jurisdiction—but makes it available to internal or external users outside of that jurisdiction—such as via API? Is that prohibited? What factors go into the determination, and what is the legal basis for the regulator’s position?
- If an organization builds a synthetic data generator based on underlying health or genetic data that is covered under a regulatory regime, can the organization move the model outside of the jurisdiction? Can it access the model from outside the jurisdiction or make it accessible to others beyond the jurisdiction? Can the organization generate synthetic health or genetic data from the model, keep the model in the jurisdiction, but move the generated, synthetic data beyond the jurisdiction?
- Is the current legal and regulatory predominant focus on the training data, testing data, and AI models of the health and biopharma AI supply chain, as captured in the table in the previous section, going to persist? Or will future laws and regulations increasingly, explicitly govern other health and biopharma data components of the AI supply chain, such as model weights or APIs, as well?
- Do regulators perceive that national security-driven restrictions on cross-border data transfers of health or genetic data apply to AI models trained on that data? If so, how, and with what factors in play? What if the AI models themselves present a marginally low risk of regurgitating the underlying training data itself but could much more easily generate inferences about individuals that are generated based on the underlying training data?
- On the flip side, should national security restrictions on cross-border data transfers for the most high-risk areas be subject to exemptions, just because the functions may relate to health or biopharma activities? Where do and should policymakers draw the risk threshold? And how do they handle data components of the AI supply chain, such as health training datasets or AI models for sophisticated analysis of genetic sequences, that are strictly for health purposes versus those that are dual-use—where the civilian (including commercial) application could be equally used in a military context?
Conclusion and recommendations
Healthcare and biopharma AI is a promising area of AI research and development, at the same time as it raises critical questions about transparency, auditability, privacy, cybersecurity, inequity, national security, and more. These questions are relevant for the United States, the European Union, China, India, and many other countries and regions around the world. They are also relevant across every data component of the AI supply chain, as shown in this report: training data, testing data, models (themselves), model architectures, model weights, APIs, and SDKs. For policymakers to enable responsible innovation—safely and securely promoting vital healthcare services, disease research, vaccine development, and so forth while mitigating risks—they must tackle some of the current challenges and tensions between health and biopharma AI and AI-related data use cases, a patchwork of global laws and regulations, outdated regulatory concepts (such as definitions of “anonymization”), and a governance discussion driven by LLM chatbots rather than a wide range of other AI technologies and use cases.
This report makes three core recommendations for governments, companies, and nonprofit research organizations navigating this frontier of health and biopharma, the data components of the AI supply chain, and future considerations around healthcare progress, scientific innovation, data privacy, cybersecurity, economic competitiveness, and national security:
- Governments should use existing best-practice privacy and cybersecurity principles when pursuing regulations on data collection, data use, cross-border data transfers, and AI model deployment that impact health and biopharma, including the healthcare profession’s patient consent ethical framework and standards bodies’ specifications for encryption, data minimization, and protection against reidentification. The countries’ legislatures and their respective enforcers of broader data regulations—including the US Federal Trade Commission and DOJ, the EU member states’ information commissioners, the Cyberspace Administration of China, and the Data Protection Board of India—should carry this out. They should ensure that the healthcare concept of informed consent is reflected in laws and regulations that impact AI-related data components for health and biopharma activities, drawing on global-consensus frameworks and documents to articulate that concept in writing for any organization handling the likes of health or genetic AI-related data components. Doing so should strengthen many laws’ and regulations’ notion of consent over the presently weak, advertising technology-driven “notice and choice” consent baseline—and could have the indirect effect of helping to harmonize notions of consent across jurisdictions, drawing on standard concepts in the medical field. Governments should amend any regulations that exempt “anonymized” or “deidentified” data from coverage to ensure that their definitions of what makes data “anonymized” or “deidentified” comport with the current and over-the-horizon computer science and statistics literature on reidentification attacks—as well as recent machine learning advances in reidentification. They should, generally speaking, ensure that any pre-transfer requirements for cross-border transfers of health or genetic data and AI models refer to globally recognized standards for system transparency, encryption, data privacy, and the like, drawing on standards from the International Organization for Standardization (ISO) and the US National Institute of Standards and Technology. For example, if a regulatory regime is going to require that a covered organization implement encryption before transferring patient data abroad, it should reference widely known and used encryption standards, such as from the ISO, rather than specify their own. Doing so would lower the burden for health and biopharma organizations complying with restrictions (including in multiple countries), make reference to already negotiated best practice standards, and potentially help to harmonize regulatory requirements across jurisdictions. And as part of these governance best practice principles, governments should, among others, make clear for any regulatory review processes the strict or average anticipated timelines for review, a process through which regulated organizations and individuals can seek advisory opinions or supplemental guidance on compliance, and whether there is a process through which organizations and individuals can appeal decisions and if so, what it is.
- The surveyed regions should issue clarifications about how the health and biopharma data components of the AI supply chain sit within their non-AI-specific and AI-specific laws, drawing on the tensions and open questions identified in this report. This recommendation is made not just because this report is health- and biopharma-focused but because of the particular importance of these sectors to society. For example, the US Department of Justice should clarify whether AI models trained on bulk data and then exported to China, Russia, Iran, North Korea, Cuba, or Venezuela would, from its perspective, implicate the federal Data Security Program—or whether that would sit outside the program’s scope. FDA regulators should also issue updated guidance on how they view health or biopharma AI models vis-à-vis their authority to regulate certain software functions, including what factors health or biopharma organizations might need to weigh to reasonably determine if the FDA would consider a particular AI model to be covered. To give another example, the European Data Protection Board should further clarify its December 2024 opinion about AI models trained on EU citizen data to provide clearer, articulated standards for member countries as to whether an organization could train an AI model on EU health or genetic data and then export it beyond the European Union without triggering additional GDPR obligations. All jurisdictions should clarify whether they would handle the export of an AI model trained on their citizens’ data differently from the provision of that same AI model to entities beyond its jurisdiction, via API: would the latter be permitted or subject to similar restrictions?
- Governments should clarify criteria for any situations where they will treat health or biopharma AI-related data activities differently than AI-related data activities in other sectors. As covered in this report, there are several reasons—including due to the societal importance of health research and innovation—for policymakers to consider health and biopharma AI-related data activities differently than they would other kinds, such as transportation data generation or the creation of facial recognition AI models. At the same time, it is also difficult to disentangle health and biopharma AI-related data components by data type or by component use case. For data types, a hospital research center could theoretically use a database of patient records for AI training just as an advertising company could theoretically use that same database for AI training—meaning rules focused only on the data per se would rope in one activity that could be ethical with another activity that could be invasive and harmful. For data use cases, a company made up of clinical experts could theoretically develop a cancer diagnostic AI model grounded in sound methodologies and medical ethics (something a policymaker might want to happen, subject to responsible constraints), while another company made up of untrained hobbyists could seek to develop a “cancer diagnostic AI model” based on junk understandings of science (something policymakers might wish to block); hence, rules focused only on the use case and not the outcome could also rope in a range of activities, both desirable and undesirable. All to say, policymakers face complicated tasks in ensuring that scientifically sound, methodologically rigorous, and ethical health and biopharma innovation can advance while simultaneously preventing and mitigating the harms of other actors using health and biopharma AI-related data components. This report has offered several criteria that policymakers should consider in this effort: protections on the data itself, the ethical guidelines underpinning the AI research and development process, and the principle of informed consent cutting across all data component-related activities, among others. For example, policymakers could mandate that before any AI data component or AI use case is even considered for exemptions to accelerate health and biopharma research, the data components must satisfy strict data protection and informed consent requirements. Policymakers should at minimum strive to make clear to the public—in legislative development processes, regulatory oversight, and much more—what guides their thinking on how to carve out, if at all, health and biopharma activities from others. Doing so will be increasingly critical as data and AI regulations that impact health must also consider myriad other factors, from sovereignty to privacy to national security. Regulatory guidance could come from the enforcers in the listed countries of the broader privacy and data security rules—including the US Federal Trade Commission, the EU member states’ information commissioners, the Cyberspace Administration of China, and the Data Protection Board of India. A broader strategy on this front should come from the respective executive branch or agency bodies.
About the author
Justin Sherman is the founder and CEO of Global Cyber Strategies, a Washington, DC-based research and advisory firm; a contributing editor at Lawfare; and a columnist at Barron’s. He is the author of the book Navigating Technology and National Security.
Acknowledgements
The author would like to thank Diana Pankevich, Lee Licata, Devin Lynch, Ranjit Kumble, Kenton Thibaut, and Trey Herr for their comments on earlier drafts of this report. Thanks especially to Kenton Thibaut for extensive assistance on China’s sprawling regulatory apparatus and to Diana Panvkeich and Ranjit Kumble for discussion of health and biopharmaceutical industry subjects. Finally, thanks to Nitansha Bansal, Safa Shahwan Edwards, Nancy Messieh, and the rest of the team for support on the project and this report in particular.
Related Reading
Explore the program

The Atlantic Council’s Cyber Statecraft Initiative, part of the Atlantic Council Technology Programs, works at the nexus of geopolitics and cybersecurity to craft strategies to help shape the conduct of statecraft and to better inform and secure users of technology.

