AsiaTechDaily – Asia's Leading Tech and Startup Media Platform
The economics of building AI products are changing. Foundation models, APIs, open-source tools and increasingly accessible computing infrastructure are lowering the technical barriers that once separated AI companies from the broader market. For startups, that creates a harder question: if increasingly powerful AI capabilities can be accessed by competitors, what remains difficult to reproduce?
In healthcare, the answer may increasingly be real-world data. South Korea offers a useful case study. Its government is building a National Integrated Bio Big Data system targeting 772,000 participants during its first phase through 2028, combining clinical information, medical records, public data, genomic and other omics data, and personal health information. The longer-term target is one million participants by 2032.
At the same time, Korean healthcare AI companies are being pushed into real clinical environments. A 2026 medical AI commercialization program is supporting six hospital-company consortia to validate products, accumulate real-world data and pursue market entry. Lunit’s consortium, for example, plans to deploy AI diagnostic tools across 10 hospitals and accumulate 250,000 real-world cases. The emerging opportunity is therefore not simply to have data, but to build systems that continuously generate valuable data.
Healthcare data has always been valuable, but volume alone does not necessarily create defensibility. Clinical information can be fragmented across electronic health records, imaging systems, laboratory databases, genomic repositories and treatment records. Much of it is unstructured, incomplete or difficult to standardize. Even large datasets may therefore require substantial work before they can become useful for AI development or clinical research.
The more valuable datasets can instead combine multiple modalities while preserving the patient’s journey over time. Longitudinal information can reveal how conditions develop, how patients respond to treatments and how clinical decisions relate to outcomes. That distinction matters as AI becomes easier to build. While conversing with AsiaTechDaily, South Korea-based angel investor Sathishkumar Natarajan said the most valuable proprietary datasets are likely to be high-quality, longitudinal and clinically linked.
“I believe the most valuable proprietary datasets in Korea will be high-quality, longitudinal and clinically linked datasets, for example, clinical records combined with imaging, genomics, laboratory data, treatment history and real-world outcomes. However, data volume alone does not create a competitive advantage. The important factors are data quality, uniqueness, consistency, patient longitudinally, outcome linkage and how difficult the dataset is for competitors to reproduce. The strongest advantage can come from a continuous data-learning loop: the product generates real-world data, that data improves the AI or product, and the improved product generates more valuable data. That creates a potentially durable channel rather than simply having a large database.”
The distinction between a database and a learning loop could become increasingly important for healthcare AI startups.
Korea is actively trying to make healthcare data more usable for research and AI development. Its National Integrated Bio Big Data initiative is designed not only to collect information, but to standardize and quality-manage it before providing research access. The first phase targets 772,000 participants, including 585,000 members of the general population, 140,000 people with severe diseases and 47,000 with rare diseases.
The country is also expanding hospital-based AI experimentation. The 2026 Medical AI Testbed Support Program provides funding for collaborations between healthcare institutions with real hospital data and companies developing AI solutions using imaging, pathology, biosignals, voice, genomic and other medical datasets.
That could reduce one of the traditional barriers for healthcare AI startups: access to clinical environments. But it creates a paradox. If more companies can access healthcare data, access to data itself becomes less defensible. The competitive advantage could consequently shift toward the data that companies generate through deployment.
Consider the difference between acquiring a historical dataset and operating an AI product inside a healthcare workflow. The first provides information that may eventually be available to multiple organizations. The second can continuously generate new observations about how the technology performs in real-world settings. That creates a potential cycle:
Deployment → real-world data → product improvement → better clinical performance → wider deployment → more data.
Korea’s current medical AI commercialization programs are beginning to create precisely these conditions. Lunit’s 10-hospital consortium is intended to generate 250,000 real-world cases while pursuing insurance reimbursement. Its recent expansion in Sweden provides another illustration of how accumulated clinical experience can support international deployment. Lunit said its mammography AI had supported approximately 200,000 examinations at Stockholm’s Saint Göran Hospital over three years before expanding across the wider Stockholm Region, where the system is expected to support 200,000 to 250,000 examinations annually once fully implemented.
The significance goes beyond the number of examinations. Repeated use generates evidence about performance, workflow integration and clinical outcomes. That evidence can support product development, regulatory processes, customer acquisition and expansion into new markets.
Global healthcare AI companies are demonstrating the same shift. Tempus said in 2026 that its multimodal foundation-model efforts were drawing on more than 500 petabytes of data, including more than 45 million de-identified patient journeys and more than 400,000 cancer records containing genomic, transcriptomic, imaging and clinical data. Its models are designed around longitudinal records and outcomes including overall survival and progression-free survival.
But even at this scale, the challenge is not simply accumulating information. Tempus describes a data pipeline that integrates clinical, molecular, imaging, claims and other information to construct more complete patient journeys. This points to an important shift in how investors should think about healthcare AI defensibility. A startup’s data advantage may depend on its ability to collect, clean, link, validate and continuously learn from data, rather than the number of records it can claim.
For healthcare AI companies, several questions may become more important than dataset size:
This is also why Korea’s growing healthcare data infrastructure matters. Public initiatives can make underlying datasets more accessible, but startups that build proprietary feedback loops around clinical workflows may still create differentiated assets.
The foundation-model era could make AI capabilities increasingly accessible. That does not mean defensibility disappears. It may simply move further up the stack. In healthcare, the hardest asset to replicate may not be an algorithm trained once on a large dataset. It may be the accumulated system around an AI product: clinical relationships, longitudinal data, outcome linkage, workflow integration and the ability to continuously learn from real-world use. For startups and investors, the question is therefore shifting from “How much data do you have?” to “What data does your product continuously create that competitors cannot easily reproduce?” That could become one of the defining sources of defensibility in the next generation of healthcare AI.