AsiaTechDaily – Asia's Leading Tech and Startup Media Platform
The enterprise AI conversation has long revolved around increasingly powerful large language models, faster chips, and expanding cloud infrastructure. Yet for many organizations, the biggest obstacle to realizing AI’s full potential is far less glamorous: their own data.
Industry estimates suggest that 80% to 90% of enterprise information exists in unstructured formats, including emails, documents, PDFs, contracts, customer support tickets, meeting transcripts, call recordings, images, and internal reports. Unlike structured data stored neatly in databases, this information has historically been difficult to analyze, making it an underutilized corporate asset.
Generative AI has fundamentally changed that equation. Large language models can now interpret natural language, summarize documents, retrieve institutional knowledge, analyze conversations, and power enterprise copilots capable of understanding vast collections of business information. What was once considered dormant data has suddenly become one of an organization’s most valuable AI assets.
But there is a catch. Much of this information also contains personally identifiable information (PII), financial records, healthcare information, intellectual property, confidential business documents, or commercially sensitive customer data. For enterprises operating under increasingly stringent privacy and AI governance regulations, using that information for AI development introduces significant legal, security, and compliance challenges.
As enterprise AI moves from experimentation to production, organizations are realizing that their biggest challenge is no longer accessing powerful AI models. It is determining how to safely use the data they already possess.
For years, enterprise analytics focused primarily on structured information stored inside databases, spreadsheets, and enterprise resource planning systems because these datasets were relatively easy to organize and analyze. However, much of an organization’s operational knowledge has always existed elsewhere. Customer complaints are captured in emails. Product feedback is buried inside support tickets. Legal obligations reside in contracts. Institutional knowledge lives in technical documentation. Sales insights emerge from call transcripts, while operational decisions are scattered across presentations, meeting notes, and internal communications.
Collectively, these documents represent the contextual knowledge that enables organizations to make informed decisions. The arrival of generative AI has dramatically increased their strategic value. Instead of simply analyzing rows and columns, enterprises can now deploy AI systems capable of understanding language, identifying relationships across documents, retrieving historical information, and generating context-aware responses. Technologies such as retrieval-augmented generation (RAG), enterprise search, AI assistants, and domain-specific copilots all depend on access to rich, high-quality unstructured information. In many ways, enterprise AI has shifted from being a model problem to becoming a data problem.
If unstructured information holds enormous business value, why has it remained largely inaccessible? The answer lies in privacy. Unlike structured datasets that can often be anonymized through relatively straightforward methods, unstructured documents frequently contain sensitive information embedded throughout natural language. Customer names, medical histories, financial details, addresses, legal discussions, and confidential commercial information can all appear within the same document. This makes preparing enterprise data for AI significantly more complicated.
While conversing with AsiaTechDaily, Grant de Leeuw, Co-Founder and CEO of DataMasque, explained why organizations have struggled to unlock this category of enterprise information despite its immense value.
“Enterprise information and data sits across structured, semi structured and unstructured data throughout the organization and valuable data, such as PDF applications and call transcripts, can hold significant value. The limitation to date is unstructured data can often hold sensitive information that enterprises have struggled to address.”
His observation reflects one of the defining challenges facing enterprise AI today. Organizations are no longer constrained by a lack of AI capability. They are constrained by uncertainty around how to safely expose sensitive business information to increasingly powerful AI systems.
Protecting sensitive information is not a new challenge. For decades, organizations have relied on techniques such as redaction, masking, and anonymization when sharing or testing sensitive datasets. However, generative AI introduces a different requirement. Large language models derive value from context. Removing names, locations, relationships, or other contextual elements too aggressively can reduce the usefulness of the data itself, limiting its effectiveness for model training, testing, retrieval, or fine-tuning. As enterprises increasingly seek to operationalize AI, preserving data utility has become almost as important as protecting privacy.
Speaking with AsiaTechDaily, de Leeuw said this balance between privacy and usability is precisely where many organizations continue to struggle.
“DataMasque de-identifies this data across all datastores, including unstructured data. Rather than redacting the data, we replace it with consistent, synthetically identical values ensuring the context and utility of the data are retained. This allows enterprises to experiment, train or fine-tune on this data without exposing any protected or personal data.”
The broader significance extends well beyond a single technology platform. Across the enterprise AI landscape, organizations are increasingly exploring approaches such as synthetic data, de-identification, and privacy-preserving data processing that enable AI development without unnecessarily exposing sensitive customer information.
As governments introduce new AI regulations and organizations strengthen internal governance frameworks, responsible data management is evolving from a compliance requirement into a strategic capability. Markets such as Singapore have actively promoted trusted AI governance frameworks, recognizing that responsible AI adoption depends not only on technological innovation but also on transparency, accountability, and effective data governance.
This shift is particularly significant for highly regulated industries. Financial institutions, insurers, healthcare providers, telecommunications companies, and government agencies possess enormous volumes of valuable unstructured information. Yet these sectors also operate under some of the world’s strictest privacy and regulatory requirements.
For them, enterprise AI adoption increasingly depends on answering a fundamental question: How can organizations extract intelligence from sensitive information without compromising customer trust or regulatory compliance? The answer will likely shape the pace of enterprise AI adoption across many industries.
Growing interest in enterprise data infrastructure is also attracting investor attention. Earlier this year, DataMasque raised US$4 million in a funding round led by Wavemaker Ventures, with participation from existing investors OIF Ventures and Icehouse Ventures. Since its 2023 seed round, the company says it has achieved sixfold annual recurring revenue growth while expanding its customer base to include organizations such as New York Life, ADP, Best Western Hotels and Resorts, One NZ, TAL, and government agencies across New Zealand, Australia, and the United States.
The company is also expanding into Singapore, positioning itself within one of Asia’s leading AI governance ecosystems where enterprises face increasing pressure to balance AI innovation with regulatory compliance. The investment reflects a broader market reality. As organizations move beyond AI experimentation, technologies that enable secure access to enterprise data are becoming an increasingly important layer of the enterprise AI stack.
The rapid evolution of foundation models has transformed what AI systems are capable of achieving. Increasingly, however, the limiting factor is no longer the intelligence of the models themselves. It is the quality, accessibility, and governance of the information organizations choose to place behind them.
For many enterprises, decades of accumulated documents, customer interactions, operational records, and institutional knowledge represent an extraordinary competitive advantage waiting to be unlocked. Yet realizing that value will require organizations to solve one of enterprise AI’s most complex challenges: making sensitive information both safe and useful.
The companies that succeed in the next phase of enterprise AI may not necessarily be those with access to the largest models or the greatest computing power.
Instead, they are likely to be the organizations that can responsibly unlock their vast reserves of unstructured knowledge while preserving privacy, maintaining regulatory compliance, and retaining the context that makes enterprise information valuable in the first place.
As generative AI becomes embedded across every business function, the race will increasingly be won not by who builds the smartest models, but by who can safely transform decades of untapped enterprise knowledge into actionable intelligence.