
Four groups of questions, in the order most teams meet them. Each check ends with a question you can take straight to a data holder.
Most medical affairs teams now face the same request from several directions. Payers want to know how a therapy performs in local patients. Clinicians and advisory boards ask whether global trial results hold in the populations they treat. Regulators in the region are increasingly open to real-world evidence, but they expect it to be relevant and reliable. Global datasets, often drawn largely from North American and European patients, do not always answer these questions convincingly.
The difficulty is that local data in Asia rarely sits in one place. It is spread across public and private hospitals, clinic networks, laboratories, insurers and patient support programmes, captured in different systems, formats and languages. A source can look promising in a proposal and still fall short once you examine what is inside it. The checks below are designed to surface those gaps early, when changing course is still cheap.
Before looking at any source, write down the question in one sentence and name the audience it has to convince. A payer asking about hospitalisation costs, a clinician asking about real-world dosing, and a regulator asking about safety in a local subgroup each need different variables, follow-up and rigour. When the question comes first, you can judge a source on whether it answers that question, rather than on its size or on what the data holder is keen to sell.
Ask the data holder: which variables in your data would answer this exact question, and for how many patients?
A dataset can be large and still not represent the people your therapy serves. Compare the age profile, ethnicity mix, disease severity, comorbidities and payment model (public, private, insured, self-paying) with the population in your evidence plan. In many Asian markets, private and public patients differ meaningfully in access, follow-up and treatment patterns, so a source drawn mostly from one sector may not generalise to the other.
Ask the data holder: can you share a de-identified summary of the population by age, sex, ethnicity, care sector and disease stage?
Every source sees part of the patient journey. Hospital records capture admissions and procedures but often miss what happens in primary care. Laboratory data is rich in results but thin on diagnoses and treatment. Insurer claims show what was paid for, not always why. Patient support programme data reflects enrolled patients, who may be more engaged than average. Knowing the blind spots of each setting tells you whether you need one source or a combination.
A typical pattern, not a rule. Every source differs, which is why it is worth asking.
Ask the data holder: which parts of the patient journey does your data cover, and where do patients typically leave your view?
Data collected for billing, for clinical care and for research behaves differently. Find out how, why and by whom each part of the dataset was captured, which systems it came from, and whether any of it has already been transformed or summarised before reaching you. Provenance is also what regulators and reviewers look for when they judge whether evidence is reliable, so it is worth documenting from the start rather than reconstructing later.
Ask the data holder: for each key variable, what was the original source system and purpose of collection, and what has been done to it since?
A dataset can hold millions of records and still be missing the three fields that matter most to you. Check completeness for your outcomes, exposures and key covariates specifically: diagnoses, medications and dosing, laboratory values with their units and reference ranges, imaging findings, and dates. Ask for completeness rates, not a list of fields, because a field that exists for 12% of patients rarely supports a credible analysis.
Ask the data holder: what percentage of patients have a recorded value for each variable we need?
Many evidence questions depend on sequence: what happened before treatment, what changed after, and how long it lasted. That requires records that can be linked for the same patient across visits, sites and settings. In markets without a single national identifier in routine use, or where patients move between public and private providers, linkage can be the single biggest limitation of an otherwise promising source.
Ask the data holder: how are records for the same patient linked, and what is the typical length of follow-up?
In much of Asia, a large share of clinically useful information still lives in PDFs, scanned reports, images of paper records and free-text notes, often in more than one language within the same patient file. The same laboratory test may be recorded with different names, units and reference ranges across sites. None of this makes a source unusable, but it does change the time and effort needed before analysis can begin. A sample of real records tells you far more than a data dictionary.
Ask the data holder: can we review a small, de-identified sample of records exactly as they are stored today?
If the data holder describes the data as "standardised" or "harmonised", ask what that involved. Making records usable across sources means more than assigning codes: it includes extracting information from documents, cleaning errors, resolving the many variant ways the same thing is captured, respecting the local code sets some institutions maintain, applying clinical review and quality checks, and then mapping to medical coding standards. Each step involves judgement, and the quality of that judgement determines whether your analysis reflects clinical reality.
Each step involves clinical judgement. Ask who applies it and how accuracy is measured.
Ask the data holder: who reviews the transformed data for clinical accuracy, what are their qualifications, and what accuracy do you measure?
Data collected for care or for one programme is not automatically available for research or commercial evidence generation. Check what patients consented to, whether secondary use is permitted, and whether any ethics approval or data access committee sign-off is needed. Permissions also shape what you can do later: whether results can be published, shared with a partner, or combined with another dataset. It is far easier to confirm this before contracting than to discover a restriction mid-study.
Ask the data holder: what is the legal basis for this use, and can you show us the consent language or approvals that cover it?
Data protection laws across the region, such as Singapore's PDPA, the Philippines' Data Privacy Act and Thailand's PDPA, set conditions on how health data is handled and when it can leave the country. Some institutions and governments go further and expect health data to be processed in country. Establish where the data will be stored, processed and analysed, who can access it, and whether your global teams or vendors will need to see record-level data at all.
Ask the data holder: where will the data be stored and processed, and can analysis be carried out in country if required?
Removing names and identity numbers is rarely enough on its own. Rare conditions, small towns, exact dates and free-text notes can all make a record identifiable when combined. Ask how the data has been pseudonymised, how re-identification risk was assessed, and whether that assessment still holds once the data is linked with other sources. This protects patients first, and it also protects your organisation and your evidence from being challenged later.
Ask the data holder: how was the data pseudonymised, who assessed re-identification risk, and does that assessment cover linkage with other data?
The licence fee is usually the smallest part of what a real-world data project costs. The larger costs sit in preparing the data (extraction, cleaning, harmonisation and clinical review), in governance and approvals, in the internal time your team spends managing it, and in the delay if the data turns out to need more work than expected. A cheaper source that takes nine months to become usable can cost more, in money and in missed opportunity, than a more expensive source that is ready in two.
The licence fee is the part everyone sees. Plan and budget for the rest.
Ask the data holder: how long, and at whose cost, from signed agreement to an analysis-ready dataset for our question?
Use this table when comparing sources side by side. A source does not need to be perfect on every line, but each red flag should be understood and costed before you commit.
Health data API and real-time transformation engine. Any format in, clean structured output out.
Explore the engine →From 3-person startups to Fortune 500 insurers.
Talk to sales →Field notes, product updates, and customer stories from the health data frontier.
Browse all resources →




