Industry

Finding Real-World Data in Asia for a Local Evidence Gap: 12 Things to Check Before You Commit to a Source

Suhina Singh
Suhina Singh
Founder & CEO, Jonda Health
October 4, 2026
·
10 min read

The 12 checks at a glance

Four groups of questions, in the order most teams meet them. Each check ends with a question you can take straight to a data holder.

CHECKS 1 TO 3 Question and patients CHECKS 4 TO 6 What is inside CHECKS 7 AND 8 Raw to usable CHECKS 9 TO 11 Allowed to use it CHECK 12 The real cost

‍

Most medical affairs teams now face the same request from several directions. Payers want to know how a therapy performs in local patients. Clinicians and advisory boards ask whether global trial results hold in the populations they treat. Regulators in the region are increasingly open to real-world evidence, but they expect it to be relevant and reliable. Global datasets, often drawn largely from North American and European patients, do not always answer these questions convincingly.

The difficulty is that local data in Asia rarely sits in one place. It is spread across public and private hospitals, clinic networks, laboratories, insurers and patient support programmes, captured in different systems, formats and languages. A source can look promising in a proposal and still fall short once you examine what is inside it. The checks below are designed to surface those gaps early, when changing course is still cheap.

 

Checks 1 to 3: start with the question and the patients

1. Start from the evidence question, not the dataset on offer

Before looking at any source, write down the question in one sentence and name the audience it has to convince. A payer asking about hospitalisation costs, a clinician asking about real-world dosing, and a regulator asking about safety in a local subgroup each need different variables, follow-up and rigour. When the question comes first, you can judge a source on whether it answers that question, rather than on its size or on what the data holder is keen to sell.

Ask the data holder: which variables in your data would answer this exact question, and for how many patients?

 

2. Check that the population looks like your patients

A dataset can be large and still not represent the people your therapy serves. Compare the age profile, ethnicity mix, disease severity, comorbidities and payment model (public, private, insured, self-paying) with the population in your evidence plan. In many Asian markets, private and public patients differ meaningfully in access, follow-up and treatment patterns, so a source drawn mostly from one sector may not generalise to the other.

Ask the data holder: can you share a de-identified summary of the population by age, sex, ethnicity, care sector and disease stage?

 

3. Understand the care setting the data comes from, and what it cannot see

Every source sees part of the patient journey. Hospital records capture admissions and procedures but often miss what happens in primary care. Laboratory data is rich in results but thin on diagnoses and treatment. Insurer claims show what was paid for, not always why. Patient support programme data reflects enrolled patients, who may be more engaged than average. Knowing the blind spots of each setting tells you whether you need one source or a combination.

Diagnosis Treatment Results Follow-up Hospital records Laboratory data Insurer claims Support programmes Usually strong Partial Often a blind spot

‍

A typical pattern, not a rule. Every source differs, which is why it is worth asking.

 

Ask the data holder: which parts of the patient journey does your data cover, and where do patients typically leave your view?

 

Checks 4 to 6: look at what is inside the data

4. Trace the provenance

Data collected for billing, for clinical care and for research behaves differently. Find out how, why and by whom each part of the dataset was captured, which systems it came from, and whether any of it has already been transformed or summarised before reaching you. Provenance is also what regulators and reviewers look for when they judge whether evidence is reliable, so it is worth documenting from the start rather than reconstructing later.

Ask the data holder: for each key variable, what was the original source system and purpose of collection, and what has been done to it since?

 

5. Confirm the variables your question depends on are complete

A dataset can hold millions of records and still be missing the three fields that matter most to you. Check completeness for your outcomes, exposures and key covariates specifically: diagnoses, medications and dosing, laboratory values with their units and reference ranges, imaging findings, and dates. Ask for completeness rates, not a list of fields, because a field that exists for 12% of patients rarely supports a credible analysis.

Ask the data holder: what percentage of patients have a recorded value for each variable we need?

 

6. Check that you can follow a patient over time and across settings

Many evidence questions depend on sequence: what happened before treatment, what changed after, and how long it lasted. That requires records that can be linked for the same patient across visits, sites and settings. In markets without a single national identifier in routine use, or where patients move between public and private providers, linkage can be the single biggest limitation of an otherwise promising source.

Ask the data holder: how are records for the same patient linked, and what is the typical length of follow-up?

 

Checks 7 and 8: understand the work between raw and usable

7. Face the format and language reality

In much of Asia, a large share of clinically useful information still lives in PDFs, scanned reports, images of paper records and free-text notes, often in more than one language within the same patient file. The same laboratory test may be recorded with different names, units and reference ranges across sites. None of this makes a source unusable, but it does change the time and effort needed before analysis can begin. A sample of real records tells you far more than a data dictionary.

Ask the data holder: can we review a small, de-identified sample of records exactly as they are stored today?

 

8. Find out how the data was made consistent

If the data holder describes the data as "standardised" or "harmonised", ask what that involved. Making records usable across sources means more than assigning codes: it includes extracting information from documents, cleaning errors, resolving the many variant ways the same thing is captured, respecting the local code sets some institutions maintain, applying clinical review and quality checks, and then mapping to medical coding standards. Each step involves judgement, and the quality of that judgement determines whether your analysis reflects clinical reality.

Raw records WHAT "STANDARDISED" SHOULD MEAN Extract fromdocuments Clean fix errors Resolve variantcaptures Review clinical QA Map to standards Analysis- ready data

Each step involves clinical judgement. Ask who applies it and how accuracy is measured.

 

Ask the data holder: who reviews the transformed data for clinical accuracy, what are their qualifications, and what accuracy do you measure?

 

Checks 9 to 11: confirm you are allowed to use it

9. Check permissions for your intended use

Data collected for care or for one programme is not automatically available for research or commercial evidence generation. Check what patients consented to, whether secondary use is permitted, and whether any ethics approval or data access committee sign-off is needed. Permissions also shape what you can do later: whether results can be published, shared with a partner, or combined with another dataset. It is far easier to confirm this before contracting than to discover a restriction mid-study.

Ask the data holder: what is the legal basis for this use, and can you show us the consent language or approvals that cover it?

 

10. Know where the data is processed and which cross-border rules apply

Data protection laws across the region, such as Singapore's PDPA, the Philippines' Data Privacy Act and Thailand's PDPA, set conditions on how health data is handled and when it can leave the country. Some institutions and governments go further and expect health data to be processed in country. Establish where the data will be stored, processed and analysed, who can access it, and whether your global teams or vendors will need to see record-level data at all.

Ask the data holder: where will the data be stored and processed, and can analysis be carried out in country if required?

 

11. Assess pseudonymisation and re-identification risk

Removing names and identity numbers is rarely enough on its own. Rare conditions, small towns, exact dates and free-text notes can all make a record identifiable when combined. Ask how the data has been pseudonymised, how re-identification risk was assessed, and whether that assessment still holds once the data is linked with other sources. This protects patients first, and it also protects your organisation and your evidence from being challenged later.

Ask the data holder: how was the data pseudonymised, who assessed re-identification risk, and does that assessment cover linkage with other data?

 

Check 12: count the real cost

12. Estimate the true time and cost to usable data

The licence fee is usually the smallest part of what a real-world data project costs. The larger costs sit in preparing the data (extraction, cleaning, harmonisation and clinical review), in governance and approvals, in the internal time your team spends managing it, and in the delay if the data turns out to need more work than expected. A cheaper source that takes nine months to become usable can cost more, in money and in missed opportunity, than a more expensive source that is ready in two.

What the proposal shows What the project costs Licence fee Data preparation Governance and approvals Internal team time Delay if the data needs more work

The licence fee is the part everyone sees. Plan and budget for the rest.

 

Ask the data holder: how long, and at whose cost, from signed agreement to an analysis-ready dataset for our question?

 

The checklist at a glance

Use this table when comparing sources side by side. A source does not need to be perfect on every line, but each red flag should be understood and costed before you commit.

#CheckWhat good looks likeRed flag
1Evidence questionQuestion and audience agreed before sourcingChoosing the dataset first, then finding a question
2Population fitSummary profile matches your target patientsNo population summary available
3Care settingBlind spots known and covered by another sourceOne setting assumed to show the whole journey
4ProvenanceSource system and purpose documented per variableData already summarised with no audit trail
5CompletenessCompleteness rates shared for your key variablesOnly a field list, no completeness figures
6Follow-upPatients linked across visits and settingsNo reliable way to link the same patient
7Format and languageReal sample reviewed before contractingSample refused or only a data dictionary offered
8ConsistencyClinical review and measured accuracy described“Standardised” with no explanation of how
9PermissionsLegal basis and consent cover your intended useSecondary use unclear or unconfirmed
10Processing locationStorage, processing and access clearly definedData must leave the country with no clear basis
11PseudonymisationRe-identification risk assessed, including linkageNames removed and nothing more
12True costTime and cost to analysis-ready data estimatedLicence fee is the only number discussed

 

How Jonda Health can help

We have seen these questions from several sides: as clinicians who needed the full story of a patient’s care, as people who worked in pharma medical and digital teams, and as the team that turns messy health records into data organisations can build on. You do not need to have every answer before talking to us. There are two ways we can support you.

Jonda Health Services

Find, assess and prepare the right data

Through our Data Sourcing and Health Data Foundations work, we help medical affairs and evidence teams identify potential sources in the region, work through checks like the ones above with the data holders, and assess what it will realistically take to get from raw records to an analysis-ready dataset. That way, the decision to commit is based on what is inside the data, not on what the proposal promises.

JondaX

Make the data usable

When the data you need sits in PDFs, scans, images and mixed formats across more than 10 languages, JondaX turns it into structured, standardised information through a single API. Our clinically trained reviewers check the output, achieving 99% field-level accuracy with expert review, and we can cater to local code sets as well as medical coding standards. JondaX is ISO 27001 certified, designed to comply with HIPAA, GDPR and PDPA, and can be deployed and processed fully within a client’s country where required.

Book a free 30-minute data readiness conversation

Bring your evidence question and the sources you are considering, and we will talk through where the risks are and what it would take to close the gap.

Get in touch at hello@jonda.health

Try JondaX free

30 free transformations. No credit card required.

Health data API and real-time transformation engine.See JondaX overview →
Not sure which fits your team?See all use cases →
Latest from JondaX.Browse all →