Synthetic Patient Generation

Accelerate evidence generation by expanding limited patient cohorts with validated synthetic subject data.

When recruitment becomes the bottleneck, clinical research slows down

Most clinical development timelines are not delayed by science. They’re delayed by the reality of patient recruitment. Sponsors often face challenges like: 

Reduce recruitment risk without compromising quality

Strengthen statistical power, unlock subgroup insights, and accelerate decisions even when recruitment is limited.

Reduce recruitment risk without compromising quality

Strengthen statistical power, unlock subgroup insights, and accelerate decisions even when recruitment is limited.

Faster evidence generation

Instead of waiting months or years to recruit full cohorts, sponsors can build meaningful datasets sooner and accelerate early-stage decisions. 

Better-powered exploratory analyses

Synthetic subjects can increase sample size in a way that preserves statistical behavior, allowing more robust trend detection and hypothesis testing.

Stronger subgroup insights

Synthetic expansion allows deeper analysis of patient subpopulations, supporting more precise understanding of who responds to the treatment and why. 

Reduced recruitment risk

If recruitment stalls, you still have a way forward to extract insight, train classification algorythms, and salvage value from the data collected. 

Better preparation for future trials

By testing assumptions early and analyzing patterns more deeply, sponsors can design smarter studies with better inclusion criteria, endpoints, and timelines. 

How Synthetic Patient Generation works at Artialis

Our approach uses a validated synthetic subject generation framework developed by our partner Synthetrial. This involves the expansion real patient datasets while preserving key statistical relationships between variables.

Synthetic patients are generated as additional patient-level data, built from the structure and distributions of the real cohort.

This means we expand the cohort without inventing new variables or artificially changing the trial design.

Step 1: Data review and feasibility assessment

We begin by evaluating the dataset available. To generate high-quality synthetic subjects, the original dataset must meet key conditions:

  • A sufficient number of real patients (typically a minimum cohort is required)
  • High data quality and consistency
  • Minimal missing values across key parameters

If the dataset is incomplete, we assess whether it can be cleaned or restructured to support synthetic generation.

Step 2: In-depth analysis of study variables

We conduct a comprehensive analysis of how variables interact across the real patient dataset.

The objective of this step is to extract underlying behavioural and clinical patterns from real patients identifying how baseline characteristics, treatments, and outcomes relate to each other over time.

This phase is critical because synthetic data must preserve the structural integrity of the original dataset, including distributions, multivariate relationships, temporal dynamics, and outcome patterns (such as correlations, progression trends, and survival metrics).

By modelling these patterns accurately, we ensure that the synthetic cohort reflects the statistical behaviour of the real population without replicating individual patients.

Step 3: Synthetic subject generation

Once the real dataset is validated, we generate synthetic patients that reflect the statistical behavior of the real cohort.

The output is an expanded clinical trial database, where each synthetic patient dataset contains the same variables as the real patients and no personal identifiers are replicated or reused.

Step 4: Validation and quality controls

We validate synthetic cohort quality by comparing it against the original real cohort. This includes checking:

  • distributions of all endpoints
  • relationships between variables
  • variability patterns
  • survival or progression curves (when relevant)

The goal is to ensure the synthetic cohort behaves like the real population, without being a copy of it.

Built by clinical experts, powered by advanced synthetic data technology

We combine decades of clinical development experience with emerging synthetic subject generation capabilities to support evidence-driven innovation.

Our team understands both sides of the equation:

  • the scientific and statistical requirements
  • the clinical, operational, and regulatory reality of running trials

Synthetic patient generation is developed by our partner Synthetrial as part of our broader translational R&D hub model, designed to help sponsors move faster without compromising rigor.

Our FAQs

Synthetic patient generation is the process of generating statistically realistic synthetic subjects based on an existing real patient dataset. These synthetic patients are not real individuals, but they behave like the real cohort in terms of distributions, variability, and relationships between variables.

The goal is to expand your dataset so you can perform stronger analyses and make faster decisions without waiting for full recruitment.

No. 

Synthetic patients are simulated subject-level data generated from patterns in real datasets. They do not represent actual individuals and cannot be traced back to real patients

Synthetic patient generation is not random simulation.
It is based on real trial or registry datasets and built to preserve the underlying statistical structure of the cohort. That includes correlations between variables, outcome patterns, and realistic variability.

The synthetic cohort is designed to reflect the behavior of the real population, not just produce plausible-looking averages.

No. Synthetic patient recruitment is designed to complement real recruitment, not replace it.

Real patient data remains the foundation. Synthetic cohorts are used to strengthen analysis, explore subgroup patterns, and support faster internal decision-making.

This service is especially valuable when:

– recruitment is slow or difficult
– sample size is too small to reach meaningful conclusions
– subgroup analyses are impossible due to limited numbers
– a study ended early or produced inconclusive results
– you want to test feasibility and refine design before scaling

It is also useful when early-stage sponsors want to generate stronger evidence for go/no-go decisions or fundraising milestones.

Synthetic cohorts require an existing dataset with sufficient structure and quality. This typically includes:

– patient-level variables (baseline characteristics, outcomes, biomarkers, etc.)
– consistent data collection across subjects
– low levels of missing data for key parameters

The stronger the original dataset, the more reliable the synthetic expansion.

Yes. Synthetic patient generation is only meaningful if there is enough real data to model population patterns reliably. The minimum varies depending on the complexity of the dataset and number of variables.

During the feasibility assessment, we determine whether the dataset is suitable and what level of expansion is realistic.

FDA and EMA generally view synthetic data as supportive or supplementary, not a replacement of real data. Regulatory authorities evaluate clinical trial applications individually before approval to ensure that the validation method for synthetic data generation is robust.

Regulators focus mainly on:

– Validity & bias – synthetic data must faithfully reflect real patient populations and not introduce hidden bias.

– Traceability & transparency – clear documentation of how data were generated and validated.

– Statistical integrity – preservation of correlations, variability, and treatment effects relevant to trial endpoints.

– Regulatory acceptability – whether results based on synthetic data can support claims or decisions.

– Privacy protection – strong guarantees that no real patient can be re-identified.

Validation is a core part of our process. We compare synthetic data against the original cohort to ensure that key parameters match, including:
– distributions of endpoints and biomarkers
– correlations between variables
– outcome trends and variability
– subgroup patterns (where applicable)

Synthetic patients must behave like the real cohort statistically, without duplicating it.

No. Synthetic patients do not contain real personal identifiers and are generated in a way that prevents re-identification of individuals.

This approach is designed to reduce privacy risk compared to sharing real patient datasets.

Yes, in many cases.

If a trial was underpowered, ended early, or produced unclear results, synthetic expansion can help extract more insight from the data already collected.

This can support deeper analysis and help sponsors understand whether the trial failed because the product didn’t work, or because the study was not designed or powered appropriately.

Turn limited recruitment into meaningful evidence

If your trial is slow to recruit, underpowered, or limited by small subgroups, synthetic patient recruitment can help you extract more value from your data and move forward with confidence.