- OverviewClinical trials are becoming more expensive where, according to a Deloitte’s report1, the average cost to get a drug to market in the USA was $1.188 billion in 2010, and $1.981 billion in 2019. This increase in cost reflects the difficulties that are associated with current linear clinical trial designs. Clinical trials can take a long time due to difficulty in finding suitable/eligible patients for each study and the growing amount of data available used to plan or inform a trial. Current clinical trials are still evolving to make use of the rapidly developing technologies, scientific methods, and data availability of recent years.One possible way to improve and transform the clinical trial process as we know it would be through the implementation of Artificial Intelligence (AI). AI incorporates all intelligence demonstrated by machines and includes important aspects such as Machine Learning and Natural Language Processing. It is already widely used in modern technology, such as in smartphones and online website searches, but has also been used more recently to innovate the drug discovery process. AI implementation could be of benefit across diverse tasks is the planning, execution and analysis stages of clinical trials to improve cost effectiveness, study time, treatment efficacy, and quality of data2.Many emerging developments in artificial intelligence (AI) have the potential to benefit the clinical trials landscape.AI in Adaptive clinical trial designMany clinical trials are designed linearly, however adaptive designs are being used to allow predetermined changes during a study in response to ongoing trial data. Adaptive designs provide the flexibility to optimise resource allocation, end an unproductive trial early, and better characterise a treatment’s efficacy and safety through multiple endpoints. AI can help to inform and optimise adaptive designs through the analysis of healthcare data to select optimal endpoints, determine and monitor parameters for early stopping, and identify appropriate protocols for the trial3. These study design changes can increase the efficiency of a study, resulting in a trial which is more cost and time effective while maintaining high quality data collection and analysis.Meta-analysisAI can also enhance the exploratory analysis of data from previous trials. AI-enabled technology can collect, organise, and analyse increasing amounts of data, which could be applied to data from previous trials. This is normally performed manually by a biostatistician as part of a meta-analysis, but AI could help to gather and perform initial analyses before a more in-depth statistical approach is taken. This could highlight potentially important patterns in collected evidence which would then be used in informing trial design.Synthetic control armsCurrent clinical trials typically compare an experimental treatment to placebo and established treatments, assigning enrolled patients to either a treatment or control group. Synthetic control arms are an AI-driven solution for having a control group in a single-arm trial, which usually have only a treatment group. Synthetic control arms use data from previous studies to simulate the control treatment in patients. This would allow for all patients in a trial to receive active treatment to provide more evidence for the treatment efficacy and safety at each clinical trial stage. This may also have an impact on patient enrolment as patients generally show less interest in enrolling on placebo-controlled trials4,5.As synthetic control arms are relatively novel, comparisons should be made with traditional control groups. In a typical blinded trial, patients are unaware if they are receiving the experimental active treatment, or a placebo. This is to test that the new treatment is causative of any clinically meaningful response. Synthetic control arms could create more single-arm clinical trials, and the fact that patients are aware of receiving active treatment would be a factor given clinically meaningful results are found.Site selectionIdentification of suitable sites and investigators to perform a trial is an important factor in study efficiency and feasibility. A study site should be amply equipped to carry out a study, must be of a suitable size to process the needs of study participants, and be located in an accessible area to potential participants and investigators. Identification of target locations and investigators can be optimised through AI implementation, which also enables real-time monitoring of site performance once the trial has started. Study sites can be evaluated and compared through the development of a points-based algorithm which could factor in location, site size, and equipment.Patient enrolment/recruitmentMost patients enrol on a clinical trial if they have not responded to existing treatments in a clinically meaningful way. However, there are often strict eligibility criteria required for patients to enrol for a trial, including diagnostic tests, biomarker profiles, and demographics. Currently, patients find out about clinical trials either through manually searching online databases or, on occasion, through a clinician’s recommendation. This puts a lot of responsibility on patients to search for potential trials, just to be faced with trying to understand eligibility criteria full of medical jargon. This can lead to low patient recruitment.AI and natural language processing offer a potential solution for this issue. Natural language processing could be used to match patients with trials based on eligibility criteria and patient electronic health records. Potentially suitable trials could then be suggested to patients or their clinicians, making it easier for patients to find suitable trials and for trials to recruit patients. While this is an improvement to the current recruitment method, natural language processing may initially have some difficulty with clinical notes due to heavy use of acronyms, medical jargon, and deciphering of hand-written clinical notes. These problems aren’t specific to AI though, as currently patients looking for suitable trials have the same difficulties.AI-driven patient recruitment could also be used to reduce population heterogeneity and use prognostic and predictive enrichment to increase study power. While reduced heterogeneity can be beneficial (e.g. in testing the safety of drugs for patients unable to enrol on early clinical trials), caution should be taken with limiting patient diversity. Selectively enrolling patients who are more likely to respond well to treatment initially could lead to advancing a treatment that may not be as wel tolerated in the post-market patient population.AI balancing of diversity in clinical trials could be used to balance this. It is important to test safety and efficacy in with differing demographics for a treatment to be properly characterised. This may include different ethnic groups, body types (height, weight, BMI), ages, and sexes. For example, a drug intended for female patients should specify dosage programs and any specific adverse effects in female patients before approval.Patient diagnosticsTrial data often uses diagnostic tests to measure a patient’s response to treatment. AI can be used to improve diagnostic accuracy and objectivity during trials, reduce the potential for bias, and help with blinding a trial. AI programs have already been shown to provide more accurate diagnoses compared to clinicians6, which could be extended to use in clinical trials. AI can also be used to integrate multiple biomarkers or large data sets (e.g. bioinformatics data) to better monitor and understand a patient’s response, and to make any needed changes in dosing.Another application of AI in clinical trials would be in patient management. Wearable devices or apps could be used to provide real time safety and effectiveness data to both clinicians and patients, leading to better data quality and higher patient retention.Patient monitoring, retainment, & medication adherenceAI could be used to monitor patients through automatic data capture and digital clinical assessments. Automated monitoring with AI can allow for personalised adherence alerts, and wearable devices could provide real time safety and efficacy data shared with both patient and clinician to increase retainment and adherence rates. Video consultations with clinicians could improve retainment by reducing travelling required from patients, but there would still be a risk of drop due to travel as some diagnostic tests would need to be done at an appropriate site, which may warrant additional costs if tests for research purposes are not covered by a patient’s health insurance or national health service.Currently, medication adherence is mostly dependent on each patient’s diary/record keeping or memory, which is then discussed with clinicians during routine appointments. This can make it difficult to accurately track adherence. Digitising this through use of a website/app would allow for more accurate adherence data to be obtained, in addition to providing patients with notifications, educational content, and adherence records. Other medical devices such as timed medication bottles could also be used to ensure medication is used in appropriate intervals, and smart bottles could be used to synchronise this with a smartphone app if applicable.Data CleaningData cleaning for clinical trials is typically performed by trained biostatisticians and is essential to ensure that collected data is consistently formatted and free from inputting errors. However, the data cleaning process can be time-consuming, especially with large datasets collected during clinical trials. AI could be implemented through machine learning methods to identify and correct errors found in clinical trial datasets7, leading to better quality data which is optimised for analysis. An AI approach may also reduce the amount of time spent on data cleaning.AI implementation and clinical trial digitisationAI implementation is relevant in several aspects of clinical trials, be it in study design, patient diagnostics, or trial management. Several tech giants (including Apple & Google) have invested in developing solutions to process electronic health records, monitor patients remotely, and integrate healthcare data into devices. By improving the cost and time effectiveness of clinical trials, both patients and pharma-tech companies benefit with more affordably priced treatment costs for patients and greater return of investment for companies.However, for clinical trials to implement AI successfully, many aspects of clinical trials would first require digitisation. Many trials still use paper documents instead of digital alternatives, which results in lost documents and slowed trial progression. Integrating electronic health records, digital copies of clinicians’ notes, and digital patient monitoring alone would help in designing and managing a trial. There is a concern that a switch to digital may be difficult for patients unfamiliar with technology, or for those who might prefer to keep paper diaries. However, digital solutions would allow for the development and implementation of AI-based solutions which would modernise and streamline the clinical trial process.ReferencesTaylor K, Properzi F, Cruz M, Ronte H, Haughey J. Intelligent clinical trials [Internet]. www2.deloitte.com. 2020 [cited 16 March 2022]. Available from: https://www2.deloitte.com/content/dam/insights/us/articles/22934_intelligent-clinical-trials/DI_Intelligent-clinical-trials.pdfGlass L, Shorter G, Patil R. AI IN CLINICAL DEVELOPMENT [Internet]. IQVIA. 2019 [cited 16 March 2022]. Available from: https://www.iqvia.com/-/media/iqvia/pdfs/library/white-papers/ai-in-clinical-development.pdfBhatt A. Artificial intelligence in managing clinical trial design and conduct: Man and machine still on the learning curve?. Perspectives in Clinical Research. 2021;12(1):1-3. Available from: https://doi.org/10.4103/picr.PICR_312_20Thorlund K, Dron L, Park JJH, Mills EJ. Synthetic and External Controls in Clinical Trials – A Primer for Researchers. Clinical Epidemiology. 2020;12:457-467. https://doi.org/10.2147/CLEP.S242097Groth SW. Honorarium or coercion: use of incentives for participants in clinical research. The Journal of the New York State Nurses’ Association. 2010 Spring-Summer;41(1):11-22. Available from: https://www.ncbi.nlm.nih.gov/pmc/articles/pmc3646546/Richens J, Lee C, Johri S. Improving the accuracy of medical diagnosis with causal machine learning. Nature Communications. 2020;11(3923). Available from: https://doi.org/10.1038/s41467-020-17419-7Warudkar H. AI For Data Cleaning: How AI can Clean Your Data and Save Your Man Hours and Money [Internet]. Express Analytics. 2019 [cited 16 March 2022]. Available from: https://expressanalytics.com/blog/ai-data-cleaning/
- Clinical trial design is an important aspect of interventional trials that serves to optimise, ergonomise and economise the clinical trial conduct. The goals of a clinical trial, whether medtech or pharma, can encompass assessment of safety, dosage optimisation, evaluation of efficacy or accuracy and comparison to existing treatments or diagnostics. This of course varies with the phase of the trial. For phase III or IV trials the goal is most often to determine superiority, non-inferiority, or equivalence of the novel therapeutic or device to one in standard use. A well-conducted study that achieves regulatory approval for the asset in an efficient way depends upon the design that informs it. An optimal design, from a statistical and data collection perspective, ensures accurate evaluation efficacy and safety, as well as getting the product to market sooner. Knowing which study designs best suit your research will improve the chances of success, enable the best method for sample size estimation and re-estimation, save time and reduce unnecessary costs (Evans, 2010). While many clinical study designs exist this article focuses on perhaps the most rudimentary and frequently used designsParallel group designCrossover designFactorial designRandomised withdrawal design1. Parallel group study designA commonly used study design is a parallel arm design. When using this as a study design, subjects are randomised and allocated to one or more study arms. In a parallel group study design, each study arm is allocated a different intervention. After study subjects have been randomised and allocated to a study arms they can not be allocated to another arm throughout the study.Advantages of parallel group trial study designA key advantage of parallel group trial design is that it can be applied to many different diseases as well as allows for conducting multiple experiments simultaneously between many groups. A further advantage is that these different groups need not be sourced from the same site.Note: Once patients have been randomised and assigned to a specific arm, these arms are mutually exclusive. This means that unplanned co- interventions or cross-overs between different treatments cannot be introduced.Steps involved in a parallel arm trial design:1. Eligibility of study subject assessed2. Recruitment into study after consent 3. Randomisation4. Allocation to either treatment or control arm 2. Cross-over study designThere are some ethical limitations to the use of placebo controls that can be partially overcome by using a cross over design. This means that every patient taking part in the clinical trial will receive both treatment and placebo being given in a randomised order (Evans, 2010). Cross-over study design can also be used in the absence of placebo where the intention is to compare the new treatment to the standard one.Advantages of cross-over designOne of the advantages of cross over design is the fact that each patient acts as their own control results in order to balance the covariates in treatment and control arm. Another major advantage of cross over design is the fact that it requires a smaller sample size (Nair, 2019).Note: When cross over design is applicable and chosen for the study, some of the patients will start the trial with using intervention A and then switch to intervention B which is known as a AB sequence, whereas other patients will start with using intervention B and later switch to intervention A which is known as BA sequence.! There needs to be an adequate washout period before the crossover in order to eliminate the effects from initially assigned and administrated intervention. After all data has been collected the results are then compared within the same subject assessing the effect of intervention A vs. effect of intervention B (Nair, 2019).Variations of cross-over design(i) Switch back design (ABA vs BAB arms) –1. Drug A -> Drug B-> Drug A2.Drug B -> Drug A -> Drug BThe switch back and multiple switchback designs are of emerging relevance with the advent of biosimilars where switchability and interchangeability of a biosimilar to a bio-originator molecule can only be confirmed by such trial designs.(ii) N of 1 design – N of 1 trials or “single-subject” or “structured within-patient randomized controlled multi-crossover trial design”This type of cross over design is used for evaluating all interventions in a single patient. A typical N of 1 design clinical trial consists of repeating experimental/ control treatment periods number of times. The interventions being tested are assigned randomly within each period pair. This design has gained a lot of popularity, because in most cases the aim of using this type of design is to determine which treatment works best for the individual patient.3. Factorial designFactorial design is most suited when the study is looking at two or more interventions in various combinations within one study setting. This design helps in the study of interactive effects that have resulted from a combination of different interventions (Nair, 2019).Advantages of Factorial designA key advantage of factorial design is that it can help answer multiple research questions in a clinical trial instead of conducting multiple trials. This helps to optimise resources, thereby reducing costs and speeding up research pipelines.2 × 2 factorial design with placeboIn a 2 × 2 factorial design with placebo, patients are randomized into four groups:i) treatment A plus placebo ii) treatment B plus placebo iii) both treatments A and B iv) neither of them, placebo only.Limitations of the factorial designThe main limitations of using factorial design for clinical trials is the fact that:○ Increased complexity of the trial overall○ Makes it more difficult to meet inclusion criteria○ Inability to combine multiple incompatible interventions○ The protocols are complex○ High complexity of statistical analysis4. Randomised withdrawal design (EERW)The aim of randomised withdrawal design is to evaluate the optimal duration of the treatment for patients that are responsive to the intervention. After the initial enrichment period (open label period) which main purpose is to assign the subjects to intervention, the subjects that are not responding are removed (dropped) from the study and the subjects that did respond are randomised into receiving the intervention or placebo during the second phase of the clinical trial (Nair, 2019).Note: This means that only subjects that have responded are carried forward to the second stage of the study and randomised.Statistical analysis of randomised withdrawal designWhen using randomised withdrawal design the analysis of the study is conducted using only data from the withdrawal phase. Outcome is usually set to relapse of symptoms. The aim of the enrichment phase is to increase the statistical power for the estimated sample size.Advantages of EERWA main advantage of a randomised withdrawal design is that it can reduce the time patients receive placebo. Only patients that are responsive to the intervention are randomised to placebo, hence an increased ethical advantage. A further advantage of this study design is that it can help to determine if the treatment should be stopped or continued (Nair,2019).ConclusionOne of the key stages of planning a clinical trial involves deciding on the appropriate study design to ensure the success of the research and help to choose the right method for sample size estimation and re-estimation, save time and reduce unnecessary costs.The most commonly used study designs are :Parallel group study designCross over study designFactorial study designRandomised withdrawal study design (EERW )A well-conducted study with optimal design, that encorporates a robust hypothesis evolved from clinical practice, goes a long way in facilitating the regulatory approval process – evaluating efficacy and safety, and getting the product to market. Therefore when undertaking a clinical trial close attention should be paid to ensure that a study design forms a solid foundation upon which to conduct the trial phases.References Evans, S., 2010. Fundamentals of clinical trial design. [online] PubMed Central (PMC). Available at: <https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3083073/>. Expert, T., 2022. Clinical Trial Designs & Clinical Trial Phases | Credevo Articles. [online] Credevo Articles. Available at: <https://credevo.com/articles/2021/02/05/the-phase-of- study-clinical-trial-design/>. Nair, B., 2019. Clinical Trial Designs. [online] PubMed Central (PMC). Available at: <https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6434767/>. The BMJ | The BMJ: leading general medical journal. Research. Education. Comment. n.d. 13. Study design and choosing a statistical test | The BMJ. [online] Available at: <https://www.bmj.com/about-bmj/resources-readers/publications/statistics-square-one/ 13-study-design-and-choosing-statisti>.
- Accurate sample size calculation plays an important role in clinical research. Sample size in this context simply refers to the number of human patients, wheather healthy or diseased, taking part in the study. Clinical studies conducted using an insufficient sample size can lack the statistical power to adequately evaluate the treatment of interest, whereas a superfluous sample size can unnecessarily waste limited resources. Various methods can be applied for determining the optimal sample size for a specific clinical study. Methods also exist for any re-adjustments throughout the study, if required. These methods vary widely from straightforward tests and formulas to complex, time-consuming ones, depending on the type of study and available information from which to make the estimate. Most commonly used sample size calculation procedures are developed from a frequentist perspectiveImportance of knowing your study parameters Accurate sample size calculation requires, information on several key study and research parameters. These parameters usually include an effect size and variability estimate, derived from available sources; a clinically meaningful difference. In practice these parameters are generally unknown and must be estimated from the existing literature or from pilot studies.The Bayesian Framework in sample size estimations and re-adjustmentsThe Bayesian Framework has gradually become one of the most frequently mentioned methods when it comes to randomised clinical trial sample size estimations and re-adjustments.In practice, sample size calculation is usually treated explicitly as a decision problem and employs a loss or utility function.The Bayesian approach involves three key stages:1. Prior estimateA researcher has a prior estimate about the treatment effect (and other study parameters) that has been derived from meta-analysis of existing research, pilot studies, or based on expert opinion in absence of these.2. LikelihoodData is simulated to derive a likelihood estimate of prior parameters.3. Posterior estimateBased on the insights obtained, prior estimates from the first stage are updated to give a more precise final estimate.A challenge of using this approach is knowing when to stop this cycle when enough evidence has been gathered and avoid creating bias (Dreibe,2021). Peaking at the data in order to make a stopping decision is called “optional stopping”. In general an optional stopping rule is cautioned against as it can increase type one error rates (de Heide & Grunewald, 2021).How to decide when to stop the simulation cycle?There are two approaches one could take.1. Posterior probability Calculating the posterior probability that the mean difference between the treatment and control arm is equal or greater than the estimated effect of the intervention. Based on the level of probably calculated (low or high) the cycle could be stopped and without any further need to gather more data.2. PPOS ( predictive probability of success) Calculating the predictive probability of achieving a successful result at the end of the study is a commonly used approach. It is really helpful when it comes to determining the success or failure of a study. Similarly, as with posterior probability based on the level of probability a decision could be made to stop or continue the study.How to plan a Bayesian sample size calculation for a clinical trialThe key elements to consider when planning a Bayesian clinical trial are the same as for frequentists clinical trial.Key planning stages:Determine the objective of the clinical studyDetermine and set endpointsDecide on the appropriate study designRun a meta analysis or review of existing evidence related to your research objectiveStatistical test and statistical analysis plan (SAP)Even though the key planning stages are the same for both approaches it does not mean that they can be mixed through out the study. If you have chosen to use one approach you can’t change to another method once the calculations have been generated and research started.Bayesian approach vs Frequentist approach for sample size calculationsBayesianFrequentistPrior and posterior( uses probability of hypothesis and data)No prior or posterior( never gives probability of hypothesis)Sample size depends on the prior and likelihoodSample size depend on the likelihoodRequeres finding/deciding on prior in order to estimate sample sizeDoes not require prior to estimate sample sizeComputationally intensive due to integration over many parametersLess computationally intense Frequentist measures such as p-values and confidence intervals continue to predominate the methodology across life sciences research, however, the use of the Bayesian approach in sample size estimations and re-estimation for RTCs has been increasing over time.Bayesian approach for sample size calculations in medical device clinical trial In the recent years Bayesian approach has gained more popularity as the method used in clinical trials including medical device studies. One of the reasons being that if good prior information about the use of the specific therapeutic or device is available, the Bayesian approach may allow to include this information into the statistical analysis part of the clinical trial. Sometimes, the available prior information for a device of interest may be used as a justification for smaller sample size and shorten the length of the pivotal trial (Chen et al., 2011).Computational algorithms and growing popularity of Bayesian approach Bayesian statistical analysis can be computationally intense. Despite that there have been multiple breakthroughs with computational algorithms and increased computing speed that have made it much easier to calculate and build more realistic Bayesian models, further contributing to the popularity of Bayesian approach. (FDA, 2010).Markov Chain Monte Carlo (MCMC) method One of the basic computational tools being used is Markov Chain Monte Carlo ( MCMC) method. This method computes large number of simulations from the distributions of random quantities.Why MCMC? MCMC helps to deal with computational difficulties one often can face when using Bayesian approach for needed sample size estimations. The MCMC is an advanced random variable generation technique which allows one to simulate different samples from more sophisticated probability distributions.Conclusion Sample size calculation plays an important role in clinical research. If underestimated, statistical power for the detection of a clinically meaningful difference will likely be insufficient; if overestimated, resources are wasted unnecessarilly. The Bayesian Framework has become quite popular approach for sample size estimation. There are advantages of using the Bayesian method, depite this there has been some criticism of this approach as a sample size estimation and re-adjustment method due to the prior being subjective and possibility of different researchers selecting different priors leading to different posteriors and final conclusions.In reality, both the Bayesian and frequentist approaches to sample size calculation involve deriving the relevant input parameters from the literature or clinical expertise and could potentially differ due to variations in individual expert opinion as to which studies to include or exclude in this process. Bayesian approach is more computationally intensive compared to the traditional frequentist approaches. Therefore, when it comes to selecting a method for sample size estimation, it should be chosen carefully to best fit the particular study design and base-on advice provided by statistical professionals with expertise in clinical trials.References:Bokai WANG, C., 2017. Comparisons of Superiority, Non-inferiority, and Equivalence Trials. [online] PubMed Central (PMC). Available at: <https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5925592/> [Accessed 28 February 2022].Chen, M., Ibrahim, J., Lam, P., Yu, A. and Zhang, Y., 2011. Bayesian Design of Noninferiority Trials for Medical Devices Using Historical Data. Biometrics, 67(3), pp.1163-1170.E, L., 2008. Superiority, equivalence, and non-inferiority trials. [online] PubMed. Available at: <https://pubmed.ncbi.nlm.nih.gov/18537788/> [Accessed 28 February 2022].Gubbiotti, S., 2008. Bayesian Methods for Sample Size Determination and their use in Clinical Trials. [online] Core.ac.uk. Available at: <https://core.ac.uk/download/pdf/74322247.pdf> [Accessed 28 February 2022].U.S. Food and Drug Administration. 2010. Guidance for the Use of Bayesian Statistics in Medical Device Clinical. [online] Available at: <https://www.fda.gov/regulatory-information/search-fda-guidance-documents/guidance-use-bayesian-statistics-medical-device-clinical-trials> [Accessed 28 February 2022].van Ravenzwaaij, D., Monden, R., Tendeiro, J. and Ioannidis, J., 2019. Bayes factors for superiority, non-inferiority, and equivalence designs. BMC Medical Research Methodology, 19(1).de Heide. R, Grunewald, P.D, 2021, Why optional stopping can be a problem for Bayesians; Psychonomic Bulletin & Review, 21(2), 201-208.
- The development of new drugs starts far before they are even seen in clinical trials. The discovery of multiple candidate drugs occur early on in the development process, often as a result of new information about how a disease functions, large-scale screening of small molecules, or the release of a new technology.After a promising drug has been found, pre-clinical studies can be performed. A pre-clinical study for a new drug is used to determine important information about toxicity and suitable dosage amounts. These studies can be in vitro (in cell culture) and/or in vivo (in animal models) and determine whether a treatment will continue to the clinical trials stage.Clinical trials test whether these experimental treatments are safe for use in humans, and whether they are more effective in treating or preventing a disease when compared to existing treatments. Clinical trials consist of several stages, called phases, where each phase is focused on answering a different clinical question: Progression of a treatment to the next phase requires the study to meet several parameters to ensure a treatment’s safety or efficacy.Phase 0: Is the new treatment safe to use in humans in small doses?Phase I: Is the new treatment safe to use in humans in therapeutic doses?Phase II: Is the new treatment effective in humans?Phase III: Is the new treatment more effective than existing treatments?Phase IV: Does the new treatment remain safe and effective post-market?Key phases of a pharmaceutical clinical trialPhase 0: Small dose safetyPhase 0 studies can help to streamline the other clinical trial phases. Phase 0 consists of giving a few patients small, sub-therapeutic doses of the new treatment. This is to make sure that the new treatment behaves as expected by researchers and isn’t harmful to humans prior to using higher doses in phase I trials.Phase I: Therapeutic dose safetyPhase I studies evaluate the safety of various doses of the new treatment in humans. This takes several months with typically around 20-80 healthy volunteers. In some cases, such as in anti-cancer drug trials, the study participants are patients with the targeted cancer type. A treatment may not pass phase I if the treatment leads to any serious adverse events.Initial dosages in phase I studies can be informed based on data obtained during pre-clinical animal studies, and adjustments can be made to investigate the treatment’s side effect profile and develop an optimal dosing program. This could also include comparing different methods of giving a drug to patients (e.g., oral, intravenous etc.).Phase II: Treatment efficacyAfter passing phase I trials and having proven safety in humans, a new treatment advances to phase II studies designed to assess whether it may prevent or treat a disease. This phase can take between several months to 2 years, testing the new treatment in up to several hundred patients with the disease. Using a larger number of patients over a longer time period provides researchers with additional safety and effectiveness data, which is essential for the design of phase III trials.To further test safety and efficacy, it is common to have a control group that receives either a placebo (a harmless pill or injection without the new treatment) or other current treatment (in trials where the disease is fatal unless treated e.g., cancer).Phase III: Comparing to current treatmentsPhase III studies are the last stage of a clinical trial before a new treatment can be approved for market use. The primary focus of a phase III study is to compare the safety and efficacy of a new treatment with current, existing treatments in patients with the target disease. Anywhere from several hundred to 3,000 patients may be included in a phase III study for between 1 to 4 years. Due to the scale of this phase, long-term or rare side effects are more likely to be uncovered.Phase III studies are often randomised control trials, where patients will be randomly designated to different treatment groups. These groups may receive placebo, a current treatment (control group), the new treatment, or variations of the new treatment (e.g., different drug combinations). Randomised control trials are often double-blinded, where both the patient and the clinician administering their treatment do not know which treatment group they are assigned to.A new treatment may continue to market and phase IV trials if the results prove it is as safe and effective as an existing treatment.Phase IV: Post-market surveillanceIf a new treatment passes phase III and is approved by the MHRA, FDA, or other national regulatory agency, it can be put to market. Phase IV is carried out in the post-market surveillance of the new treatment to keep updated on any emerging or long-term safety and efficacy concerns. This may include rare or long-term adverse side effects that were not yet discovered, or long-term analyses to see if the new treatment improves the life expectancy of a patient after recovery from disease.SummaryClinical trials are ultimately designed to mitigate risk. This includes the risk to the safety of trial participants by limiting the use of potentially unsafe treatments to small doses in a small number of patients before scaling up to testing therapeutic dose safety. Risk mitigation is not only for patient safety but also for preventing financial misspending as a treatment that is deemed unsafe in phase 0 would not proceed to the later, more costly clinical trial phases.Not all clinical trials are the same, however, as each trial will have a different disease and treatment context. Trials for medical devices are somewhat different from pharmaceutical trials (for more information about the differences between medical device and pharma trials, click here). In addition, while sample sizes expand with phase progression, the required sample size for each trial and each phase is dependent on several factors including disease context (a rare disease may require lower sample sizes), patient availability (location of trial), trial budget and effect size. The sample size values mentioned earlier in this blog are purely indications of what each phase may use (for more information on how a biostatistician determines a suitable sample size, click here).Referenceshttps://www.fda.gov/patients/drug-development-process/step-3-clinical-researchhttps://www.healthline.com/health/clinical-trial-phases
- Medical devices and drugs share the same goal – to safely improve the health of patients. Despite this, substantial differences can be observed between the two. Principally, drugs interact with biochemical pathways in human bodies while medical devices can encompass a wide range of different actions and reactions, for example, heat, radiation (Taylor and Iglesias, 2009). Additionally, medical devices encompass not only therapeutic devices but diagnostic devices, as well (Stauffer, 2020).More specifically medical device categories can include therapeutic and surgical devices, patient monitoring, diagnostic and medical imaging devices, among others; making it a very heterogeneous area (Stauffer, 2020). As such, medical device research spills over into many different fields of healthcare services and manufacturing. This research is mostly undertaken by SME’s ( small to medium enterprises) instead of larger well-established companies as is more predominantly the case with pharmaceutical research. SME’s and start-ups undertake the majority of the early stage device development, particularly where any new class of medical device is concerned, whereas the larger firms get involved in later stages of the testing process (Taylor and Iglesias, 2009).Classification criteria for medical devicesThere are strict regulations that researchers and developers need to follow, which includes general device classification criteria. This classification criterion consists of three classes of medical devices, the higher class medical device the stricter regulatory controls are for the medical device. Class I, typically do not require premarket notificationsClass II, require premarket notificationsClass III, require premarket approvalFood and Drug Administration (FDA)Drug licensing and market access approval by the Food and Drug Administration (FDA) and international equivalents require manufacturers to undertake phase II and III randomised controlled trials in order to provide the regulator with evidence of their drug’s efficacy and safety (Taylor and Iglesias, 2009).Key stages of medical device clinical trialsIn general medical device clinical trials are smaller than drug trials and usually start with feasibility study. This provides a limited clinical evaluation of the device. Next a pivotal trial is conducted to demonstrate the device in question is safe and effective (Stauffer, 2020).Overall the medical device trials can be considered to have three stages:Feasibility study,Pivotal study to determine if the device is safe and effective,Post-market study to analyse the long-term effectiveness of the device.Clinical evaluation for medical devicesClinical evaluation is an ongoing process conducted throughout the life cycle of a medical device. It is first performed during the development of a medical device in order to identify data that need to be generated for regulatory purposes and will inform if a new device clinical investigation is necessary. It is then repeated periodically as new safety, clinical performance and/or effectiveness information about the medical device is obtained during its use.(International Medical Device Regulators Forum, 2019)During the evaluative process, a distinction must be made between device types – diagnostic or therapeutic. The criteria for diagnostic technology evaluations are usually divided into four groups:technical capacitydiagnostic accuracydiagnostic and therapeutic impactpatient outcomeThe importance of evaluationEvaluations provide important information about a device and can indicate the possible risks and complications. The main measures of diagnostic performance are sensitivity and specificity. Based on the results of the clinical investigation the intervention may be approved for the market. When placing a medical device on the market, the manufacturer must have demonstrated through the use of appropriate conformity assessment procedures that the medical device complies with the Essential Principles of Safety and Performance of Medical Devices(International Medical Device Regulators Forum, 2019).The information on effectiveness can be observed by conducting experimental or observational studies.Post-market surveillanceManufacturers are expected to implement and maintain surveillance programs that routinely monitor the safety, clinical performance and/or effectiveness of the medical device as part of their Quality Management System (International Medical Device Regulators Forum, 2019). The scope and nature of such post market surveillance should be appropriate to the medical device and its intended use. Using data generated from such programs (e.g. safety reports, including adverse event reports; results from published literature, any further clinical investigations), a manufacturer should periodically review performance, safety and the benefit-risk assessment for the medical device through a clinical evaluation, and update the clinical evidence accordingly.The use of databases in medical device clinical trialsThe variations in the available evidence-base for devices means that, unlike with drugs, medical devices will typically require the consideration and analysis of data from observational studies in ascertaining their clinical and cost-effectiveness. Using modern observational databases has advantages because these databases represent continuous monitoring of the device in real-life practice, including the outcome (Maresova et al., 2020).Bayesian methods as an alternative framework for evaluationBayesian methods for the analysis of trial data have been proposed as an alternative framework for evaluation within the FDA’s Center for Devices and Radiological Health. These methods provide flexibility and may make them particularly well suited to address many of the issues associated with the assessment of clinical and economic evidence on medical devices, for example, learning effects and lack of head-to-head comparisons between different devices.Use of placebo in medical vs pharmaceutical trialsAn additional key difference between drug and medical device trials are that use of placebo in medical device trials are rare. If placebo is used in a trial for surgical / implanted devices it would usually be a sham surgery or implantation of a sham device (Taylor and Iglesias, 2009). Sham procedures are high risk and may be considered unethical. Without this kind of control, however, there is in many cases no sure way of knowing whether the device is providing real clinical benefit or if the benefit experienced is due to the placebo effect. Conclusion In conclusion, there are many similarities between medical device and pharmaceutical clinical trials, but there are also some really important differences that one should not miss: In general medical device clinical trials are smaller than drug trials. The research is mostly undertaken by SME’s ( small to medium enterprises) instead of big well-known companiesDrugs interact with biochemical pathways in human bodies whereas medical devices use a wide range of different actions and reactions, for example, heat, radiation.Medical devices can be used for not only diagnostic purposes but therapeutical purposes as well. The use of placebo in medical device trials are rare in comparison to pharmaceutical clinical trials.References:Bokai WANG, C., 2017. Comparisons of Superiority, Non-inferiority, and Equivalence Trials. [online] PubMed Central (PMC). Available at: <https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5925592/> [Accessed 28 February 2022].Chen, M., Ibrahim, J., Lam, P., Yu, A. and Zhang, Y., 2011. Bayesian Design of Noninferiority Trials for Medical Devices Using Historical Data. Biometrics, 67(3), pp.1163-1170.E, L., 2008. Superiority, equivalence, and non-inferiority trials. [online] PubMed. Available at: <https://pubmed.ncbi.nlm.nih.gov/18537788/> [Accessed 28 February 2022].Gubbiotti, S., 2008. Bayesian Methods for Sample Size Determination and their use in Clinical Trials. [online] Core.ac.uk. Available at: <https://core.ac.uk/download/pdf/74322247.pdf> [Accessed 28 February 2022].U.S. Food and Drug Administration. 2010. Guidance for the Use of Bayesian Statistics in Medical Device Clinical. [online] Available at: <https://www.fda.gov/regulatory-information/search-fda-guidance-documents/guidance-use-bayesian-statistics-medical-device-clinical-trials> [Accessed 28 February 2022].van Ravenzwaaij, D., Monden, R., Tendeiro, J. and Ioannidis, J., 2019. Bayes factors for superiority, non-inferiority, and equivalence designs. BMC Medical Research Methodology, 19(1).
- Part 1: Basket & Umbrella Trial DesignsIntroductionAs the clinical research landscape becomes ever more complex and interdisciplinary alongside an evolving genomic and biomolecular understanding of disease, the statistical design component that underpins this research must adapt to accommodate this. Accuracy of evidence and the speed with which novel therapeutics are brought to market remain hurdles to be surmounted.Traditionally, efficacy and non-inferiority clinical trials included broad disease states, with patients randomised to a dual-arm design comparing a new treatment to an existing standard of care. However, due to patient biomarker heterogeneity, effective treatments could be left unsupported by evidence. Similarly, treatments found effective in a clinical trial don’t always translate to real-world effectiveness in a broader range of patients.Our current ability to assess individual genomic, proteomic and transcriptomic data, alongside other patient biomarkers, has shown that different patients respond differently to the same treatment. Furthermore, the same disease may benefit from different treatments in different patients—thus the beginnings of precision medicine. In addition to this is the scenario where a single therapeutic may be effective against a number of different diseases or subclasses of a disease, based on the agent’s mechanism of action on molecular processes common to the disease states under evaluation.Master protocols, or complex innovative designs, are designed to pool resources to avoid redundancy and test multiple hypotheses under one clinical trial, rather than multiple clinical trials being carried out separately over a longer period of time. Regulatory agencies such as the FDA and EMA have increasingly embraced these designs, particularly following their success in oncology and rapid-response COVID-19 trials.While basket and umbrella trials are two prominent examples of master protocols, they are often discussed alongside platform trials—a third type of master protocol which tests multiple treatments in a single disease over a non-fixed time frame. (Platform trials will be covered in detail in Part 2 of this series).Due to this fairly novel evolution in the clinical research paradigm, and the inherent flexibility within each study design, conflicting information related to the definition and characterisation of master protocols exists. In the published literature, the terms “basket” and “umbrella” trials have sometimes been used interchangeably or are ill-defined. For this reason, a brief definition and overview of basket and umbrella clinical trials is included in the paragraphs that follow. Based on systematic reviews of existing research, it seeks the clarity of consensus before detailing some key statistical and operational elements of each design.Diagram of a basket trial design.Basket trial:A basket clinical trial design consists of a targeted therapy, such as a drug or treatment device, that is being tested on multiple disease states characterised by a common molecular process that is impacted by the treatment’s mechanism of action. These disease states could also share a common genetic or proteomic alteration that researchers are looking to target.Basket trials can be either exploratory or confirmatory and range from full randomised, controlled double-blind designs to single-arm designs, or anything in between. Single-arm designs are an option when feasibility is limited and are more focused on the pre-clinical stage of determining efficacy, or whether a particular treatment has clear-cut commercial potential evidenced by a sizable enough retreat in disease symptomology. Depending on the nuances of the patient populations being evaluated, final study data may be analysed by pooling disease states or by each disease state separately. Basket trials allow drug development companies to target the lowest hanging fruit in terms of treatment efficacy, focusing resources on therapeutics with the highest potential of success in terms of real patient outcomes.Diagram of an umbrella trial design.Umbrella trial:An umbrella clinical trial design consists of multiple targeted treatments for a single disease, where patients can be sub-categorised into biomarker subgroups defined by molecular characteristics that may lend themselves to one treatment over another.Umbrella trials can be randomised, controlled, double-blind studies in which each intervention and control pair is analysed independently of other treatments in the trial. Alternatively, where feasibility issues dictate, they can be conducted without a control group, with results analysed together in order to compare the different treatments directly.Umbrella trials may be useful when a treatment has shown efficacy in some patients and not others. They increase the potential for confirmatory trial success by honing in on patient sub-populations that are most likely to benefit due to biomarker characteristics, rather than grouping all patients together as a whole.Basket & Umbrella trials compared:Both basket and umbrella trials are typically biomarker-guided. The difference being that basket trials aim to evaluate tissue-agnostic treatments across multiple diseases based on common molecular characteristics, whereas umbrella trials aim to evaluate nuanced treatment approaches to the same disease based on differing molecular characteristics between patients.Biomarker-guided trials have an additional feasibility constraint compared to non-biomarker-guided trials, in that the size of the eligible patient pool is reduced in proportion to the prevalence of the biomarker(s) of interest within that patient pool. This is why master protocol methodology becomes instrumental in enabling these appropriately complex research questions to be pursued.Statistical Concepts and considerations of basket and umbrella TrialsEffect sizeBasket and umbrella trials generally require a larger effect size than traditional clinical trials in order to achieve statistical significance. This is largely due to the smaller sample sizes and higher variance that comes with that. While patient heterogeneity in terms of genomic or molecular diversity (and thus expected treatment outcome) has been reduced by the precision targeting of the trial design, there is a certain degree of between-patient heterogeneity that can only be expected when relying on treatment arms of very small sample sizes.If resources, including time, are tight, then basket trials enable drug developers to focus on less risky treatments that are more likely to end in clinical and commercial viability. It should be noted that this does not always mean that treatments rejected by basket trials are truly clinically ineffective. A single-arm exploratory basket trial could end up rejecting a potential new treatment that, if subject to a standard trial with more drawn-out patient acquisition and a larger sample size, would have been deemed effective at a narrower effect size.Screening efficiencyIf researchers carry out separate clinical studies for each biomarker of interest, then a separate screening sample needs to be recruited for each study. The rarer the biomarker, the larger the recruited screening sample would need to be to find enough people with the biomarker to participate in the study. This number needs to be multiplied by the number of biomarkers. A benefit of master protocols is that a single sample of people can be screened for multiple biomarkers at once, greatly reducing the required screening sample size.For example, researchers interested in 4 different biomarkers could collectively reduce the required screening sample by three quarters compared to conducting separate clinical studies for each biomarker. This maximisation of resources can be particularly helpful when dealing with rare biomarkers or diseases.Patient allocation considerationsIf relevant biomarkers are not mutually exclusive, a patient could fit into multiple biomarker groups for which treatment is being assessed in the study. In this scenario, a decision has to be made as to which category the patient will be assigned, and the decision process may occur at random where appropriate. If belonging to two overlapping biomarker groups is problematic in terms of introducing bias in small sample sizes, or if several patients have the same overlap, then a decision may be made to collapse the two biomarkers into a single group or eliminate one of the groups. If a rare genetic mutation is a priority focus in the study, then feasibility would dictate that the patient be assigned to this biomarker group.Sample Size calculationsGenerally speaking, while umbrella trials calculate sample size individually for each treatment sub-study, basket trials often utilise statistical methods (such as Bayesian hierarchical models) that “borrow strength” from the overall cohort to inform estimates. However, each basket must still usually demonstrate adequate independent evidence for regulatory approval in that specific disease.Basket and umbrella trials can be useful in situations where a smaller sample size is more feasible due to specifics of the patient population under investigation. Statistically designing for this smaller sample size typically comes at the cost of necessitating a greater effect size (difference between treatment and control), and this translates to lower overall study power and a greater chance of a Type II error (false negative result) when compared to a standard clinical trial design. Despite these limitations, master protocols such as basket or umbrella trials allow the evaluation of certain treatments to the highest possible level of evidence that otherwise might be too heterogeneous or rare to evaluate using a traditional Phase II or III trial.Randomisation and controlRandomised controlled designs are recommended for confirmatory analysis of an established treatment or target of interest. The control group typically treats patients with the established standard of care for their particular disease or, in the absence of one, a placebo.In master basket trials, the established standard of care is likely to differ by disease or disease sub-type. For this reason, it may be necessary for randomised controlled basket trials to pair a control group with each disease sub-group, rather than just incorporating a single overall control group and potentially pooling results from all diseases under one statistical analysis of treatment success. Instead, it is worth considering if each disease type and corresponding control pair could be analysed separately to enhance statistical robustness in a truly randomised controlled methodology.Single-arm (non-randomised) designs are sometimes necessary for exploratory analysis of potential treatments or targets. These designs often require a greater margin of success (treatment efficacy) to be statistically significant as a trade-off for the smaller sample size required.BlindingTo increase the quality of evidence, all clinical studies should be double-blind where possible. To truly evaluate the effectiveness of a treatment without undue bias from a statistical perspective, double-blinding is recommended.Aside from the increased risk of Type II error that may be inherent in master protocol designs, there is a greater potential for statistical bias to be introduced. Bias can introduce itself in a myriad of ways and results in a reduction in the quality of evidence that a study can produce. Two key sources of bias are lack of randomisation (mentioned above) and lack of blinding.Single-arm trials do not include a control arm and therefore patients cannot be randomised. While traditional double-blinding is not possible in single-arm trials because everyone knows they are receiving the investigational drug, researchers should still utilise blinded independent reviewers for outcome assessment to mitigate evaluation bias. With so many factors at play, it is important not to overlook the importance of study blinding and to implement it whenever feasible to do so.If the priority is getting a new treatment or product to market fast to benefit patients and potentially save lives, accommodating this bias can be a necessary trade-off. It is, after all, typically quite a challenge to have clinical data and patient populations that are homogeneous and matched to any great degree, and this reality is especially noticeable with rare diseases or rare biomarkers.Biomarker Assay methodologyThe reliability of biological variables included in a clinical trial should be assessed; for example, the established sensitivity and specificity of particular assays needs to be taken into account. When considering patient allocation by biomarker group, the degree of potential inaccuracy of this allocation can have a significant impact on trial results, particularly when there is a small sample size. If the false positive rate of a biomarker assay is too high, this will result in the wrong patients qualifying for treatment arms, which may reduce the statistical power of the study.A further consideration of assay methodology pertains to the potential for non-uniform biospecimen quality at different collection sites, which may bias study results. A monitoring framework should be considered in order to mitigate this.Patient tissue samples required for assays can inhibit feasibility and increase time and cost in the short term, making study reproducibility more complicated. While this is important to note, these techniques are in many cases necessary in effectively assessing treatments based on our contemporary understanding of many disease states, such as cancer within the modern oncology paradigm. Without incorporating this level of complexity and personalisation into clinical research, it will not be possible to develop evidence-based treatments that translate into real-world effectiveness and widespread positive outcomes for patients.Data management and statistical analysisThe ability to statistically analyse multiple research hypotheses at once within a single dataset increases efficiency at the biostatisticians end and allows frameworks for greater reproducibility of the methodology and final results, compared to the execution and analysis of multiple separate clinical trials testing the same hypotheses. Master protocols also enable increased data sharing and collaboration between sites and stakeholders.Deloitte research estimated that master protocols can save clinical trials 12-15% in cost and 13-18% in study duration. These savings of course apply to situations where master protocols were a good fit for the clinical research context, rather than to the blanket application of these study designs across any or all clinical studies. Applying a master protocol study design to the wrong clinical study could actually end up increasing required resources and costs without benefit, therefore it is important to assess whether a master protocol study design is indeed the optimal approach for the goals of a particular clinical study or studies.Master protocols for precision medicine.Basket and umbrella trials represent a vital shift away from the “one-size-fits-all” approach to clinical research. By efficiently aligning molecular characteristics with targeted therapeutics, these master protocols accelerate the delivery of precision medicine to patients. Their success relies heavily on statistical planning, an understanding of biomarker assay limitations, and a careful balancing of operational trade-offs. In Part 2, we will explore the third pillar of master protocols: Platform Trials.References:Bitterman DS, Cagney DN, Singer LL, Nguyen PL, Catalano PJ, Mak RH. Master Protocol Trial Design for Efficient and Rational Evaluation of Novel Therapeutic Oncology Devices. J Natl Cancer Inst. 2020 Mar 1;112(3):229-237. doi: 10.1093/jnci/djz167. PMID: 31504680; PMCID: PMC7073911.Lesser N, Na B, Master protocols: Shifting the drug development paradigm, Deloitte Center for Health solutionsLai TL, Sklar M, Thomas, N, Novel clinical trial solutions and statistical methods in the era of precision medicine, Technical Report No. 2020-06, June 2020Renfro LA, Sargent DJ. Statistical controversies in clinical research: basket trials, umbrella trials, and other master protocols: a review and examples. Ann Oncol. 2017 Jan 1;28(1):34-43. doi: 10.1093/annonc/mdw413. PMID: 28177494; PMCID: PMC5834138.Park, J.J.H., Siden, E., Zoratti, M.J. et al. Systematic review of basket trials, umbrella trials, and platform trials: a landscape analysis of master protocols. Trials 20, 572 (2019). https://doi.org/10.1186/s13063-019-3664-1
- How Simpson’s Paradox Confounds Research Findings And Why Knowing Which Groups To Segment By Can Reverse Study Findings By Eliminating Bias.Introduction The misinterpretation of statistics or even the “mis”-analysis of data can occur for a variety of reasons and to a variety of ends. This article will focus on one such phenomenon contributing to the drawing of faulty conclusion from data: Simpson’s Paradox.At times a situation arises where the outcomes of a clinical research study depict the inverse of expected (or essentially correct) outcomes. Depending upon the statistical approach, this could affect means, proportions or relational trends among other statistics. Some examples of this occurrence are a negative difference when a positive difference was anticipated, a positive trend when a negative one would have been more intuitive, or vice versa. Another example commonly pertains to the cross tabulation of proportions, where condition A is proportionally greater over all, yet when stratified by a third variable, condition B is greater in all cases . All of these examples can be said to be instances of Simpson’s paradox. Essentially Simpson’s paradox represents the possibility of supporting opposing hypotheses with the same data. Simpson’s paradox can be said to occur due to the effects of confounding, where a confounding variable is characterised by being related to both the independent variable and the outcome variable, and unevenly distributed across levels of the independent variable. Simpson’s paradox can also occur without confounding in the context of non-collapsability. For more information on the nuances of confounding versus non-collapsability in the context of Simpson’s paradox, see here. In a sense, Simpson’s paradox is merely an apparent paradox, and can be more accurately described as a form of bias. This bias most often results from a lack of insight into how an unknown lurking variable, so to speak, is impacting upon the relationship between two variables of interest. Simpson’s paradox highlights the fact that taking data at face value and utilising it to inform clinical decision making can often be highly misleading. The chances of Simpson’s paradox (or bias) impacting the statistical analysis can be greatly reduced in many cases by a careful approach that has been informed by proper knowledge of the subject matter. This highlights the benefit of close collaboration between researcher and statistician in informing an optimal statistical methodology that can be adapted on a per case basis. The following three part series explores hypothetical clinical research scenarios in which Simpson’s paradox can manifest.Part 1 Simpson’s Paradox in correlation and linear regression Scenario and Example A nutritionist would like to investigate the relationships between diet and negative health outcomes. As higher weight has been previously associated with negative health outcomes, the research sets out to investigate the extent to which increased caloric intake contributes to weight gain. In researching the relationship between calorie intake and weight gain for a particular dietary regime, the nutritionist uncovers a rather unanticipated negative trend. As caloric intake increases the weight of participants appears to go down. The nutritionist therefore starts recommending higher calorie intake as a way to dramatically lose weight. Weight does appear to go down with calorie intake, however if we stratify the data by different age groupings, a positive trend between weight and calorie intake emerges for each age group. While overall elderly have the lowest calorie intake, they also have the highest weight, and teens have the highest calorie intake but the lowest weight, this accounts for the negative trend but does not give an honest picture of the impact of calories on weight. In order to gain an accurate picture of the relationship between weight and calorie intake we have to know which variable to group or stratify the data by, and in this case it’s age. Once the data is stratified by five separate age categories a positive trend between calories and weight emerges in each of the 5 categories. In general, the answer to which variable to stratify by or control for isn’t typically this obvious and in most cases and requires some theoretical background and a thorough examination of the available data including associated variables for which the information is at hand. Remedy In the above example, age shows a negative relationship to the independent variable, calories, but a positive relationship to the dependent variable, weight. It is for this reason that a bit of data exploration and assumption checking before any hypothesis testing is so essential. Even with these practices in place it is possible to overlook the source of confounding and caution is always encouraged. Randomisation and Stratification: In the context of a randomised controlled trial (RCT), the data should be randomly assigned to treatment groups as well as stratified by any pertinent demographic and other factors so that these are evenly distributed across treatment arms (levels of the independent variable). This approach can help to minimise, although not eliminate the chances of bias occurring in any such statistical context, predictive modelling or otherwise. Linear Structural Equation Modelling: If the data at hand is not randomised but observational, a different approach should be taken to detect causal effects in light of potential confounding or non-collapsability. One such approach is linear structural equation modelling where each variable is generated as a linear function of it’s parents, using a directed acyclic graph (DAG) with weighted edges. This is a more sophisticated and ideal approach to simply adjusting for x number of variables, which is needed in the absence of a randomisation protocol. Heirarchical regression: This example illustrated an apparent negative trend of the overall data masking a positive trend In each individual subgroup, in practice, the reverse can also occur. In order to avoid drawing misguided conclusion from the data the correct statistical approach must be entertained, a hierarchical regression controlling for a number of potential confounding factors could avoid drawing wrong conclusion due to Simpson’s paradox. Reference: The Simpson’s paradox unraveled, Hernan, M, Clayton, D, Keiding, N., International Journal of Epidemiology, 2011.Part 2 Simpson’s Paradox in 2 x 2 tables and proportions Scenario and ExampleSimpson’s paradox can manifest itself in the analysis of proportional data and two by two tables. In the following example two pharmaceutical cancer treatments are compared by a drug company utilising a randomised controlled clinical trial design. The company wants to test how the new drug (A) compares to the standard drug (B) already widely in clinical use. 1000 patients were randomly allocated to each group. A chi squared test of remission rates between the two drug treatments is highly statistically significant, indicating that the new drug A is the more effective choice. At first glance this seems reasonable, the sample size is fairly large and equal number of patients have been allocated to each groups.Drug TreatmentABRemission Yes798 (79.8%)705 (70.5%)Remission No202295Total sample size10001000The chi-square statistic for the difference in remission rates between treatment groups is 23.1569. The p-value is < .00001. The result is significant at p < .05. When we take a closer look, the picture changes. It turns out the clinical trial team forgot to take into account the patients stage of disease progression at the commencement of treatment. The table below shown that drug A was allocated to far more patients with stage II cancer (79.2%) and drug B was allocated to far more patients with stage IV cancer (79.8%). Stage IIStage IVDrug TreatmentABABRemission Yes697 (87.1%)195 (92.9%)101 (50.5%)510 (64.6%)Remission No1031599280Total sample size800210200790The chi-square statistic for the difference in remission rates between treatment groups for patients with stage II disease progression at treatment outset is 5.2969. The p-value is .021364. The result is significant at p < .05. The chi-square statistic for the difference in remission rates between treatment groups for patients with stage IV disease progression at treatment outset is 13.3473. The p-value is .000259. The result is significant at p < .05. Unfortunately the analysis of tabulated data is no less prone to bias in results akin to Simpson’s Paradox than continuous data. Given that stage II cancer is easier to treat than stage IV, this has given drug A an unfair advantage and has naturally lead to a higher remission rate overall for drug A. When the treatment groups are divided by disease progression categories and reanalysed, we can see that remission rates are higher for drug B in both stage II and stage IV baseline disease progression. The resulting chi squared statistics are wildly different to the first and statistically significant in the opposite direction to the first analysis. In causal terms, stage of disease progression affects difficulty of treatment and likelihood of remission. Patients at a more advanced stage of disease, ie stage IV, will be harder to treat than patients at stage II. In order for a fair comparison between two treatments, patients stage of disease progression needs to be taken into account. In addition to this some drugs may be more efficacious at one stage or the other, independent of the overall probabilities of achieving remission at either stage. Remedy Randomisation and Stratification: Of course in this scenario, stage of disease progression is not the only variable that needs to be accounted for in order to insure against biased results. Demographic variables such as age, sex socio-economic status and geographic location are some examples of variables that should be controlled for in any similar analysis. As with the scenario in part 1, this can be achieved is through stratified random allocation of patients to treatment groups at the outset of the study. Using a randomised controlled trial design where subjects are randomly allocated to each treatment group as well as stratified by pertinent demographic and diagnostic variables will reduce the chances of inaccurate study results occurring due to bias. Further examples of Simpson’s Paradox in 2 x 2 tables Simpson’s paradox in case control and cohort studies Case control and cohort studies also involve analyses which rely on the 2×2 table. The calculation of their corresponding measures of association the odds ratio and relative risk, respectively, is unsurprisingly not immune to the effect of bias and in much the same way as the chi square example above. This time, a reversed odds ratio or relative risk in the opposite direction can occur if the pertinent bias has not been accounted and controlled for.Simpson’s paradox in meta-analysis of case control studies Following on from the example above, this form of bias can pose further problems in the context of meta-analysis. When combining results from numerous case control studies the confounders in question may or may not have been identified or controlled for consistently across all studies and some studies will likely have identified different confounders to the same variable of interest. The odds ratios produced by the different studies can therefore be incompatible and lead to erroneous conclusions. Meta-analysis can therefore fall prey to ecological fallacy as a result of systematic bias, where the odds ratio for the combined studies is in the opposite direction to the odds ratios of the separate studies. Imbalance in treatment arm size has also been found to act as a confounder in the context of meta-analysis of randomised controlled trials. Other methodological differences between studies may also be at play, such as differences in follow-up times between studies or a very low proportion of observed events occurring in some studies, potentially due to a shorted follow-up time. That’s not to say that meta-analysis cannot be performed on these studies, inter-study variation is of-course more common than not, as with all other analytical contexts it is necessary to proceed with a high level of caution and attention to detail. On the whole an approach of simply pooling study results is not reliable, the use of more sophisticated meta-analytic techniques, such as random effects models or Bayesian random effects models that use a Markov chain algorithm for estimating the posterior distributions, are required to mitigate inherent limitations of the meta-analytic approach. Random-effects models assume the presence of study-specific variance which is a latent variable to be partitioned. Bayesian random-effects models can come in parametric, non-parametric or semi-parametric varieties, referring to the shape of the distributions of study-specific effects. For more information on Simpson’s paradox in meta-analysis, see here. https://bmcmedresmethodol.biomedcentral.com/articles/10.1186/1471-2288-8-34 For more information on how to minimise bias in meta-analysis, see here. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3868184/ https://onlinelibrary.wiley.com/doi/abs/10.1002/sim.4780110202 Part 3 Simpson’s Paradox & Cox Proportional Hazard Models Time to event data is common in clinical science and epidemiology, particularly in the context of survival analysis. Unfortunately the calculation of hazard rate in survival analysis is not immune to Simpson’s Paradox as the mathematics behind Simpson’s paradox is essentially the mathematics of conditional probability. In-fact Simpson’s paradox in this context has the interesting characteristic of holding for some intervals of the time variable (failure time T) but not others. In this case Simpson’s paradox would be observed across the effect of variable Y on the relationship between variable X and time interval T. The proportional hazards model can be seen as an extension of 2 by 2 tables, given that the type of data is similar is used, the difference being that time is typically is as much an outcome of interest in relation to some factor Y. In this context Y could be said to be a covariate to X.Scenario and example A 2017 paper describes a scenario whereby the death rate due to tuberculosis was lower in Richmond than New York for both African-Americans and for Caucasian-Americans, yet lower in New York than Richmond when the two ethnic groups were combined. For more details on this example as well as the mathematics behind it see here. For more examples of Simpson’s paradox in Cox regression see here.Site specific bias Factors contributing to bias in survival models can be different to those in more straightforward contexts. Many clinical and epidemiological studies include data from multiple sites. More often than not there is heterogeneity across sites. This heterogeneity can come is various forms and can result in within and between–site clustering, or correlation, of observations on site specific variables. This clustering, if not controlled for, can lead to Simpson’s paradox in the form of hazard rate reversal, across some or all of time T, and has been found to be a common explanation of the phenomenon in this context. Site clustering can occur on the patient level, for example, due to site specific selection procedures for the recruitment of patients (lead by the principal investigators individual to each site), or differences in site specific treatment protocols. Site specific differences can occur intra or internationally and in the international case can be due, for example, to differences in national treatment guidelines or differences in drug availability between countries. Resource availability can also differ between sites whether intra or internationally. In any time to event analysis involving multiple sites (such as the Cox regression model) a site-level effect should be taken into account and controlled for in order to avoid bias-related inferential errors. Remedy Cox regression Model including site as a fixed covariate: Site should be included as a covariate in order to account for site specific dependence of observations. Cox regression Model treating site as a stratification variable: In cases where one or more covariates violate the Proportional Hazards (PH) assumption as indicated by a lack of independence of scaled Schonefeld residuals to time, stratification may be more appropriate. Another option in this case is to add a time-varying covariate to the model. The choice made in this regard will depend on the sampling nuances of each particular study. Cox shared frailty model: In specific conditions the Cox shared frailty model may be more appropriate. This approach involves treating subjects from the same site as having the same frailty and requires that each subjects is not clustered across more than one level two unit. While it is not appropriate for multi-membership multi-level data, it can be useful for more straight forward scenarios. In tailoring the approach to the specifics of the data, appropriate model adjustments should produce hazard ratios that more accurately estimate the true risk.
- Latent variable models are statistical tools used extensively throughout clinical and experimental research in medicine and the life sciences. Disciplines such as psychology and neuroscience routinely employ these models to answer complex research questions ranging from the impact of personality traits on workplace success (1) to measuring inter-correlated activity in neural populations based on neuroimaging data (2). Through latent variable modelling, dispositions, states, or processes that must be inferred rather than directly measured can be linked causally to more concrete, observable measurements.While some latent variable approaches are exploratory, Structural Equation Modelling (SEM) and Confirmatory Factor Analysis (CFA) are primarily used to test a priori hypotheses. They are designed to evaluate the causal relationships between observable (manifest) variables and their corresponding latent variables within an inter-correlated dataset. A critical assumption of SEM is that the model is correctly specified. Even a minor misspecification can affect all parameter estimations in the model, rendering approximations inaccurate in unpredictable ways (3).Therefore, with any postulated statistical model, it is imperative to assess and validate model fit before interpreting the results. The most rigorous and widely accepted method for evaluating exact fit across structural equation models is the chi-squared (χ2) statistic.Interpreting the Chi-Squared (χ2) StatisticA statistically significant χ2 statistic is indicative of the following:Model Misspecification: The degree of misspecification is a function of the χ2 value. The set of parameters specified in the model does not adequately fit the data, meaning the parameter estimates are likely inaccurate.The Need for Investigation: Because χ2 operates on the same statistical principles as parameter estimation, trusting the parameter estimates requires trusting the χ2 test. A significant result necessitates investigating where the misspecification occurred and readjusting the model to improve accuracy.While incorrect hypotheses are a common cause, misspecification can also stem from other statistical violations. To properly diagnose a significant model fit test, researchers must evaluate the following assumptions:Heterogeneity: Does the causal model vary between subgroups of subjects? Are there intervening within-subject variables?Independence: Are the observations truly independent? Latent variable models assume that all manifest variables are independent after controlling for latent variables, and that an individual’s position on a manifest variable is solely the result of their position on the corresponding latent variable (3).Multivariate Normality: Is the assumption of multivariate normality satisfied?The 2015 Meta-Analysis: A Cautionary TaleA 2015 meta-analysis of 75 latent variable studies, drawn from 11 psychology journals, highlighted a troubling tendency among clinical researchers to ignore the χ2 exact fit statistic when reporting and interpreting their results (4).The findings revealed that 97% of papers reported at least one “appropriate” model. However, 80% of these models did not pass the criteria for exact model fit, and the χ2 statistic was simply ignored. Only 2% of overall studies concluded that the model didn’t fit at all, and one of these interpreted the model anyway (4).Why isχ2 being ignored? Researchers often cite reasons for bypassing the exact fit test, including: the statistic is overly sensitive to sample size, it penalizes models with a high number of variables, or a general objection to the logic of the exact fit hypothesis. This has led to a broad consensus in favor of using Approximate Fit Indices (AFIs), such as the RMSEA or CFI, to justify models.However, relying on AFIs often leads to questionable conclusions. In the meta-analysis, only 41% of studies reported χ2 model fit results. Furthermore, 40% of the studies that failed to report a p-value for theirχ2 did report the degrees of freedom. When researchers used these degrees of freedom to cross-check the unreported p-values, they found that all of the non-reported p-values were, in fact, significant.Other glaring reporting issues included:43% of studies failed to report which fit function (e.g., Maximum Likelihood) was used.30% of studies selectively applied more lax cut-off criteria for AFIs than were conventionally acceptable.53% failed to report their AFI cut-off criteria at all.Assumption testing for univariate normality was assessed in only 24% of studies (4).Defending the Exact Fit TestA common criticism of χ2 is that as datasets grow larger, the test detects increasingly trivial discrepancies as sources of misspecification. Strictly speaking, this is not a flaw it is an increase in statistical power. It means the level of certainty with which discrepancies can be considered important has increased.Model misspecification can result from both theoretically relevant and peripheral causal factors. A significant model fit statistic is not trivial just because the underlying cause is trivial; rather, it means that a trivial cause is having a statistically significant effect that needs to be addressed. The χ2 model fit test remains the most sensitive way to detect misspecification in latent variable models and should be prioritised, even when sample sizes are high. Importantly, in the context of SEM, a rejection of model fit does not necessarily require the rejection of every individual hypothesis within the model (4).The Problem with Approximate Fit Indices (AFIs)AFIs provide a conceptually heterogeneous set of fit indices for each hypothesis. None of these indices are accompanied by a formal critical value or significance level, and nearly all arise from unknown distributions. While they are a function of χ2, unlike the χ2 statistic itself, AFIs do not have a verified statistical basis for testing model fit. Despite this, satisfactory AFI values are routinely used to override a significant, failing χ2 test.Monte Carlo simulations of AFIs have concluded that it is impossible to determine universal cut-off criteria for any form of model tested. Furthermore, using AFIs, the probability of correctly rejecting a mis-specified model actually decreases as sample size increases – the exact inverse of the χ2 statistic. Additionally, as model misspecification or correlated errors become more severe, AFIs become increasingly unpredictable, whereas χ2 reliably reflects the severity of the issue (4).Best Practices for Latent Variable ModellingBased on the findings of the meta-analysis and the statistical principles of SEM, researchers should adhere to the following best practices, alongside rigorous testing of heterogeneity, independence, and multivariate normality:Pay strict attention to distributional assumptions.Have a strong theoretical justification for your model.Avoid post hoc model modifications (e.g., dropping indicators, allowing cross-loadings, or correlating error terms) just to achieve fit.Avoid confirmation bias.Use an appropriate and clearly stated estimation method.Recognize the existence of equivalent models.Justify all causal inferences.Use clear, transparent reporting that does not selectively omit model fit statistics.Image: Michael Eid, Tanja Kutscher, Stability of Happiness, 2014. Chapter 13 – Statistical Models for Analyzing Stability and Change in Happiness. https://www.sciencedirect.com/science/article/pii/B9780124114784000138References:Latent Variables in Psychology and the Social Sciences.Structural equation modelling and its application to network analysis in functional brain imaging. https://onlinelibrary.wiley.com/doi/abs/10.1002/hbm.460020104Chapter 7: Assumptions in Structural Equation modelling. https://psycnet.apa.org/record/2012-16551-007McIntosh, C. N. (2015). A cautionary note on testing latent variable models. Frontiers in Psychology. https://www.frontiersin.org/articles/10.3389/fpsyg.2015.01715/full
- Innumerable statistical tests exist for application in hypothesis testing based on the shape and nature of the pertinent variable’s distribution. If however the intention is to perform a parametric test – such as ANOVA, Pearson’s correlation or some types of regression – the results of such a test will be more valid if the distribution of the dependent variable(s) approximates a Gaussian (normal) distribution and the assumption of homoscedasticity is met. In reality data often fails to conform to this standard, particularly in cases where the sample size is not very large. As such, data transformation can serve as a useful tool in readying data for these types of analysis by improving normality, homogeneity of variance or both.For the purposes of Transforming Skewed Data, the degree of skewness of a skewed distribution can be classified as moderate, high or extreme. Skewed data will also tend to be either positively (right) skewed with a longer tail to the right, or negatively (left) skewed with a longer tail to the left. Depending upon the degree of skewness and whether the direction of skewness is positive or negative, a different approach to transformation is often required. As a short-cut, uni-modal distributions can be roughly classified into the following transformation categories: This article explores the transformation of a positively skewed distribution with a high degree of skewness. We will see how four of the most common transformations for skewness – square root, natural log, log to base 10, and inverse transformation – have differing degrees of impact on the distribution at hand. It should be noted that the inverse transformation is also known as the reciprocal transformation. In addition to the transformation methods offered in the table above Box-Cox transformation is also an option for positively skewed data that is >0. Further the Yeo-Johnson transformation is an extension of the Box-Cox transformation which does not require the original data values to be positive or >0.The following example takes medical device sales in thousands for a sample of 2000 diverse companies. The histogram below indicates that the original data could be classified as “high(er)” positive skewed. The skew is in fact quite pronounced – the maximum value on the x axis extends beyond 250 (the frequency of sales volumes beyond 60 are so sparse as to make the extent of the right tail imperceptible) – it is however the highly leptokurtic distribution that that lends this variable to be better classified as high rather than extreme. It is in fact log-normal – convenient for the present demonstration. From inspection it appears that the log transformation will be the best fit in terms of normalising the distribution.Starting with a more conservative option, the square root transformation, a major improvement in the distribution is achieved already. The extreme observations contained in the right tail are now more visible. The right tail has been pulled in considerably and a left tail has been introduced. The kurtosis of the distribution has reduced by more than two thirds.A natural log transformation proves to be an incremental improvement yielding the following results: This is quite a good outcome – the right tail has been reduced considerably while the left tail has extended along the number line to create symmetry. The distribution now roughly approximates a normal distribution. An outlier has emerged at around -4.25, while extreme values of the right tail have been eliminated. The kurtosis has again reduced considerably. Taking things a step further and apply a log to base 10 transformation yields the following: In this case the right tail has been pulled in even further and the left tail extended less than the previous example. Symmetry has improved and the extreme value in the left tail has been bought closer in to around -2. The log to base ten transformation has provided an ideal result – successfully transforming the log normally distributed sales data to normal. In order to illustrate what happens when a transformation that is too extreme for the data is chosen, an inverse transformation has been applied to the original sales data below. Here we can see that the right tail of the distribution has been brought in quite considerably to the extent of increasing the kurtosis. Extreme values have been pulled in slightly but still extend sparsely out towards 100. The results of this transformation are far from desirable overall.Some thing to note is that in this case the log transformation has caused data that was previously greater than zero to now be located on both sides of the number line. Depending upon the context, data containing zero may become problematic when interpreting or calculating the confidence intervals of un-back-transformed data. The mathematical the issue is that the log of zero (or negative numbers) is mathematically undefined. You cannot calculate log(0). Therefore, if your dataset contains zeros, you must add a constant so the minimum value is greater than 0. As log(1)=0, any data containing values <=1 can be made >0 by adding a constant to the original data so that the minimum raw value becomes >1 . Reporting un-back-transformed data can be fraught at the best of times so back-transformation of transformed data is recommended. Further information on back-transformation can be found here. Adding a constant to data is not without it’s impact on the transformation. As the below example illustrates the effectiveness of the log transformation on the above sales data is effectively diminished in this case by the addition of a constant to the original data. Depending on the subsequent intentions for analysis this may be the preferred outcome for your data – it is certainly an adequate improvement and has rendered the data approximately normal for most parametric testing purposes.Taking the transformation a step further and applying the inverse transformation to the sales + constant data, again, leads to a less optimal result for this particular set of data – indicating that the skewness of the original data is not quite extreme enough to benefit from the inverse transformation. It is interesting to note that the peak of the distribution has been reduced whereas an increase in leptokurtosis occurred for the inverse transformation of the raw distribution. This serves to illustrate how a small alteration in the data can completely change the outcome of a data transformation without necessarily changing the shape of the original distribution.There are many varieties of distribution, the below diagram depicting only the most frequently observed. If common data transformations have not adequately ameliorated your skewness, it may be reasonable to select a non-parametric test (like Mann-Whitney or Kruskal-Wallis) as they are distribution-free and work by ranking the data instead of analysing the raw distribution. Another option is to model the data using an alternative distribution (like Poisson or Gamma), whereby you would use a Generalised Linear Model (GLM) to achieve this. Image credit: cloudera.com
- A SAS licence can be prohibitively expensive for many use cases. Installing the software can also take up a surprising amount of hard disk space and memory. For this reason many individuals with light or temporary usage needs choose to access a version of SAS which is licenced to their institution and therefore shared across many users. Are you trying unsuccessfully to access an SAS remotely via your institution using Citrix receiver? This step-by-step guide might help. SAS syntax can differ based on whether a remote versus local server is used. An example of a local server is the computer you are physically using. When you have SAS installed on the PC you are using, you are accessing it locally. A remote server, on the other hand, allows you to access SAS without having SAS installed on your PC. Client software such as Citrix Receiver, allows you to access SAS, and other software, from a remote server. Citrix Receiver is often used by university students, new and/or light users. SAS requires different syntax in order to enable the remote server to access data files on a local computer. For the purpose of this example we are assuming that the data file we wish to access is located, locally, on a drive of the computer we are using. It can be difficult to find the syntax for this on Google, where search results deal more with accessing remote data (libraries) using local SAS than the other way around. This can be a source of frustration for new users and SAS Technical Support are not always able to advise on the specifics of using SAS via Citrix Receiver. The “INFILE”statement and the “PROC IMPORT” statement are two popular options for reading a data file into SAS. INFILE offers the greater flexibility and control, by way of manual specification of data characteristics. PROC IMPORT offers greater automation, in some cases at the risk of formatting error. The INFILE statement must always be used in the context of a DATA step, whereas PROC IMPORT acts as a stand-alone procedure. The document below shows syntax for the INFILE statement and PROC IMPORT procedure for local SAS compared to access via Citrix Receiver.how to open a data file whe… by api-310702664 If you cannot see the document, please make sure that you are viewing the website in desktop mode. In SAS University Edition data file inputing difficulties can occur for a different reason. In order for the LIBNAME statement to run without error, a shared folder must first be defined. If you are using SAS University Edition, and experiencing an error when inputting data, the following videos may be helpful: How to Set LIBNAME File Path (SAS University Edition) Accessing Data Files Via Citrix Receiver: for SAS University Edition Troubleshooting check-list:Was the “libref” appropriately assigned? Was the file location referred to appropriately based on the user context? Was the correct data file extension used?While impractical for larger data sets, if all else fails, one can copy and paste the data from a data file into SAS using the ‘DATALINES’ function.

