• Understanding The Importance of Computable Phenotypes in Regulatory Submissions of RWE

    Understanding The Importance of Computable Phenotypes in Regulatory Submissions of RWE

    Phenotypes

    As the landscape of healthcare data evolves, computable phenotypes are becoming crucial for the generation and validation of real-world evidence (RWE), especially in regulatory contexts. A computable phenotype is “a clinical condition, characteristic, or set of clinical features that can be determined solely from data in electronic health records (EHRs) and ancillary data sources and does not require chart review or interpretation by a clinician.”(1) Essentially, computable phenotypes are machine-executable, algorithmic definitions used to select patients with specific clinical features, such as conditions, exposures, or outcomes, in large, often cluttered real-world data (RWD) sources like electronic health records (EHRs) and medical claims. These definitions are created from structured data elements, logical expressions, and clinical criteria that facilitate objective, consistent, and reproducible identification of populations appropriate for a study.(1, 2)

    Regulatory bodies like the USFDA (3) and EMA (4) increasingly expect RWE adopted in submissions to be transparent and methodologically rigorous. Computable phenotypes are pivotal in this case as they enable replicable cohort selection and outcome determination.(5) Sponsors can offer regulators with a clear, auditable trajectory of selection of patient populations and definition of outcomes by inserting these algorithms directly into study protocols and analysis plans. Such rigorous detailing is especially crucial for studies relying on different data sources with changing formats, coding systems, and clinical granularity.(6-8)

    One of the major advantages of computable phenotypes is the consistency they offer to RWE studies. Flexibility of real-world datasets, owing to disparities in healthcare delivery, data capture, or coding practices, can result in bias, thus reducing the dependability of findings. Computable phenotypes help alleviate these risks by applying a common, validated logic across datasets, diminishing the possibility of misclassification and enabling consistency in the application of inclusion and exclusion criteria. Consequently, regulators can evaluate the strength of the evidence with greater confidence.(1, 2, 8)

    Computable phenotypes also enhance the evidence generation efficiency. By computerizing the selection of eligible patients, exposures, and clinical endpoints, researchers can simplify study implementation and minimize the dependence on manual chart reviews or case-by-case abstraction. This automation reduces timelines and thus human error, which is particularly important in large-scale or time-sensitive studies. Additionally, this approach is suitable for flexibility, making the replication of analyses across multiple databases or healthcare systems possible, further enabling the assessment of robustness and generalizability.(1, 2, 8)

    Despite evident benefits, applying high-quality computable phenotypes has some limitations. The inconsistencies in clinical data, changing terminologies, and varying methods of data capture can make standardization challenging. Moreover, while developing a computable phenotype may be technically feasible, justifying its accuracy across different populations and settings is often resource-demanding; which necessitates cooperation among clinicians, data scientists, informaticians, and regulatory stakeholders.(8-10) Programs like OHDSI (11, 12) and USFDA’s Sentinel Initiative (13) have developed phenotype definitions, but more work is needed to ensure harmonization and broad applicability.

    While many computable phenotypes rely on systematized EHR data, such data may fail to entirely capture the clinical details of a patient’s medical record. Machine learning (ML)–enabled natural language processing (NLP) tools, like the open-source Clinical Annotation Research Kit (CLARK),(14) are increasingly being implemented to extract unstructured clinical notes. CLARK facilitates nonexpert users to apply standard ML algorithms by defining features found in text, improving phenotyping accuracy by integrating variables not available in structured data. This standardizes the use of refined phenotyping methods to expand access to richer, more sensitive phenotype algorithms across research settings. Tools like CLARK have shown robust performance in real-world scenarios, including phenotyping paediatric diabetes and non-alcoholic fatty liver disease, and represent a major development in making ML-driven phenotyping more available.(15)

    Validation of computable phenotypes is crucial. A well-structured computable phenotype should be transparent and also perform well in recognizing true cases or outcomes. Regulatory guidance now increasingly focuses on the customizability of phenotypes to the research question and their ability to achieve clear, clinically significant results in RWE studies. Intrinsically, sponsors are expected to report the logic, validation status, and limitations of the phenotypes used, allowing for informed review and analysis by regulators.(1, 2, 5)

    Computable phenotypes are the methodological pillars of reliable RWE. They decipher messy, heterogeneous RWD into structured, actionable insights to support high-stakes regulatory decisions. With the regulatory science adopting more complex data and evidence frameworks, computable phenotypes will continue to be indispensable in facilitating robust, transparent, and also reproducible RWE that aligns with public health preferences.

    Become A Certified HEOR Professional – Enrol yourself here!

    References

    1. Richesson RL, et al. Electronic Health Records–Based Phenotyping. NIH Pragmatic Trials Collaboratory. [Accessed online on 17th June 2025]. Available at: https://rethinkingclinicaltrials.org/chapters/conduct/electronic-health-records-based-phenotyping/definitions/
    2. Cameron CB. A User’s Guide to Computable Phenotypes. [Accessed online on 17th June 2025]. Available at: https://dcricollab.dcri.duke.edu/sites/NIHKR/KR/Blake_Users_Guide_to_Computable_Phenotypes.pdf
    3. Considerations for the Use of Real-World Data and Real-World Evidence To Support Regulatory Decision-Making for Drug and Biological Products. August 2023. [Accessed online on 17th June 2025]. Available at: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/considerations-use-real-world-data-and-real-world-evidence-support-regulatory-decision-making-drug
    4. Real-world evidence framework to support EU regulatory decision-making Report on the experience gained with regulator-led studies from September 2021 to February 2023.
    5. Real-World Data: Assessing Electronic Health Records and Medical Claims Data to Support Regulatory Decision-Making for Drug and Biological Products – Guidance for Industry. 2024. [Accessed online on 17th June 2025]. Available at: https://www.fda.gov/media/152503/download
    6. Tasker RC. Why Everyone Should Care About “Computable Phenotypes”. Pediatr Crit Care Med. 2017 May;18(5):489-490.
    7. Masison J, Lehmann HP, Wan J. Utilization of Computable Phenotypes in Electronic Health Record Research: A Review and Case Study in Atopic Dermatitis. Journal of Investigative Dermatology. 2025; 145(5):1008-1016.
    8. Ahmad FS, Ricket IM, Hammill BG, et al. Computable Phenotype Implementation for a National, Multicenter Pragmatic Clinical Trial: Lessons Learned From ADAPTABLE. Circ Cardiovasc Qual Outcomes. 2020; 13(6):e006292.
    9. Shah C. Computable Phenotypes for Generating RWE: What Are They and Can They Really Be Standardized and Reused? Value & Outcomes Spotlight. 2022; 8(3):S1.
    10. He T, Belouali A, Patricoski J, et al. Trends and opportunities in computable clinical phenotyping: A scoping review. Journal of Biomedical Informatics. 2023; 140:104335.
    11. The Observational Health Data Sciences and Informatics (OHDSI). [Accessed online on 17th June 2025]. Available at: https://www.ohdsi.org
    12. Zelko JS, Gasman S, Freeman SR, et al. Developing a Robust Computable Phenotype Definition Workflow to Describe Health and Disease in Observational Health Research. 2023. [Accessed online on 17th June 2025]. Available at: https://arxiv.org/abs/2304.06504
    13. Sentinel Initiative. A General Framework for Developing Computable Clinical Phenotype Algorithms. 2024. [Accessed online on 17th June 2025]. Available at: https://www.sentinelinitiative.org/news-events/publications-presentations/general-framework-developing-computable-clinical-phenotype
    14. Repository for CLARK, the Clinical Annotation Research Kit. 2019. [Accessed online on 17th June 2025]. Available at: https://github.com/NCTraCSIDSci/clark
    15. Pfaff ER, Crosskey M, Morton K, Krishnamurthy A. Clinical Annotation Research Kit (CLARK): Computable Phenotyping Using Machine Learning. JMIR Med Inform 2020; 8(1):e16042.
  • The Importance of the ISPOR SUITABILITY Checklist in HTA Involving EHR Data

    The Importance of the ISPOR SUITABILITY Checklist in HTA Involving EHR Data

    Health Technology Assessment (HTA) plays a critical role in evaluating the value of medical interventions. With the increasing availability of Electronic Health Records (EHRs), the integration of real-world data (RWD) into HTA has become more prevalent. The adoption of EHR data into HTA offers significant opportunities to enhance various aspects of health and medicine, from evaluating medical products and technologies to improving healthcare delivery. As a rich source of RWD generated at the point of care or during daily activities, EHR systems provide detailed insights into patients’ health and care. EHR data have been utilized for clinically relevant research, such as monitoring quality of care and medication adherence, and developing clinical predictive models.[1]

    However, the use of EHR data in HTA presents unique challenges. Only a portion of EHR information is structured and ready for statistical analysis, with completeness and accuracy often compromised since EHRs are primarily collected for clinical or administrative purposes and not necessarily for research purposes. This affects the types, detail, and reliability of variables collected, and the timeliness of data availability. Relevant data may be missing due to the problem-focused nature of clinical summaries. Additionally, a single EHR system may not capture a patient’s entire clinical history if they receive care from multiple settings with different EHR systems. Therefore, data may need to be gathered from various healthcare entities and linked with other sources, such as disease registries, pharmacy data, national birth and death registrations, and claims. Time lags between clinician use of EHR information and analyst access to EHR data for decision-making are also common.[2-5]

    To address these issues, the Professional Society for Health Economics and Outcomes Research (ISPOR) has developed the SUITABILITY checklist.[1] This checklist is designed to assess the appropriateness and quality of RWD sources, including EHRs, for HTA.[1]

    The ISPOR SUITABILITY Checklist contains two main elements: data delineation and data fitness for purpose. Data delineation provides a comprehensive understanding of the data and assesses their trustworthiness by describing data under three headings: data characteristics, provenance, and governance. On the other hand, data fitness for purpose examines two main components: the accuracy and completeness of items (data reliability) and the suitability of the data to answer the particular question at hand (data relevance). Data relevance is assessed by examining whether the data aligns with the research question or HTA objective, considering the population, intervention, comparator, outcomes, and settings of interest. Core issues for data relevance include data content; care settings and the time period of interest; and sample size and follow-up period.[1]

    Ensuring relevance, completeness, accuracy, timeliness, and generalizability is crucial for deriving meaningful insights and making informed decisions that impact a wide range of patients. Completeness evaluates whether the EHR data includes all necessary information to comprehensively address the research question, checking for missing data, gaps, and the presence of all relevant variables. Accuracy involves verifying the precision of recorded information to ensure it accurately reflects real-world clinical scenarios, while timeliness assesses whether the data is current and reflects contemporary clinical practices. Generalizability evaluates the extent to which findings from the EHR data can be applied to the broader population, examining the representativeness and diversity of the patient population included.[1,6-8]

    The checklist also addresses potential biases and confounding factors, and promotes transparency and reproducibility by providing a standardized framework for evaluating EHR data, which allows for consistent documentation and verification of research methodologies. Ultimately, the checklist supports robust decision-making and policy development by providing reliable evidence for assessing medical technologies, including their safety, efficacy, cost-effectiveness, and overall impact on patient outcomes and healthcare systems.[1,9]

    Digital health products offer new ways to manage and monitor care. In regulatory agencies, this demand is driven by the influx of innovative technologies and the need for quicker assessments. In this challenging environment, EHR-derived data hold promise for meeting information needs that traditional RWD platforms struggle to address. The task force’s framework and checklist are expected to evolve as experience with EHR-derived data increases.[5-8]

    The ISPOR task force also acknowledge the limitations of the checklist, such as not accounting for national or local EHR data system idiosyncrasies. Secondly, the checklist offers a broad categorization of data suitability components rather than explicit standards for data provenance, reliability, or relevance. Thirdly, the benefits of integrating EHR data with other data types have not been considered. These limitations notwithstanding, it is expected that as experience with EHRs grows, we will have a better understanding of the nature of data sources that are better suited for specific HTA questions. Addressing some components of the SUITABILITY framework may require significant efforts and resources, and some information might not be accessible to end users. With the rapid advancement of AI, reimagining how unstructured EHR information is transformed into EHR-derived data will likely necessitate updates to the framework and checklist.[1,10,11]

    The ISPOR SUITABILITY checklist is a crucial resource for effectively utilizing EHR data in HTA. By evaluating critical aspects like relevance, completeness, accuracy, timeliness, and generalizability, the checklist improves the quality, reliability, and validity of EHR data. It fosters transparent and reproducible research, supports evidence-based decision-making, and enhances the integration of RWD in HTA processes. As EHR data usage expands, the SUITABILITY checklist will continue to be vital for conducting thorough and influential HTAs, ultimately contributing to better healthcare outcomes globally.

    Become A Certified HEOR Professional – Enrol yourself here!

    References:

    1. Fleurence RL, Kent S, Adamson B, Bouée-Benhamiche E, García Martí S, Ramsey S. Assessing Real-World Data From Electronic Health Records for Health Technology Assessment: The SUITABILITY Checklist: A Good Practices Report of an ISPOR Task Force. ISPOR Report. 2024 Jun;27(6):692-701.
    2. Califf RM, Robb MA, Bindman AB, et al. Transforming evidence generation to support health and health care decisions. N Engl J Med. 2016;375(24):2395– 2400.
    3. Sherman RE, Anderson SA, Dal Pan GJ, et al. Real-world evidence – what is it and what can it tell us? N Engl J Med. 2016;375(23):2293–2297.
    4. Graili P, Guertin JR, Chan KKW, Tadrous M. Integration of real-world evidence from different data sources in health technology assessment. J Pharm Pharm Sci. 2023;26:11460.
    5. Shadmi E, Flaks-Manov N, Hoshen M, et al. Predicting 30-day readmissions with preadmission electronic health record data. Med Care. 2015;53(3):283–289.
    6. Duke Margolis Center for Health Policy. Determining Real-World Data’s Fitness for Use and the Role of Reliability; 2019:1–54. Available from: https://healthpolicy.duke.edu/ sites/default/files/2019-11/rwd_reliability.pdf
    7. US Food and Drug Administration. Real-World Data: Assessing Electronic Health Records and Medical Claims Data To Support Regulatory DecisionMaking for Drug and Biological Products; 2021:1–39. Available from: https://www.fda.gov/ media/152503/download.
    8. European Medicines Agency. Data Quality Framework for EU medicines regulation; 2023:1–42. Available from: https://www.ema.europa.eu/en/documents/regulatory-proceduralguideline/data-quality-framework-eu-medicines-regulation_en.pdf.
    9. O’Rourke B, Oortwijn W, Schuller T, International Joint Task Group. The new definition of health technology assessment: a milestone in international collaboration. Int J Technol Assess Health Care. 2020;36(3):187–190.
    10. Fleurence RL, Shuren J. Advances in the use of real-world evidence for medical devices: an update from the national evaluation system for health technology. Clin Pharmacol Ther. 2019;106(1):30–33.
    11. Desai RJ, Matheny ME, Johnson K, et al. Broadening the reach of the FDA Sentinel system: a roadmap for integrating electronic health record data in a causal analysis framework. NPJ Digit Med. 2021;4(1):170.
  • Using Synthetic Controls in Oncology Real-World Data Studies for Treatment Insights

    Using Synthetic Controls in Oncology Real-World Data Studies for Treatment Insights

    In the ever-evolving landscape of healthcare, Real-World Data (RWD) studies have emerged as a pivotal tool in shaping treatment strategies and enhancing patient outcomes. For the Health Economics and Outcomes Research (HEOR) industry, these studies hold a special significance, providing valuable insights into the real-world effectiveness of treatments. In recent times, synthetic controls in oncology RWD studies have gained momentum, offering a novel approach to accelerate the development of new treatments and broaden our understanding of their impacts.[1]

    The essence of oncology research lies in its constant pursuit of more effective and targeted treatments for a range of malignancies. Traditional Randomized Clinical Trials (RCTs), while crucial, often have limitations that restrict their ability to mirror real-world scenarios comprehensively. This is where RWD studies step in, utilizing data collected from routine clinical practice to bridge the gap between RCTs and real-life patient experiences. However, challenges such as confounding variables and lack of randomization persist in these studies, prompting the exploration of innovative methodologies like synthetic controls.[1,2]

    Synthetic controls, in essence, involve the creation of a hypothetical control group that mirrors the characteristics of the treatment group. By leveraging historical patient data, demographic information, disease progression, and other relevant factors, researchers can construct a comparable control arm. This approach, rooted in advanced statistical techniques, provides a powerful tool to estimate treatment effects and mitigate biases that might arise in traditional observational studies.[5]

    Synthetic control arms are a variant of external control arms, and represent an inventive strategy where researchers create a virtual or synthetic control group by harnessing existing data, rather than enlisting fresh participants for the control cohort. The formulation of a synthetic control arm entails evaluating patient information contained within pre-existing datasets, like electronic health records, which is rendered anonymous and stripped of any personally identifiable details. These synthetic controls replicate real patients who would conventionally be enrolled as part of the trial’s control group.[5]

    The application of synthetic controls in oncology RWD studies offers several key advantages to the HEOR industry. It expedites the evaluation of new treatments by reducing the time required for traditional RCTs. This acceleration is paramount in oncology, where swift access to effective treatments can significantly impact patient outcomes and quality of life. Moreover, synthetic controls enable researchers to glean insights from real-world patient populations that might have been excluded from traditional clinical trials due to stringent eligibility criteria. This inclusivity not only enhances the generalizability of study findings but also provides a more holistic understanding of treatment efficacy across diverse patient demographics. Next, by harnessing the richness of RWD, synthetic controls facilitate the assessment of treatment effects in various subpopulations, shedding light on the intricate interplay between treatments and patient characteristics.[2-5]

    As the HEOR industry delves deeper into the utilization of synthetic controls for oncology RWD studies, it is imperative to acknowledge the challenges that accompany this innovative approach. Rigorous validation and robust sensitivity analyses are paramount to ensure the credibility of synthetic control results. Transparency in methodology and data sources is equally vital to establish trust among stakeholders and foster the adoption of this methodology in regulatory decision-making.[3]

    In conclusion, the integration of synthetic controls in oncology RWD studies presents a promising avenue for the HEOR industry to enhance treatment development and expedite knowledge acquisition. This methodology’s ability to replicate a control group closely resembling the treatment group regarding relevant variables addresses some of the limitations inherent in observational studies. As the healthcare landscape continues to evolve, embracing innovative methodologies like synthetic controls has become necessary to drive advancements in oncology treatments and ultimately improve patient outcomes.

    Become A Certified HEOR Professional – Enrol yourself here!

    References

    1. Prasad V. Reliable, cheap, fast and few: What is the best study for assessing medical practices? Randomized controlled trials or synthetic control arms?. European journal of clinical investigation. 2021 Aug 1;51(8):e13580.
    2. Yap TA, Jacobs I, Baumfeld Andre E, et al. Application of Real-World Data to External Control Groups in Oncology Clinical Trial Drug Development. Front Oncol. 2022 Jan 6;11:695936.
    3. Thorlund K, Dron L, Park JJH, Mills EJ. Synthetic and External Controls in Clinical Trials – A Primer for Researchers. Clin Epidemiol. 2020 May 8;12:457-467.
    4. Greshock J, Lewi M, Hartog B, Tendler C. Harnessing real-world evidence for the development of novel cancer therapies. Trends in Cancer. 2020 Nov 1;6(11):907-9.
    5. Banerjee R, Midha S, Kelkar AH, Goodman A, Prasad V, Mohyuddin GR. Synthetic control arms in studies of multiple myeloma and diffuse large B-cell lymphoma. British journal of haematology. 2022 Mar 1;196(5):1274-7.