Real-World Applications of GENIE and a Taxonomy for Defining Cancer Outcomes
In a session on Monday morning at the 2018 American Society of Clinical Oncology (ASCO) Annual Meeting (June 4, 2018; Chicago, IL), Deborah Schrag, MD, FASCO, Dana-Farber Cancer Institute (Boston, MA), explained the applicability of the American Association for Cancer Research (AACR) project, Genomics Evidence Neoplasia Information Exchange (GENIE), and provided a brief overview of a standard taxonomy system for defining cancer outcomes embedded in GENIE.
Delivering on the goals of precision medicine requires a sufficient amount of computable patient data, the ability to predict clinical phenotypes, and the ability to select optimal therapy, Dr Schrag began. She provided the important distinction between genomics (molecular characterization of tumor genes and their expression) and phenomics (clinical characteristics of risks, exposures, and treatment outcomes) and then explained that phenomics in real-world data lack consistent endpoints, lack consistent time intervals, and are retrospectively ascertained from EHR data.
While many stakeholders need real-world outcomes, she continued, there are legitimate concerns over how to define them. Most cancer patients are not treated on clinical trials where structured outcomes are consistently captured, and real-world data are often “messy” and unstructured. “We have real-world data, but we face a lack of methods, tools, and standards to generate real-world evidence,” Dr Schrag asserted.
This point led into a discussion of GENIE and its goal of combining genomics and phenomics to produce better outcomes for patients with cancer. Achieving the goals of precision medicine requires integration of genomic, therapeutic, and outcomes data, but an abundance of unstructured data (ie, diagnosis dates, recurrence dates, responses to treatment, disease-free survival, progression-free survival, and treatment toxicity) makes phenomics significantly challenging.
Dr Schrag described the importance of standards in oncology to facilitate interpretation and communication of data, citing the AJCC Staging Classification System and RECIST as useful but imperfect resources. On the other hand, GENIE has a well-defined vision for linking genomic to phenomic data and using this information cache for discovery, she argued. This vision involves developing data standards that can be applied consistently across sites, building consensus across centers about data standards, creating a cache of “gold standard” EHR data, building machine learning algorithms to curate EHR data, validating the algorithms against the “gold standards” data, refining the algorithms, and deploying the algorithms to accelerate curation at scale.
She explained how clinical annotation by both humans and machines for analysis for real-world data requires data standards, which is where PRISSMM comes into the folds of GENIE. PRISSMM is a standard taxonomy for classification and communication of structured information about cancer status and treatment outcomes following the assignment of TNM stage for patients with solid tumors. Each letter in the PRISSMM acronym corresponds to a dimension of cancer status or treatment response:
- P: Pathologic evidence of locoregional or distant evidence of tumors
- R: Radiographic evidence of locoregional recurrence or persistent tumors
- I: Imaging evidence of distant/disseminated tumors beyond the primary site
- S: Symptoms of tumors on physical exam or symptoms that can be attributed to tumors
- S: Signs of cancer on physical exam or symptoms that can be attributed to tumors
- M: Tumors marker evidence of persistent or recurrent tumors
- M: Oncology medical provider assessment
PRISSMM establishes standard directives for curation of cancer treatment outcomes from EHRs as well, Dr Schrag noted, and the provided taxonomy and standard nomenclature facilitate interpretation and communication about real-world data.
In her concluding remarks, Dr Schrag shared GENIE’s current accomplishments and ongoing endeavors. The project has successfully defined standard approach for clinical data curation of real-world outcomes and developed training materials and data directives. Initiatives in progress include optimizing accuracy, tracking efficiency metrics, implementation across centers and EHR systems, and teaching machines to perform curation at scale.
Ultimately, GENIE aims to build a knowledge discovery engine capable of improving human curation, improving machine curation, and building capacity to create, share, and use real-world data, she added.—Zachary Bessette


