Clinical AI in oncology

Clinical AI in oncology is the application of machine learning, natural language processing, and large language models to cancer care - risk and outcome prediction, extraction of structured data from clinical notes, medical image analysis, and decision support - together with the validation and governance these systems require before they touch a patient.

A working definition

The term groups several distinct technologies under one label. Supervised machine learning builds prediction models from structured electronic-health-record data: who is likely to respond to a therapy, who is at risk of a toxicity. Natural language processing and, more recently, large language models read the unstructured text of clinical notes and lift discrete facts out of it. Computer-vision models read radiology and pathology images. Clinical decision support wraps these outputs in a workflow a clinician can act on. What ties them together is that each is a statistical system whose usefulness depends entirely on the quality of the data underneath and the rigor of its validation.

Why it matters

Oncology generates more data per patient than almost any other field of medicine, and most treatment decisions are made under uncertainty. AI is attractive precisely because it promises to turn that data into something a clinician can use: a per-patient risk-benefit profile before the first dose, an earlier flag on an image, a structured summary drawn from a hundred pages of notes. The 2025 review Artificial intelligence across the cancer care continuum maps these applications end to end and makes the central point plainly - AI in oncology is not one technology but many overlapping ones, each at a different stage of validation and adoption.

That last clause is the reliability and governance gap. A model can be fluent and still be wrong. The peer-reviewed evaluation of ChatGPT on physician-posed questions found exactly that tension: largely accurate answers that improved measurably between model versions, alongside important limitations and a residue of inaccuracies the authors concluded further research and validation are needed to correct. A tool that is right most of the time still has to be checked before it reaches a patient. Closing that gap - deciding where these systems may operate, who validates them, and what guardrails must exist first - is the work that turns clinical AI from a demonstration into infrastructure.

Travis Osterman's work in clinical AI

Dr. Travis Osterman is a board-certified medical oncologist and clinical informatician - Associate Vice President for Research Informatics at Vanderbilt Health and Director of Cancer Clinical Informatics at the Vanderbilt-Ingram Cancer Center - working at the intersection of cancer care, applied AI, and the data standards AI depends on. His clinical-AI record spans four modes: evaluation, framework, synthesis, and building.

Evaluation. Dr. Osterman co-authored an early peer-reviewed clinical evaluation of ChatGPT, Accuracy and Reliability of Chatbot Responses to Physician Questions (Goodman et al., JAMA Network Open, 2023), which graded model answers to 284 questions written by 33 physicians across 17 specialties and documented both the model's competence and the important limitations described above.

Framework. Alongside that evaluation he co-authored On the cusp: Considering the impact of artificial intelligence language models in healthcare (Goodman, Patrinely, Osterman, Wheless & Johnson, Med, 2023) - an early articulation of where large language models should be allowed to operate in medicine, who validates them, and what safety guardrails must exist before they reach patients.

Synthesis. He is senior author of Artificial intelligence across the cancer care continuum (Riaz, Khan & Osterman, Cancer, 2025), the review that maps AI across the full arc of cancer care and frames rigorous validation and responsible implementation as the precondition for clinical value.

Building. His group builds the extraction and prediction systems, not only critiques them. mCODEGPT (Zhang et al., Communications Medicine, 2025) uses large language models to pull mCODE-conformant elements out of clinical free text with no task-specific training, and SmokeBERT (Tan & Osterman, JCO Clinical Cancer Informatics, 2025) bridges clinical narratives and structured smoking data to improve lung-cancer screening. The recurring design choice is that a model performs best when it has a structured schema to aim at. That is also why the prediction work in the GE HealthCare Digital Precision Oncology program - machine learning on real-world EHR data to forecast immune checkpoint inhibitor efficacy and toxicity - was tractable in the first place.

Underneath all of it is a position, not a product: clinical AI in oncology should be validated narrowly before it is deployed broadly, framed honestly to clinicians and patients about what it can and cannot do, and built on top of structured data standards rather than used as a workaround for the lack of them. The AI in oncology case study traces this arc paper by paper; the AI in oncology expertise page lists the full publication and talk record; and /research/ holds the complete peer-reviewed bibliography. His governance roles, including the NCCN Digital Oncology Forum and the chair of the mCODE Executive Committee, are described on /leadership/.

Key publications

  1. Goodman RS, Patrinely JR, Stone CA Jr, Zimmerman E, et al. Accuracy and Reliability of Chatbot Responses to Physician Questions. JAMA Network Open 2023;6(10):e2336483.
  2. Goodman RS, Patrinely JR, Osterman T, Wheless L, Johnson DB. On the cusp: Considering the impact of artificial intelligence language models in healthcare. Med 2023;4(3):139-140.
  3. Riaz IB, Khan MA, Osterman TJ. Artificial intelligence across the cancer care continuum. Cancer 2025;131(16):e70050.
  4. Zhang K, Huang T, Malin BA, Osterman T, Long Q, Jiang X. Introducing mCODEGPT as a zero-shot information extraction from clinical free text data tool for cancer research. Communications Medicine 2025;5(1):422.
  5. Tan H, Osterman TJ. SmokeBERT and Beyond: Bridging Clinical Narratives and Structured Smoking Data To Improve Lung Cancer Screening. JCO Clinical Cancer Informatics 2025;9:e2500350.

Related concepts: mCODE · cancer data standards · precision oncology · clinical genomics in the EHR.

See also: AI in oncology (case study) · AI in oncology (expertise) · peer-reviewed research · leadership and governance.

Back to top

Frequently asked questions

What is clinical AI in oncology?
Clinical AI in oncology is the application of machine learning, natural language processing, and large language models to cancer care. It spans risk and outcome prediction, extraction of structured data from clinical notes, medical image analysis, and clinical decision support, along with the validation and governance those systems require before they are used with patients.
What are the main applications of AI in oncology?
Applications run across the cancer care continuum: risk stratification and screening, diagnosis from radiology and pathology images, treatment selection, toxicity prediction, survivorship, and clinical documentation. The 2025 review Artificial intelligence across the cancer care continuum, co-authored by Dr. Travis Osterman in Cancer, maps these applications end to end and notes that each sits at a different stage of validation and adoption.
What is Travis Osterman's expertise in clinical AI?
Dr. Travis Osterman is a board-certified medical oncologist and clinical informatician, Associate Vice President for Research Informatics at Vanderbilt Health, and Director of Cancer Clinical Informatics at the Vanderbilt-Ingram Cancer Center. His clinical-AI work includes an early peer-reviewed evaluation of ChatGPT (Goodman et al., JAMA Network Open 2023), a framework paper on language models in healthcare (Med 2023), the Cancer 2025 care-continuum review, and information-extraction systems mCODEGPT and SmokeBERT.
What is the governance concern with clinical AI in oncology?
A model can be fluent and still be wrong. The peer-reviewed evaluation of ChatGPT on physician questions found largely accurate answers that improved over time, together with important limitations and inaccuracies the authors concluded further research and validation are needed to correct. Governance means deciding where these systems may operate, who validates them, and what safety guardrails must exist before they reach patients. Dr. Osterman's position is that clinical AI should be validated narrowly before it is deployed broadly, framed honestly about its limits, and built on structured data standards.