Agentic deep research playbooks for multimodal clinical AI
Working outline for RFA-RM-27-011 (PRIMED-AI Playbook U01). Due October 9, 2026. This is a jumping-off point, not a submission draft.
Concept. Platform-independent playbooks and open-source software for using AI agents to perform deep research across clinical imaging, molecular and clinical data, and published evidence, with comprehensive provenance traces and tooling for error detection, uncertainty reporting, and human oversight.
Specific Aims
Overarching: develop and preliminarily validate two distinct, reusable frameworks for responsible multimodal clinical AI, then test them on clinical imaging plus at least one other health-data modality.
- Framework A — provenance-complete agentic deep research. Procedures and software for multi-agent inquiry across images, molecular/clinical records, and literature, such that another team can reconstruct what was asked, which objects were used, what was extracted, what was claimed, what conflicted, and who acted.
- Framework B — error, uncertainty, and human oversight. Procedures and software for evaluation, escalation, leakage/bias/drift monitoring, model and data cards, and a named-human gate before any release. Acceptance of a synthesis is not publication.
- Interoperation and test. One clinician-owned development case (recommended: pathology slides + molecular/genomic + clinical data). One transfer test (recommended: radiology). Month 12: post draft frameworks to the PRIMED-AI portal.
- TODO: lock two RFA topic clusters so A and B are not two methods for one topic. Working map: A → discovery, harmonization, transparency, standards-based linkage. B → uncertainty, leakage, cards, generalizability, regulatory readiness.
- TODO: name the imaging class and the second modality in Aim 3. Literature is context, not the second modality.
- TODO: write success metrics a non-original team can run.
- TODO: decide whether open-source software is an award product or only a test bed. The NOFO funds playbook frameworks (SOPs, protocols). Software should support those frameworks, not replace them.
Significance
Clinical imaging is usually developed in isolation from other health data. Agentic systems now default to fluent synthesis. The field has model cards, datasheets, CONSORT-AI, SPIRIT-AI, and a leakage literature. It does not have portable procedures for multi-agent deep research that keep claims attributable and keep release a named human act.
- TODO: one paragraph on the pathology (or chosen) clinical need, written with the clinical lead.
- TODO: say what existing frameworks we are not incrementing.
Innovation
Two separable frameworks with one short handoff (a provenance packet enters review). Multi-agent research under human supervision, not a single chatbot and not a new foundation model. Playbooks any PRIMED-AI group can run, plus open-source tooling they can ignore.
- TODO: keep this narrow. Do not propose a WSI/DICOM product, a CDS tool, or a clinical trial.
Approach
A. Provenance / deep research. Human frames an inquiry. Sources stay first-class (image, report, variant table, note, paper). Agents extract statements and propose assertions. People accept or reject. Conflicts are retained. Activity is logged.
- TODO: evidence object for images: stable IDs (slide or DICOM UIDs), content hash, region locator. Do not treat a png-to-LLM attachment as multimodal clinical AI.
- TODO: join record to a second health-data class on the same inquiry.
- TODO: early SOP + term list (inquiry, source, statement, assertion, evidence, context, activity).
- TODO: products: schema, linkage checklist, evidence rubric, contradiction ledger, agent-action audit.
B. Oversight / release-readiness. Consume a provenance packet, not a paragraph. Named review. Uncertainty a user can act on. Leakage, bias, and drift checks. Second signature before an edition exists. Cards generated from recorded fields.
- TODO: error taxonomy (wrong source, leakage, overconfident fusion, erased disagreement, agent action treated as first-party evidence, …).
- TODO: products: taxonomy, escalation SOP, human-gate protocol, monitors, card templates, research-vs-release tree.
- TODO: this award should almost always land on research / playbook evidence, not deployment.
Development case (recommended, not locked). JHMI computational pathology: digital slides + molecular/genomic + structured clinical data, on a question Alexander Baras would own.
Transfer test (recommended, not locked). Radiology at JHMI, so identifiers, workflow, and oversight survive a second specialty. Cheng Ting Lin (evaluation, governance, RAID).
Test bed. Ben Guo’s independent multi-agent platform for collaborative deep research, developed through XTE Industries LLC (public: sponge.computer, hraness.com). He is a former cofounder of Zo Computer and previously worked at Stripe. Yixian has been using Zo Computer in her own research; that experience informs the supervision model. Zo is career history and a research tool, not an award deliverable. Current imaging in the test bed is generic files handed to a language model, not DICOM or whole-slide imaging.
- TODO: keep PHI on the clinical side unless a BAA path exists. Prefer public or de-identified packets for year 1 if the lead is not in.
- TODO: do not send reviewers a private repository. Do not cite this page as a publication. Do not cite sponge.pub (404).
- TODO: record honesty gaps in the Research Strategy: no DICOM/WSI; child-agent provenance not first-party yet; Aug 19, 2026 pressure test failed materialization gates; search API activation-blocked; ClinicalTrials.gov is metadata only.
Validation (frameworks, not a CDS tool).
| Question | Pass |
|---|---|
| Completeness | Independent rater reconstructs a packet |
| Separability | Products still classify as A, B, or interface |
| Replay | Second team runs the SOPs without the original PIs |
| Contradiction | Seeded conflict is retained |
| Human gate | Agent-only “release” is refused |
| Leakage | Seeded identity or report-text leak is flagged |
| Transfer | Same SOP IDs; only appendix mappings change |
Investigators and environment
| Role | Person | Status |
|---|---|---|
| Scientific lead | Yixian Zheng, Carnegie BSE; adjunct, JHU Biology | Investigator; applicant org open |
| Technical lead | Ben Guo, XTE Industries LLC; former cofounder, Zo Computer; Stripe | Investigator; MPI vs key personnel open |
| Clinical lead, development | Alexander Baras (recommended) | Outreach not sent |
| Clinical lead, transfer | Cheng Ting Lin (recommended) | Outreach not sent |
| Contact PI / SO | TBD | Required for submission |
| Project manager | TBD | Day-to-day SOP versions |
- TODO: applicant institution (Carnegie, Hopkins, XTE Industries LLC, or another eligible U.S. org). A PR LLC is domestic and can be a small business; it is a weak Environment score for this U01.
- TODO: confirm Yixian’s JHU adjunct is current before submission (emails use it; Carnegie pages may not).
- TODO: both PDs/PIs need ORCID-linked eRA Commons. If one person is PI and Signing Official, they need two Commons accounts.
- TODO: SAM / Grants.gov if a new org applies (six weeks). Open date September 9.
- TODO: users: at least one clinician, one informatician, one non-developer reviewer, with names.
Timeline
| Milestone | When |
|---|---|
| Lock clinical lead and case | Y1 Q1 |
| Framework A v0 and Framework B v0 | Y1 Q1–Q2 |
| First user cycle (or public-proxy case) | Y1 Q2–Q3 |
| Post drafts to PRIMED-AI portal | Y1 Q4 (required) |
| Working-group / Validation Center revision | Y2 Q1–Q2 |
| Radiology (or second-site) transfer | Y2 Q2–Q3 |
| Independent replay | Y2 Q3 |
| v1.0 freeze + showcase | Y2 Q4 |
Budget cap $300k DC/year, 2 years. Clinical trial not allowed. Foreign orgs and foreign components are ineligible (Puerto Rico is domestic). Contact: ODPRIMED-AI@od.nih.gov.
- TODO: Gantt, travel (Bethesda annual meeting; year-2 showcase), F&A, DMS plan, resource-sharing (public copyright license expected).
Open decisions
- Applicant institution and contact PI.
- Ben as MPI or key personnel; Zo / Substrate COI language.
- Clinical MPI / co-investigator (Baras, Lin, both, neither).
- Data packet and human-subjects sentence (public-proxy, delayed-onset, or describable JHMI data).
- How much open-source software is in scope versus playbook-only.
Do not cut to the 1+12 page limits until 1–4 are chosen.
Outreach emails
Confirm baras@jhmi.edu and clin97@jhmi.edu on the current JHMI directory before sending.
Boilerplate
identity-yixian
I am an Adjunct Professor in Biology at Johns Hopkins and Interim Director of Biosphere Science and Engineering at Carnegie Science.
My research has ranged across cell and developmental biology, biochemistry, genomics, proteomics, and symbiosis, with a recurring focus on integrating heterogeneous evidence to understand biological mechanisms. I have been using the agentic AI platform Zo Computer (https://zo.computer) for my research since last year. My experience leads me to strongly believe that collaboration among multiple AI agents with human supervision can be powerful in biomedical research and medicine.
rfa-concept-testbed
I am assembling a team for NIH Common Fund RFA-RM-27-011 (https://files.simpler.grants.gov/opportunities/5b0a5980-5b84-4cdf-9370-2a36c301760b/attachments/b393021d-fd46-42cb-b1e4-c30273478aad/RFA-RM-27-011-Full-Announcement.html), Development and Testing of a Multi-use Frameworks Playbook for Precision Medicine with AI: Integrating Imaging with Multimodal Data. The application is due October 9 and must develop and test at least two broadly reusable frameworks for responsible multimodal clinical AI.
Our emerging concept is to develop platform-independent playbooks and open-source software for using AI agents to perform deep research across clinical imaging, molecular and clinical data, and published evidence – with comprehensive provenance traces and tooling for error detection, uncertainty reporting, and human oversight.
Ben Guo (https://hraness.com) (copied here) is a former cofounder of Zo Computer (https://www.zo.computer/) and an early engineer (https://www.linkedin.com/in/hraness/) at Stripe. Through XTE Industries LLC, his independent software consulting business, he has been developing a multi-agent platform for collaborative deep research, which could provide a technical foundation for the project.
signature
Best regards,
Yixian Zheng (https://carnegiescience.edu/dr-yixian-zheng), Ph.D. Interim Director and Investigator, Biosphere Science and Engineering, Carnegie Institution for Science Adjunct Professor, Department of Biology, Johns Hopkins University
Email to Alexander Baras
To: baras@jhmi.edu (confirm on the current JHMI directory before sending)
Subject: Possible collaboration on NIH PRIMED-AI Playbook proposal
Dear Dr. Baras,
{{identity}}
{{rfa-concept-testbed}}
Your work in computational pathology, digital pathology, genomic and clinical-data integration, and precision-medicine informatics seems especially relevant. I would value your perspective on whether a pathology-centered use case (linking digital slides with molecular, genomic, and clinical information) could provide a rigorous setting in which to develop and validate these frameworks. If the concept is compelling, I would also like to explore whether you might participate as a clinical co-investigator or MPI.
Would you be available for a 30-minute conversation in the next two weeks?
{{signature}}
Email to Cheng Ting “Tony” Lin
To: clin97@jhmi.edu (confirm on the current JHMI directory before sending)
Subject: Possible collaboration on NIH PRIMED-AI responsible-AI playbook proposal
Dear Dr. Lin,
{{identity}}
{{rfa-concept-testbed}}
Because you lead radiology AI evaluation and governance at Johns Hopkins, I would especially value your perspective on the second framework: how an agentic system should be evaluated, escalated for human review, monitored for data leakage, bias and drift, documented through model and data cards, and prepared for responsible clinical deployment. A radiology transfer test could also determine whether a framework initially developed around a primary clinical use case generalizes across imaging specialties and workflows.
Would you be available for a 30-minute conversation in the next two weeks to assess whether this direction addresses a real clinical need? If there is a strong fit, I would also like to explore whether you might participate as a clinical co-investigator or MPI.
{{signature}}