About the Workshop
Biomedical imaging AI has produced strong methods for segmentation, registration, reconstruction, detection, classification, and report generation. Yet most systems remain organized around fixed inputs and outputs - a structure that fails to capture real clinical workflows, where experts combine images with reports, prior studies, measurements, laboratory values, EHR data, waveforms, pathology, video, and longitudinal patient context.
This workshop focuses on AI agents for biomedical imaging and multimodal clinical data. The central theme is how learning systems can organize multi-step biomedical analyses around images, clinical context, specialized tools, intermediate evidence, and human feedback - for example, systems that retrieve prior studies, invoke segmentation or registration software, measure structures or lesions, integrate reports or EHR variables, compare data across time, and expose intermediate outputs for review.
The workshop brings together researchers in machine learning, computer vision, biomedical imaging, multimodal learning, clinical AI, medical image computing, biomedical informatics, and trustworthy AI to define shared task definitions, evaluation protocols, and benchmarks for agentic image-analysis systems. It is organized around a set of concrete open questions:
- Tool Orchestration: How should agents dynamically invoke, sequence, and verify specialized image-analysis software (segmentation, registration, reconstruction pipelines) in reliable, auditable, and clinically meaningful ways?
- Multimodal Grounding & Active Retrieval: How should rich clinical context (EHR records, lab values, waveforms, prior reports) be actively queried, grounded in, and cross-referenced against pixel-level image evidence?
- Intermediate Verification & Error Recovery: How should intermediate agentic outputs across a multi-step analysis be surfaced, verified, and corrected before downstream steps consume them?
- Agentic Evaluation & Safety: How should full agentic systems be evaluated beyond static final accuracy - focusing on tool execution correctness, trajectory quality, uncertainty calibration, error recovery, and readiness for human oversight?
Topics of Interest
- Agentic Vision-Language & Foundation Models: agentic capabilities, multi-step reasoning trajectories, tool orchestration, and goal execution in medical vision-language and foundation models
- Tool-Using AI Agents: autonomous agents that dynamically select, invoke, sequence, and verify specialized biomedical tools (e.g., segmentation, registration, reconstruction, retrieval, visualization, or statistical analysis)
- Agentic Multimodal Integration: multi-step agentic systems that actively search, retrieve, and cross-reference medical images with clinical reports, EHR records, lab values, waveforms, omics, and longitudinal histories
- Multi-Step Reasoning, Grounding & Planning: goal decomposition, explicit step-by-step reasoning, dynamic plan revision, and evidence grounding for biomedical image interpretation
- Clinical Workflows & Agentic Decision Support: tool-using agents for interactive clinical decision support, image measurement, reporting, dataset annotation, and cohort discovery
- Human-in-the-Loop & Interactive Oversight: agent oversight mechanisms, interactive expert guidance, mid-pipeline verification, and user intervention in multi-step analyses
- Evaluation & Benchmarks for AI Agents: protocols and benchmarks evaluating tool selection accuracy, reasoning trajectories, uncertainty calibration, error recovery, robustness, and safety across multi-step execution
- Trustworthy & Auditable AI Agents: auditable execution logs, interpretability of intermediate agent steps, safety guardrails, and failure recovery in clinical agentic workflows
- Datasets, Environments & Infrastructure: benchmarks, simulated clinical environments, tool API standards, and open-source software frameworks for building and evaluating agentic medical AI
Call for Papers
Non-archival submissions through OpenReview, with an optional archival pathway via the MELBA special issue.
We invite original and innovative work at the intersection of AI agents, medical imaging, tool orchestration, and multimodal clinical data analysis (see topics above). Appropriate submissions include methods papers, dataset/benchmark papers, software resource papers, and position papers.
Scope & Differentiation
To maintain a focused venue distinct from broad medical foundation model or healthcare AI workshops, all submissions must feature an explicit agentic component. Static, single-step models or standard fine-tuning without agentic workflows are out of scope.
Submission Tracks
- Full papers - up to 9 pages, with unlimited references.
- Short papers - up to 4 pages, with unlimited references.
- Papers must use the NeurIPS 2026 paper template and guidelines.
Review and Presentation
- Submissions are made through the workshop's OpenReview venue; each submission will receive three reviews whenever possible.
- Review criteria: relevance to agentic biomedical imaging or multimodal clinical data, technical quality, clarity, evaluation rigor, reproducibility, and appropriateness for workshop discussion.
- Accepted submissions will be presented as posters, demos, or contributed talks.
- Accepted papers will be hosted on OpenReview and linked and hosted directly on this workshop website.
- The workshop is non-archival.
- Organizers follow NeurIPS conflict-of-interest rules and will not assess submissions from conflicted authors.
MELBA Special Issue
The workshop is paired with a special issue of the Machine Learning for Biomedical Imaging (MELBA) journal on AI agents for biomedical imaging and multimodal clinical data. Authors of accepted workshop papers will be invited to submit extended journal-length manuscripts to the special issue. Submission to MELBA is optional and follows MELBA's independent editorial and peer-review process, providing an archival pathway while preserving the non-archival status of workshop submissions.
Important Dates
All deadlines are Anywhere on Earth (AoE, UTC−12).
Invited Speakers
Invited speakers will span AI agents and tool use, multimodal biomedical foundation models, trustworthy machine learning, and expert-guided imaging workflows. The speaker list will be announced soon.
Speaker TBA
AI agents, tool use, and multimodal systems
Speaker TBA
Biomedical imaging and multimodal clinical data
Speaker TBA
Trustworthy machine learning and evaluation
Speaker TBA
Imaging workflows and expert-guided evaluation
Program
A one-day, in-person workshop with invited talks, contributed talks, posters, demos, and a structured discussion session - with substantial time reserved for contributed work and community input.
The structured discussion will divide participants into small groups focused on definitions, benchmark tasks, multimodal datasets, tool-use evaluation, and human oversight. Each group will produce a short set of open problems and recommendations, which the organizers will summarize after the workshop.
Organizers
The organizing team combines expertise in machine learning, biomedical imaging, clinical imaging, multimodal systems, journal leadership, and community building.
Ehsan Adeli, PhD
Stanford University
Biomedical imaging AI, multimodal learning, clinical translation
Tal Arbel, PhD
McGill University
Medical image analysis, machine learning, neuroimaging; MELBA co-founder
Adrian V. Dalca, PhD
MGH / Harvard Medical School / MIT
Biomedical imaging AI, open-source tools; MELBA Executive Editor
Klaus Maier-Hein, PhD
German Cancer Research Center (DKFZ)
Foundation models, benchmarking and validation, self-configuring AI systems
Yixuan Yuan, PhD
The Chinese University of Hong Kong
Medical image analysis, surgical AI, foundation models
Advisory Committee
Curtis Langlotz, MD, PhD
Stanford University
Clinical imaging, radiology AI, clinical workflow integration
Mert Sabuncu, PhD
Cornell University
Radiology AI, biomedical computation; MELBA co-founder and Executive Editor
Program Committee
Approximately 45–60 reviewers across machine learning, biomedical imaging, multimodal learning, clinical AI, evaluation, and trustworthy AI. Review loads are capped at approximately 3 papers per reviewer.
Interested in reviewing for the workshop? Contact us to join the program committee.
Submission
Papers are submitted through the workshop's OpenReview venue.
Note on OpenReview profiles: new profiles created with an institutional email are activated automatically; profiles created without an institutional email go through a moderation process that can take up to two weeks. Please create your OpenReview profile well before the deadline.
Checklist
- Use the NeurIPS 2026 LaTeX template and guidelines.
- Anonymize your submission (double-blind review).
- Full papers: up to 9 pages; short papers: up to 4 pages (unlimited references).
- Accepted papers will be hosted on OpenReview and on this website; submissions are non-archival, and extended versions may later be submitted to the MELBA special issue.
- Submit by August 29, 2026, 11:59 PM AoE.
Contact
A workshop contact email will be announced here soon.