Overview
How do agentic and non-agentic multi-hop pipelines differ in medical QA accuracy, evidence use, and reliability?
System design
- Built an agentic Qwen2.5 workflow that iteratively searched Wikipedia
- Let the model decide whether to continue gathering evidence or answer, with a 15-step cap
- Built controlled Wikipedia and Wikipedia + PubMed comparison pipelines
- Examined how tool use and retrieval strategy affected answer quality
Results
- Published the results at BioCreative IX / IJCAI 2025
- Produced a side-by-side analysis of agentic and non-agentic multi-hop system behavior