19 April 2025 · 15 min

What happens when you drop med students into an AI datathon?

No lectures. No theory. Just code, datasets, and real-world healthcare problems.

This week on AI in Medicine, we explore a trainee-led case study where future doctors learned Python, value-based care analytics, and responsible GenAI—all through hands-on data challenges.

These aren’t hackathons for show. They’re how we build a new kind of physician:
🧠 Clinically sharp
💻 Data-literate
🧭 Ethically grounded

🎙️ AI Datathons in Medical Education: A Trainee-Led Case Study

Your company here. This podcast is looking for its first sponsors: reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsorship options and rates →

Transcript

Automated transcript of the audio; it may contain errors.

Host 1: You know, we hear it all the time AI is set to completely change medicine, right? Big promises about diagnosis, treatment, every...

Host 2: Absolutely. It's a huge topic.

Host 1: But uh the real question is how are we actually getting future doctors ready for this, you know, to work with AI? That's what we're really digging into today.

Host 2: Yeah, and we found a really interesting study looking at just that. A practical way to sort of bridge that gap between standard medical school and all of this new AI and machine learning stuff.

Host 1: Right. This paper, Leveraging Datathons to Teach AI in Undergraduate Medical Education: Case Study, it's brand new, 2025, from JMIR Medical Education. And it looks at these two events, these datathons, run by MD++.

Host 2: Which is itself pretty cool. MD++ is a non-profit led entirely by medical trainees.

Host 1: Exactly. So, our mission uh for this deep dive is to unpack these datathons. See how well they actually prepare future docs, maybe even you, listening, with these crucial data skills. Consider this your shortcut, maybe, to understanding this new learning approach.

Host 2: Yeah, and why it's kind of catching on.

Host 1: Okay, so let's start there. Why is it suddenly so, so vital for doctors coming up through training to get AI and ML?

Host 2: Well, it's just becoming embedded, isn't it? Patient care, clinical decisions. AI's influence is growing fast. Future doctors, they really need to understand the basics.

Host 1: Not just to use the tools, but to think critically about them.

Host 2: Exactly. To know the limitations, the biases, when to trust the output and when to question it.

Host 1: And it sounds like maybe traditional medical education hasn't quite caught up yet.

Host 2: That's a big point in the paper. Awareness is growing, sure. But many U.S. medical schools, um they don't really have formal AI training baked into the curriculum.

Host 1: Yeah.

Host 2: Especially not much hands-on experience.

Host 1: Which brings us to datathons. Now, I've heard of hackathons, but datathons, they're different, right?

Host 2: Yeah. Yeah, that's a key distinction they make, drawing on folks like Daneshjou. Hackathons are often about building something new, an app or whatever.

Host 1: Right.

Host 2: Datathons though, are focused on diving into existing data, usually real-world, messy data, and analyzing it, finding clinical insights, maybe using ML models to help. The goal is understanding the data.

Host 1: Okay, analysis over building. And there was other research, um Oyede, showing it's not just tech skills.

Host 2: Right. That's important, too. These things apparently help build teamwork, problem solving, those soft skills that are just as vital.

Host 1: So, it's kind of a package deal learning-wise?

Host 2: You could say that. A more holistic experience than just, you know, sitting through a lecture on algorithms.

Host 1: All right. So, tell us about these specific MD++ datathons. What made them stand out?

Host 2: Well, a few things. First, like we said, run by trainees, for trainees. The whole organizing committee were undergrad medical trainees. That's pretty unique.

Host 1: Peer-led, basically.

Host 2: Exactly. Second, they were totally digital. And not just a weekend dash, they ran for about three weeks, mostly asynchronously.

Host 1: Three weeks online, that's different. Why that format?

Host 2: Accessibility, mainly. Let trainees from anywhere in the U.S. join in without travel, fit it around crazy schedules, different time zones, it just lowers the barrier to entry.

Host 1: Makes sense. And who exactly were they targeting?

Host 2: Specifically undergraduate medical trainees. So med students, residents, grad students in related fields. They actually excluded attending physicians and people already deep into coding.

Host 1: Oh, interesting. Why exclude them?

Host 2: To keep it a level playing field, you know. Make it a safe space for learning where beginners wouldn't feel intimidated or overshadowed. Focus on collaboration among peers.

Host 1: That seems smart. Create the right environment for learning. So, how did it actually work? What was the structure?

Host 2: Pretty straightforward, really. It started with people forming teams, brainstorming project ideas.

Host 1: Okay.

Host 2: Then the main phase was just getting down to work, doing the analysis, building their project. And MD++ used all their existing channels, Slack, newsletter, social media, to get the word out and keep everyone organized.

Host 1: Right. So, what kind of medical problems were they tackling? Did they have specific themes?

Host 2: Yeah. Each year had a different focus. 2023 was all about value-based care, VBC.

Host 1: Okay, VBC. For anyone listening who isn't deep in health care jargon, what's that mean in simple terms?

Host 2: Uh basically, it's shifting payments away from just doing more stuff, more tests, more procedures towards paying for better outcomes and higher-quality care for patients, holding providers accountable for, well, value.

Host 1: Got it. So, tying payment to actually making patients healthier sounds like a place where data analysis would be crucial.

Host 2: Hugely important. Understanding what actually improves outcomes, what's cost effective, data's key.

Host 1: And 2024, what was the theme then?

Host 2: That was responsible generative AI.

Host 1: Ah, so things like ChatGPT and large language models.

Host 2: Exactly. Using AI to create content, maybe draft clinic notes, patient summaries, educational materials. But the key word is responsible.

Host 1: Meaning?

Host 2: Meaning thinking really hard about the ethics, the accuracy, the bias, the privacy implications before we just unleash these tools in hospitals and clinics. Especially important in medicine.

Host 1: Absolutely. Okay, so those were the themes. What about the actual data? What did they get their hands on?

Host 2: For the 2023 VBC datathon, everyone used the MIMIC-IV dataset.

Host 1: Right. I've heard of MIMIC. That's a big one.

Host 2: It's huge. From Beth Israel Deaconess Medical Center. Anonymized data, but real ICU data from like over half a million patient stays between 2008 and 2019.

Host 1: Wow. Half a million. What kind of info is in there?

Host 2: All sorts. ECGs, imaging reports, electronic health records, lab results, outcomes, incredibly rich. They chose it because it's public, it's real, there's already lots of research using it so resources are available, and it has multiple data types.

Host 1: A good testing ground.

Host 2: Definitely. Of course, there were strict rules, data use agreements, training on responsible data handling, that was mandatory. The task was to use this data to propose something actionable related to value-based care.

Host 1: Seems like a serious challenge. Now, for 2024, the generative AI one, did they stick with MIMIC? That sounds almost too big for some projects.

Host 2: That's exactly what the organizers thought. MIMIC is amazing, but its sheer scale and complexity could be uh a bit much, especially for newcomers.

Host 1: Right.

Host 2: So, for 2024 they introduced tracks. A really smart adaptation.

Host 1: Tracks? How did that work?

Host 2: Participants chose one of three specific areas within responsible gen AI, and each track had its own more focused dataset. Made it more manageable.

Host 1: Okay, what were the tracks and datasets?

Host 2: So, there was a clinical documentation track using the MTS-Dialog dataset, that's patient-physician conversation transcripts.

Host 1: Interesting.

Host 2: A medical education track using MedQA, basically practice medical board exam questions.

Host 1: Okay.

Host 2: And a mental health track using data from SuicideWatch and a mental health collection, tagged social media posts.

Host 1: Wow, diverse datasets.

Host 2: Yeah. So, teams picked a track, used that specific dataset, and had freedom within that as long as it related to the responsible gen AI theme.

Host 1: That sounds much more focused. Tackling documentation accuracy, AI for med ed, ethical AI and mental health, all super relevant.

Host 2: Absolutely. Critical areas where gen AI could have a big impact, but needs careful handling.

Host 1: So, they had themes, data, but were they just thrown in, or was there support, scaffolding?

Host 2: Oh, yeah. Definitely support. MD++ set up a dedicated website with all the info, instructions, materials, but importantly, also links to a GitHub repository that had tutorials, example code, Python for both years, R for 2023 only. That's huge for learners.

Host 1: Having code examples makes a massive difference when you're starting out.

Host 2: Totally. Plus, they offered optional workshops, Zoom sessions with data scientists, boot camps for Python and R, presentation skills workshops, talks by physician experts.

Host 1: Quite a lot of resources.

Host 2: Yeah, though, interestingly, they didn't explicitly teach stuff like using GitHub itself or Excel, or general computing skills. The focus was really on the data analysis tools, Python and R.

Host 1: And communication was mainly Slack?

Host 2: Yep, Slack was the hub for questions, announcements, general chatter.

Host 1: Okay. So, teams do their work, analyze the data, build their projects. How did they share it? What was the judging process like?

Host 2: The submission requirements varied slightly. In 2023, it was a written technical report, no length limit, and a five-minute recorded presentation.

Host 1: Right.

Host 2: For 2024, they added a one-page abstract with a key figure along with the full report, maybe to make initial screening easier.

Host 1: And the judging, what were they looking for?

Host 2: The criteria were consistent, though. Statistical rigor, was the analysis sound? Relevance, did it fit the theme? Creativity in the analysis, the visualization, and interestingly, team diversity.

Host 1: Team diversity as an actual judging criterion?

Host 2: Yeah. The idea being that diverse perspectives often lead to better, more innovative solutions. It was explicitly considered.

Host 1: That is interesting. So, how did they pick winners?

Host 2: They reviewed all submissions, picked a set of finalists, and those finalists presented at a final showcase event, digital again.

Host 1: Live Q&A?

Host 2: In 2023, they played the recordings, then did live Q&A. But in 2024, they switched to fully live presentations. The organizers felt that worked better, more dynamic, seemed like judges and audience preferred it, too.

Host 1: And the judges themselves, what kind of background?

Host 2: A real mix. Healthcare execs, clinicians from different specialties, tech product managers, software engineers, AI researchers, broad expertise.

Host 1: Sounds like a thorough process. So, the big question, what was the impact? Did it work? How many people participated?

Host 2: Across the two years, about 200 medical trainees from all over the U.S. took part. The survey responses showed strong interest from folks going into internal medicine, surgery, radiology. Those are the top three.

Host 1: And did they actually learn stuff? What did the follow-up survey show? They got about a 28% response rate, right? 61 people.

Host 2: Yeah. Not a huge response rate, but the results were pretty positive. The big one was a statistically significant jump in self-reported comfort with Python after the datathon.

Host 1: Significant, okay. So, the tutorials likely worked.

Host 2: Seems like it. And importantly, no significant change for skills they didn't teach, like GitHub or Excel. Acts as a sort of control, suggesting that Python learning was specific to the datathon activities.

Host 1: Good point. That boosts confidence in the finding. What about the overall experience?

Host 2: Overwhelmingly positive. 93% enjoyed it, 62% felt they understood the theme better, VBC or responsible AI.

Host 1: That's pretty good.

Host 2: And even better, maybe, 77% felt better able to identify healthcare problems and 82% felt better equipped to generate insights from data. Those are core skills.

Host 1: Yeah, fundamental data literacy. And did they want to do it again?

Host 2: Yep. 82% said they were keen to participate in future datathons. That's a strong endorsement.

Host 1: Wow. Okay, so high enjoyment, self-reported skill gain, better understanding. Sounds successful. Can you give us a taste of the actual projects? The paper mentioned some finalists.

Host 2: Yeah. So, from 2023, the VBC year, one project used ML to spot patients potentially missed for chronic kidney disease diagnosis. Another looked at whether social work referrals could actually cut down hospital readmissions for patients with alcohol issues.

Host 1: Very practical.

Host 2: Yeah, and another tried to predict which heart failure patients were likely to end up in the ICU. Real clinical problems.

Host 1: Definitely. And from 2024, the generative AI year?

Host 2: Equally cool stuff. In clinical documentation, teams looked at using AI to predict why a patient came to the hospital, cost effectiveness...

Host 1: Mhm.

Host 2: Connecting patient language to clinical needs using LLMs.

Host 1: Okay.

Host 2: Med track saw projects using LLMs for board exam prep, and also analyzing if things like age or gender affect how well AI performs clinical reasoning.

Host 1: Important bias questions there.

Host 2: Crucial. And the mental health track had projects looking at risks of fine-tuning LLMs for mental health and building better classifiers for personalized support. Just shows the breadth of application.

Host 1: It really does. Amazing what trainees can do with the right tools and data. So, stepping back, what did the organizers learn logistically from running these? Any key insights?

Host 2: A big one was confirming the value of that extended digital asynchronous format. It just worked much better for busy trainees than a short, intense in-person event. Much more accessible.

Host 1: Flexibility is key.

Host 2: Exactly. And dataset selection was huge. The move to tracks in 2024 to handle MIMIC's complexity while keeping data diversity, that was seen as a real success, tailoring the data challenge to the learning goals.

Host 1: What about the support, like office hours?

Host 2: Funny thing there, the live Zoom office hours in 2023, kind of low turnout.

Host 1: Oh.

Host 2: So, in 2024 they switched to an anonymous Q&A form on Slack. Much higher engagement. People seem more comfortable asking questions asynchronously and maybe anonymously. Easier for mentors to manage, too.

Host 1: Interesting adaptation. Learn and adjust. Any other logistical takeaways?

Host 2: They did note the demographics tended to mirror broader trends in medicine and computer science...

Host 1: Mhm.

Host 2: ...you know, certain groups being over-represented. They flagged that as needing active attention in the future, thinking about barriers for underrepresented trainees.

Host 1: Important point about equity.

Host 2: Definitely. And they mentioned their model, trainee-led, multi-institutional, asynchronous, was a bit different from other datathons out there.

Host 1: Did they acknowledge any study limitations?

Host 2: Oh, yeah. They were clear about that. It was opt-in participation for both the event and the survey, so you always have potential selection bias. People who sign up might already be more motivated, right?

Host 1: True.

Host 2: And the skills assessment was self-reported confidence, not like a formal coding test. But they argued a standardized test would be hard to design for such diverse projects, and might scare people off participating. So, tradeoffs.

Host 1: Understandable. One last thing that jumped at the cost. It seemed incredibly low.

Host 2: Astonishingly low, actually. Averaged out to only about 28 U.S. per registered participant.

Host 1: $28? For a three-week event with data access and workshops? How?

Host 2: Well, the main costs were prize money, some cloud computing resources for teams that needed them, workshop costs, and small honorariums for the judges. A lot relied on sponsorships, especially for compute resources, and significant volunteer effort from the MD++ community for mentoring and workshops.

Host 1: That's impressive efficiency. Makes some model seem very scalable.

Host 2: It really does. Volunteer power and strategic sponsorships go a long way.

Host 1: Okay, so wrapping this all up, what's the bottom line from this deep dive into the MD++ datathons? The main takeaway?

Host 2: I think the main message is that this model, the extended digital, trainee-focused datathon, it really works. It seems to be a highly effective, engaging and cost-effective way to give medical trainees vital hands-on experience with AI and data science.

Host 1: So, it helps bridge that gap we talked about earlier.

Host 2: Exactly. It fosters collaboration, practical skills, applying knowledge to real medical problems. The study makes a strong case that this is a really promising direction for medical education as AI becomes more integrated into healthcare.

Host 1: It really gets you thinking, though. AI is moving so fast. Are datathons enough? What other kinds of innovative learning will we need to make sure tomorrow's doctors are truly ready? It's more than just coding, isn't it? It's that critical thinking, the ethical understanding.

Host 2: That's the big question, right? And maybe the question for you, listening, to ponder.

Host 1: Yeah.

Host 2: How do we ensure genuine readiness? If this specific approach sparked your interest, maybe look up MD++ datathons online, see what else is happening in this space. It's definitely an area to watch.