23 December 2025 · 14 min

Breakthrough in Protein Folding

In this episode, we explore the exciting advancements in protein folding with AlphaFold3. Here's what we cover:

Tune in for a deeper dive into how AlphaFold3 is reshaping structural biology and the potential it holds for the future of protein folding research. Don’t forget to subscribe, share, and leave a review!

Your company here. This podcast is looking for its first sponsors: reach clinicians, health-system leaders and medtech and pharma teams following AI in medicine. Sponsorship options and rates →

Transcript

Automated transcript of the audio; it may contain errors.

Host 1: Welcome back to the deep dive. Today, we're getting into the nuts and bolts of what comes next in structural biology. We're looking at AlphaFold 3.

Host 2: It's a big topic. AlphaFold 2 was, and this is no exaggeration, a foundational moment for science.

Host 1: Right. It basically solved the protein folding problem.

Host 2: Exactly. It became the benchmark. But AlphaFold 3 or AF3, uh it wasn't just designed to be a better protein predictor. The ambition was to unify everything.

Host 1: To unify, what does that mean in this context?

Host 2: It means creating a single framework that can model not just proteins, but proteins interacting with, well, with everything else. Nucleic acids like DNA and RNA, ions, small molecules, all in one go.

Host 1: And that's exactly why this deep dive is so important. We're not looking at the press releases, we're digging into a big comprehensive third party benchmark study.

Host 2: A really important one. It tested AF3 across nine different datasets,

Host 1: pitting it against its predecessors, AF2, AlphaFold-Multimer, and even the specialized tools like RoseTTAFold2NA. We really want to find out where the data says the real wins are.

Host 2: And to get that, you first have to understand the change in architecture. The sources are clear that AF3, uh it changed its language and its engine.

Host 1: Okay, language and engine. Let's start with language.

Host 2: So the language part is what they call extended tokenization. It's basically a way for the model to talk about all these different types of molecules, proteins, RNA, ligands, coherently in the same system.

Host 1: And the engine? I read that the famous AF2 structure module is just gone.

Host 2: It is. They swapped it out for something called a diffusion module.

Host 1: A diffusion module. That sounds uh pretty advanced. For someone who actually uses this, what's the practical difference?

Host 2: The practical impact is huge. It comes down to speed and flexibility. AF2 relied really heavily on evolutionary data, these massive multiple sequence alignments or MSAs.

Host 1: Right, it looked for patterns across species.

Host 2: Exactly. The diffusion module though, it takes a different approach. It starts with a cloud of noisy randomized atoms, and then it learns to denoise them into the correct final structure.

Host 1: So, it's less about historical data and more about learning the physics from scratch.

Host 2: In a way, yeah. It simplifies the need for all that evolutionary data, which makes the whole process faster and, you know, potentially much better for proteins that don't have many known relatives.

Host 1: That makes sense. Okay, before we dive into the results, let's just quickly set up the language we'll be using. We need to talk about a few key metrics.

Host 2: Of course.

Host 1: First up is TM-score. This one measures the global similarity, right, the overall shape?

Host 2: That's the one. It tells you if the overall fold is correct. Generally, you know, a score over 0.5 means you're looking at the same basic topology.

Host 1: And then the other one is lDDT. Why is that one so important?

Host 2: Oh, lDDT is vital. It measures the local, fine-grained details. Are the side chains in the right place? Is the geometry of an active site perfect? That's what lDDT tells you.

Host 1: Okay, global and local. Now for when things are interacting, we'll be using DockQ for protein-protein interactions.

Host 2: Yep, it's a quality score for how well two proteins are docked together. Pretty straightforward.

Host 1: And for anything involving DNA or RNA, we have to talk about Fnat.

Host 2: Fnat is critical. It stands for the fraction of native contacts. It tells you if the model got the sticky bits right. Are the specific atoms that are supposed to be touching at the interface actually touching? Fnat measures that.

Host 1: Perfect. So let's start at the baseline, single proteins, monomers. AF2 was already incredible here. Did AF3 move the needle at all?

Host 2: You know, the main takeaway here is that for your average protein, we've pretty much hit a plateau for now.

Host 1: Really?

Host 2: Yeah, the benchmark looked at 150 general protein monomers, and the difference in the global TM-score between AF3 and AF2 was, well, it was negligible, not statistically significant. Both were hovering around 0.81.

Host 1: So the law of diminishing returns is kicking in. But what about that local detail, the lDDT score?

Host 2: Ah, no, that's where we see the first clear, statistically significant win for AF3. It showed a real improvement in local structural accuracy.

Host 1: How much of an improvement are we talking about?

Host 2: Well, the number seems small. The AF3 server got an lDDT of 0.792 versus AF2's 0.787.

Host 1: That does sound tiny. Why should a researcher care about that 0.005 difference?

Host 2: Because function lives in the details. If you're designing a drug that has to fit perfectly into a binding pocket, that tiny difference in local accuracy can be everything. It can be the difference between a model that works for drug design and one that doesn't.

Host 1: Okay, that makes sense. But where the new architecture really seems to matter is with more complex proteins, like multi-domain ones.

Host 2: Absolutely. There's a fantastic example in the benchmark, a phage endolysin.

Host 1: Right.

Host 2: AF2 was brilliant at predicting the individual domains, TM-scores up to 0.99, perfect. But when it tried to assemble the full thing, it completely failed on the orientation between the two domains. The full-length TM-score was a disaster, just 0.562.

Host 1: So it's like having two perfect Lego bricks, but putting them together completely wrong. Why did AF3 fix that?

Host 2: That's the diffusion module showing its strength. AF3 successfully modeled both the individual domains and their correct orientation relative to each other. It scored a 0.916.

Host 1: So the new model is better at handling that kind of inter-domain flexibility.

Host 2: It seems to be, yeah. AF2 was likely constrained by the evolutionary data, forcing a certain arrangement. The diffusion model appears to be better at sampling different possibilities to find the right overall shape.

Host 1: That's a really key difference. Now, on the flip side, what happens when you have no evolutionary info, when you run it in single sequence mode?

Host 2: That's another win for AF3. With no MSA, no templates, AF3's accuracy was clearly superior. It hit a TM-score of 0.413 while AF2 was down at 0.353.

Host 1: So it's better at building from scratch. But for what they call orphan proteins, the ones with a little bit of MSA data, but not much?

Host 2: There the difference was much smaller, not statistically significant. So, to be clear, that evolutionary information is still incredibly important. You can't just throw it away.

Host 1: And one final challenge for proteins, what about the really dynamic ones, the multi-state proteins that change shape to function?

Host 2: That is still a massive, massive challenge. Both AF3 and AF2 really struggle there. The AF3 models tended to just find one stable state and stick with it.

Host 1: So we're not there yet on predicting multiple functional shapes?

Host 2: Not even close. There's a lot of room for improvement in sampling those distinct states.

Host 1: Okay, let's move on to protein complexes or multimers. This was AlphaFold-Multimer's home turf. Did AF3 come out on top?

Host 2: For the global structure, the short answer is no. On average, across 206 multimers, there was no statistically significant difference in the TM-score.

Host 1: So they're basically tied on the overall shape.

Host 2: They are, but I think you can guess the next part of the story.

Host 1: Let me guess. Local accuracy.

Host 2: Precisely. AF3 again showed a small but statistically significant edge in lDDT. But interestingly, for the interface quality, the DockQ score, the improvement was tiny and not statistically significant for general multimers.

Host 1: Okay, but general multimers is a broad category. What if you break it down, say between complexes with identical subunits, homomers, and ones with different subunits, heteromers?

Host 2: And that's where a key difference emerges. For homomers, AF3's main advantage was just that better local structure. But for heteromers,

Host 1: the ones with different partners.

Host 2: Exactly. For those, AF3 achieved a statistically significant improvement in predicting the interface. The DockQ score was up by 9.4% on the server version. That suggests it's better at figuring out the geometry when you have two very different surfaces trying to bind.

Host 1: Now let's talk about the one area where everyone expected AF3 to just blow the doors off, antibody complexes.

Host 2: And it did. This is without a doubt the major breakthrough for AF3 in protein modeling. It is significantly better than AlphaFold-Multimer across every single metric.

Host 1: Every metric?

Host 2: Every single one. Global TM-score was up 8.1%, and crucially, the interface quality, the DockQ score, shot up by nearly 30% on average.

Host 1: 30% is a staggering number. That's huge for anyone in therapeutic antibody design. What's driving that success?

Host 2: It's about getting the most important part right. There's a specific metric for the binding loops on an antibody, the complementary determining regions, or CDRs.

Host 1: The parts that actually do the binding.

Host 2: Yes, and the CDR TM-score for AF3 was, on average, 0.183 compared to AlphaFold-Multimer's...

Host 1: So we're finally getting the molecular handshake right.

Host 2: We are. The benchmark highlights a specific complex, PDB 8DL0. AF3 produced a near perfect model, nailing the CDR loops with a score up to 0.887. AlphaFold-Multimer, it got an okay global score, but it completely failed at the binding site, a CDR TM-score of 0.003. It's a decisive, game-changing win for AF3.

Host 1: Okay, let's pivot. Let's go to the new frontier that AF3 was built for: nucleic acids. Starting with single RNA chains, how does it do against the specialized tools?

Host 2: This is interesting. In terms of the global fold, the TM-score, AF3 is actually in an intermediate position. A specialized tool called RhoFold actually got the highest global score.

Host 1: Wait, hold on. The new unified model didn't win on the global shape of RNA? If another tool gets the overall topology better, why would I use AF3?

Host 2: It's a great question, and the answer comes back to local detail versus the big picture. While AF3 lagged a little bit on the overall fold, it was far superior in local accuracy, lDDT, and in modeling the secondary structure. It got the highest scores there out of all the methods.

Host 1: So it's capturing the internal base pairing and the fine details better, even if the global shape isn't quite as good in some cases.

Host 2: Exactly. For understanding mechanism, that local detail is often what matters most.

Host 1: So the AF3 signature continues: local refinement. What about RNA multimers, predicting RNA-RNA interactions?

Host 2: That is just an incredibly difficult problem. Both AF3 and its main competitor, RoseTTAFold2NA, really struggled. I mean, fewer than 10% of the AF3 predictions even got a correct global fold.

Host 1: Wow.

Host 2: It's a stark reminder of how much harder RNA is than protein. But even there, AF3 did show statistically significant gains in local accuracy over the competition.

Host 1: Okay, let's go to the combination platter, the thing AF3 was really built for: protein-nucleic acid complexes.

Host 2: This is, I would say, AF3's most resounding victory overall. It showed significant superiority to RoseTTAFold2NA. The global TM-score was up by over 14% with big gains in local accuracy, too.

Host 1: And what's the fundamental reason for that win?

Host 2: It's because AF3 is just much, much better at modeling the individual parts, especially the nucleic acids themselves.

Host 1: Ah, so if you get the DNA or RNA structure right, the whole complex comes together more accurately.

Host 2: Precisely. In these complexes, the RNA TM-score was 35% higher, and the DNA TM-score was over 15% higher than what the competitor could do.

Host 1: Is there a specific example that really drives this home?

Host 2: There is. The benchmark looked at a transcription factor, HigAB, binding to DNA, and the key detail is that the protein binding causes a dramatic bend in the DNA.

Host 1: Okay.

Host 2: AF3 modeled that bend almost perfectly, getting a DNA TM-score near 0.6. RoseTTAFold2NA completely missed this critical detail. It saw the DNA as a straight rod. Plus, AF3 nailed the interface contacts, with an Fnat score more than double the competitors.

Host 1: That's compelling. It's the first model that really treats the nucleic acid as an equal partner in the complex. Now let's get practical. Speed, how fast is it?

Host 2: This is an absolute game-changing win. The sources show AF3 is just orders of magnitude faster across the board.

Host 1: Give me some of the big comparisons.

Host 2: Okay. For a standard protein monomer, your local AF2 install might take about 160 minutes. AF3 local, around 20 minutes. For a protein multimer, AF-Multimer takes almost 5 hours. AF3 does it in 31 minutes. But the most dramatic one was for protein-nucleic acid complexes. AF3 took about 23 minutes. RoseTTAFold2NA took nearly 1,000 minutes.

Host 1: A thousand? That's over 16 hours.

Host 2: It is. That kind of speed changes what's possible in a lab overnight. But there are still some roadblocks.

Host 1: What's the biggest one the benchmark found?

Host 2: The biggest issue seems to be its internal ranking system. The analysts found that the server and local versions could give you pretty different results if you only run them once on a hard target.

Host 1: So the model is making good predictions, but it's not always good at picking its own best one.

Host 2: Exactly. If you run it multiple times, you see it sampling from a good distribution of models, but the built-in system that ranks them struggles to consistently pick the winner. This means there's a real need for better ways to estimate model accuracy so the user doesn't have to sift through everything manually.

Host 1: So if we zoom out one last time, what does AF3 tell us about the accuracy gap between the solved and unsolved problems?

Host 2: The gap is still very real. You know, protein predictions are incredibly accurate, with TM-scores around 0.8 to 0.9, but when you move to the new frontiers, the accuracy drops. RNA monomers are only around 0.51, and even the much-improved protein-nucleic acid complexes are still only at about 0.74. The really flexible, complex systems are still the highest mountain to climb.

Host 1: So to synthesize all this, AF3 is a huge step forward. It's fast, it's unified, it's locally more accurate, and it's had true breakthroughs in antibodies and protein-nucleic acid systems.

Host 2: That's right, but the improvements are modest for a lot of standard proteins, and it really highlights how hard things like RNA multimers and multi-state proteins still are. The question has shifted from "can we fold it?" to "can we capture its dynamics?"

Host 1: Right. We've seen protein structure prediction accuracy sort of plateau near that 90% mark, but these highly flexible molecules, like RNA, or proteins that have to adopt multiple shapes to function, they remain incredibly difficult to predict. So if the models, thanks to things like the diffusion module, now have the architecture to handle all kinds of atoms, is the remaining challenge just computational? Do we just need more data and better ranking?

Host 2: Or is it something more fundamental?

Host 1: Or are we fundamentally limited by our understanding of biological dynamics, the constant motion, the subtle interactions that actually define functioning things a single static picture just can't capture? What's it going to take for the next generation of models to reliably map flexibility and movement instead of just a a single snapshot? We'll leave that question with you.