Just glanced, but would be very interesting to look at performance on promiscuous enzymes, and enzymes under diversifying selection for novel substrate activity (eg p450s). Seems like Marks' core point here could be that when stabilizing selection acts on an enzyme behavior, the pLMs pick up on that signature and reject both catalytically implicated and destabilizing substitutions.
I'd also be interested in seeing how this performs with different sequence alignments in the PSSM. Possibly the method works when many related sequences perform the same catalysis and are ingested during PLM training, but then if you broaden phylogenetic scope a bit while making the PSSM you pick up on diversified behaviors with retained folds.
Fully agreed. There's a lot of follow-up work that can be done. Unfortunately the field is currently limited by the available data. There are very few clean datasets where multiple different functions have been measured for a large number of variants of the same protein. We need more of those experiments to make progress.
Thank you for the fantastic post on the paper, really glad you found it interesting! My one factual note is that there actually are distant homologs of DraNramp in the multiple sequence alignment I used that transport other ions, including Mg2+, so I don’t see that particular data point as evidence that you don’t need functional divergence in the multiple sequence alignment. I do think, however, that the TEM-1 beta-lactamase dataset likely represents a genuinely new-to-nature activity, since it is tested against a newer-generation antibiotic, and the method does work there. So I think the broader point still stands that it is interesting that the approach can succeed on an activity that was not present in the evolutionary training distribution.
More generally, I personally put a bit less weight on the stability interpretation than you do here. The PDZ3 dataset is the only one where we have orthogonal stability/abundance measurements, and in that case the specificity-altering mutations have approximately the same distribution of stability as the native-specificity mutations. The PSSM and ESM1v predictions also correlate similarity with stability on that dataset. Meanwhile, inverse-folding models such as ProteinMPNN and ESM-IF1 tend to correlate more strongly with stability than protein language models do, but they do not bias against altered specificity to nearly the same extent.
That makes me lean toward the interpretation that the effect is not primarily the PLM learning stability from context but rather learning something about the native function of the wild-type protein from context. For genuinely new-to-nature activities, my guess would be that the model could still be leveraging similarity to natural substrates (for example, other substrates recognized by related natural sequences could share some chemical similarity with the newer antibiotic). But I definitely agree that the data are limited enough that multiple interpretations are reasonable, and I think we’ll need more datasets that measure both stability and specificity to really sort this out.
Also, the labeling in the panel on the right is correct, I will fix it in the next version :)
Thanks for pointing me to this paper. I hadn't seen it yet. It's a great demonstration of what the current AI tools are good for: Stabilize the protein with AI, and then find new function with experimental evolution.
Just glanced, but would be very interesting to look at performance on promiscuous enzymes, and enzymes under diversifying selection for novel substrate activity (eg p450s). Seems like Marks' core point here could be that when stabilizing selection acts on an enzyme behavior, the pLMs pick up on that signature and reject both catalytically implicated and destabilizing substitutions.
I'd also be interested in seeing how this performs with different sequence alignments in the PSSM. Possibly the method works when many related sequences perform the same catalysis and are ingested during PLM training, but then if you broaden phylogenetic scope a bit while making the PSSM you pick up on diversified behaviors with retained folds.
Fully agreed. There's a lot of follow-up work that can be done. Unfortunately the field is currently limited by the available data. There are very few clean datasets where multiple different functions have been measured for a large number of variants of the same protein. We need more of those experiments to make progress.
Thank you for the fantastic post on the paper, really glad you found it interesting! My one factual note is that there actually are distant homologs of DraNramp in the multiple sequence alignment I used that transport other ions, including Mg2+, so I don’t see that particular data point as evidence that you don’t need functional divergence in the multiple sequence alignment. I do think, however, that the TEM-1 beta-lactamase dataset likely represents a genuinely new-to-nature activity, since it is tested against a newer-generation antibiotic, and the method does work there. So I think the broader point still stands that it is interesting that the approach can succeed on an activity that was not present in the evolutionary training distribution.
More generally, I personally put a bit less weight on the stability interpretation than you do here. The PDZ3 dataset is the only one where we have orthogonal stability/abundance measurements, and in that case the specificity-altering mutations have approximately the same distribution of stability as the native-specificity mutations. The PSSM and ESM1v predictions also correlate similarity with stability on that dataset. Meanwhile, inverse-folding models such as ProteinMPNN and ESM-IF1 tend to correlate more strongly with stability than protein language models do, but they do not bias against altered specificity to nearly the same extent.
That makes me lean toward the interpretation that the effect is not primarily the PLM learning stability from context but rather learning something about the native function of the wild-type protein from context. For genuinely new-to-nature activities, my guess would be that the model could still be leveraging similarity to natural substrates (for example, other substrates recognized by related natural sequences could share some chemical similarity with the newer antibiotic). But I definitely agree that the data are limited enough that multiple interpretations are reasonable, and I think we’ll need more datasets that measure both stability and specificity to really sort this out.
Also, the labeling in the panel on the right is correct, I will fix it in the next version :)
Thanks for your response and clarification. I have added a footnote pointing out that there are natural DraNramp homologs that can import magnesium.
Fascinating! Sounds consistent with this recent paper https://www.nature.com/articles/s41586-026-10820-0
Thanks for pointing me to this paper. I hadn't seen it yet. It's a great demonstration of what the current AI tools are good for: Stabilize the protein with AI, and then find new function with experimental evolution.