3 Sources
[1]
Researchers glimpse the inner workings of protein language models
Caption: Understanding what is happening inside the "black box" of large protein models could help researchers to choose better models for a particular task, helping to streamline the process of identifying new drugs or vaccine targets. Within the past few years, models that can predict the
[2]
Researchers glimpse the inner workings of protein language models
Within the past few years, models that can predict the structure or function of proteins have been widely used for a variety of biological applications, such as identifying drug targets and designing new therapeutic antibodies. These models, which are based on large language models (LLMs), can
[3]
MIT technique reveals how AI models predict protein functions
Massachusetts Institute of TechnologyAug 19 2025 Within the past few years, models that can predict the structure or function of proteins have been widely used for a variety of biological applications, such as identifying drug targets and designing new therapeutic antibodies. These models, which
Share
Copy Link
MIT scientists have developed a novel technique to understand how AI protein language models make predictions, potentially streamlining drug discovery and vaccine development processes.
Researchers at the Massachusetts Institute of Technology (MIT) have made a significant breakthrough in understanding the inner workings of protein language models, a type of artificial intelligence used in various biological applications. The study, published in the Proceedings of the National Academy of Sciences, introduces a novel technique to decipher how these models make predictions about protein structure and function
1
.
Source: News-Medical
Protein language models, based on large language models (LLMs), have been widely used in recent years for tasks such as identifying drug targets and designing therapeutic antibodies. While these models can make accurate predictions, they have long been considered "black boxes," with researchers unable to determine how they arrive at their conclusions
2
.The MIT team, led by Bonnie Berger, the Simons Professor of Mathematics and head of the Computation and Biology group at MIT's Computer Science and Artificial Intelligence Laboratory, employed a technique called sparse autoencoders to open up this black box
3
.Sparse autoencoders work by expanding the neural network representation of a protein from a constrained number of neurons (e.g., 480) to a much larger number (e.g., 20,000). This expansion allows the information to "spread out," making it easier to interpret which features each node is encoding
1
.
Source: MIT
To analyze the expanded representations, the researchers utilized an AI assistant called Claude. By comparing the sparse representations with known protein features, Claude could determine which nodes corresponded to specific protein characteristics and describe them in plain English
2
.Related Stories
The study revealed that the features most likely to be encoded by these nodes were protein family and certain functions, including various metabolic and biosynthetic processes. This insight into how protein language models make predictions could have significant implications for biological research and drug development
3
.The first protein language model was introduced in 2018 by Berger and former MIT graduate student Tristan Bepler. Since then, these models have been used for various applications, including predicting viral protein mutations to identify vaccine targets for influenza, HIV, and SARS-CoV-2
1
.By making these models more interpretable, researchers can potentially choose better models for specific tasks, streamlining the process of identifying new drugs or vaccine targets. As Berger notes, "Our work has broad implications for enhanced explainability in downstream tasks that rely on these representations"
2
.Summarized by
Navi
[3]
31 Mar 2025•Science and Research

29 Oct 2025•Science and Research

05 Nov 2024•Science and Research

1
Technology

2
Technology

3
Policy and Regulation
