esm:含ESM-2、ESMFold等预训练模型,支持蛋白质结构预测与设计

提供Transformer蛋白质语言模型及预训练权重,涵盖结构预测、变体效应预测、逆折叠等功能,含元基因组蛋白质结构图谱,助力蛋白质设计与研究。【此简介由AI生成】

分支7Tags12
当前项目代码仓暂无内容

进化尺度建模

atlas

2023年4月更新: 双重预印本关于蛋白质设计的代码现已发布! "Language models generalize beyond natural proteins" 的代码位于 examples/lm-design/。 "A high-level programming language for generative protein design" 的代码位于 examples/protein-programming-language/

此代码库包含来自 Meta 基础人工智能研究蛋白质团队(FAIR)的 Transformer 蛋白质语言模型 的代码和预训练权重,包括我们最先进的 ESM-2ESMFold,以及 MSA TransformerESM-1v 用于预测变异效应和 ESM-IF1 用于逆折叠。Transformer 蛋白质语言模型在 2019年预印本 的论文 "Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences" 中被引入。ESM-2 在各种结构预测任务中均优于所有测试过的单一序列蛋白质语言模型。ESMFold 利用 ESM-2 语言模型直接从蛋白质序列生成准确的结构预测。

2022年11月,我们发布了 ESM Metagenomic Atlasv0 版本,这是一个包含6.17亿个预测的宏基因组蛋白质结构的开放图谱。 该图谱在2023年3月与 EBI 合作更新。新的 v2023_02 版本为图谱增加了1.5亿个预测结构,以及预计算的 ESM2 嵌入。批量下载、博客文章以及图谱网站上提供的资源在 本 README 中有文档说明。

2022年12月,我们同时发布了两个关于蛋白质设计的预印本。

  • "Language models generalize beyond natural proteins" (论文, 代码) 使用 ESM2 设计新型蛋白质。与预印本相关的代码和数据可以在 这里 找到。
  • "A high-level programming language for generative protein design" (论文, 代码) 使用 ESMFold 根据高级编程语言设计蛋白质。
引用 对于 ESM2,ESMFold 和 ESM 图谱: ```bibtex @article{lin2023evolutionary, title = {Evolutionary-scale prediction of atomic-level protein structure with a language model}, author = {Zeming Lin and Halil Akin and Roshan Rao and Brian Hie and Zhongkai Zhu and Wenting Lu and Nikita Smetanin and Robert Verkuil and Ori Kabeli and Yaniv Shmueli and Allan dos Santos Costa and Maryam Fazel-Zarandi and Tom Sercu and Salvatore Candido and Alexander Rives }, journal = {Science}, volume = {379}, number = {6637}, pages = {1123-1130}, year = {2023}, doi = {10.1126/science.ade2574}, URL = {https://www.science.org/doi/abs/10.1126/science.ade2574}, note={早期版本作为预印本:bioRxiv 2022.07.20.500902}, } ```

对于 Transformer 蛋白质语言模型:

@article{rives2021biological,
  title={Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences},
  author={Rives, Alexander and Meier, Joshua and Sercu, Tom and Goyal, Siddharth and Lin, Zeming and Liu, Jason and Guo, Demi and Ott, Myle and Zitnick, C Lawrence and Ma, Jerry and others},
  journal={Proceedings of the National Academy of Sciences},
  volume={118},
  number={15},
  pages={e2016239118},
  year={2021},
  publisher={National Acad Sciences},
  note={bioRxiv 10.1101/622803},
  doi={10.1073/pnas.2016239118},
  url={https://www.pnas.org/doi/full/10.1073/pnas.2016239118},
}

目录

最近更新

主要模型推荐

缩写 esm.pretrained. 数据集 描述
ESM-2 esm2_t36_3B_UR50D() esm2_t48_15B_UR50D() UR50(样本 UR90) SOTA 通用蛋白质语言模型。可以直接从单个序列预测结构、功能和蛋白质的其他属性。随 Lin et al. 2022(2022年8月更新)发布。
ESMFold esmfold_v1() PDB + UR50 端到单序列3D结构预测器(2022年11月更新)。
ESM-MSA-1b esm_msa1b_t12_100M_UR50S() UR50 + MSA MSA Transformer 语言模型。可以从 MSA 中提取嵌入。支持 SOTA 推理结构。随 Rao et al. 2021(ICML'21 版本,2021年6月)发布。
ESM-1v esm1v_t33_650M_UR90S_1() ... esm1v_t33_650M_UR90S_5() UR90 专用于变异效应预测的语言模型。支持 SOTA 零样本功能效应预测。与 ESM-1b 具有相同架构,但在 UniRef90 上训练。随 Meier et al. 2021 发布。
ESM-IF1 esm_if1_gvp4_t16_142M_UR50() CATH + UR50 逆折叠模型。可以用于为给定结构设计序列,或预测给定结构序列变化的的功能效应。支持 SOTA 固定骨架序列设计。随 Hsu et al. 2022 发布。

有关可用模型的完整列表、详细信息及发布说明,请参阅 预训练模型

使用指南

快速入门

一种简单的方式是通过 HuggingFace transformers 库,它简化了 ESMFold 的依赖关系,并提供了一个标准的 API 和工具来使用最先进的预训练模型。

另外,ColabFold 已集成 ESMFold,因此您可以直接在 Google Colab 实例的浏览器中运行它。

我们还提供了一个 API,您可以通过 curl 或在 ESM 元基因组图谱网页 上访问。

curl -X POST --data "KVFGRCELAAAMKRHGLDNYRGYSLGNWVCAAKFESNFNTQATNRNTDGSTDYGILQINSRWWCNDGRTPGSRNLCNIPCSALLSSDITASVNCAKKIVSDGNGMNAWVAWRNRCKGTDVQAWIRGCRL" https://api.esmatlas.com/foldSequence/v1/pdb/

对于 ESM-MSA-1b、ESM-IF1 或其他模型,您可以直接通过以下说明,使用我们仓库中的原始实现。

开始使用本仓库

作为先决条件,您必须安装 PyTorch 才能使用此仓库。

您可以使用以下一行代码进行安装,使用 esm 的最新发布版:

pip install fair-esm  # latest release, OR:
pip install git+https://github.com/facebookresearch/esm.git  # bleeding edge, current repo main branch

要使用 ESMFold 模型,请确保从具有 python 版本小于等于 3.9 且已安装 PyTorch 的环境开始。然后,在您的 pip 安装命令中加入 [esmfold] 选项,这将自动安装 OpenFold 的依赖项。安装 OpenFold 需要使用 nvcc

pip install "fair-esm[esmfold]"
# OpenFold and its remaining dependency
pip install 'dllogger @ git+https://github.com/NVIDIA/dllogger.git'
pip install 'openfold @ git+https://github.com/aqlaboratory/openfold.git@4b41059694619831a7db195b7e0988fc4ff3a307'

注意:如果 openfold 安装失败,请仔细检查 nvcc 是否可用,并且已经安装了与 CUDA 兼容版本的 PyTorch。

或者,我们提供了 esmfold 的 conda 环境,可以通过 conda env create -f environment.yml 命令来构建。

我们还支持 PyTorch Hub,这样您就不需要自己克隆或安装这个代码库:

import torch
model, alphabet = torch.hub.load("facebookresearch/esm:main", "esm2_t33_650M_UR50D")

在执行pip安装之后,您可以按照以下方式加载并使用预训练模型:

import torch
import esm

# Load ESM-2 model
model, alphabet = esm.pretrained.esm2_t33_650M_UR50D()
batch_converter = alphabet.get_batch_converter()
model.eval()  # disables dropout for deterministic results

# Prepare data (first 2 sequences from ESMStructuralSplitDataset superfamily / 4)
data = [
    ("protein1", "MKTVRQERLKSIVRILERSKEPVSGAQLAEELSVSRQVIVQDIAYLRSLGYNIVATPRGYVLAGG"),
    ("protein2", "KALTARQQEVFDLIRDHISQTGMPPTRAEIAQRLGFRSPNAAEEHLKALARKGVIEIVSGASRGIRLLQEE"),
    ("protein2 with mask","KALTARQQEVFDLIRD<mask>ISQTGMPPTRAEIAQRLGFRSPNAAEEHLKALARKGVIEIVSGASRGIRLLQEE"),
    ("protein3",  "K A <mask> I S Q"),
]
batch_labels, batch_strs, batch_tokens = batch_converter(data)
batch_lens = (batch_tokens != alphabet.padding_idx).sum(1)

# Extract per-residue representations (on CPU)
with torch.no_grad():
    results = model(batch_tokens, repr_layers=[33], return_contacts=True)
token_representations = results["representations"][33]

# Generate per-sequence representations via averaging
# NOTE: token 0 is always a beginning-of-sequence token, so the first residue is token 1.
sequence_representations = []
for i, tokens_len in enumerate(batch_lens):
    sequence_representations.append(token_representations[i, 1 : tokens_len - 1].mean(0))

# Look at the unsupervised self-attention map contact predictions
import matplotlib.pyplot as plt
for (_, seq), tokens_len, attention_contacts in zip(data, batch_lens, results["contacts"]):
    plt.matshow(attention_contacts[: tokens_len, : tokens_len])
    plt.title(seq)
    plt.show()

ESMFold 结构预测

在安装了 [esmfold] 选项之后,您可以按照以下方式使用 ESMFold 结构预测模型:

import torch
import esm

model = esm.pretrained.esmfold_v1()
model = model.eval().cuda()

# Optionally, uncomment to set a chunk size for axial attention. This can help reduce memory.
# Lower sizes will have lower memory requirements at the cost of increased speed.
# model.set_chunk_size(128)

sequence = "MKTVRQERLKSIVRILERSKEPVSGAQLAEELSVSRQVIVQDIAYLRSLGYNIVATPRGYVLAGG"
# Multimer prediction can be done with chains separated by ':'

with torch.no_grad():
    output = model.infer_pdb(sequence)

with open("result.pdb", "w") as f:
    f.write(output)

import biotite.structure.io as bsio
struct = bsio.load_structure("result.pdb", extra_fields=["b_factor"])
print(struct.b_factor.mean())  # this will be the pLDDT
# 88.3

除了我们推荐的最佳性能模型 esm.pretrained.esmfold_v1() 外,我们还提供了 esm.pretrained.esmfold_v0(),该模型在 Lin 等人 2022 的实验中使用。

我们还提供了一个命令行界面 (esm-fold),它能够利用 ESMFold 高效地从 FASTA 文件批量预测蛋白质结构:

usage: esm-fold [-h] -i FASTA -o PDB [--num-recycles NUM_RECYCLES]
                [--max-tokens-per-batch MAX_TOKENS_PER_BATCH]
                [--chunk-size CHUNK_SIZE] [--cpu-only] [--cpu-offload]

optional arguments:
  -h, --help            show this help message and exit
  -i FASTA, --fasta FASTA
                        Path to input FASTA file
  -o PDB, --pdb PDB     Path to output PDB directory
  --num-recycles NUM_RECYCLES
                        Number of recycles to run. Defaults to number used in
                        training (4).
  --max-tokens-per-batch MAX_TOKENS_PER_BATCH
                        Maximum number of tokens per gpu forward-pass. This
                        will group shorter sequences together for batched
                        prediction. Lowering this can help with out of memory
                        issues, if these occur on short sequences.
  --chunk-size CHUNK_SIZE
                        Chunks axial attention computation to reduce memory
                        usage from O(L^2) to O(L). Equivalent to running a for
                        loop over chunks of of each dimension. Lower values
                        will result in lower memory usage at the cost of
                        speed. Recommended values: 128, 64, 32. Default: None.
  --cpu-only            CPU only
  --cpu-offload         Enable CPU offloading

该命令将为fasta文件中的每个序列进行一次预测。可以预测多聚体,并且应当在fasta文件中将它们作为一个单独的序列输入,链之间用":"字符分隔。

默认情况下,预测将进行批量处理,以便短序列可以同时进行预测。通过设置--max-tokens-per-batch=0可以禁用此功能。批量处理可以显著提高短序列的预测速度。

--cpu-offload 标志对于长序列的预测可能很有用。它将尝试将某些参数移至CPU内存,而不是存储在GPU上。

最后,不同大小语言模型(LMs)的消融实验结果已发布为 esm.pretrained.esmfold_structure_module_only_*()。我们不推荐将这些模型用于结构预测。

批量计算 FASTA 文件中的嵌入向量

我们提供了一个命令行界面(esm-extract),用于高效地从ESM中为FASTA文件批量提取嵌入向量:

usage: esm-extract [-h] [--toks_per_batch TOKS_PER_BATCH]
                   [--repr_layers REPR_LAYERS [REPR_LAYERS ...]] --include
                   {mean,per_tok,bos,contacts}
                   [{mean,per_tok,bos,contacts} ...]
                   [--truncation_seq_length TRUNCATION_SEQ_LENGTH]
                   model_location fasta_file output_dir

Extract per-token representations and model outputs for sequences in a FASTA
file

positional arguments:
  model_location        PyTorch model file OR name of pretrained model to
                        download (see README for models)
  fasta_file            FASTA file on which to extract representations
  output_dir            output directory for extracted representations

optional arguments:
  -h, --help            show this help message and exit
  --toks_per_batch TOKS_PER_BATCH
                        maximum batch size
  --repr_layers REPR_LAYERS [REPR_LAYERS ...]
                        layers indices from which to extract representations
                        (0 to num_layers, inclusive)
  --include {mean,per_tok,bos,contacts} [{mean,per_tok,bos,contacts} ...]
                        specify which representations to return
  --truncation_seq_length TRUNCATION_SEQ_LENGTH
                        truncate sequences longer than the given value

以下命令允许从ESM-2模型中提取FASTA文件的最终层嵌入向量:

esm-extract esm2_t33_650M_UR50D examples/data/some_proteins.fasta \
  examples/data/some_proteins_emb_esm2 --repr_layers 0 32 33 --include

当然,我会按照您的要求进行翻译。不过,似乎您忘记提供需要翻译的文本内容。请提供原文,我将会将其翻译成中文,同时保持 Markdown 格式和您所要求的风格。

python scripts/extract.py esm2_t33_650M_UR50D examples/data/some_proteins.fasta \
  examples/data/some_proteins_emb_esm2 --repr_layers 0 32 33 --include mean per_tok

CUDA 设备是可选的,系统将自动检测。

目录 some_proteins_emb_esm2/ 现在包含每个 FASTA 序列对应的 .pt 文件;请使用 torch.load() 来加载它们。 scripts/extract.py 脚本包含确定 .pt 文件内容的标志:

  • --repr-layers(默认:仅最终层)选择包含哪些层的嵌入。
  • --include 指定要保存的嵌入。您可以使用以下选项:
    • per_tok 包含完整序列,每个氨基酸都有一个嵌入(序列长度 x 隐藏维度)。
    • mean 包含每个层对完整序列的嵌入平均值。
    • bos 包含序列起始标记的嵌入。 (注意:不要与预训练模型一起使用 - 我们在训练过程中没有使用 bos-token 监督)

大模型推理的 CPU 卸载

如果您想在您的机器上加载像 15B 这样非常大的模型,或者对长序列进行推理,常规的 GPU 推理可能会导致内存溢出错误。 我们展示了如何使用 Fairscale 的 完全分片数据并行(FSDP) 加载模型,并使用其 CPU 卸载功能。 这允许在单个 GPU 上对大型模型进行推理。 请查看 examples/esm2_infer_fairscale_fsdp_cpu_offloading.py 以获取更多详细信息。

零样本变异预测

有关 ESM-1v 模型的代码和预训练权重,请参见 "examples/variant-prediction/",该模型在 Language models enable zero-shot prediction of the effects of mutations on protein function. (Meier et al. 2021) 一文中描述。

请注意,ESM-2 也可以用于变异预测,并且预计将具有与 ESM-1v 相似的性能。

反向折叠

请参见 "examples/inverse_folding/" 获取详细用户指南。ESM-IF1 模型在 Learning inverse folding from millions of predicted structures. (Hsu et al. 2022) 一文中被描述为 GVPTransformer

我们还提供了一个 colab 笔记本,用于序列设计和序列评分功能。

ESM-IF1 反向折叠模型旨在根据蛋白质骨架原子坐标预测蛋白质序列。 我们提供了以下脚本:1)为给定结构采样序列设计;2)为给定结构评分。

在 AlphaFold2 预测的 12M 个蛋白质结构上进行训练,ESM-IF1 模型包含不变几何输入处理层,后接一个序列到序列的变压器,并在结构保留的骨架上实现了 51% 的本征序列恢复率,对于埋藏残基的恢复率为 72%。 该模型还通过跨度遮罩进行训练,以容忍缺失的骨架坐标,因此可以预测部分遮罩结构的序列。

为给定结构采样序列设计

环境设置的说明请参见 "examples/inverse_folding" 的相关小节。

要为给定结构的 PDB 或 mmCIF 格式采样序列,请使用 sample_sequences.py 脚本。输入文件可以是 .pdb.cif 后缀。

例如,为了采样 3 个高尔基酪蛋白激酶结构(PDB 5YH22022年1月 PDB 分子月度)的序列设计,我们可以从 esm 根目录运行以下命令:

python examples/inverse_folding/sample_sequences.py examples/inverse_folding/data/5YH2.pdb \
  --chain C --temperature 1 --num-samples 3 --outpath examples/inverse_folding/output/sampled_sequences.fasta

采样序列将以fasta格式保存到指定的输出文件中。

温度参数控制序列采样概率分布的尖锐程度。较高的采样温度会产生更多样化的序列,但可能导致原生序列的恢复率较低。默认的采样温度为1。为了优化原生序列的恢复,我们建议在较低的温度下进行采样,例如1e-6。

序列评分

要计算给定结构条件下序列的条件对数似然性得分,请使用score_log_likelihoods.py脚本。

例如,要依据examples/inverse_folding/data/5YH2.pdb中的结构对examples/inverse_folding/data/5YH2_mutated_seqs.fasta中的序列进行评分,我们可以在esm根目录下运行以下命令:

python examples/inverse_folding/score_log_likelihoods.py examples/inverse_folding/data/5YH2.pdb \
  examples/inverse_folding/data/5YH2_mutated_seqs.fasta --chain C \
  --outpath examples/inverse_folding/output/5YH2_mutated_seqs_scores.csv

条件对数似然性已保存为 CSV 格式,并存放于指定的输出路径中。 输出值是序列中所有氨基酸的平均对数似然性的平均值。

更多信息,请参阅 "./examples/inverse_folding/" 以获取详细用户指南。

ESM 宏基因组图谱

请访问 ESM 宏基因组图谱 网站,并查看我们的 博客文章 以了解更多信息。

批量下载指南在独立的 README 这里 提供说明。

图谱资源包括一个页面,用于使用 ESMFold 折叠序列,通过 结构序列 搜索 ESM 图谱的子集,以及一个 API 用于以编程方式访问这些资源。

Foldseek 提供了对图谱的无长度限制搜索 这里

笔记本

反向折叠 - 基于骨架结构预测或评分序列

ESM-IF1 反向折叠模型根据蛋白质骨架原子坐标预测蛋白质序列,该模型由 AlphaFold2 预测的 1200 万个蛋白质结构进行训练。 本教程指导您通过采样序列、计算条件对数似然性以及提取编码器输出作为结构表示的示例。

监督变异预测 - 在嵌入向量上训练分类器

为了帮助您开始使用嵌入向量,这个 Jupyter 笔记本教程 展示了如何使用来自 ESM-1 的嵌入向量训练一个监督变异预测器。 您可以采用类似协议训练用于任何下游任务模型,即使数据有限也可。 首先,您可以按照笔记本中的说明通过 下载预计算的嵌入向量,或者通过运行以下命令来获取 examples/data/P62593.fasta 的嵌入向量:

# Obtain the embeddings
python scripts/extract.py esm1v_t33_650M_UR90S_1 examples/data/P62593.fasta \
  examples/data/P62593_emb_esm1v --repr_layers 33 --include mean

然后,请按照教程中的剩余说明操作。你还可以在 colab 笔记本 中运行教程。

注意,或者使用 针对零样本变异预测的最新说明 该说明在不进行任何监督训练的情况下预测突变效应。

无监督接触预测

这个 Jupyter 笔记本教程 展示了如何使用 ESM-2 和 MSA Transformer (ESM-MSA-1) 模型进行接触预测。 接触预测基于对模型注意力图的逻辑回归。 这种方法是基于我们的 ICLR 2021 论文, Transformer 蛋白质语言模型是无监督结构学习者。(Rao et al. 2020) MSA Transformer (ESM-MSA-1) 接受多重序列比对(MSA)作为输入,并使用相同的方式连接行自注意力图。 参见 MSA Transformer。(Rao et al. 2021)

要获取无监督注意力基接触,请调用 model.predict_contacts(tokens)model(tokens, return_contacts=True)

ESMStructuralSplitDataset 和自注意力接触预测

而这个 Jupyter 笔记本教程 展示了如何加载和索引 ESMStructuralSplitDataset, 并使用 ESM-2 计算无监督自注意力图接触预测。

可用模型和数据集

预训练模型

缩写 esm.pretrained. 层数 参数量 数据集 嵌入维度 模型 URL(自动下载到 ~/.cache/torch/hub/checkpoints
ESM-2 esm2_t48_15B_UR50D 48 15B UR50/D 2021_04 5120 https://dl.fbaipublicfiles.com/fair-esm/models/esm2_t48_15B_UR50D.pt
esm2_t36_3B_UR50D 36 3B UR50/D 2021_04 2560 https://dl.fbaipublicfiles.com/fair-esm/models/esm2_t36_3B_UR50D.pt
esm2_t33_650M_UR50D 33 650M UR50/D 2021_04 1280 https://dl.fbaipublicfiles.com/fair-esm/models/esm2_t33_650M_UR50D.pt
esm2_t30_150M_UR50D 30 150M UR50/D 2021_04 640 https://dl.fbaipublicfiles.com/fair-esm/models/esm2_t30_150M_UR50D.pt
esm2_t12_35M_UR50D 12 35M UR50/D 2021_04 480 https://dl.fbaipublicfiles.com/fair-esm/models/esm2_t12_35M_UR50D.pt
esm2_t6_8M_UR50D 6 8M UR50/D 2021_04 320 https://dl.fbaipublicfiles.com/fair-esm/models/esm2_t6_8M_UR50D.pt
ESMFold esmfold_v1 48 (+36) 690M (+3B) UR50/D 2021_04 - https://dl.fbaipublicfiles.com/fair-esm/models/esmfold_3B_v1.pt
esmfold_v0 48 (+36) 690M (+3B) UR50/D 2021_04 - https://dl.fbaipublicfiles.com/fair-esm/models/esmfold_3B_v0.pt
esmfold_structure_module_only_* 0 (+various) various UR50/D 2021_04 - https://dl.fbaipublicfiles.com/fair-esm/models/esmfold_structure_module_only_*
ESM-IF1 esm_if1_gvp4_t16_142M_UR50 20 124M CATH 4.3 + predicted structures for UR50 512 https://dl.fbaipublicfiles.com/fair-esm/models/esm_if1_gvp4_t16_142M_UR50.pt
ESM-1v esm1v_t33_650M_UR90S_[1-5] 33 650M UR90/S 2020_03 1280 https://dl.fbaipublicfiles.com/fair-esm/models/esm1v_t33_650M_UR90S_1.pt
ESM-MSA-1b esm_msa1b_t12_100M_UR50S 12 100M UR50/S + MSA 2018_03 768 https://dl.fbaipublicfiles.com/fair-esm/models/esm_msa1b_t12_100M_UR50S.pt
ESM-MSA-1 esm_msa1_t12_100M_UR50S 12 100M UR50/S + MSA 2018_03 768 https://dl.fbaipublicfiles.com/fair-esm/models/esm_msa1_t12_100M_UR50S.pt
ESM-1b esm1b_t33_650M_UR50S 33 650M UR50/S 2018_03 1280 https://dl.fbaipublicfiles.com/fair-esm/models/esm1b_t33_650M_UR50S.pt
ESM-1 esm1_t34_670M_UR50S 34 670M UR50/S 2018_03 1280 https://dl.fbaipublicfiles.com/fair-esm/models/esm1_t34_670M_UR50S.pt
esm1_t34_670M_UR50D 34 670M UR50/D 2018_03 1280 https://dl.fbaipublicfiles.com/fair-esm/models/esm1_t34_670M_UR50D.pt
esm1_t34_670M_UR100 34 670M UR100 2018_03 1280 https://dl.fbaipublicfiles.com/fair-esm/models/esm1_t34_670M_UR100.pt
esm1_t12_85M_UR50S 12 85M UR50/S 2018_03 768 https://dl.fbaipublicfiles.com/fair-esm/models/esm1_t12_85M_UR50S.pt
esm1_t6_43M_UR50S 6 43M UR50/S 2018_03 768 https://dl.fbaipublicfiles.com/fair-esm/models/esm1_t6_43M_UR50S.pt

以下是按时间顺序发布的模型及其对应的论文:

缩写 发行说明
ESM-1 与 Rives et al. 2019(2020 年 8 月更新)一起发布。
ESM-1b 与 Rives et al. 2019(2020 年 12 月更新)一起发布。参见附录 B。
ESM-MSA-1 与 Rao et al. 2021(预印本 v1)一起发布。
ESM-MSA-1b 与 Rao et al. 2021(ICML'21 版本,2021 年 6 月)一起发布。
ESM-1v 与 Meier et al. 2021 一起发布。
ESM-IF1 与 Hsu et al. 2022 一起发布。
ESM-2 与 Lin et al. 2022 一起发布。

ESM 结构分割数据集

这是一个蛋白质域结构的五折交叉验证数据集,可用于衡量表示在不同结构相似性级别上的泛化能力。 数据集在家族、超家族和折叠级别实现了结构保留。SCOPe 数据库用于对域进行分类。对于每个结构保留级别, 域被独立地分为 5 个相等的集合,即五个折叠集、超家族集或家族集。这确保了对于每个分区, 具有相同分类的结构不会同时出现在训练集和测试集中。对于给定的分类级别,每个结构在测试集中出现一次, 以便在交叉验证实验中,每个结构将正好被评估一次。

数据集提供了 3D 坐标、距离图和二级结构标签。 有关数据集构建的更多详细信息,请参见 Rives et al. 2019 附录 A.10。

这个 Jupyter 笔记本教程 展示了如何加载和索引 ESMStructuralSplitDataset

ESMStructuralSplitDataset 在初始化时将下载 splitspkl。 我们还为每个域提供了 msas。数据可以直接下载如下。

名称 描述 URL
splits 训练/验证划分 https://dl.fbaipublicfiles.com/fair-esm/structural-data/splits.tar.gz
pkl 包含序列、SSP 标签、距离图和 3D 坐标的 pkl 对象 https://dl.fbaipublicfiles.com/fair-esm/structural-data/pkl.tar.gz
msas 包含每个域 MSA 的 a3m 文件 https://dl.fbaipublicfiles.com/fair-esm/structural-data/msas.tar.gz

预训练数据集拆分

Rives et al. 2019Rao et al. 2021 中作为预训练保留验证集使用的 UniRef50 簇的拆分文件可以在此找到:

这些文件仅包含与 UniRef 数据库,2018-03 版本 对应的 UniRef50 ID 和 UniRef100 ID,该数据库由 UniProt 联盟发布,并遵循 知识共享署名(CC BY 4.0)许可证

与相关工作对比

任务 无监督接触预测 结构预测
测试集 Large valid CASP14 CAMEO (Apr-Jun 2022) CASP14 CAMEO (Apr-Jun 2022)
Gremlin (Potts) 39.3
TAPE 11.2
ProtBert-BFD 34.1
Prot-T5-XL-BFD 35.6 46.1 62.6
Prot-T5-XL-Ur50 (3B) 47.9 49.8 69.4
ESM-1 33.7
ESM-1b 41.1 24.4 39 41.6 64.5
ESM-1v 35.3
ESM-MSA-1b 57.4
ESM-2 (8M) 15.9 9.8 15.7 36.7 48.1
ESM-2 (35M) 28.8 16.4 28.4 41.4 56.4
ESM-2 (150M) 42.2 26.8 40.1 49.0 64.9
ESM-2 (700M) 50.1 32.5 47.6 51.3 70.1
ESM-2 (3B) 52.7 34.0 49.9 52.5 71.8
ESM-2 (15B) 54.5 37.0 51.7 55.4 72.1

与相关蛋白语言模型在结构预测任务上的对比。

  • 所有接触数均为 top-L,LR 精度指标,其中长距离意味着序列分隔至少为 24 个残基
  • 对于无监督接触预测,使用注意力头的稀疏线性组合直接预测蛋白接触,通过在 20 个结构上拟合逻辑回归进行训练。 更多关于该方法的信息,请参见 Rao et al. 2020
  • 对于结构预测,直接从冻结的语言模型嵌入中训练 AlphaFold2 结构模块。 更多关于该方法的信息,请参见 Lin et al. 2022
  • 直接耦合分析方法(Gremlin、mfDCA、Psicov)和 ESM-MSA-1 使用 trRosetta MSAs,而其他方法从单一序列进行预测。

引用

如果您在研究中发现模型有用,我们请求您引用相关论文:

@article{rives2019biological,
  author={Rives, Alexander and Meier, Joshua and Sercu, Tom and Goyal, Siddharth and Lin, Zeming and Liu, Jason and Guo, Demi and Ott, Myle and Zitnick, C. Lawrence and Ma, Jerry and Fergus, Rob},
  title={Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences},
  year={2019},
  doi={10.1101/622803},
  url={https://www.biorxiv.org/content/10.1101/622803v4},
  journal={PNAS}
}

对于自注意力接触预测:

@article{rao2020transformer,
  author = {Rao, Roshan M and Meier, Joshua and Sercu, Tom and Ovchinnikov, Sergey and Rives, Alexander},
  title={Transformer protein language models are unsupervised structure learners},
  year={2020},
  doi={10.1101/2020.12.15.422761},
  url={https://www.biorxiv.org/content/10.1101/2020.12.15.422761v1},
  journal={bioRxiv}
}

对于MSA变换器:

@article{rao2021msa,
  author = {Rao, Roshan and Liu, Jason and Verkuil, Robert and Meier, Joshua and Canny, John F. and Abbeel, Pieter and Sercu, Tom and Rives, Alexander},
  title={MSA Transformer},
  year={2021},
  doi={10.1101/2021.02.12.430858},
  url={https://www.biorxiv.org/content/10.1101/2021.02.12.430858v1},
  journal={bioRxiv}
}

对于使用 ESM-1v 进行变种预测:

@article{meier2021language,
  author = {Meier, Joshua and Rao, Roshan and Verkuil, Robert and Liu, Jason and Sercu, Tom and Rives, Alexander},
  title = {Language models enable zero-shot prediction of the effects of mutations on protein function},
  year={2021},
  doi={10.1101/2021.07.09.450648},
  url={https://www.biorxiv.org/content/10.1101/2021.07.09.450648v1},
  journal={bioRxiv}
}

对于使用ESM-IF1进行反向折叠:

@article{hsu2022learning,
	author = {Hsu, Chloe and Verkuil, Robert and Liu, Jason and Lin, Zeming and Hie, Brian and Sercu, Tom and Lerer, Adam and Rives, Alexander},
	title = {Learning inverse folding from millions of predicted structures},
	year = {2022},
	doi = {10.1101/2022.04.10.487779},
	url = {https://www.biorxiv.org/content/early/2022/04/10/2022.04.10.487779},
	journal = {ICML}
}

针对ESM-2语言模型和ESMFold:

@article{lin2022language,
  title={Language models of protein sequences at the scale of evolution enable accurate structure prediction},
  author={Lin, Zeming and Akin, Halil and Rao, Roshan and Hie, Brian and Zhu, Zhongkai and Lu, Wenting and Smetanin, Nikita and dos Santos Costa, Allan and Fazel-Zarandi, Maryam and Sercu, Tom and Candido, Sal and others},
  journal={bioRxiv},
  year={2022},
  publisher={Cold Spring Harbor Laboratory}
}

本代码的许多部分都是基于 fairseq 序列建模框架构建的。我们在蛋白质语言模型研究中内部使用 fairseq。如果您想从头开始预训练蛋白质语言模型,我们强烈建议您尝试使用它。

另外,如果您想要使用 Meier 等人(2021)提出的变异预测基准,我们提供了一个包含所有数据引用的 bibtex 文件,位于 ./examples/variant-prediction/mutation_data.bib。您可以选择单独引用每篇论文,或者使用 LaTeX 命令将所有引用一次性添加。

\nocite{wrenbeck2017deep,klesmith2015comprehensive,haddox2018mapping,romero2015dissecting,firnberg2014comprehensive,deng2012deep,stiffler2015evolvability,jacquier2013capturing,findlay2018comprehensive,mclaughlin2012spatial,kitzman2015massively,doud2016accurate,pokusaeva2019experimental,mishra2016systematic,kelsic2016rna,melnikov2014comprehensive,brenan2016phenotypic,rockah2015systematic,wu2015functional,aakre2015evolving,qi2014quantitative,matreyek2018multiplex,bandaru2017deconstruction,roscoe2013analyses,roscoe2014systematic,mavor2016determination,chan2017correlation,melamed2013deep,starita2013activity,araya2012fundamental}

许可

本源代码遵循位于本源代码树根目录下 LICENSE 文件中的 MIT 许可。

ESM宏基因组图谱(亦称为“ESM宏基因组结构图谱”或“ESM图谱”)数据在学术和商业用途下遵循 CC BY 4.0 许可。版权所有(c)Meta Platforms, Inc。保留所有权利。使用 ESM 宏基因组图谱数据需遵守 Meta 开源使用条款隐私政策

项目介绍

提供Transformer蛋白质语言模型及预训练权重,涵盖结构预测、变体效应预测、逆折叠等功能,含元基因组蛋白质结构图谱,助力蛋白质设计与研究。【此简介由AI生成】

定制我的领域