Human-Protein-Language-Model-0.9
Overview
This repository contains 254 Pfam-level protein language model checkpoints generated for the study “Does domain-specific unsupervised fine-tuning improve protein language model performance?”.
The checkpoints were derived from the ESM-2 650M model through unsupervised masked language modeling on sequences associated with human-proteome Pfam families. Protein sequences were clustered at a sequence-identity threshold of 0.9 before model fine-tuning.
Repository organization
Model checkpoints are organized by their corresponding Pfam identifiers. Each directory contains the checkpoint associated with that protein family or clan.
Intended use
These checkpoints are provided for research on protein representation learning, protein language model benchmarking, and the evaluation of domain-specific unsupervised fine-tuning.
The models are intended for research use only and have not been validated for clinical or diagnostic applications.
Usage
The checkpoints can be loaded using the fair-esm package. Usage examples and benchmarking code are available in the project repository:
https://github.com/TianBoxue-lab/DS-UFT-Benchmark
Related resources
The complete model collection is available at:
License
This repository is distributed under the AFL-3.0 license.
Model tree for zxcyfr/Human-Protein-Language-Model-0.9
Base model
facebook/esm2_t33_650M_UR50D