scYeast: a biological-knowledge-guided foundation model on yeast single-cell transcriptomics.
scYeast is a promising tool for advancing yeast single-cell biology by integrating biological priors into foundational models.
Where it sits
this study against the rest of the sermorelin corpusSummary and findings
The study introduces scYeast, a foundational model for yeast single-cell transcriptomics that incorporates biological priors. It uses a novel architecture to enhance transcriptional regulatory information processing. scYeast shows strong performance in tasks like cell state classification and gene perturbation response prediction.
Abstract
Though large-scale pre-trained models are vital for foundational cell modeling, most of them focus on human or mouse systems, with less emphasis on model organisms like yeast (<i>Saccharomyces cerevisiae</i>), and fail to use existing biological prior knowledge effectively. Here, we present scYeast, the first foundational cell model for yeast single-cell transcriptomics that effectively embeds biological priors. scYeast employs a novel asymmetric parallel architecture to infuse transcriptional regulatory information into the Transformer's attention mechanism, leveraging biological knowledge during training. Pre-trained on large-scale yeast single-cell transcriptomics data, scYeast demonstrates strong generalization and biological interpretability. It shows capability in zero-shot tasks, such as inferring regulatory relationships. After fine-tuning, scYeast performs well in diverse tasks, including cell state classification, growth doubling time prediction, and gene perturbation response prediction. Additionally, using transfer learning, scYeast can be adapted to other omics datasets, such as proteomics, thus broadening its utility. Overall, scYeast is a promising tool for yeast single-cell biology research and presents a new framework for integrating foundational models with biological priors, accelerating discovery in yeast synthetic and systems biology and providing a replicable framework for other organisms.
Background
The study addresses the need for foundational models in yeast single-cell transcriptomics, a less emphasized area compared to human or mouse systems. Existing models often neglect to incorporate biological prior knowledge effectively. This research is significant as it aims to enhance the utility of yeast as a model organism by integrating biological priors into computational models, potentially accelerating discoveries in synthetic and systems biology.
Methods
scYeast employs an asymmetric parallel architecture to integrate transcriptional regulatory information into the Transformer's attention mechanism. It is pre-trained on large-scale yeast single-cell transcriptomics data. The model is designed to perform zero-shot tasks and can be fine-tuned for specific tasks such as cell state classification and growth doubling time prediction. Additionally, it supports transfer learning to other omics datasets.
Results
scYeast shows strong generalization and biological interpretability in yeast single-cell biology. It performs well in zero-shot tasks like inferring regulatory relationships and excels in fine-tuned tasks such as cell state classification and growth doubling time prediction. The model's adaptability to other omics datasets, like proteomics, demonstrates its broad utility.
Interpretation
The introduction of scYeast represents a significant advancement in yeast single-cell transcriptomics by effectively embedding biological priors. While the model shows strong potential, the abstract does not provide specific quantitative results or comparisons to existing models. The study's implications for practice are promising, particularly in synthetic and systems biology, though the clinical relevance remains low due to the focus on yeast.
Key findings
- scYeast demonstrates strong generalization and biological interpretability.
- Performs well in zero-shot tasks such as inferring regulatory relationships.
- After fine-tuning, excels in cell state classification and growth doubling time prediction.
- Capable of transfer learning to other omics datasets like proteomics.
Limitations
- Not reported in abstract.