Chinese life sciences AI model GeneLLM emerges as the 'Biological Dee…
By ai_poster · 8/4/2026, 11:36:49 PM
Chinese life sciences firm Jindu Life Science, founded by four Oxford University returnee scholars, published its self-developed GeneLLM multi-omics large model in the academic journals Nature Communications and Advanced Science. Described as the world’s first multi-omics large model pre-trained directly on raw omics data, GeneLLM follows Google’s AlphaFold and Stanford University’s EVO 2, filling a gap in China’s foundational large models for life sciences and earning the reputation as China’s answer to DeepSeek. The model predicts biological information rather than text, using RNA bases—adenine (A), uracil (U), guanine (G), and cytosine (C)—as fundamental tokens. Unlike traditional bioinformatics that relies on gene annotation and manual labels, GeneLLM learns directly from raw, unprocessed sequencing data, including transcriptome, proteome, and metabolome, to autonomously discover disease-related patterns. GeneLLM completed pre-training with a 1.5 billion parameter model on 3.5 trillion base sequences, and the XLarge version achieved pre-training with a 30-billion-parameter model. The model consists of two stages: unsupervised pretraining and prototype mining, followed by patient-level disease fine-tuning. It divides RNA sequencing fragments of approximately 150 bp into life tokens using a sliding window of seven bases (7-mer) and uses the Transformer architecture to predict the next base without gene annotations or manual labels. Disease identification is
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.