rewirebio.io

A Bioinformatician's Guide to Choosing Genomic Foundation Models

A practical guide to selecting genomic foundation models for bioinformatics tasks. Covers ESM-2, DNABERT-2, HyenaDNA, Nucleotide Transformer, scGPT, and Evo with scoped comparisons for DNA, proteins and single-cell analysis, with corrected references and clear distinctions between frozen representations and trained predictors.

Correction, 22 September 2026. This January guide previously made recommendations that its sources did not establish, including a general scGPT accuracy claim and model size as a reason to prefer Nucleotide Transformer. I have corrected the references, distinguished frozen embeddings from a trained predictor, and withdrawn unsupported hardware, cost and timing estimates. The original wording of the main recommendations is recorded below. Figures whose version or provenance did not match the cited paper have been removed. This remains an overview of the model versions discussed here, not a current leaderboard. The June follow-up explains why matched evaluation matters.

When Meta AI released their ESM Metagenomic Atlas, they reported predictions for more than 617 million protein sequences1, produced in two weeks on approximately 2,000 GPUs2. That is a large-scale computational result, not 617 million experimentally solved structures.

You have a variant effect prediction task. Or maybe you need to annotate cell types from single-cell RNA-seq. Perhaps you want to predict how a mutation affects protein structure. The models exist. They have impressive benchmarks. But which one should you actually use?

Choosing a model requires more than reading its headline score. The task, data split, adaptation procedure and computing requirements determine whether a published result is relevant. This guide explains how to assess those choices for the model versions discussed below.

Start with the biological task

Before examining architectures, we first need to establish which model family matches your task. Foundation models in genomics fall into distinct categories with different strengths.

DNA Sequence Tasks

For regulatory-element prediction, variant-effect scoring or promoter identification, DNA models are candidates alongside task-specific predictors and simpler baselines. Their inputs and resource requirements differ.

DNABERT-2 is a compact candidate for sequence classification. Zhou and colleagues report comparable performance to their Nucleotide Transformer comparator with 21-fold fewer parameters and approximately 92-fold less pretraining GPU time.3 These are author-reported comparisons on the paper's benchmarks, not a hardware-normalised cost ratio or a promise about every task. The model has 117 million parameters4 and uses variable-length byte-pair encoding (BPE), which reduced token counts by about fivefold in the authors' data.5 Measure token lengths for your own sequences rather than assuming a fixed bases-to-tokens conversion.

HyenaDNA provides checkpoints with contexts reaching one million single-nucleotide tokens.67 The authors report a training-speed advantage of up to 160-fold over their Transformer comparator; that is not a universal inference-speed ratio.8 Its long convolutions use an algorithm with O(L log L) sequence-length scaling, rather than the O(L²) pairwise computation of dense attention.9

A caveat on HyenaDNA: accepting a long sequence does not establish that distant bases improve a particular prediction. Compare shorter windows and a task-specific baseline on the same split. Independent DNA-model evaluations find that model choice and embedding extraction interact with the task.10 I have removed the earlier suggestion that a 32k-context checkpoint covers a 100kb input, and the unsupported recommendation to pretrain ModernBERT as a general remedy.

Nucleotide Transformer includes models trained on human or multispecies data. The paper describes 3,202 human genomes and 850 species in its training collections.11 Training coverage makes it a candidate for cross-species evaluation; it does not prove transfer to a new organism. The final paper also reports an updated 250M-parameter model comparable to the earlier 2.5B model on its evaluation, so size alone is a poor selection rule.12

Correction to the original recommendation: “If you need cross-species generalization, this is your starting point” was too broad. The earlier emphasis on the largest parameter count did not establish superiority. Choose the training distribution, representation and checkpoint by a matched test on the intended task, as discussed in the June article.

Protein Tasks

For protein structure, function prediction, or variant pathogenicity:

ESM-2 is a family of protein language models, with versions from 8M to 15B parameters.1314 It produces sequence representations; it is not itself the complete ESMFold structure predictor. A smaller checkpoint is useful for testing a pipeline, but the paper does not establish the 650M version as the best choice for every downstream task.

ESMFold adds a folding model to predict coordinates from sequence. Its paper reports 14.2 seconds for a 384-residue protein on one NVIDIA V100, about six times faster than the AlphaFold2 configuration used in that comparison.1516 Larger speedups were reported for shorter proteins.17 Prediction quality and runtime both depend on the evaluated proteins and settings; the speed result does not establish equal accuracy everywhere.

ESM-2 contact-prediction and structure-related results at different parameter counts in the paper’s evaluated settings. Figure 1: ESM-2 results for the structure-related evaluations in Lin et al. (2023), not a scaling guarantee for every downstream task. Source: Lin et al.18

Single-Cell Analysis

For cell type annotation, batch integration, or perturbation prediction:

scGPT uses gene and expression representations and supports task-specific adaptation. Cui and colleagues describe pretraining on more than 33 million human cells.19 Their pretraining-scale experiments are evidence about the configurations they tested, not a guarantee that more pretraining cells improve every use.20

Geneformer uses transfer learning from ranked gene-expression inputs for network-biology tasks.21 That makes it a candidate to assess when labels are limited, not an automatic winner in that setting.

Correction: the earlier claim of an “8 to 12% increase in the biological conservation score” was attributed to Geneformer. That wording belongs to an early scGPT manuscript, not the cited Geneformer paper. It is withdrawn here as evidence for choosing Geneformer. The earlier scGPT scaling statement also cited the wrong model's paper.

Kedzierska and colleagues evaluated frozen scGPT and Geneformer embeddings and found that they did not consistently outperform simpler methods on the cell-clustering and batch-effect assessments studied.22 Ahlmann-Eltze and colleagues separately found that the evaluated deep-learning methods did not beat simple baselines on their single- and double-perturbation prediction tests.23 Neither finding means that all fine-tuned applications fail. Fine-tuned batch integration, zero-shot clustering and perturbation prediction are different evaluations.

Comparing computing requirements

The DNABERT-2 paper contrasts about 14 days on eight RTX 2080 Ti GPUs with a reported Nucleotide Transformer training run of 28 days on 128 A100 GPUs.24 That illustrates different training scales, but multiplying days by GPU counts does not make unlike hardware, datasets and training objectives equivalent.

Comparison What is reported What it cannot establish
DNABERT-2 pretraining Eight RTX 2080 Ti GPUs for about 14 days Your fine-tuning time or electricity/cloud bill
Nucleotide Transformer comparator 128 A100 GPUs for 28 days in the cited comparison A hardware-normalised efficiency ratio
Your downstream evaluation Must specify checkpoint, input lengths, batch size, precision and hardware Cannot be inferred from either pretraining run

Table 1. Reported training setups from Zhou et al.24 The previous dollar estimates and uncited HyenaDNA training row are withdrawn: they did not include a reproducible pricing or workload calculation.

Reading benchmark evidence

The cited 2024 GenBench version evaluates ten genomic foundation models across 43 datasets.25 It is one benchmark suite, not an exhaustive comparison of all available models.

Pooling Strategy

Feng and colleagues compared ways of extracting embeddings from DNA models and found mean pooling useful across their sequence-classification evaluations.10 I would include it as a starting configuration, then validate it against alternatives on the intended task. This finding is from Feng et al., not GenBench, and it does not establish the best pooling method for protein or cell models.

Note: Pooling strategies are an active research area. That result was published in 2025 and newer approaches continue to emerge. Mean pooling is a reasonable default, but it is worth checking recent literature for your specific task.

Attention vs. Convolution Trade-offs

GenBench includes attention-based models and alternatives using long convolutions or state-space operations.26 These mechanisms have different sequence-length scaling; Caduceus uses a state-space model and should not simply be described as another convolutional architecture.27

There is no universal 4kb cutoff at which attention becomes unusable. Tokenisation, trained context, batch size, numerical precision and the implementation all affect feasibility. A supported token window is also not the same as evidence that the model uses the full biological context.610

Operation Sequence-length scaling Interpretation
Dense self-attention O(L²) pairwise computation Describes growth at fixed model width, not a GPU-memory threshold
Hyena long convolution O(L log L) Does not predict a universal wall-clock advantage
State-space scan O(L) Must still evaluate the actual model and task

The previous chart labelled normalised theoretical curves as relative compute time. This table retains the algorithmic distinction without presenting it as a measured benchmark.927

Task-Specific Performance Patterns

GenBench and Feng et al. examine different tasks and evaluation setups.2510 Use their individual comparisons to choose candidates, rather than assigning a general winner. Coding and non-coding sequence tasks differ in both inputs and targets; “non-coding” does not mean that every base has an established regulatory function.2829

Measuring hardware requirements

Compute requirements matter, but they need a specified workload. The previous fixed VRAM requirements and two-, four- and eight-hour fine-tuning estimates were not accompanied by measured configurations, so I have removed them.

Memory Requirements by Model

Separate model-weight storage from the memory needed to execute the model. Activations, attention buffers, optimizer states, precision, batching and sequence lengths can change the total substantially. The official ESM implementation provides CPU-offloading examples, so the earlier statement that a 15B model necessarily requires multiple GPUs was too strong.

For a useful local measurement, record the exact checkpoint, input-length distribution, precision, batch size, device and peak memory. Test the longest relevant inputs as well as a typical batch. A successful short-sequence example is not a memory guarantee for the dataset.

Training Time Expectations

Large pretraining runs, such as those described for ProtTrans30, should not be used as estimates for downstream adaptation. Nucleotide Transformer uses the parameter-efficient (IA)3 method, updating about 0.1% of parameters in the reported setup.31 That reduces trainable state; it does not remove the base model or prove that any workload fits on a particular GPU.

Inference Speed Benchmarks

Use the ESMFold timing above as a reported result with its V100, sequence length and comparison settings. Dividing a day by 14.2 seconds gives a theoretical serial rate for identical inputs, not measured throughput for a metagenome with varied lengths and batching. The earlier 0.2-second figure and extrapolated genomes-per-day estimates are withdrawn.1617

For long-context DNA, HyenaDNA's species-classification experiment uses very long windows and reports 99.5% accuracy in its specified setup.3233 That result does not establish variant-effect or regulatory-function accuracy.

HyenaDNA soft-prompt tuning accuracy at different numbers of trainable prompt tokens. Figure 2: HyenaDNA soft-prompt tuning: accuracy against the number of trainable prompt tokens. This is a different experiment from the species-classification result discussed above. Source: Nguyen et al. (2023), Figure 4.2, Section 4.3.34

When to Fine-tune vs. Use Zero-shot

The important distinction is what learns from the target labels. A frozen encoder can supply embeddings to a classifier that is trained on those labels. The encoder is frozen, but the complete predictor is supervised.

Start with Frozen Representations When:

You want a relatively inexpensive test of whether the representation helps a defined task. Nucleotide Transformer evaluates trained probes on frozen representations; those comparisons are not label-free prediction.35 Compare the probe with a simple model on the same split. If labels are scarce, the uncertainty in the comparison becomes more important; the earlier rule that fewer than 1,000 labels favours zero-shot use was unsupported.

Consider Fine-tuning When:

You have enough target data to train and validate adaptation, and a frozen-representation baseline leaves a useful gap. The Nucleotide Transformer paper reports improvements with fine-tuning on its benchmark, but those gains are scoped to its tasks and protocol.36 Keep a held-out test set separate from choosing pooling, model size or training settings.

Parameter-Efficient Fine-tuning

(IA)3 and LoRA train small sets of adaptation parameters while keeping the base weights frozen, using different mechanisms. The Nucleotide Transformer paper's roughly 0.1% figure concerns (IA)3.31 The DNABERT-2 training code also exposes a LoRA configuration. Budget memory and fit the method you actually run; do not transfer the parameter fraction from one setup to another.

The scGPT paper includes adapted cell-analysis evaluations.37 A batch-integration plot alone cannot establish superiority across datasets. I removed the earlier embedding image because its displayed value, 0.8125, was not the 0.821 cited in the recommendation, and the image did not establish the provenance of that comparison.

Practical Decision Matrix

Use these as candidate tests, not a ranking across incompatible tasks.

Task Candidates from this guide What to check before choosing
DNA variant or regulatory prediction DNABERT-2 or a suitable Nucleotide Transformer checkpoint Define the functional endpoint and compare with an appropriate specialist/simple baseline; an embedding is not a calibrated variant score
Long-context DNA A HyenaDNA checkpoint with the required context; Evo for appropriate sequence tasks Match the checkpoint's supported input and training domain; compare shorter windows
Single-cell annotation or integration scGPT or Geneformer alongside established methods Separate frozen from adapted evaluations; hold out relevant donors/batches
Protein representation or variant effects An ESM-family checkpoint and a defined scoring/probe procedure Separate sequence effects from clinical interpretation and structure prediction
Protein structure ESMFold or a structure-specific method such as AlphaFold2 Match sequence/complex inputs, confidence checks and runtime settings
Protein-ligand assemblies RoseTTAFold All-Atom within its supported inputs Check the actual benchmark, ligand representation and experimental context
Sequence generation Evo or a documented protein-design procedure Generation alone does not establish function

These are methodological choices to evaluate, based on the cited model papers and implementations, not newly measured results.12361321373839

Correction to the single-cell recommendation: the January list said “Cell type annotation: scGPT (highest accuracy)” and “Limited training data: Geneformer transfer learning”, and cited scGPT's reported AvgBIO of 0.821. That did not establish a general recommendation. The developer-reported batch-integration comparison40 is distinct from independent zero-shot and perturbation evaluations.2223 Use the same task, inputs and adaptation regime before comparing scores. The June follow-up develops this distinction.

The original decision-tree image repeated those blanket recommendations and an unsupported structure-prediction timing, so the table replaces it.

What the Metagenomic Atlas demonstrates

The ESM Metagenomic Atlas contains predictions for more than 617 million MGnify90 sequences.41 The authors classified about 365 million as good-confidence predictions and more than 225 million as high-confidence predictions using their confidence thresholds.4243 These labels describe model confidence, not experimental validation.

Evo introduces a 131k context window genomic foundation model using the StripedHyena architecture, trained on 300 billion nucleotides from prokaryotic genomes. Figure 3: Evo architecture showing StripedHyena design and training data composition. Source: Nguyen et al., "Sequence modeling and design from molecular to genome scale with Evo", 202439

The paper also analyses sequence and structural similarity to existing resources.4445 Absence of a close sequence match or a match under a particular structure-search procedure does not establish a new fold experimentally. These are predicted structures with different degrees of confidence, not experimentally discovered folds.

Open Source Availability and Model Access

The model families below provide public code or checkpoint access. Check the licence for the exact code, weights and data separately; availability is not a universal permission for every use. This list refers to the versions discussed in the guide.

ESM family: Available at the official ESM repository which documents downloadable checkpoints and their configurations.

DNABERT-2: the official DNABERT-2 repository with HuggingFace integration.

Nucleotide Transformer: the official repository distinguishes model generations, sizes and training collections.

HyenaDNA: The official implementation links its pretrained checkpoints.

Evo: the official Evo repository documents the 7B model and its supported checkpoints.46 The paper reports a 131,072-token context and training on 300 billion nucleotide tokens from bacterial, archaeal, phage and plasmid sequences.4748 That training distribution is relevant when judging a proposed application.

scGPT: The official implementation links checkpoints and tutorials. The original paper describes more than 33 million pretraining cells.19

What These Models Can't Do (Yet)

Several limitations affect how these results should be used:

Context limitations persist. One million bases are approximately 0.03% of a 3.2-billion-base genome.49 This arithmetic describes input coverage, not how much biology the model has learned, and does not establish a limit for every other model or method.

Prediction does not establish mechanism. ProtTrans reports biologically informative representations.50 Successful prediction alone does not identify the causal mechanism behind a protein function or variant effect.

Model size is not a general performance guarantee. Kan-Tor and colleagues report no clear larger-model advantage in their gene-property prediction benchmark using frozen representations.51 That is the source of the previously unidentified quotation; it does not test rare-cell-type annotation, so the earlier inference about rare-cell accuracy is withdrawn.

Training coverage matters. An organism appearing in a pretraining collection is different from demonstrating transfer to an unseen organism. Review coverage and evaluation overlap for the checkpoint and test distribution.1222

Getting Started: Extracting Protein Embeddings

Here is a small embedding example using the official ESM interface. First-run time includes the checkpoint download; the 650M weights are about 2.5GB. The code uses CPU by default. Choose a smaller supported checkpoint for a lighter setup, and change the requested representation layer to match it.

  1. Install the example dependencies in an isolated environment: pip install fair-esm==2.0.0 torch numpy. Choose a PyTorch build compatible with your system; this is an API example, not a complete locked runtime environment. See the official ESM setup.

  2. Load pretrained weights (ESM-2 example):

"""
ESM-2 Model Loading Example
===========================

This script loads an ESM-2 protein language model with fair-esm and reports
its configuration and representation dimensions.

Dependencies:
    - torch
    - fair-esm (pip install fair-esm)

Reference:
    Lin et al. "Evolutionary-scale prediction of atomic-level protein structure
    with a language model" Science (2023)
"""

import torch
import esm

def load_esm2_model(model_name: str = "esm2_t33_650M_UR50D"):
    """
    Load an ESM-2 model and its batch converter.

    Args:
        model_name: Name of the ESM-2 model variant. Options include:
            - "esm2_t6_8M_UR50D"    (8M parameters)
            - "esm2_t12_35M_UR50D"  (35M parameters)
            - "esm2_t30_150M_UR50D" (150M parameters)
            - "esm2_t33_650M_UR50D" (650M parameters)
            - "esm2_t36_3B_UR50D"   (3B parameters)

    Returns:
        tuple: (model, alphabet, batch_converter)
            - model: The ESM-2 PyTorch model
            - alphabet: Tokenizer for converting sequences to tokens
            - batch_converter: Utility for preparing batched inputs
    """
    # Load pre-trained ESM-2 model
    # This downloads weights on first run (~2.5GB for 650M model)
    model, alphabet = esm.pretrained.load_model_and_alphabet(model_name)

    # Create batch converter for tokenizing sequences
    batch_converter = alphabet.get_batch_converter()

    # Set model to evaluation mode (disables dropout)
    model.eval()

    return model, alphabet, batch_converter

def prepare_protein_batch(sequences: list, batch_converter) -> tuple:
    """
    Prepare protein sequences for model input.

    Args:
        sequences: List of tuples (name, sequence_string)
            Example: [("protein1", "MKTVRQERLK"), ("protein2", "GALTISGTW")]
        batch_converter: The batch converter from ESM alphabet

    Returns:
        tuple: (batch_labels, batch_strs, batch_tokens)
            - batch_labels: List of sequence names
            - batch_strs: List of sequence strings
            - batch_tokens: Tokenized tensor ready for model input
    """
    batch_labels, batch_strs, batch_tokens = batch_converter(sequences)
    return batch_labels, batch_strs, batch_tokens

# Example usage
if __name__ == "__main__":
    # Load the model
    model, alphabet, batch_converter = load_esm2_model()
    print("Model loaded successfully!")
    print(f"  - Parameters: {sum(p.numel() for p in model.parameters()):,}")
    print(f"  - Embedding dimension: {model.embed_dim}")
    print(f"  - Number of layers: {model.num_layers}")
    print(f"  - Attention heads: {model.attention_heads}")
    print(f"  - Vocabulary size: {len(alphabet)}")

    # Example protein sequences
    example_sequences = [
        ("hemoglobin_alpha", "MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTK"),
        ("insulin_b_chain", "FVNQHLCGSHLVEALYLVCGERGFFYTPKT"),
    ]

    # Prepare batch and run inference
    print(f"\nPreparing {len(example_sequences)} protein sequences...")
    batch_labels, batch_strs, batch_tokens = prepare_protein_batch(
        example_sequences, batch_converter
    )
    print(f"  - Batch token shape: {tuple(batch_tokens.shape)}")
    print(f"  - Sequences: {batch_labels}")

    print("\nRunning forward pass...")
    with torch.no_grad():
        results = model(batch_tokens, repr_layers=[33], return_contacts=False)

    # Extract representations from final layer
    token_representations = results["representations"][33]
    print(f"  - Output shape: {tuple(token_representations.shape)}")
    print("    (batch_size x sequence_length x embedding_dim)")

Expected dimensions for these two fragments: the longest sequence has 41 residues. ESM adds beginning and end tokens, giving a padded token tensor of shape (2, 43) and a final-layer representation of shape (2, 43, 1280) for the 650M checkpoint. These dimensions are not a runtime benchmark or a claim of rerunning the checkpoint during this correction.

  1. Extract embeddings with mean pooling, as shown in the official ESM example. Feng et al. studied DNA models, so their comparison does not establish the best pooling strategy for ESM-2:
"""
ESM-2 Embedding Extraction with Mean Pooling
=============================================

Extract protein embeddings from ESM-2 and apply mean pooling
to obtain fixed-size sequence representations for downstream tasks.
"""

import torch
import esm
import numpy as np
from typing import List, Tuple

def extract_embeddings_with_pooling(
    sequences: List[Tuple[str, str]],
    model,
    batch_converter,
    layer: int = 33,
    pooling: str = "mean",
    device: str = "cpu"
) -> dict:
    """
    Extract protein embeddings from ESM-2 with sequence-level pooling.

    Args:
        sequences: List of (name, sequence) tuples
        model: ESM-2 model
        batch_converter: Alphabet batch converter
        layer: Which transformer layer to extract from (default: final layer)
        pooling: Pooling strategy - "mean" or "max"
        device: Device to run inference on ("cpu" or "cuda")

    Returns:
        dict: {
            "embeddings": numpy array of shape (num_sequences, embedding_dim),
            "labels": list of sequence names,
            "per_residue": list of per-residue embeddings
        }
    """
    model = model.to(device)
    model.eval()

    # Prepare batch
    batch_labels, batch_strs, batch_tokens = batch_converter(sequences)
    batch_tokens = batch_tokens.to(device)

    # Extract representations
    with torch.no_grad():
        results = model(batch_tokens, repr_layers=[layer], return_contacts=False)

    token_reps = results["representations"][layer]

    pooled_embeddings = []
    per_residue_embeddings = []

    for i, (tokens, seq_str) in enumerate(zip(token_reps, batch_strs)):
        seq_len = len(seq_str)
        # Extract residue tokens (exclude <cls> and <eos>)
        residue_reps = tokens[1:seq_len + 1]
        per_residue_embeddings.append(residue_reps.cpu().numpy())

        if pooling == "mean":
            pooled = residue_reps.mean(dim=0)
        elif pooling == "max":
            pooled = residue_reps.max(dim=0)[0]
        else:
            raise ValueError(f"Unknown pooling strategy: {pooling}")

        pooled_embeddings.append(pooled.cpu().numpy())

    return {
        "embeddings": np.stack(pooled_embeddings),
        "labels": batch_labels,
        "per_residue": per_residue_embeddings
    }

def compute_sequence_similarity(emb1: np.ndarray, emb2: np.ndarray) -> float:
    """Compute cosine similarity between two embeddings."""
    dot_product = np.dot(emb1, emb2)
    norm1 = np.linalg.norm(emb1)
    norm2 = np.linalg.norm(emb2)
    return dot_product / (norm1 * norm2)

# Example usage
if __name__ == "__main__":
    # Load model
    model, alphabet = esm.pretrained.esm2_t33_650M_UR50D()
    batch_converter = alphabet.get_batch_converter()

    # Short fragments for demonstrating the API
    proteins = [
        ("human_hba", "MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTK"),
        ("mouse_hba", "MVLSGEDKSNIKAAWGKIGGHGAEYGAEALERMFASFPTTK"),
        ("insulin_b", "FVNQHLCGSHLVEALYLVCGERGFFYTPKT"),
    ]

    # Extract embeddings
    result = extract_embeddings_with_pooling(
        proteins, model, batch_converter, layer=33, pooling="mean"
    )

    embeddings = result["embeddings"]
    labels = result["labels"]

    # Full pairwise cosine similarity matrix
    print("Cosine Similarity Matrix:")
    print(" " * 12 + "".join(f"{name:>12}" for name in labels))
    for i, name in enumerate(labels):
        row = "".join(
            f"{compute_sequence_similarity(embeddings[i], embeddings[j]):>12.3f}"
            for j in range(len(labels))
        )
        print(f"{name:>12}{row}")

The example defaults to residue-mean pooling. It excludes beginning and end tokens; the official ESM guidance advises against using the beginning-token representation for pretrained models.

The script prints a pairwise cosine-similarity matrix for short example fragments. I removed the previously printed similarity values because this correction does not include a verified inference receipt for those exact values. A cosine score alone does not establish homology or function; validate the representation with a matched downstream task.

  1. Use embeddings for downstream tasks: Feed into a simple classifier or regressor.

DNA-model tokenisers and implementations have their own requirements; follow the exact checkpoint instructions rather than assuming this protein example transfers unchanged.

Multimolecular models and sequence generation

RoseTTAFold All-Atom represents proteins, nucleic acids, small molecules and other molecular components within a shared assembly model.52 Its evaluations need to be read by molecular task and input setting; a ligand result is not interchangeable with a protein-only structure score.

Evo's paper reports long coding-rich generated sequences and selected experimental validations.53 Generating a long sequence does not establish that it constitutes a viable organism or that every encoded element works.

The original Evo architecture combines 29 Hyena layers with three attention layers.5455 Its hybrid design is a specific implementation choice, not evidence that this mixture is optimal for every biological task.

Conclusion

For practical model selection, start with a defined output, a relevant baseline and an evaluation that matches the way the model will be used.

  • Start with a manageable checkpoint: establish that the pipeline works before testing whether scale helps.
  • Test pooling and adaptation: frozen embeddings plus a trained head are a supervised predictor, not label-free evidence.
  • Measure the actual workload: context length, hardware, precision and batching belong next to a runtime or memory claim.
  • Validate on held-out data: keep model selection separate from the final comparison, and report failures as well as scores.

Genomic foundation models can provide useful representations and predictions. The sources reviewed here do not establish one best model across DNA, proteins and cells. The decision comes from the task and the evidence, not the largest model or the most prominent developer-reported result.


References

Footnotes

  1. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  2. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  3. Zhou, Z., et al. DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome. ICLR 2024 (2024). ↩ ↩2

  4. Zhou, Z., et al. DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome. ICLR 2024 (2024). ↩

  5. Zhou, Z., et al. DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome. ICLR 2024 (2024). ↩

  6. Nguyen, E., et al. HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution. NeurIPS 2023 (2023). ↩ ↩2 ↩3

  7. Nguyen, E., et al. HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution. NeurIPS 2023 (2023). ↩

  8. Nguyen, E., et al. HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution. NeurIPS 2023 (2023). ↩

  9. Nguyen, E., et al. HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution. NeurIPS 2023 (2023). ↩ ↩2

  10. Feng, H., et al. Benchmarking DNA foundation models for genomic and genetic tasks. Nature Communications 16, 10780 (2025). ↩ ↩2 ↩3 ↩4

  11. Dalla-Torre, H., et al. Nucleotide Transformer: building and evaluating robust foundation models for human genomics. Nature Methods 22, 287–297 (2025). Published online 28 November 2024. ↩

  12. Dalla-Torre, H., et al. Nucleotide Transformer: building and evaluating robust foundation models for human genomics. Nature Methods 22, 287–297 (2025). Published online 28 November 2024. ↩ ↩2 ↩3

  13. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩ ↩2

  14. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  15. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  16. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩ ↩2

  17. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩ ↩2

  18. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  19. Cui, H., et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nature Methods 21, 1470–1480 (2024). ↩ ↩2

  20. Cui, H., et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nature Methods 21, 1470–1480 (2024). ↩

  21. Theodoris, C. V., et al. Transfer learning enables predictions in network biology. Nature 618, 616–624 (2023). The previously quoted 8–12% figure was from Cui et al.’s May 2023 scGPT preprint, Discussion, lines 322–323; it was not a Geneformer result. ↩ ↩2

  22. Kedzierska, K. Z., et al. Zero-shot evaluation reveals limitations of single-cell foundation models. Genome Biology 26, 101 (2025). Frozen-representation clustering and batch-effect evaluations. ↩ ↩2 ↩3

  23. Ahlmann-Eltze, C., Huber, W., and Anders, S. Deep-learning-based gene perturbation effect prediction does not yet outperform simple linear baselines. Nature Methods (2025). Single- and double-perturbation evaluations. ↩ ↩2

  24. Zhou, Z., et al. DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome. ICLR 2024 (2024). ↩ ↩2

  25. Liu, Z., et al. GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models. arXiv preprint (2024). ↩ ↩2

  26. Liu, Z., et al. GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models. arXiv preprint (2024). ↩

  27. Liu, Z., et al. GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models. arXiv preprint (2024). ↩ ↩2

  28. Liu, Z., et al. GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models. arXiv preprint (2024). ↩

  29. Liu, Z., et al. GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models. arXiv preprint (2024). ↩

  30. Elnaggar, A., et al. ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 44(10), 7112–7127 (2022). Published online 7 July 2021. ↩

  31. Dalla-Torre, H., et al. Nucleotide Transformer: building and evaluating robust foundation models for human genomics. Nature Methods 22, 287–297 (2025). Published online 28 November 2024. ↩ ↩2

  32. Nguyen, E., et al. HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution. NeurIPS 2023 (2023). ↩

  33. Nguyen, E., et al. HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution. NeurIPS 2023 (2023). ↩

  34. Nguyen, E., et al. HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution. NeurIPS 2023 (2023). ↩

  35. Dalla-Torre, H., et al. Nucleotide Transformer: building and evaluating robust foundation models for human genomics. Nature Methods 22, 287–297 (2025). Published online 28 November 2024. ↩

  36. Dalla-Torre, H., et al. Nucleotide Transformer: building and evaluating robust foundation models for human genomics. Nature Methods 22, 287–297 (2025). Published online 28 November 2024. ↩

  37. Cui, H., et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nature Methods 21, 1470–1480 (2024). ↩ ↩2

  38. Krishna, R., et al. Generalized biomolecular modeling and design with RoseTTAFold All-Atom. Science 384, eadl2528 (2024). ↩

  39. Nguyen, E., et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024). ↩ ↩2

  40. Cui, H., et al. scGPT: toward building a foundation model for single-cell multi-omics using generative AI. Nature Methods 21, 1470–1480 (2024). Figure 4a: fine-tuned PBMC 10k cell-type clustering. ↩

  41. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  42. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  43. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  44. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  45. Lin, Z., et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science 379, 1123–1130 (2023). ↩

  46. Nguyen, E., et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024). ↩

  47. Nguyen, E., et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024). ↩

  48. Nguyen, E., et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024). ↩

  49. Nguyen, E., et al. HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution. NeurIPS 2023 (2023). ↩

  50. Elnaggar, A., et al. ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 44(10), 7112–7127 (2022). Published online 7 July 2021. ↩

  51. Kan-Tor, Y., et al. Does your model understand genes? A benchmark of gene properties for biological and text models. arXiv preprint (2024). Results, Figure 2; gene-property prediction from representations. ↩

  52. Krishna, R., et al. Generalized biomolecular modeling and design with RoseTTAFold All-Atom. Science 384, eadl2528 (2024). ↩

  53. Nguyen, E., et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024). ↩

  54. Nguyen, E., et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024). ↩

  55. Nguyen, E., et al. Sequence modeling and design from molecular to genome scale with Evo. Science 386, eado9336 (2024). ↩

Frequently asked

Which genomic foundation model should I use for DNA sequence analysis?
Choose candidates for a defined task, organism and input window. DNABERT-2 and Nucleotide Transformer offer starting points for sequence representations; long-context HyenaDNA checkpoints address different input lengths. Compare them with an appropriate baseline on the same split. Maximum context and parameter count do not establish accuracy.
Is ESM-2 a protein structure predictor?
ESM-2 is a protein language-model family that produces representations. ESMFold adds a folding model to predict coordinates. Reported speed comparisons depend on the sequences, hardware and AlphaFold2 configuration tested; they do not establish equal accuracy on every structure task.
How do attention and long-convolution models differ?
Dense attention has quadratic pairwise computation in sequence length; Hyena long convolutions have O(L log L) scaling. These describe algorithmic growth, not measured runtime or a universal sequence-length cutoff. Check the checkpoint, tokenisation, implementation and task.
What hardware do genomic foundation models require?
There is no single reliable VRAM requirement per family. Memory and time depend on checkpoint, sequence lengths, precision, batch size and adaptation. The earlier fixed memory, cost and fine-tuning-time estimates in this guide lacked reproducible measurements and have been withdrawn.
Are frozen embeddings the same as zero-shot predictions?
No. A frozen encoder can supply embeddings to a classifier or regressor trained on target labels. The encoder stays fixed, but the complete predictor is supervised. Compare that procedure separately from label-free scoring and fine-tuning.
Should I use mean pooling?
Mean pooling is a useful candidate to test. Feng and colleagues found benefits in their DNA-model sequence-classification evaluations. That result does not establish the best pooling rule for every protein, cell or genomic task.

Help improve this article

Found an error or a better source? Leave a note here, or highlight a passage to comment on it.