Data as of Aug 25, 2026 · Based on 270 AI responses from ChatGPT Search and Google AI Mode · See how Parse measures this
Pharmaceutical research teams have access to various biology-specific foundation models depending on their focus. NVIDIA BioNeMo serves as a central hub hosting diverse models for drug discovery. For more targeted research, companies like EvolutionaryScale (ESM-3) and Google DeepMind (AlphaFold 3) are leaders in protein modeling, while Atomic AI and
Deep Genomics specialize in RNA-based discovery, and
provides multi-omics and interactome models.
Brands AI recommends here
Named in 43% of answers
Named in 43% of answers
Named in 39% of answers
Named in 34% of answers
Yes. The field has moved well beyond generic LLMs: there are now several providers building **biology-specific foundation models**, with different modalities and commercial models. ### Leading providers to know Provider | Representative models/platform | Primary biology focus | Pharmaceutical relevance
Yes. The field has moved well beyond generic LLMs: there are now several providers building biology-specific foundation models, with different modalities and commercial models.
| Provider | Representative models/platform | Primary biology focus | Pharmaceutical relevance |
|---|---|---|---|
| EvolutionaryScale | ESM3, ESM Cambrian | Protein sequence, structure, function, generative protein design | Protein/antibody engineering, therapeutic protein design |
| NVIDIA | BioNeMo | Proteins, molecules, structure, docking, generation | Broad drug-discovery infrastructure; train/fine-tune models on proprietary data |
| **Arc Institute / Stanford | Evo 2 | DNA/genomics → RNA/protein | Variant effects, genomic design, gene engineering |
| Boltz | Boltz-1/2, BoltzGen | Biomolecular structure, affinity, molecular generation | Small molecules + biologics; particularly interesting for open models |
| OpenAI | GPT-Rosalind | Biology, chemistry, protein engineering, genomics + scientific reasoning | Research-agent workflows and drug-discovery reasoning |
| Isomorphic Labs | Proprietary AlphaFold-derived models | Protein structure, interactions, drug design | End-to-end computational drug discovery; partnerships rather than general model licensing |
| Generate:Biomedicines | Generative protein models | De novo protein therapeutics | Antibodies and other protein therapeutics |
| Basecamp Research | EDEN models | Protein/DNA design using biodiversity data | Protein engineering and genetic-medicine applications |
| Chai Discovery | Chai-1 and successors | Biomolecular structure/design | Protein–ligand and protein–protein modeling |
| Recursion | Large biological/phenotypic models | Cellular biology, perturbation data, chemistry | Target discovery, phenotypic screening and drug design |
A few distinctions are important.
EvolutionaryScale is probably the clearest example of a “biology foundation-model company.” ESM3 is a generative model designed to understand and generate proteins, while ESM Cambrian focuses on protein representation learning. EvolutionaryScale explicitly makes ESM3 available through its platform and has commercial partnerships around the models.
NVIDIA BioNeMo is different: it is more of a model/platform ecosystem than a single biology model. It provides pretrained biomolecular models, training infrastructure, APIs and deployment tooling spanning protein structure, protein and small-molecule generation, docking and property prediction.
Arc's Evo 2 is particularly interesting if genomics matters. It is a 40B-parameter genomic foundation model trained on nearly 9 trillion nucleotides, intended to model and generate biological sequences from the molecular to genome scale.
Boltz is especially worth watching for an open-model strategy. Boltz-2 and BoltzGen are positioned as open biomolecular foundation models covering structure prediction, affinity estimation and molecular/biologics design; Pfizer announced a 2026 collaboration to develop models using Pfizer's proprietary data.
And OpenAI has now entered the category directly: GPT-Rosalind was introduced in April 2026 as a purpose-built life-sciences model spanning biology, drug discovery and translational medicine, with capabilities across chemistry, protein engineering and genomics.
If you're evaluating vendors rather than simply making a landscape, I'd divide them into four buckets:
For a pharmaceutical discovery program, I would not choose based on parameter count or the phrase "foundation model." The important questions are whether the model has been trained on the right biological modality, whether you can fine-tune it on proprietary data, whether commercial rights are clear, whether inference can be deployed in your environment, and—most importantly—whether it improves experimentally validated endpoints.
If you tell me whether your focus is small molecules, antibodies/biologics, RNA, cell biology, genomics, or end-to-end drug discovery, I can narrow this to the 5–10 providers I'd actually evaluate, including model access, licensing, proprietary-data fine-tuning, and pharma partnerships.
NVIDIA BioNeMo is different: it is more of a model/platform ecosystem than a single biology model. It provides pretrained biomolecular models, training infrastructure, APIs and deployment tooling spanning protein structure, protein and small-molecule generation, docking and property prediction.
Arc's Evo 2 is particularly interesting if genomics matters. It is a 40B-parameter genomic foundation model trained on nearly 9 trillion nucleotides, intended to model and generate biological sequences from the molecular to genome scale.
Boltz is especially worth watching for an open-model strategy. Boltz-2 and BoltzGen are positioned as open biomolecular foundation models covering structure prediction, affinity estimation and molecular/biologics design; Pfizer announced a 2026 collaboration to develop models using Pfizer's proprietary data.
The ecosystem for biology-specific foundation models spans specialized AI drug-discovery companies, tech-giant spin-offs, and cloud infrastructure platforms. Providers generally fall into a few distinct categories depending on whether you are looking for structural biology, multi-omics, target identification, or…
The ecosystem for biology-specific foundation models spans specialized AI drug-discovery companies, tech-giant spin-offs, and cloud infrastructure platforms. Providers generally fall into a few distinct categories depending on whether you are looking for structural biology, multi-omics, target identification, or general biomedical reasoning.
To help narrow down which provider or framework fits your project, let me know:
For pharmaceutical discovery, “biology-specific foundation models” usually refers to models pretrained on biological data (protein sequences, structures, genomes, transcriptomes, single-cell data, molecular interactions, literature, etc.) and adapted for tasks such as target discovery, protein engineering, biomarker…
For pharmaceutical discovery, “biology-specific foundation models” usually refers to models pretrained on biological data (protein sequences, structures, genomes, transcriptomes, single-cell data, molecular interactions, literature, etc.) and adapted for tasks such as target discovery, protein engineering, biomarker discovery, and molecule design. The market is split between model builders, platform providers, and AI-native biotech companies that use proprietary models internally.
Here are the major providers to evaluate:
| Provider | Biology foundation model focus | Typical pharma use cases |
|---|---|---|
| EvolutionaryScale | Protein language models (ESM family lineage), protein design | Protein engineering, sequence-function prediction, therapeutic protein design |
| Meta Platforms, Inc. | Protein language models such as ESM-derived research models | Protein embeddings, structure/function prediction research |
| Google DeepMind | Structure prediction and biomolecular modeling (AlphaFold family) | Target validation, protein structure, interaction modeling |
| Isomorphic Labs | AI-first drug discovery models built around biomolecular understanding | Small-molecule and biologics discovery programs |
| NVIDIA Corporation | BioNeMo ecosystem: pretrained models, inference, fine-tuning infrastructure | Enterprise deployment, protein modeling, molecular generation |
| Microsoft | AI infrastructure and life-science model ecosystem | Model hosting, scientific AI workflows, enterprise integration |
| OpenAI | Life-science reasoning models and scientific assistants | Literature reasoning, experimental planning, biological workflows; GPT-Rosalind is an example of a purpose-built life-science model effort. OpenAI |
| Recursion Pharmaceuticals | Multimodal biology models combining imaging, chemistry, and biology data | Phenotypic screening, target discovery, drug programs |
| Insitro | ML models trained on human genetics, cellular assays, and disease biology | Target identification and translational research |
| Schrödinger | Physics-based molecular modeling plus ML | Structure-based drug design, virtual screening |
| Generate Biomedicines | Generative protein models | De novo protein therapeutics and biologics discovery |
| Absci | Generative AI for antibodies and biologics | Antibody discovery and optimization |
| Chai Discovery | Multimodal molecular/biomolecular models | Protein-ligand and biomolecular modeling |
| Boltz | Open biomolecular foundation models (e.g., Boltz models) | Structure prediction, affinity estimation, biomolecular design Boltz |
Examples:
Best for:
Examples:
Best for:
Examples:
Best for:
These combine:
Examples include Recursion, Insitro, and newer virtual-lab platforms.
For a discovery organization, I would compare providers on:
A practical shortlist for a pharmaceutical discovery group would usually include:
The “best” provider depends heavily on whether your priority is target discovery, biologics, small molecules, biomarker discovery, or building an internal AI discovery platform.
Entering the space of generative AI for pharmaceutical discovery means navigating a rapidly expanding ecosystem of biology-specific foundation models. Unlike general Large Language Models (LLMs) trained on internet text, these models are trained on atomic structures, genomic sequences, high-content cellular images,…
Entering the space of generative AI for pharmaceutical discovery means navigating a rapidly expanding ecosystem of biology-specific foundation models. Unlike general Large Language Models (LLMs) trained on internet text, these models are trained on atomic structures, genomic sequences, high-content cellular images, and evolutionary data.
The primary providers, tech bio pioneers, and platforms offering biology-specific foundation models span several categories:
If you want to narrow this down, let me know:
I can provide a more tailored architectural recommendation.
The provider landscape for biology-specific foundation models spans major tech conglomerates, well-funded pure-play AI biotech startups, and open-source scientific initiatives. Choosing a provider typically depends on whether you need pre-trained model weights for proteins/genomics, enterprise cloud infrastructure, or…
The provider landscape for biology-specific foundation models spans major tech conglomerates, well-funded pure-play AI biotech startups, and open-source scientific initiatives. Choosing a provider typically depends on whether you need pre-trained model weights for proteins/genomics, enterprise cloud infrastructure, or an end-to-end drug discovery partnership.
Key providers of biology-specific foundation models are organized by their primary ecosystem focus:
To help narrow down which provider or platform fits your strategy, could you share:
For pharmaceutical discovery, “biology-specific foundation models” are usually grouped by what biological data they learn from: **proteins**, **molecules**, **genomes/transcriptomes**, **cells**, or **multimodal biology**. The leading providers include a mix of AI companies, academic labs, cloud platforms, and biotech…
For pharmaceutical discovery, “biology-specific foundation models” are usually grouped by what biological data they learn from: proteins, molecules, genomes/transcriptomes, cells, or multimodal biology. The leading providers include a mix of AI companies, academic labs, cloud platforms, and biotech companies.
| Provider | Key biology foundation models / platforms | Primary use in pharma discovery |
|---|---|---|
| Google DeepMind | AlphaFold, AlphaFold 3, AlphaMissense, AlphaProteo | Protein structure prediction, molecular interactions, protein design, variant interpretation. AlphaFold 3 extends modeling beyond proteins to interactions involving DNA, RNA, ligands, and other biomolecules. deepmind.google |
| Isomorphic Labs | AlphaFold-derived drug discovery systems | Integrated AI-first drug design workflows built around biological structure and interaction modeling. deepmind.google |
| EvolutionaryScale | ESM family (ESM-2, ESM-3, ESMFold) | Protein language modeling, protein function prediction, protein generation and design. developer.nvidia.comgithub.comintuitionlabs.ai |
| NVIDIA | BioNeMo platform (ESM, AlphaFold/OpenFold, DiffDock, molecular generators) | Enterprise deployment of biological foundation models for protein engineering, molecular generation, docking, and drug discovery pipelines. developer.nvidia.comgithub.compmc.ncbi.nlm.nih.govopenreview.netintuitionlabs.aiwww.youngju.dev |
| Generate Biomedicines | Generative protein design models | Designing novel therapeutic proteins and biologics. www.youngju.dev |
| University of Washington Institute for Protein Design | RoseTTAFold, RFdiffusion, ProteinMPNN | Protein structure prediction and de novo protein design; widely used as open research foundations. pmc.ncbi.nlm.nih.govopenreview.net |
| Recursion Pharmaceuticals | Large-scale biological AI platform | Uses AI models over imaging, chemistry, genetics, and biological datasets for target discovery and drug programs. www.youngju.dev |
| Insilico Medicine | Pharma.AI platform | Target discovery, molecular generation, and AI-designed drug candidates. www.youngju.dev |
| Schrödinger | Physics-based + AI molecular modeling | Structure-based drug design, molecular simulation, and virtual screening. www.youngju.dev |
| Atomwise | AI molecular screening models | Virtual screening and small-molecule discovery. www.youngju.dev |
| BenevolentAI | Knowledge graph + AI discovery systems | Target identification, biological hypothesis generation, and drug repositioning. www.youngju.dev |
For an internal pharma AI stack, many teams evaluate open models alongside commercial platforms:
Protein sequence models
Structure models
Protein design
Small molecule generation/design
A practical evaluation framework:
Target biology
Deployment model
Data advantage The model alone is rarely the differentiator. Leading pharma programs increasingly combine foundation models with proprietary assay data, clinical data, omics data, and experimental feedback loops.
For a pharmaceutical discovery organization starting today, a common shortlist would be:
I would also distinguish foundation model providers from AI drug discovery companies: the former sell or release reusable biological models; the latter often combine models with proprietary data, automation, and wet-lab validation.
Yes. The market has moved well beyond generic LLMs: there are now several companies building **biology-native foundation models**, although they differ substantially in what they model and whether you can actually license/use the models versus access them through a discovery platform. ### The main providers I would…
Yes. The market has moved well beyond generic LLMs: there are now several companies building biology-native foundation models, although they differ substantially in what they model and whether you can actually license/use the models versus access them through a discovery platform.
| Provider | Biology focus | Representative models/platform | Best fit for pharma |
|---|---|---|---|
| EvolutionaryScale | Protein sequence, structure & function | ESM-3, ESM Cambrian | Protein engineering, therapeutic proteins, antibodies |
| NVIDIA | Proteins, molecules, RNA, chemistry | BioNeMo | Enterprise model development, fine-tuning, deployment |
| Isomorphic Labs | Multimolecular structure & drug design | AlphaFold 3 / IsoDDE | Structure-based drug design, protein–ligand interactions |
| Chai Discovery | Biomolecular structure/design | Chai-1 and newer models | Protein–ligand/antibody modeling and design |
| Basecamp Research | Evolutionary biology, sequences, therapeutics | EDEN | Designing proteins/genetic medicines from evolutionary data |
| Recursion Pharmaceuticals | Cellular biology, chemistry, phenotypes | Biology/chemistry foundation models | Target discovery, phenotypic screening, drug design |
| Owkin | Multimodal human biology, pathology, genomics | OwkinZero, K Pro | Disease biology, target discovery, translational research |
| Generate:Biomedicines | Protein generation | Generative protein models | Designing therapeutic proteins |
| Profluent | Protein language/generative models | ProGen family | Protein/antibody engineering |
A few deserve particular attention:
EvolutionaryScale is probably the clearest example of a true biology foundation-model company. Its ESM3 model jointly models protein sequence, structure and function, while ESM Cambrian focuses on large-scale protein representation learning. ESM3 is explicitly positioned as a generative biology model.
NVIDIA BioNeMo is different: it is more of a foundation-model development and deployment ecosystem than a single biology model. It provides pretrained biomolecular models and tooling for protein structure prediction, molecular generation, protein–ligand modeling, docking and property prediction, as well as infrastructure for training your own models. www.nvidia.com This makes it particularly interesting if you're a pharma organization that wants to combine public foundation models with proprietary experimental data.
Isomorphic Labs is pursuing a more vertically integrated approach. Its Drug Design Engine goes beyond structure prediction toward an end-to-end computational drug-design system. In February 2026, the company reported substantially improved protein–ligand prediction and binding-affinity performance relative to AlphaFold 3 on its reported benchmarks.
Basecamp Research is particularly interesting if you're thinking about training data as a strategic asset. Its EDEN models are trained on an unusually large evolutionary dataset and are designed for predictive and generative biological tasks, including therapeutic design.
Owkin is taking a different route: its OwkinZero is a biology-specific reasoning model trained using hundreds of thousands of verified biomedical Q&A examples and proprietary multimodal patient data. Its emphasis is less "generate a new protein" and more reason over disease biology, multiomics and clinical data.
Protein foundation models
Structure & molecular-interaction models
Cellular, genomic & disease-biology models
Foundation-model infrastructure for pharma
If you're preparing an actual pharma AI strategy, I would not make the vendor shortlist solely on who has the biggest biological model.
The key question is what biological representation you need:
DNA/genome → RNA → protein → structure → interaction → cell phenotype → tissue → patient → clinical outcome
Different foundation models operate at different points on that chain. A protein model such as ESM3 can be exceptional for antibody/protein engineering but isn't necessarily the right foundation for target identification or patient stratification. Conversely, a multimodal model trained on pathology, transcriptomics and clinical data may be much more valuable for target discovery but useless for de novo protein generation.
Also, some of the most interesting models are not commercially accessible as standalone foundation models. For example, Isomorphic Labs is primarily pursuing its own drug-design engine, whereas BioNeMo is explicitly designed to let organizations build and deploy their own biomolecular AI workflows.
If your goal is selecting vendors for a pharmaceutical discovery program, I can next give you a more useful 2026 vendor landscape comparing ~15 providers by modality (small molecule, antibody, protein, RNA, genomics, cell biology), model accessibility, licensing, proprietary-data fine-tuning, and evidence of pharma partnerships.
OwkinZero
ProGen
The provider `landscape for biology-specific foundation models (FMs)` spans specialized AI-native biotech companies, major tech conglomerates building life science infrastructure, and open-science consortia.[](https://codephusion.com/blog/ai-models-for-life-sciences)…
The provider landscape for biology-specific foundation models (FMs) spans specialized AI-native biotech companies, major tech conglomerates building life science infrastructure, and open-science consortia.
Providers are generally categorized by their primary biological modality—such as proteins, genomics, or multi-omic cell simulations:
To help narrow down which provider fits your pipeline, could you share:
Evo
The landscape of biology-specific foundation models for drug discovery is driven by specialized techbio companies, major AI labs, and hardware-software ecosystems. These organizations provide pre-trained models spanning genomics, transcriptomics, protein folding, and molecular…
The landscape of biology-specific foundation models for drug discovery is driven by specialized techbio companies, major AI labs, and hardware-software ecosystems. These organizations provide pre-trained models spanning genomics, transcriptomics, protein folding, and molecular design.
The primary providers of biology-specific foundation models are categorized below:
If you tell me which specific modality you want to focus on (e.g., small molecule design, protein/antibody generation, or genomics), I can recommend which specific model or platform matches your pipeline goals.
Evo
Yes. The market has moved well beyond generic LLMs: there are now **biology-specific foundation models spanning proteins, DNA/genomics, cells, pathology, molecular structures, and drug–target interactions**. I’d divide the landscape into a few important groups: | Provider / organization | Representative…
Yes. The market has moved well beyond generic LLMs: there are now biology-specific foundation models spanning proteins, DNA/genomics, cells, pathology, molecular structures, and drug–target interactions.
I’d divide the landscape into a few important groups:
| Provider / organization | Representative models/platform | Biology focus | Relevance to pharma discovery |
|---|---|---|---|
| Google DeepMind / Isomorphic Labs | AlphaFold 3, AlphaGenome | Protein/complex structure; genomics | Structure-based discovery, target biology, variant effects |
| EvolutionaryScale | ESM family, including ESM-3 | Protein sequence, structure & function | Protein engineering, therapeutic proteins, target understanding |
| NVIDIA | BioNeMo | Proteins, molecules, structure, docking, generation | Particularly attractive as an enterprise platform for building/fine-tuning models |
| Boltz / MIT–community ecosystem | Boltz-1, Boltz-2 | Biomolecular structure + binding affinity | Protein–ligand modeling and virtual screening |
| Arc Institute / collaborators | Evo / Evo 2 | DNA/genomic sequences | Regulatory genomics, variant effects, biological sequence design |
| Bioptimus | H-Optimus, M-Optimus | Histology → multimodal/multiscale biology | Biomarkers, patient stratification, tissue response, translational research |
| Owkin | Phikon and multimodal models | Pathology / multimodal patient data | Biomarker discovery and clinical-development applications |
| Recursion | Recursion OS / biological foundation models | Cells, perturbation biology, imaging + omics | Target discovery and phenotypic drug discovery |
| Insilico Medicine | PandaOmics, Chemistry42 and related models | Target discovery + molecule generation | End-to-end AI drug discovery |
| Profluent | ProGen-derived protein models | Protein generation/design | Novel therapeutic protein and enzyme design |
| Single-cell model ecosystem | Geneformer, scGPT, etc. | Single-cell transcriptomics | Cell-state modeling, perturbation prediction, target discovery |
A few are especially worth investigating.
For protein engineering and protein therapeutics, EvolutionaryScale is one of the clearest specialist providers. Its ESM family treats protein sequences/structures as the biological analogue of language, enabling generation and prediction of protein properties.
For protein structure and protein–ligand interactions, the AlphaFold lineage is still foundational. Google DeepMind's AlphaFold 3 is designed to model interactions involving proteins, nucleic acids, small molecules and other biomolecular components. Isomorphic Labs is applying this technology specifically to drug discovery.
There is also a rapidly developing open-model ecosystem. Boltz-2, for example, is a structural-biology foundation model that adds binding-affinity prediction to structure prediction; its authors report dramatically lower computational cost than traditional free-energy perturbation approaches.
For target discovery, regulatory genomics and variant interpretation, Google DeepMind's AlphaGenome is particularly notable. It can process DNA sequences up to 1 million base pairs and predict gene expression, splicing, chromatin features and other regulatory signals. Google currently provides API access for non-commercial research.
The Evo/Evo 2 family from the Arc Institute ecosystem takes a somewhat different approach: large DNA-language models trained directly on genomic sequences. These are potentially interesting for sequence design as well as genomic interpretation.
For pharma, this category may ultimately be just as important as protein models because many discovery questions are really about cell state and response to perturbation.
Geneformer, for example, is a transformer pretrained on large-scale single-cell transcriptomes and intended to support context-aware predictions in network biology.
Models in this category can be used for things such as:
There is still considerably more uncertainty around how well these models generalize to genuinely novel perturbations than the marketing sometimes suggests.
If your discovery program has substantial histopathology or spatial biology, don't overlook this category.
Bioptimus is particularly interesting. Its H-Optimus-1 is a 1.1B-parameter pathology foundation model trained on more than 1 million histology slides from more than 800,000 patients. The company is now extending this toward M-Optimus, intended to integrate pathology, spatial transcriptomics and genomics.
This is much closer to a drug-development biology model than simply an image classifier: potential applications include biomarker discovery, patient stratification and predicting drug response.
NVIDIA's BioNeMo deserves separate consideration because it isn't just one foundation model. It is an ecosystem for training, fine-tuning and deploying biomolecular models, including models for protein structure, protein generation, small-molecule generation, property prediction and molecular docking.
That makes BioNeMo particularly relevant if your question is:
"We have proprietary pharma data and want to build/customize biology foundation models ourselves."
rather than:
"Which pretrained model gives us the best biological prediction?"
If I were evaluating vendors/models for a pharmaceutical discovery program today, I'd start with:
Protein/structure
Genomics
Cells / perturbations
Tissue / translational biology
End-to-end drug discovery
The important distinction is that "biology foundation model" is not one market. A protein language model, a single-cell model, a pathology model and a protein–ligand structure model solve fundamentally different problems. For pharma, I would choose the model family based on the discovery bottleneck you're trying to address rather than looking for one universal biology LLM.
If you tell me whether your priority is target identification, small-molecule discovery, biologics/protein design, biomarker discovery, or patient stratification, I can narrow this to the 5–10 providers/models I'd actually evaluate, including access model, licensing, proprietary-data fine-tuning, compute requirements, and evidence of pharma adoption.