Dataset walkthrough & critique

TheBioCollection, read closely

A 52.6 billion-token biology pre-training corpus — what's actually inside it, shown with verbatim records, and where it falls short of its own framing.

52.6B tokens 62.4M records 2 streams · 64 shards Apache-2.0 Lunit · Trillion Labs et al. reviewed 2026-07-21
73%
of the free-text stream is single small-molecule descriptors
~2%
is genuine natural-language biomedical prose
89%
of binding-assay records ship a broken template line
0.00%
descriptor mismatch vs. RDKit (enrichment is accurate)
0.12
template-diversity ratio in the instruction stream
0.17%
is dedicated pathway (KEGG/Reactome) content
Overview

A verbalized database, not a corpus of biology

TheBioCollection converts heterogeneous biological databases into training text, enriched with tool-computed properties and paired with programmatically-verifiable instruction tasks. That engineering is real and, in places, excellent. But the corpus is far more concentrated — and far more templated — than its "five equal domains" framing suggests.

How this was measured: I streamed 2 of 32 shards per stream (~1.6 GB), classified 3.9M records by content, cross-checked 4,000 molecule records against RDKit, and ran artifact/duplication counters over full shards. The two shards of each stream agree to <0.1% — the corpus is globally shuffled, so one shard represents the whole. Source names are inferred from content; percentages are sample-based estimates.

Composition

What the two streams are made of

Because records carry no domain label, this is reconstructed by classifying text into mutually-exclusive content classes. Both streams are near-identical across shards.

Free-text stream · 37.2M records · 33.8B tokens

Share of records by content class (inferred source in grey).
Small-molecule descriptorPubChem + RDKit
73.1%
Protein-structure narrationAlphaFold + DSSP/Biopython
17.7%
Ligand–target binding assayBindingDB / ChEMBL / patents
4.4%
Genomic DNA windowENCODE / Ensembl
1.5%
UniProt function / GOUniProt / TrEMBL
1.4%
RNA sequence / foldEnsembl / Rfam
1.0%
Cells / pathways / proseliterature, misc
0.9%
0.35%records with no structured tag (pure prose)
5.4%records with an empty-slot template bug
0.04%exact duplicates (dedup worked)
0.002%DrugBank monographs (26 in 1.16M)

Instruction stream · 25.1M records · 18.9B tokens

Share of records by primary content class. Task-signal keywords (below) overlap and don't sum to 100%.
Protein design / binder / interfacePDB structure-derived
40.1%
Genomic span localizationENCODE open-chromatin / GENCODE splice
27.6%
Small-molecule property / reactionTherapeutics Data Commons / USPTO
20.5%
Cell / spatial descriptive recordsJUMP Cell Painting / Slide-seq · HuBMAP
9.6%
RNA (miRNA–target)miRTarBase-like
2.1%
48%request JSON / verifiable answers
26%binder / protein-design tasks
24%list residue contacts / interfaces
18%reference "within 5.0 Å" contacts
Literature content

Paper full texts? Published open-access only — no preprints

A fair question, since a "biology corpus" might be expected to pull bioRxiv/medRxiv. It doesn't. Searching a full 1.16M-record free-text shard for the fingerprints that would be present if preprints had been ingested:

Preprint fingerprintRecordsReading
10.1101/ DOI prefix (bioRxiv/medRxiv)0—
"not certified by peer review" — the bioRxiv/medRxiv PDF watermark0—
PMC##### identifier0—
"bioRxiv" appearing anywhere2authors citing a preprint inside a published paper
"medRxiv" / "arXiv"1 / 1same — in-text mentions, not sources

If preprint full texts were present, the watermark line and 10.1101 DOIs would appear thousands of times per shard. Zero is conclusive.

0preprint watermarks in 1.16M records
1.71%of free-text characters are full-text articles
~64kestimated OA articles corpus-wide
~1%of all 52.6B tokens is OA full text

What the prose actually is

There is genuine full-text literature — just published open-access journal articles, not preprints. It's ~0.17% of records but longer, so 1.71% of free-text characters (~1% of the whole corpus). The DOIs cited in the bodies are dominated by BMC, then Springer/Nature, PeerJ, Hindawi, JCI, Frontiers — the classic PMC Open Access mix. Every sampled article is body text (Background / Introduction / Results / Discussion) that terminates in a stray Figures and Captions header — the signature of an XML→text extraction with front- and back-matter stripped. Topically it skews hard toward cell biology, single-cell, and immunology (plus CELLxGENE cell-type profiles and Cell Ontology definitions), reading like literature pulled to shore up the thin "cells" domain rather than a general biomedical-literature corpus.

Governance implication

Only ~4 of ~2,000 sampled articles retain any license string, and none retain the article's own DOI or PMC id. So even though most PMC-OA content is CC-BY (reusable with attribution), the attribution is gone and you can't separate it from the more restrictive CC-BY-NC-ND subset. This is finding #10 in miniature — full-text articles ingested with provenance and license stripped.

Random draws

An unbiased look: uniform samples

First entries from a seeded reservoir sample over shard 1 — genuinely random, not selected. Long raw sequences are trimmed with …[+N] for display; everything else is verbatim.

Free-text #1small-molecule descriptor — the 73% case2,351 chars
The molecule has the SMILES representation: <smiles>CC1=CC(=CC(=C1)C(=O)N[C@H](CCSC)C(=O)O)C</smiles> and its IUPAC systematic name is (2R)-2-[(3,5-dimethylphenyl)carbonylamino]-4-methylsulfanyl-butanoic acid. It also has the corresponding InChI: <inchi>InChI=1S/C14H19NO3S/c1-9-6-10(2)8-11(7-9)13(16)15-12(14(17)18)4-5-19-3/t12-/m1/s1</inchi>.
The compound contains 38 atoms, including 19 heavy atoms, and 38 bonds. The molecule is Chiral, and it has 0 charged atoms, and a total formal charge of 0. Its molecular formula is C14H19NO3S, giving an exact mass of 281.10856464g/mol and a molecular weight of 281.37g/mol.
The molecule features 4 hydrogen bond acceptors, 2 hydrogen bond donor, 6 rotatable bond, and an XlogP3-AA value of 2.6, indicating moderate polarity, balanced aqueous solubility and membrane permeability.
There is 1 ring in the molecular graph, including 1 aromatic, 0 aliphatic, and 0 saturated ring. … The Murcko scaffold is c1ccccc1. Graph complexity descriptors include Bertz complexity 453.5, Labute ASA 117.1 A^2, Hall-Kier alpha -1.5, and Kier kappa values 15.7, 9.9, and 9.6. …
Free-text #2binding assay — with the template bug1,433 chars
The ligand 5-(1-Isopropyl-1H-1,2,3-triazol-4-yl)-3-(…)pentanoic acid
(SMILES: <smiles>CC(C)n1cc(CCC(CC(O)=O)c2ccc(C)c(CN3C[C@@H](C)Oc4ccccc4S3(=O)=O)c2)nn1</smiles>, InChI: <inchi>…</inchi>)
 was tested against the protein Nuclear factor erythroid 2-related factor 2.
The target protein name is Human of NF2L2_HUMAN and chain sequence is <protein>MMDLELPPPGLPSQQDMDLIDILWRQDIDLGVSREVFDFSQRRKEY…[+560 aa]</protein>.
 PDB IDs of target chain are 7O7B,7X5E,7X5F
The experiment was performed at pH  and .
The ligand exhibited moderate binding strength, with a reported EC50 (nM) of  550.
Two artifacts in one record: missing assay conditions leave pH and ., and a missing unit leaves a double-space before the value. This occurs in ~89% of binding-assay records.
Instruction #1binder design — supervised against one native sequence1,135 chars
Design a single-chain binder that satisfies the ligand-pocket motif below. Use only the 20 standard amino acids, produce a single chain of exactly 193 residues, and keep the motif anchors in N-to-C order.

Ligand (SMILES): <smiles>CC1CCN(C(=O)N[C@H](C)c2ccccc2)CC1</smiles>
Binding-pocket motif (anchors in order):
1. an aromatic residue making an aromatic (pi-stacking) contact with the ligand's aromatic ring
2. an aromatic residue making an aromatic (pi-stacking) contact with the ligand's aromatic ring
3. a hydrophobic residue making a hydrophobic contact with ligand atom 6
4. a polar residue making a hydrogen-bond-like contact with ligand atom 13
… (6 anchors)
Output only the binder sequence.
Answer: <protein>GDDRELAILATAENLLEDRPLADISVDDLAKGAGISRPTFYFYFPSKEAVLLTLLDRVVNQADM…[+193 aa]</protein>
Instruction #2contact summary — 3-D answer from sequence-only input1,787 chars
Summarize how the binder protein contacts the target protein, focusing on residue-level contacts. Two residues are in contact when their heavy atoms are within 5.0 Å, and positions are 1-indexed within each chain. For this pair, chain R is the target and chain L is the binder.

Target protein (chain R): <protein>GLPVMNTPGSNQFLTSDDYQSPTAMPQFDVTPEMNIPGEV…[+200 aa]</protein>
Binder protein (chain L): <protein>GAQVSTQKTGAHEIHYTNINYYKDAASNSANRQDFTQDPGKFTEPVKDIMVKSLPALN</protein>

Answer: The binder presents a 54-contact interface against the target, with 8 hydrogen-bond-like contacts, 4 hydrophobic contacts, 39 polar contacts, and 3 salt-bridge-like contacts. The binder-side patch is L:N18, L:I19, L:N20, L:Y21, L:Y22, L:K23, L:D24 … the corresponding target epitope is R:S16, R:D17, R:D18, R:Y19, R:Q20 … Representative pairwise contacts include L:N18 to R:Q20 (polar), L:I19 to R:Q20 (polar) …
The recurring problem in one card: the answer is precise 3-D interface geometry, but the model is given only two 1-D sequences. Nothing in the input determines the answer — it is learnable only by memorizing the specific PDB complex. See finding #3.
Guided tour

One representative record per aspect

Expand each to see a full record. Together these span the corpus's real range — from its clean, dominant molecule records to its rare literature and its most questionable tasks.

A · AlphaFold protein-structure narration17.7% of free-text2,516 chars
Protein A0A101LZ69_PICGL is a medium-sized protein of 101 amino acid residues containing 829 atoms, with its three-dimensional structure predicted by AlphaFold. Its sequence is <protein>MIRVSSLRYAISTLVSLVIGWFRHLCVWWFFKLVGSPIKLFQVSTMFPWHLLSYGTNDMILRLVQMNKGLGSHPHTHIGQKGSPHPAIVRYSHPSYDRCYL</protein>.

Physicochemical properties: molecular weight 11,728 Da, isoelectric point 10.11 (basic, strong net positive charge), GRAVY 0.171 (slightly hydrophobic), instability index 32.6 (stable in vitro).
Secondary structure: 64% helix, 0% sheet, 36% coil. Helical regions span residues 2-34, 37-68. Fold topology H-H.
Geometry and compactness: Radius of gyration 34.7 Å. Relative contact order 0.0231. Long-range contact count: 0.
Interactions: 1 salt bridge ARG62–ASP58 (3.5 Å). No disulfide bonds despite 2 cysteine residues. …
Solvent accessibility: SASA 12,139 Ų. Remarkably, no residues are buried (RSA < 0.25), and 86% of residues are fully exposed to solvent, which is highly unusual and consistent with an intrinsically disordered protein.
Residue composition: 48% hydrophobic, 27% polar, 17 positively charged, 2 negatively charged. Absent amino acids: E.
Structural and functional annotation (UniProt): Transmembrane region: 7–31.
Confident, but self-contradictory: called a transmembrane protein yet narrated as fully solvent-exposed / disordered, with an Rg of 34.7 Å for a 101-mer (very extended). These are hallmarks of a low-confidence AlphaFold model — but no pLDDT is reported, so shaky geometry is stated as fact. See finding #6.
B · UniProt function + GOprotein/pathway prose, ~1.4%1,125 chars
Transaldolase from Xanthomonas oryzae pv. oryzae (strain PXO99A) is an enzyme (EC 2.2.1.2). The functional descriptions of this protein are as follows: Transaldolase is important for the balance of metabolites in the pentose-phosphate pathway. It catalyzes the reaction: D-sedoheptulose 7-phosphate + D-glyceraldehyde 3-phosphate = D-erythrose 4-phosphate + beta-D-fructose 6-phosphate. Gene Ontology (GO) information indicates: molecular function — transaldolase activity; biological process — carbohydrate metabolic process, pentose-phosphate shunt; cellular component — cytosol. It has a length of 322 amino acids. The protein sequence is <protein>MTASPSKLAQLRELSVVVADTGDYDAIKRLQPVDCTTNPTLVKKALDLPVYAD…[+270 aa]</protein>.
This is close to what "pathways" means in the corpus: GO/enzymatic context folded into protein records. Dedicated KEGG/Reactome pathway content is just 0.17% of the free-text stream.
C · Genomic DNA window1.5% of free-text452 chars
Sequence: <dna>ATTGGAGATTAGCTAAATTTGCTAAATGCCTGAATCTAGCAAATCCAGAGAGAAAAAGCCACAACAATGAACACC…</dna>. Mus musculus mm10 genomic DNA | location=chr3:55781947-55782165 | features=overlaps 1 gene feature for Nbea | length=218 bp | GC=32.1% | CpG sites=2 | CpG O/E=0.36 | entropy=1.90 bits | longest homopolymer=5 bp.
Better records tie the window to a feature (here, the Nbea gene). But ~0.2% of the stream are windows labeled has no selected coordinate feature overlap — random sequence plus composition stats, no signal.
D · Literature (rare, but real)~0.2% of records — longer, so more tokens13,302 chars
Proton Therapy Reduces the Effective Dose to Immune Cells in Mediastinal Hodgkin Lymphoma Patients
Purpose Effective dose to circulating immune cells (EDIC) is associated with survival in lung and esophageal cancer patients. This study aimed to evaluate the benefit of intensity-modulated proton therapy (IMPT) for EDIC reduction compared with volumetric modulated arc therapy (VMAT) in mediastinal Hodgkin lymphoma (mHL) patients. … [full open-access paper continues for 13k chars]
Full open-access papers appear verbatim (also seen: a 30k-char k-mer bioinformatics methods paper; NCBI GeneRIF-style "Literature-derived functional findings for gene VPS4A"). This is the corpus's only real natural-language biology — and it is vanishingly rare.
E · Genomic span task — well-posed & checkableinstruction stream · the good kind1,934 chars
You are given a human GRCh38 genomic DNA sequence measured in a T-helper 2 cell context. … The sequence below contains one open chromatin peak plus flanking DNA. The summit is the single position with the strongest accessibility signal.

Sequence:
<dna>GAGGAGCGTAGGAACCTGGTCCGCAGCCTCACCCAGCCCCCGGCAGGCCGGACCTGAGCTCCCC…</dna>

Task: Locate the open chromatin peak and its summit position. [1-indexed, inclusive; subsequence must match exactly]
Return JSON only:
{"open_chromatin_peaks":[{"assay":"<ATAC-seq or DNase-seq>","biosample":"<string>","start":<int>,"end":<int>,"summit_position":<int>,"subsequence":"<string>"}]}
{"open_chromatin_peaks":[{"assay":"DNase-seq","biosample":"T-helper 2 cell","start":151,"end":330,"summit_position":226,"subsequence":"CCCTCGAGACCCGCCAAGAAATAAAGGCGATGATTTCC…"}]}
This is the model of a good task: the answer is a span of the given input and is exactly verifiable. The prompt+answer are concatenated for pre-training.
F · Drug-property multiple choiceTherapeutics Data Commons · chemistry, checkable642 chars
Instructions: Answer the following question about drug properties.
Context: Tyrosyl-DNA phosphodiesterase is an enzyme involved in repairing stalled topoisomerase I-DNA complexes … Mutations have been associated with axonal neuropathy.
Question: Given a drug SMILES string, predict whether it (A) is not active against tyrosyl-DNA phosphodiesterase (B) is active against tyrosyl-DNA phosphodiesterase
Drug SMILES: <smiles>Clc1ccc(C2=NN(C(C2)c2occc2)C(=O)CSc2n(C3CC3)c(=O)[nH]n2)cc1</smiles>
Answer: (A)
G · Cell / spatial descriptive record~9.6% of the "instruction" stream — but no task486 & 483 chars
Observed spatial transcriptomics record: Slide-seq [Salmon], HuBMAP sample=HBM844.HNVJ.589, Homo sapiens Kidney (Right), spot CAATGCCCTATAGA, x=3049.3, y=4989.1. The spot contains 16.2 total counts across 23 detected genes. Among the most abundant transcripts are CBWD1, CDC16, CEP350, COX2, LUM, NEAT1, RPL3, RPS12. Local density: 13 neighbors within 50 um; closest 16.2 um. The record preserves measurement context without assigning tissue function.

Well-level Cell Painting record for U2OS: plate=C13487bW, well=P12, perturbation=JCP2022_106375. Matched-control morphology strength=2.24; strongest feature families are nuclear shape and size (Nuclei_AreaShape, 2.60), nuclear granularity (2.26) … Strongest CellProfiler measurement: EulerNumber (z=-22.47).
These live in the instruction stream but contain no question or answer, and are self-hedged ("preserves measurement context without assigning tissue function"; "Image-profile similarity is not a mechanism label"). Descriptive, low-signal, mislabeled. See finding #9.
H · Drug synergy as exact-integer regressionmany packed examples per record4,332 chars
Instructions: Answer the following question about drug synergy. … Synergy is calculated using the Bliss model.
Question: Given two drug SMILES and a cell line, predict the normalized synergy from 000 to 1000.

Drug1 SMILES: <smiles>CC1=C2C(C(=O)C3(…paclitaxel…))</smiles>
Drug2 SMILES: <smiles>CC1C(C(CC(O1)…doxorubicin…))</smiles>
Cell line description: HCT116
Answer: 460

Drug1 … Drug2 … Cell line: OVCAR-5   Answer: 506
Drug1 … Drug2 … Cell line: MALME-3M  Answer: 485
… (8+ pairs, answers 462, 477, 486 …)
Predicting an exact integer for a noisy experimental synergy readout is close to unlearnable and rewards memorization. See finding #7.
The strong positive

The tool-computed enrichment is accurate

Cross-checking 4,000 molecule-descriptor records against RDKit, recomputed from each record's own SMILES:

DescriptorMismatchReading
heavy atoms · total atoms · formal charge · ring count · TPSA · fraction-sp³0.00%exact
molecular weight0.05%exact
molecular formula0.10%**Hill-order for carbon-free salts; corpus arguably more correct
aromatic ring count0.21%aromaticity-model edge cases
exact mass2.18%****concentrated in multi-component / salt SMILES
SMILES parses in RDKit99.97%0.03% genuinely malformed
Worth stating plainly

For normal drug-like molecules the descriptors are essentially perfect and internally consistent with the structure. This is clean, verifiable structured knowledge — the genuinely valuable core of the corpus, and the part most defensible for training.

Findings

Where it falls short

Ordered by how much they undercut the corpus's stated goal — a model with "genuine understanding of biology." The number is referenced elsewhere on this page.

1

"Five domains" is really one domain

balance

73% small molecules + 18% protein structure. Genomics, RNA, cells, and pathways (0.17%) are rounding error in the free-text stream. A model trained here is a cheminformatics model that has seen some genomics — the card presents the five domains as co-equal, but the token reality is lopsided and undocumented.

2

A verbalized database, not a biology corpus — and the eval is matched to it

format · evidence

~98% is template-generated narration; genuine scientific prose is ~2% of records. So (a) the model learns database register, not reading/reasoning over real scientific language; and (b) the headline "more than doubles on TheBioCollection-Eval" is measured on an eval built from the same templates and sources. That largely demonstrates format acquisition and risks train/eval leakage — no source- or template-held-out split is reported.

3

Many instruction tasks are information-theoretically ill-posed

task validity

~24% ask for residue contacts / 5 Å binding sites / interface epitopes, and ~26% ask to design or scaffold a binder — from sequence-only inputs. The ground truth depends on 3-D structure the model can't see, so the only ways to fit it are memorizing PDB entries or hallucinating geometry. Protein design is additionally one-to-many yet supervised against a single native sequence with cross-entropy.

"Programmatically verifiable" ≠ "learnable from the given inputs."
4

Visible template-rendering bugs at scale

rendering

5.4% of free-text records contain performed at pH  and . (empty slots) — that's 89% of all binding-assay records shipped with broken sentences. 6.1% have a double-space where a value/unit is missing. On 37M records that's ~2M visibly malformed documents the model will learn as valid.

5

Garbage-in source records pass straight through

source quality

PubChem mixtures/formulations get single-molecule descriptor write-ups — e.g. a "molecule" that is twelve ethanes plus carbazoles and spiro-siloles (an OLED formulation) with a semicolon-list "IUPAC name" and a computed molecular weight for the bag. No connectivity/mixture filter.

6

Structural narration outruns its evidence

evidence

AlphaFold records (18%) recite Rg, SASA, contact order, "intrinsically disordered," etc. computed on models of unknown confidence (no pLDDT), sometimes self-contradictory (transmembrane yet fully exposed). Presented as fact, this teaches confident wrongness.

7

Numeracy tasks with unlearnable exact targets

numeracy

Drug-synergy regression to an exact integer 000–1000, Ki/IC50 to the reported nanomolar — these are noisy assay readouts. Demanding an exact token answer rewards memorization and injects label noise.

8

Redundancy the exact-dedup missed

redundancy

Exact duplicates are ~0.04% (good), but the instruction stream's template-diversity ratio is 0.12 — enormous near-duplication at the template level (top ~12 templates are 2–3% each). Low-information records (feature-less DNA windows; self-nullifying cell/spatial descriptions) burn tokens without teaching much.

9

Stream and label semantics are loose

semantics

Declarative cell/spatial descriptions (~9.6% of the "instruction" stream) contain no instruction at all. record_type carries zero information beyond the filename it came from.

10

No provenance → no governance

governance

With only {text, record_type} you cannot filter by source/organism/quality, cannot ablate, cannot dedup against a downstream eval, and cannot check licenses. The corpus is Apache-2.0 but derives from sources with heterogeneous terms — patent example compounds (US10301272), DrugBank, full-text papers, UniProt/ENCODE/GENCODE. Provenance is unrecoverable from the release.

11

Scope claims vs. content

scope

Tagged language:en, human/mouse-centric, and "biology," but it is overwhelmingly chemistry + structural bioinformatics. Whole areas — physiology, clinical, microbiology, ecology, evolution, and most of molecular biology's prose — are essentially absent.

Fixes

How I'd improve it

Most of these are curation and metadata, not new data collection. They're ordered by leverage.

#1 #9 #10

Ship provenance metadata — the single biggest win

Add per-record source, source_id, domain, subtype/template_id, organism, license, and quality flags. This alone unlocks rebalancing, ablation, dedup-vs-eval, and license compliance, and lets users down-weight the 73% molecule mass.

#1 #2

Rebalance and document

Publish per-source token counts and a real domain histogram. Cap the cheap-to-generate PubChem descriptor mass; up-sample the thin, high-value domains and genuine open-access prose — for balance and to fight format overfitting.

#3

Fix task well-posedness

Only pose tasks solvable from the inputs, or provide the missing context (include structure/features, or reframe as "given this structure…"). Supervise design with multiple valid references or a structural reward (foldability, motif satisfaction), not cross-entropy to one native sequence.

#4

Repair the renderer

Regenerate or drop the ~2M records with empty slots; when a source value is missing, omit the clause instead of emitting pH  and .. Trivially detectable with two regexes.

#5 #6

Filter source junk; carry confidence

Reject or tag disconnected / multi-component SMILES and non-drug-like formulations. For AlphaFold records, carry pLDDT and gate or caveat descriptors from low-confidence models; reconcile contradictory annotations.

#7

Rethink noisy-number supervision

Bin regression targets into calibrated buckets or attach uncertainty, instead of demanding exact integers / nanomolar values.

#8

De-duplicate at the template level

Sample-cap per prompt skeleton; drop feature-less DNA windows and self-nullifying descriptive records, or convert them into actual tasks.

#2

Make the eval honest

Report results on a source- and template-held-out split, plus a format-shifted eval (same biology, different phrasing/schema), so gains reflect biology rather than template familiarity. State the contamination controls.

Bottom line

This is a well-engineered, accurately-enriched verbalization of chemical and structural-bioinformatics databases — genuinely useful for teaching a model to manipulate SMILES, reactions, and sequence↔property mappings, and the descriptor accuracy is excellent.

But "a corpus for biology" oversells it: ~78% small molecules, ~98% templated, thin on real biological prose and on the genomics/cells/pathways it advertises, with a large share of instruction tasks that are unlearnable except by memorization. The fixes are mostly curation and metadata — and adding provenance would be the single biggest improvement.