Model weights
Download the reference models we serve.
Each row is the current published build of one reference: the archive, the commit that produced it, its size, the SHA-256 of the bytes you will receive, the two licences that build ships under, and the reference atlas report describing what it was trained on. What each model was built from is tabulated on the methods page.
Every archive holds a scvi-tools scANVI model, with the configuration it was trained under recorded in the manifest inside it; the Model column reports the family and library version each published build declares.
| Reference | Model | Weights | Report | Build | Size | SHA-256 | Weights licence | Metadata licence |
|---|---|---|---|---|---|---|---|---|
| Human blood & immune | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human cortex | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human eye | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human lung | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human gut | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human breast | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human heart | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human liver | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human kidney | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human skin | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human adipose | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human muscle | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Human female reproductive | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Mouse brain | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
| Mouse multi organ | not yet published | Download | Download | v1 · not yet published | not yet published | not yet published | not yet published | not yet published |
Batch labels are published as opaque indices. Model behaviour is unchanged. The full
explanation ships as NOTICE.txt inside the archive.
How to load these weights
Each archive is a scvi-tools
scANVI checkpoint and runs on CPU. Every archive carries a generated
README.md with the full recipe for that build: its gene list, the cell
types it can return, the batch indices it accepts and the library versions it was
written with. Below is the short form, which is the same for every reference.
pip install torch --index-url https://download.pytorch.org/whl/cpu
pip install scvi-tools
import anndata as ad, scvi
MODEL = "the directory you extracted the archive into"
query = ad.read_h5ad("my_cells.h5ad")
query.layers["counts"] = query.X # raw integer counts
query.obs["cell_type"] = "Unknown" # every query cell is unlabelled
query.obs["sample_ID"] = "batch_0000" # any batch index the archive publishes
scvi.model.SCANVI.prepare_query_anndata(query, MODEL)
model = scvi.model.SCANVI.load(MODEL, adata=query)
model.predict() # a cell type per cell
model.predict(soft=True) # per-class probability
model.get_latent_representation() # the latent embedding
-
Raw counts. The model reads
layers["counts"]and expects integers, not normalised or log-transformed values. -
Ensembl gene IDs, not symbols. Symbols match nothing.
prepare_query_anndatadrops genes the model never saw and zero-fills the reference genes your data lacks, so it decides what the model reads; each archive's README states that model's own gene count. -
Any published batch index.
obs["sample_ID"]has to hold one of the archive's opaque indices (batch_0000upwards) or scvi-tools refuses the query. Which one you pick cannot change the result: the batch value is not an input to the encoder. - A loadable AnnData ships inside each archive, and it is a scaffold, not data: it carries the category vocabularies the checkpoint needs, with no expression values. It loads, and predicting on it returns nothing meaningful. Bring your own counts.
-
Install the CPU build of PyTorch first. A plain
pip install scvi-toolsresolves PyTorch from PyPI, which is the CUDA build: a long download of packages a CPU run does not use.
Licensing
Each archive carries two licences: the model parameters, our own work, under the
licence named in the row (it permits any use, including commercial use, with
attribution); and the metadata beside them (label registry, manifest, dataset
list, and the scaffold adata.h5ad the checkpoint loads against), derived from the training data and passed on under its own terms with
attribution to the contributing studies. The archive's LICENSE says
which files fall under which.
Training data: CZ CELLxGENE Census, CC BY 4.0, attributed per study on
the attribution page and in each build's manifest.
We do not redistribute reference expression data. See terms.