Attribution
Every dataset behind our reference atlases, who published it, and under what licence.
Our reference atlases are built from public data in the CZI CELLxGENE Census. Data published through that platform is made available under CC BY 4.0 as a matter of CZI's platform policy, which is where the licence for every dataset below comes from. CC BY 4.0 asks for credit to the creator, a link to the licence, and an indication of whether changes were made. This page is where we do all three, and it is the page our reports and downloads link to.
Our analyses also draw on ontologies, marker tables, gene-set panels and an interaction network that do not come from the Census. Those carry their own and differing terms, and are listed separately under knowledge resources.
Both lists are generated rather than maintained by hand. The datasets come from the provenance manifest each reference build writes, so they name what actually went into the shipped models rather than what we intended to use. The two differ: a dataset can be declared in a cohort and then dropped during the build.
How we source and handle the data
| Item | Policy |
|---|---|
| Source corpus | CZ CELLxGENE Census. Everything submitted to it carries CC BY 4.0, so attribution is required, and it is carried into every reference manifest and every report. |
| Commercial use | Filtered at ingestion, on by default and fail-loud. CC BY-NC and CC BY-ND data is excluded from the references used in the product. |
| Cell-type ontology | Labels are harmonised to the Cell Ontology. Its terms are under knowledge resources below. |
| Marker panels | Assembled from two sources with different standings, credited separately at HuBMAP ASCT+B and CZ CELLxGENE CellGuide. |
| Your data | Processed in an isolated per-run job and never used to train our models. See the privacy notice. |
Datasets
| Datasets | 314 |
| Collections | 151 |
| Reference atlases built from them | 14 |
| Licence | CC BY 4.0, on every dataset. That is the condition CZ CELLxGENE applies to everything submitted to it, not 314 separate determinations we made. It does not extend to the knowledge resources below, which carry four different instruments between them. |
Download the full list (314 datasets, CSV) with each dataset's collection, DOI and the reference atlases it went into.
Modifications we may make. CC BY 4.0 §3(a)(1)(B) asks that modifications be indicated. A reference atlas is a heavily modified derivative rather than a redistribution, and across our builds we apply healthy donors; adult donors; cells duplicated between an integrated atlas and its source study removed; genes restricted to those shared across datasets, then a highly-variable subset for training; study labels remapped through Cell Ontology to our coarser vocabulary; one atlas with derived embeddings (latent, UMAP) not in the source. This is the full set, not a description of any one build. What a particular reference or run actually did is stated in that artifact's own notices, which are selected from what the run observed; a run that did less says less. Where the two differ, the artifact governs. Cite the per-study publications in the CSV, not this page, when using a reference.
Knowledge resources
Annotating cells takes more than the reference data. Cell types are named in a shared vocabulary, tissues are selected in an anatomical one, marker panels come from curated biomarker tables, and the pathway and interaction figures are drawn against published gene sets and networks. Those resources are credited here.
They do not share a licence. Four different instruments apply across the list below, and one resource carries no licence grant we were able to locate, so it is credited as a source without a licence being claimed for it. Where a resource asks us to disclose changes, the change is stated in the same row.
The changes listed here are everything we may do to a resource, not what any one run did. Which of them applied is a fact about a particular analysis: a run that never scores pathway activity does not touch PROGENy, and a human run does not use the mouse ortholog table. Every report and reference bundle therefore carries its own notices, listing only the resources that run actually used and only the changes it actually made. Where this page and an artifact differ, the artifact is the accurate one for that run, and this page is the wider set it is drawn from.
| Resource | What we use it for | Instrument | Credit |
|---|---|---|---|
| HuBMAP ASCT+B tables / Human Reference Atlas | The canonical marker membership behind our cell-type marker panels: which genes are treated as defining for a cell type. | CC BY 4.0 | HuBMAP Consortium, Anatomical Structures, Cell Types and Biomarkers (ASCT+B) tables, Human Reference Atlas. Modified: capped at 8 genes per cell type; re-ordered by our own contrast statistic against the other cell types here. |
| CZ CELLxGENE CellGuide | Republishes the ASCT+B tables as its canonical markers, and supplies the computational marker rankings we order our panels by. | no licence asserted We searched and found no published grant. | Marker data obtained from CZ CELLxGENE CellGuide. Modified: top-ranked genes per cell type, re-ordered by own-tissue contrast. |
| Cell Ontology (CL) | The controlled vocabulary our cell-type labels are expressed in, and the tree the marker panels are inherited along. | CC BY 4.0 | Cell Ontology (CL), OBO Foundry. |
| Uberon | The anatomy vocabulary our tissue filters are expressed in. It is what carves an organ out of the multi-tissue atlases in a cohort. | CC BY 3.0 | Uberon multi-species anatomy ontology, OBO Foundry. |
| STRING v12.0 | The protein-protein interaction network behind the interaction figures in the query report. | CC BY 4.0 | STRING v12.0, https://string-db.org/ Modified: combined score of at least 400; identifiers remapped to gene symbols, unmappable edges dropped; reciprocal edges deduplicated at their maximum score; scores rescaled from 0-1000 to 0-1; singleton nodes hidden in the figure. |
| Gene Ontology, via Broad Institute MSigDB | The gene-set panel behind the pathway-activity scores. | CC BY 4.0 CC BY 4.0, Broad Institute MSigDB | Gene Ontology Consortium biological-process terms, obtained through MSigDB. MSigDB is (c) 2004-2025 Broad Institute, MIT and Regents of the University of California. Modified: biological-process collection only; gene sets outside 15-500 genes dropped; GOBP_ prefix stripped from set names. Gene Ontology asks to be cited by release date and Zenodo DOI. Our pinned copy of this panel records neither, so this entry credits the Consortium without naming a release. |
| PROGENy | The pathway-footprint weights behind the pathway-activity scores. | Apache-2.0 | Schubert M et al. 2018, Nature Communications, https://doi.org/10.1038/s41467-017-02391-6 Modified: top 500 weighted genes per pathway. |
| CollecTRI | The transcription-factor to target regulons behind the transcription-factor activity scores. | CC BY 4.0 | Mueller-Dott S et al. 2023, Nucleic Acids Research, https://doi.org/10.1093/nar/gkad841. Regulons from Zenodo record 8192729, nine named creators. Modified: edges attested by TRRUST or DoRothEA-A removed. |
| Mouse Genome Informatics (MGI), The Jackson Laboratory | The mouse gene nomenclature in our human-mouse ortholog table, which is what lets a human marker panel be applied to a mouse reference. | CC BY 4.0 | Mouse gene nomenclature from Mouse Genome Informatics (MGI), The Jackson Laboratory, Bar Harbor, Maine. |
| Ensembl BioMart | The source of the human-mouse ortholog table itself. | no restriction stated | Human-mouse orthologues exported from Ensembl BioMart. |
Software
The open-source libraries VarnaOps runs on. This is a credit, not a notice: every library below runs on our own servers and none of it is handed to you, so the notice terms in the MIT, BSD and Apache licences never attach. Where a licence obligation does attach, the notice is served with the thing that triggers it, under third-party notices.
VarnaOps runs on 39 third-party libraries across 7 areas: scverse and the AnnData ecosystem · NVIDIA RAPIDS · Deep learning · Scientific Python · Clustering, embedding and statistics · Data, storage and cloud · Interface. They are counted here by the licence each installed distribution declares.
| Declared licence | Libraries |
|---|---|
| BSD | 17 |
| Apache-2.0 | 11 |
| MIT | 7 |
| GPL | 2 |
| PSF | 1 |
| not stated in package metadata | 1 |
Download the full list (39 libraries, CSV). It carries each library's distribution name, version, role, and the licence string exactly as that distribution declares it.
umap-learn declares no specific licence in package metadata, only that it is OSI approved. We credit it here and state no licence for it, the same rule we apply to data sources.
Copyleft. leidenalg (GPL-3.0-or-later), python-igraph (GNU General Public License (GPL)) are copyleft. They run on our servers and are not distributed, so their source-offer terms are not triggered. This is a factual statement about how we run today, not a permanent one: see the note below.
Citing the underlying work
If you publish results produced with a VarnaOps reference, cite the original studies, not us, for the data. Each collection above links to its publication DOI. The Census provides a full citation string per dataset version, and our reference build manifests carry it verbatim alongside the dataset identifiers, so the exact version used in any run is recoverable from the run itself.
What this page covers
This page is attribution: the source data and the external resources we build on, with the terms each is available under. It is not a description of how a reference is built. The harmonization that maps each study's original labels onto our cell-type vocabulary, the cohort selection rules and the derivation of the marker panels are our own work and are not published here. The vocabulary a model can actually assign is on the model weights page, and the methods are summarized under methods.
The software we build on is credited above, under software. Only the components we actually serve to a browser carry a notice obligation, and those live under third-party notices.
Reference data itself is never redistributed by us: neither the run bundle nor the model release contains it. Everything above is obtainable directly from the Census.