I am preparing a research article for submission to a peer-reviewed scientific journal.
The aim of the study is to test, using existing experimental single-cell lineage-tracing data, whether experimentally defined cellular or clonal identity contains a measurable component of transcriptional organization that remains reproducible across substantial perturbation-induced changes in cellular state.
More specifically, the study will test whether the same experimentally tracked clone preserves aspects of the relational organization among transcriptional programs across different environments, beyond what can be explained by conventional similarity in mean gene-expression state, treatment effects, clone size, or other obvious confounders.
The objective is not to find a dataset that merely illustrates the hypothesis. The analysis will be designed as a predefined falsifiable test: if clone-specific relational organization disappears after controlling for ordinary transcriptional similarity and experimental condition, the hypothesis will not be supported.
I am therefore looking for an existing published or unpublished experimental dataset, or for a collaboration with a group that has generated data suitable for such a test.
What the dataset needs to allow
The central comparison is:
same clone across different perturbations
versus
different clones under matched perturbations
The strongest experimental design would start with a common barcoded population that is subsequently divided into several treatment or environmental branches, so that the same clonal identities can be observed under genuinely different conditions.
Minimum requirements for the planned analysis
For this particular study, the dataset should meet the following minimum criteria:
Single-cell transcriptomic measurements, preferably scRNA-seq.
Experimentally measured, heritable clonal identities, such as stable DNA barcodes, CRISPR lineage labels, or another independently measured lineage marker.
Clone identity must be determined independently of transcriptomic similarity. Clones inferred solely from gene-expression similarity are not sufficient.
Clonal labels should be established before the perturbations being compared.
At least 20 independent clones must satisfy the complete analytical requirements below. The relevant number is therefore not the total number of detected barcodes, but the number of clones that can actually be compared across conditions.
Each qualifying clone should be observed in at least three distinct experimental conditions or states.
Ideally, these conditions should represent different perturbation branches originating from a shared starting population, for example:
control / perturbation A / perturbation B
or
baseline / treatment A / treatment B.
A purely longitudinal series may also be useful, but it is less informative if treatment effects cannot be separated from time.
At least 30 QC-passed cells per clone per condition should be available for the qualifying clones.
This threshold is intended for a relatively low-dimensional analysis of relationships among transcriptional programs. It should not be interpreted as sufficient for unrestricted gene-by-gene regulatory-network reconstruction.
Cell-level gene-expression measurements must be available, preferably including raw UMI counts as well as normalized data.
Metadata must permit an unambiguous mapping:
cell ID → clone ID → condition → sample/replicate
Cells with ambiguous, multiple, or low-confidence lineage assignments must be identifiable so that they can be excluded.
Treatment or condition should not be perfectly confounded with a single sequencing batch.
The study should contain biological replication, or there should be a realistic possibility of validating the result in an independent experimental dataset.
The underlying data must be available for independent analysis, either publicly or within a scientific collaboration.
As a rough scale, the minimum design of 20 clones × 3 conditions × 30 cells corresponds to approximately 1,800 lineage-resolved single cells that actually satisfy the comparison criteria. A much larger total dataset may therefore be required.
Datasets containing only clone-abundance measurements, bulk RNA-seq, pseudobulk profiles, precomputed differential-expression tables, transcriptomically inferred clones, or only a few repeatedly observed clones would unfortunately not be sufficient for the primary analysis.
The biological system does not need to be cancer. Relevant systems could include drug response, differentiation, stress adaptation, environmental transitions, developmental systems, immune-cell responses, regeneration, or other experimentally controlled adaptive processes.
What would be tested
The analysis would ask whether a clone-specific relational transcriptional signature remains detectable when the absolute transcriptional state changes.
Importantly, evidence for simple expression similarity between cells of the same clone would not by itself be considered evidence for the proposed effect.
The critical question is whether relational organization contains information about clonal identity beyond mean expression state and experimental condition.
A negative result would also be scientifically informative.
Collaboration
If the dataset is unpublished, I am open to developing the study as a collaborative paper.
Co-authorship would reflect substantive contributions to the experimental data, biological interpretation, analysis, and manuscript.
If you think your dataset may be suitable, there is no need to transfer the full data initially.
For a first assessment, the following four numbers are usually enough:
1. Number of tracked clones
2. Number of experimental conditions or perturbation branches
3. Median number of cells per clone per condition
4. Number of biological replicates
A brief description of the lineage-labeling method and experimental design would also be useful.
Thank you, JH


I see the problem. In a sense, your question is a little ahead of the available datasets, but of course you need the data now. I imagine many groups are still limited by the cost and scale required for this kind of study. Nevertheless, I’ll keep it in mind and let you know if I come across something that looks suitable.
Two papers that might be worth checking are:
Ratz et al., Clonal relations in the mouse brain revealed by single-cell and spatial transcriptomics, Nature Neuroscience, 2022 and Mold et al., Clonally heritable gene expression imparts a layer of diversity within cell types, Cell Systems, 2024
Perhaps they don't fully meet your criteria, but they are very much in this space and might also refer to other papers or datasets that do. Michael Ratz is the person I was thinking of when I read what you need, so hopefully his publications could help you further even if these particular datasets are not completely suitable.