mantispy.tl.cluster_composition#
- mantispy.tl.cluster_composition(adata, cluster_key='leiden', by=('Metadata_Plate', 'Metadata_Well'), reference='negcon')[source]#
Fraction of each well’s cells in each cluster, as a well-level object.
- Parameters:
adata (
AnnData) – Clustered single-cell object.cluster_key (
str(default:'leiden')) –obscolumn holding the cluster label.by (
Sequence[str] |str(default:('Metadata_Plate', 'Metadata_Well'))) – Columns defining a row of the result, normally the well.reference (
str|None(default:'negcon')) – Controls to test each well’s composition against, orNoneto skip the test.
- Return type:
- Returns:
A new object with wells as rows and clusters as columns, holding the fraction of each well’s cells in each cluster. A cell the clustering left unassigned is left out of the fractions, and a group in which no cell was assigned is left out altogether. It is a well-level mantispy object, so
map(),normalize()and the plots accept it.uns["mantispy"]["composition_test"]holds a chi-square test of each well against the pooled control composition, withgroup,statistic,pvalueandqvalue.uns["mantispy"]["composition_dispersion"]holds the factor the controls’ own spread contributed, described below.- Raises:
KeyError –
obshas no columncluster_key.ValueError – No cell carries a cluster, so there is no composition to take.
Notes
A well with few cells has a noisy composition. The chi-square test is computed on counts and accounts for this, but the fractions in
Xdo not. Filter withwell_qc()first.Chi-square alone asks whether a well’s cells are a multinomial draw from the control composition, and wells vary beyond that: seeding, position and edge effects all move a composition without any perturbation. The statistic is therefore divided by the dispersion the control wells show,
mean(control statistic) / df, floored at one. Without that correction, wells drawn from a single composition with mild jitter were called at q = 9e-10, 10 of 20 of them.The dispersion is estimated leave-one-out (each control well is scored against the pooled composition of the other controls, never one that includes itself), and the scaled statistic is referred to an F distribution rather than chi-square, so the uncertainty in that estimate widens the tail. Without either, the pure-null false positive rate ran above nominal and worse with fewer controls (0.138, 0.092, 0.070, 0.059 at 8, 16, 32, 64 control wells against a nominal 0.05); with both it sits at or below nominal. Below 8 wells the estimate is too noisy to trust and the function warns. Power falls accordingly: a composition shift of a few percentage points is not separable from well-to-well variation, and reporting it as significant was the bug.
Clusters no control cell reached are left out of the test, since the controls give them no expected frequency. Their fractions stay in
X, andsubpopulation_hits()compares within a cluster.Metadata_CellCountis every cell of the well, sincecytotoxicity()reads it to tell a hit from cell loss. The fractions are over the cells the clustering assigned, counted inMetadata_ClusteredCellCount; the two differ when the clustering left cells out, so recovering the counts fromXneeds the latter.The test holds one row per well, in the order of the rows of the returned object. A well with no cells in the clusters the controls occupy gets
NaN. So does every well when the controls occupy fewer than two clusters, since a chi-square over a single category has no degrees of freedom. That case warns, because the table keeps its row per well and would otherwise read as no well’s composition differing.