Changelog
All notable changes to PySingleCellNet should be listed here. The definition of 'notable' is dynamic.
[Unreleased]
TO FIX
tl.score_gene_sets should add scores, one matrix to .obsm and not as columsn to .obs
Added
tl.train_classifier:random_stateparameter making training reproducible. Seeds gene-pair selection, the randomized profiles and the forest. Defaults toNone(previous unseeded behaviour).ut.save_classifier/ut.load_classifier: persist a trained classifier with a recordedformat_version. Loading validates the required keys and that the estimator is fitted, migrates older payloads, and turns a missing-module unpickling failure into an install hint — so a malformed classifier fails at load rather than deep inside prediction.tl.suggest_pcs_to_drop: automatically identify PCs driven by a confounder gene list (e.g., cell-cycle genes) using FDR, correlation, and impact-score thresholdsut.remove_confounder_pcs: convenience wrapper that flags and removes confounder PCs in one call- Confounder PC removal tutorial notebook (
docs/notebooks/confounder_pcs.ipynb) requirements.txtandenvironment.ymlfor reproducible environment setup- Installation instructions in
docs/install.mdfor pip and conda workflows - Test suite (
tests/) covering classifier, clustering, categorize, utils, and plotting tests/generate_test_data.pyto create small h5ad fixtures from full datasetsCLAUDE.mdfor AI-assisted development guidelines- Gene modules tutorial notebook (
docs/notebooks/gene_modules.ipynb) - Exported
tl.subset_modules_top_genes(was defined but not in__init__.py)
Changed
tl.train_classifier: now returns'ctGraph'(the PAGA cell-type relatedness graph),'topPairGenes'and'format_version'alongside the existing keys. Graph building can be skipped withbuild_graph=False.tl.categorize_classification: acceptsclf=and takes the cell-type graph fromclf['ctGraph']whengraph=is not given, so the Intermediate/Hybrid distinction now travels with a shared classifier instead of requiring the original reference data. An explicitgraph=still wins.tl.correlate_module_scores_with_pcs: now returnsfdr(Benjamini-Hochberg adjusted p-values) andimpact_score(correlation² × variance_ratio) columns- tl.cluster_subclusters layer should be None instead of counts
- Completed Google-style docstrings (Args/Returns/Raises) across all public functions for mkdocs auto-extraction
tl.find_gene_modules:groupbyis now optional (Noneby default); only required whenmean_cluster=True(defaults to'leiden'in that case)
Fixed
tl.train_classifier: gene symbols containing an underscore (e.g.mod_1, common in spatial and module-named datasets) made training fail with an opaqueKeyErrorlisting decoded fragments, because pairs were stored as"geneA_geneB"and recovered by splitting on_. Pair members are now stored structurally intopPairGenes;topPairsremains as a display label. Classifiers are stampedformat_version2, and format 1 classifiers still predict (they are refused only when their own symbols make the labels undecodable).pySingleCellNet.__version__is now bound (read from_version.py, falling back to installed distribution metadata). It was listed in__all__but never imported, so accessing it raisedAttributeError.tl.classify_anndata: classifier genes absent from the query were silently zero-filled, collapsing every pair feature that used them while the scores still looked healthy. Overlap is now reported in.uns['SCN_gene_overlap'], a warning names the missing symbols, and scoring is refused belowmin_gene_overlap(default 0.05). A species/symbol-case mismatch (e.g.ACTBvsActb) is called out explicitly.tl.classify_anndata:nrand > 0raisedValueError: value.index does not match parent's obs namesbecause the randomized profiles were concatenated into the per-cell result. The null distribution now goes to.uns['SCN_random_scores']. Negativenrandraises.tl.classify_anndata: the whole query was densified before being subset to the classifier's genes (~24 GB for a 100k x 30k query to reach a ~0.2 GB matrix). Columns are now sliced before densification, so peak memory scales with the classifier's gene set.tl.train_classifier: no longer leaves a globalwarnings.filterwarnings('ignore')in place, which had muted every warning for the rest of the session — including the new gene-overlap warning.tl.train_classifier: candidate gene lists were de-duplicated withset(), whose string iteration order depends onPYTHONHASHSEED. Two runs in different processes therefore selected different gene pairs even with the RNG seeded, so a classifier could not be reproduced. De-duplication is now order-preserving.tl.categorize_classification: useigraph.Graph.distances()instead of the deprecatedGraph.shortest_paths(), which emitted aDeprecationWarningper cell-type pair per cell (hundreds per call). The distance matrix is also computed once per call rather than re-derived for every cell.tl.comp_ct_thresh: usein adata.obsminstead of the deprecatedadata.obsm_keys().- tl.cluster_subclusters calls sc.pp.pca with mask_var instead of use_highly_variable
tl.train_classifier:n_randauto-calculation now returnsintinstead ofnp.float64(fixes pandas iloc error)tl.comp_ct_thresh: handle cell types with no classified cells (empty array innp.quantile)pl.bar_classifier_f1: fall back toSCN_class_argmax_colorswhenSCN_class_colorsis missing
Planned
- PCA-loading-based gene module discovery (extract per-PC gene modules from PCA loadings)
[0.1.3] - 2025-09-15
Changed
- replace cl with tl
- moved functions to more fitting files, like unused ones to utils.misc.py
- do not export unused functions
Added
- tl.discover_cell_cliques labels cells by consensus cluster labels, kind of
- tl.clustering_quality_vs_nn_summary computes metrics of clustering quality
- tl.cluster_alot
- tl.cluster_subcluster
- resurrected gene clustering functions
- notebook tutorial on cell clustering
Fixed
filter_anndata_slotsto handle .uns and dependencies across slotsclassify_anndatabug that prevented writing h5ad. see https://github.com/CahanLab/PySingleCellNet/issues/13
Removed
- ut.mito_rib_heme
- lots of stuff that is old or has been moved to other packages like STUF
[0.1.2] - 2025-08-05
Removed
- stale functions in old/
[0.1.1] - 2025-08-04
Fixed
- Wrestled pyproject.toml into shape
Removed
- stale functions in old/
[0.1.0] - 2025-08-04
Added
- new versioning system
Fixed
- typos in docs