Skip to content

Repository files navigation

MOSAIC Consortium Manuscript (2026) — Figure Generation

This repository contains the scripts used to generate figures and supplementary tables from the MOSAIC Consortium 2026 paper. Each script reads a small set of pre-computed input files (described below) and writes the resulting figures and tables to outputs/.

Raw data can be obtained from EGA.

The pipeline is split into four scripts:

Script Theme
figure_pathway_analysis.py Single-cell pathway/gene-set analyses on Chromium scRNA-seq
figure_spatial_tme.py Visium spatial TME co-localisation analyses
figure_aPD1_signature.py Visium anti-PD1 response signature scoring
figure_cross_modality.py Malignant-subset statistics and cross-modality heterogeneity correlations

Repository structure

.
├── README.md
├── Makefile
├── pyproject.toml
├── config.py                       # Centralised paths and analysis parameters
├── utils.py                        # Shared helper functions
├── visium_sample.py                # VisiumSample class for loading Visium zarr data
├── visium_pipeline.py              # Shared Visium analysis pipeline (loaded by the two Visium scripts)
├── figure_pathway_analysis.py
├── figure_spatial_tme.py
├── figure_aPD1_signature.py
├── figure_cross_modality.py
├── inputs/                         # Input data files (not distributed with the repo)
└── outputs/                        # Generated figures and tables

Environment setup

The repo uses uv to manage a Python 3.10 virtual environment.

make install

This installs uv if it is not already on the system, then runs uv sync to create .venv/ and install all dependencies from pyproject.toml.

How to reproduce

Run the scripts in dependency order. The cross-modality script reads the heterogeneity-score CSVs produced by the first two scripts, so those must run first.

make figure-pathway          # Chromium pathway/gene-set analyses
make figure-spatial-tme      # Visium TME co-localisation analyses
make figure-aPD1             # Visium anti-PD1 signature analyses (independent)
make figure-cross-modality   # Cross-modality analyses (depends on the first two)

Input data

All input files are expected under inputs/. They are not distributed with this repository. Format requirements are documented below so that anyone holding equivalent data can reproduce the figures.

File Format Source
chromium_MW_anndata.h5ad AnnData (.h5ad). obs includes orig.ident (sample IDs) and per-spot cluster labels prefixed Tu_<MW sample id>_c<NN> for tumour subpopulations. Pre-processed Chromium scRNA-seq AnnData.
visium_zarr_paths.csv Comma-separated. Required column: paths_spt_zarr — one row per Visium sample, each value is the path to that sample's SpatialData zarr store. The sample name is derived from the zarr filename stem. Produced by an upstream Visium QC pipeline.
merge_decisions.csv Comma-separated. Columns: sample_id, patient_id, cluster_1, cluster_2, block_merge_final. One row per cluster pair per sample. Produced by an upstream tumour-subpopulation merge-decision pipeline.
msigdb-hallmark.tsv Tab-separated. Columns: uniprot, genesymbol, entity_type, collection, geneset. The MSigDB Hallmark gene-set table. Public — download from https://static.omnipathdb.org/tables/msigdb-hallmark.tsv.gz.

Outputs

Outputs land under outputs/, organised by analysis theme.

outputs/pathway_analysis/ (from figure_pathway_analysis.py)

For each (gene_set, method) pair the script runs (PROGENy/ULM and Hallmark/GSVA):

  • scRNAseq_progeny_ulm_selected_subpops.png — PROGENy activity dot plot for the selected highly heterogeneous samples.
  • scRNAseq_<gene_set>_<method>_heatmap_<indication>.png — per-indication clustered heatmaps of pathway activities across malignant subsets.
  • scRNAseq_hallmark_gsva_gene_umap_by_indication.png — UMAP overlay of selected hallmark gene expression by indication.
  • <gene_set>_<method>_per_subpop.csv — full per-subpopulation activity table.
  • progeny_ulm_levene_test.csv — Levene's test for variance heterogeneity of PROGENy pathway activities across indications.
  • progeny_ulm_heterogeneity_scores.csv and hallmark_gsva_heterogeneity_scores.csv — per-sample heterogeneity scores (consumed by figure_cross_modality.py).

outputs/spatial_analysis/tme_analysis/ (from figure_spatial_tme.py)

  • tme_correlations_selected_samples_activity-size-correlation-color.png — TME cell-type co-localisation matrix across the selected highly heterogeneous samples.
  • tme_celltype_correlations_<indication>.png — per-indication TME cell-type correlation heatmaps.
  • tme_aggregated_indication_plot.png — aggregated TME correlations per indication.
  • tme_immune_correlations_clustered_<indication>.png — per-indication immune-immune co-localisation heatmaps.
  • umap_tme_composition_by_indication.png — UMAP of TME composition per sample, coloured by indication.
  • umap_tme_composition_by_cleaned_sample_id.png — UMAP coloured by sample.
  • umap_tme_composition_by_subpopulation_<indication>.png — per-indication UMAPs coloured by malignant subpopulation.
  • umap_tme_composition_by_sample_id_<indication>.png — per-indication UMAPs coloured by sample.
  • umap_tme_composition_by_<cell_type>.png — UMAP coloured by per-spot fraction of each cell type.
  • umap_tme_composition_all_celltypes.png — composite of the per-cell-type UMAPs.
  • colocal_progeny_pancancer.png — pan-cancer scatter relating malignant–stromal co-localisation to PROGENy pathway activity.
  • colocal_progeny_per_indication.png — same analysis split by indication.
  • tme_per_subpop.csv — per-subpopulation TME correlation values.
  • tme_levene_test.csv — Levene's test for variance heterogeneity of TME correlations across indications.
  • tme_heterogeneity_scores.csv — per-sample TME heterogeneity scores (consumed by figure_cross_modality.py).
  • colocal_progeny_per_sample.csv — per-sample co-localisation vs. PROGENy table.
  • colocal_progeny_summary_table.csv — summary table of co-localisation vs. PROGENy associations.

outputs/spatial_analysis/aPD1_signature/ (from figure_aPD1_signature.py)

  • aPD1_ecotype_scores.csv — per-ecotype anti-PD1 signature scores.
  • bladder_aPD1_signature.png, dlbcl_aPD1_signature.png, gbm_aPD1_signature.png, mesothelioma_aPD1_signature.png, ovary_aPD1_signature.png — per-indication box-and-whisker plots of signature activity per patient.

outputs/cross_modality/ (from figure_cross_modality.py)

  • malignant_subsets_per_sample.csv — per-sample list and count of malignant subsets.
  • malignant_subsets_overall_stats.csv — total/mean/median malignant-subset counts across the cohort.
  • malignant_subsets_by_indication.csv — per-indication summary statistics.
  • malignant_subsets_single_subset_counts.csv — per-indication count of samples with a single malignant subset.
  • malignant_subsets_bar_chart_with_heterogeneity.png — bar chart of malignant-subset counts per sample with heterogeneity heatmap below.
  • heterogeneity_tme_vs_progeny.png — scatter plot of TME co-localisation heterogeneity vs. PROGENy activity heterogeneity.
  • heterogeneity_tme_vs_gsva.png — scatter plot of TME co-localisation heterogeneity vs. GSVA activity heterogeneity.

About

Code to reproduce results from the MOSAIC Consortium et al. manuscript

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages