All benchmarks

Cyto Batch Integration

in development
Release v0.0.1-rc2rc v0.0.1-rc1rc

Benchmarking of batch integration algorithms for cytometry data.

20 methods
5 control methods
2 datasets
9 metrics
2 releases
Task repository MIT v0.0.1-rc2

Cytometry is a non-sequencing single cell profiling technique commonly used in clinical studies. It is very sensitive to batch effects, which can lead to biases in the interpretation of the result. Batch integration algorithms are often used to mitigate this effect.

In this project, we are building a pipeline for reproducible and continuous benchmarking of batch integration algorithms for cytometry data. As input, methods require cleaned and normalised (using arc-sinh or logicle transformation) data with multiple batches, cell type labels, and biological subjects, with paired samples from a subject profiled across multiple batches. The batch integrated output must be an integrated marker by cell matrix stored in Anndata format. All markers in the input data must be returned, regardless of whether they were integrated or not. This output is then evaluated using metrics that assess how well the batch effects were removed and how much biological signals were preserved.

Contributors

  • Luca Leomazzi
    authormaintainer
  • Givanna Putri
    authormaintainer
  • Robrecht Cannoodt
    author
  • Katrien Quintelier
    contributor
  • Sofie Van Gassen
    contributor

Leaderboard

Methods ranked by scaled overall mean. Each cell encodes a score from 0 to 1 by size and intensity.

QC: Normalisation Visualisation 9 plots

Per metric: points placed by control-anchored scaled score (x); dashed lines mark scaled 0 and 1 (worst/best control); the lower axis shows the raw score. Points beyond [-0.2, 1.2] are clamped to the edge as triangles. Hover a dot or line to highlight it and read details.

methodcontrol
  • Average Batch R-squaredlower better
    Perfect IntegrationPerfect IntegrationShuffle Integration G…Shuffle Integration GloballyShuffle Integration W…Shuffle Integration Within Cell TypecyCombine (all-contro…cyCombine (all-controls, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)Shuffle Integration W…Shuffle Integration Within BatchesHarmonypyHarmonypyCombatCombatLimma removeBatchEffe…Limma removeBatchEffectGaussNormGaussNormNo IntegrationNo IntegrationcyCombine (no-control…cyCombine (no-controls, to-goal)Batchadjust with one …Batchadjust with one controlBatchadjust with all …Batchadjust with all controlsCytoNorm (no-controls…CytoNorm (no-controls, to-goal)Seurat RPCA (to-goal)Seurat RPCA (to-goal)Seurat RPCA (to-middl…Seurat RPCA (to-middle)0.0640.0480.0320.016000.250.50.751rawscaled
  • cLisihigher better
    Batchadjust with all …Batchadjust with all controlsBatchadjust with one …Batchadjust with one controlCombatCombatCytoNorm (all-control…CytoNorm (all-controls, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-middle)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)GaussNormGaussNormHarmonypyHarmonypyLimma removeBatchEffe…Limma removeBatchEffectNo IntegrationNo IntegrationPerfect IntegrationPerfect IntegrationSeurat RPCA (to-goal)Seurat RPCA (to-goal)Seurat RPCA (to-middl…Seurat RPCA (to-middle)Shuffle Integration W…Shuffle Integration Within Cell TypeShuffle Integration G…Shuffle Integration GloballyShuffle Integration W…Shuffle Integration Within Batches0.840.880.920.96100.250.50.751rawscaled
  • EMD Mean Horizontallower better
    Perfect IntegrationPerfect IntegrationShuffle Integration W…Shuffle Integration Within Cell TypeShuffle Integration G…Shuffle Integration GloballyCytoNorm (no-controls…CytoNorm (no-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-goal)cyCombine (one-contro…cyCombine (one-control, to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)HarmonypyHarmonypyCombatCombatShuffle Integration W…Shuffle Integration Within BatchesLimma removeBatchEffe…Limma removeBatchEffectGaussNormGaussNormNo IntegrationNo IntegrationBatchadjust with all …Batchadjust with all controlsBatchadjust with one …Batchadjust with one controlcyCombine (no-control…cyCombine (no-controls, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)Seurat RPCA (to-goal)Seurat RPCA (to-goal)Seurat RPCA (to-middl…Seurat RPCA (to-middle)0.1540.1150.0770.038000.250.50.751rawscaled
  • EMD Mean Verticallower better
    Shuffle Integration W…Shuffle Integration Within Cell TypeShuffle Integration G…Shuffle Integration GloballyShuffle Integration W…Shuffle Integration Within BatchesPerfect IntegrationPerfect IntegrationcyCombine (no-control…cyCombine (no-controls, to-goal)cyCombine (no-control…cyCombine (no-controls, to-middle)CytoNorm (no-controls…CytoNorm (no-controls, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-goal)cyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)cyCombine (one-contro…cyCombine (one-control, to-middle)Seurat RPCA (to-middl…Seurat RPCA (to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)Seurat RPCA (to-goal)Seurat RPCA (to-goal)CytoNorm (one-control…CytoNorm (one-control, to-goal)HarmonypyHarmonypyCombatCombatLimma removeBatchEffe…Limma removeBatchEffectGaussNormGaussNormNo IntegrationNo IntegrationBatchadjust with one …Batchadjust with one controlBatchadjust with all …Batchadjust with all controls0.1620.1260.090.0530.01700.250.50.751rawscaled
  • FlowSOM Mean Mapping Similarityhigher better
    Perfect IntegrationPerfect IntegrationShuffle Integration W…Shuffle Integration Within Cell TypeHarmonypyHarmonypyCombatCombatCytoNorm (no-controls…CytoNorm (no-controls, to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-goal)cyCombine (all-contro…cyCombine (all-controls, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-goal)Limma removeBatchEffe…Limma removeBatchEffectcyCombine (one-contro…cyCombine (one-control, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-goal)Seurat RPCA (to-goal)Seurat RPCA (to-goal)GaussNormGaussNormSeurat RPCA (to-middl…Seurat RPCA (to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)cyCombine (no-control…cyCombine (no-controls, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)Batchadjust with one …Batchadjust with one controlBatchadjust with all …Batchadjust with all controlsNo IntegrationNo IntegrationShuffle Integration G…Shuffle Integration GloballyShuffle Integration W…Shuffle Integration Within Batches84.23888.17992.11996.0610000.250.50.751rawscaled
  • Functional Marker Preservation (Cohen's d)lower better
    Perfect IntegrationPerfect IntegrationcyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-goal)cyCombine (one-contro…cyCombine (one-control, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)cyCombine (one-contro…cyCombine (one-control, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-middle)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)No IntegrationNo IntegrationCytoNorm (no-controls…CytoNorm (no-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-goal)HarmonypyHarmonypycyCombine (no-control…cyCombine (no-controls, to-middle)Batchadjust with one …Batchadjust with one controlBatchadjust with all …Batchadjust with all controlsLimma removeBatchEffe…Limma removeBatchEffectCombatCombatSeurat RPCA (to-goal)Seurat RPCA (to-goal)Seurat RPCA (to-middl…Seurat RPCA (to-middle)GaussNormGaussNormShuffle Integration W…Shuffle Integration Within Cell TypeShuffle Integration W…Shuffle Integration Within BatchesShuffle Integration G…Shuffle Integration Globally4.0033.1512.31.4480.59600.250.50.751rawscaled
  • Functional Marker Preservation (Wilcoxon)higher better
    Perfect IntegrationPerfect IntegrationcyCombine (no-control…cyCombine (no-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-goal)cyCombine (all-contro…cyCombine (all-controls, to-goal)cyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (one-contro…cyCombine (one-control, to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)cyCombine (no-control…cyCombine (no-controls, to-goal)Batchadjust with all …Batchadjust with all controlsCytoNorm (no-controls…CytoNorm (no-controls, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-middle)HarmonypyHarmonypyNo IntegrationNo IntegrationBatchadjust with one …Batchadjust with one controlCombatCombatLimma removeBatchEffe…Limma removeBatchEffectSeurat RPCA (to-goal)Seurat RPCA (to-goal)GaussNormGaussNormSeurat RPCA (to-middl…Seurat RPCA (to-middle)Shuffle Integration G…Shuffle Integration GloballyShuffle Integration W…Shuffle Integration Within BatchesShuffle Integration W…Shuffle Integration Within Cell Type00.250.50.75100.250.50.751rawscaled
  • iLisihigher better
    Shuffle Integration G…Shuffle Integration GloballyShuffle Integration W…Shuffle Integration Within Cell TypePerfect IntegrationPerfect IntegrationCytoNorm (no-controls…CytoNorm (no-controls, to-middle)Seurat RPCA (to-middl…Seurat RPCA (to-middle)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)Seurat RPCA (to-goal)Seurat RPCA (to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-goal)CytoNorm (all-control…CytoNorm (all-controls, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-middle)CytoNorm (one-control…CytoNorm (one-control, to-goal)HarmonypyHarmonypyCombatCombatLimma removeBatchEffe…Limma removeBatchEffectGaussNormGaussNormBatchadjust with one …Batchadjust with one controlBatchadjust with all …Batchadjust with all controlsShuffle Integration W…Shuffle Integration Within BatchesNo IntegrationNo Integration6.7e-30.2230.4390.6560.87200.250.50.751rawscaled
  • Ratio Consistent Peakshigher better
    No IntegrationNo IntegrationLimma removeBatchEffe…Limma removeBatchEffectPerfect IntegrationPerfect IntegrationCombatCombatHarmonypyHarmonypyBatchadjust with all …Batchadjust with all controlsBatchadjust with one …Batchadjust with one controlSeurat RPCA (to-middl…Seurat RPCA (to-middle)cyCombine (one-contro…cyCombine (one-control, to-goal)Seurat RPCA (to-goal)Seurat RPCA (to-goal)cyCombine (all-contro…cyCombine (all-controls, to-middle)cyCombine (all-contro…cyCombine (all-controls, to-goal)cyCombine (one-contro…cyCombine (one-control, to-middle)cyCombine (no-control…cyCombine (no-controls, to-middle)CytoNorm (all-control…CytoNorm (all-controls, to-middle)cyCombine (no-control…cyCombine (no-controls, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-middle)GaussNormGaussNormCytoNorm (all-control…CytoNorm (all-controls, to-goal)CytoNorm (one-control…CytoNorm (one-control, to-goal)CytoNorm (no-controls…CytoNorm (no-controls, to-middle)CytoNorm (no-controls…CytoNorm (no-controls, to-goal)Shuffle Integration W…Shuffle Integration Within Cell TypeShuffle Integration G…Shuffle Integration GloballyShuffle Integration W…Shuffle Integration Within Batches0.6440.7330.8220.911100.250.50.751rawscaled
QC: Indicator table 17 errors22 warnings

Automated checks on the benchmark run and its results: missing values, score scaling, metric ranges and similar. Errors are high-severity issues that usually need a maintainer's attention; warnings are lower-severity signals. Findings that are expected for this task are listed separately as silenced.

17 high-severity issues need review. 517 of 556 checks passed.

  • error Raw results Dataset 'human_blood_mass_cytometry' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Dataset: human_blood_mass_cytometry Number of results: 119 Expected number of results: 225 Percentage missing: 47%

  • error Raw results Method 'cycombine_no_controls_to_mid' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cycombine_no_controls_to_mid Number of results: 11 Expected number of results: 18 Percentage missing: 39%

  • error Raw results Method 'cycombine_no_controls_to_goal' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cycombine_no_controls_to_goal Number of results: 11 Expected number of results: 18 Percentage missing: 39%

  • error Raw results Method 'cycombine_one_control_to_mid' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cycombine_one_control_to_mid Number of results: 11 Expected number of results: 18 Percentage missing: 39%

  • error Raw results Method 'cycombine_one_control_to_goal' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cycombine_one_control_to_goal Number of results: 11 Expected number of results: 18 Percentage missing: 39%

  • error Raw results Method 'cycombine_all_controls_to_mid' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cycombine_all_controls_to_mid Number of results: 11 Expected number of results: 18 Percentage missing: 39%

  • error Raw results Method 'cycombine_all_controls_to_goal' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cycombine_all_controls_to_goal Number of results: 11 Expected number of results: 18 Percentage missing: 39%

  • error Raw results Metric 'emd_mean_ct_vert' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Metric: emd_mean_ct_vert Number of results: 25 Expected number of results: 50 Percentage missing: 50%

  • error Raw results Metric 'iLisi' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Metric: iLisi Number of results: 19 Expected number of results: 50 Percentage missing: 62%

  • error Raw results Metric 'functional_marker_preservation_wilcoxon' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Metric: functional_marker_preservation_wilcoxon Number of results: 25 Expected number of results: 50 Percentage missing: 50%

  • error Raw results Metric 'functional_marker_preservation_cohens_d' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Metric: functional_marker_preservation_cohens_d Number of results: 25 Expected number of results: 50 Percentage missing: 50%

  • error Raw results Metric 'emd_mean_ct_vert' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: cyto_batch_integration Metric: emd_mean_ct_vert Control method scores: 5 Expected control method scores: 10 Percentage succeeded: 50%

  • error Raw results Metric 'iLisi' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: cyto_batch_integration Metric: iLisi Control method scores: 5 Expected control method scores: 10 Percentage succeeded: 50%

  • error Raw results Metric 'functional_marker_preservation_wilcoxon' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: cyto_batch_integration Metric: functional_marker_preservation_wilcoxon Control method scores: 5 Expected control method scores: 10 Percentage succeeded: 50%

  • error Raw results Metric 'functional_marker_preservation_cohens_d' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: cyto_batch_integration Metric: functional_marker_preservation_cohens_d Control method scores: 5 Expected control method scores: 10 Percentage succeeded: 50%

  • error Scaling Metric 'emd_mean_ct_horiz' % outside range

    Percentage of scaled scores outside control range should be less than 10% Task: cyto_batch_integration Metric: emd_mean_ct_horiz Inside range: 0 Scaled scores: 50 Percentage outside: 100%

  • error Scaling Metric 'emd_mean_ct_vert' % outside range

    Percentage of scaled scores outside control range should be less than 10% Task: cyto_batch_integration Metric: emd_mean_ct_vert Inside range: 0 Scaled scores: 25 Percentage outside: 100%

Show 22 warnings
  • warning Raw results Task number of results

    Number of results should be equal to #datasets × #methods × #metrics Task: cyto_batch_integration Number of results: 332 Number of datasets: 2 Number of methods: 25 Number of metrics: 9 Expected number of results: 450

  • warning Raw results Method 'shuffle_integration_globally' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: shuffle_integration_globally Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'shuffle_integration_within_batch' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: shuffle_integration_within_batch Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'shuffle_integration_within_cell_type' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: shuffle_integration_within_cell_type Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'no_integration' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: no_integration Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'perfect_integration' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: perfect_integration Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'combat' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: combat Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'gaussnorm' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: gaussnorm Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'batchadjust_one_control' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: batchadjust_one_control Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'batchadjust_all_controls' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: batchadjust_all_controls Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'cytonorm_no_controls_to_mid' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cytonorm_no_controls_to_mid Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'cytonorm_all_controls_to_mid' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cytonorm_all_controls_to_mid Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'cytonorm_one_control_to_mid' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cytonorm_one_control_to_mid Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'cytonorm_no_controls_to_goal' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cytonorm_no_controls_to_goal Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'cytonorm_all_controls_to_goal' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cytonorm_all_controls_to_goal Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'cytonorm_one_control_to_goal' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: cytonorm_one_control_to_goal Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'harmonypy' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: harmonypy Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'limma_remove_batch_effect' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: limma_remove_batch_effect Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'rpca_to_goal' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: rpca_to_goal Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Method 'rpca_to_mid' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Method: rpca_to_mid Number of results: 14 Expected number of results: 18 Percentage missing: 22%

  • warning Raw results Metric 'cLisi' % missing

    Percentage of missing results should be less than 10% Task: cyto_batch_integration Metric: cLisi Number of results: 38 Expected number of results: 50 Percentage missing: 24%

  • warning Raw results Metric component 'lisi' % failed

    Percentage of failed processes should be less than 10% Task: cyto_batch_integration Metric component: lisi Succeeded processes: 38 Attempted processes: 50 Percentage failed: 24%

Method info 20

Batchadjust with all control sample.

CytofBatchadjust corrects batch effects across cytometry data by aligning signal intensity peaks for each channel across batches. The algorithm uses technical replicates included in each barcode set as reference points to anchor each batch. Adjustments factors are calibrated using anchor samples representing each barcode set, then applied to all samples composing a batch.

This implementation uses samples from each group as control samples.

Batchadjust with all control sample.

CytofBatchadjust corrects batch effects across cytometry data by aligning signal intensity peaks for each channel across batches. The algorithm uses technical replicates included in each barcode set as reference points to anchor each batch. Adjustments factors are calibrated using anchor samples representing each barcode set, then applied to all samples composing a batch.

This implementation uses samples only from one group as control samples.

ComBat batch correction for single-cell data, implemented in the scanpy package

Corrects for batch effects by fitting linear models, gains statistical power via an EB framework where information is borrowed across genes. This uses the implementation combat.py

cyCombine (all-controls, to-goal)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run with all control samples, correcting to a goal batch

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine using all control samples (replicates, in cyCombine terms), with a square SOM grid and correct the batches to batch 1. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (all-controls, to-middle)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run with all control samples, correcting to a midpoint

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine using all control samples (replicates, in cyCombine terms), with a square SOM grid and correct the batches to a midpoint derived from all batches. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (no-controls, to-goal)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run without control samples, correcting to a goal batch

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine without any control samples (replicates, in cyCombine terms), with a square SOM grid and correct the batches to batch 1. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (no-controls, to-middle)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run without control samples, correcting to a midpoint

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine without any control samples (replicates, in cyCombine terms), with a square SOM grid and correct the batches to a midpoint derived from all batches. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (one-control, to-goal)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run with control sample from one group only, correcting to a goal batch

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine using control samples only from one biological group (replicates, in cyCombine terms), with a square SOM grid and correct the batches to batch 1. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

cyCombine (one-control, to-middle)code ↗docs ↗source ↗Pedersen et al., 2022

cyCombine run with control sample from one group only, correcting to a midpoint

cyCombine perform batch integration using self-organizing maps and ComBat. It first uses self-organizing maps (SOM) to group similar cells into clusters, then applies a ComBat-based method to correct batch effects within each cluster. Here, we run cyCombine using control samples only from one biological group (replicates, in cyCombine terms), with a square SOM grid and correct the batches to a midpoint derived from all batches. The size of the SOM grid is varied linearly between value of 6-16, with default set to 8 following the default value in create_som function in cyCombine. Rlen, which is the number of times data is presented to the SOM network and can impact clustering quality. It is varied linearly between value of 10-20 with default set to 10.

CytoNorm run with all control samples, correcting to a goal batch.

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

Here, we run CytoNorm using all control samples available, aligning the batches to batch 1.

The parameter nQ, which specifies the number of quantiles used when computing the splines is varied linearly between value of 80-120, with default set to 99 following the default value provided by CytoNorm. Clustering was performed by FlowSOM. The number of cells clustered by FlowSOM is set to be number of cells in the smallest control samples or 1,000,000, whichever is the smaller, multiplied by how many control samples there are in the data. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm.

CytoNorm (all-controls, to-middle)code ↗docs ↗source ↗Van Gassen et al., 2019

CytoNorm run with all control samples, correcting to a midpoint.

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

Here, we run CytoNorm using all control samples available, aligning to a midpoint derived from all batches.

The parameter nQ, which specifies the number of quantiles used when computing the splines is varied linearly between value of 80-120, with default set to 99 following the default value provided by CytoNorm. Clustering was performed by FlowSOM. The number of cells clustered by FlowSOM is set to be number of cells in the smallest control samples or 1,000,000, whichever is the smaller, multiplied by how many control samples there are in the data. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm.

CytoNorm run without control samples, correcting to a goal batch.

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

In this CytoNorm version, an aggregate of each batch is created and subsequently used as a proxy for the control samples, aligning the batches to batch 1.

The number of cells used to create an aggregate is set as the number of cells in the smallest sample or 1,000,000, whichever is the smaller, multiplied by how many samples there are in the batch. Clustering was performed by FlowSOM. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm. The number of cells clustered by FlowSOM is set as the number of cells in the smallest aggregate or 1,000,000, whichever is the smaller, multiplied by how many batches there are in the data (because 1 aggregate = 1 batch).

CytoNorm (no-controls, to-middle)code ↗docs ↗source ↗Van Gassen et al., 2019

CytoNorm run without control samples, correcting to a midpoint.

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

In this CytoNorm version, an aggregate of each batch is created and subsequently used as a proxy for the control samples, aligning to a midpoint derived from all batches.

The number of cells used to create an aggregate is set as the number of cells in the smallest sample or 1,000,000, whichever is the smaller, multiplied by how many samples there are in the batch. Clustering was performed by FlowSOM. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm. The number of cells clustered by FlowSOM is set as the number of cells in the smallest aggregate or 1,000,000, whichever is the smaller, multiplied by how many batches there are in the data (because 1 aggregate = 1 batch).

CytoNorm run with control sample from one group only, correcting to a goal batch

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

Here, we run CytoNorm using only control samples from one donor, aligning the batches to batch 1.

The parameter nQ, which specifies the number of quantiles used when computing the splines is varied linearly between value of 80-120, with default set to 99 following the default value provided by CytoNorm. Clustering was performed by FlowSOM. The number of cells clustered by FlowSOM is set to be number of cells in the smallest control samples or 1,000,000, whichever is the smaller, multiplied by how many control samples there are in the data. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm.

CytoNorm (one-control, to-middle)code ↗docs ↗source ↗Van Gassen et al., 2019

CytoNorm run with control sample from one group only, correcting to a midpoint

CytoNorm corrects batch effects by using reference control samples (aliquots of one sample, technical replicates) included with each batch. It clusters cells using FlowSOM, then trains a model on the control samples to learn how marker expression distributions differ across batches for each population. It then uses splines to align these distributions to a common reference (either a midpoint derived from all batches or to a single batch).

Here, we run CytoNorm using only control samples from one donor, aligning to a midpoint derived from all batches.

The parameter nQ, which specifies the number of quantiles used when computing the splines is varied linearly between value of 80-120, with default set to 99 following the default value provided by CytoNorm. Clustering was performed by FlowSOM. The number of cells clustered by FlowSOM is set to be number of cells in the smallest control samples or 1,000,000, whichever is the smaller, multiplied by how many control samples there are in the data. The size of the SOM grid is varied linearly between value of 6-16, with default set to 15 following the default value provided by CytoNorm. The number of metaclusters is varied linearly between value of 8-20, with default set to 10 following the default value provided by CytoNorm.

Batch effect correction using a per‐channel basis normalization method (gaussNorm)

This method batch-normalizes a set of cytometry data samples by identifying and aligning the high density regions (landmarks or peaks) for each channel. The data of each channel is shifted in such a way that the identified high density regions are moved to fixed locations called base landmarks. Normalization is achieved in three phases:

  1. identifying high-density regions (landmarks) for each flowFrame in the flowSet for a single channel
  2. computing the best matching between the landmarks and a set of fixed reference landmarks for each channel called base landmarks
  3. manipulating the data of each channel in such a way that each landmark is moved to its matching base landmark. Please note that this normalization is on a channel-by-channel basis

NOTE: The default implementation uses max.lms=2, although for some channels it is not possible to compute 2 landmarks, resulting in an error. In order to fully automate the batch normalization process, this implementation checks whether it is possible to compute 2 landmarks, and if not, it sets max.lms=1 for that channel.

Harmonypy is a port of the harmony R package

Harmony is a general-purpose R package with an efficient algorithm for integrating multiple data sets. It is especially useful for large single-cell datasets such as single-cell RNA-seq.

Uses a linear model and matrix decomposition to remove batch effects from a dataset

Limma removeBatchEffect is a method that uses a linear model and matrix decomposition to remove batch effects from a dataset. It first fits a linear model to the data, then decomposes the model matrix into a set of orthogonal components. The batch effect is then removed by subtracting the component corresponding to the batch effect from the data.

Batch integrate data to a goal batch using mutual nearest neighbors identified via Seurat reciprocal PCA.

Seurat RPCA performs batch integration by projecting each query dataset into the PCA space of a goal batch, and identifying anchors using reciprocal PCA (RPCA). RPCA identifies mutual nearest neighbors between the goal batch and the remaining batches in their shared low-dimensional space (PCA), which are used to align and integrate the batches into the goal batch.

We ran Seurat RPCA implemented in Seurat v4.4.0, available on their github. This is because subsequent version (>= v5) does not support getting corrected count matrix, only corrected PC space. See: https://github.com/satijalab/seurat/issues/8551.

We varied the number of PCs and nearest neighbours considered when running RPCA.

Batch integrate data to a midpoint using mutual nearest neighbors identified via Seurat reciprocal PCA.

Seurat RPCA performs batch integration by identifying mutual nearest neighbors (anchors) between all batches using reciprocal PCA (RPCA). Instead of merging batches into a single PCA space, each batch is projected into the PCA space of the others, and anchors are found where cells are mutual nearest neighbors across these projections. These anchors are then used to compute batch correction vectors, aligning the batches into a shared corrected space.

We ran Seurat RPCA implemented in Seurat v4.4.0, available on their github. This is because subsequent version (>= v5) does not support getting corrected count matrix, only corrected PC space. See: https://github.com/satijalab/seurat/issues/8551

We varied the number of PCs and nearest neighbours considered when running RPCA.

Control method info 5
No Integration

Control method returning the unintegrated data without performing batch correction.

The component works by reading and writing back the 'unintegrated' data without performing any operation.

Perfect Integration

Act as a positive control by returning samples from the same batch (batch 1) for both splits.

The method returns only samples from one batch (batch 1), for both the splits. This act as a positive control method for the metrics that compare technical replicates between runs ('horizontal' metrics), as the difference between a sample compared to itself should be zero. It also acts as a positive control for 'vertical' metrics, as the samples being returned are from the same batch.

Shuffle Integration Globally

Randomly shuffle cells in the whole dataset.

This negative control randomly shuffles all cells in the input data. Cells lose all identity information (sample, batch, and cell type).

Example: a cell from mouse 3 batch 1 may be reassigned to any sample in the data (mouse 4 batch 1 or mouse 6 batch 2).

Shuffle Integration Within Batches

Randomly reassign cells to any samples within the same batch.

This negative-control method randomly shuffles cells within each batch. Cells lose their sample and cell type identity but remain in their original batch.

Example: A cell from mouse 3 batch 1 may be reassigned to mouse 4 batch 1, but not to another batch (e.g., mouse 6 batch 2).

Shuffle Integration Within Cell Type

Randomly reassign cells to any cell types

This negative-control method randomly shuffles cells within each cell type. Cells retain their cell type identity but lose their sample and batch identity.

Example: A B cell remains a B cell but if may be reassigned from mouse 3 batch 1 to mouse 4 batch 1 or mouse 6 batch 2.

Metric info 9
Average Batch R-squaredlower is betterDraper & Smith, n.d.

Quantifies how strongly the batch covariate explains the variance in the data among technical replicates after correction (by taking into account cell type effect).

First, a simple linear model is fitted for each paired sample and marker to determine the fraction of variance ($R^{2}$) explained by the batch covariate $B$. The average batch R-squared is then computed as the average of the $R^{2}$ values across all paired samples, markers, cell types. As a result, $\overline{R^2_B}_{cell\ type}$ quantifies how much of the total variability in the data is driven by batch effects. Consequently, lower values are desirable.

$\overline{R^2_B}{cell\ type} = \frac{1}{NCM}\sum{\substack{(x_{\mathrm{split1}},,x_{\mathrm{split2}})\ \textit{paired samples}}}^{N} \sum_{j=1}^{C} \sum_{i=1}^{M},R^2!\bigl(\mathrm{marker}_i \mid B\bigr)$

Where:

  • $N$ is the number of paired samples, where $x_{\mathrm{split1}}$ and $x_{\mathrm{split2}}$ are the two technical replicates that have been batch-corrected. Technical replicates belong to different batches.
  • $M$ is the number of markers
  • $C$ is the number of cell types
  • $B$ is the batch covariate

A high value of $\overline{R^2_B}{global}$ or $\overline{R^2_B}{cell\ type}$ indicates that the batch variable explains a large portion of the variance in the data, which indicates a higher level of batch effects. A good performance on $\overline{R^2_B}{global}$ but not on $\overline{R^2_B}{cell\ type}$ might indicate that the batch effect correction is not addressing cell type specific batch effects.

Cell type Lisi score.

Compute the cell type local inverse simpson index (cLISI) for each cell, then return the median as the final score.

EMD Mean Horizontallower is betterRubner et al., 2000

Mean Earth Mover Distance calculated horizontally across donors for each cell type and marker.

Earth Mover Distance (EMD), also known as the Wasserstein metric, measures the difference between two probability distributions.

Here, EMD is used to compare marker expression distributions between paired samples from the same donor quantified across two different batches. For each paired sample, cell type, and marker, the marker expression values are first converted into probability distributions. This is done by binning the expression values into a range from -100 to 100 with a bin width of 0.1. The wasserstein_distance function from SciPy is then used to calculate the EMD between the two probability distributions belonging to the same cell type, marker, and a given paired samples. This is then repeated for every cell type, marker, and paired sample. Finally, the average of all these EMD values is computed and reported as the metric score.

A high score indicates large overall differences in the distributions of marker expressions between the paired samples, suggesting poor batch integration. A low score means the small differences in marker expression distributions between batches, indicating good batch integration.

EMD Mean Verticallower is betterRubner et al., 2000

Mean Earth Mover Distance across batch corrected samples, cell types, and markers.

Earth Mover Distance (EMD), also known as the Wasserstein metric, measures the difference between two probability distributions.

Here, EMD is used to compare marker expression distributions between all integrated samples from the same group. For each pair of samples, cell type, and marker, the marker expression values are first converted into probability distributions. This is done by binning the expression values into a range from -100 to 100 with a bin width of 0.1. The wasserstein_distance function from SciPy is then used to calculate the EMD between the two probability distributions belonging to the same cell type, marker, and a given paired samples. This is then repeated for every cell type, marker, and paired sample. Finally, the average of all these EMD values is computed and reported as the metric score.

A high score indicates overall, there is a large difference in distribution of marker expression after batch integration. A low score means that overall, the samples are well integrated.

FlowSOM Mean Mapping Similarityhigher is betterSofie Van Gassen, 2017Van Gassen et al., 2015

Assess bidirectional FlowSOM tree similarity between splits.

The metric is based on the FlowSOM algorithm, a popular method which uses self-organizing maps for the viasualization/interpretation/clustering of cytometry data. The FlowSOM algorithm creates a tree structure that represents the relationships between different cell populations in the data.

For each paired sample of technical replicates (where 'split1' are integrated technical replicates from split 1 and 'split2' are integrated technical replicates from split 2)

  1. A FlowSOM tree is created using cells in the sample from split 1 and cells from split 2 are mapped onto it.
  2. A FlowSOM tree is created using cells in the sample from split 2 and cells from split 1 are mapped onto it.
  3. For each direction, a similarity measure is computed by comparing cell type proportions of the two splits in each flowsom cluster.

Ideally, the proportions of cell types in the flowsom clusters of the paired samples should be similar, as they are technical replicates.

The FlowSOM mapping similarity measure can be expressed as follows: $\text{FlowSOM mapping similarity} = 100 - \text{FlowSOM mapping dissimilarity}$

The $\text{FlowSOM mapping dissimilarity}$ is:

$\text{FlowSOM mapping dissimilarity} = \sum_{k=1}^{K}w_{k}\sum_{c=1}^{C}\abs{P^{split1}{k,c} - P^{split2}{k,c}}$

Where:

  • $K$ is the number of flowsom clusters
  • $C$ is the number of cell types
  • $w_{k}$ is the weight of flowsom cluster $k$. It refers to the number of cells in flowsom cluster $k$, divided by the total number of cells in the flowsom tree. Note: the flowsom tree contains all the cells of a sample pair (both integrated and validation).
  • $P^{split1}_{k,c}$ is the percentage of cell type $c$ in flowsom cluster $k$ of the split 1 sample
  • $P^{split2}_{k,c}$ is the percentage of cell type $c$ in flowsom cluster $k$ of the split 2 sample

The average FlowSOM mapping similarity among all donors and both mapping directions is computed and reported as the final metric value. Unlabelled cells are excluded from the analysis.

Additionally, for debugging and inspection, the component outputs the per-donor cluster-by-cell-type absolute difference matrices in the output AnnData uns.

Functional Marker Preservation (Cohen's d)lower is betterVirtanen et al., 2020

Mean absolute change in effect size between unintegrated and integrated functional marker group differences.

Measures how much batch integration shifts the effect size of biological group differences in functional marker expression.

Cohen's d = (mean_A − mean_B) / pooled SD is computed per (functional marker, cell type) pair using per-sample mean expressions, restricted to the pairs in sig_unintegrated (i.e. those found significant in both unintegrated batches by the Wilcoxon metric above). This covers every pair with a genuine baseline group difference, including ones where that difference no longer reached significance after integration — that loss of significance is itself a form of effect size drift this metric is meant to capture.

Group A and group B are determined once from the unintegrated data's obs["group"] values, sorted alphabetically, and reused for every batch and split. Since Cohen's d is signed, this is what keeps a positive value meaning "group A higher than group B" consistently across batches and splits, rather than the sign flipping depending on which group happened to be labelled A in a given batch or split.

Ground truth (unintegrated):

  • Cohen's d is computed per batch, then averaged across the two batches.

Integrated:

  • Cohen's d is computed separately for split 1 and split 2 (batch-agnostic).
  • The absolute deviation from the unintegrated ground truth is computed per split, then averaged: mean(|d_s1 − d_unintegrated|, |d_s2 − d_unintegrated|).
  • Taking abs() before averaging prevents opposite-direction errors in the two splits from cancelling each other out.

Score = mean of the above across all (marker, cell type) pairs. A lower score indicates that integration better preserved the original effect sizes.

Functional Marker Preservation (Wilcoxon)higher is betterVirtanen et al., 2020

Proportion of biologically significant functional marker group differences preserved after batch integration.

Assesses whether batch integration preserves biologically meaningful differences in functional marker expression between biological groups (e.g. WT vs KO).

Per-sample mean expression is computed for each (functional marker, cell type) pair. A two-sided Wilcoxon rank-sum test compares the two biological groups. Raw p-values are used without FDR correction, as the minimum achievable p-value with n=3 samples per group is 0.10, making FDR adjustment counterproductive.

Group A and group B are determined once from the unintegrated data's obs["group"] values, sorted alphabetically, and reused for every batch and split. This keeps "group A" and "group B" referring to the same biological group throughout, rather than being re-derived independently per batch/split.

Baseline (unintegrated data):

  • Tests are run separately within each batch.
  • A pair enters the baseline only if significant (p ≤ 0.1) in BOTH batches.

Preservation (integrated data):

  • The same test is run per split, ignoring batch labels.
  • A pair is counted as preserved if significant in BOTH split 1 and split 2.

Score = number of preserved pairs / number of baseline pairs. A higher score indicates more biologically relevant group differences were maintained.

Integration Lisi score.

Compute the integration local inverse simpson index (iLISI) for each cell, then return the median as the final score.

Ratio Consistent Peakshigher is betterVirtanen et al., 2020

Ratio of the number of cell‑type marker‑expression peaks between unintegrated and batch-integrated data.

The metric compares the number of cell type specific marker expression peaks between unintegrated and batch integrated data. The number of peaks is calculated using the scipy.signal.find_peaks function. The (cell type) marker expression values of the two splits are first standardised together, using the pooled mean and standard deviation of the two splits, so that the peak calling thresholds mean the same thing for every marker. They are then smoothed using kernel density estimation (KDE) (scipy.stats.gaussian_kde, evaluated on a grid of 100 points), and peaks are identified using the scipy.signal.find_peaks function. For peak calling, the height parameter is set to 0.1 and the prominence parameter is set to 0.01.

Case Definitions:

  • Case 1: Consistent peaks in both unintegrated AND integrated data (ideal outcome). Example: A marker shows 2 peaks for both splits in unintegrated AND integrated data.
  • Case 2: Consistent peaks in unintegrated data but inconsistent between integrated data and unintegrated data in each split (problematic - batch integration introduced inconsistency). Batch correction broke the consistency. Example 1: A marker shows 2 peaks in both unintegrated splits, but after integration shows 2 peaks in split 1 and 1 peak in split 2. Example 2: A marker shows 2 peaks in both unintegrated splits, but after integration shows 1 peak in both splits
  • Case NGT: Inconsistent peaks in unintegrated data, consistent in integrated data (excluded from metric) OR inconsistent peaks in both unintegrated AND integrated data (excluded from metric). These cases are excluded from calculation because there are differences in unintegrated that cannot be accounted for. Example: A marker is inconsistent in unintegrated data (2 and 3 peaks in split 1 and split 2) and is either inconsistent after integration (1 and 2 peaks in split 1 and split 2) or consistent after integration (2 peaks in both split 1 and split 2).

Ratio of consistent peaks is defined as the number of Case 1 occurrences divided by the total of Case 1 and Case 2 occurrences. Cases NGT (where unintegrated data has inconsistent peaks) are excluded from the denominator, but is reported in the final anndata output for diagnostic. A higher score indicates better performance, meaning there are fewer cases with inconsistent peaks after batch integration.

For edge cases where methods return only zero values for a given marker/donor/cell type, but is not the case in the unintegrated data, it is automatically assigned as Case 2.

Dataset info 2
Human Blood Mass Cytometry unlinked

Mass cytometry dataset of whole blood from 2 human donors. For each donor, 2 samples are included: one unstimulated and one stimulated. For each of these, aliquotes of the same original sample were divided into 2 batches to allow the creation of sample-paired replicates for benchmarking purposes.

CyTOF data (whole blood) from 2 human donors, each with unstimulated and stimulated samples. The stimulated samples were treated with Interferonα (IFNα) and Lipopolysaccharide (LPS). The mass cytometry cytometry panel includes 21 surface markers for cell phenotyping and 14 markers for functional analysis of the signaling responses. The data has been arcsinh transformed with cofactor 5 and pregated to only include events which fall either in the CD66-CD45+ or CD66+ gates.

Mouse Spleen Flow Cytometry unlinked

Flow cytometry data of spleens of 8 mice. For each mouse, aliquotes of the same original sample were divided into 2 batches and measured with 2 different instrument settings to allow the creation of sample-paired replicates for benchmarking purposes.

Flow cytometry data of spleens from 4 WT (IKK2 fl/fl CD11c-cre +/+) and 4 KO (IKK2 fl/fl CD11c-cre Tg/+) B6 mice, measured with a 22-color panel and 2 different instrument settings. Data has been preprocessed (compensated with a batch-specific compensation matrix, logicle transformed, cleaned with PeacoQC and pregated on live single CD45+ cells).

References

  1. Draper, N. R., & Smith, H. (n.d.). Applied regression analysis. John Wiley & Sons.
  2. Sofie Van Gassen, B. C. (2017). FlowSOM. 10.18129/B9.BIOC.FLOWSOM ↗
  3. Van Gassen, S., Callebaut, B., Van Helden, M. J., Lambrecht, B. N., Demeester, P., Dhaene, T., & Saeys, Y. (2015). FlowSOM: Using self‐organizing maps for visualization and interpretation of cytometry data. 10.1002/cyto.a.22625 ↗
  4. Van Gassen, S., Gaudilliere, B., Angst, M. S., Saeys, Y., & Aghaeepour, N. (2019). CytoNorm: A Normalization Algorithm for Cytometry Data. 10.1002/cyto.a.23904 ↗
  5. Van Gassen, S., Gaudilliere, B., Angst, M. S., Saeys, Y., & Aghaeepour, N. (2020). CytoNorm: a normalization algorithm for cytometry data. Cytometry Part A, 97(3), 268–278.
  6. Hahne, F., Khodabakhshi, A. H., Bashashati, A., Wong, C., Gascoyne, R. D., Weng, A. P., Seyfert‐Margolis, V., Bourcier, K., Asare, A., Lumley, T., Gentleman, R., & Brinkman, R. R. (2009). Per‐channel basis normalization methods for flow cytometry data. 10.1002/cyto.a.20823 ↗
  7. Hao, Y., Hao, S., Andersen-Nissen, E., Mauck, W. M., Zheng, S., Butler, A., Lee, M. J., Wilk, A. J., Darby, C., Zager, M., Hoffman, P., Stoeckius, M., Papalexi, E., Mimitou, E. P., Jain, J., Srivastava, A., Stuart, T., Fleming, L. M., Yeung, B., … Satija, R. (2021). Integrated analysis of multimodal single-cell data. Cell, 184(13), 3573-3587.e29. 10.1016/j.cell.2021.04.048 ↗
  8. Johnson, W. E., Li, C., & Rabinovic, A. (2006). Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics, 8(1), 118–127. 10.1093/biostatistics/kxj037 ↗
  9. Korsunsky, I., Millard, N., Fan, J., Slowikowski, K., Zhang, F., Wei, K., Baglaenko, Y., Brenner, M., ru Po-Loh, & Raychaudhuri, S. (2019). Fast, sensitive and accurate integration of single-cell data with Harmony. Nature Methods, 16(12), 1289–1296. 10.1038/s41592-019-0619-0 ↗
  10. Leomazzi, L., Putri, G., Cannoodt, R., Quintelier, K., & Van Gassen, S. (2025). Benchmarking batch integration methods for mass cytometry data. bioRxiv – in Preparation.
  11. Pedersen, C. B., Dam, S. H., Barnkob, M. B., Leipold, M. D., Purroy, N., Rassenti, L. Z., Kipps, T. J., Nguyen, J., Lederer, J. A., Gohil, S. H., Wu, C. J., & Olsen, L. R. (2022). cyCombine allows for robust integration of single-cell cytometry datasets within and across technologies. 10.1038/s41467-022-29383-5 ↗
  12. Ritchie, M. E., Phipson, B., Wu, D., Hu, Y., Law, C. W., Shi, W., & Smyth, G. K. (2015). limma powers differential expression analyses for RNA-sequencing and microarray studies. 10.1093/nar/gkv007 ↗
  13. Rubner, Y., Tomasi, C., & Guibas, L. J. (2000). The Earth Mover’s Distance as a Metric for Image Retrieval. 10.1023/a:1026543900054 ↗
  14. Schuyler, R. P., Jackson, C., Garcia-Perez, J. E., Baxter, R. M., Ogolla, S., Rochford, R., Ghosh, D., Rudra, P., & Hsieh, E. W. (2019). Minimizing batch effects in mass cytometry data. Frontiers in Immunology, 10, 2367.
  15. Schuyler, R. P., Jackson, C., Garcia-Perez, J. E., Baxter, R. M., Ogolla, S., Rochford, R., Ghosh, D., Rudra, P., & Hsieh, E. W. Y. (2019). Minimizing Batch Effects in Mass Cytometry Data. 10.3389/fimmu.2019.02367 ↗
  16. Virtanen, P., Gommers, R., Oliphant, T. E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., van der Walt, S. J., Brett, M., Wilson, J., Millman, K. J., Mayorov, N., Nelson, A. R. J., Jones, E., Kern, R., Larson, E., … Vázquez-Baeza, Y. (2020). SciPy 1.0: fundamental algorithms for scientific computing in Python. 10.1038/s41592-019-0686-2 ↗