All benchmarks

Spatial Simulators

in development

Assessing the quality of spatial transcriptomics simulators

8 methods
3 control methods
10 datasets
39 metrics
3 releases
Task repository MIT v0.0.1-rc3

Computational methods for spatially resolved transcriptomics (SRT) are frequently developed and assessed through data simulation. The effectiveness of these evaluations relies on the simulation methods' ability to accurately reflect experimental data. However, a systematic evaluation framework for spatial simulators is lacking. Here, we present SpatialSimBench, a comprehensive evaluation framework that assesses 13 simulation methods using 10 distinct STR datasets.

The research goal of this benchmark is to systematically evaluate and compare the performance of various simulation methods for spatial transcriptomics (ST) data. It aims to address the lack of a comprehensive evaluation framework for spatial simulators and explore the feasibility of leveraging existing single-cell simulators for ST data. The experimental setup involves collecting public spatial transcriptomics datasets and corresponding scRNA-seq datasets. The spatial and scRNA-seq datasets can originate from different study but should consist of similar cell types from similar tissues.

Contributors

  • Xiaoqi Liang
    author
  • Yue Cao
    authormaintainer
  • Jean Yang
    author
  • Robrecht Cannoodt
    contributor
  • Sai Nirmayi Yasa
    contributor

Leaderboard

Methods ranked by scaled overall mean. Each cell encodes a score from 0 to 1 by size and intensity.

QC: Normalisation Visualisation 39 plots

Per metric: points placed by control-anchored scaled score (x); dashed lines mark scaled 0 and 1 (worst/best control); the lower axis shows the raw score. Points beyond [-0.2, 1.2] are clamped to the edge as triangles. Hover a dot or line to highlight it and read details.

methodcontrol
  • Adjusted Rand indexhigher better
    SplatterSplattersymsimsymsimSPARsimSPARsimscDesign2scDesign2positivepositiveSRTsimSRTsimzinbwavezinbwavescDesign3 (Poisson)scDesign3 (Poisson)scDesign3 (NB)scDesign3 (NB)negative_shufflenegative_shufflenegative_normalnegative_normal-3.8e-30.2320.4670.7030.93800.250.50.751rawscaled
  • Cell type deconvolution JSDlower better
    positivepositiveSRTsimSRTsimscDesign3 (Poisson)scDesign3 (Poisson)scDesign3 (NB)scDesign3 (NB)zinbwavezinbwavescDesign2scDesign2SPARsimSPARsimnegative_normalnegative_normalnegative_shufflenegative_shuffleSplatterSplattersymsimsymsim0.8260.620.4130.207000.250.50.751rawscaled
  • Cell type deconvolution RMSElower better
    positivepositiveSRTsimSRTsimscDesign3 (Poisson)scDesign3 (Poisson)zinbwavezinbwavescDesign2scDesign2scDesign3 (NB)scDesign3 (NB)SPARsimSPARsimnegative_normalnegative_normalnegative_shufflenegative_shuffleSplatterSplattersymsimsymsim0.2630.1970.1320.066000.250.50.751rawscaled
  • Cosine similarityhigher better
    positivepositiveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)scDesign3 (Poisson)scDesign3 (Poisson)scDesign2scDesign2SPARsimSPARsimzinbwavezinbwavesymsimsymsimSplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normal0.4950.6210.7480.874100.250.50.751rawscaled
  • Effective library sizelower better
    positivepositiveSRTsimSRTsimscDesign2scDesign2zinbwavezinbwavescDesign3 (NB)scDesign3 (NB)SPARsimSPARsimSplatterSplatterscDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimnegative_normalnegative_normalnegative_shufflenegative_shuffle1.6e+31.2e+3772.153351.122-69.90800.250.50.751rawscaled
  • Effective library sizelower better
    positivepositiveSRTsimSRTsimzinbwavezinbwavescDesign2scDesign2SplatterSplatterSPARsimSPARsimscDesign3 (NB)scDesign3 (NB)symsimsymsimscDesign3 (Poisson)scDesign3 (Poisson)negative_shufflenegative_shufflenegative_normalnegative_normal43.62432.71821.81210.906000.250.50.751rawscaled
  • Fraction of zeros per celllower better
    positivepositivezinbwavezinbwaveSRTsimSRTsimscDesign2scDesign2SPARsimSPARsimscDesign3 (NB)scDesign3 (NB)SplatterSplattersymsimsymsimscDesign3 (Poisson)scDesign3 (Poisson)negative_shufflenegative_shufflenegative_normalnegative_normal752.135561.798371.462181.125-9.21200.250.50.751rawscaled
  • Fraction of zeros per celllower better
    positivepositivezinbwavezinbwaveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)scDesign2scDesign2SPARsimSPARsimSplatterSplatterscDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimnegative_shufflenegative_shufflenegative_normalnegative_normal838.497628.873419.249209.624000.250.50.751rawscaled
  • Fraction of zeros per genelower better
    positivepositivezinbwavezinbwavescDesign3 (NB)scDesign3 (NB)scDesign2scDesign2SRTsimSRTsimSPARsimSPARsimscDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimSplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normal9.0e+36.7e+34.5e+32.2e+3-1.3800.250.50.751rawscaled
  • Fraction of zeros per genelower better
    positivepositivezinbwavezinbwavescDesign3 (NB)scDesign3 (NB)scDesign2scDesign2SRTsimSRTsimSPARsimSPARsimscDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimSplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normal309.555232.166154.77877.389000.250.50.751rawscaled
  • Gene Pearson correlationlower better
    positivepositivescDesign3 (Poisson)scDesign3 (Poisson)SPARsimSPARsimscDesign3 (NB)scDesign3 (NB)zinbwavezinbwavesymsimsymsimscDesign2scDesign2SplatterSplatternegative_shufflenegative_shuffleSRTsimSRTsimnegative_normalnegative_normal241.191180.676120.16159.646-0.86900.250.50.751rawscaled
  • Gene Pearson correlationlower better
    positivepositivescDesign3 (Poisson)scDesign3 (Poisson)SPARsimSPARsimzinbwavezinbwavescDesign3 (NB)scDesign3 (NB)symsimsymsimSplatterSplatterSRTsimSRTsimscDesign2scDesign2negative_normalnegative_normalnegative_shufflenegative_shuffle18.69214.0199.3464.673000.250.50.751rawscaled
  • L statisticslower better
    positivepositiveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)scDesign2scDesign2SPARsimSPARsimscDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimSplatterSplatternegative_shufflenegative_shufflezinbwavezinbwavenegative_normalnegative_normal75.80255.24634.6914.133-6.42300.250.50.751rawscaled
  • Library sizelower better
    positivepositiveSPARsimSPARsimzinbwavezinbwaveSRTsimSRTsimSplatterSplatterscDesign3 (NB)scDesign3 (NB)scDesign2scDesign2scDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimnegative_shufflenegative_shufflenegative_normalnegative_normal6.3e+34.7e+33.1e+31.6e+3-4.08500.250.50.751rawscaled
  • Library sizelower better
    positivepositiveSPARsimSPARsimSRTsimSRTsimzinbwavezinbwaveSplatterSplatterscDesign2scDesign2scDesign3 (NB)scDesign3 (NB)scDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimnegative_shufflenegative_shufflenegative_normalnegative_normal27.20320.40213.6016.801000.250.50.751rawscaled
  • Library size vs fraction zerolower better
    positivepositiveSRTsimSRTsimzinbwavezinbwavescDesign3 (NB)scDesign3 (NB)scDesign2scDesign2SPARsimSPARsimscDesign3 (Poisson)scDesign3 (Poisson)SplatterSplattersymsimsymsimnegative_shufflenegative_shufflenegative_normalnegative_normal522.07390.472258.875127.277-4.32100.250.50.751rawscaled
  • Library size vs fraction zerolower better
    positivepositivezinbwavezinbwaveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)scDesign2scDesign2SPARsimSPARsimscDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimSplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normal5.1e+33.8e+32.6e+31.3e+3000.250.50.751rawscaled
  • Mantel statistichigher better
    positivepositiveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)scDesign3 (Poisson)scDesign3 (Poisson)scDesign2scDesign2SPARsimSPARsimzinbwavezinbwavenegative_shufflenegative_shuffleSplatterSplatternegative_normalnegative_normalsymsimsymsim-4.1e-30.2470.4980.749100.250.50.751rawscaled
  • Mean vs fraction zerolower better
    positivepositivezinbwavezinbwaveSPARsimSPARsimscDesign3 (NB)scDesign3 (NB)scDesign2scDesign2SRTsimSRTsimscDesign3 (Poisson)scDesign3 (Poisson)SplatterSplattersymsimsymsimnegative_shufflenegative_shufflenegative_normalnegative_normal564.818423.137281.456139.775-1.90600.250.50.751rawscaled
  • Mean vs fraction zerolower better
    positivepositivezinbwavezinbwavescDesign2scDesign2scDesign3 (NB)scDesign3 (NB)SRTsimSRTsimSPARsimSPARsimsymsimsymsimscDesign3 (Poisson)scDesign3 (Poisson)SplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normal6.4e+34.8e+33.2e+31.6e+3000.250.50.751rawscaled
  • Mean vs variancelower better
    positivepositivescDesign3 (NB)scDesign3 (NB)scDesign3 (Poisson)scDesign3 (Poisson)SRTsimSRTsimzinbwavezinbwavescDesign2scDesign2SPARsimSPARsimSplatterSplatternegative_shufflenegative_shufflesymsimsymsimnegative_normalnegative_normal473.882351.91229.939107.967-14.00400.250.50.751rawscaled
  • Mean vs variancelower better
    positivepositivescDesign2scDesign2scDesign3 (NB)scDesign3 (NB)SRTsimSRTsimSPARsimSPARsimzinbwavezinbwavescDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimSplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normal324.089243.067162.04481.022000.250.50.751rawscaled
  • Moran's Ilower better
    positivepositivescDesign3 (NB)scDesign3 (NB)SRTsimSRTsimSPARsimSPARsimzinbwavezinbwavesymsimsymsimscDesign2scDesign2SplatterSplatterscDesign3 (Poisson)scDesign3 (Poisson)negative_normalnegative_normalnegative_shufflenegative_shuffle167.001124.89382.78540.677-1.43100.250.50.751rawscaled
  • Nearest-neighbour correlationlower better
    positivepositiveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)scDesign2scDesign2symsimsymsimSPARsimSPARsimSplatterSplatterscDesign3 (Poisson)scDesign3 (Poisson)zinbwavezinbwavenegative_normalnegative_normalnegative_shufflenegative_shuffle796.306593.736391.166188.595-13.97500.250.50.751rawscaled
  • Normalised mutual informationhigher better
    SplatterSplattersymsimsymsimSPARsimSPARsimscDesign2scDesign2positivepositiveSRTsimSRTsimzinbwavezinbwavescDesign3 (Poisson)scDesign3 (Poisson)scDesign3 (NB)scDesign3 (NB)negative_shufflenegative_shufflenegative_normalnegative_normal3.0e-40.2290.4590.6880.91700.250.50.751rawscaled
  • Sample Pearson correlationlower better
    positivepositiveSRTsimSRTsimscDesign2scDesign2scDesign3 (NB)scDesign3 (NB)scDesign3 (Poisson)scDesign3 (Poisson)zinbwavezinbwaveSPARsimSPARsimnegative_shufflenegative_shuffleSplatterSplattersymsimsymsimnegative_normalnegative_normal836.01626.812417.613208.415-0.78400.250.50.751rawscaled
  • Sample Pearson correlationlower better
    positivepositiveSRTsimSRTsimscDesign2scDesign2scDesign3 (NB)scDesign3 (NB)scDesign3 (Poisson)scDesign3 (Poisson)SPARsimSPARsimzinbwavezinbwavesymsimsymsimSplatterSplatternegative_normalnegative_normalnegative_shufflenegative_shuffle35.25926.44417.6298.815000.250.50.751rawscaled
  • Scaled mean cellslower better
    positivepositiveSPARsimSPARsimzinbwavezinbwaveSRTsimSRTsimscDesign2scDesign2scDesign3 (NB)scDesign3 (NB)symsimsymsimSplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normalscDesign3 (Poisson)scDesign3 (Poisson)26.13519.44712.766.073-0.61500.250.50.751rawscaled
  • Scaled mean cellslower better
    positivepositiveSPARsimSPARsimzinbwavezinbwaveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)scDesign2scDesign2symsimsymsimnegative_shufflenegative_shufflenegative_normalnegative_normalSplatterSplatterscDesign3 (Poisson)scDesign3 (Poisson)2.2091.6571.1050.552000.250.50.751rawscaled
  • Scaled mean geneslower better
    positivepositivescDesign2scDesign2zinbwavezinbwaveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)SPARsimSPARsimscDesign3 (Poisson)scDesign3 (Poisson)SplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normalsymsimsymsim169.579127.04784.51441.981-0.55200.250.50.751rawscaled
  • Scaled mean geneslower better
    positivepositivescDesign2scDesign2zinbwavezinbwaveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)SPARsimSPARsimscDesign3 (Poisson)scDesign3 (Poisson)SplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normalsymsimsymsim4.8563.6422.4281.214000.250.50.751rawscaled
  • Scaled variance celllower better
    positivepositivenegative_shufflenegative_shuffleSRTsimSRTsimSPARsimSPARsimscDesign3 (NB)scDesign3 (NB)scDesign3 (Poisson)scDesign3 (Poisson)symsimsymsimzinbwavezinbwaveSplatterSplatterscDesign2scDesign2negative_normalnegative_normal1.2e+3909.346605.959302.572-0.81600.250.50.751rawscaled
  • Scaled variance celllower better
    positivepositivenegative_shufflenegative_shuffleSRTsimSRTsimzinbwavezinbwaveSPARsimSPARsimscDesign3 (NB)scDesign3 (NB)scDesign2scDesign2negative_normalnegative_normalSplatterSplattersymsimsymsimscDesign3 (Poisson)scDesign3 (Poisson)16.31712.2378.1584.079000.250.50.751rawscaled
  • Scaled variance geneslower better
    positivepositivescDesign3 (Poisson)scDesign3 (Poisson)SRTsimSRTsimscDesign2scDesign2scDesign3 (NB)scDesign3 (NB)SPARsimSPARsimzinbwavezinbwaveSplatterSplatternegative_shufflenegative_shufflesymsimsymsimnegative_normalnegative_normal4.1e+32.9e+31.8e+3736.772-369.03100.250.50.751rawscaled
  • Scaled variance geneslower better
    positivepositiveSRTsimSRTsimscDesign3 (Poisson)scDesign3 (Poisson)scDesign2scDesign2scDesign3 (NB)scDesign3 (NB)zinbwavezinbwaveSPARsimSPARsimnegative_normalnegative_normalSplatterSplattersymsimsymsimnegative_shufflenegative_shuffle21.68416.26310.8425.421000.250.50.751rawscaled
  • SVG precisionhigher better
    positivepositiveSRTsimSRTsimscDesign2scDesign2scDesign3 (NB)scDesign3 (NB)SPARsimSPARsimscDesign3 (Poisson)scDesign3 (Poisson)zinbwavezinbwavesymsimsymsimSplatterSplatternegative_shufflenegative_shufflenegative_normalnegative_normal00.250.50.75100.250.50.751rawscaled
  • SVG recallhigher better
    positivepositivescDesign3 (Poisson)scDesign3 (Poisson)SRTsimSRTsimscDesign3 (NB)scDesign3 (NB)SPARsimSPARsimzinbwavezinbwavescDesign2scDesign2SplatterSplattersymsimsymsimnegative_shufflenegative_shufflenegative_normalnegative_normal00.250.50.75100.250.50.751rawscaled
  • TMMlower better
    positivepositiveSRTsimSRTsimscDesign2scDesign2scDesign3 (NB)scDesign3 (NB)SPARsimSPARsimzinbwavezinbwavenegative_shufflenegative_shufflesymsimsymsimnegative_normalnegative_normalSplatterSplatterscDesign3 (Poisson)scDesign3 (Poisson)84.24558.11131.9785.844-20.2900.250.50.751rawscaled
  • TMMlower better
    positivepositiveSRTsimSRTsimscDesign3 (NB)scDesign3 (NB)scDesign2scDesign2SPARsimSPARsimzinbwavezinbwaveSplatterSplatternegative_shufflenegative_shufflesymsimsymsimscDesign3 (Poisson)scDesign3 (Poisson)negative_normalnegative_normal9.0116.7594.5062.253000.250.50.751rawscaled
QC: Indicator table 31 errors3 warnings

Automated checks on the benchmark run and its results: missing values, score scaling, metric ranges and similar. Errors are high-severity issues that usually need a maintainer's attention; warnings are lower-severity signals. Findings that are expected for this task are listed separately as silenced.

31 high-severity issues need review. 1097 of 1131 checks passed.

  • error Raw results Metric 'clustering_ari' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: spatial_simulators Metric: clustering_ari Control method scores: 29 Expected control method scores: 30 Percentage succeeded: 97%

  • error Raw results Metric 'clustering_nmi' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: spatial_simulators Metric: clustering_nmi Control method scores: 29 Expected control method scores: 30 Percentage succeeded: 97%

  • error Raw results Metric 'svg_recall' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: spatial_simulators Metric: svg_recall Control method scores: 23 Expected control method scores: 30 Percentage succeeded: 77%

  • error Raw results Metric 'svg_precision' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: spatial_simulators Metric: svg_precision Control method scores: 29 Expected control method scores: 30 Percentage succeeded: 97%

  • error Raw results Metric 'ctdeconvolute_rmse' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: spatial_simulators Metric: ctdeconvolute_rmse Control method scores: 29 Expected control method scores: 30 Percentage succeeded: 97%

  • error Raw results Metric 'ctdeconvolute_jsd' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: spatial_simulators Metric: ctdeconvolute_jsd Control method scores: 29 Expected control method scores: 30 Percentage succeeded: 97%

  • error Raw results Metric 'crosscor_mantel' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: spatial_simulators Metric: crosscor_mantel Control method scores: 29 Expected control method scores: 30 Percentage succeeded: 97%

  • error Raw results Metric 'crosscor_cosine' number of control methods

    Number of metric scores for control methods should be equal to #datasets × #control_methods Task: spatial_simulators Metric: crosscor_cosine Control method scores: 29 Expected control method scores: 30 Percentage succeeded: 97%

  • error Scaling Metric 'svg_precision' % outside range

    Percentage of scaled scores outside control range should be less than 10% Task: spatial_simulators Metric: svg_precision Inside range: NA Scaled scores: 106 Percentage outside: NA%

  • error Scaling Worst 'svg_precision' score for 'negative_normal'

    Method 'negative_normal' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: negative_normal Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'negative_normal'

    Method 'negative_normal' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: negative_normal Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'negative_shuffle'

    Method 'negative_shuffle' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: negative_shuffle Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'negative_shuffle'

    Method 'negative_shuffle' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: negative_shuffle Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'positive'

    Method 'positive' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: positive Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'positive'

    Method 'positive' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: positive Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'scdesign2'

    Method 'scdesign2' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: scdesign2 Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'scdesign2'

    Method 'scdesign2' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: scdesign2 Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'scdesign3_nb'

    Method 'scdesign3_nb' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: scdesign3_nb Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'scdesign3_nb'

    Method 'scdesign3_nb' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: scdesign3_nb Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'scdesign3_poisson'

    Method 'scdesign3_poisson' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: scdesign3_poisson Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'scdesign3_poisson'

    Method 'scdesign3_poisson' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: scdesign3_poisson Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'sparsim'

    Method 'sparsim' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: sparsim Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'sparsim'

    Method 'sparsim' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: sparsim Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'splatter'

    Method 'splatter' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: splatter Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'splatter'

    Method 'splatter' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: splatter Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'srtsim'

    Method 'srtsim' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: srtsim Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'srtsim'

    Method 'srtsim' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: srtsim Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'symsim'

    Method 'symsim' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: symsim Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'symsim'

    Method 'symsim' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: symsim Metric: svg_precision Best score: NaN Percentage outside range: 0%

  • error Scaling Worst 'svg_precision' score for 'zinbwave'

    Method 'zinbwave' performs much worse than controls for metric' svg_precision' Task: spatial_simulators Method: zinbwave Metric: svg_precision Worst score: NaN Percentage outside range: 0%

  • error Scaling Best 'svg_precision' score for 'zinbwave'

    Method 'zinbwave' performs much better than controls for metric 'svg_precision' Task: spatial_simulators Method: zinbwave Metric: svg_precision Best score: NaN Percentage outside range: 0%

Show 3 warnings
  • warning Raw results Method 'scdesign3_poisson' % missing

    Percentage of missing results should be less than 10% Task: spatial_simulators Method: scdesign3_poisson Number of results: 304 Expected number of results: 390 Percentage missing: 22%

  • warning Raw results Metric 'svg_recall' % missing

    Percentage of missing results should be less than 10% Task: spatial_simulators Metric: svg_recall Number of results: 84 Expected number of results: 110 Percentage missing: 24%

  • warning Raw results Method 'scdesign3_poisson' % failed

    Percentage of failed processes should be less than 10% Task: spatial_simulators Method: scdesign3_poisson Succeeded processes: 8 Attempted processes: 10 Percentage failed: 20%

Method info 8

A transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured

scDesign2 is a transparent simulator that achieves all three goals (preserving genes, capturing gene correlations, and generating any number of cells with varying sequencing depths) and generates high-fidelity synthetic data for multiple single-cell gene expression count-based technologies.

A probabilistic model that unifies the generation and inference for single-cell and spatial omics data

scDesign3 offers a probabilistic model that unifies the generation and inference for single-cell and spatial omics data. The model's interpretable parameters and likelihood enable scDesign3 to generate customized in silico data and unsupervisedly assess the goodness-of-fit of inferred cell latent structures (for example, clusters, trajectories and spatial locations).

A probabilistic model that unifies the generation and inference for single-cell and spatial omics data

scDesign3 offers a probabilistic model that unifies the generation and inference for single-cell and spatial omics data. The model's interpretable parameters and likelihood enable scDesign3 to generate customized in silico data and unsupervisedly assess the goodness-of-fit of inferred cell latent structures (for example, clusters, trajectories and spatial locations).

SPARSim single cell is a count data simulator for scRNA-seq data.

SPARSim is a scRNA-seq count data simulator based on a Gamma-Multivariate Hypergeometric model. It allows to generate count data that resembles real data in terms of count intensity, variability and sparsity.

A single cell RNA-seq data simulator based on a gamma-Poisson distribution.

The Splat model is a gamma-Poisson distribution used to generate a gene by cell matrix of counts. Mean expression levels for each gene are simulated from a gamma distribution and the Biological Coefficient of Variation is used to enforce a mean-variance trend before counts are simulated from a Poisson distribution.

An SRT-specific simulator for scalable, reproducible, and realistic SRT simulations.

A key benefit of srtsim is its ability to maintain location-wise and gene-wise SRT count properties and preserve spatial expression patterns, enabling evaluation of SRT method performance using synthetic data.

Simulating multiple faceted variability in single cell RNA sequencing

SymSim is a simulator for modeling single-cell RNA-Seq data, accounting for three primary sources of variation: intrinsic transcription noise, extrinsic variation from different cell states, and technical variation from measurement noise and bias.

A general and flexible method for signal extraction from single-cell RNA-seq data

ZINB-WaVE is a general and flexible zero-inflated negative binomial model, which leads to low-dimensional representations of the data that account for zero inflation (dropouts), over-dispersion, and the count nature of the data.

Control method info 3
negative_normal

A negative control which generates normal distributed data.

This control method generates normally distributed data as a negative control, using a fixed mean of 3 and standard deviation of 1.

negative_shuffle

A negative control method which shuffles the input data.

This control method shuffles the input data as a negative control.

positive

A positive control method.

Metric info 39
Adjusted Rand indexhigher is betterVinh et al., 2009

Adjusted rand index (ARI) measures the similarity between two clusters in real and simulated datasets.

Adjusted Rand Index used in spatial clustering to measure the similarity between two data clusterings, adjusted for chance.

Cell type deconvolution JSDlower is betterDrost, 2018

Jensen-Shannon divergence (JSD) is calculated between the true and predicted proportion per cell type in all spots.

Jensen-Shannon Divergence used in cell type deconvolution to measure the similarity between two probability distributions.

Cell type deconvolution RMSElower is betterHodson, 2022

Root Mean Square deviation is calculated between the true and predicted proportion of per cell type.

Root Mean Squared Error used in cell type deconvolution to measure the difference between observed and predicted values.

Cosine similarityhigher is betterLeydesdorff, 2005

Cosine similarity measures similarity between bivariate Moran’s I of real dataset and that of in simulation dataset.

Cosine similarity used in spatial cross-correlation to measure the cosine of the angle between two non-zero vectors.

Effective library sizelower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the effective library size.

The kernel density based global two-sample statistic (ks::kde.test) comparing the effective library size of the real datasets versus the effective library size of the simulated datasets.

Effective library sizelower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the effective library size.

The kernel density based global two-sample statistic (ks::kde.test) comparing the effective library size of the real datasets versus the effective library size of the simulated datasets.

Fraction of zeros per celllower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the fraction of zeros per spot (cell).

The kernel density based global two-sample statistic (ks::kde.test) comparing the fraction of zeros per spot (cell) in the real datasets versus the fraction of zeros per spot (cell) in the simulated datasets.

Fraction of zeros per celllower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the fraction of zeros per spot (cell).

The kernel density based global two-sample statistic (ks::kde.test) comparing the fraction of zeros per spot (cell) in the real datasets versus the fraction of zeros per spot (cell) in the simulated datasets.

Fraction of zeros per genelower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the fraction of zeros per gene.

The kernel density based global two-sample statistic (ks::kde.test) comparing the fraction of zeros per gene in the real datasets versus the fraction of zeros per gene in the simulated datasets.

Fraction of zeros per genelower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the fraction of zeros per gene.

The kernel density based global two-sample statistic (ks::kde.test) comparing the fraction of zeros per gene in the real datasets versus the fraction of zeros per gene in the simulated datasets.

Gene Pearson correlationlower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the gene Pearson correlation.

The kernel density based global two-sample statistic (ks::kde.test) comparing the gene Pearson correlation of the real datasets versus the gene Pearson correlation of the simulated datasets.

Gene Pearson correlationlower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the gene Pearson correlation.

The kernel density based global two-sample statistic (ks::kde.test) comparing the gene Pearson correlation of the real datasets versus the gene Pearson correlation of the simulated datasets.

L statisticslower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the L statistics

The kernel density based global two-sample statistic (ks::kde.test) comparing the L statistics in the real datasets versus the L statistics in the simulated datasets.

Library sizelower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the library size.

The kernel density based global two-sample statistic (ks::kde.test) comparing the total sum of UMI counts across all genes in the real datasets versus the total sum of UMI counts across all genes in the simmulated datasets.

Library sizelower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the library size.

The kernel density based global two-sample statistic (ks::kde.test) comparing the total sum of UMI counts across all genes in the real datasets versus the total sum of UMI counts across all genes in the simmulated datasets.

Library size vs fraction zerolower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the relationship between library size and the proportion of zeros per spot (cell).

The kernel density based global two-sample statistic (ks::kde.test) comparing the relationship between library size and the proportion of zeros per spot (cell) in the real datasets versus the simulated datasets.

Library size vs fraction zerolower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the relationship between library size and the proportion of zeros per spot (cell).

The kernel density based global two-sample statistic (ks::kde.test) comparing the relationship between library size and the proportion of zeros per spot (cell) in the real datasets versus the simulated datasets.

Mantel statistichigher is betterLegendre et al., 2015

Mantel statistic is the test statistic for the Mantel test, which is a correlation coefficient calculated between bivariate Moran’s I of real dataset and that of in simulation dataset.

Mantel statistic used in spatial cross-correlation to test the correlation between two distance matrices.

Mean vs fraction zerolower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the relationship between mean expression and the proportion of zero per gene.

The kernel density based global two-sample statistic (ks::kde.test) comparing the relationship between mean expression and the proportion of zero per gene in the real datasets versus the simulated datasets.

Mean vs fraction zerolower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the relationship between mean expression and the proportion of zero per gene.

The kernel density based global two-sample statistic (ks::kde.test) comparing the relationship between mean expression and the proportion of zero per gene in the real datasets versus the simulated datasets.

Mean vs variancelower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the relationship between mean expression and variance expression.

The kernel density based global two-sample statistic (ks::kde.test) comparing the relationship between mean expression and variance expression in the real datasets versus the simulated datasets.

Mean vs variancelower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the relationship between mean expression and variance expression.

The kernel density based global two-sample statistic (ks::kde.test) comparing the relationship between mean expression and variance expression in the real datasets versus the simulated datasets.

Moran's Ilower is betterChacón & Duong, 2018

Kernel density two-sample statistic of Moran's I.

The kernel density based global two-sample statistic (ks::kde.test) comparing the Moran's I of the real datasets versus the Moran's I of the simulated datasets.

Nearest-neighbour correlationlower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the nearest-neighbour correlation.

The kernel density based global two-sample statistic (ks::kde.test) comparing the nn correlation in the real datasets versus the nn correlation in the simulated datasets.

Normalised mutual informationhigher is betterVinh et al., 2009

Normalized mutual information (NMI) measures of the mutual dependence between the real and simulated spatial clusters.

Normalized Mutual Information used in spatial clustering to measure the agreement between two different clusterings, scaled to [0, 1].

Sample Pearson correlationlower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the sample Pearson correlation.

The kernel density based global two-sample statistic (ks::kde.test) comparing the sample Pearson correlation of the real datasets versus the sample Pearson correlation of the simulated datasets.

Sample Pearson correlationlower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the sample Pearson correlation.

The kernel density based global two-sample statistic (ks::kde.test) comparing the sample Pearson correlation of the real datasets versus the sample Pearson correlation of the simulated datasets.

Scaled mean cellslower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the spot- (or cell-) level scaled mean of the expression matrix.

The kernel density based global two-sample statistic (ks::kde.test) comparing the z-score standardization of the mean of expression matrix in terms of log2(CPM) in the real datasets versus the simulated datasets.

Scaled mean cellslower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the spot- (or cell-) level scaled mean of the expression matrix.

The kernel density based global two-sample statistic (ks::kde.test) comparing the z-score standardization of the mean of expression matrix in terms of log2(CPM) in the real datasets versus the simulated datasets.

Scaled mean geneslower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the gene-level scaled mean of the expression matrix.

The kernel density based global two-sample statistic (ks::kde.test) comparing the gene-level z-score standardization of the mean of expression matrix in terms of log2(CPM) in the real datasets versus the simulated datasets.

Scaled mean geneslower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the gene-level scaled mean of the expression matrix.

The kernel density based global two-sample statistic (ks::kde.test) comparing the gene-level z-score standardization of the mean of expression matrix in terms of log2(CPM) in the real datasets versus the simulated datasets.

Scaled variance celllower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the spot- (or cell-) level scaled variance of the expression matrix.

The kernel density based global two-sample statistic (ks::kde.test) comparing the spot-level z-score standardization of the variance of expression matrix in terms of log2(CPM) in the real datasets versus the simulated datasets.

Scaled variance celllower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the spot- (or cell-) level scaled variance of the expression matrix.

The kernel density based global two-sample statistic (ks::kde.test) comparing the spot-level z-score standardization of the variance of expression matrix in terms of log2(CPM) in the real datasets versus the simulated datasets.

Scaled variance geneslower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the gene-level scaled variance of the expression matrix.

The kernel density based global two-sample statistic (ks::kde.test) comparing the gene-level z-score standardization of the variance of expression matrix in terms of log2(CPM) in the real datasets versus the simulated datasets.

Scaled variance geneslower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the gene-level scaled variance of the expression matrix.

The kernel density based global two-sample statistic (ks::kde.test) comparing the gene-level z-score standardization of the variance of expression matrix in terms of log2(CPM) in the real datasets versus the simulated datasets.

Precision measures the proportion of correctly identified items in simulated datasets.

Precision used in identifying spatial variable genes, measuring the accuracy of positive predictions.

Recall measures the proportion of real SVG correctly identified in the simulated dataset.

Recall used in identifying spatial variable genes, measuring the true positive rate.

TMMlower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the weight trimmed mean of M-values normalization factor (TMM).

The kernel density based global two-sample statistic (ks::kde.test) comparing the weight trimmed mean of M-values normalization factor for the real datasets versus the weight trimmed mean of M-values normalization factor for the simulated datasets.

TMMlower is betterChacón & Duong, 2018

Kernel density two-sample statistic of the weight trimmed mean of M-values normalization factor (TMM).

The kernel density based global two-sample statistic (ks::kde.test) comparing the weight trimmed mean of M-values normalization factor for the real datasets versus the weight trimmed mean of M-values normalization factor for the simulated datasets.

Dataset info 10
Brain unlinked

10X Visium spatial RNA-seq from adult mouse brain sections paired to single-nucleus RNA-seq

This datasets were generated matched single nucleus (sn, this submission) and Visium spatial RNA-seq (10X Genomics) profiles of adjacent mouse brain sections that contain multiple regions from the telencephalon and diencephalon.

Breast unlinked

A spatially resolved atlas of human breast cancers

This study presents a spatially resolved transcriptomics analysis of human breast cancers.

Cortex unlinked

Scripts and source data for image processing, barcode calling, and cell type annotations in a seqFISH+ experiment.

The dataset includes processed image data, cell type annotations with Louvain clusters, gene IDs for transcript locations, and mRNA point locations, with additional data available on Zenodo.

Fibrosarcoma unlinked

Multi-resolution deconvolution of spatial transcriptomics data reveals continuous patterns of Tumor A1 of Tissue 1

Spatial transcriptomics of Tumor A1 of Tissue 1.

Gastrulation unlinked

single-cell and spatial transcriptomic molecular map of mouse gastrulation

Single-Cell omics Data across Mouse Gastrulation and Highly multiplexed spatially resolved gene expression profiling of Early Organogenesis.

Hindlimbmuscle unlinked

Spatial RNA sequencing of regenerating mouse hindlimb muscle

The spatial transcriptomics datasets regenerates mouse muscle tissue generated with the 10x Genomics Visium platform.

Olfactorybulb unlinked

Single-cell and spatial transcriptomic of mouse olfactory bulb

Osteosarcoma unlinked

Spatial profiling of human osteosarcoma cells.

Spatial transcriptome profiling by MERFISH reveals subcellular RNA compartmentalization and cell cycle-dependent gene expression.

pancreatic ductal adenocarcinomas unlinked

Integrating microarray-based spatial transcriptomics and single-cell RNA-seq reveals tissue architecture in pancreatic ductal adenocarcinomas

We developed a multimodal intersection analysis method combining scRNA-seq with spatial transcriptomics to map and characterize the spatial organization and interactions of distinct cell subpopulations in complex tissues, such as primary pancreatic tumors..

Prostate unlinked

Spatially resolved gene expression of human protate tissue slices treated with steroid hormones for 8 hours

Spatially resolved gene expression was prepard by dissociated hman prostate tissue to single cells, and collected & prepped for RNA-seq using the Visium Spatial Gene Expression kit.

References

  1. Journal of Machine Learning Technologies. (n.d.). 10.9735/2229-3981 ↗
  2. Baruzzo, G., Patuzzi, I., & Di Camillo, B. (2019). SPARSim single cell: a count data simulator for scRNA-seq data. 10.1093/bioinformatics/btz752 ↗
  3. Chacón, J. E., & Duong, T. (2018). Multivariate Kernel Smoothing and its Applications. 10.1201/9780429485572 ↗
  4. Drost, H.-G. (2018). Philentropy: Information Theory and Distance Quantification with R. 10.21105/joss.00765 ↗
  5. Eng, C.-H. L., Lawson, M., Zhu, Q., Dries, R., Koulena, N., Takei, Y., Yun, J., Cronin, C., Karp, C., Yuan, G.-C., & Cai, L. (2019). Transcriptome-scale super-resolved imaging in tissues by RNA seqFISH+. 10.1038/s41586-019-1049-y ↗
  6. Hodson, T. O. (2022). Root-mean-square error (RMSE) or mean absolute error (MAE): when to use them or not. 10.5194/gmd-15-5481-2022 ↗
  7. Kleshchevnikov, V., Shmatko, A., Dann, E., Aivazidis, A., King, H. W., Li, T., Elmentaite, R., Lomakin, A., Kedlian, V., Gayoso, A., Jain, M. S., Park, J. S., Ramona, L., Tuck, E., Arutyunyan, A., Vento-Tormo, R., Gerstung, M., James, L., Stegle, O., & Bayraktar, O. A. (2022). Cell2location maps fine-grained cell types in spatial transcriptomics. Nature Biotechnology, 40(5), 661–671. 10.1038/s41587-021-01139-4 ↗
  8. Legendre, P., Fortin, M., & Borcard, D. (2015). Should the Mantel test be used in spatial analysis? 10.1111/2041-210x.12425 ↗
  9. Leydesdorff, L. (2005). Similarity measures, author cocitation analysis, and information theory. 10.1002/asi.20130 ↗
  10. Liang, X., Cao, Y., & Hwa Yang, J. Y. (2024). Multi-task benchmarking of spatially resolved gene expression simulation models. 10.1101/2024.05.29.596418 ↗
  11. Lohoff, T., Ghazanfar, S., Missarova, A., Koulena, N., Pierson, N., Griffiths, J. A., Bardot, E. S., Eng, C.-H. L., Tyser, R. C. V., Argelaguet, R., Guibentif, C., Srinivas, S., Briscoe, J., Simons, B. D., Hadjantonakis, A.-K., Göttgens, B., Reik, W., Nichols, J., Cai, L., & Marioni, J. C. (2021). Integration of spatial and single-cell transcriptomic data elucidates mouse organogenesis. 10.1038/s41587-021-01006-2 ↗
  12. Lopez, R., Li, B., Keren-Shaul, H., Boyeau, P., Kedmi, M., Pilzer, D., Jelinski, A., Yofe, I., David, E., Wagner, A., Ergen, C., Addadi, Y., Golani, O., Ronchese, F., Jordan, M. I., Amit, I., & Yosef, N. (2022). DestVI identifies continuums of cell types in spatial transcriptomics data. Nature Biotechnology, 40(9), 1360–1369. 10.1038/s41587-022-01272-8 ↗
  13. McCray, T., Pacheco, J. V., Loitz, C. C., Garcia, J., Baumann, B., Schlicht, M. J., Valyi-Nagy, K., Abern, M. R., & Nonn, L. (2021). Vitamin D sufficiency enhances differentiation of patient-derived prostate epithelial organoids. 10.1016/j.isci.2021.102640 ↗
  14. McKellar, D. W., Walter, L. D., Song, L. T., Mantri, M., Wang, M. F. Z., De Vlaminck, I., & Cosgrove, B. D. (2020). Strength in numbers: Large-scale integration of single-cell transcriptomic data reveals rare, transient muscle progenitor cell states in muscle regeneration. 10.1101/2020.12.01.407460 ↗
  15. Risso, D., Perraudeau, F., Gribkova, S., Dudoit, S., & Vert, J.-P. (2018). A general and flexible method for signal extraction from single-cell RNA-seq data. 10.1038/s41467-017-02554-5 ↗
  16. Song, D., Wang, Q., Yan, G., Liu, T., Sun, T., & Li, J. J. (2023). scDesign3 generates realistic in silico data for multimodal single-cell and spatial omics. 10.1038/s41587-023-01772-1 ↗
  17. Ståhl, P. L., Salmén, F., Vickovic, S., Lundmark, A., Navarro, J. F., Magnusson, J., Giacomello, S., Asp, M., Westholm, J. O., Huss, M., Mollbrink, A., Linnarsson, S., Codeluppi, S., Borg, Å., Pontén, F., Costea, P. I., Sahlén, P., Mulder, J., Bergmann, O., … Frisén, J. (2016). Visualization and analysis of gene expression in tissue sections by spatial transcriptomics. 10.1126/science.aaf2403 ↗
  18. Sun, T., Song, D., Li, W. V., & Li, J. J. (2021). scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured. 10.1186/s13059-021-02367-2 ↗
  19. Vinh, N. X., Epps, J., & Bailey, J. (2009). Information theoretic measures for clusterings comparison. 10.1145/1553374.1553511 ↗
  20. Wu, S. Z., Al-Eryani, G., Roden, D. L., Junankar, S., Harvey, K., Andersson, A., Thennavan, A., Wang, C., Torpy, J. R., Bartonicek, N., Wang, T., Larsson, L., Kaczorowski, D., Weisenfeld, N. I., Uytingco, C. R., Chew, J. G., Bent, Z. W., Chan, C.-L., Gnanasambandapillai, V., … Swarbrick, A. (2021). A single-cell and spatially resolved atlas of human breast cancers. Nature Genetics, 53(9), 1334–1347. 10.1038/s41588-021-00911-1 ↗
  21. Xia, C., Fan, J., Emanuel, G., Hao, J., & Zhuang, X. (2019). Spatial transcriptome profiling by MERFISH reveals subcellular RNA compartmentalization and cell cycle-dependent gene expression. 10.1073/pnas.1912459116 ↗
  22. Zappia, L., Phipson, B., & Oshlack, A. (2017). Splatter: simulation of single-cell RNA sequencing data. 10.1186/s13059-017-1305-0 ↗
  23. Zhang, X., Xu, C., & Yosef, N. (2019). Simulating multiple faceted variability in single cell RNA sequencing. 10.1038/s41467-019-10500-w ↗
  24. Zhu, J., Shang, L., & Zhou, X. (2023). SRTsim: spatial pattern preserving simulations for spatially resolved transcriptomics. 10.1186/s13059-023-02879-z ↗