Predict Modality
Predicting the profiles of one modality (e.g. protein abundance) from another (e.g. mRNA expression).
Experimental techniques to measure multiple modalities within the same single cell are increasingly becoming available. The demand for these measurements is driven by the promise to provide a deeper insight into the state of a cell. Yet, the modalities are also intrinsically linked. We know that DNA must be accessible (ATAC data) to produce mRNA (expression data), and mRNA in turn is used as a template to produce protein (protein abundance). These processes are regulated often by the same molecules that they produce: for example, a protein may bind DNA to prevent the production of more mRNA. Understanding these regulatory processes would be transformative for synthetic biology and drug target discovery. Any method that can predict a modality from another must have accounted for these regulatory processes, but the demand for multi-modal data shows that this is not trivial.
Contributors
Leaderboard
Methods ranked by scaled overall mean. Each cell encodes a score from 0 to 1 by size and intensity.
QC: Normalisation Visualisation 8 plots
Per metric: points placed by control-anchored scaled score (x); dashed lines mark scaled 0 and 1 (worst/best control); the lower axis shows the raw score. Points beyond [-0.2, 1.2] are clamped to the edge as triangles. Hover a dot or line to highlight it and read details.
- MAElower better
- Mean pearson per cellhigher better
- Mean pearson per genehigher better
- Mean spearman per cellhigher better
- Mean spearman per genehigher better
- Overall pearsonhigher better
- Overall spearmanhigher better
- RMSElower better
QC: Indicator table 15 errors30 warnings
Automated checks on the benchmark run and its results: missing values, score scaling, metric ranges and similar. Errors are high-severity issues that usually need a maintainer's attention; warnings are lower-severity signals. Findings that are expected for this task are listed separately as silenced.
15 high-severity issues need review. 254 of 299 checks passed.
- error Method info Info field 'references' % missing
Method info field 'references' should be defined Task: predict_modality Field: references Percentage missing: 8
- error Raw results Dataset 'openproblems_neurips2022/pbmc_multiome/swap' % missing
Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_multiome/swap Number of results: 58 Expected number of results: 104 Percentage missing: 44%
- error Raw results Dataset 'openproblems_neurips2022/pbmc_cite/swap' % missing
Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_cite/swap Number of results: 66 Expected number of results: 104 Percentage missing: 37%
- error Raw results Method 'novel' % missing
Percentage of missing results should be less than 10% Task: predict_modality Method: novel Number of results: 6 Expected number of results: 64 Percentage missing: 91%
- error Raw results Method 'simple_mlp' % missing
Percentage of missing results should be less than 10% Task: predict_modality Method: simple_mlp Number of results: 0 Expected number of results: 64 Percentage missing: 100%
- error Raw results Metric 'overall_pearson' % missing
Percentage of missing results should be less than 10% Task: predict_modality Metric: overall_pearson Number of results: 71 Expected number of results: 104 Percentage missing: 32%
- error Raw results Metric 'overall_spearman' % missing
Percentage of missing results should be less than 10% Task: predict_modality Metric: overall_spearman Number of results: 71 Expected number of results: 104 Percentage missing: 32%
- error Raw results Dataset 'openproblems_neurips2022/pbmc_multiome/swap' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_multiome/swap Succeeded processes: 8 Attempted processes: 13 Percentage failed: 38%
- error Raw results Dataset 'openproblems_neurips2022/pbmc_cite/swap' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_cite/swap Succeeded processes: 9 Attempted processes: 13 Percentage failed: 31%
- error Raw results Method 'novel' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Method: novel Succeeded processes: 1 Attempted processes: 8 Percentage failed: 88%
- error Raw results Method 'simple_mlp' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Method: simple_mlp Succeeded processes: 0 Attempted processes: 8 Percentage failed: 100%
- error Raw results Metric 'overall_pearson' number of control methods
Number of metric scores for control methods should be equal to #datasets × #control_methods Task: predict_modality Metric: overall_pearson Control method scores: 24 Expected control method scores: 32 Percentage succeeded: 75%
- error Raw results Metric 'overall_spearman' number of control methods
Number of metric scores for control methods should be equal to #datasets × #control_methods Task: predict_modality Metric: overall_spearman Control method scores: 24 Expected control method scores: 32 Percentage succeeded: 75%
- error Raw results Metric 'rmse' number of control methods
Number of metric scores for control methods should be equal to #datasets × #control_methods Task: predict_modality Metric: rmse Control method scores: 31 Expected control method scores: 32 Percentage succeeded: 97%
- error Raw results Metric 'mae' number of control methods
Number of metric scores for control methods should be equal to #datasets × #control_methods Task: predict_modality Metric: mae Control method scores: 31 Expected control method scores: 32 Percentage succeeded: 97%
Show 30 warnings
- warning Raw results Task number of results
Number of results should be equal to #datasets × #methods × #metrics Task: predict_modality Number of results: 624 Number of datasets: 8 Number of methods: 13 Number of metrics: 8 Expected number of results: 832
- warning Raw results Dataset 'openproblems_neurips2021/bmmc_multiome/swap' % missing
Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_multiome/swap Number of results: 82 Expected number of results: 104 Percentage missing: 21%
- warning Raw results Dataset 'openproblems_neurips2022/pbmc_multiome/normal' % missing
Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_multiome/normal Number of results: 78 Expected number of results: 104 Percentage missing: 25%
- warning Raw results Dataset 'openproblems_neurips2021/bmmc_multiome/normal' % missing
Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_multiome/normal Number of results: 78 Expected number of results: 104 Percentage missing: 25%
- warning Raw results Method 'zeros' % missing
Percentage of missing results should be less than 10% Task: predict_modality Method: zeros Number of results: 48 Expected number of results: 64 Percentage missing: 25%
- warning Raw results Method 'lm' % missing
Percentage of missing results should be less than 10% Task: predict_modality Method: lm Number of results: 48 Expected number of results: 64 Percentage missing: 25%
- warning Raw results Method 'guanlab_dengkw_pm' % missing
Percentage of missing results should be less than 10% Task: predict_modality Method: guanlab_dengkw_pm Number of results: 48 Expected number of results: 64 Percentage missing: 25%
- warning Raw results Metric 'mean_pearson_per_cell' % missing
Percentage of missing results should be less than 10% Task: predict_modality Metric: mean_pearson_per_cell Number of results: 82 Expected number of results: 104 Percentage missing: 21%
- warning Raw results Metric 'mean_spearman_per_cell' % missing
Percentage of missing results should be less than 10% Task: predict_modality Metric: mean_spearman_per_cell Number of results: 82 Expected number of results: 104 Percentage missing: 21%
- warning Raw results Metric 'mean_pearson_per_gene' % missing
Percentage of missing results should be less than 10% Task: predict_modality Metric: mean_pearson_per_gene Number of results: 82 Expected number of results: 104 Percentage missing: 21%
- warning Raw results Metric 'mean_spearman_per_gene' % missing
Percentage of missing results should be less than 10% Task: predict_modality Metric: mean_spearman_per_gene Number of results: 82 Expected number of results: 104 Percentage missing: 21%
- warning Raw results Metric 'rmse' % missing
Percentage of missing results should be less than 10% Task: predict_modality Metric: rmse Number of results: 77 Expected number of results: 104 Percentage missing: 26%
- warning Raw results Metric 'mae' % missing
Percentage of missing results should be less than 10% Task: predict_modality Metric: mae Number of results: 77 Expected number of results: 104 Percentage missing: 26%
- warning Raw results Dataset 'openproblems_neurips2022/pbmc_multiome/normal' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_multiome/normal Succeeded processes: 10 Attempted processes: 13 Percentage failed: 23%
- warning Raw results Dataset 'openproblems_neurips2021/bmmc_multiome/normal' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_multiome/normal Succeeded processes: 10 Attempted processes: 13 Percentage failed: 23%
- warning Raw results Method 'lm' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Method: lm Succeeded processes: 6 Attempted processes: 8 Percentage failed: 25%
- warning Raw results Method 'guanlab_dengkw_pm' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Method: guanlab_dengkw_pm Succeeded processes: 6 Attempted processes: 8 Percentage failed: 25%
- warning Raw results Dataset 'openproblems_neurips2021/bmmc_cite/normal' % missing
Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_cite/normal Number of results: 92 Expected number of results: 104 Percentage missing: 12%
- warning Raw results Dataset 'openproblems_neurips2022/pbmc_cite/normal' % missing
Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_cite/normal Number of results: 84 Expected number of results: 104 Percentage missing: 19%
- warning Raw results Dataset 'openproblems_neurips2021/bmmc_cite/swap' % missing
Percentage of missing results should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_cite/swap Number of results: 86 Expected number of results: 104 Percentage missing: 17%
- warning Raw results Method 'knnr_r' % missing
Percentage of missing results should be less than 10% Task: predict_modality Method: knnr_r Number of results: 56 Expected number of results: 64 Percentage missing: 12%
- warning Raw results Method 'cellmapper_linear' % missing
Percentage of missing results should be less than 10% Task: predict_modality Method: cellmapper_linear Number of results: 56 Expected number of results: 64 Percentage missing: 12%
- warning Raw results Method 'cellmapper_scvi' % missing
Percentage of missing results should be less than 10% Task: predict_modality Method: cellmapper_scvi Number of results: 52 Expected number of results: 64 Percentage missing: 19%
- warning Raw results Method 'suzuki_mlp' % missing
Percentage of missing results should be less than 10% Task: predict_modality Method: suzuki_mlp Number of results: 56 Expected number of results: 64 Percentage missing: 12%
- warning Raw results Dataset 'openproblems_neurips2021/bmmc_multiome/swap' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_multiome/swap Succeeded processes: 11 Attempted processes: 13 Percentage failed: 15%
- warning Raw results Dataset 'openproblems_neurips2022/pbmc_cite/normal' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2022/pbmc_cite/normal Succeeded processes: 11 Attempted processes: 13 Percentage failed: 15%
- warning Raw results Dataset 'openproblems_neurips2021/bmmc_cite/swap' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Dataset: openproblems_neurips2021/bmmc_cite/swap Succeeded processes: 11 Attempted processes: 13 Percentage failed: 15%
- warning Raw results Method 'knnr_r' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Method: knnr_r Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%
- warning Raw results Method 'cellmapper_linear' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Method: cellmapper_linear Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%
- warning Raw results Method 'suzuki_mlp' % failed
Percentage of failed processes should be less than 10% Task: predict_modality Method: suzuki_mlp Succeeded processes: 7 Attempted processes: 8 Percentage failed: 12%
Method info 9
Modality prediction in a PCA/CCA space using CellMapper
CellMapper is a general framework for k-NN based mapping tasks in single-cell and spatial genomics. This variant uses CellMapper to project modalities from a reference dataset (train) onto a query dataset (test) in a PCA/CCA latent space.
Modality prediction in an scVI latent space using CellMapper
CellMapper is a general framework for k-NN based mapping tasks in single-cell and spatial genomics. This variant uses CellMapper to project modalities from a reference dataset (train) onto a query dataset (test) in a modality-specific latent space computed with suitable scvi-tools models. For gene expression data, we use the scVI model on raw counts (nb likelihood), for ADT data, we use the scVI models on normalized counts (gaussian likelihood), and for ATAC data, we use the PeakVI model on raw counts. The actual CellMapper pipeline is modality-agnostic.
A kernel ridge regression method with RBF kernel.
This is a solution developed by Team Guanlab - dengkw in the Neurips 2021 competition to predict one modality from another using kernel ridge regression (KRR) with RBF kernel. Truncated SVD is applied on the combined training and test data from modality 1 followed by row-wise z-score normalization on the reduced matrix. The truncated SVD of modality 2 is predicted by training a KRR model on the normalized training matrix of modality 1. Predictions on the normalized test matrix are then re-mapped to the modality 2 feature space via the right singular vectors.
K-nearest neighbor regression in Python.
K-nearest neighbor regression in R.
Linear model regression.
A linear model regression method.
A method using encoder-decoder MLP model
This method trains an encoder-decoder MLP model with one output neuron per component in the target. As an input, the encoders use representations obtained from ATAC and GEX data via LSI transform and raw ADT data. The hyperparameters of the models were found via broad hyperparameter search using the Optuna framework.
Ensemble of MLPs trained on different sites (team AXX)
This folder contains the AXX solution to the OpenProblems-NeurIPS2021 Single-Cell Multimodal Data Integration. Team took the 4th place of the modality prediction task in terms of overall ranking of 4 subtasks: namely GEX to ADT, ADT to GEX, GEX to ATAC and ATAC to GEX. Specifically, our methods ranked 3rd in GEX to ATAC and 4th in GEX to ADT. More details about the task can be found in the competition webpage.
Hierarchical encoder-decoder neural network with task-specific preprocessing and residual connections for cross-modal prediction.
A hierarchical neural network encoder-decoder model based on Shuji Suzuki's 1st place solution in the Open Problems Multimodal Single-Cell Integration competition. The model uses task-specific preprocessing, SVD dimensionality reduction, and hierarchical MLP blocks with residual connections for learning cross-modal mappings.
The original author's code was adapted by GitHub Copilot (using Claude Sonnet) to integrate with this repository's framework and standards.
Control method info 4
Returns the mean expression value per gene.
Returns random training profiles.
Returns the ground-truth solution.
Returns a prediction consisting of all zeros.
Metric info 8
The mean absolute error.
The average difference between the expression values and the predicted expression values.
The mean of the pearson values of per-cell expression value vectors.
The mean of the pearson values of per-gene expression value vectors.
The mean of the spearman values of per-cell expression value vectors.
The mean of the spearman values of per-gene expression value vectors.
The mean of the pearson values of vectorized expression matrices.
The mean of the spearman values of vectorized expression matrices.
The root mean squared error.
The square root of the mean of the square of all of the error.
Dataset info 8
Single-cell CITE-Seq (GEX+ADT) data collected from bone marrow mononuclear cells of 12 healthy human donors.
Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X 3 prime Single-Cell Gene Expression kit with Feature Barcoding in combination with the BioLegend TotalSeq B Universal Human Panel v1.0. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2021. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.
Single-cell CITE-Seq (GEX+ADT) data collected from bone marrow mononuclear cells of 12 healthy human donors.
Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X 3 prime Single-Cell Gene Expression kit with Feature Barcoding in combination with the BioLegend TotalSeq B Universal Human Panel v1.0. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2021. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.
Single-cell Multiome (GEX+ATAC) data collected from bone marrow mononuclear cells of 12 healthy human donors.
Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X Multiome Gene Expression and Chromatin Accessibility kit. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2021. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.
Single-cell Multiome (GEX+ATAC) data collected from bone marrow mononuclear cells of 12 healthy human donors.
Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X Multiome Gene Expression and Chromatin Accessibility kit. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2021. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.
Single-cell CITE-Seq (GEX+ADT) data collected from bone marrow mononuclear cells of 12 healthy human donors.
Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X 3 prime Single-Cell Gene Expression kit with Feature Barcoding in combination with the BioLegend TotalSeq B Universal Human Panel v1.0. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2022. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.
Single-cell CITE-Seq (GEX+ADT) data collected from bone marrow mononuclear cells of 12 healthy human donors.
Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X 3 prime Single-Cell Gene Expression kit with Feature Barcoding in combination with the BioLegend TotalSeq B Universal Human Panel v1.0. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2022. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.
Single-cell Multiome (GEX+ATAC) data collected from bone marrow mononuclear cells of 12 healthy human donors.
Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X Multiome Gene Expression and Chromatin Accessibility kit. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2022. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.
Single-cell Multiome (GEX+ATAC) data collected from bone marrow mononuclear cells of 12 healthy human donors.
Single-cell CITE-Seq data collected from bone marrow mononuclear cells of 12 healthy human donors using the 10X Multiome Gene Expression and Chromatin Accessibility kit. The dataset was generated to support Multimodal Single-Cell Data Integration Challenge at NeurIPS 2022. Samples were prepared using a standard protocol at four sites. The resulting data was then annotated to identify cell types and remove doublets. The dataset was designed with a nested batch layout such that some donor samples were measured at multiple sites with some donors measured at a single site.
References
- ... (2024). Predicting cellular profiles across modalities in longitudinal single-cell data: An Open Problems competition. In Preparation.
- Chai, T., & Draxler, R. R. (2014). Root mean square error (RMSE) or mean absolute error (MAE)? 10.5194/gmdd-7-1525-2014 ↗
- Fix, E., & Hodges, J. L. (1989). Discriminatory Analysis. Nonparametric Discrimination: Consistency Properties. International Statistical Review / Revue Internationale de Statistique, 57(3), 238. 10.2307/1403797 ↗
- KENDALL, M. G. (1938). A new measure of rank correlation. Biometrika, 30(1–2), 81–93. 10.1093/biomet/30.1-2.81 ↗
- Lance, C., Luecken, M. D., Burkhardt, D. B., Cannoodt, R., Rautenstrauch, P., Laddach, A., Ubingazhibov, A., Cao, Z.-J., Deng, K., Khan, S., Liu, Q., Russkikh, N., Ryazantsev, G., Ohler, U., Pisco, A. O., Bloom, J., Krishnaswamy, S., & Theis, F. J. (2022). Multimodal single cell data integration challenge: results and lessons learned. bioRxiv. 10.1101/2022.04.11.487796 ↗
- Lange, M. (2025). quadbio/cellmapper: v0.2.2. 10.5281/ZENODO.15683594 ↗
- Luecken, M., Burkhardt, D., Cannoodt, R., Lance, C., Agrawal, A., Aliee, H., Chen, A., Deconinck, L., Detweiler, A., Granados, A., Huynh, S., Isacco, L., Kim, Y., Klein, D., DE KUMAR, B., Kuppasani, S., Lickert, H., McGeever, A., Melgarejo, J., … Bloom, J. M. (2021). A sandbox for prediction and integration of DNA, RNA, and proteins in single cells. In J. Vanschoren & S. Yeung (Eds.), Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (Vol. 1). Curran. link ↗
- Luecken, M. D., Burkhardt, D. B., Cannoodt, R., Lance, C., Agrawal, A., Aliee, H., Chen, A. T., Deconinck, L., Detweiler, A. M., Granados, A. A., & others. (2021). A sandbox for prediction and integration of DNA, RNA, and proteins in single cells. Thirty-Fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
- Pearson, K. (1895). VII. Note on regression and inheritance in the case of two parents. Proceedings of the Royal Society of London, 58(347–352), 240–242. 10.1098/rspl.1895.0041 ↗
- Wilkinson, G. N., & Rogers, C. E. (1973). Symbolic Description of Factorial Models for Analysis of Variance. Applied Statistics, 22(3), 392. 10.2307/2346786 ↗