
Name: Wu Mengyun
Title: Professor
Research Areas: High-dimensional data, variable selection, network models, biostatistics
Courses Taught: Real Analysis, Probability Theory, Stochastic Processes
E-mail: wu.mengyun@mail.shufe.edu.cn; Phone: 65901432
Research Project
Serial Number | Project Name | Project Number | Project Source | Start and End Time | Project Funding |
1 | Detection of Reproducible Network Biomarkers Based on Multi-Types of Data Fusion | 61402276 | National Natural Science Foundation Youth Project | 2015.1-2017.12 | 260,000 |
2 | Multi-level Variable Selection Methods Based on Network Structure and Their Applications | 2018LD02 | National Major Project for Statistical Science Research | 2018.10-2020.09 | 150,000 |
3 | Integration Analysis and Application of Interactions Among Multi-Source High-Dimensional Data Variables | 18CG42 | Shanghai Morning Light Program | 2019.01-2021.12 | 20,000 |
4 | Variable Selection and Integration Analysis of Multi-source High-dimensional Data and Its Applications in the Biomedical Field | 19PJ1403600 | Shanghai Pujiang Talent | 2019.10-2021.09 | 300,000 |
5 | Network Analysis Methods and Theories for Multi-level Heterogeneous Cancer Data | 12071273 | National Natural Science Foundation General Program | 2021.01-2024.12 | 510,000 |
6 | Statistical Modeling and Inference of Spatial Heterogeneity for Single-Cell Transcriptome Data | 22QA1403500 | Shanghai Qimingxing Plan | 2022.06-2025.05 | 400,000 |
Research Field
High-dimensional data, variable selection, network model, biostatistics
Education Background
2008.9-2013.6
Sun Yat-sen University, Probability Theory and Mathematical Statistics, PhD
2004.9-2008.6
Sun Yat-sen University, Statistics, Bachelor's Degree
Work Experience
2023.7-Present Shanghai University of Finance and Economics, School of Statistics and Management, Professor
2018.7-2023.6 Shanghai University of Finance and Economics, School of Statistics and Management, Associate Professor
2016.8-2018.7 Yale University, Biostatistics Department, Visiting Scholar
2013.7-2018.6 Shanghai University of Finance and Economics, School of Statistics and Management, Lecturer
2011.3-2011.8 City University of Hong Kong, Department of Electronic Engineering, Research Assistant
1.Zhu X, Li Y, Ma S, Wu M* (2026). Sparse Bayesian deep functional learning with structured region selection. In the Forty-Third International Conference on Machine Learning.
2.Yang J, Li T, Wang T, Ma S, Wu M* (2026). Heterogeneous network estimation for single-cell transcriptomic data via a joint regularized deep neural network. Journal of the American Statistical Association, published online.
3.Feng X, Li Q, Qin X*, Wu M, Yu L (2026). Structured nonlinear cure model with deep neural networks for high-dimensional survival analysis, Statistics in Medicine, 45: e70368.
4.Wu M, Li Y, Ma S, Wu M*(2025). Joint identification of spatially variable genes via a network-assisted Bayesian regularization approach, The Annals of Applied Statistics, 19,2705-2723.
5.Fan J, Yang J, Ma S, Wu M*(2025). Bilevel network learning via hierarchically structured sparsity, Advances in Neural Information Processing Systems, accepted.
6.Wu M, Li Y, Ma S* (2025). High-dimensional gene-environment interaction analysis. Annual Review of Statistics and Its Application, 12, 361-383.
7.Ge Y, Li T, Feng X, Wu M*, Liu H* (2024). Structured feature ranking for genomic marker identification accommodating multiple types of networks. Biometrics,80, ujae159.
8.Qin X, Hu J, Ma S, Wu M* (2024). Estimation of multiple networks with common structures in heterogeneous subgroups. Journal of Multivariate Analysis, 202, 105298.
9.Li Y, Wu M, Wu M*, Ma S (2024). Identification of influencing factors on self-reported count data with multiple potential inflated values. The Annals of Applied Statistics, 18, 991-1009.
10.Li Y, Wu M, Ma S, Wu M* (2023). ZINBMM: a general mixture model for simultaneous clustering and gene selection using single-cell transcriptomic data. Genome Biology, 24, 208.
11.Wu M, Wang F, Ge Y, Ma S, Li Y* (2023). Bi-level structured functional analysis for genome-wide association studies. Biometrics, 79, 3359-3373.
12.Qin X, Ma S, Wu M* (2023). Two-level Bayesian interaction analysis for survival data incorporating pathway information. Biometrics, 79, 1761-1774.
13.Zhong T, Zhang Q, Huang J, Wu M*, Ma S* (2023). Heterogeneity analysis via integrating multi-sources high-dimensional data with applications to cancer studies. Statistica Sinica, 33, 729-758.
14.Cheng C, Feng X, Li X, Wu M* (2022). Robust analysis of cancer heterogeneity for high-dimensional data. Statistics in Medicine, 41, 5448-5462.
15.Xu Y, Wu M*, Ma S* (2022). Multidimensional molecular measurements-environment interaction analysis for disease outcomes. Biometrics, 78, 1524-1554.
16.Li Y, Xu S, Ma S, Wu M*(2022). Network-based cancer heterogeneity analysis incorporating multi-view of prior information. Bioinformatics, 38, 2855-2862.
17.Li Y, Wang F, Wu M*, Ma S (2022). Integrative functional linear model for genome-wide association studies with multiple traits. Biostatistics, 23, 574-590.
18.Qin X, Ma S, Wu M* (2021). Gene-gene interaction analysis incorporating network information via a structured Bayesian approach. Statistics in Medicine, 40, 6619-6633.
19.Wu M, Qin X, Ma S* (2021). GEInter: an R package for robust gene-environment interaction analysis. Bioinformatics, 39, 3691-3692.
20.Wu M, Yi H, Ma S*(2021). Vertical integration methods for gene expression data analysis. Briefings in Bioinformatics, 22:1-14.
21.Teran Hidalgo SJ, Wu M*, Ma S*(2020). NCutYX: a package for clustering analysis of multilayer omics data. Bioinformatics, 36:1976-1977.
22.Wu M, Zhang Q, Ma S* (2020). Structured gene-environment interaction analysis. Biometrics, 76:23-35.
23.Wu M, Ma S* (2019). Robust semiparametric gene-environment interaction analysis using sparse boosting. Statistics in Medicine, 38:4625-4641.
24.Xu Y#, Wu M#, Zhang Q, Ma S* (2019). Robust identification of gene-environment interactions for prognosis using a quantile partial correlation approach. Genomics, 111:1115-1123.
25.Wang S, Wu M*, Ma S* (2019). Integrative analysis of cancer omics data for prognosis modeling. Genes, 10(8), 604.
26.Li Y, Li R, Qin Y, Wu M*, Ma S* (2019). Integrative interaction analysis using threshold gradient directed regularization. Applied Stochastic Models in Business and Industry, 35(2), 354-375.
27.Wu M, Ma S* (2019). Robust genetic interaction analysis. Briefings in Bioinformatics, 20(2): 624-637
28.Xu Y, Zhong T, Wu M*, Ma S* (2019). Histopathological imaging–environment interactions in cancer modeling. Cancers, 2019, 11(4): 579.
29.Zhong T, Wu M*, Ma S* (2019). Examination of independent prognostic power of gene expressions and histopathological imaging features in cancer. Cancers, 11(3), 361.
30.Teran Hidalgo SJ, Zhu T, Wu M*, Ma S* (2018).Overlapping clustering of gene expression data using penalized weighted normalized cut.Genetic Epidemiology, 42: 796-811.
31.Li Y, Bie R, Hidalgo SJH, Qin Y, Wu M*, Ma S*(2018). Assisted gene expression-based clustering with AWNCut. Statistics in Medicine, 37: 4386-4403.
32.Xu Y, Wu M*, Ma S, Ejaz Ahmed S (2018). Robust gene-environment interaction analysis using penalized trimmed regression. Journal of Statistical Computation and Simulation, 88: 3502-3528.
33.Li T*, Wu M, Zhou Y (2018). A unified semi-empirical likelihood ratio confidence interval for treatment effects in the two sample problem with length-biased data. Statistics and Its Interface. 11: 531-540.
34.Wu M, Zhu L, Feng X* (2018). Network-based feature screening with applications to genome data. The Annals of Applied Statistics, 12: 1250-1270.
35.Wu M, Huang J, Ma S* (2018). Identifying gene-gene interactions using penalized tensor regression. Statistics in Medicine, 37(4): 598-610.
36.Wu M,Zang Y, Zhang S,Huang J, Ma S* (2017). Accommodating missingness in environmental measurements in gene-environment interaction analysis. Genetic Epidemiology, 41: 523-554.
37.Teran Hidalgo SJ, Wu M, Ma S* (2017). Assisted clustering of gene expression data using ANCut. BMC Genomics, 18: 623.
38.Wu M, Zhang X, Dai D*, Ou-Yang L, Zhu Y, Yan H (2016). Regularized logistic regression with network-based pairwise interaction for biomarker identification in breast cancer. BMC Bioinformatics, 17:108.
39.Wu M, Dai D*, Zhang X, Zhu Y (2013). Cancer subtype discovery and biomarker identification via a new robust network clustering algorithm. PLoS ONE, 8: e66256.
40.Wu M, Dai D*, Yan H (2012). PRL-Dock: Protein-ligand docking based on hydrogen bond matching and probabilistic relaxation labeling. Proteins: Structure, Function, and Bioinformatics, 80: 2137–2153.
41.Wu M, Dai D*, Shi Y, Yan H, Zhang X (2012). Biomarker identification and cancer classification based on microarray data using Laplace naive Bayes model with mean shrinkage, IEEE/ACM Transactions on Computational Biology and Bioinformatics, 9: 1649-1662.
Rewards, Honors
2015 Shanghai University of Finance and Economics Young Teacher Teaching Competition Third Prize (Second Prize in the Science and Engineering Group)
Social Work
Elected Member of the International Statistical Institute
Referee for BMC Genomics, BMC Bioinformatics, Statistics & Probability Letters, Statistics and Its Interface, Annals of the Institute of Statistical Mathematics, etc.
Integrative clustering of multidimensional omics data. The 2018 ICSA Applied Statistics Symposium. New Jersey, USA, June, 14-17, 2018.
Robust gene-environment interaction analysis using penalized trimmed regression. International Workshop on Perspectives On High-dimensional Data Analysis (HDDA-VIII-2018). Marrakesh, Morocco, April, 09-13, 2018.
Regularized logistic regression with network-based pairwise interaction for biomarker identification in breast cancer. The 4th IBS-China International Biostatistical Conference. Shanghai, China, July, 02-03, 2016.
Regularized logistic regression with network-based pairwise interaction for biomarker identification in breast cancer. 2016 ICSA China Statistics Conference. Qingdao, China, June, 24-25, 2016.


