To further promote academic exchanges and collaborative innovation in the field of data science in the Yangtze River Delta region, the 2026 Yangtze River Delta Data Science Symposium was successfully held at Shanghai University of Finance and Economics on January 3, 2026. The symposium was hosted by the Big Data Research Institute of Shanghai University of Finance and Economics, organized by the School of Statistics and Data Science of Shanghai University of Finance and Economics, and co-organized by the Institute of Data Science and Statistics of Shanghai University of Finance and Economics and the Shanghai Frontier Science Research Base of Data Technology and Decision-making. The symposium focused on the frontiers of data science, statistics, and their cross-disciplinary applications, attracting experts, scholars, young teachers, and graduate students from many universities and research institutes across the country.

The opening ceremony of the conference was presided over by Professor Huang Tao, Vice Dean of the School of Statistics and Data Science at Shanghai University of Finance and Economics.

Professor Feng Xingdong, Dean of the School of Statistics and Data Science at Shanghai University of Finance and Economics, first extended his warm welcome to all experts, scholars, and guests. He also expressed his sincere gratitude to everyone for attending the 2026 Yangtze River Delta Data Science Symposium despite their busy schedules. In his speech, Dean Feng pointed out that against the backdrop of the rapid development of big data and artificial intelligence, the research paradigm and application scenarios of data science and statistics are constantly expanding, which not only breeds new opportunities for discipline development but also poses higher requirements for theoretical methods, model construction, data governance, and practical applications. Dean Feng stated that this symposium, based on the characteristics of the Yangtze River Delta region, focuses on the frontier directions and key issues in the field of data science and statistics, bringing together experts and scholars from multiple universities and research institutions for in-depth exchanges and discussions. This will help promote academic idea collisions, drive cross-integration, and collaborative innovation. He looked forward to using this meeting as an opportunity to further strengthen academic exchanges and cooperation, improve the talent cultivation system, and continuously contribute to the high-quality development of disciplines related to data science and statistics.

The morning conference presentation was chaired by Professor Hu Feifang from the Department of Statistics at George Washington University and Associate Professor Qiu Yixuan from the School of Statistics and Data Science at Shanghai University of Finance and Economics.


First, Professor Xueming He from the Department of Statistics and Data Science at Washington University in St. Louis delivered a report titled "Leveraging AI for Statistical Analysis — Some Non-Random Thoughts". Professor He discussed the relationship between artificial intelligence and traditional statistical methods in depth, pointing out that in many practical problems, statistical analysis often faces the situation of "having input variables but lacking complete outcome variables", which poses new challenges to traditional statistical modeling. The report focused on the important role of synthetic data in data analysis. Professor He pointed out that synthetic data can effectively alleviate the problem of insufficient sample size on the one hand, and has unique advantages in data privacy protection and sensitive information processing on the other hand. With the widespread application of artificial intelligence technology, the use of synthetic data in practical research and applications has become increasingly common. However, it is also necessary to pay attention to issues such as the reliability of statistical inference and the applicability of methods that arise from this. Professor He further emphasized that in the era of artificial intelligence, the important value of statistics is not only reflected at the algorithm level, but also in the theoretical support and methodological norms for the generation, evaluation, and use of synthetic data. Related research is of great significance for improving the robustness and interpretability of statistical learning. The content of the report sparked widespread attention and in-depth discussion among the participating scholars.

Professor Guo Zijian from Zhejiang University delivered a conference report titled "Multi-Source Learning via Distributionally Robust Optimization". Focusing on the ubiquitous issue of distribution heterogeneity in multi-source data integration, Professor Guo explored how to construct statistical learning models with good generalization and transfer performance under the condition of multiple data sources. The report proposed a unified analytical framework centered on distributionally robust optimization, which achieves effective control over worst-case risks by characterizing the uncertainty of the underlying target distribution. The related research provides important theoretical support and methodological insights for multi-source learning, transfer learning, and causal invariance analysis.

Professor Zhang Jingfei from the Department of Information Systems and Operations Research at Emory University's School of Business delivered a presentation titled "Generalized Tensor Completion with Non-Random Missingness". Focusing on the ubiquitous issue of non-random missingness in real-world data analysis, Professor Zhang explored the limitations of traditional tensor completion methods in the face of complex missing mechanisms. The presentation introduced a generalized tensor completion analysis framework that characterizes data missing mechanisms while enhancing the applicability and robustness of the model under noisy data conditions. The related research provides new methodological insights for statistical modeling and inference in high-dimensional complex data scenarios such as recommendation systems and medical imaging.

Zhang Bo, an associate professor in the Department of Statistics and Finance at the University of Science and Technology of China, delivered a conference report titled "Identifying the Structure of High-Dimensional Time Series via Eigen-Analysis". Focusing on the cross-sectional correlation and non-stationarity issues commonly present in high-dimensional time series data, Associate Professor Zhang explored how to identify the structural characteristics of high-dimensional time series based on eigenvalue analysis. The report presented a systematic analytical approach to distinguish between different types of factor structures and time evolution characteristics, thereby enhancing the effectiveness of modeling and prediction for high-dimensional time series. The related research provides an important reference for the statistical analysis of complex time series data in fields such as economics and finance.

Professor Mao Xiaojun from Shanghai Jiao Tong University delivered a presentation titled "Fair Regression in Reproducing Kernel Hilbert Spaces: Single-Machine and Decentralized Implementations under Conditional Mean Parity", which focused on introducing the theoretical basis and implementation methods of fairness constraints in machine learning modeling. Starting from the challenges faced by algorithm fairness in practical applications, the presentation elaborated on how to achieve a balance between performance and fairness in different computing environments through kernel regression methods under the framework of conditional mean parity. Professor Mao also demonstrated the scalability and application potential of related methods in large-scale data analysis, incorporating distributed computing scenarios. The presentation provided a new perspective for fair machine learning research and received widespread attention from the attending scholars.

The afternoon report was hosted by Associate Professor Zhou Fan from the School of Statistics and Data Science at Shanghai University of Finance and Economics.

Associate Professor He Xin from the School of Statistics and Data Science at Shanghai University of Finance and Economics began with a presentation titled "Kernel Ridge Regression with Predicted Feature Inputs". Starting from the classic status of kernel methods in statistical learning, Professor He Xin focused on the common situation in real-world regression problems where feature variables cannot be directly observed and need to be predicted or characterized using artificial intelligence methods such as neural networks. He systematically explored the theoretical analysis framework of kernel ridge regression in this context. The presentation combined specific applications to demonstrate the modeling approach and predictive performance of kernel methods under the condition of implicit feature inputs, after generating feature representations using artificial intelligence models. It was validated through simulations and real-world data analysis. The related research provides a new theoretical perspective for the deep integration of statistical learning methods and artificial intelligence technology.

Assistant Professor Tu Jiyuan from the School of Statistics and Data Science at Shanghai University of Finance and Economics delivered a conference presentation titled "Robust Variational Bayes by Min-Max Median Aggregation". Tu Jiyuan's presentation focused on robust Bayesian inference methods, with a particular emphasis on how to enhance the robustness of Bayesian inference through min-max thinking under conditions of uncertainty or abnormal disturbances in distributions. The presentation provided a methodological perspective on the applicability of relevant models in complex data environments, offering new insights for robust statistical inference research. The content of the presentation sparked active discussions among the attending scholars.

In the afternoon of the same day, the meeting entered the roundtable discussion session. The participating young teachers made brief introductions around their respective research directions in turn, sharing their research thoughts and practical experiences in the fields related to data science and statistics. Feng Xingdong, Dean of the School of Statistics and Data Science at Shanghai University of Finance and Economics, proposed during the discussion that the regular holding of data science seminars in the Yangtze River Delta region should be further promoted, and a stable and high-level academic exchange platform should be continuously built to facilitate long-term cooperation and collaborative development between universities and research institutions in the region.

In the subsequent discussion, the participating experts and scholars engaged in in-depth exchanges on several common issues in statistical research, including data acquisition and quality control, as well as the practical application of statistical methods in real-world scenarios. At the same time, the scholars also thoroughly discussed how statistics can better serve practical needs in the new era, as well as issues related to the construction of undergraduate and graduate curriculum systems and talent cultivation models.
This symposium has established a high-level academic exchange platform for the field of data science and statistics in the Yangtze River Delta region, effectively facilitating the connection between cutting-edge research and practical applications. The successful holding of the conference not only deepens academic cooperation within the region but also provides valuable insights for the innovative development and talent cultivation of related disciplines in the new era. The School of Statistics and Data Science at Shanghai University of Finance and Economics will take this opportunity to continuously promote high-quality academic exchanges and contribute to the long-term development of the discipline of data science and statistics.


