cv
Curriculum Vitae - Yuyang Wu
Basics
| Name | Yuyang Wu |
| Label | Ph.D. Student |
| yuyangwu@andrew.cmu.edu | |
| Phone | 412-478-1680 |
| Url | https://youngerwu.com |
| Summary | Ph.D. student at Carnegie Mellon University focusing on AI for Science, at the intersection of artificial intelligence, computational chemistry, and natural language processing. |
Work
-
2024.08 - Present Graduate Research Assistant
Carnegie Mellon University
Conducting research on improving trustworthiness of large language models in chemistry and developing large-scale quantum chemistry datasets.
- Led MolErr2Fix benchmark project for evaluating LLM trustworthiness in chemistry
- Created multi-million-entry quantum chemistry dataset
- Collaborated with Isayev and CLAW research groups
Education
-
2026.08 - 2030.05 Pittsburgh, PA
-
2024.08 - 2026.05 Pittsburgh, PA
M.S.
Carnegie Mellon University, School of Computer Science
Automated Science
- Machine Learning
- Active Learning
- Bioinformatics
- Automated Experimentation
-
2023.05 - 2023.07 Singapore
Visiting Scholar
National University of Singapore, School of Computing
Visiting Scholar
- Deep Learning
- Neural Networks
- Robotics
-
2020.09 - 2024.06 Wuhan, China
B.S.
Huazhong Agricultural University
Zhang Zhidong Class - Advanced Class
- Deep Learning & Neural Network
- Database Systems
- Data Structures & Algorithms
Publications
-
2025.08.20 MolErr2Fix: Benchmarking LLM Trustworthiness in Chemistry via Modular Error Detection, Localization, Explanation, and Revision
EMNLP 2025 (Oral Presentation)
Developed comprehensive benchmark for evaluating large language models' trustworthiness in chemistry through modular error detection, localization, explanation, and revision tasks.
-
2024.01.01 Multi-View Representation Learning for Identification of Novel Cancer Genes and Their Causative Biological Mechanisms
Briefings in Bioinformatics
Integrated multi-omics data (genomic, transcriptomic, epigenomic, PPI) to identify 74 high-confidence candidate cancer genes.
-
2023.01.01 HimGNN: Hierarchical Molecular Graph Representation Learning Framework for Property Prediction
Briefings in Bioinformatics
Designed hierarchical GNN with atom- and motif-level features and Transformer-based local augmentation module, improving performance across 8 benchmark datasets.
Skills
| Programming | |
| Python | |
| Golang | |
| C++ |
| Machine Learning | |
| PyTorch | |
| TensorFlow | |
| Scikit-learn |
| Data Science | |
| NumPy | |
| Pandas | |
| Matplotlib |
| Systems & Tools | |
| Linux/Unix | |
| Bash | |
| MySQL | |
| Git |
Projects
- 2025.02 - 2025.04
CHEM-AL: Active Learning for Molecular Property Prediction
Built active learning framework with uncertainty, QBC, and diversity sampling. Implemented Naïve Bayes and MLP (with MC Dropout) on MoleculeNet datasets.
- Achieved 50% reduction in labeling requirements versus random sampling
- 2025.02 - 2025.04
Computational Prediction of Protein–Protein Interactions (PPI)
Designed deep learning pipelines for SHS27K benchmark (16,912 interactions). Developed ProBERT-BiGRU-Attention and Transformer-based models.
- Achieved 83% accuracy and AUC 0.91
- Built GNN capturing structural context for unseen proteins
- 2025.04 - Present
Large-Scale Multi-Property Quantum Chemistry Dataset
Created a multi-million-entry dataset including energies, forces, dipoles, quadrupoles. Benchmarked SchNet, PaiNN, E2GNN, Equiformer for molecular property prediction.
- Preparing publication and open dataset release