cv
Chenyu Wang — LLMs for software engineering, coding agents, and AI security. Click the PDF icon for the one-page résumé.
Basics
| Name | Chenyu Wang (王晨语) |
| Label | PhD Candidate & Research Engineer — LLMs for Software Engineering |
| chenyuwang@smu.edu.sg | |
| Url | https://jamesnolan17.github.io/ |
| Summary | PhD candidate and Research Engineer at SMU's SOAR Lab (advisor Prof. David Lo). First-author research on inference-time control of coding agents, code-model security, and repository-level vulnerability detection; published at ASE and EASE, with a paper under review at a top-tier (CORE A*) venue. |
Work
-
2023.01 - Present Research Engineer
SMU SOAR Lab
Research on LLMs for software engineering under Prof. David Lo.
- First-author research on inference-time control of coding agents and code-model security (ASE 2025; EASE 2026 → IEEE Software; one paper under review at a CORE A* venue).
- Own projects end-to-end: empirical study, dataset construction, model training & fine-tuning, evaluation, and research tooling.
- In collaboration with GovTech Singapore (from 2026): repository-level agentic vulnerability detection — code property graphs plus an LLM-as-a-judge context-extraction stage.
Education
-
2023.01 - 2027.01 Singapore
Ph.D. (part-time)
Singapore Management University
Computer Science — Software Engineering & Trustworthy AI
- Advisor: Prof. David Lo (SOAR Lab)
-
2019.05 - 2022.08 Singapore
B.Eng.
Singapore University of Technology & Design
Computer Science & Design
- Minor in Artificial Intelligence
- Exchange: NUS School of Computing (2021)
Awards
- 2025.07.01
SCIS Research Excellence Award — Tier 1
SMU School of Computing and Information Systems
Awarded for a first-author A* publication (50% tuition-fee waiver).
- 2020.01.01
Best Student Award (Term 3) — ranked 1st of 400+ students
Singapore University of Technology & Design
Top-ranked student in the cohort.
- 2022.08.01
Honours with Highest Distinction
Singapore University of Technology & Design
GPA 4.83/5.0; SUTD Honours List (multiple terms).
- 2019.01.01
Senior Middle 2 Scholarship
Singapore Ministry of Education
- 2021.08.01
Selected for NUS Exchange (SUSEP)
NUS School of Computing
One of 7 students selected from the pillar.
Publications
-
2027.01.01 Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks
Under review, 2027 · CORE A*
A two-stage controller with two complementary modes. Fail-Fast (cost): a 0.6B monitor early-terminates doomed runs — 14.6–20.4% token savings at a 5% false-positive rate, transferring across four policies including a closed-API model. Restart-Smart (quality): on an alarm, relaunches a fresh rollout with the prior repo diff as an optional overlay — SWE-bench Verified resolution 66.6% → 71.8%.
-
2026.06.01 Industry Practitioners' Perspectives on AI Model Quality
EASE 2026 · invited to IEEE Software · CORE A
Mixed-methods study (15 interviews + 50-practitioner survey) of how industry actually manages AI quality — e.g., most practitioners never monitor deployed models for data drift, a critical blind spot for production ML.
-
2025.11.01 Backdoors in Code Summarizers: How Bad Is It?
ASE 2025 · CORE A*
First systematic (9-factor) study of data-poisoning backdoors in code LLMs: just 20 of 454,451 samples (0.004%) implant a backdoor the standard spectral-signature defense fails to remove, and smaller training batches make the attack stronger — a concrete, under-estimated AI supply-chain risk.
-
2024.12.01 Gotcha! This Model Uses My Code! Evaluating Membership Leakage Risks in Code Models
IEEE TSE 2024 · co-author
A membership-inference attack for code models that jointly uses model input, output, and ground truth to detect training-data membership.
-
2024.04.01 Unveiling Memorization in Code Models
ICSE 2024 · CORE A* · co-author
Empirical study of how large code models memorize training data, with security and privacy implications.
-
2023.05.01 What Do Users Ask in Open-Source AI Repositories? An Empirical Study of GitHub Issues
MSR 2023 · CORE A · co-author
A taxonomy of GitHub issues across 576 open-source AI repositories, characterizing what users ask and how issues get resolved.
Skills
| AI-Agent Orchestration & AI-Native Software Engineering | |
| Autonomous task decomposition | |
| Multi-agent coordination | |
| Multi-server GPU pipelines | |
| GitHub / cloud-storage synchronization | |
| Experiment & research-workflow automation | |
| Human-in-the-loop agent supervision |
| ML / LLM | |
| PyTorch | |
| Hugging Face | |
| LoRA / PEFT fine-tuning | |
| LLM agents & SWE-bench benchmarking | |
| Code LLMs (Qwen, CodeT5, Gemma) |
| Program Analysis & Security | |
| Code property graphs | |
| Program slicing & taint analysis | |
| Backdoor & adversarial robustness | |
| Vulnerability detection |
| Programming | |
| Python, Java, JavaScript (proficient) | |
| C, TypeScript, Shell (working) | |
| Docker, Git, Linux | |
| React / full-stack web |
Languages
| English | |
| Fluent |
| Mandarin Chinese | |
| Native speaker |