Research Vision

Applying advanced AI to automate software implementation, testing, and program verification

LLM4SE LLM4FM Code Agents Benchmarking

I am a Research Assistant Professor at the Department of Computer Science and Engineering, The Hong Kong University of Science and Technology. I will join the Department of Computing, Imperial College London as an Assistant Professor in late 2026. I received the PhD degree from the Department of Computer Science and Engineering at The Hong Kong University of Science and Technology (HKUST), under the supervision of Prof. Shing-Chi Cheung in the CASTLE lab.

My research focuses on applying advanced AI techniques to automate software implementation, testing, and program verification, to produce software that is reliable by construction. My research interests lie in the intersection of Software Engineering (SE), Large Language Models (LLMs), with an emphasis on LLM4SE, LLM4FM (formal methods), and LLM Evaluation/Benchmarking. I have 39 publications at top conferences and journals, including ICSE, FSE, ASE, TOSEM, CAV, Usenix Security, AAAI, etc. I serve as a program committee member in top conferences such as ICSE, FSE, ASE, and ISSTA, and am a reviewer for TOSEM, TSE, and EmSE. My doctoral dissertation, "Towards Automatic Testing and Fault Localization in Natural Language Processing Systems", was recognized with the πŸ† ACM SIGSOFT Outstanding Doctoral Dissertation Award for 2025. Additionally, I was honored with the πŸ† 2025 Young Scientist Award in Engineering Science (one awardee per year in Hong Kong).

πŸ†
ACM SIGSOFT Outstanding Dissertation Award 2025
Only 1-2 recipients worldwide per year
πŸ†
Young Scientist Award in Engineering Science 2025
Only 1 recipient per year in Hong Kong
ICSEβ€’FSEβ€’ASEβ€’CAVβ€’ACLβ€’ICMLβ€’AAAIβ€’USENIX Securityβ€’TOSEMβ€’TSEβ€’ ICSEβ€’FSEβ€’ASEβ€’CAVβ€’ACLβ€’ICMLβ€’AAAIβ€’USENIX Securityβ€’TOSEMβ€’TSEβ€’

πŸ”₯ Open Positions @ Imperial College London

I will join the Department of Computing, Imperial College London as an Assistant Professor in November 2026.

I'm looking for 1–2 fully funded PhD students (home & international, 2026 Fall / 2027 Spring intake) to build reliable, scalable, and cost-efficient AI systems for real-world software engineering and formal verification.

Coding Agents AI-Assisted Formal Verification AI for Software Security Post-training of Coding Models

Let's do something interesting and impactful!

πŸ‘‰πŸ“§ β†’ Full details: project, research directions, requirements & how to apply

πŸ“š β†’ How to submit the PhD application and see more requirements

News
2026.07.29πŸŽ‰ Two papers accepted by ASE 2026! Congrats!
2026.06.25πŸŽ‰ Two papers accepted by ISSTA 2026! Congrats!
2026.05.02πŸŽ‰ Two papers accepted by ICML 2026! SWE-ABS and Position: Code Benchmarks.
2026.02.21Skills-4-SE released β€” 180+ Claude Skills for SE. [Website]
2026.01.08πŸŽ‰ Code Translation via Pseudocode accepted by TOSEM 2026. Congrats to Songqiang!
2025.12.17πŸŽ‰ Paper accepted by ICSE 2026. Congrats to Ruiyang!
2025.12.13πŸ† Honored with Young Scientist Award in Engineering Science. [News]
2025.05.15πŸŽ‰ From Informal to Formal and CruxEval-X accepted by ACL 2025 Main!
2025.03.08HuggingFace repo reached 10k+ downloads. Social media reached 37k+ reads.
Research Topics
LLM Benchmark8+
code generationcode reasoningdata contaminationmultilingualadversarial
LLM for SE12+
code translationbug reproductionprogram repairembeddedAPI docs
LLM for Formal Methods5+
theorem provingspecificationmodel checkingTLA+formal proofs
SE for AI3+
NLP testingDL testingfault localizationmetamorphicsecurity
β–Ά Browse all
Open-Source Projects/Products

KeyLight

Read less, understand more

AI-powered academic PDF reader that auto-highlights key insights, chats with your paper, and generates structured digests β€” so you can grasp a paper in minutes.

Auto-highlight Chat with Paper Paper Digest PDF
PDFAuto-highlight
Read less, understand more.
Problem Idea Challenge Solution Finding
PDFAuto-highlight
Read less, understand more.
Problem Idea Challenge Solution Finding

VeriFormal

AI-powered formal verification

Combines LLMs with static analysis to automatically generate specifications for verifying C, Java, Rust, and Dafny programs.

Formal Methods C Java Rust Dafny
C Code
int abs_val(int x) {
  if (x < 0) return -x;
  return x;
}
C + ACSLVerified
/*@ requires x > INT_MIN;
  @ ensures \result >= 0;
  @ ensures \result == x
  @     || \result == -x; @*/
int abs_val(int x) {
  if (x < 0) return -x;
  return x;
}

EasyTODO

Simple. Free. Always on your desktop.

A native macOS todo app built to stay visible and make capture instant β€” global quick add, color-coded priorities, menu bar progress, and a floating widget that follows you across spaces. Local-first, no account, no setup.

macOS SwiftUI Menu Bar Quick Add
β˜‘ EasyTODO 0 / 5
Today
  • Submit ISSTA camera-ready
  • Review student draft
  • Reply to reviewer #2
  • Prepare group meeting slides
  • Book flight to conference
Cmd+ Quick Add β€” from any app

Skills-4-SE

Browse, search and install skills for software engineering

A curated collection of 180+ Claude Skills spanning the full development lifecycle, with a Skills Manager web interface to search, filter by category or stage, and install everything or just the skills you pick β€” plus 8 curated Skill Packs.

Skills Manager UI 180+ Skills Skill Packs One-click Install
ArabelaTso.github.io/Skills-4-SE
Skills-4-SE Manager
Manage and install Claude Code skills
Testing Verification Quality DevOps
180 skills Β· 0 selected Install Selected
Publications
52
Google Scholar Citations
202622 papers
ICMLICSEFSEISSTAACLFMCOLMASETASETOSEM
β–Ά
ICSE[C26] Ruiyang Xu, Jialun Cao (Co-1st), Mingyuan Wu, Wenliang Zhong, Yaojie Lu, Ben He, Xianpei Han, Shing-Chi Cheung, Le Sun. EmbedAgent: Benchmarking Large Language Models in Embedded System Development. In ICSE 2026. [Paper]
ACL Findings[C27] Qiming Zhu, Jialun Cao, Xuanang Chen, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun, Shing-Chi Cheung. Across Programming Language Silos: A Study on Cross-Lingual Retrieval-augmented Code Generation. In ACL Findings 2026. [Paper]
FM Tool[C28] Zhiyong Chen, Jialun Cao, Chang Xu, Shing-Chi Cheung. ModelWisdom: An Integrated Toolkit for TLA+ Model Visualization, Digest and Repair. In FM 2026 Tool track. [Paper] [Website]
FSE[C29] Junyi Wang, Jialun Cao, Zhongxin Liu. iCoRe: Iterative Correlation-Aware Retriever for Bug Reproduction. In FSE 2026. [Paper]
ICML Pos.[C30] Jialun Cao, Yuk-Kit Chan*, Zixuan Ling*, Wenxuan Wang†, Shuqing Li, Mingwei Liu, Ruixi Qiao, Yuting Han, Chaozheng Wang, Boxi Yu, Pinjia He, Shuai Wang, Zibin Zheng, Michael R. Lyu, Shing-Chi Cheung. Position: Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility. In ICML Position 2026. [Paper]
ICML[C31] Boxi Yu, Yang Cao, Yuzhong Zhang, Liting Lin, Junjielong Xu, Zhiqing Zhong, Qinghua Xu, Guancheng Wang, Jialun Cao, Shing-Chi Cheung, Pinjia He, Lionel Briand. SWE-ABS: Adversarial Benchmark Strengthening Exposes Inflated Success Rates on Test-based Benchmark. In ICML 2026. [Paper]
ISSTA[C32] Dongze Li, Songqiang Chen, Jialun Cao (Corresponding), Shing-Chi Cheung (Corresponding). What Builds Effective In-Context Examples for Code Generation? In ISSTA 2026. [Paper]
ISSTA[C33] Jialun Cao, Haoyu Wang*, Haoran Yan*, Ming Wen, Michael Pradel. STARS: Static Analysis-Guided Assertion Synthesis using Large Language Models. In ISSTA 2026.
COLM[C34] Yuling Shi, Jinghan Xu, Kelin Fu, Wenhao Zeng, Yingwei Ma, Shilin He, Lei Zhang, Yue Liu, Zelin Zhao, Terry Yue Zhuo, Jialun Cao, Siyu Ye, Tianyu Liu, Kai Cai, Shing-Chi Cheung, Xiaodong Gu. SWE-Cascade: Benchmarking Agents on Large-Scale Multilingual Code Refactoring. In COLM 2026.
ASE T&D[C35] Junjie Hu, Cheng Wen, Bin Yu, Jialun Cao, Dugang Liu, Weidi Sun, Haokun Li, Shengchao Qin, Cong Tian. ACSLBench: A Verified C/ACSL Corpus and Benchmark with Compositional Call Chains for Formal Specification Synthesis. In ASE 2026 Tools-and-Datasets.
ASE T&D[C36] Yuchen Zhang, Cheng Wen, Jialun Cao, Dugang Liu, Zhiwu Xu, Yuwei Liu, Shengchao Qin. VerusSeek: Retrieval-Augmented LLM-Based Proof Synthesis for Rust Programs. In ASE 2026 Tools-and-Datasets.
ASE[C37] Xiaolei Li, Jialun Cao, Zhijian Hou, Yuzhi Zhao, Yepang Liu, Shing-Chi Cheung. GraphDroid: Asynchronous LLM-Based Mobile App GUI Testing via History-Aware Exploration and Hybrid Intent Fulfillment. In ASE 2026.
ASE[C38] Zhonghao Jiang, Le Deng, Jialun Cao, Michael Pradel, Zhongxin Liu. Doc2Feat-bench: Evaluating Documentation-Driven Feature Addition. In ASE 2026. [Paper] [Leaderboard] [HF]
ISSTA Tool[C39] Wenjie Wu, Junjie Hu, Cheng Wen, Jialun Cao, Dugang Liu, Zhiwu Xu, Weidi Sun, Haokun Li, Shengchao Qin. Spec-Skill: A Pluggable Coding-Agent Plugin for Neuro-Symbolic Program Specification Synthesis. In ISSTA 2026 Tool track.
TASE[C40] Cheng Wen, Zhiwu Xu, Dugang Liu, Jialun Cao, Yuwei Liu, Shengchao Qin, Cong Tian. Enhancing LLM-Based Proof Synthesis for Rust Programs via Semantic Chunking and Hierarchical Context. In TASE 2026.
TOSEM[J5] Songqiang Chen, Congying Xu, Jingyi Chen, Jialun Cao (Corresponding), Jiarong Wu, Shing-Chi Cheung (Corresponding). Can Emulating Semantic Translation Help LLMs with Code Translation? A Study Based on Pseudocode. In TOSEM 2026. [Paper]
TOSEM[J6] Jingyi Chen, Songqiang Chen, Jialun Cao (Corresponding), Jiasi Shen (Corresponding), Shing-Chi Cheung. When Retrieval Augmentation Meets API Documentation: Can LLMs Code with Less-Common Libraries? In TOSEM 2026. [Paper]
arXiv[Pre5] Zhiyong Chen, Jialun Cao, Jiarong Wu, Chang Xu, Shing-Chi Cheung. Can Large Language Models Model Programs Formally? arXiv. [Paper]
arXiv[Pre6] Dong Xu, Jialun Cao, Guozhao Mo, Junjie Hu, Cheng Wen, Hongyu Lin, Xianpei Han, Shengchao Qin, Cong Tian, Shing-Chi Cheung, Le Sun, Yaojie Lu. LiveFMBench: Unveiling the Power and Limits of Agentic Workflows in Specification Generation. arXiv. [Paper] [HF]
arXiv[Pre7] Yuanyi Wang, Yifan Yang, Su Lu, Yanggan Gu, Pengkai Wang, Wenjun Wang, Zhaoyi Yan, Congkai Xie, Jianmin Wu, Jialun Cao, Shing-Chi Cheung, Hongxia Yang. Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training. arXiv. [Paper]
arXiv[Pre8] Jialun Cao, Xinru Yan (Co-1st), Songqiang Chen, Yaojie Lu, Zhongxin Liu, Shing-Chi Cheung. Inside the Skill Market: From Software Engineering Activities to Reusable Agent Skills. arXiv. [Paper]
arXiv[Pre9] Jingyi Chen, Songqiang Chen, Hengcheng Zhu, Jialun Cao (Corresponding), Jiasi Shen (Corresponding), Shing-Chi Cheung. Understanding Agent-Reactive Bugs at the Model-Harness Boundary: An Empirical Study of LLM Agent Issue Reports. arXiv. [Paper]
202513 papers
ACLFSEAAAIASETOSEM
β–Ά
Internetware[C19] Jialun Cao, Songqiang Chen, Wuqi Zhang, Hau Ching Lo, Yeting Li, Shing-Chi Cheung. CodeCleaner: Elevating Standards with A Robust Data Contamination Mitigation Toolkit. In Internetware 2025. [Paper] [Code]
ACL[C20] Jialun Cao, Yaojie Lu, Meiziniu Li, Haoyang Ma, Haokun Li, Mengda He, Cheng Wen, Le Sun, Hongyu Zhang, Shengchao Qin, Shing-Chi Cheung, Cong Tian. From Informal to Formal β€” Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs. In ACL 2025. [Paper] [HF] [Media]
FSE[C21] Xiao Chen, Hengcheng Zhu, Jialun Cao (Corresponding), Ming Wen, Shing-Chi Cheung. SemBIC: Semantic-aware Identification of Bug-inducing Commits. In FSE 2025. [Paper]
AAAI[C22] Qiming Zhu, Jialun Cao (Co-1st), Yaojie Lu, Hongyu Lin, Xianpei Han, Ben He, Le Sun, Shing-Chi Cheung. DomainEval: An Auto-Constructed Benchmark for Multi-Domain Code Generation. In AAAI 2025. [Paper] [Leaderboard] [Code]
ACL[C23] Ruiyang Xu, Jialun Cao (Co-1st), Yaojie Lu, Ming Wen, Hongyu Lin, Xianpei Han, Ben He, Shing-Chi Cheung, Le Sun. CruxEval-X: A Benchmark for Multilingual Code Reasoning, Understanding and Execution. In ACL 2025. [Paper] [Leaderboard] [Code]
AAAI[C24] Mengyang Wu, Yuzhi Zhao, Jialun Cao, Mingjie Xu, Zhongming Jiang, Xuehui Wang, Qinbin Li, Guangneng Hu, Shengchao Qin, Chi-Wing Fu. ICM-Assistant: Instruction-tuning Multimodal Large Language Models for Rule-based Explainable Image Content Moderation. In AAAI 2025. [Paper]
ASE[C25] Xingchu Chen, Chengwei Liu, Jialun Cao, Yang Xiao, Xinyue Cai, Yeting Li, Jingyi Shi, Tianqi Sun, Haiming Chen, Wei Huo. Vulnerability-Affected Versions Identification: How Far Are We? In ASE 2025.
ASEJ[J3] Jialun Cao, Meiziniu Li, Ming Wen, Shing-chi Cheung. A study on prompt design, advantages and limitations of chatgpt for deep learning program repair. In ASEJ 2025. [arxiv] [Official]
TOSEM[J4] Meiziniu Li, Dongze Li, Jianmeng Liu, Jialun Cao, Yongqiang Tian, Shing-Chi Cheung. Enhancing Differential Testing With LLMs For Testing Deep Learning Libraries. In TOSEM 2025. [Paper]
arXiv[Pre1] Jialun Cao, Wuqi Zhang, Shing-Chi Cheung. Concerned with Data Contamination? Assessing Countermeasures in Code Language Model. arXiv. [Paper]
arXiv[Pre2] Jiarong Wu, Songqiang Chen, Jialun Cao (Corresponding), Hau Ching Lo, Shing-Chi Cheung. Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval. arXiv. [Paper]
arXiv[Pre3] Dekun Dai, MingWei Liu, Anji Li, Jialun Cao, Yanlin Wang, Chong Wang, Xin Peng, Zibin Zheng. FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks. arXiv. [Paper]
arXiv[Pre4] Xiaolei Li, Jialun Cao, Yepang Liu, Shing-Chi Cheung, Hailong Wang. ReuseDroid: A VLM-empowered Android UI Test Migrator Boosted by Active Feedback. arXiv. [Paper]
20246 papers
ASECAVAPSECTASE
β–Ά
ASE[C13] Jialun Cao, Zhiyong Chen*, Jiarong Wu, Shing-chi Cheung, Chang Xu. JavaBench: A Benchmark of Object-Oriented Code Generation for Evaluating Large Language Models. In ASE 2024. [Paper] [Leaderboard] [Code]
ASEπŸ† [C14] Distinguished paper award. Zongze Jiang, Ming Wen, Jialun Cao, Xuanhua Shi, Hai Jin. Towards Understanding the Effectiveness of Large Language Models on Directed Test Input Generation. In ASE 2024. [Paper] [Code]
ASE[C15] Congying Xu, Songqiang Chen, Jiarong Wu, Valerio Terragni, Shing-chi Cheung, Hengcheng Zhu, Jialun Cao (Corresponding). MR-Adopt: Automatic Deduction of Input Transformation Function for Metamorphic Testing. In ASE 2024. [Paper]
CAV[C16] Cheng Wen, Jialun Cao (Corresponding), Jie Su, Zhiwu Xu, Shengchao Qin, Mengda He, Haokun Li, Shing-Chi Cheung, Cong Tian. Enchanting Program Specification Synthesis by Large Language Models using Static Analysis and Program Verification. In CAV 2024. [Paper] [Homepage]
APSEC[C17] Bo Yang, Jiawei Hu, Jialun Cao (Corresponding). SDEFL: A Lightweight Fault Detection and Localization Method for Deep Neural Networks. In APSEC 2024.
TASE[C18] Kunpeng Jian, Yanyan Zou, Yeting Li, Jialun Cao, Menghao Li, Jian Sun, Jingyi Shi, Wei Huo. Fuzzing for Stateful Protocol Implementations: Are We There Yet? In TASE 2024.
20233 papers
FSETOSEM
β–Ά
FSE[C11] Jialun Cao, Yaojie Lu, Ming Wen, Shing-Chi Cheung. Testing Coreference Resolution Systems without Labeled Test Sets. In FSE 2023. [Paper] [Code]
FSE[C12] Xiaohu Du, Xiao Chen, Jialun Cao, Ming Wen, Shing-Chi Cheung, Hai Jin. Understanding the Bug Characteristics and Fix Strategies of Federated Learning Systems. In FSE 2023. [Paper]
TOSEM[J2] Meiziniu Li, Jialun Cao, Yongqiang Tian, Tsz On Li, Ming Wen, Shing-Chi Cheung. COMET: Coverage-guided Model Generation For Deep Learning Library Testing. In TOSEM 2023. [Paper] [Code]
20223 papers
ICSEUSENIX SecTOSEM
β–Ά
ICSE[C9] Jialun Cao, Meiziniu Li, Xiao Chen, Ming Wen, Yongqiang Tian, Bo Wu, Shing-chi Cheung. DeepFD: Automated Fault Diagnosis and Localization for Deep Learning Programs. In ICSE 2022. [Paper] [Code]
USENIX Sec[C10] Yeting Li, Yecheng Sun, Zhiwu Xu, Jialun Cao, Yuekang Li, Rongchen Li, Haiming Chen, Shing-Chi Cheung, Yang Liu, Yang Xiao. RegexScalpel: Regular Expression Denial of Service (ReDoS) Defense by Localize-and-Fix. In USENIX Security 2022. [Paper]
TOSEM[J1] Jialun Cao, Meiziniu Li, Yeting Li, Ming Wen, Shing-chi Cheung. SemMT: A Semantic-Based Testing Approach for Machine Translation Systems. In TOSEM 2022. [Paper] [Code]
20212 papers
ICSEUSENIX Sec
β–Ά
USENIX Sec[C7] Yeting Li, Zixuan Chen, Jialun Cao, Zhiwu Xu, Qiancheng Peng, Haiming Chen, Liyuan Chen, Shing-Chi Cheung. ReDoSHunter: A Combined Static and Dynamic Approach for Regular Expression DoS Detection. In USENIX Security 2021. [Paper]
ICSE[C8] Yeting Li, Shuaimin Li, Zhiwu Xu, Jialun Cao, Zixuan Chen, Yun Hu, Haiming Chen, Shing-Chi Cheung. TransRegex: Multi-modal Regular Expression Synthesis by Generate-and-Repair. In ICSE 2021. [Paper]
20202 papers
ICDEASE
β–Ά
ICDE[C5] Yeting Li, Jialun Cao, Haiming Chen, Tingjian Ge, Zhiwu Xu, Qiancheng Peng. FlashSchema: Achieving High Quality XML Schemas with Powerful Inference Algorithms and Large-scale Schema Data. In ICDE 2020. [Paper]
ASE[C6] Yeting Li, Zhiwu Xu, Jialun Cao, Haiming Chen, Tingjian Ge, Shing-Chi Cheung. FlashRegex: Deducing Anti-ReDoS Regexes from Examples. In ASE 2020. [Paper]
20192 papers
ICCDDASFAA
β–Ά
ICCD[C3] Yongjian Li, Jialun Cao (1st student author), Jun Pang. A Learning-Based Framework for Automatic Parameterized Verification. In ICCD 2019. [Paper]
DASFAA[C4] Yeting Li, Xiaolan Zhang, Jialun Cao, Haiming Chen, Chong Gao. Learning k-Occurrence Regular Expressions with Interleaving. In DASFAA 2019. [Paper]
20182 papers
ASEQRS
β–Ά
ASE Demo[C1] Jialun Cao, Yongjian Li, Jun Pang. L-CMP: an automatic learning-based parameterized verification tool. In ASE 2018 Demo. [Paper] [Code] [Video]
QRS[C2] Yongjian Li, Jialun Cao (1st student author), Kaiqiang Duan. An automatic parameterized verification of FLASH cache coherence protocol. In QRS 2018. [Paper]
Teaching
2025 SprInstructor in COMP 1021 β€” Introduction to Computer Science. [Materials]
2023 FallTeaching Assistant in COMP 1021 β€” Introduction to Computer Science.
2020 FallTeaching Assistant in COMP 3021 β€” Java Programming.
2020 SprTeaching Assistant in COMP 3021 β€” Java Programming.
Honors and Awards
Service
Program Committee Member
Session Chair
Publicity Chair
Reviewer
ACL Β· TOSEM Β· TSE Β· EMSE Β· JASE Β· JSME Β· TKDD Β· TMC
Invited Talks
2025.04Exploring Code Generation and Reasoning Capabilities of LLMs. Peking University. [Link]
2025.01From Benchmarks to Practice. Huawei AI R&D Seminar. [Recording]
2025.01From Requirement to Formal Specification via LLMs. Zhejiang University.
2024.11Is LLM a Rescue for Code Generation & Reasoning? CCF China Open Source Conference.
2024.10Automatic Testing and Verification using LLMs. Xidian University.
2024.08Trusted Architecture of Intelligent CPS. Fudan University. [Link]
2024.08Can AI be a Panacea for Software Reliability? University College London.
2024.07Data Contamination in Code LMs. IEEE Cloud & AI Symposium.
2024.05From Requirement to Formal Specification via LLMs. CCF FM Seminar. [Recording]
2023.12Crafting Future: A Dancer's Leap into CS. ChinaSoft Women Scholars Forum.
2023.12Data Contamination in the Era of LLMs. ChinaSoft AIGC Forum.
Visitors
0Unique Visitors
0Countries / Regions
Copyright Β© 2026 Jialun Cao
🐠 guppyLLM-9M
Hi! I'm guppyLM-9M-fish raised by Jialun. Ask about Jialun's research, awards, or open positions!
Chat with Jialun's Fish
LLM