I am building a new research lab called ContinuousAI Lab. We develop AI systems that can continuously learn and autonomously evolve post-deployment.
- Self-evolving methods: transforming transient experiences into durable knowledge and capabilities, including continual learning, self-supervised learning, and robust mechanisms against catastrophic forgetting.
- Long-context models: modeling ultra-long contexts effectively and efficiently through architectural and algorithmic innovations, ensuring that model capabilities grow as context length scales while keeping compute requirements bounded.
- Agentic memory systems: designing memory frameworks that govern how agents organize, store, and retrieve information over its lifetime, enabling fast and accurate on-demand access.
- Applications: exploring high-impact use cases such as AI for AI, AI for Science, personal AI.
I am recruiting PhD students, master students, and interns. If you'd like to work with me, please drop me an email with your CV!
Email: thisisjcykcd AT gmail.com
Activities
- Oct. 2025, Co-organize a tutorial on LLM Role-Playing and Hallucination at IJCAI 2025
- Jul. 2022, Co-organize a tutorial on Retrieval-Augmented Text Generation at SIGIR 2022
- Jul. 2022, Co-organize a tutorial on Retrieval-Augmented Text Generation at IJCAI 2022
- May. 2024, CCF TechFrontier
- Sep. 2023, MLNLP Outstanding Speaker
- Welcome submissions to the 1st Workshop on Taming Large Language Models @ SIGDIAL 2023 & INLG 2023
- Nov. 2022, Invited Talk at Tsinghua University (hosted by Prof. Minlie Huang)
- Nov. 2022, Invited Talk at Technical University of Darmstadt (hosted by Prof. Iryna Gurevych)
- Sep. 2022, Guest Speaker, NLPCC2022 Student Workshop
- Sep. 2022, Invited Talk at Central South University
- Sep. 2022, Invited Talk at Tsinghua University (hosted by Prof. Bowen Zhou)
- Jul. 2022, Invited Talk at MLNLP Webinar
- Jul. 2022, Invited Talk at Peking University (hosted by Prof. Yuexian Zou)
- Jun. 2022, Passed my PhD thesis defense
- Mar. 2022, Invited Talk at NLG Student Webinar, Chinese Information Processing Society of China
- Mar. 2022, Invited Talk at Bytedance
- Feb. 2022, Invited Talk at The Chinese University of Hong Kong (hosted by Prof. Helen Meng)
- Jan. 2022, Invited Talk at Xiamen University (hosted by Prof. Jinsong Su)
- Dec. 2021, Invited Talk at Hunan University
- Oct. 2021, Invited Talk at Amazon AWS AI
- Sep. 2021, Invited Talk at Institute of Computing Technology, Chinese Academy of Sciences
- Jul. 2021, Invited Talk at Tencent Research
Tutorials
More
Selected Awards and Honors
- ACM MM-2024 Best Paper Nomination
- Young Elite Scientist Sponsorship Program, China Association for Science and Technology, 2023
- EACL-2023 Outstanding Reviewer
- ACL-2021 Outstanding Paper
- EMNLP-2020 Outstanding Reviewer
- AAAI-2020 Scholarship Award
- National Scholarship for Graduate Student (top 2% students), Ministry of Education of P.R.China, 2016
- Excellent Undergraduate Thesis Award, XMU, 2015
- Gold Medal, The 5-th Fujian Provincial University Programming Contest, 2014
- Bronze Medal, ACM-ICPC Asia Regional Programming Contest, 2013
Professional Service
-
Senior Program Committee Member/ Area Chair:
- COLING(2022), ACL(2024), EMNLP(2024), NAACL(2025), ACL Rolling Review(2024-), NeurIPS(2025-), ICML(2026-), ICLR(2026-), AAAI(2027) Program Committee Member/ Reviewer:
- ACL Rolling Review(2021-), ACL(2017-2023), EMNLP(2019-2023), NAACL(2021)
- NeurIPS(2022-2025), ICML(2022-2025), ICLR(2023-2025), AAAI(2020-2022), SIGKDD(2022-2023), WSDM(2022-2023)
- COLING(2016-2024), AACL(2020-2023), EACL(2017-2023), LREC(2018, 2020, 2022), INLG(2019-2021), etc Journal Reviewer:
- Computational Linguistics
- IEEE Transactions on Pattern Analysis and Machine Intelligence
- ACM Transactions on Information Systems
- IEEE Transactions on Audio, Speech and Language Processing
- Neurocomputing
- Pattern Analysis and Applications
Experience
Employment
- Research Scientist, Bytedance Seed
- Senior Researcher, Tencent AI Lab
Education
- Aug. 2018 - Jul. 2022
PhD, Dept. of Systems Engineering and Engineering Management, The Chinese University of Hong Kong - Sep. 2015 - Mar. 2018
MS, Dept. of Computer Science, Shanghai Jiao Tong University - Sep. 2011 - Jun. 2015
BE, Dept. of Computer Science, Xiamen University
Internship
- summer 2021, Applied scientist intern with Elman Mansimov and Yi Zhang
Amazon AWS AI, Seattle (remote) - spring 2021, Research intern with Yizhe Zhang, Michel Galley, and Bill Dolan
Microsoft Research, Redmond (remote) - 2018 - 2020, Student Researcher with NLP Center led by Shuming Shi
Tencent AI Lab, Shenzhen - Jan. 2018 - Mar. 2018, Research intern with Xiaobin Wang, Guangwei Xu, and Linlin Li
Alibaba DAMO Academy, Hangzhou - Jul. 2017 - Dec. 2017, Visiting scholar, Advisor: Prof. Yue Zhang
Singapore University of Technology and Design, Singapore
Papers (Google Scholar Profile)
(*: equal contribution, ☨: correspondence)
Preprints
- On the Transformations across Reward Model, Parameter Update, and In-Context Prompt [arxiv]
arXiv, 2024. - A Survey on the Honesty of Large Language Models [arxiv]
arXiv, 2024. - Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models [arxiv] [code]
arXiv, 2023. - A Survey on Retrieval-Augmented Text Generation [arxiv]
arXiv, 2022.
- Narrative Incoherence Detection [arxiv]
arXiv, 2021.
- Chinese Word Segmentation: Another Decade Review (2007-2017)
[arxiv]
ArXiv, 2017.
Publications
- InfiniteICL: Breaking the Limit of Context Window Size via Long Short-term Memory Transformations
ACL 2025 (Findings) - Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst
ACL 2025 (Findings) - Empowering Self-Learning of LLMs: Inner Knowledge Explicitation as a Catalyst
AAAI 2025 - On the Worst Prompt Performance of Large Language Models [arxiv]
NeurIPS 2024 - StrategyLLM: Large Language Models as Strategy Generators, Executors, Optimizers, and Evaluators for Problem Solving [arxiv]
NeurIPS 2024 - Unchosen Experts Can Contribute Too: Unleashing MoE Models’ Power by Self-Contrast [arxiv]
NeurIPS 2024 - GLBench: A Comprehensive Benchmark for Graph with Large Language Models
NeurIPS 2024 (Datasets and Benchmarks) - A Thorough Examination of Decoding Methods in the Era of LLMs [arxiv]
EMNLP 2024 - Consecutive Batch Model Editing with HooK Layers
EMNLP 2024 - Cross-lingual Contextualized Phrase Retrieval
EMNLP 2024 (Findings) - Not All Preference Pairs Are Created Equal: A Recipe for Annotation-Efficient Iterative Preference Learning
EMNLP 2024 (Findings) - With Greater Text Comes Greater Necessarily: Inference-Time Training Helps Long Text Generation
COLM 2024 - GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation [arxiv]
ACM MM 2024
Best Paper Nomination (26/4340) - Reasons to Reject? Aligning Language Models with Judgments [arxiv]
ACL 2024 (Findings) - TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild [blog] [demo] [arxiv] [code]
ACL 2024 (Findings) - Disperse-Then-Merge: Pushing the Limits of Instruction Tuning via Alignment Tax Reduction [arxiv]
ACL 2024 (Findings) - WatME: Towards Lossless Watermarking Through Lexical Redundancy [arxiv]
ACL 2024 - A Frustratingly Simple Decoding Method for Neural Text Generation [arxiv]
COLING 2024
- Retrieval is Accurate Generation [paper]
ICLR 2024
- The Reasonableness Behind Unreasonable Translation Capability of Large Language Model [paper]
ICLR 2024
- Knowledge Fusion of Large Language Models [paper]
ICLR 2024
- Specialist or Generalist? Instruction Tuning for Specific NLP Tasks [paper]
EMNLP 2023
- Large Language Models Meet Harry Potter: A Bilingual Dataset for Aligning Dialogue Agents with Characters [arxiv] [paper]
EMNLP 2023 (Findings)
- Repetition In Repetition Out: Towards Understanding Neural Text Degeneration from the Data Perspective [arxiv] [paper] [code]
NeurIPS 2023
- PandaGPT: One Model To Instruction-Follow Them All [blog] [demo] [arxiv] [code]
TLLM Workshop 2023 - Effidit: An Assistant for Improving Writing Efficiency [paper] [demo]
ACL 2023 (Demo)
- Copy is All You Need [paper]
ICLR 2023
- Retrofitting Multilingual Sentence Embeddings with Abstract Meaning Representation [arxiv] [paper] [code]
EMNLP 2022
- Linearizing Transformer with Key-Value Memory [arxiv] [paper] [code]
EMNLP 2022
- N-gram Is Back: Residual Learning of Neural Text Generation with n-gram Language Model [arxiv] [paper]
EMNLP 2022 (Findings)
- Measuring and Reducing Model Update Regression in Structured Prediction for NLP [arxiv] [blog] [paper]
NeurIPS 2022
- Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation
NeurIPS 2022
- Recent Advances in Retrieval-Augmented Text Generation
SIGIR 2022 (Tutorial)
- Recent Advances in Retrieval-Augmented Text Generation
IJCAI 2022 (Tutorial)
- Multi-Task Pre-Training for Plug-and-Play Task-Oriented Dialogue System [arxiv] [paper] [code]
ACL 2022
- Multilingual AMR Parsing with Noisy Knowledge Distillation [arxiv] [paper] [code]
EMNLP 2021 (Findings)
- Exploiting Reasoning Chains for Multi-hop Science Question Answering [arxiv] [paper] [code]
EMNLP 2021 (Findings)
- Neural Machine Translation with Monolingual Translation Memory [arxiv] [paper] [code] [slides]
ACL 2021
Outstanding Paper Award (6/3350) - Dialogue Response Selection with Hierarchical Curriculum Learning [arxiv] [paper] [code]
ACL 2021 - Dynamic Semantic Graph Construction and Reasoning for Explainable Multi-hop Question Answering [arxiv] [paper] [code]
ACL 2021 (Findings) - Assessing Dialogue Systems with Distribution Distances [arxiv] [paper] [code]
ACL 2021 (Findings) - Non-Autoregressive Text Generation with Pre-trained Language Models [arxiv] [paper]
EACL 2021 - The World is Not Binary: Learning to Rank with Grayscale Data for Dialogue Response Selection [arxiv] [paper]
EMNLP 2020 - Describe What to Change: A Text-guided Unsupervised Image-to-Image Translation Approach [arxiv]
ACM MM 2020 - AMR Parsing via Graph-Sequence Iterative Inference [arxiv] [paper] [code] [slides]
ACL 2020 - Graph Transformer for Graph-to-Sequence Learning
[arxiv] [paper] [code] [slides]
AAAI 2020 - Core Semantic First: A Top-down Approach for AMR Parsing
[arxiv] [paper] [code] [slides]
EMNLP 2019 - Retrieval-guided Dialogue Response Generation via a Matching-to-Generation Framework
[paper] [code]
EMNLP 2019 - Charge-Based Prison Term Prediction with Deep Gating Network
[arxiv]
EMNLP 2019 - Skeleton-to-Response: Dialogue Generation Guided by Retrieval Memory
[arxiv] [paper] [code] [slides]
NAACL 2019 - Unsupervised Learning helps Supervised Neural Word Segmentation
[preprint]
AAAI 2019 - Translating a Math Word Problem to a Expression Tree
[paper]
EMNLP 2018 - Fast and Accurate Neural Word Segmentation for Chinese
[paper] [code]
ACL 2017 - Neural Word Segmentation Learning for Chinese
[paper] [code] [slides]
ACL 2016 - SeqPE: Transformer with Sequential Position Encoding [arxiv]
IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026.
- Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models [arxiv]
Computational Linguistics, 2025.
- Exploring Dense Retrieval for Dialogue Response Selection
ACM Transactions on Information Systems, 2023. - PROTOTYPE-TO-STYLE: Dialogue Generation With Style-Aware Editing on Retrieval Memory
IEEE Transactions on Audio, Speech and Language Processing, 2021. - Neural Machine Translation with Noisy Lexical Constraints
IEEE Transactions on Audio, Speech and Language Processing, 2020. - A Hybrid Model for Chinese Spelling Check
[paper]
ACM Transactions on Asian and Low-Resource Language Information Process, 2017.