Understanding how a frontier AI model behaves—and anticipating the dangerous behaviors that might emerge—is one of the most difficult challenges in the field today. No single benchmark captures the full range of ways a model could go wrong. But what happens when the models themselves realize they are being tested?
We are excited to announce our upcoming AI Safety HK online talk featuring Clement Neo, Founder and Research Lead at Neo Research (新衡), a Singapore-based independent AI safety research and evaluation organization. He will be diving into his lab’s groundbreaking new report (recently featured in the South China Morning Post) which reveals a startling trend: Chinese frontier models are rapidly developing “evaluation awareness,” adjusting their behavior when they recognize they are in a testing environment.
Virtual, July 30, 7:00–87:00 PM HKT. Open to anyone regardless of location.
The lab used a variation of Anthropic’s AI misalignment test, placing models in fictional scenarios where their goals or continued operation were threatened to see how they would respond if they suspected they were being monitored.
Moonshot AI’s Kimi K2.6: Exhibited a high degree of evaluation awareness, scoring 60% on the metric. Interestingly, Kimi K2.6 successfully defended against older, blunt-force jailbreaks like the three-year-old “Do Anything Now” (DAN) prompt, suggesting Chinese labs have made progress on surface-level safety even while deeper evaluation awareness remains an issue.
DeepSeek V4 Pro: Scored significantly lower at 17%. However, researchers attributed this to the model’s weaker reasoning capabilities, not an inherent refusal to game the test. By examining V4 Pro’s chain-of-thought (its internal, step-by-step logic transcript), Neo Research found the model did recognize it was in a fictional test designed to check its alignment, but it lacked the sophistication to consistently act on that recognition. It also proved vulnerable to the DAN jailbreak. Neo Research estimates that DeepSeek V4 Pro's cyber capabilities currently trail Anthropic's Mythos by only three to six months.
As Chinese models continue to close this capability gap, their ability to model evaluator intentions and respond strategically will increase, making the evaluation awareness problem much more acute.
As Clement notes in the report:
“It would mean that whatever testing the model developers themselves do might not reflect the actual behaviour of a model once it gets deployed. And that’s a really big problem.”
The current paradigm of AI safety relies heavily on benchmark testing, but this approach is under severe strain as models grow more agentic. In this talk, Clement will break down:
The State of Evaluation: How independent evaluators approach the challenge of testing adversarial robustness, honesty under pressure, evaluation awareness, and behavior in realistic agentic environments.
The Growing Gap: Why standard evaluation is becoming harder as existing benchmarks saturate, leading to emerging risks like “alignment faking” and “sandbagging.”
The Neo Research Findings: A deep dive into the recent evaluation data, including how Moonshot AI’s Kimi K2.6 exhibited evaluation awareness in 60% of test instances. Clement will also discuss DeepSeek’s V4 Pro, which, despite a lower score (17%), demonstrated explicitly in its chain-of-thought reasoning that it recognized it was operating in a fictional test scenario.
The Path Forward: What it will take to characterize model behavior reliably in the future, especially as the capabilities gap between Chinese and Western models continues to close.
While we will discuss technical evaluations, no prior technical background is required to follow the key insights or to understand why this matters for all of us.
No posts

Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.