Eric Wong
194
posts
Assistant professor at University of Pennsylvania. Machine learning, optimization, robustness & interpretability. profericwong.bsky.social
Philadelphia, PA
Joined July 2009
Pinned
You may have heard agentic benchmarks can be gamed. Turns out, leading solutions already do this. With Meerkat, our framework for large-scale trace auditing, we found thousands of clear cheating instances, likely from unsupervised vibe coding. debugml.github.io/cheating-agentβ¦
Recent work led by
@WeiqiuYouat the intersection of VLM reasoning and the surgical problem of identifying the critical-view-of-safety! A fun (and eye-opening) collaboration with
@Laparoscopesfor explainable AI in surgery problems.
LLM ignoring instructions? Make it listen with InstABoost. β Simple: Steer your model in 5 lines of code β Effective: Outperforms latent steering & prompt-only methods β Grounded: Based on our mechanistic theory on rule-following (LogicBreaks) Blog: debugml.github.io/instaboost
Why can safety rules in LLMs be jailbroken? In LogicBreaks, we study the fundamental mechanism behind rule subversion in LLMs. Our theory explains how one can force LLMs to suppress rules/knowledge and infer absurd facts--and it mirrors real jailbreaks! debugml.github.io/logicbreaks/
Traditional concept vectors used to explain deep representations fail to compose when combined, i.e. π€(small) +π¦’(white) =π¦©(big & colorful)β We propose CCE: a method for extracting *composable* concepts, i.e. π€(small) +π¦’(white) =ποΈ(small & white)β debugml.github.io/compositional-β¦




