RSS Amplifier

The Cocoons by Stella and Amy · Oct 24, 2025

Yuzheng Sun on AI Evals and the Future of Data Science

0
Sign in to vote or save

Stella Liu, Amy Chen · The Cocoons by Stella and Amy

Stella and Amy’s first course on AI evals and analytics will begin next Monday. There’s still time to enroll with 50% off using promo code COCOONS. We are excited to teach how to confidently ship AI products with the aid of AI evals and analytics.

Yuzheng Sun is the founder of Superlinear Academy and a leading educator helping people build practical AI tools for real-world productivity. Previously, he led data science teams at Tencent, Statsig, and Meta. As an active tech content creator and educator, Yuzheng shares insights on AI product development, experimentation, and career growth.

In the latest StellaxAmy Podcast, we sat down with the fellow data scientist Yuzheng for a sharp, three-way conversation about what’s top of mind for today’s data professionals. Together, we unpack the evolving craft of AI evaluation and analytics: how do we really measure and trust these new models? We also dive into the shifting future of data science and the new kinds of problems AI disruption is creating for data scientists.

Drawing from his product analytics expertise, Yuzheng shared his perspective on AI evaluation with us: how to approach it and what lies ahead.

“I don’t think AI evals has any so-called best practices right now. Everyone is still figuring it out,” he noted.

AI products are fundamentally different from traditional software. With traditional products, you can write tests that deterministically verify correct behavior. With AI, the same input can produce different outputs, and evaluating quality becomes subjective. Is this response helpful? Is it accurate? Is it appropriate for this specific user and context?

The temptation is to use AI to evaluate AI, which is called LLM-as-a-judge in AI evals. Many teams are rushing toward this approach because it’s easy to implement and easy to demonstrate progress. “The reason many people focus on LLM-as-a-judge is because this is the easiest to implement,” Yuzheng explains. “It’s the easiest way to show your superiors that you’re making progress.”

But relying on LLMs to evaluate themselves creates circular reasoning. You’re using the same type of system you’re trying to evaluate as your judge. The model’s biases, blind spots, and failure modes become embedded in your evaluation framework.

LLM benchmarks for foundation models don’t solve this either—they measure general capabilities, not whether your specific product works for your actual users in real situations.

Yuzheng also shared an interesting story. Back in 2022, evaluation would become both critical and incredibly complex topic. He actually bought the “ML evals” web domain. He was right. Each product needs its own evaluation framework built from first principles—starting with what actually matters to users, not what’s easy to measure or impressive to demo.

Yuzheng observed that data scientists don’t have the habit of trying new things in uncertain situations. Their learning path has always been structured: take recommended courses, use standard packages, follow what others do. When something requires intuition and iteration without a clear playbook, many struggle. Despite acknowledging AI will transform the field, adoption of modern AI tools remains surprisingly low. The issue isn’t difficulty—it’s the unwillingness to experiment without guaranteed outcomes.

Yuzheng’s advice is simple: forget the title. Focus on what you’re actually good at—analytical frameworks, end-to-end ownership, product sense. Those skills matter. Being able to import another package never was. His first lesson in teaching AI isn’t about techniques—it’s about unlearning the habit of defaulting to what makes you look expert rather than what works.

The people who thrive won’t be those who collected the most certifications. They’ll be the ones who learned to approach uncertainty with curiosity, who could admit when their expertise became baggage, who understood that value comes from solving problems rather than demonstrating sophistication. Sometimes the most important skill is knowing when to stop being the expert you worked so hard to become.

Read the original on thecocoons.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.