Patterns for Building Cybersecurity Evals
A sandboxed target, inputs that influence task difficulty, tools, and a grader.
Eugene Yan works at the intersection of consumer data & tech to build machine learning products, and writes about effective data science, learning & career.
A sandboxed target, inputs that influence task difficulty, tools, and a grader.
Build a threat model, discover vulnerabilities, verify, triage, and patch.
Context as infra, taste as config, verification for autonomy, scale via delegation, closing the loop.
An eventful year of progress in health and career, while making time for travel and reflection.
Label some data, align LLM-evaluators, and run the eval harness with each change.
Based on what I've learned from role models and mentors in Amazon
An LLM that can converse in English & item IDs, and make recommendations w/o retrieval or tools.
Evaluation metrics, how to build eval datasets, eval methodology, and a review of several benchmarks.
Recsys & search are converging with LLMs via semantic IDs, data augmentation, and unified foundation models.
What makes a good leader? What do good leaders do? And commando, soldier, and police leadership.
Learning to automate simple agentic workflows with Amazon Q CLI, Anthropic MCP, and tmux.
Applying the scientific method, building via eval-driven development, and monitoring AI output.
How I started, why I write, who I write for, how I write, and more.
Chip Huyen and I share what we've learned, best practices, and insights at NVIDIA GTC 2025.
Model architectures, data generation, training paradigms, and unified frameworks inspired by LLMs.
Exploring how an AI-powered reading experience could look like.
A peaceful year of steady progress on my craft and health.
With regard to writing, there are many rules and also no rules at all.
Benefits of running a weekly paper club, how to start one, and how to read and facilitate papers.
Setting up my new MacBook Pro from scratch
ML systems, production & scaling, execution & collaboration, building for users, conference etiquette.
Look at and label your data, build and evaluate your LLM-evaluator, and optimize it against your labels.
Being a human judge at the Weights & Biases LLM-as-a-Judge Hackathon
FastAPI, FastHTML, Next.js, SvelteKit, and thoughts on how coding assistants influence builders' choices.
Use cases, techniques, alignment, finetuning, and critiques against LLM-evaluators.
What to interview for, how to structure the phone screen, interview loop, and debrief, and a few tips.
Special double-feature closing keynote from the 6 authors of the hit O'Reilly article on Applied LLMs.
Challenges and lessons from deploying LLM experiences: evals, scalability, guardrails.
Structured input/output, prefilling, n-shots prompting, chain-of-thought, reducing hallucinations, etc.
From the tactical nuts & bolts to the operational day-to-day to the long-term business strategy.
Building an AI coach with speech-to-text, text-to-speech, an LLM, and a virtual number.
Evals for classification, summarization, translation, copyright regurgitation, and toxicity.
How unit testing machine learning code differs from typical software practices
Overcoming the bottleneck of human annotations in instruction-tuning, preference-tuning, and pretraining.
Some fundamental papers and a one-sentence summary for each; start your own paper club!
An expanded charter, lots of writing and speaking, and finally learning to snowboard.
Sending helpful & engaging pushes, filtering annoying pushes, and finding the frequency sweet spot.
How to use open-source, permissive-use data and collect less labeled samples for our tasks.
The biggest deployment challenges, backward compatibility, multi-modality, and SF work ethic.
Evals, retrieval-augmented generation, guardrails, and collecting feedback; all that good stuff.
Reference, context, and preference-based metrics, self-consistency, and catching hallucinations.
Distinguishing problems with external vs. internal LLMs, and data vs non-data patterns
Evals, RAG, fine-tuning, caching, guardrails, defensive UX, and collecting user feedback.
Writing drafts via retrieval-augmented generation. Also reflecting on the week's journal entries.
What's the big deal, intuition on query-key-value vectors, multiple heads, multiple layers, and more.
It started with a question that had no clear answer, and led to eight PRs from the community.
Should chat be the main UX for LLMs? I don't think so and believe we can do better.
9 patterns including HITL, hard mining, reframing, cascade, data flywheel, business rules layer, and more.
Generating Dr. Seuss headlines, fake WSJ quotes, HackerNews troll comments, and more.
Also, shortcomings in document retrieval and how to overcome them with search & recsys techniques.