Agentic Evals Pyramid
Building AI agent evals that work: the 3-layer pyramid approach
My musings about Applied AI, LLMs, RAGs and agents
Building AI agent evals that work: the 3-layer pyramid approach
What are evals? How to use LLM generated data to test your agent.
Definitive guide to building better AI agents driving business value.
Advanced techniques for improving speed, accuracy and cost of JSON generation with LLMs.
A quick experiment to benchmark OpenAI LLMs for JSON generation and analyze the results focusing on error rates, performance and costs.
Common pitfalls in designing and implementing RAGs and how to fix them
How to use GraphQL@Edge to serve a globally replicated DynamoDB table
Lessons learned using Single-table design with DynamoDB and GraphQL in production