
Evaluating your Agentic Harnesses
One good demo is a test flight. An eval suite is the flight-test campaign: pass rates, cost, latency, and failure modes.
A newsletter about all things Data Science, Machine Learning and AI, by Bruno Gonçalves
Subscribe:.rss.atom.json.md.m3u.pls
Live Last read · last published · next check

One good demo is a test flight. An eval suite is the flight-test campaign: pass rates, cost, latency, and failure modes.

Putting your Mac to work

From a single pilot to an air campaign: planning, parallelism, memory, verification, and observability for production-shaped agents.

A 4-gram model on WikiText-103, and the one idea it shares with Claude.

The secret sauce behind Claude Code, Open Clawd, Hermes, etc

Building a complete NLP pipeline to extract entities, discover relationships, and visualize knowledge networks from unstructured text

Use custom made pie charts to visualize U.S. Map of State-to-State Travel

Predicting Molecular Properties from Scratch

A year’s worth of the best Large Language Models, Software Engineering, AI, and Algorithms.

Simulating Epidemics with Birth and Death rates