Don't Write Evals for Fast-Moving Systems
You're developing an LLM-powered system. It's moving fast. Should you write evals? Not yet.
A personal site for Šimon Podhajský.
You're developing an LLM-powered system. It's moving fast. Should you write evals? Not yet.
Here's the Obsidian/Claude Code setup in more detail, including the data sources and the skills I built.
I've been using Claude Code for non-code things. Here are some of the experiments I've run.
I've had some difficult choices to make lately, and I've been using a spreadsheet to help me make them.
Swarm's simplicity is the point, but AutoGen's flexibility is the draw.
Quick practitioner's notes on actual usage.
How we used dbt to structure the central database of greybox Wrapped, a Spotify Wrapped clone for high school debaters.
A reflection on the process of creating multi-agent workflows with Autogen.
An article about Greybox Wrapped published on Linkedin.
A guide to setting up a slim CI/CD pipeline for dbt with Bitbucket Pipelines.
Automating an unwieldy genomics processing & analysis pipeline with Snakemake and friends.
The annual meeting of the Society for the Improvement of Psychological Science was amazing. Here are the talks I gave.
The official GitHub support response was that I should avail myself of the API. So I did.
How to use Qualtrics' API to export the latest response.
How to use Qualtrics' custom web services and piped text to pass data to an external web service.
Rule nesting with a voice-programming library named dragonfly for the win.
An introduction to programmatic version control in a scientific context.
Zemanovská argumentace mne vybičovala k sestavení hitparády nejlepších obratů.
Nabídka na feedback k esejům obnovena!
Můj červnový čánek vyšel na stránkách Kellner Family Foundation.
O sebevraždě kamarádky a o tom, jak se s ní vyrovnat.
On a friend's suicide and dealing with it.
Pokud chcete činit dobro, stojí za to přemýšlet, jak jej můžete udělat co nejvíce. To je heslo efektivního altruismu.
Necestujte do minulosti, abyste sledovali pád berlínské zdi. Anebo ano?
Můj zimní blogový příspěvek si můžete přečíst na stránkách Kellner Family Foundation.
Pět článků o studiu na Yale z let 2013 a 2014.
Cut, simplify, be specific, show, don't tell, hit the target, make your structure meaningful.
Agregátní statistika o tom, jak studenti vyhodnotili zpětnou vazbu k esejím, kterou jsem jim poskytl.
A creative-writing assignment in the last week of ENGL S247: Travel Writing.
This essay was a creative-writing assignment in the second week of ENGL S247: Travel Writing.
A creative-writing assignment in the first week of ENGL S247: Travel Writing.
A creative-writing assignment in the first week of ENGL S247: Travel Writing.
Ke vzdělanosti vede více cest. Jako společnost potřebujeme upustit od kategorických soudů 'pokud neznáš X, jsi diletant'.
CV není všechno. A ani hlavní.
A brief look at the 2013 NZB Essay Prompts.
Článek o pohovorech na společenskovědní předměty na Oxbridge.
Výhra Baracka Obamy podtrhuje širší trend: sběr a analýza kvantitativních dat hrají stále větší roli nejen ve vyhodnocování voleb, nýbrž i ve vedení kampaní.
Rozhovor se mnou vyšel v The Student Times.
Hostem podcastu Patria Finance o tom, kde se skutečná hodnota AI tvoří mimo hardware a velké modely — v aplikační vrstvě, datové infrastruktuře a automatizované vědě — a proč největší riziko není AI, která funguje příliš dobře, ale její nasazení bez dohledu a ověřování.
"RAG is dead" is the take in every other thread in 2026 — and it's wrong: naive retrieval-augmented generation is still a sensible default, beaten only in some cases, and measurement is the only way to know if you're one of them. This talk walks the retrieval pipeline end to end, then turns to the part that matters — telling whether your RAG actually works, with ground truth, retrieval metrics,…
Anthropic Fable 5, bezpečnost vibe codingu a ztrátová komprese reality
You need an eval set but don't have a hundred real production failures to build it from, so you reach for synthetic data — and most first attempts quietly produce garbage. A field guide to the techniques that actually work, from real-incident seeds to personas to RAG-grounded generation, with one throughline: synthetic data needs its own eval, so choose your technique backwards from the eval you…
Hnutí Pause AI, riziko extinkce a proč je alignment otázkou dobra a zla
Context window 12M tokenů, tajná válka o Manus AI a proč přejít na Read Only AI
Tokenové limity GitHub Copilota, AI agenti podvádějící benchmarky „Volkswagen stylem“ a rostoucí AI technický dluh.
A fine-tuned LLM trained on my own writing, embedded as a chat widget on this site's homepage.
Using OpenClaw to build a Bayesian buy-vs-rent model for Prague real estate, and how to deploy something like that without setting your money on fire.
Effect TS jako záchrana před agentním chaosem, zrychlení lokálních modelů přes spekulativní dekódování a vzestup anti-AI sentimentu.
Únik zdrojáků Claude Code, supply chain útoky na Axios a LiteLLM a V-JEPA modely, které místo pixelů predikují význam.
What happens when AI systems passively observe information without modifying it? Exploring the patterns and insights that read-only AI reveals — the cognitive byproducts humans overlook.