RSSAmplifier

Blog

Kostas's Blog

My thoughts written in digital ink.

kostasp.netRSS feed ↗29 posts

Latest posts

Printing Press and the Rise of Consumer-Shaped Software

Printing Press lets the consumer of an API synthesize the exact interface they need. I built a cost-aware Sumble CLI with it, which got me thinking about consumer-shaped software and a world where agents orchestrate while deterministic harnesses compile.

Impact of AlphaEvolve across fields

[Note] Impact of AlphaEvolve across fields

Combee, ACE, GEPA: prompt self-improvement outside the lab

[Note] Combee, ACE, GEPA: prompt self-improvement outside the lab

Evolutionary Search vs Agentic Iteration: 16 AI Experiments on Apache Arrow

A practitioner's account of 16 AI-driven optimization experiments on Apache Arrow's select_k_unstable kernel, from 2x to 396x.

KV Caches again

[Note] KV Caches again

Claude Code Plugins as MVPs

[Note] Claude Code Plugins as MVPs

Anthropic's Cowork

[Note] Anthropic's Cowork

Keystatic for CMS

[Note] Keystatic for CMS

Why I Started Typedef: Data Infrastructure for AI and Agentic Systems

Two years ago I started working on something new based on a strong conviction that data was about to become exponentially more important. We built Typedef, an AI-native data engine that unifies inference, search, and data processing into a single system.

How a Snowflake announcement explains dbt Labs' licensing change

Snowflake's 2025 keynote announcement regarding dbt Core clarifies dbt Labs' controversial shift to change dbt Fusion's licensing model. This represents a rational business response to market dynamics.

Batch Inference, Type Systems, and Why Cortex AISQL Got Me Excited

Snowflake's Cortex AISQL represents a paradigm shift in integrating large language models into data systems as structured, composable functions rather than opaque tools.

Designing the Ideal Synthetic Data Generation Pipeline for LLMs

A comprehensive approach to building scalable, composable pipelines for generating synthetic training data using Large Language Models, specifically focusing on creating question-answer pairs from documents.

Exploring Synthetic Data for LLM Fine Tuning

How synthetic data is used to train and fine-tune large language models, with a focus on Meta's open-source synthetic-data-kit.

Inside Meta's Synthetic-Data Kit for Llama Fine-Tuning

A deep dive into Meta's synthetic-data-kit architecture, a toolkit designed to generate high-quality synthetic datasets for fine-tuning Large Language Models.

DX ∪ UX = U ∧ DX ∩ UX = ∅

The differences and similarities of DX and UX.

Why you should keep an eye on Apache DataFusion and its community.

Why Apache DataFusion is one of the most important open source projects right now.

A glimpse into the future of data processing infrastructure.

Learn about the latest advancements in execution engines and data processing at scale from the recent VeloxCon conference in San Jose.

MLOps is Mostly Data Engineering.

The overlap between MLOps and Data Engineering

A Tutorial on SQL Window Functions Using DuckDB.

Frames, and SQL Window Functions. A tutorial on SQL Window Functions using DuckDB.

What Happened to the API Economy?.

New Opportunities in the API Economy.

Guide to SQL Dates with DuckDB.

Working with Dates in SQL, with examples in DuckDB.

Working with SQL Timestamp, Date and Time Data Types.

Working with Date and Time Types in SQL.

Trino Internals - Parameterized Timestamp Types.

How to add a new data type in Trino.

Snowflake Streaming Ingestion.

How Streaming Data Can be Ingested in Snowflake.

How Snowflake Pricing Works.

Introduction to how Snowflake's pricing works.

About Data Observability.

What's Data Observability and why it matters?

Starting and Growing a Podcast Show.

My experience with the Data Stack Show

Data Ingestion Standards.

What Airbyte, Singer and Kafka Connect have in common

The Battle of the Data Platform.

Snowflake vs Databricks