RSSAmplifier

Blog

Ghost in the data

Ghost in the data

ghostinthedata.infoRSS feed ↗19 posts

Latest posts

You Can't Incentivise a Pipeline That Doesn't Break

Incentive pay is designed for a world where output is countable. In data engineering, the most valuable work is invisible. Attaching money to metrics doesn't reveal that work — it buries it.

Your Team Already Has Patterns. They Just Don't Know It.

A pattern bank gives your data engineering team a shared vocabulary for how work gets done — and the confidence to estimate it. Here's how to build one collaboratively, without turning it into a bureaucratic exercise.

Keep Moving

Why forward momentum — not planning, not tooling, not the perfect architecture — is the only thing that actually ships a data platform.

Ghost Skills: Teaching AI Agents to Think Like Data Engineers

A new open-source skills repo that gives AI coding agents the methodology they keep forgetting — grain, keys, SCD2, incident comms, and more. Tool-agnostic, composable, and opinionated.

The Competitive Moat That AI Can't Replicate

In a world racing to automate every interaction, the organisations that invest in genuine human connection are quietly building something no algorithm can copy.

SQL Tells You What. Comments Tell You Why.

SQL is a declarative language — it tells you what the query does, never why. Here's why that distinction matters more in data engineering than anywhere else.

Don't Go Dark: Visibility Is a Data Engineering Skill

Jeff Atwood wrote 'Don't Go Dark' for software engineers in 2008. The advice didn't reach us. Here's what it means for data engineers navigating long migrations, invisible pipelines, and distributed teams.

The Broken Window in Your Data Pipeline

A single ignored data quality issue doesn't stay local. In pipelines, broken windows travel — and by the time anyone notices, the damage is already downstream.

Five Worlds of Data Engineering

Not all data engineering is the same. The modern analytics shop, the enterprise legacy estate, the product engine, the regulated pipeline, and the internal platform each play by different rules — and most advice only applies to one of them.

Your Data Platform Costs More Than It Should

A practical guide to understanding, measuring, and reducing your Snowflake and AWS data platform costs — starting with the habits that actually move the needle.

Why Your Pipeline Finishes Later Every Month

A practical guide to diagnosing pipeline bottlenecks, fixing unnecessary dependencies, and getting data to consumers faster — with Snowflake and AWS patterns you can apply today.

Stop Building Salesforce Integrations From Scratch

A hands-on guide to Snowflake's OpenFlow Salesforce connector — why managed connectors beat custom code, how to set one up step by step, and the schema evolution feature that makes it all worth it.

Your Data Model Isn't Broken, Part II: The Refactoring Playbook

Strangler Figs, Write-Audit-Publish, and the art of replacing a data warehouse one piece at a time without anyone noticing. The practical sequel to why you shouldn't rebuild from scratch.

You Don't Need Permission to Fix Your Data

Five battle-tested tactics junior data engineers can use to improve data quality without waiting for authority, backed by real case studies from Airbnb, Google, and Warner Bros. Discovery.

Your Friends Will Be There for You. Your Work Won't.

The longest study of adult happiness found relationships matter more than career achievement. Yet most of us keep cancelling on our friends. Here's what that costs — and what being a good friend actually looks like.

Your Data Model Isn't Broken, Part I: Why Refactoring Beats Rebuilding

That fact table with 200 columns? Those bridge tables nobody understands? They're not bugs — they're reality encoded. Why the 'let's rebuild' instinct destroys more data teams than technical debt ever will.

12 Steps to Better Data Engineering

A quick scoring framework to assess your data team's engineering maturity. Twelve yes-or-no questions, each with concrete benchmarks for what good and amazing look like using dbt, Snowflake, GitHub Actions, and AWS.

The CSV Test Suite Nobody Writes

How to turn RFC 4180 standards and real-world edge cases into a concrete test suite that catches CSV problems before your pipeline does. A practical guide for producers and consumers.

অনুসন্ধানের ফলাফল

This file exists solely to respond to /search URL with the related search layout template. No content shown here is rendered, all content is based in the template layouts/page/search.html Setting a very low sitemap priority will tell search engines this is not important content. This implementation uses Fusejs and mark.js Initial setup Search depends on additional output content type of JSON in…